Data warehouse modeling method, device and equipment and computer storage medium

Through the feature-based data modeling method, the data model is simplified, and the data redundancy and complexity problems in traditional data warehouse modeling methods are solved, development and maintenance costs are reduced, and a good operation and maintenance foundation is provided for machine learning systems.

CN119938798APending Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311455426.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional data warehouse modeling methods have problems in data redundancy, data consistency, query operation complexity, etc., resulting in high development and maintenance costs.

Method used

Adopt a data modeling method with features as the core to simplify data models, reduce data warehouse development and maintenance costs, avoid dimensional data redundancy problems, and lay a good foundation for machine learning system operation and maintenance.

Benefits of technology

By using features as the core of data modeling, the data model is simplified, the development and maintenance costs of data warehouses are reduced, the redundancy of dimensional data is avoided, and a good operation and maintenance foundation is provided for machine learning systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938798A_ABST
    Figure CN119938798A_ABST
Patent Text Reader

Abstract

The invention provides a data warehouse modeling method and device, electronic equipment and a computer storage medium, and relates to the technical field of data processing. The method comprises the following steps: analyzing data reported by first equipment, and generating a first data table; according to data of a first feature included in the first data table and a feature registry, a first basic feature data table needing to be updated is determined, the first basic feature data table comprises at least one set of feature data of the first feature, and each set of feature data in the at least one set of feature data comprises time, an entity and a measurement value; the first feature is a registered feature, and the first data table is a unique dependent data table of the at least one basic feature data table; and updating the first basic feature data table according to the data of the first feature. According to the method, the features are taken as the core of data modeling, and concepts such as dimensions, metrics and indexes are not distinguished any more, so that the data model is simplified, and the development and maintenance cost of a data warehouse is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular, to a method, apparatus, device and computer storage medium for data warehouse modeling. Background Art

[0002] In an ever-growing and changing data environment, data can only be better utilized if it is organized and stored in an orderly manner. A data warehouse is a large data storage collection that screens and integrates different source data, and ultimately completes the transformation of the source data organization form in a reasonable modeling method in order to analyze the data.

[0003] The development and maintenance costs of traditional data warehouse modeling methods, such as star schema, snowflake schema, fact constellation schema, etc., are relatively high, and there are certain problems in data redundancy, data consistency, query operation complexity, etc.

[0004] How to better organize and store data in an orderly manner so that the data can be better used is an issue that needs to be addressed urgently. Summary of the invention

[0005] The embodiments of the present application provide a method, apparatus, device and computer storage medium for data warehouse modeling, which simplifies the data model, reduces the development and maintenance costs of data warehouses, avoids the dimensional data redundancy problem caused by dimensional data modeling, and lays a good foundation for the operation and maintenance of machine learning systems.

[0006] In a first aspect, an embodiment of the present application provides a method for data warehouse modeling, including:

[0007] Parsing a first data table reported by a first device;

[0008] Determine the first basic feature data table that needs to be updated according to the data of the first feature included in the parsed first data table and the feature registration table, wherein the first basic feature data table includes at least one set of feature data of the first feature, each set of feature data in the at least one set of feature data includes time, entity and measurement value, the first feature is a registered feature, and the first data table is the only dependent data table of the at least one basic feature data table;

[0009] The first basic feature data table is updated according to the data of the first feature included in the parsed first data table.

[0010] In a second aspect, an embodiment of the present application provides a device for data warehouse modeling, including:

[0011] A processing unit, which parses the data reported by the first device and generates a first data table;

[0012] The processing unit is further used to determine the first basic feature data table that needs to be updated according to the data of the first feature included in the first data table and the feature registration table, wherein the first basic feature data table includes at least one set of feature data of the first feature, each set of feature data in the at least one set of feature data includes time, entity and measurement value, the first feature is a registered feature, and the first data table is the only dependent data table of the at least one basic feature data table;

[0013] The processing unit is further configured to update the first basic feature data table according to the data of the first feature.

[0014] In a third aspect, the present application provides an electronic device, including:

[0015] a processor adapted to implement computer instructions; and,

[0016] The memory stores computer instructions, where the computer instructions are suitable for being loaded by the processor and executing the method of the first aspect.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are read and executed by a processor of a computer device, the computer device executes the method of the first aspect above.

[0018] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of the first aspect above.

[0019] Through the above technical solution, features are used as the core of data modeling, and concepts such as dimensions, metrics, and indicators are no longer distinguished, which greatly simplifies the data model and reduces the cost of data warehouse development and maintenance. The feature data table only contains time, entities, and measurements, avoiding the dimensional data redundancy problem caused by dimensional data modeling. And because the feature data table has a simple structure, the processed features can be directly applied to machine learning model training and prediction, laying a good foundation for machine learning system operation and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 An optional schematic structural block diagram of the system architecture involved in the embodiments of the present application;

[0021] Figure 2An optional schematic structural block diagram of the system architecture involved in the embodiments of the present application;

[0022] Figure 3 A schematic structural block diagram of another optional system architecture involved in an embodiment of the present application;

[0023] Figure 4 A schematic flow chart of a method for data warehouse modeling provided in an embodiment of the present application;

[0024] Figure 5 A schematic structural block diagram of a basic feature data table provided in an embodiment of the present application that solely relies on one upstream data table;

[0025] Figure 6 is a schematic flow chart of a method for generating derived features provided in an embodiment of the present application;

[0026] Figure 7 is a schematic structural diagram of a feature splicing provided in an embodiment of the present application;

[0027] Figure 8 is a schematic flow chart of a feature stitching method provided by the present application;

[0028] Fig. 9 It is a schematic structural block diagram of a data warehouse hierarchical structure proposed in an embodiment of the present application;

[0029] Fig.10 is a schematic block diagram of a device provided in an embodiment of the present application;

[0030] Fig.11 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0031] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0032] First of all, in order to more clearly understand the embodiments of the present application, the relevant terms involved in the embodiments of the present application are specifically described below.

[0033] Feature: A measure of a characteristic of an entity, such as a user or a product, at a specific point in time (or time period). For example, the behavioral characteristics of a user in a certain time period, such as the number of logins, online time, number of clicks, etc. It is particularly important to point out that gender, age, and device platform (iOS, Android) are generally treated as portrait features of dimensions and are also considered as features in this application.

[0034] Fundamental Feature: A direct measurement of a certain characteristic of an entity, which is generally additive. The data source of the fundamental feature is generally the original log data reported by the server or client.

[0035] Derived Feature: A feature that is processed using basic features as input data through mathematical calculations (such as sum, mean, extreme value), vectorization, and other operations.

[0036] Feature Derivation: The process of deriving features based on basic features. Feature derivation is not limited to numerical calculations, and can even be the prediction results of a machine learning model. In short, all operations that can be encapsulated into functions can be used as a means of feature derivation.

[0037] In an ever-growing and changing data environment, only by organizing and storing data in an orderly manner can data be better utilized. For example, a library stores a large number of books. If these books are placed in a disorderly manner, readers will spend a lot of time looking for the books they need. However, if these books are placed in different categories on the bookshelf, readers will not spend a lot of time looking for the books they need. The data warehouse model is a data organization and storage method that emphasizes the reasonable storage of data from the perspective of business, data access and use. Only after the data warehouse model organizes and stores data in an orderly manner can big data be used with high performance, low cost, high efficiency and high quality.

[0038] Figure 1 An optional schematic structural block diagram of a data warehouse system architecture 100 provided in an embodiment of the present application. Figure 1 As shown, the data warehouse system architecture can be divided into three layers: operation data storage layer (Operate datastore, ODS) 110, data warehouse layer (Data warehouse, DW) 120, and data application layer (Application Data Service, ADS) 130.

[0039] The operational data storage layer refers to the data source of the data warehouse, which stores the unprocessed raw data. The data in the operational data storage layer can be obtained through business libraries, embedded point logs, message queues, etc.

[0040] The data warehouse layer builds various data models from the data obtained from the ODS layer. The data warehouse layer is the destination of data. All data coming from the ODS is stored here for a long time, and this data will not be modified. The data in the DW layer should be consistent, accurate, and clean data, that is, data after the source system data has been cleaned (impurities removed).

[0041] The DW data is layered, from bottom to top, into the data warehouse detail layer (DWD), the basic data layer (DWB) and the data warehouse service layer (DWS). The DWD layer mainly stores the original data after cleaning, parsing and conversion. The DWB stores objective data and is generally used as an intermediate layer. It can be considered as a data layer for a large number of indicators. At the DWS layer, the data is further processed and organized to facilitate subsequent modeling and analysis.

[0042] The data application layer mainly provides data products and data analysis. It is usually stored in systems such as MySQL for online systems to use, or in Hive for data analysis and data mining. For example, the report data or wide tables we often talk about are usually placed in the data application layer.

[0043] The following is a brief description of the traditional data warehouse modeling approach.

[0044] Traditional data warehouse modeling methods mainly include the following modes:

[0045] Star Schema mainly consists of fact table and dimension table. Fact table is used to record the measurement value (measure) in the business process, such as sales, number of clicks, etc. Dimension table is used to store the attributes describing these measurement values, such as time, location, product, etc. The main feature of star schema is to directly associate fact table and dimension table through foreign key, thus forming a star structure.

[0046] Snowflake Schema is an extension of the star model. Compared with the star model, the snowflake schema normalizes the dimension table, splits some dimension tables, and then associates them with the main dimension table, thereby establishing a hierarchical relationship between dimension tables.

[0047] Fact Constellation Schema, also known as Galaxy Schema, is also an extension of the star model, but is more complex than the star and snowflake schemas. Its main features are: according to the fact subject, the original star schema is decomposed into several smaller star schemas, namely fact constellations. The fact tables in each fact constellation share the dimension table. The fact constellation schema is generally suitable for cross-business process analysis scenarios.

[0048] Data Vault Modelling is an extension of the Entity-Relation model. The model contains three types of tables: Hub, Link, and Satellite. The Hub table records the main entities in the business process; the Link table records the relationship between two or more Hub tables and the changes of these relationships over time; the Satellite table is used to record descriptive or contextual information about the Hub and Link tables, such as status, type, etc. These three tables can effectively cope with the rapid changes in business needs, while ensuring the traceability, auditability, and scalability of data to a certain extent.

[0049] Anchor modeling is an improvement on the data evolution model, which mainly includes four types of tables: Anchors, Attributes, Ties, and Knots. Anchors are equivalent to Hubs in the data evolution model and are used to record entities. Attributes are similar to Satellites in the data evolution model and are used to record the characteristics of anchors. Ties are equivalent to Links in the data evolution model and are used to record the relationship between anchors. Finally, Knots are used to record stable attributes shared by multiple anchors, such as enumeration constants such as gender and nationality.

[0050] The following describes the problems and shortcomings of the above five methods respectively:

[0051] Each dimension in the star schema is directly associated with the fact table, so there is a certain degree of data redundancy. Since the dimensions are not normalized, data inconsistency and incompleteness may occur. In addition, the star schema is not flexible enough for data analysis tasks, especially for many-to-many relationships between entities, which often requires the establishment of additional bridge tables.

[0052] Although the snowflake schema reduces data redundancy by normalizing the dimension tables, it increases the query join operations and complexity, and still cannot guarantee data integrity. In addition, the snowflake schema is more complex than the star schema, so the development and maintenance costs are also higher.

[0053] Fact constellation schemas are more complex than snowflake schemas, so they require more time and expertise to design and implement. Fact constellation schemas consume more storage space than star schemas. Data queries often involve more join operations, making query statements difficult to understand and query performance degraded.

[0054] Compared with the first three models, the data evolution model is more complex and difficult to master. Designing and implementing a complete data evolution model may require relevant training for the data development team first. In addition, since the data evolution model saves all historical data, it consumes a huge amount of storage space. Finally, due to its highly standardized design, more join operations are involved in data queries.

[0055] The anchor model and the data evolution model belong to the ensemble modeling category, and the anchor model also has high requirements for developers. The large number of connection operations brought about by high normalization puts higher requirements on the database, and may require support from capabilities such as table elimination. In addition, the anchor model also involves the processing of time interval data, so it is necessary to use a dedicated temporal database or write additional programs to maintain time interval data, such as deleting records with the same value for the same ID in two consecutive time events.

[0056] In addition, the physical implementation of the data model usually stratifies the data tables according to the principles of data reuse and computational reuse for use in downstream tasks such as machine learning, causal inference, and online analytical processing (OLAP). However, data stratification strategies are difficult to standardize and often appear as Figure 1 The data stratification requires long-term management and control, otherwise the data warehouse is at risk of structural degradation. On the other hand, because the data has been calculated in multiple layers, its lineage is often difficult to trace. The logical definition changes slowly due to the adjustment of the intermediate calculation layer, and the data quality is difficult to control.

[0057] Therefore, an embodiment of the present application proposes a method for data warehouse modeling, which takes features as the core of data modeling and no longer distinguishes between concepts such as dimensions, metrics, and indicators. This greatly simplifies the data model, thereby reducing the cost of data warehouse development and maintenance. The feature data table only contains time, entities, and measurements, avoiding the dimensional data redundancy problem caused by dimensional data modeling. Moreover, since the feature data table has a simple structure, the processed features can be directly applied to machine learning model training and prediction, laying a good foundation for machine learning system operation and maintenance.

[0058] In order to more clearly understand the embodiments of the present application, the system architecture involved in the embodiments of the present application is described below.

[0059] Figure 2 2 is an optional schematic structural block diagram of the system architecture 200 involved in the embodiment of the present application. Figure 2As shown, the system architecture 200 includes a business layer service module 210, a backend service module 220, and a storage and computing engine module 230. The business layer service module 210 includes a web user interface (WEB UI) module 211 and a client module 212. The business layer service module 210 is user-oriented. The user can retrieve the features of the data warehouse model, determine the blood relationship of the features, define derived features, and perform feature detection through the WEB UI module 211 included in the business layer service module 210; the user can also register the features of the data warehouse model, materialize the features, determine the derived features, etc. through the client module 212 included in the business layer service module 210.

[0060] The backend service module 220 and the storage and computing engine module 230 are used to provide basic services of the data warehouse system, that is, the business layer service module 210 needs to call certain services in the backend service module 220 and / or the storage and computing engine module 230 to implement business layer services.

[0061] The business layer service module 210, the backend service module 220 and the storage and computing engine module 230 are described in detail below.

[0062] The business layer service module 210 includes a WEB UI module 211, which provides graphical feature operation capabilities for users. Its operation capabilities may include feature retrieval, feature materialization, feature lineage and feature detection. Among them, feature retrieval is used to search for existing features in the data warehouse through subject domains and feature definitions; feature materialization is used to combine multiple features selected by developers into one storage unit and store them in a distributed file system (Hadoop Distributed File System, hdfs), relational database management system (My structured query language, mysql), star rocks or remote dictionary server (Remote Dictionary Server, redis); feature lineage is used to view the calculation logic of the feature, if it is a derived feature, check which basic features the derived feature is generated by, and how the derived feature is calculated and generated; feature detection is used to configure feature detection items, such as feature distribution detection, feature outlier detection, etc.

[0063] The client module 212 included in the business layer service module 210 includes Jupyter Lap, code editor (VisualStudio Code, VS Code), feature registration, derived features, feature materialization, and feature service Serving. It supports the use of multiple integrated development environments (IDEs), such as Jupyter Lab, VSCode, etc.; feature registration is used to register feature metadata to feature services; derived features are used to derive features from existing feature definitions by users; feature materialization is used to derive features by defining calculations and write the calculated features into corresponding storage units.

[0064] The backend service module 220 is used to provide backend service capabilities for various feature processing, including feature retrieval service, feature calculation service, feature quality detection service, feature lineage service, feature storage service, feature materialization service, feature registration service, data source management service and permission service. Among them, the feature retrieval service is used to search for existing features in the data warehouse; the feature calculation service is used to perform relevant calculations on the features, such as finding the mean of the features, summing the features, etc.; feature quality detection service; feature lineage service is used to view the calculation logic of the features, if it is a derived feature, to view which basic features the derived feature is generated by, and how the derived feature is calculated and generated; feature storage service is used to store the features in the corresponding storage unit; feature materialization service is used to combine multiple features selected by the developer into a storage unit, and store it in HDFS, MySQL, Star Rocks or Redis; feature registration service is used to register a new feature; data source management service is used to manage data sources; permission service is used to limit the objects that can use the data warehouse model.

[0065] The storage and computing engine module 230 is used to actually take charge of the basic services of feature computing and feature storage, including spark, hdfs, mysql, flink, starrocks, redis, kafka, etc.

[0066] Among them, Spark is a fast and general computing engine designed for large-scale data processing; HDFS refers to a distributed file system designed to run on general-purpose hardware. HDFS is a highly fault-tolerant system suitable for deployment on cheap machines. HDFS can provide high-throughput data access and is very suitable for applications on large-scale data sets. MySQL is a relational database management system; Flink is an open source stream processing framework that executes any stream data program in a data parallel and pipeline manner; StarRocks is a new generation of extremely fast full-scenario (Massively Parallel Processing, MPP) database. It can make users' data analysis simpler and more agile. Users can use StarRocks to support extremely fast analysis of various data analysis scenarios without complex preprocessing; Redis is a key-value storage system that supports relatively more value types, including string, list, set, and hash; Kafka is a distributed, publish / subscribe-based messaging system.

[0067] Figure 3 FIG. 3 is a schematic structural block diagram of another optional system architecture 300 involved in an embodiment of the present application. Figure 3 As shown, the system architecture 300 includes a feature application module 310, a feature processing module 320, a feature management module 330 and a feature storage module 340. The feature application module 210 accesses the data warehouse model through an application programming interface (Application Programming Interface, API). API is a set of predefined functions. The purpose is to provide applications and developers with the ability to access a set of routines based on certain software or hardware without having to access the source code or understand the details of the internal working mechanism.

[0068] The feature application module 310 includes feature applications such as Machine Leaning Operations (MLOps), OLAP, and causal inference. MLOps is developed from DevOps in the field of machine learning and software engineering, and refers to a series of technologies for deploying and maintaining machine learning models in production; OLAP is a fast and flexible multidimensional data analysis method.

[0069] The feature processing module 320 includes basic feature calculation, feature derivation, feature splicing, flink and spark. It relies on flink and spark to realize the basic feature calculation, feature derivation, splicing and other functions mentioned above. Users can compile feature processing pipelines (pipeline choreography) in a configurable manner.

[0070] The feature management module 330 includes functions such as feature registration, quality inspection, and lineage analysis to ensure that feature data is controllable, reliable, traceable, and reusable.

[0071] The feature storage module 340 includes Hive, MySQL, Star Rocks, Redis, etc. In addition to storing feature data, the feature storage layer also has a feature anchoring mechanism, which can anchor basic features to different data sources, such as Hive, MySQL, Redis, etc., thereby shielding the differences between streams and batches, so that data scientists and machine learning engineers can focus on feature processing and use.

[0072] It should be understood that the above modules are divided by function. These modules can be deployed on an independent physical server, or on a server cluster or distributed system consisting of multiple physical servers. They can also be deployed on cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0073] The solution provided by the embodiments of the present application is described below in conjunction with the accompanying drawings.

[0074] Figure 4 The present invention provides a schematic flow chart of a method 400 for data warehouse modeling. The method 400 can be executed by any electronic device with data processing capability. For example, the electronic device can be implemented as a server or a computer. The following description is made by taking the electronic device as a joint scheduling device as an example. Figure 4 As shown, method 400 may include steps S410 to S430.

[0075] S410, the terminal device parses the data reported by the first device and generates a first data table.

[0076] S420, the terminal device determines the first basic feature data table that needs to be updated based on the data of the first feature included in the first data table and the feature registration table, wherein the first basic feature data table includes at least one set of feature data of the first feature, each set of feature data in the at least one set of feature data includes time, entity and measurement value, the first feature is a registered feature, and the first data table is the only dependent data table of the at least one basic feature data table.

[0077] S430: The terminal device updates the first basic feature data table according to the data of the first feature included in the first data table.

[0078] The data warehouse modeling method uses features as the core of data modeling and no longer distinguishes between concepts such as dimensions, metrics, and indicators. This greatly simplifies the data model and reduces the cost of data warehouse development and maintenance. The feature data table only contains time, entities, and measurements, avoiding the dimensional data redundancy problem caused by dimensional data modeling. And because the feature data table has a simple structure, the processed features can be directly applied to machine learning model training and prediction, laying a good foundation for machine learning system operation and maintenance.

[0079] In order to more clearly understand the embodiment of the present application, the method 400 is described in steps below.

[0080] Optionally, in step S410, the terminal device parses the data reported by the first device to generate a first data table, including:

[0081] The terminal device performs format conversion on the data reported by the first device, processes missing values ​​and abnormal values ​​of the data reported by the first device, and generates the first data table.

[0082] Optionally, the first device includes a client device and / or a server.

[0083] In step S420, the terminal device determines the first basic feature data table that needs to be updated based on the data of the first feature included in the first data table and the feature registration table, wherein the first basic feature data table includes at least one set of feature data of the first feature, each set of feature data in the at least one set of feature data includes time, entity and measurement value, the first feature is a registered feature, and the first data table is the only dependent data table of the at least one basic feature data table.

[0084] Specifically, the first basic feature data table includes at least one set of feature data of the first feature, and each set of feature data includes time, entity and measurement value, that is, the first basic feature data table includes multiple sets of time + entity + measurement value, and the table structure is simple, thereby greatly reducing the complexity of data warehouse modeling.

[0085] For example, the structure of the basic feature data table is explained by taking the message receiving and sending behaviors of users as an example:

[0086] CREATE TABLE dws_feature_chat_1di (

[0088] m_date STRING COMMENT "Measurement date"

[0089] ,user_id BIGINT COMMENT "User ID"

[0090] ,send_msg_cnt_1d BIGINT COMMENT "Number of messages sent on that day"

[0091] ,receive_msg_cnt_1d BIGINT COMMENT "Number of messages received on that day" )

[0093] COMMENT "Data table of basic characteristics of receiving and sending message behaviors"

[0094] PARTITION BY LIST(m_date) (

[0096] PARTITION default )

[0098] STORED AS ORCFILE COMPRESS;

[0099] Among them, m_date is the generation time of the measurement value, which can be a date here, user_id is an entity, which is the unique identifier of the user here, send_msg_cnt_1d and receive_msg_cnt_1d are the number of messages sent and received by the user on the same day respectively.

[0100] In the data warehouse modeling method proposed in the embodiment of the present application, the basic features correspond to the DWS layer of the traditional data warehouse. The simplification of the basic feature structure realizes the simplification of the DWS layer and makes the physical layering of the data warehouse below the DWS layer thinner.

[0101] In the data warehouse modeling method proposed in the embodiment of the present application, unless necessary, a basic feature table depends on only one upstream table and is not self-dependent. That is, the basic feature data table is solely dependent on one upstream data table. For example, the basic feature data table 1 and the basic feature data table 2 can both depend on the upstream table first data table, and the basic feature data table 1 cannot both depend on the upstream table first data table and the first data table. The basic feature data table is solely dependent on one upstream data table, which can avoid single point failures and data delays caused by complex task dependencies, reduce the computational cost of the Extract-Transform-Load (ETL) process, improve data quality, and lay the foundation for data lineage analysis.

[0102] For example, Figure 5 A schematic structural block diagram of a basic feature data table provided in an embodiment of the present application that solely relies on one upstream data table, such as Figure 5 As shown, the DWD layer stores the parsing results of the data reported by the first device, which are the parsing results of the data reported by the client and the parsing results of the data reported by the server. The parsing results of the data reported by the client and the parsing results of the data reported by the server are the upstream tables. The terminal device needs to update the chat behavior basic feature table and the comment behavior basic feature table according to the parsing results of the data reported by the client, and update the gift behavior basic feature table according to the parsing results of the data reported by the server. The basic feature data table will not rely on the two upstream data tables at the same time, such as the chat behavior basic feature will not rely on the parsing results of the data reported by the server and the parsing results of the data reported by the client at the same time.

[0103] It should be understood that Figure 5 The chat behavior basic feature data table, comment behavior basic feature data table, and gift behavior basic feature data table shown in are basic feature data tables constructed based on feature themes, which is conducive to subsequent data application analysis of different themes.

[0104] It should also be understood that the embodiment of the present application does not limit the number of topics that can be included in a basic feature data table. A basic feature data table can include one topic or multiple topics. Generally speaking, the features and topics included in a basic feature data table can be defined when registering the features.

[0105] Optionally, the method 400 further includes:

[0106] If the first data table also includes data of a second feature, and the second feature is an unregistered feature, register the second feature,

[0107] The registering the second feature includes:

[0108] In the feature registration table, a storage relationship between the data of the second feature and the first basic feature data table is established; or,

[0109] In the feature registration table, a storage relationship between the data of the second feature and the second basic feature data table is established.

[0110] Specifically, when the second feature included in the first data table is an unregistered feature, the second feature needs to be registered. The second feature is registered through the feature registration service, and a basic feature storage table corresponding to the second feature is established. The data of the second feature can be recorded in the first basic feature data table, and the data of the second feature can be recorded in the second basic feature data table. Through feature registration, the basic features can be clearly defined, and redundancy problems such as repeated storage of basic features will not occur.

[0111] Optionally, determining the first basic feature data table to be updated according to the data of the first feature included in the first data table and the feature registration table includes:

[0112] Determining a subject corresponding to the first feature according to the first feature and the feature registration table;

[0113] Determine the first basic feature data table associated with the subject corresponding to the first feature.

[0114] Specifically, the subject corresponding to different features and the association relationship between the basic feature data table corresponding to the subject can be recorded in the feature registration table. When the terminal device obtains the first feature, the subject corresponding to the first feature is determined in the feature registration table, and the first basic feature data table associated with the subject corresponding to the first feature is determined.

[0115] Dividing features into different topics, such as social behavior topics and commercial behavior topics, facilitates the application of data analysis and machine learning tasks.

[0116] It should be understood that the process of dividing features into different topics can be performed in the feature registration service.

[0117] Optionally, the method 400 further includes:

[0118] The terminal device acquires derived feature configuration information according to the list of derived features to be derived;

[0119] The terminal device acquires basic feature data to be derived according to the derived feature configuration information;

[0120] The terminal device generates a derived feature data table according to the feature aggregation function and the basic feature data to be derived.

[0121] Specifically, the terminal device can derive derived features based on the basic feature data. In order to more clearly understand the process of deriving derived features based on the basic feature data, the following is a Figure 6 Please specify. Figure 6 As shown, Figure 6 It is a schematic flow chart of a method 500 for generating a derived feature provided by the present application, and the execution subject of the method 500 is the same as the aforementioned method 400. The method 500 may include steps S510 to S570.

[0122] S510, the program starts.

[0123] S520: The terminal device traverses the list of derived features to be derived and obtains derived feature configuration information.

[0124] S530: The terminal device obtains basic feature data to be derived according to the derived feature configuration information.

[0125] S540: The terminal device generates a derived feature data table according to the feature aggregation function and the basic feature data to be derived.

[0126] S550: The terminal device determines whether the list of derived features to be derived has been traversed to completion.

[0127] S560: When the traversal of the derived feature list to be derived is not completed, continue to execute S520.

[0128] S570, when the traversal of the derived feature list to be derived is completed, the feature derivation procedure ends.

[0129] Optionally, the derived feature configuration information includes at least one of the following parameters:

[0130] The basic features of the deduction, the target entity of the deduction, the time of the deduction and the place of the deduction.

[0131] Optionally, the feature aggregation function includes any one of the following functions:

[0132] Functions corresponding to mathematical operations, vectorization functions, and machine learning models.

[0133] The following is a specific explanation using the example that the derived feature configuration information is time information such as time window aggregation. For example, the derived feature configuration information obtained by the terminal device is a feature for deriving the total number of messages sent and received by the user in the past 7 days and 14 days. The terminal device obtains the number of messages sent and received by the user in the past 7 days and 14 days that need to be derived based on the derived feature configuration information. The terminal device derives the feature of the total number of messages sent and received by the user in the past 7 days and 14 days based on the feature aggregation function and the number of messages sent and received by the user in the past 7 days and 14 days. The specific procedure is as follows:

[0134]

[0135]

[0136] This program describes how to automatically derive features such as the number of messages sent and received by a user in the past 7 days and 14 days based on the derived feature configuration information. It should be understood that the feature aggregation function used here is sum, that is, the summation function, and as mentioned above, any operation that can be encapsulated in the form of a function can be used to process derived features.

[0137] The embodiment of the present application provides a configurable derivation mechanism for derivative features, which provides capabilities such as time window aggregation, social behavior subject and object aggregation, and the amount of SQL code can be reduced to one tenth of the original at most.

[0138] Based on high-quality, traceable, and well-defined basic features, business developers can efficiently create derivative features quickly and conveniently through development tools to achieve customized rich feature data usage scenarios. At the same time, the calculation method of basic features naturally avoids the problem of cross-layer data retrieval, which is a good data warehouse layering specification and avoids the problem of data warehouse degradation.

[0139] Optionally, the derived feature data table structure is generated by the feature derivation function module, and there is no need to manually define the table structure.

[0140] Optionally, the method 400 further includes:

[0141] The terminal device acquires a target variable and an update time of the target variable;

[0142] The terminal device determines at least one feature associated with the target variable;

[0143] The terminal device arranges data corresponding to each feature of the at least one feature in chronological order;

[0144] The terminal device obtains the last feature measurement value whose data update time corresponding to each feature is earlier than the update time of the target table variable;

[0145] The terminal device predicts the target variable based on the last feature measurement value of each feature.

[0146] Specifically, after the data warehouse model is established, it can be used for applications such as machine learning or data analysis. When performing applications such as machine learning or data analysis, feature splicing is required. The purpose of feature splicing is to bring together the features of a certain type of entity that is related to the machine learning training goal or analysis problem. For example, when predicting user churn, it is often necessary to combine multiple features such as the user's age, gender, weekly active days, and daily average online time into a feature vector. Since features are measures of the characteristics of an entity at a certain point in time, feature splicing that is misplaced in time will lead to feature leakage (Feature Leakage) and training-serving skew problems (Training-Serving Skew). Especially when there are differences in feature update frequencies, feature splicing is more prone to errors. With Figure 7 For example, Figure 7 is a schematic structural diagram of a feature splicing provided in an embodiment of the present application. Figure 7 As shown in the figure, the daily features of the entity on the third day have not been updated, so the daily features of the previous day are used, while the real-time features can use the measurement values ​​of the current day. In short, during the feature splicing process, it is necessary to ensure that all features are earlier than the target variables in training or analysis and are the latest (Latest), otherwise, future information may be leaked into the machine model training data, resulting in inflated machine model performance. The feature splicing capability ensures the consistency of online and offline computing logic, thereby avoiding the problem of training-service skew.

[0147] In order to understand the feature splicing process more clearly, Figure 8 Please specify. Figure 8 As shown, Figure 8 It is a schematic flow chart of a method 600 for feature stitching provided by the present application, and the execution body of the method 600 is the same as the aforementioned methods 400 and 500. The method 600 may include steps S610 to S670.

[0148] S610, the program starts.

[0149] S620: The terminal device obtains a target variable and an update time of the target variable.

[0150] S630, the terminal device traverses the feature list to determine at least one feature associated with the target variable. The feature list refers to the feature list corresponding to the target variable, and the feature list can be a basic feature data table or a derived feature data table. For example, the target variable is user churn, and the update time of the target variable is the same day. When predicting user churn, it is often necessary to combine multiple features such as the user's age, gender, weekly active days, and daily average online time into a feature vector. Then the feature list corresponding to the target variable is a list of multiple features including the user's age, gender, weekly active days, and daily average online time. The multiple features may be in the same basic feature data table or in different basic feature data tables.

[0151] S640: The terminal device arranges the data corresponding to each feature of the at least one feature in chronological order.

[0152] S650: The terminal device obtains the last feature measurement value whose data update time corresponding to each feature is earlier than the update time of the target table variable.

[0153] S660: The terminal device determines whether the traversal of the feature list is completed. If the traversal of the feature list is not completed, S630 is continued.

[0154] S670, when the traversal of the feature list is completed, the feature splicing program ends.

[0155] The feature splicing process provided in the embodiment of the present application can be encapsulated into an application programming interface (API). The API is a set of predefined functions whose purpose is to provide applications and developers with the ability to access a set of routines based on certain software or hardware without having to access the source code or understand the details of the internal working mechanism.

[0156] Data analysis and machine learning model developers can complete feature splicing in a configurable way. Taking churn prediction as an example, the following code means: taking churn as the target variable, the latest values ​​of gender, age, and weekly active days are spliced ​​into a feature vector.

[0157]

[0158] Optionally, the method 400 further includes:

[0159] The terminal device combines the first basic feature data table and the second basic feature data table into a wide table according to the materialized view.

[0160] Specifically, the data warehouse modeling method provided in this application uniformly treats existing dimensions and metrics as entity features. Therefore, column storage, materialized views and other technologies can be used to optimize the execution efficiency of connection operations, and there is no need for support from capabilities such as table pruning.

[0161] Different from the traditional data warehouse modeling method, the data warehouse modeling method proposed in the embodiment of the present application takes features as the core of data modeling, and no longer distinguishes between concepts such as dimensions, metrics, and indicators, thereby greatly simplifying the data model. Different from the traditional data warehouse model system structure, the data warehouse structure formed by the data warehouse modeling method proposed in the embodiment of the present application is as follows: Fig. 9 As shown, Fig. 9 It is a schematic structural block diagram of a data warehouse hierarchical structure proposed in an embodiment of the present application. It includes an ODS layer, a DWD layer and a DWS layer. The ODS layer stores the original data reported by different devices. The DWD layer is responsible for parsing and cleaning the reported data of the ODS layer. The main tasks include data format conversion, missing value and abnormal value processing, etc. The feature-oriented data warehouse modeling method in the embodiment of the present application can make the DWD layer thinner. The DWS layer is simplified into two sub-layers, namely the basic feature sub-layer and the derived feature sub-layer. As mentioned above, the basic feature is a direct measurement of a certain attribute of an entity; and the derived feature is generated by the basic feature through complex mathematical operations, vectorization, machine learning model prediction and other operations. Derived features are a wide range of feature construction methods that can further refine and enhance the amount of basic feature information and describe entities at a higher level of abstraction. From the perspective of the data model, the basic feature sub-layer directly depends on the ODS layer data source, and the derived feature sub-layer is the downstream task of the basic feature sub-layer. Finally, to facilitate data analysis and machine learning task applications, basic features and derived features can be divided into different topics. For example, basic features can be divided into basic features of chat behavior, basic features of comment behavior, and basic features of gift behavior. Derived features can be divided into time window aggregation features and social behavior derived features.

[0162] In summary, a method of data warehouse modeling proposed in the embodiment of the present application can integrate various data sources, provide a more comprehensive and in-depth perspective for data governance, and thus help enterprises improve data quality. The method can also be used in the field of data analysis and causal inference to help enterprises better understand data and discover the connections and meanings between data. In addition, the method can also assist in the development of machine learning algorithms, improve development efficiency, and reduce development costs. Specifically, the method can be widely used in various mobile games, online communities, online videos, online music, news, and information flow businesses, helping enterprises analyze user behavior habits, social tendencies, content preferences, reading interests, etc. The following are illustrated by case studies:

[0163] Data analysis: The purpose of data analysis is to discover patterns and trends in data through statistical methods, so as to provide a basis for decision-making. However, due to the uneven quality of data, data analysis often involves a lot of tedious data collection and cleaning work. On the other hand, complex data models, hierarchical relationships between dimension tables and other factors also increase the burden of data analysis. The use of feature-oriented data models can greatly simplify data extraction, cleaning and conversion. Taking the data analysis work of an online game community as an example, the number of user logins, online time, number of posts on the forum, number of replies, number of likes and other data can all be processed into features, while retention, churn and other data can be regarded as derived features of login behavior. In addition, using the feature derivation capability provided by this solution, time window aggregation features can also be quickly processed, such as weekly active days, monthly active days, etc. The process does not involve code writing, but is implemented in a configured zero-code manner. With rich feature support, statistical analysis tools can be used to complete tasks such as exploratory data analysis, relationship analysis, and churn attribution analysis.

[0164] Causal inference: quantitative analysis of the impact of one variable (independent variable) on another variable (dependent variable). For example, the impact of a new game item or new level on the user's active time. The key challenge of causal inference is to identify and deal with confounding variables. The so-called confounding variable refers to a variable that may affect both the independent variable and the dependent variable. For example, factors such as the user's age, game rank, consumption level, interest preferences, etc. may simultaneously affect the player's purchase and use of new items and their active time. Therefore, causal inference work often involves a lot of confounding variable development work. Without a concise and efficient data model and high-quality data support, confounding variable development will become difficult.

[0165] Machine learning: In traditional machine learning, feature engineering plays a vital role. Its goal is to process raw data into features suitable for model processing, thereby improving the model's prediction accuracy, robustness, and interpretability, and reducing model complexity. With the development of deep learning technology, some of the work in feature engineering has been automated, but the raw data still needs to be cleaned and processed before it can be fed into the model for training. Therefore, entity-level feature construction is still an important part of the machine learning workflow. On the other hand, feature leakage and training-serving skew are also common problems encountered in the deployment and operation of machine learning models. The feature-oriented data model proposed in this solution is very suitable for the needs of machine learning model training and prediction, and almost no additional processing is required. In addition, capabilities such as feature inference and temporal feature splicing are helpful in solving problems such as feature leakage and training-serving skew.

[0166] The specific implementation modes of the present application are described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the specific details in the above implementation modes. Within the technical concept of the present application, the technical solution of the present application can be subjected to a variety of simple modifications, and these simple modifications all belong to the protection scope of the present application. For example, the various specific technical features described in the above specific implementation modes can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present application will not further explain various possible combinations. For another example, the various different implementation modes of the present application can also be arbitrarily combined, and as long as they do not violate the ideas of the present application, they should also be regarded as the contents disclosed in the present application.

[0167] It should also be understood that in the various method embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. It should be understood that these sequence numbers can be interchanged where appropriate, so that the embodiments of the present application described can be implemented in an order other than those shown or described.

[0168] Combination of the above Figures 1 to 9 , describes in detail the method embodiment of the present application, and the following is combined with Fig.10 and Fig.11 , describe in detail the device embodiments of the present application.

[0169] Fig.10 is a schematic block diagram of an apparatus 700 provided in an embodiment of the present application, and the apparatus 700 can implement the function of the joint scheduling device in the above method. Fig.10 As shown, the apparatus 700 may include a processing unit 710 .

[0170] The processing unit 710 is configured to parse a first data table reported by a first device.

[0171] The processing unit 710 is also used to determine the first basic feature data table that needs to be updated based on the data of the first feature included in the parsed first data table and the feature registration table, wherein the first basic feature data table includes at least one set of feature data of the first feature, each set of feature data in the at least one set of feature data includes time, entity and measurement value, the first feature is a registered feature, and the first data table is the only dependent data table of the at least one basic feature data table.

[0172] The processing unit 710 is further configured to update the first basic feature data table according to the data of the first feature included in the parsed first data table.

[0173] In some embodiments, the processing unit 710 is further configured to:

[0174] If the first data table after parsing also includes data of a second feature, and the second feature is an unregistered feature, register the second feature,

[0175] The registering the second feature includes:

[0176] In the feature registration table, a storage relationship between the data of the second feature and the first basic feature data table is established; or,

[0177] In the feature registration table, a storage relationship between the data of the second feature and the second basic feature data table is established.

[0178] In some embodiments, determining the first basic feature data table that needs to be updated according to the data of the first feature included in the parsed first data table and the feature registration table includes:

[0179] Determining a subject corresponding to the first feature according to the first feature and the feature registration table;

[0180] The first basic feature data table having the same subject as that corresponding to the first feature is determined.

[0181] In some embodiments, the method, the processing unit 710 is further configured to:

[0182] According to the derived feature list, obtain the derived feature configuration information;

[0183] According to the derived feature configuration information, basic feature data to be derived is obtained;

[0184] A derived feature data table is generated according to the feature aggregation function and the basic feature data to be derived.

[0185] In some embodiments, the derived feature configuration information includes at least one of the following parameters: a derived basic feature, a derived target entity, a derived time, and a derived location.

[0186] In some embodiments, the feature aggregation function includes any one of the following functions:

[0187] Functions corresponding to mathematical operations, vectorization functions, and machine learning models.

[0188] In some embodiments, the processing unit 710 is further configured to:

[0189] Obtaining a target variable and an update time of the target variable;

[0190] determining at least one feature associated with the target variable;

[0191] Arrange the data corresponding to each feature of the at least one feature in chronological order;

[0192] Obtain the last feature measurement value whose data update time corresponding to each feature is earlier than the update time of the target table variable;

[0193] The target variable is predicted based on the last feature measurement value of each feature.

[0194] In some embodiments, the processing unit 710 is further configured to:

[0195] According to the materialized view, the first basic feature data table and the second basic feature data table are combined into a wide table.

[0196] In some embodiments, the first device comprises a client device and / or a server.

[0197] In some embodiments, the apparatus 700 may include a receiving unit 720, and the receiving unit 720 is configured to receive a first data table reported by the first device.

[0198] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, no further description is given here. Specifically, when the data processing device 700 in this embodiment may correspond to the execution subject of the method 400 of the embodiment of the present application, the aforementioned and other operations and / or functions of each module in the device 700 are respectively to implement Figure 4 For the sake of brevity, the corresponding processes of each method in are not repeated here.

[0199] The above describes the device and system of the embodiment of the present application from the perspective of the functional module in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software form instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or a combination of hardware and software modules in the decoding processor to perform. Optionally, the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in conjunction with its hardware.

[0200] like Fig.11 It is a schematic block diagram of an electronic device 800 provided in an embodiment of the present application.

[0201] like Fig.11 As shown, the electronic device 800 may include:

[0202] The memory 810 and the processor 820, the memory 810 is used to store the computer program and transmit the program code to the processor 820. In other words, the processor 820 can call and run the computer program from the memory 810 to implement the method in the embodiment of the present application.

[0203] For example, the processor 820 may be configured to execute the steps of each execution subject in the above method 300 according to the instructions in the computer program.

[0204] In some embodiments of the present application, the processor 820 may include but is not limited to:

[0205] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0206] In some embodiments of the present application, the memory 810 includes but is not limited to:

[0207] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0208] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 810 and executed by the processor 820 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device 800.

[0209] Optionally, the electronic device 800 may further include:

[0210] The communication interface 830 may be connected to the processor 820 or the memory 810 .

[0211] The processor 820 may control the communication interface 830 to communicate with other devices, specifically, to send information or data to other devices, or to receive information or data sent by other devices. Exemplarily, the communication interface 830 may include a transmitter and a receiver. The communication interface 830 may further include an antenna, and the number of antennas may be one or more.

[0212] It should be understood that the various components in the electronic device 800 are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0213] According to one aspect of the present application, a communication device is provided, including a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory, so that the encoder executes the method of the above method embodiment.

[0214] According to one aspect of the present application, a computer storage medium is provided, on which a computer program is stored, and when the computer program is executed by a computer, the computer can perform the method of the above method embodiment. In other words, the present application embodiment also provides a computer program product containing instructions, and when the instructions are executed by a computer, the computer can perform the method of the above method embodiment.

[0215] According to another aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of the above method embodiment.

[0216] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server, or data center. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (digital video disc, DVD)), or a semiconductor medium (e.g., a solid state drive (solid state disk, SSD)), etc.

[0217] It should be understood that in the embodiment of the present application, "B corresponding to A" means that B is associated with A. In one implementation, B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based only on A, and B can also be determined based on A and / or other information.

[0218] In the description of the present application, unless otherwise specified, "at least one" means one or more, and "plurality" means two or more than two. In addition, "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0219] It should also be understood that the first, second, etc. descriptions appearing in the embodiments of the present application are only used for illustration and distinction of the description objects, without any distinction of order, nor do they indicate any special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.

[0220] It should also be understood that the specific features, structures or characteristics related to the embodiments in the specification are included in at least one embodiment of the present application. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.

[0221] In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or device.

[0222] It is understandable that in the specific implementation of this application, user information and other related data may be involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0223] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0224] In the several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the module is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0225] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. For example, each functional module in each embodiment of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0226] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for data warehouse modeling, characterized in that: The method comprises: Parsing data reported by the first device to generate a first data table; Determine a first basic feature data table that needs to be updated according to the data of the first feature included in the first data table and the feature registration table, wherein the first basic feature data table includes at least one set of feature data of the first feature, each set of feature data in the at least one set of feature data includes time, entity and measurement value, the first feature is a registered feature, and the first data table is the only dependent data table of the at least one basic feature data table; The first basic feature data table is updated according to the data of the first feature.

2. The method according to claim 1, characterized in that The method further comprises: If the first data table also includes data of a second feature, and the second feature is an unregistered feature, register the second feature, The registering the second feature includes: In the feature registration table, a storage relationship between the data of the second feature and the first basic feature data table is established; or, In the feature registration table, a storage relationship between the data of the second feature and the second basic feature data table is established.

3. The method according to claim 1, characterized in that The determining the first basic feature data table to be updated according to the data of the first feature included in the first data table and the feature registration table includes: Determining a subject corresponding to the first feature according to the first feature and the feature registration table; A basic feature data table having the same subject as that corresponding to the first feature is determined as the first basic feature data table.

4. The method according to claim 1, characterized in that: The method further comprises: According to the derived feature list, obtain the derived feature configuration information; According to the derived feature configuration information, basic feature data to be derived is obtained; A derived feature data table is generated according to the feature aggregation function and the basic feature data to be derived.

5. The method according to claim 4, characterized in that The derived feature configuration information includes at least one of the following parameters: The basic features of the deduction, the target entity of the deduction, the time of the deduction and the place of the deduction.

6. The method according to claim 4, characterized in that The feature aggregation function includes any one of the following functions: Functions corresponding to mathematical operations, vectorization functions, and machine learning models.

7. The method according to claim 1, characterized in that The method further comprises: Obtaining a target variable and an update time of the target variable; determining at least one feature associated with the target variable; Arrange the data corresponding to each feature of the at least one feature in chronological order; Obtain the last feature measurement value whose data update time corresponding to each feature is earlier than the update time of the target table variable; The target variable is predicted based on the last feature measurement value of each feature.

8. The method according to claim 2, characterized in that: The method further comprises: According to the materialized view, the first basic feature data table and the second basic feature data table are combined into a wide table.

9. The method according to claim 1, characterized in that: The first device includes a client device and / or a server.

10. A data modeling device, characterized in that: include: A processing unit, which parses the data reported by the first device and generates a first data table; The processing unit is further used to determine a first basic feature data table that needs to be updated according to the data of the first feature included in the first data table and the feature registration table, wherein the first basic feature data table includes at least one set of feature data of the first feature, each set of feature data in the at least one set of feature data includes time, entity and measurement value, the first feature is a registered feature, and the first data table is the only dependent data table of the at least one basic feature data table; The processing unit is further configured to update the first basic feature data table according to the data of the first feature.

11. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores instructions, and when the processor runs the instructions, the processor executes the method according to any one of claims 1 to 9.

12. A computer storage medium, characterized in that: The method comprises instructions which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 9.

13. A computer program product, characterized in that The method comprises a computer program code, and when the computer program code is executed by an electronic device, the electronic device executes the method according to any one of claims 1 to 9.