Feature generation system and method, computer equipment and computer readable storage medium

By combining generation modules, synchronization modules and sharing modules in the feature production system, the problems of low feature generation efficiency and waste of storage resources in the existing system are solved, and efficient feature generation, synchronization and sharing are achieved.

CN120196942APending Publication Date: 2025-06-24GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311794268.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing feature production system only distinguishes off-line features and online features, lacks finer-grained feature division and feature association management, resulting in low feature generation efficiency and waste of storage resources.

Method used

By generating the features based on the meta information of the feature group and the meta information of at least two features, it is ensured that the similarity difference between any two features is smaller than the threshold value, thereby improving the feature generation efficiency. The synchronization module synchronizes the generated features to the repository, and the sharing module shares features through the repository to reduce waste of storage resources.

Benefits of technology

It improves the efficiency of feature generation, reduces the waste of storage resources, and realizes efficient synchronization and sharing of features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196942A_ABST
    Figure CN120196942A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a feature generation system and method, computer equipment and a computer readable storage medium, which are used for generating at least two features according to meta-information of a feature group and meta-information of the at least two features, improving the feature generation efficiency, synchronizing and sharing the at least two features, and reducing the waste of storage resources. The feature generation system comprises a generation module used for generating at least two features according to meta-information of a feature group and meta-information of at least two features, the feature group comprising the at least two features, and a difference value of similarity of any two features in the at least two features being smaller than a difference value threshold value; the synchronization module is used for synchronizing the at least two features into a storage library; and the sharing module is used for carrying out feature sharing through the storage library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of deep learning, and particularly to a feature generation system, method, computer device, and computer-readable storage medium. Background Art

[0002] In one implementation, a machine learning feature production system is provided, including: a feature management system that determines whether there are duplicate features, and writes the offline features to be accessed into the feature data warehouse or writes the online features to be accessed into the feature message queue when determining that there are no duplicates; a feature distribution cluster, where each feature distributor listens to the feature message queue respectively, distributes the features to be distributed to each feature subscription repository after listening to the message, regularly loads feature metadata from the feature metadata repository to detect whether there are changes, and processes the feature metadata to be changed when there are changes; a task scheduling system that regularly synchronizes the features in the feature data warehouse to the feature message queue; a feature access software development kit (SDK) that provides an access port for online feature producers.

[0003] However, the above-mentioned feature production system only distinguishes between offline features and online features, the division of features is relatively general, and the association between features is not managed. Summary of the Invention

[0004] Embodiments of the present application provide a feature generation system, method, computer device, and computer-readable storage medium, which can generate at least two features according to the meta-information of a feature group and the meta-information of at least two features, and the difference between the similarities of any two features among the at least two features is less than a difference threshold, thereby improving the efficiency of generating features, and can also synchronize and share at least two features to reduce waste of storage resources.

[0005] The first aspect of the present application provides a feature generation system, which may include:

[0006] A generation module, configured to generate at least two features according to the meta-information of a feature group and the meta-information of at least two features, the feature group includes the at least two features, and the difference between the similarities of any two features among the at least two features is less than a difference threshold;

[0007] A synchronization module, configured to synchronize the at least two features to a repository;

[0008] A sharing module, configured to perform feature sharing through the repository.

[0009] The second aspect of the present application provides a feature generation method, which is applied to a feature generation system, and the method may include:

[0010] Generate at least two features according to the meta-information of a feature group and the meta-information of at least two features, where the feature group includes the at least two features, and the difference between the similarities of any two of the at least two features is less than a difference threshold;

[0011] Synchronize the at least two features to a repository;

[0012] Perform feature sharing through the repository.

[0013] A third aspect of the present application provides a computer device, which may include:

[0014] A memory storing executable program code;

[0015] A processor coupled to the memory;

[0016] The processor is configured to correspondingly execute the method described in the second aspect of the present application.

[0017] Another aspect of the embodiments of the present application provides a computer-readable storage medium, including instructions, which when running on a processor, cause the processor to execute the method described in the second aspect of the present application.

[0018] Another aspect of the embodiments of the present application discloses a computer program product, which when running on a computer, causes the computer to execute the method described in the second aspect of the present application.

[0019] Another aspect of the embodiments of the present application discloses an application publishing platform, which is used to publish a computer program product, wherein when the computer program product runs on a computer, it causes the computer to execute the method described in the second aspect of the present application.

[0020] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:

[0021] In the embodiments of the present application, the provided feature generation system may include: a generation module, configured to generate at least two features according to the meta-information of a feature group and the meta-information of at least two features, where the feature group includes the at least two features, and the difference between the similarities of any two of the at least two features is less than a difference threshold; a synchronization module, configured to synchronize the at least two features to a repository; a sharing module, configured to perform feature sharing through the repository. It is possible to generate at least two features according to the meta-information of a feature group and the meta-information of at least two features, improve the efficiency of generating features, and can also synchronize and share at least two features, reducing the waste of storage resources. Description of the Drawings

[0022] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments and the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained based on these drawings.

[0023] Figure 1A A schematic diagram of a feature generation system applied to an embodiment of the present application;

[0024] Figure 1B Another schematic diagram of a feature generation system applied to an embodiment of the present application;

[0025] Figure 1C Another schematic diagram of a feature generation system applied to an embodiment of the present application;

[0026] Figure 1D Another schematic diagram of a feature generation system applied to an embodiment of the present application;

[0027] Figure 1E Another schematic diagram of a feature generation system applied to an embodiment of the present application;

[0028] Figure 1F Another schematic diagram of a feature generation system applied to an embodiment of the present application;

[0029] Figure 2 A schematic diagram of an embodiment of a feature generation method in an embodiment of the present application;

[0030] Figure 3A A schematic diagram of a meta-information registration page for a feature group in an embodiment of the present application;

[0031] Figure 3B A schematic diagram of a meta-information registration page for a feature in an embodiment of the present application;

[0032] Figure 3C A schematic diagram of a task data pipeline for creating a new or selecting an existing created feature group in an embodiment of the present application;

[0033] Figure 3D A schematic diagram of the interface of a task data pipeline pipeline in an embodiment of the present application;

[0034] Figure 3E A schematic diagram of the configuration information of a sub-task in an embodiment of the present application;

[0035] Figure 3F A schematic diagram of the process of a feature synchronization task and a verification task in an embodiment of the present application;

[0036] Figure 3G A schematic diagram of monitoring and data embedding for a sub-task in an embodiment of the present application;

[0037] Figure 3H It is a schematic diagram of the operation duration and the number of internal restarts of a job in task monitoring in an embodiment of the present application;

[0038] Figure 3I It is a schematic diagram of the Source metric in task monitoring in an embodiment of the present application;

[0039] Figure 3J It is another schematic diagram of the Source metric in task monitoring in an embodiment of the present application;

[0040] Figure 3K It is a schematic diagram of the Sink metric in task monitoring in an embodiment of the present application;

[0041] Figure 3L It is another schematic diagram of the Sink metric in task monitoring in an embodiment of the present application;

[0042] Figure 4 It is a schematic diagram of an embodiment of a computer device in an embodiment of the present application;

[0043] Figure 5 It is another schematic diagram of an embodiment of a computer device in an embodiment of the present application. Detailed implementation manners

[0044] The embodiment of the present application provides a feature generation system, method, computer device and computer-readable storage medium, which can generate at least two features according to the meta-information of a feature group and the meta-information of at least two features, and the difference between the similarities of any two features among the at least two features is less than a difference threshold, so that the efficiency of generating features can be improved, and at least two features can be synchronized and shared, reducing the waste of storage resources.

[0045] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all should belong to the scope protected by the present application.

[0046] Flink, a new generation of big data processing engine (A new generation of big data processing engine), is also called a distributed streaming data flow engine (Distributed streaming data flow engine).

[0047] Feature, an individual measurable property or characteristic of a phenomenon

[0048] Feature metadata, the relevant data (Metadata) that describes the feature, such as the type of the feature, the type of feature identity (Identity, ID), the repository, the business classification, etc.

[0049] Message queue, a container that saves messages during message transmission

[0050] Parker, a key-value storage system developed by OPPO

[0051] Kafka, a high-throughput distributed publish-subscribe messaging system (Distributed event streaming platform), is an open-source stream processing platform developed by the Apache Software Foundation and written in Scala and Java. Kafka is a high-throughput distributed publish-subscribe messaging system that can process all action stream data of consumers on websites. Such actions (web page browsing, searching, and other user actions) are a key factor in many social functions on modern networks. These data are usually solved by processing logs and log aggregation due to throughput requirements. For log data and offline analysis systems like Hadoop, but with the limitation of real-time processing requirements, this is a feasible solution. The purpose of Kafka is to unify online and offline message processing through Hadoop's parallel loading mechanism and also to provide real-time messages through a cluster.

[0052] Redis (Remote Dictionary Server), that is, remote dictionary service, is an open-source log-type, key-Value database written in ANSI C language, supporting networks, and can be memory-based or persistent, and provides application programming interfaces (Application Programming Interface, API) in multiple languages.

[0053] The Hadoop Distributed File System (HDFS) is a distributed file system designed to run on commodity hardware. It has many commonalities with existing distributed file systems. At the same time, its differences from other distributed file systems are also obvious. HDFS is a highly fault-tolerant system suitable for deployment on inexpensive machines. HDFS can provide high-throughput data access and is very suitable for applications on large-scale datasets. HDFS relaxes some of the Portable Operating System Interface (POSIX) constraints to achieve the purpose of streaming access to file system data.

[0054] Hive is a data warehouse tool based on Hadoop for data extraction, transformation, and loading. It is a mechanism for storing, querying, and analyzing large-scale data stored in Hadoop. The Hive data warehouse tool can map structured data files into a database table and provide Structured Query Language (SQL) query capabilities. It can transform SQL statements into programming model (MapReduce) tasks for execution. The advantage of Hive is its low learning cost. It can achieve fast MapReduce statistics through SQL-like statements, making MapReduce simpler without having to develop specialized MapReduce applications. Hive is very suitable for statistical analysis of data warehouses.

[0055] Structured Query Language (SQL) is a database language with multiple functions such as data manipulation and data definition. This language has an interactive feature, which provides great convenience for users. Database management systems should make full use of the SQL language to improve the working quality and efficiency of computer application systems. SQL language can not only be independently applied to terminals but also serve as a sublanguage to provide effective assistance for other programming. In this program application, SQL can optimize program functions together with other programming languages, thereby providing users with more comprehensive information.

[0056] MySQL is a relational database management system. Relational databases store data in different tables instead of putting all the data in a large warehouse, which increases speed and flexibility. The SQL language used by MySQL is the most commonly used standardized language for accessing databases. MySQL software adopts a dual-licensing policy, divided into community edition and commercial edition. Due to its small size, fast speed, low total cost of ownership, especially the feature of open source, MySQL is generally chosen as the website database for the development of small and medium-sized and large websites.

[0057] Spark, a distributed data engine.

[0058] PySpark, Spark APIs for Python developers.

[0059] StreamGraph, a dataflow graph, a Flink task execution plan. StreamGraph is a dataflow graph generated by the client according to the Flink-API and is an encapsulation of the Flink task execution process topology graph. When the Environment object calls the execute method, the written program (data processing process) will be transformed into a StreamGraph.

[0060] JobGraph, a job graph, a Flink task execution plan (Represents a Flink dataflow execution plan). JobGraph is optimized based on StreamGraph (including setting checkpoints, subset (slot) grouping strategies, memory occupancy, etc.). The most important thing is to chain multiple eligible dataflow nodes (StreamNode) together as one node, reducing the consumption of serialization, deserialization, and transmission required for data flow between nodes.

[0061] JobManager, a job management node. Generally refers to the Flink job management node.

[0062] Grafana, a data detection alarm system for querying, visualizing, and alerting. Through Grafana, user permissions can be managed, data analyzed, viewed, exported, and alerts set, etc.

[0063] QPS (Queries Per Seconds), the number of requests per second. Usually, it refers to the ability to process requests, that is, it represents the number of requests processed by a server within a unit of time, with the unit of times / second. Through QPS, the maximum access traffic that a system can withstand under different configurations can be roughly estimated, and it is one of the metrics used to evaluate the performance of the backend server. When the load is high, a single server may not be able to meet the business requirements, and load balancing and other methods can be used to expand the capacity to solve the problem.

[0064] Metrics is a Java library that can provide you with unparalleled insights into the running of your code. It was developed by Yammer and is used to detect the health of backend services on the JVM. Metrics provides a powerful set of tools for measuring the behavior of key components in your production environment.

[0065] In one implementation, the feature production system only provides the SDK for users to access features, and finally imports online and offline features into the repository that subscribes to this feature. This feature production system does not have the functions of feature production and feature loading samples. In the machine learning scenario, data mining engineers use big data-related components to develop and produce online and offline features, and then import the features into various repositories. When generating online or offline samples, the corresponding features are retrieved from the repository and loaded into the samples. Online and offline feature development is the upstream of the feature source, and feature loading samples is one of the downstream of feature usage.

[0066] The feature production system only differentiates between offline features and online features, without further finer-grained division. At the same time, the association between features is not managed. For example, data mining engineers produce 3 user features, namely the top 3 (TOP3) of the article first-level categories clicked by users in the recent 6 hours, the TOP3 of the article second-level categories clicked in the recent 6 hours, and the TOP3 of the article keywords clicked in the recent 6 hours. These 3 features have specific similar meanings and production logics and can be grouped into one feature group.

[0067] The feature verification of the feature production system is to verify whether the feature data accessed by the SDK is consistent with the feature metadata defined by the feature, and there is no further comparison of the matching rate between the feature data in the repository and the feature data accessed by the SDK. Because during the process of importing features into the repository through the message queue, message loss may occur.

[0068] The above technology realizes feature sharing by synchronizing the features to be shared to the repository that subscribes to this feature, which will cause duplicate storage in the underlying repository and waste resources.

[0069] The feature generation system provided by the embodiments of the present application can be applied to a recommendation algorithm platform. The feature generation system provides the capabilities of metadata management, feature sharing, and feature generation for features. It supports users to quickly process data, generate features, synchronize features, share features, and use features, etc.

[0070] As Figure 1A shown, it is a schematic diagram of the feature generation system applied in the embodiments of the present application. The feature generation system may include a generation module 101, a synchronization module 102, and a sharing module 103.

[0071] The generation module 101 is configured to generate at least two features according to the metadata of a feature group and the metadata of at least two features, where the feature group includes the at least two features, and the difference between the similarities of any two features in the at least two features is less than a difference threshold;

[0072] The synchronization module 102 is configured to synchronize the at least two features to a repository;

[0073] The sharing module 103 is configured to perform feature sharing through the repository.

[0074] The brief introduction of the calculation of similarity is as follows:

[0075] Regarding the calculation of similarity, several existing basic methods are based on vectors (Vectors), which actually means calculating the distance between two vectors. The closer the distance, the greater the similarity. In the recommendation scenario, in the two-dimensional matrix of user and item preferences, the preferences of a user for all items can be used as a vector to calculate the similarity between users, or the preferences of all users for a certain item can be used as a vector to calculate the similarity between items. Commonly used similarity calculation methods may include: Pearson Correlation Coefficient, Euclidean Distance, Cosine Similarity, Spearman's rank correlation coefficient, log-likelihood similarity, Manhattan distance, etc.

[0076] The feature generation system provided by the embodiments of the present application can generate at least two features according to the metadata of a feature group and the metadata of at least two features, improve the efficiency of generating features, and can also synchronize and share at least two features, reducing the waste of storage resources.

[0077] Optionally, the meta information of the feature group includes at least one of: feature group identifier, feature group name, service classification, feature type, real-time type, online storage information, offline storage information, and whether the feature information is public;

[0078] The meta information of the at least two features further includes at least one of: feature identifier, feature name, feature group, service classification, feature type, real-time type, online storage information, offline storage information, data type, feature identifier type, and whether it is public.

[0079] Optionally, as Figure 1B shown, it is another schematic diagram of the feature generation system applied in the embodiment of the present application. The feature generation system may further include a registration module 104,

[0080] The registration module 104 is configured to respond to an operation of a user registering the meta information of a feature group and features, and obtain the meta information of the feature group and the meta information of the at least two features.

[0081] Optionally, the generation module 101 is specifically configured to determine a task data pipeline of the feature group according to the meta information of the feature group, where the task data pipeline includes at least one subtask; obtain configuration information of the at least one subtask; and run the at least one subtask according to the configuration information of the at least one subtask to generate the at least two features; where the configuration information of the at least one subtask includes a feature identifier or a feature name, and the meta information of the at least two features includes at least one of the feature name and the feature name.

[0082] In the embodiment of the present application, after the meta information of the feature group is registered, the task data pipeline (pipeline) of the feature group can be determined because the task data pipeline (pipeline) of the feature group includes at least one subtask; the user can configure at least two feature subtasks to obtain the configuration information of the at least one subtask; and then run at least one subtask according to the configuration information of at least one subtask to automatically generate at least two features, improving the efficiency of generating features. The subtask can also be referred to as a feature development subtask.

[0083] Optionally, the configuration information of the at least one subtask further includes at least one of: task name, task type, engine type, input address, output address, and task model.

[0084] Optionally, the input address includes the address of the input data, and the output address includes the address of the output data. The embodiment of the present application provides the configuration information of the subtask, and each subtask includes the address of the input data and the address of the output data when running the subtask, improving the feasibility of the solution.

[0085] Optionally, the generation module 101 is specifically configured to obtain input data according to the configuration information of the at least one subtask; run the at least one subtask according to the input data to generate at least two features.

[0086] Optionally, as Figure 1C shown, it is another schematic diagram of the feature generation system applied in the embodiment of the present application. The feature generation system may further include: a cleaning module 105,

[0087] The cleaning module 105 is configured to clean the input data to obtain the cleaned input data;

[0088] The generation module 101 is specifically configured to run the at least one subtask according to the cleaned input data to generate at least two features.

[0089] In the embodiment of the present application, the input data can be cleaned to ensure the reliability of the input data.

[0090] Optionally, the generation module 101 is specifically configured to create a task data pipeline for the feature group according to the meta-information of the feature group, or select an already created task data pipeline for the feature group.

[0091] In the embodiment of the present application, two implementation manners for determining a task data pipeline according to the meta-information of a feature group are provided. One is to create a task data pipeline corresponding to the feature group after the meta-information of the feature group is registered, and the other is to select an already created task data pipeline corresponding to the feature group after the meta-information of the feature group is registered.

[0092] Optionally, as Figure 1D shown, it is another schematic diagram of the feature generation system applied in the embodiment of the present application. The feature generation system may further include: a monitoring module 106,

[0093] The monitoring module 106 is configured to monitor the running of the at least one subtask to obtain a monitoring result, where the monitoring result includes at least one of job running duration, internal restart times, input metrics, output metrics, exception parsing metrics, resource usage metrics, and job delay metrics;

[0094] The display module 107 is configured to display the monitoring result.

[0095] In the embodiment of the present application, the running of the at least one subtask can be monitored to facilitate the user to understand various metrics during the running of the subtask.

[0096] Optionally, as Figure 1EAs shown, it is another schematic diagram of the feature generation system applied in the embodiment of the present application. The feature generation system further includes: a verification module 108 and a display module 107.

[0097] The verification module 108 is configured to verify at least two features stored in the repository with the at least two features to obtain a verification result.

[0098] The display module 107 is configured to display the verification result, and the verification result is used to indicate the accuracy rate of data synchronization.

[0099] In the embodiment of the present application, at least two features in the repository and at least two generated features can be verified to obtain a verification result, and then the verification result is displayed to facilitate the user to understand the accuracy rate of data synchronization.

[0100] Optionally, the sharing module 103 is specifically configured to load the target feature from the repository into the sample data pipeline for model training, and the at least two features include the target feature.

[0101] In the embodiment of the present application, the use of generating at least two features is provided, and at least two features can be loaded into the sample data pipeline (pipeline) for model training. Generally speaking, the larger the amount of data in the sample, the greater the reliability.

[0102] The feature generation system in the embodiment of the present application may include a feature center part, a task center part, and a task monitoring part. As Figure 1F shown, it is another schematic diagram of the feature generation system applied in the embodiment of the present application. The feature center part may include a registration module 104; the task center part may include a cleaning module 105, a generation module 101, a synchronization module 102, a sharing module 103, a verification module 108; the task monitoring part may include a monitoring module 106 and a display module 107.

[0103] The registration module 104 can register the meta-information of the feature group and can also register the meta-information of the features included in the feature group. After the meta-information of the feature group is registered, the cleaning module 105 can perform a data cleaning task, that is, clean the input data; the cleaning module 105 generates a module 101 through a logical connection, and the generating module 101 can perform a feature development task, that is, create or select a created task data pipeline (pipeline), and the task data pipeline includes at least one subtask (which can also be called a feature development task). After associating features, features are generated. The synchronization module 102 can automatically or passively initiate a feature synchronization task after the feature development task is completed; the verification module 108 can automatically or passively initiate a feature verification task after the feature synchronization task is completed. The monitoring module 106 can monitor the feature development task and can also perform task alerts through policy configuration according to the monitoring results; the display module 107 can display at least one of the verification results of the feature verification task and the monitoring results.

[0104] The following combines Figures 1A - 1F The feature generation system shown in Figure 2 is used to further illustrate the method embodiments in this application. As

[0105] 201. Generate at least two features according to the meta-information of the feature group and the meta-information of at least two features, where the feature group includes the at least two features, and the difference between the similarities of any two features in the at least two features is less than a difference threshold.

[0106] Optionally, before generating at least two features according to the meta-information of the feature group and the meta-information of at least two features, the method may further include: responding to an operation of a user registering the meta-information of the feature group and the feature, and obtaining the meta-information of the feature group and the meta-information of the at least two features.

[0107] Optionally, the meta-information of the feature group includes at least one of: a feature group identifier, a feature group name, a service classification, a feature type, a real-time type, online storage information, offline storage information, and whether the feature information is public;

[0108] The meta-information of the feature includes at least one of: a feature identifier, a feature name, a feature group, a service classification, a feature type, a real-time type, online storage information, offline storage information, a data type, a feature identifier type, and whether it is public.

[0109] It can be understood that the online storage information may include the address of online data storage, and the offline storage information may include the address of offline data storage.

[0110] Optionally, the feature types may include, but are not limited to: user features, material features, and semantic features.

[0111] Optionally, the real-time types include: offline type and real-time type.

[0112] Optionally, the data types include but are not limited to the following:

[0113] (1) Integer types: byte (byte type), short (short integer type), int (integer type), long (long integer type);

[0114] (2) Floating-point types: float (single-precision floating-point type), double (double-precision floating-point type);

[0115] (3) Character type: char (character type);

[0116] (4) Boolean type: boolean. The boolean type has only two values: true and false, and generally defaults to false.

[0117] Optionally, the feature identification types may include, but are not limited to: user identity (Identity, ID), material ID, and International Mobile Equipment Identity (IMEI).

[0118] The registration of the meta-information of the feature group and features is the basis of feature development. The user registers the meta-information of the feature group and features, and then specifies a task pipeline to implement the development of specific features. Subsequently, the feature group and features loading samples, feature screening, feature sharing, etc. will all use the meta-information of the feature group and features. The embodiments of the present application abstract the concepts of feature group and features, and all similar features can be generated in the task data pipeline (pipeline) of a feature group.

[0119] As Figure 3A shown, it is a schematic diagram of the meta-information registration page of the feature group in the embodiments of the present application. The user can fill in the feature group identifier, feature group name (in Chinese or other languages), business classification, feature type, real-time type, online or offline storage information, feature group description information, and whether the feature information is public.

[0120] Among them, the number of characters of the feature group identifier and feature group name is not limited. Figure 3ATaking the word count limit of the feature group identifier and the feature group name within the range of 0 - 128 as an example for illustration. The business classification indicates the feature group of what kind of business. The feature type can include, but is not limited to, more fine-grained divisions such as user features, material features, semantic features, etc. The real-time type is divided into offline and real-time. Whether the feature information is public can correspond to display the public identifier and the non-public identifier. If the non-public identifier is selected, other services cannot query the meta-information of this feature group. If the public identifier is selected, other services can query the meta-information of this feature group. By selecting whether to open it for other services to use through the public option, the function of feature sharing is realized. At this time, all features of this feature group are stored in a repository. The feature groups selected to be public will be synchronized to the user whether stored online or offline. Whether the user retrieves features or loads feature samples, etc., they are all retrieved from this repository. This avoids the waste of storage resources caused by storing the same feature multiple times.

[0121] As Figure 3B shown, it is a schematic diagram of the meta-information registration page of features in an embodiment of the present application. The user can fill in the feature identifier, feature name (in Chinese or other languages), feature group, business classification, feature type, real-time type, online or offline storage information, data type, feature identifier (Identity, ID) type, feature time to live (Time To Live, TTL), offline feature column, feature default value, feature description information, and whether the feature is public. The difference from the meta-information registered for the feature group is that it is necessary to select the feature group associated with the feature, specify the feature ID type, and select the data type. There is no limit to the word count of the feature identifier and the feature name. Figure 3B Taking the word count limit of the feature identifier and the feature name within the range of 0 - 128 as an example for illustration.

[0122] Optionally, generating at least two features according to the meta-information of the feature group and the meta-information of at least two features, where the feature group includes the at least two features, may include: determining the task data pipeline of the feature group according to the meta-information of the feature group, where the task data pipeline includes at least one subtask; obtaining the configuration information of the at least one subtask; running the at least one subtask according to the configuration information of the at least one subtask to generate the at least two features; where the configuration information of the at least one subtask includes a feature identifier or a feature name, and the meta-information of the at least two features includes at least one of the feature name and the feature name.

[0123] Optionally, running the at least one subtask according to the configuration information of the at least one subtask to generate the at least two features may include: in response to the user's generation operation, running the at least one subtask according to the configuration information of the at least one subtask to generate the at least two features.

[0124] Optionally, the configuration information of the at least one subtask further includes at least one of a task name, a task type, an engine type, an input address, an output address, and a task model.

[0125] Optionally, determining the task data pipeline of the feature group according to the meta information of the feature group may include: creating the task data pipeline of the feature group according to the meta information of the feature group; or, selecting the task data pipeline of the created feature group.

[0126] After registering the meta information of the feature group, or after registering the meta information of the feature group and the feature, the task data pipeline of the feature group can be determined, which can also be called the feature development task data pipeline (pipeline). As Figure 3C shown, it is a schematic diagram of creating or selecting the task data pipeline of the created feature group in an embodiment of the present application. The user can input an identifier, a name, description information, etc. for the task data pipeline of the newly created feature group. Selecting the task data pipeline of the created feature group here can also be understood as associating the task data pipeline of the created feature group.

[0127] Optionally, obtaining the configuration information of the at least one subtask may include: responding to an operation of the user inputting the configuration information of the at least one subtask, and obtaining the configuration information of the at least one subtask; or, receiving the configuration information of the at least one subtask sent by other devices.

[0128] Optionally, the input address includes the address of the input data, that is, the address of the input data before running the target subtask, and the output address includes the address of the output data, that is, the address of the output data after running the target subtask.

[0129] Optionally, the engine type may include but is not limited to Flink1.13, Flink1.14, Spark3.1.2, PySpark3.1.2 types.

[0130] Optionally, the method may further include: responding to a debugging operation of the user on the at least one subtask, and debugging the at least one subtask.

[0131] Optionally, the method may further include: responding to a data preview operation of the user on the at least one subtask, and displaying the relevant data of the at least one subtask.

[0132] After registering the meta-information of the feature group and features, the task pipeline can be used for feature development. The task pipeline includes subtasks of feature development logic. The task pipeline can also include a batch of related subtasks such as data cleaning and dimension table association. The subtasks in the task pipeline are the logical tasks for specifically implementing feature development.

[0133] As Figure 3D shown, it is a schematic diagram of the interface of the task data pipeline in an embodiment of the present application. In Figure 3D the shown, through the subtasks of cleaning the source data (which can be understood as input data) and cleaning and associating the subtasks of associating the data source with material information, and then connecting three subtasks of feature development to produce 3 features. For example: the 5 subtasks are respectively: browser_data_source_parse (browser source data parsing), browser_data_source_join_item (browser source data associated item), 6h_click_top3_category (the top 3 of the first-level categories of the articles clicked in the recent 6 hours), 6h_click_top3_keywords (the top 3 of the keywords of the articles clicked in the recent 6 hours), 6h_click_top3_secondcategory (the top 3 of the second-level categories of the articles clicked in the recent 6 hours). The process template and task template in the interface can support users to drag and select, support creating a task data pipeline or subtasks through the template for rapid iterative development. The subtasks can also support users to drag and connect to represent the logical relationship between the subtasks.

[0134] As Figure 3EAs shown in the figure, it is a schematic diagram of the configuration information of the subtask in the embodiment of the present application. The task name is 6h_click_top3_category, the task type is real-time, the engine type is Flink1.13.2, the input address is Kafka / ads_cpd_ctr_relate_join_clean_hdfs, and the output address is Kafka / game_load_oneplus_hotsearch_sample. That is, the user can select the registered input address and output address, and here it supports repositories of various message queues such as Kafka, Redis, Parker, HDFS, and Hive. The user can also select the engine type, specifically supporting various engine types such as Flink1.13, Flink1.14, Spark3.1.2, and PySpark3.1.2. It is also possible to confirm whether the subtask is a task for generating features. If so, then the associated features to be generated can be selected, such as user_app_download_seqs. After the subtask configuration is completed, the development button can be clicked to enter the task logic development interface.

[0135] In the subtask development interface, the source module (which can also be called the input data module) and the sink module (which can also be called the output data module) can be automatically generated according to the input address and output address in the configuration information of the subtask after the subtask configuration is completed. Since the construction of the entire engine highly abstracts the entire data flow, each module has been built with plug-ins and configuration, and at the same time supports Hive functions and various plug-ins adapted to the business. Therefore, the user only needs to write the logical SQL in the transform module. After development, the online operation can be performed, that is, it can be submitted to the computing cluster to start running at least one subtask. At the same time, the subtask development supports online debugging (debug) and data preview capabilities, which can solve the difficulty of difficult data positioning during the development process.

[0136] In the embodiment of the present application, the multi-engine and highly abstract data processing flow support the user to develop using SQL logic and quickly produce data into features. The development iteration efficiency of a single feature can be quickly improved.

[0137] 202. Synchronize the at least two features to the repository.

[0138] Optionally, the method may further include: synchronizing the at least two features to the repository. In the embodiment of the present application, after generating at least two features, the at least two features can be synchronized to the repository to avoid waste of storage resources. This feature generation system can achieve sharing by sharing the meta-information of the features, and only one copy of the feature data is stored in the underlying repository, saving resources.

[0139] Optionally, synchronizing the at least two features to the repository may include: synchronizing the at least two features to Kafka, and writing the at least two features from Kafka to the repository. In the embodiments of the present application, feature synchronization can be automatically triggered after feature production is completed, or can be initiated by the user. It is first written to Kafka and then synchronized to the online repository, which can effectively protect the online repository and avoid direct interaction between the user and the online repository.

[0140] Optionally, synchronizing the at least two features to Kafka and writing the at least two features from Kafka to the repository may include: using a Flink batch task to write the at least two features to Kafka, and using a Flink streaming task to write the at least two feature data from Kafka to the repository.

[0141] Optionally, the method may further include: storing the start time and end time of synchronizing the at least two features to the repository. In the embodiments of the present application, verification is performed after the feature data is completely imported into the repository, and the start time and end time of the current synchronization to the repository can also be updated, which is convenient for users to monitor the efficiency of data synchronization.

[0142] 203. Share features through the repository.

[0143] Optionally, sharing features through the repository may include: in response to a user's access operation, retrieving a target feature from the repository, where the at least two features include the target feature. An implementation manner of feature sharing is provided in the embodiments of the present application.

[0144] Optionally, the method further includes: loading the target feature from the repository into a sample for model training, where the at least two features include the target feature. Loading the sample in the embodiments of the present application is also an implementation manner of feature sharing.

[0145] Optionally, the method may further include: in the case of using a Flink batch task to write the at least two features to Kafka, appending a first preset number of end messages at the end of the transmission; in the case where the Flink streaming task reads an end message, counting the number of received end messages; and in the case where the number of received end messages is greater than a second preset number, performing the step of verifying the at least two features stored in the repository with the at least two features, where the second preset number is less than the first preset number.

[0146] Optionally, when the Flink streaming task reads an end message, the number of received end messages is counted; when the number of received end messages is greater than a second preset number, the step of verifying at least two features stored in the repository with the at least two features may include: when the Flink streaming task reads an end message, updating the end message to the MySQL database, and when the number of end messages in the MySQL database is greater than the second preset number, performing the step of verifying at least two features stored in the repository with the at least two features.

[0147] Optionally, the method may further include: verifying at least two features stored in the repository with the at least two features to obtain a verification result; displaying the verification result, where the verification result is used to indicate the accuracy rate of the synchronized data.

[0148] Exemplarily, after the features are synchronized to the repository, a feature verification task can be automatically triggered or initiated by the user. As Figure 3F shown, it is a schematic flowchart of the feature synchronization task and the verification task in an embodiment of the present application. After the features are generated, a Flink batch task is driven to write the features into Kafka, and at the same time, a first preset number (for example, 100) of end messages can be appended at the end of the transmission. The downstream Flink streaming task reads the features in Kafka and then writes them into the repository. When the read message is an end message, the end message is updated to the MySQL database. After the Flink batch task is executed, a script task (for example, a Python script task) is triggered to periodically detect the number of end messages received in the MySQL database. Assume that the second preset number is 95. If 96 end messages have been received, a verification task will be executed, that is, at least two features are compared with at least two features stored in the repository, and then the verification result is written into the MySQL database. After that, the verification result can be displayed on the front end. If 96 end messages have not been received, the verification will be delayed. All updates and verification results can be stored in the MySQL database.

[0149] Optionally, the method may further include: monitoring the at least one subtask to obtain a monitoring result, where the monitoring result includes at least one of the job running duration, the number of internal restarts, the input (Source) metrics, the output (Sink) metrics, the exception parsing metrics, the resource usage metrics, and the job delay metrics; displaying the monitoring result.

[0150] Optionally, the input indicators include at least one of the input data volume, QPS, number of column null values, and proportion of column null values; the output indicators include at least one of the output data volume, QPS, number of column null values, and proportion of column null values.

[0151] Task monitoring mainly involves monitoring the subtasks in the above task data pipeline. Figure 3G As shown, it is a schematic diagram of monitoring and tracking points for subtasks in an embodiment of the present application. SQL files (text) can be parsed to obtain operations (Operations). Before the Flink backend engine generates a work graph (JobGraph) and submits it to the job manager (JobManager), it will generate a data flow graph (StreamGraph) according to different transformations (Transformation), parse all Transformations to get operators (Operators), and determine whether it is an instance of an output stream (StreamSink) or an input stream (StreamSource). If it is, a new Operator is generated after adding tracking points (such as Metrics). If not, the Transformation is regenerated. In this way, the JobGraph submitted last will report the data indicators to Grafana at the Source and Sink. In this way, a variety of monitoring alarms can be configured, such as input (Source) indicators (such as the amount of data, QPS, number of empty values ​​in the column, and the proportion of empty values ​​in the column), output (Sink) indicators (such as the amount of data, QPS, number of empty values ​​in the column, and the proportion of empty values ​​in the column) of the Sink, etc.

[0152] Among them, Source indicators are used to monitor and evaluate the quality of input data. Sink indicators are used to monitor and evaluate the quality of output data. Both the null value rate and the percentage of null value columns can be used to monitor data quality.

[0153] For example, online task monitoring mainly includes the above-mentioned input (Source) indicators, output (Sink) indicators, exception analysis indicators, resource usage indicators, job delay and job custom monitoring indicators, etc. The custom monitoring indicator system here provides indicator plug-ins and indicator functions, and users can customize them in SQL logic. Exception analysis indicators are used to monitor and evaluate the quality of upstream input data. Figure 3H As shown, it is a schematic diagram of the job running time and the number of internal restarts in the task monitoring in the embodiment of the present application. Figure 3I As shown, it is a schematic diagram of the Source indicator in the task monitoring in the embodiment of the present application. Figure 3J As shown, it is another schematic diagram of the Source indicator in the task monitoring in the embodiment of the present application. Figure 3KAs shown in the figure, it is a schematic diagram of the Sink metric in task monitoring in the embodiment of the present application. As Figure 3L shown in the figure, it is another schematic diagram of the Sink metric in task monitoring in the embodiment of the present application.

[0154] Exemplarily, the resource usage metrics may include the JobManager heap memory usage rate, the overall average Central Processing Unit (CPU) usage rate of the JobManager, the TaskManager heap memory usage rate, the overall average CPU usage rate of the TaskManager, etc.

[0155] The job latency metrics may include job latency (offset), kafka consumption latency time, kafka write latency time, Hadoop Distributed File System (HDFS) write latency, Redis / Plusar write latency, Redis Lookup, etc.

[0156] The job custom metrics may include counter metrics, QPS metrics, etc.

[0157] The offline task monitoring may include metrics such as the feature name, English name, status of the offline feature data, the start time of the synchronization task, the completion time of writing to kafka, the start time of writing to the repository (abbreviation: storage), the completion time of writing to the storage, the verification matching degree, and the update time consumption. Its data comes from the data reported by the verification task in the task pipeline.

[0158] In the embodiment of the present application, obtain the meta-information of the feature group and the meta-information of at least two features, where the feature group includes the at least two features; determine the task data pipeline of the feature group according to the meta-information of the feature group; generate the at least two features according to the task data pipeline and the meta-information of the at least two features. The task data pipeline of the feature group can be determined according to the meta-information of the feature group, and at least two features can be generated according to the task data pipeline, improving the efficiency of feature generation.

[0159] The feature generation system in the embodiment of the present application can not only generate features, but also load the features into the feature samples for model training, which can improve the construction of the feature generation system. User features, material features, semantic features, etc. can be abstracted in the feature type dimension. In the feature association dimension, a feature group is abstracted. The feature group can include features, providing functions such as a consistent feature task data pipeline (pipeline) for development and loading samples into the feature group.

[0160] The feature generation system in this application is a feature capability system serving the recommendation algorithm, providing comprehensive feature metadata management, feature sharing, and feature production capabilities, supporting recommendation algorithm engineers to quickly process data, produce features, share features, and use features. Starting from feature groups and feature registration, to production, and finally entering the repository, the platform monitors the entire link to ensure the consistency of feature data online and offline.

[0161] Feature development supports multiple engines such as Flink 1.13, Flink 1.14, Spark 3.1.2, and PySpark 3.1.2. At the same time, the engine highly abstracts the data processing process. Users can choose the engine by themselves, configure the Source and Sink modules, and then use SQL for logical development to quickly produce features from data. The same feature group and feature will only be synchronized to one repository, and feature sharing is achieved through the meta-information data of the feature group and feature, avoiding the waste of resources caused by storing the same feature in multiple repositories. By adding buried points by rewriting the engine source code, automatically invoking the verification task to verify the metadata and the repository, and providing personalized monitoring configuration metric plugins and functions, etc., comprehensive full-link data quality detection and verification are provided.

[0162] As Figure 4 shown, it is a schematic diagram of an embodiment of a computer device in an embodiment of this application, which may include a feature generation system as Figures 1A - 1F shown in any one.

[0163] As Figure 5 shown, it is a schematic diagram of another embodiment of a computer device in an embodiment of this application, which may include:

[0164] A memory 501 storing executable program code;

[0165] A processor 502, a display 503 coupled to the memory 501;

[0166] The processor 502 is used to correspondingly execute the following steps:

[0167] Generate at least two features according to the meta-information of the feature group and the meta-information of at least two features, where the feature group includes the at least two features, and the difference between the similarities of any two features in the at least two features is less than a difference threshold;

[0168] Synchronize the at least two features to the repository;

[0169] Perform feature sharing through the repository.

[0170] Optionally, the processor 502 is specifically used to correspondingly execute the following steps:

[0171] In response to an operation on the meta-information of the user registration feature group and features, obtain the meta-information of the feature group and the meta-information of the at least two features.

[0172] Optionally, the processor 502 is specifically configured to perform the following steps correspondingly:

[0173] Determine the task data pipeline of the feature group according to the meta-information of the feature group, where the task data pipeline includes at least one subtask;

[0174] Obtain the configuration information of the at least one subtask;

[0175] Run the at least one subtask according to the configuration information of the at least one subtask to generate the at least two features;

[0176] Wherein, the configuration information of the at least one subtask includes a feature identifier or a feature name, and the meta-information of the at least two features includes at least one of the feature name and the feature name.

[0177] Optionally, the configuration information of the at least one subtask further includes at least one of: a task name, a task type, an engine type, an input address, an output address, and a task model.

[0178] Optionally, the processor 502 is specifically configured to perform the following steps correspondingly:

[0179] Create the task data pipeline of the feature group according to the meta-information of the feature group, or select the task data pipeline of the created feature group.

[0180] Optionally, the processor 502 is further configured to perform the following steps correspondingly:

[0181] Monitor the running of the at least one subtask to obtain a monitoring result, where the monitoring result includes at least one of: job running duration, internal restart times, input metrics, output metrics, exception parsing metrics, resource usage metrics, and job delay metrics;

[0182] The display 503 is further configured to perform the following steps correspondingly:

[0183] Display the monitoring result.

[0184] Optionally, the processor 502 is further configured to perform the following steps correspondingly:

[0185] Verify the at least two features stored in the repository with the at least two features to obtain a verification result;

[0186] The display 503 is further configured to perform the following steps correspondingly:

[0187] Display the verification result, which is used to indicate the accuracy rate of data synchronization.

[0188] Optionally, the processor 502 is further configured to perform the following steps correspondingly:

[0189] Load the target feature from the repository into the sample for model training, where the at least two features include the target feature.

[0190] Optionally, the meta-information of the feature group includes at least one of: feature group identifier, feature group name, service classification, feature type, real-time type, online storage information, offline storage information, and whether the feature information is public;

[0191] The meta-information of the at least two features further includes at least one of: feature identifier, feature name, feature group, service classification, feature type, real-time type, online storage information, offline storage information, data type, feature identifier type, and whether it is public.

[0192] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a Solid State Disk (SSD)), etc.

[0193] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0194] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0195] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0196] In addition, in each embodiment of the present application, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0197] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0198] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A feature generation system, characterized in that, Including: A generation module, configured to generate at least two features according to the meta-information of a feature group and the meta-information of at least two features, where the feature group includes the at least two features, and the difference between the similarities of any two of the at least two features is less than a difference threshold; A synchronization module, configured to synchronize the at least two features to a repository; A sharing module, configured to perform feature sharing through the repository.

2. The feature generation system according to claim 1, wherein The feature generation system further includes: A registration module, configured to obtain the meta-information of the feature group and the meta-information of the at least two features in response to an operation of a user registering the meta-information of the feature group and the features.

3. The feature generation system according to claim 1 or 2, wherein the generation module is specifically configured to determine a task data pipeline of the feature group according to the meta-information of the feature group, where the task data pipeline includes at least one subtask; obtain configuration information of the at least one subtask; and run the at least one subtask according to the configuration information of the at least one subtask to generate the at least two features; wherein the configuration information of the at least one subtask includes a feature identifier or a feature name, and the meta-information of the at least two features includes at least one of the feature name and the feature name.

4. The feature generation system according to claim 3, wherein The configuration information of the at least one subtask further includes at least one of: a task name, a task type, an engine type, an input address, an output address, and a task model.

5. The feature generation system according to claim 3, wherein the generation module is specifically configured to create a task data pipeline of the feature group according to the meta-information of the feature group, or select a task data pipeline of the created feature group.

6. The feature generation system according to claim 3, wherein The feature generation system further includes: A monitoring module, configured to monitor the running of the at least one subtask to obtain a monitoring result, where the monitoring result includes at least one of: a job running duration, an internal restart count, an input metric, an output metric, an exception parsing metric, a resource usage metric, and a job latency metric.

7. The feature generation system according to claim 1 or 2, characterized in that, The feature generation system further includes: A verification module, configured to verify the at least two features stored in the repository with the at least two features to obtain a verification result; A display module, configured to display the verification result, and the verification result is used to indicate the accuracy rate of data synchronization.

8. The feature generation system according to claim 1 or 2, wherein the sharing module is specifically configured to load a target feature from the repository into a sample data pipeline for model training, and the at least two features include the target feature.

9. The feature generation system according to claim 1 or 2, characterized in that, The meta-information of the feature group includes at least one of: a feature group identifier, a feature group name, a service classification, a feature type, a real-time type, an online storage information, an offline storage information, and whether the feature information is public; The meta-information of the at least two features further includes at least one of: a feature identifier, a feature name, a feature group, a service classification, a feature type, a real-time type, an online storage information, an offline storage information, a data type, a feature identifier type, and whether it is public.

10. A feature generation method, characterized in that, The method is applied to a feature generation system, and the method includes: Generate at least two features based on the meta-information of a feature group and the meta-information of at least two features, where the feature group includes the at least two features, and the difference between the similarities of any two of the at least two features is less than a difference threshold; Synchronize the at least two features to a repository; Perform feature sharing through the repository.

11. The method according to claim 10, wherein The generating at least two features based on the meta-information of a feature group and the meta-information of at least two features includes: Determine a task data pipeline of the feature group according to the meta-information of the feature group, where the task data pipeline includes at least one subtask; Obtain configuration information of the at least one subtask; Run the at least one subtask according to the configuration information of the at least one subtask to generate the at least two features; Wherein, the configuration information of the at least one subtask includes a feature identifier or a feature name, and the meta-information of the at least two features includes at least one of the feature name and the feature name.

12. The method according to claim 11, wherein The method further includes: Monitor the running of the at least one subtask to obtain a monitoring result, where the monitoring result includes at least one of a job running duration, an internal restart count, an input metric, an output metric, an exception resolution metric, a resource usage metric, and a job delay metric.

13. The method according to any one of claims 10-12, characterized in that, The method further includes: Verify the at least two features stored in the repository with the at least two features to obtain a verification result; Display the verification result, where the verification result is used to indicate the accuracy rate of data synchronization.

14. A computer device, characterized in that, Includes: A memory storing executable program code; A processor coupled to the memory; The processor is configured to correspondingly execute the method according to any one of claims 10-13.

15. A computer-readable storage medium, characterized in that, Includes instructions that, when running on a processor, cause the processor to execute the method according to any one of claims 10-13.