Feature production method and system, electronic equipment and storage medium
By managing the numbering and attribute information configuration of features, the problems of low consistency and reuse rate in feature production are solved, efficient production and management of features are achieved, and the accuracy and development efficiency of the recommended model are improved.
Patent Information
- Application Number
- CN202311861506.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
The existing technology lacks unified management of features, resulting in abnormal problems such as inconsistency between offline and online features, missing features, and feature failure during feature extraction, reducing feature production efficiency and reuse rate, and increasing the difficulty of developing recommended models.
By obtaining the attribute information and configuration information of the feature to number, generating feature numbers, and performing feature production based on the feature number, attribute information and configuration information, the unified management of features and data set generation is realized.
It improves the reuse rate and production efficiency of features, can quickly check and avoid abnormal features, and improves the accuracy and development efficiency of the recommended model.
Smart Images

Figure CN120234585A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of feature data processing, and particularly to a feature production method, system, electronic device, and storage medium. Background Art
[0002] With the continuous development of Internet technology, recommendation algorithms are particularly important in advertising, search, and recommendation services. Recommendation models mainly rely on extracting useful features from massive data for personalized recommendation and precise positioning. Through feature extraction, raw data can be transformed into sample data that can be used and learned by the recommendation model.
[0003] However, due to the lack of unified management of features in related technologies, abnormal problems such as inconsistent offline and online features, feature missing, and feature invalidation are likely to occur during the feature extraction process. On the one hand, it reduces the efficiency of feature production and affects the accuracy of the recommendation model. On the other hand, due to the lack of feature management, it directly leads to a low reuse rate of features and increases the development difficulty of the recommendation model. Summary of the Invention
[0004] To solve or partially solve the problems existing in related technologies, this application provides a feature production method, system, electronic device, and storage medium, which can achieve unified management of features and improve the reuse rate of features and the efficiency of feature production.
[0005] In the first aspect of this application, a feature production method is provided, and the method includes:
[0006] Create a new feature and obtain the attribute information and configuration information corresponding to the feature;
[0007] Number the feature according to the attribute information and the configuration information to obtain a feature number for the feature;
[0008] Perform feature production according to the feature number, the attribute information, and the configuration information to obtain a feature data set.
[0009] In an embodiment, the attribute information at least includes the feature value range and feature caliber of the feature, and the configuration information includes binding relationship configuration and feature warning configuration. The step of numbering the feature according to the attribute information and the configuration information to obtain a feature number for the feature includes:
[0010] When automatically numbering the feature, associate the feature with the feature value range, the feature caliber, the binding relationship configuration, and the feature warning configuration to generate a feature number for the feature.
[0011] In an embodiment, the method further includes:
[0012] Store the feature number as the representation of the feature in a preset database, where each feature number corresponds to a feature;
[0013] Among them, the feature number is used to identify and extract the configuration information and attribute information corresponding to the feature.
[0014] In one embodiment, the feature production includes offline feature production, the feature dataset includes an offline feature dataset, and the feature production according to the feature number, the attribute information, and the configuration information to obtain a feature dataset includes:
[0015] Obtain historical data according to the attribute information and configuration information corresponding to the feature number;
[0016] Perform the offline feature production on the historical data using the attribute information and the configuration information to generate the offline feature dataset.
[0017] In one embodiment, the feature production includes real-time feature production, the feature dataset includes a real-time feature dataset, and the feature production according to the feature number, the attribute information, and the configuration information to obtain a feature dataset includes:
[0018] Obtain real-time data according to the attribute information and configuration information corresponding to the feature number;
[0019] Perform the real-time feature production on the real-time data using the attribute information and the configuration information to generate the real-time feature dataset.
[0020] In one embodiment, the method further includes:
[0021] Receive a batch processing request and obtain the feature number corresponding to the batch processing request;
[0022] Call the target feature data corresponding to the batch processing request through the feature number, and print the target feature data using a preset feature printing tool.
[0023] In one embodiment, the method further includes:
[0024] Obtain buried point data and the target feature data;
[0025] Integrate the buried point data with the target feature data to generate a training sample.
[0026] A second aspect of the present application provides a feature production system, and the feature production system at least includes a feature number module and a feature production module;
[0027] The feature numbering module is used to create a new feature, obtain the attribute information and configuration information corresponding to the feature, number the feature according to the attribute information and the configuration information, and obtain a feature number for the feature;
[0028] The feature production module is used to perform feature production according to the feature number, the attribute information, and the configuration information, and obtain a feature data set.
[0029] In one embodiment, the feature production system further includes a feature recall module;
[0030] The feature recall module is used to receive a batch processing request; call the target feature data corresponding to the batch processing request through the feature number, and print the target feature data by using a preset feature printing tool.
[0031] In one embodiment, the feature production system further includes a training sample generation module;
[0032] The training sample generation module is used to obtain buried point data and the target feature data; integrate the buried point data and the target feature data to generate a training sample.
[0033] A third aspect of the present application provides an electronic device, including:
[0034] A processor; and
[0035] A memory having executable code stored thereon, which when executed by the processor, causes the processor to execute the method as described above.
[0036] A fourth aspect of the present application provides a computer-readable storage medium having executable code stored thereon, which when executed by a processor of an electronic device, causes the processor to execute the method as described above.
[0037] The technical solution provided by the present application may include the following beneficial effects:
[0038] The solution provided by the present application creates a new feature, obtains the attribute information and configuration information corresponding to the feature, numbers the feature according to the attribute information and the configuration information, obtains a feature number for the feature, and performs feature production according to the feature number, the attribute information, and the configuration information to obtain a feature data set. By numbering according to the attribute information and configuration information corresponding to the feature, the present application realizes the concretization of the abstract feature into a feature number. The representation form of "feature number" not only facilitates the management of features and improves the reuse rate of features, but also can avoid abnormal problems or quickly troubleshoot abnormal features during the feature production process, thereby improving the efficiency of feature production.
[0039] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] By describing the exemplary embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more apparent. Among them, in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.
[0041] Figure 1 is a schematic flowchart of a feature production method shown in an embodiment of the present application;
[0042] Figure 2 is another schematic flowchart of a feature production method shown in an embodiment of the present application;
[0043] Figure 3 is a schematic structural diagram of a feature production system shown in an embodiment of the present application;
[0044] Figure 4 is another schematic structural diagram of a feature production system shown in an embodiment of the present application;
[0045] Figure 5 is a schematic structural diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0047] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0048] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, the meaning of "a plurality" is two or more, unless otherwise specifically defined.
[0049] The current feature extraction and production follow a roughly unified process: first, extract features, then store the features in a cache. When performing online prediction, the features are extracted from the cache and then online inference prediction is carried out. When printing formatted feature samples offline, the samples are input into the model for training and evaluation, and finally model iteration is performed.
[0050] However, in the context of multi-person collaboration, the current feature extraction and production methods are prone to problems such as abnormal feature extraction, inconsistent feature printing, time travel, feature invalidation, etc. In practical applications, on the one hand, relevant technical personnel need to spend a lot of time troubleshooting problems, resulting in low efficiency of feature production and affecting the accuracy of the recommendation algorithm. On the other hand, due to the lack of feature management, it directly leads to a low reuse rate of features and increases the development difficulty of the recommendation model.
[0051] To address the above problems, an embodiment of this application provides a feature production method. By numbering according to the attribute information and configuration information corresponding to the features, the abstract features are concretized into feature numbers. The representation form of "feature numbers" not only facilitates the management of features and improves the reuse rate of features, but also can avoid or quickly troubleshoot abnormal features during the feature production process, thereby improving the efficiency of feature production.
[0052] The technical solutions of the embodiments of this application are described in detail below with reference to the accompanying drawings.
[0053] See Figure 1 , Figure 1 which is a schematic flowchart of a feature production method shown in an embodiment of this application, including at least the following steps:
[0054] Step 101, create a new feature and obtain the attribute information and configuration information corresponding to the feature;
[0055] In the embodiment of this application, applied to a feature production system, the feature production system can create a new feature and obtain the attribute information and configuration information corresponding to the feature.
[0056] Feature production refers to the process of processing and transforming raw data to extract feature data that can be used to train machine learning models, which can convert raw data into feature representations suitable for model training.
[0057] The process of processing and transforming raw data may include: first, collecting raw data, such as user behavior data, device information, text data, image data, etc., and then performing operations such as cleaning and preprocessing on the raw data for subsequent feature extraction and transformation. In the feature extraction stage, information representing the data features can be extracted from the preprocessed data.
[0058] Among them, the attribute information can be the features or properties possessed by the feature itself, including but not limited to type, format, range, precision, etc.
[0059] The configuration information can be various parameters, settings, etc. related to the system or module for management, mainly used to reflect the functions, relationships, or behaviors of the system or module, etc. In this application, the configuration information can include binding relationship configuration and feature warning configuration, and the configuration information is entered into a preset database.
[0060] Step 102: Number the features according to the attribute information and configuration information to obtain a feature number for the feature;
[0061] In the embodiments of this application, the features can be numbered according to the attribute information and configuration information to obtain a feature number for the feature.
[0062] The feature number can be a unique identifier or number generated automatically, used to identify and extract the corresponding attribute information and configuration information of the feature. All subsequent processes (feature production, printing, use, and monitoring) adopt the form of the feature number. Using the feature number can not only achieve transparency of the feature externally, ensure the security of the feature, but also save storage space and facilitate unified printing and monitoring.
[0063] Step 103: Perform feature production according to the feature number, attribute information, and configuration information to obtain a feature dataset.
[0064] In the embodiments of this application, feature production can be performed according to the feature number, attribute information, and configuration information to obtain a feature dataset.
[0065] During the feature production process, it is necessary to select features useful for model training from the extracted features and remove redundant or irrelevant features to improve the generalization ability and effect of the model.
[0066] It should be noted that the features newly created in step 101 belong to the original data, and the feature dataset in step 103 is a data set obtained after processing and processing the original data, such as after offline feature production and real-time feature production, such as historical data, statistical data, aggregated data, real-time monitoring data, real-time calculation results, user features, context features, etc. There is a corresponding relationship between the feature dataset and the feature number, and it is stored in a preset storage system.
[0067] The solution provided by this application first creates a feature, obtains the attribute information and configuration information corresponding to the feature, then numbers the feature according to the attribute information and configuration information to obtain a feature number for the feature, and performs feature production according to the feature number, attribute information, and configuration information to obtain a feature dataset. By numbering according to the attribute information and configuration information corresponding to the feature, this application realizes the concretization of abstract features into feature numbers. The representation form of "feature number" is not only conducive to the management of features and the improvement of the reuse rate of features, but also in the process of feature production, it can avoid abnormal problems or quickly troubleshoot abnormal features, thereby improving the efficiency of feature production.
[0068] Figure 2 It is another process schematic diagram of a feature production method shown in an embodiment of this application. Figure 2 Relative Figure 1 It describes the technical solution of the embodiment of this application in more detail. The method may include the following steps:
[0069] Step 201, create a feature, and obtain the attribute information and configuration information corresponding to the feature;
[0070] In the embodiment of this application, all fields are abstracted as features, which are suitable for a variety of different platforms, such as recommendation systems, metric systems, and various data dashboards, and can also be used for business analysis, data warehouse mining presentation, and aggregation of various business data. For example, quickly produce an ES (Elasticsearch, a distributed search engine) index for a product aggregation.
[0071] Optionally, the attribute information at least includes the feature value range and feature caliber of the feature, and the configuration information at least includes the binding relationship configuration and feature warning configuration.
[0072] Among them, the feature value range refers to the set of all possible values of the feature in the dataset. For different features, the feature value range is also different. For example, for categorical features, the feature value range can be a set of discrete categories or labels, such as the gender of the visiting user (male, female), the product categories browsed or searched by the visiting user (electronic products, clothing, food), etc. For numerical features, the feature value range can be an interval or a continuous set of numerical values, such as age, price, temperature, etc.
[0073] The feature caliber refers to the rules for describing the data sources and calculation methods required for offline feature production and real-time feature production. For example, the feature caliber is "determine whether to read the db or hive table according to the source_type".
[0074] The binding relationship configuration refers to setting and managing the association relationships between different components, modules, or services. Such association relationships can include data streams, call relationships, dependency relationships, etc. Through configuration, the parameters, options, and behaviors of the association relationships can be specified. In this application, the binding relationship configuration at least includes the binding relationship configuration between features and resource positions and the binding relationship configuration between resource positions and the data logging system.
[0075] By configuring the binding relationship between features and resource positions, when creating a new feature, it is possible to determine which pages of the client the feature is used for, and then select the corresponding page numbers to complete the binding. The binding relationship can be represented by a piece of data in the feature table. And by configuring the binding relationship between resource positions and the data logging system, when a user operates, it is possible to correctly trigger the data logging system to collect user behavior data and record the user's click, browsing, submission, and other operation behaviors on the application or website, which is beneficial to monitoring, analyzing, and statistically analyzing user behaviors.
[0076] The feature warning configuration can be the configuration of the feature alarm level and threshold. For example, the first-level alarm level corresponds to the first threshold, and the second-level alarm level corresponds to the second threshold. This application does not limit the actual setting content of the feature alarm level and threshold.
[0077] As another example, the embodiments of this application provide other attribute information and configuration information of features. The specific content is shown in Feature Table 1, including the feature number, the target metadata information stored after production, etc. Features can be configured in a one-to-many relationship with resource positions and storage objects, enabling one-time production, multiple reuse, and storage.
[0078] As shown in Feature Table 1, in addition to the above-mentioned feature value range, feature caliber, binding relationship configuration, and feature alarm configuration, it at least further includes:
[0079] a) The unique number of the feature, which is subsequently used for the entire life cycle of feature production, use, and monitoring;
[0080] b) The data source type. The currently set data sources include the data warehouse and the business production domain;
[0081] c) The feature type, which identifies what type of feature it is. A real-time feature is first of all an offline feature. For example, 1 identifies an offline feature, and 2 identifies a real-time feature;
[0082] d) Identify the specific type of the material, which involves reading the main material table, such as the material number, 1 for commodity, 2 for diary, etc.;
[0083] e) The key name (key name) of the feature input into Redis (an in-memory data structure storage system). There can be multiple keys, which are expanded into separate tables;
[0084] f) Monitoring type, which can be a phone alarm or a separate message notification according to the level, 1 indicates an IM alarm, and 2 indicates a phone alarm;
[0085] g) The ES to which the features are output can be multiple and expanded into separate tables;
[0086] h) The Doris (distributed column storage and query) to which the features are output can be multiple and extended into separate tables;
[0087] i) MQ (Message Queue) addresses for real-time feature reading, which can be multiple and expanded into a separate table;
[0088] j) A detailed description of the characteristics.
[0089]
[0090] Table 1
[0091] Step 202, numbering the features according to the attribute information and the configuration information to obtain a feature number for the feature;
[0092] In an embodiment of the present application, when automatically numbering features, the features are associated with feature value ranges, feature calibers, binding relationship configurations, and feature warning configurations to generate feature numbers for the features.
[0093] As an example, the feature number is used to classify and identify the value range and caliber of the feature, which facilitates the subsequent management, monitoring and use of the feature. For example, the feature number of a feature can indicate the value range and caliber type to which it belongs, so that when using the feature, the attributes and purpose of the feature can be quickly determined.
[0094] Similarly, when binding features and resource bits, the feature number is associated with the corresponding resource bit so that the relationship between the feature and the resource bit can be quickly located through the feature number in the subsequent process. The feature number can also be used to monitor and warn each feature. When an abnormality occurs, alarm information is generated in time and fed back to relevant technical personnel to improve the efficiency and maintainability of the system. At the same time, the security of the feature can also be ensured, thereby effectively improving work efficiency, achieving decoupling of engineering and algorithms, and permanently precipitating the features as platform assets. The features can be reused by multiple businesses and effectively avoid business losses.
[0095] Step 203: Perform feature production based on the feature number, attribute information, and configuration information to obtain a feature dataset, and store it in a preset storage system.
[0096] In the embodiment of the present application, after completing the feature numbering, the feature number stored in the preset database is called, and according to the association relationship between the feature number and the attribute information and configuration information, the attribute information and configuration information are accurately and quickly extracted. Then, the attribute information and configuration information are used for feature production to obtain a feature dataset.
[0097] The preset storage system may include Redis (a data structure storage system in memory), Elasticsearch (ES, a distributed search engine), and Doris (a distributed columnar storage and query), etc. The purpose of storing the produced features in these storage systems is to facilitate fast data access and query, and at the same time support different data processing and analysis requirements, such as real-time query, offline analysis, data visualization, etc.
[0098] As an example, feature production includes offline feature production, and the feature dataset includes an offline feature dataset. In an offline environment, the attribute information and configuration information are read through the feature number, and at the same time, a preset timing scheduler is used to obtain historical data. The historical data is used for offline feature production with the attribute information and configuration information to generate an offline feature dataset, which is then stored in the preset storage system.
[0099] Among them, offline feature production can be the process of feature extraction and processing of historical data in an offline environment, usually including operations such as analyzing, extracting, processing, and transforming historical data to generate an offline feature dataset. The offline feature dataset can be used for offline model training and evaluation, and is usually used for batch processing tasks, such as feature processing in an offline recommendation system.
[0100] In the time dimension, historical data belongs to data that has already occurred and been recorded. Historical data can be features obtained from storage systems such as databases, data warehouses, or log files. For example, user features such as user behavior data and transaction data, item features such as recommended article or recommended video, tags of recommended content, and context features such as device information and time.
[0101] For example, in an offline environment, the attribute information and configuration information are read through the feature number, and then these attribute information and configuration information are used to determine the data source or screening conditions, etc. The preset timing scheduler further collects the historical data that needs to be used for offline feature production based on the attribute information and configuration information, and then performs offline feature production on the historical data to generate an offline feature dataset, which is stored in the preset storage system. During the offline feature production process, the feature number is used to identify and manage features, while the attribute information and configuration information are used to guide feature production.
[0102] As another example, feature production includes real-time feature production, and the feature dataset includes a real-time feature dataset. In a real-time environment, the corresponding attribute information and configuration information can be called by the feature number, and the real-time data is read from the subscribed message list using the attribute information and configuration information, and real-time feature production is performed on the real-time data to generate a real-time feature dataset.
[0103] Real-time feature production can be the real-time configuration and processing of real-time data features in an offline environment, including the process of real-time feature extraction, processing, and transformation of the real-time generated data to generate a real-time feature dataset. The real-time feature dataset can be used for tasks such as fast response and real-time inference, such as online advertising systems and real-time recommendation systems.
[0104] The message list can be a Kafka (distributed stream processing platform and message queue system) message middleware. The MQ address (Message Queue) of the message list can be determined according to the attribute information and configuration information corresponding to the feature number, and the message queue is connected through the MQ address, thereby realizing message subscription and obtaining real-time data from the subscribed messages.
[0105] In the time dimension, real-time data belongs to the data that is currently occurring and being recorded. Real-time data can be features obtained in real time from the message list. For example, user features such as user behavior data and transaction data, item features such as recommended article or recommended video and tags of recommended content, and context features such as device information and time.
[0106] For example, real-time feature production first reads the feature configuration (configuration information configured through the management background, including the MQ address, message format, and captured feature caliber. MQ is middleware for asynchronous communication. By specifying the message queue address to be subscribed, the message is sent to the queue to achieve decoupling between different systems. Then, the message format is obtained from the message queue, including information such as the fields, data types, and encodings of the messages, and data is captured simultaneously, including parameters such as capture frequency, capture conditions, and capture volume, to achieve real-time feature update and ensure the synchronization of feature data and real-time data.
[0107] As another example, the feature production system has pre-set a data tracking tool to collect data tracking data. On the one hand, when conducting real-time feature production, the data tracking data can be used as the data source for extracting real-time data. On the other hand, during model training, the data tracking data can be used as training data to participate in model training. Among them, the data tracking tool can be a tool that inserts data tracking code into an application or website to collect user behavior data, providing the function of customizing the collection of specific data and specific events. The data tracking data collected by the data tracking tool includes various behaviors and operations of users in the application or website. For example, the time of browsing a page, submitting a form, etc.
[0108] As another example, data maintenance is carried out through offline feature production and real-time feature production: read the offline feature configuration every day and conduct offline feature production. The produced features will be stored in Redis, Doris, and ES in the form of numbers. Read the real-time feature configuration in real time, read the messages of feature changes from the specified message list for real-time feature production, and write the results into Redis, Doris, and ES. At the same time, the real-time features will also parse real-time data tracking to extract item features (specific data or events) and user features and store them in Redis and Doris. The data maintenance function is realized by combining offline feature production and real-time feature production.
[0109] Step 204, call the target feature data corresponding to the batch processing request through the feature number, and use a preset feature printing tool to print the target feature data;
[0110] In the embodiment of the present application, first receive a batch processing request, obtain the feature number corresponding to the batch processing request, then call the target feature data corresponding to the batch processing request from a preset storage system through the feature number, and finally use a feature printing tool to print the target feature data.
[0111] Among them, the batch processing request can be a Pipeline request. Pipeline is a batch processing technology provided by the client, which can be used to process multiple Redis commands at one time to improve the performance of interaction.
[0112] The target feature data is the feature data corresponding to the batch processing request, including at least user features, item features, and context features.
[0113] The feature printing tool is used to print all the features called after a single pipeline request. For example, when a user accesses the product list on the home page, a pipeline call will be generated at this time. A single call will return many items, and the features of each item and the features of the user for this visit, etc., will be printed out in the form of a log exactly as they occurred at that time, such as the gender of the user who initiated the request and the number of orders placed for the product in the last 30 days, etc.
[0114] As an example, item features used to describe items can be extracted from the recommendation model, such as product category, release time, author, etc., and these features can be used to assist the recommendation model in performing personalized recommendations. Data recording various behaviors and operations of users in the application or website can also be extracted from the recommendation model, such as user operations like clicks, views, purchases, searches, etc. By analyzing user operations, user preferences and behaviors can be analyzed, thereby enabling advertising placement and user profiling.
[0115] As another example, specific features or data can be extracted from the offline feature dataset or the real-time feature dataset through a pipeline request. ES features are used for recall in the rough recall stage, and real-time interception of the rough recall is performed through Redis features. After entering the fine ranking service, features are used again for Ctr (ClickThroughRate) prediction (predicting the click situation of each advertisement, predicting whether the user will click or not), and finally the recommended data is intercepted and returned. At the same time, all the recalled features, including user features, context features, and item features, will be printed through the feature printing SDK (Software Development Kit).
[0116] Step 205: Integrate the buried point data with the target feature data to generate training samples.
[0117] In the embodiment of the present application, the feature printing tool is used to automatically obtain the target feature data, and the buried point data in the preset storage system is integrated with the target feature data to generate training samples, and these training samples can be used to train various models.
[0118] As an example, the feature printing SDK is used to automatically collect features, and then training samples are automatically generated after connecting with the buried point data. Relevant technical personnel can use the training samples to train corresponding models, such as recommendation models, and deploy or iterate them to the online environment automatically through the model deployment service. Throughout the process, only the training of the model itself needs to be concerned, without having to pay attention to the engineering process, thereby improving the production efficiency.
[0119] The solution provided by this application can materialize abstract features into feature numbers. The representation form of "feature numbers" not only facilitates the management of features and improves the reuse rate of features, but also can avoid abnormal problems or quickly troubleshoot abnormal features during the feature production process, thereby improving the efficiency of feature production.
[0120] Corresponding to the foregoing method embodiments for implementing application functions, this application also provides a feature production system, an electronic device, and corresponding embodiments.
[0121] See Figure 3 , Figure 3 FIG. is a schematic structural diagram of a feature production system shown in an embodiment of this application. The feature production system in the embodiment of this application at least includes a feature number module 301 and a feature production module 302; the feature number module 301 is used to create a feature, obtain attribute information and configuration information corresponding to the feature, number the feature according to the attribute information and configuration information, and obtain a feature number for the feature; the feature production module 302 is used to perform feature production according to the feature number, attribute information, and configuration information to obtain a feature dataset.
[0122] In some embodiments, the process of the feature number module 301 numbering features is as follows: First, create a feature. The types of features created include numerical features, categorical features, text features, image features, etc., and the newly created features can be converted and combined according to user needs to generate new features. Then, obtain the attribute information and configuration information corresponding to each feature, number according to the attribute information and configuration information to obtain a feature number for each feature, and enter it into the feature production system. The feature production system uniformly manages each feature according to the feature number, and other modules can flexibly retrieve and use it. And when an abnormal problem occurs, the abnormal feature data can be quickly determined through the feature number and a warning can be issued, avoiding the need to spend a lot of time troubleshooting and maintaining, and improving work efficiency.
[0123] In some embodiments, during the feature production process, it is necessary to select features useful for model training from the extracted features and remove redundant or irrelevant features to improve the generalization ability and effect of the model. In this application, the feature production module 302 can call the feature numbers stored in the preset database by the feature number module 301, accurately and quickly extract the attribute information and configuration information based on the association between the feature number and the attribute information and configuration information, and then use the attribute information and configuration information to perform feature production to obtain a feature dataset.
[0124] As an example, feature production includes offline feature production, and the feature dataset includes an offline feature dataset. In an offline environment, the feature production module 302 can read attribute information and configuration information by feature number, and at the same time use a preset timing scheduler to obtain historical data. The historical data is used for offline feature production with the attribute information and configuration information to generate an offline feature dataset, which is then stored in a preset storage system.
[0125] As another example, the feature production system also presets a data tracking tool. Feature production includes real-time feature production, and the feature dataset includes a real-time feature dataset. In a real-time environment, the feature production module 302 can read the corresponding attribute information and configuration information by feature number, use the attribute information and configuration information to read real-time data from the message list, perform real-time feature production on the real-time data to generate a real-time feature dataset, use the data tracking tool to collect data tracking data, and store the real-time feature dataset and the data tracking data in a preset storage system.
[0126] In addition, the feature production module 302 can perform offline and online feature data maintenance, and write the features produced once into Redis, Elasticsearch, and Doris for use by multiple consumers.
[0127] For example, the feature production module 302 can read the offline feature configuration every day and perform offline feature production. The produced features will be stored in Redis, Doris, and ES in the form of numbers. It reads the real-time feature configuration in real time, reads the messages of feature changes from the specified message list for real-time feature production, and writes the results into Redis, Doris, and ES. At the same time, the real-time features will also parse real-time data tracking to extract item features and user features and store them in Redis and Doris for use by multiple consumers.
[0128] This application can effectively improve work efficiency, decouple engineering and algorithms, permanently precipitate features as the assets of the platform, and can be reused by multiple services. Through alarm monitoring, problems existing in the features can be discovered in time, and the prediction effect of the model can be effectively discovered to avoid business losses.
[0129] Refer to Figure 4 , Figure 4 , which is another structural schematic diagram of the feature production system shown in the embodiments of this application. Through this system, the entire process of online feature configuration, production, use, and monitoring can be realized. Among them, the feature management and monitoring part corresponds to the feature number module of this application, the feature production part corresponds to the feature production module of this application, and both the recall pipeline part and the algorithm training part belong to the consumers of the feature number module and the feature production module.
[0130] The feature management and monitoring part mainly includes new features, feature numbering, real-time feature configuration, ES index configuration, resource location configuration, page embedding system, feature monitoring, and IM (Instant Messaging) / telephone alarms, and stores these data in the database.
[0131] The feature production part mainly includes offline feature production, feature dashboard, and point analysis through feature scheduling and Hive / DB, as well as real-time feature production through Kafka messages, and the data of offline feature production and real-time feature production (including item features and user data) are stored in Redis, Elasticsearch, and Doris.
[0132] The recall pipeline mainly includes recommendation data, search and promotion engine, fine ranking feature service, and material rough recall service, which reflects the overall sorting list process. Feature data is used in both rough recall and fine ranking.
[0133] The algorithm training part mainly includes feature printing SDK, training sample integration, model training services, and deployment of new models. It is responsible for the entire process of automated production of feature samples and redeployment of features to the model.
[0134] In some implementations, the feature production system further includes a feature recall module 303 .
[0135] The feature recall module 303 is one of the users of the feature production module, and is mainly used to receive batch processing requests, obtain the feature number corresponding to the batch processing request, call the target feature data corresponding to the batch processing request from the preset storage system through the feature number, and use the preset feature printing tool to print the target feature data.
[0136] The batch request may be a Pipeline request. The batch processing technology provided by Pipeline to the client can be used to process multiple Redis commands at a time to improve the performance of the interaction.
[0137] The feature printing tool is used to print all the features called after a pipeline request. For example, when a user visits the product list on the homepage, a pipeline call will be generated, and a call will return many items. The features of each item and the features of the user who visited this time are printed out in the form of a log according to what happened at the time, such as the gender of the user who initiated the request and the number of orders for the product in the last 30 days.
[0138] As an example, the feature recall module 303 can extract item features used to describe items from the recommendation model, such as product category, release time, author, etc., and use these features as specific event data to assist the recommendation model in making personalized recommendations. The feature recall module 303 can also extract data recording various behaviors and operations of users in the application or website from the recommendation model, such as user operations like clicks, views, purchases, searches, etc., and analyze the preferences and behaviors of users through user operations, so as to carry out advertisement placement and user profiling.
[0139] As another example, the feature recall module 303 can extract specific features or data from the offline feature dataset or the real-time feature dataset, use ES features for recall in the rough recall stage, and perform real-time interception of the rough recall through redis features. After entering the fine ranking service, features are used again for ctr (ClickThroughRate) prediction (predict the click situation of each advertisement, predict whether the user clicks or not), and finally the recommended data is intercepted and returned. At the same time, the feature printing SDK (Software Development Kit) will be used to print all the recalled features, including user features, context features, and item features.
[0140] In some embodiments, the feature production system further includes a training sample generation module 304.
[0141] The training sample generation module 304 also belongs to one of the users of the feature production module, and is mainly used to automatically obtain target feature data using the feature printing tool, integrate the buried point data in the preset storage system with the target feature data, and generate training samples, which can be used to train various models.
[0142] As an example, the training sample generation module 304 automatically collects features through the feature printing SDK, and automatically generates training samples after connecting with the buried point data. Relevant technical personnel can use the training samples to train corresponding models, such as the recommendation model, and automatically deploy or iterate to the online environment through the model deployment service. Throughout the process, only the training of the model itself needs to be concerned, without having to pay attention to the engineering process, thus improving the production efficiency.
[0143] It should be noted that the feature recall module 303 and the training sample generation module 304 belong to one of the users of the feature numbering module 301 and the feature production module 302. In actual use, the users include but are not limited to the feature recall module 303 and the training sample generation module 304.
[0144] The solution provided by this application, the feature production system at least includes a feature numbering module and a feature production module. The feature numbering module is used to create features, obtain attribute information and configuration information corresponding to the features, number the features according to the attribute information and configuration information, and obtain a feature number for the features. The feature production module is used to produce features according to the feature number, attribute information and configuration information, and obtain a feature data set. By numbering according to the attribute information and configuration information corresponding to the features, this application realizes the concretization of abstract features into feature numbers. The representation form of "feature number" not only facilitates the management of features and improves the reuse rate of features, but also can avoid abnormal problems or quickly troubleshoot abnormal features during the feature production process, thereby improving the efficiency of feature production.
[0145] Regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the system, and will not be elaborated here.
[0146] Figure 5 It is a schematic structural diagram of an electronic device shown in an embodiment of this application.
[0147] See Figure 5 , the electronic device 500 includes a memory 510 and a processor 520.
[0148] The processor 520 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0149] The memory 510 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 520 or other modules of the computer. The permanent storage device can be a readable and writable storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device employs a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 510 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 510 can include removable storage devices that are readable and / or writable, such as compact discs (CDs), read-only digital versatile discs (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray discs, high-density discs, flash memory cards (such as SD cards, mini SD cards, Micro-SD cards, etc.), magnetic floppy disks, etc. Computer-readable storage media do not include carrier waves and instantaneous electronic signals transmitted wirelessly or wired.
[0150] Executable code is stored on the memory 510, and when the executable code is processed by the processor 520, it can cause the processor 520 to execute some or all of the methods described above.
[0151] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the above steps of the method according to the present application.
[0152] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When the executable code (or the computer program or computer instruction code) is executed by a processor of an electronic device (or a server, etc.), it causes the processor to execute some or all of the steps of the method according to the present application.
[0153] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for producing features, characterized in that, The method includes: Creating a new feature and obtaining the attribute information and configuration information corresponding to the feature; Numbering the feature according to the attribute information and the configuration information to obtain a feature number for the feature; Performing feature production according to the feature number, the attribute information, and the configuration information to obtain a feature data set.
2. The method according to claim 1, wherein The attribute information at least includes the feature value range and feature caliber of the feature, and the configuration information includes binding relationship configuration and feature warning configuration. The step of numbering the feature according to the attribute information and the configuration information to obtain a feature number for the feature includes: When automatically numbering the feature, associating the feature with the feature value range, the feature caliber, the binding relationship configuration, and the feature warning configuration to generate a feature number for the feature.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Storing the feature number as the representation form of the feature in a preset database, and each feature number corresponds to one feature; Wherein, the feature number is used to identify and extract the configuration information and attribute information corresponding to the feature.
4. The method according to claim 1, wherein The feature production includes offline feature production, and the feature data set includes an offline feature data set. The step of performing feature production according to the feature number, the attribute information, and the configuration information to obtain a feature data set includes: Obtaining historical data according to the attribute information and configuration information corresponding to the feature number; Performing the offline feature production on the historical data by using the attribute information and the configuration information to generate the offline feature data set.
5. The method according to claim 1, wherein The feature production includes real-time feature production, and the feature data set includes a real-time feature data set. The step of performing feature production according to the feature number, the attribute information, and the configuration information to obtain a feature data set includes: Obtaining real-time data according to the attribute information and configuration information corresponding to the feature number; Performing the real-time feature production on the real-time data by using the attribute information and the configuration information to generate the real-time feature data set.
6. The method according to claim 5, wherein The method further includes: Receiving a batch processing request and obtaining the feature number corresponding to the batch processing request; Invoking the target feature data corresponding to the batch processing request through the feature number and printing the target feature data by using a preset feature printing tool.
7. The method according to claim 6, wherein The method further includes: Obtaining buried point data and the target feature data; Integrating the buried point data with the target feature data to generate a training sample.
8. A feature production system, characterized in that: The feature production system at least includes a feature number module and a feature production module; The feature number module is used to create a new feature, obtain the attribute information and configuration information corresponding to the feature, and number the feature according to the attribute information and the configuration information to obtain a feature number for the feature; The feature production module is used to perform feature production according to the feature number, the attribute information, and the configuration information to obtain a feature data set.
9. The system according to claim 8, characterized in that: The feature production system further includes a feature recall module; The feature recall module is configured to receive a batch request, obtain a feature number corresponding to the batch request; call target feature data corresponding to the batch request through the feature number, and print the target feature data by using a preset feature printing tool.
10. The system according to claim 9, wherein: The feature production system further includes a training sample generation module; The training sample generation module is configured to obtain buried point data and the target feature data; integrate the buried point data with the target feature data to generate a training sample.
11. An electronic device, characterized in that, Comprising: A processor; And A memory storing executable code thereon, which when executed by the processor causes the processor to execute the method according to any one of claims 1-7.
12. A computer-readable storage medium storing executable code thereon, which when executed by a processor of an electronic device causes the processor to execute the method according to any one of claims 1-7.