Priori knowledge migration-based recommendation fine arrangement model construction method

By adopting prior knowledge transfer and Huber loss function processing in the recommended fine layout model, the problems of data sparseness and target imbalance are solved, the training accuracy and stability of the model are improved, and better recommendation results are achieved.

CN119990368APending Publication Date: 2025-05-13CHEZHI HULIAN BEIJING SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510081262.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The traditional multi-objective fine-scheduling model faces sparse data and imbalance in the training process, resulting in poor recommendation results.

Method used

The prior knowledge migration method is adopted to use the data-rich link front-end link output as the input of the link back-end link, and the Huber loss function is used to process sparse targets, prevent the backward transmission of the front-end link output and reduce the impact of outliers.

Benefits of technology

It improves the training accuracy and stability of the model in the data sparse link, improves the overall recommendation effect, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990368A_ABST
    Figure CN119990368A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing and analysis, machine learning and recommendation systems, and discloses a priori knowledge migration-based recommendation fine arrangement model construction method, which specifically comprises the following steps of: 1, training sample production; step 2, feature engineering; step 3, model training; and 4, reducing the influence on the front-end link. Through model training optimization, rich front-end data is utilized to improve the training accuracy of a rear-end sparse link, the model preferentially stabilizes a front-end link target at the initial stage, then the rear-end link loss is gradually increased, the overall performance is improved, in order to reduce the interference of the front end on rear-end training, the reverse transmission of the front-end output is blocked before the front-end output is taken as the rear-end input, and meanwhile, the rear-end sparse link training accuracy is improved. The sparse target is processed by adopting functions such as huberl oss, and the influence of abnormal values is reduced by setting hyper-parameters, so that the model can quickly adapt to non-mainstream distribution data, the interference of abnormal data is reduced, and the prediction effect is further enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing and analysis, machine learning and recommendation system, and specifically is a method for constructing a recommendation ranking model based on prior knowledge transfer. Background Art

[0002] With the development of the Internet, recommendation systems have been widely used in e-commerce, video, music, news and other industries. Among them, the fine ranking model, as an important means to improve the accuracy of recommendation, has achieved a high recommendation effect by learning a large amount of data. However, the traditional multi-objective fine ranking model often faces the problem of data sparsity during the training process, which limits the accuracy of the target. The existing model training process includes data collection and processing, determination of multiple targets, feature engineering, model construction, model training, model evaluation and model optimization, but there are shortcomings such as data bias and target imbalance. The data bias problem stems from the imbalance of data distribution between multiple targets. For example, the browsing and clicking behavior of users of Autohome is much greater than the likes and comments, which may cause the model to overfit some target data and ignore other targets. The target imbalance problem stems from the competitive relationship between multiple targets. The model may ignore other targets due to excessive focus on the targets of some rare samples, thus affecting the overall recommendation effect. To address the above issues, this solution proposes a prior knowledge transfer method, which aims to improve the accuracy of data-sparse targets in the recommendation model by using the output of the front-end link with sufficient data as the input of the back-end link with sparse data. Summary of the invention

[0003] The purpose of the present invention is to provide a method for constructing a recommendation ranking model based on prior knowledge transfer to solve the problems raised in the above background technology.

[0004] In order to achieve the above-mentioned object, the present invention provides the following technical solution: a method for constructing a recommendation ranking model based on prior knowledge transfer, wherein the specific steps of the method for constructing the recommendation ranking model include:

[0005] Step 1: Training sample production: Collect user click data, behavior data, and material data, perform data cleaning, deduplication, and preprocessing to obtain a data set for building a recommendation model;

[0006] Step 2, feature engineering: according to the target characteristics and processing requirements, feature extraction, feature selection and feature construction are performed to obtain effective features related to multiple targets;

[0007] Step 3: Model training: Train the multi-objective model and use the output of the link front-end environment as the input of the link back-end environment. The initial model l oss tends to the front-end link target with rich training data. As the number of training steps increases, the front-end link target gradually stabilizes, and the back-end data sparse environment l oss increases.

[0008] Step 4: Reduce the impact on the front-end link: Before using the front-end link output as the back-end link input, prevent the reverse transmission of the front-end link output, use the Huber Loss function to process sparse targets, reduce the impact of outliers, and make the training more robust.

[0009] Preferably, the data cleaning in step one includes data collection and understanding, statistical description and visualization, missing value processing, outlier processing, data type conversion, data standardization / normalization, internal consistency check, external consistency verification, documentation, version control, and regular review.

[0010] Preferably, the data deduplication in step 1 refers to: using a unique identifier or combining multiple fields to determine whether the data is duplicated; retaining one record and deleting the remaining duplicates, or merging the information of duplicate records according to business logic.

[0011] Preferably, the data preprocessing in step 1 refers to further sorting of data after data cleaning and deduplication to ensure that the data format is unified and meets the requirements of subsequent analysis or model training.

[0012] Preferably, the feature extraction in step 2 includes:

[0013] Statistical methods: Calculate the chi-square value and information gain of each word to select features that have a significant impact on the target variable;

[0014] Algorithmic methods: including image feature extraction, speech feature extraction and text feature extraction.

[0015] Preferably, the feature selection in step 2 includes: generating subsets, subset evaluation, stopping criteria and result verification; the feature selection methods include filtering methods, encapsulation methods and embedding methods.

[0016] Preferably, the feature construction in step 2 includes: based on business logic, based on data transformation, feature intersection and missing value processing; the feature construction methods include sorting features, discrete features, counting features, missing value features and intersection features.

[0017] The beneficial effects of the present invention are as follows:

[0018] The present invention utilizes the data-rich front-end link as input through the model training step, effectively improving the training accuracy of the back-end sparse link. In the early stage of training, the model l oss tends to the front-end link target, and quickly stabilizes these targets. As the number of training steps increases, the front-end link target is stabilized, and the l oss of the back-end link gradually increases, thereby achieving the improvement of the overall model performance. In addition, by reducing the influence of the front-end link on the back-end link training, the reverse transmission of the front-end link output is prevented before the front-end link output is used as the back-end link input, and at the same time, functions such as huber_l oss are used to process sparse targets, and hyperparameters are set to reduce the influence of outliers, so that the model can be quickly updated when facing data with non-mainstream distribution in abnormal time periods, reducing the influence of abnormal situations on data, and further improving the model prediction effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A simplified diagram of the steps of the refined model construction method recommended by the present invention;

[0020] Figure 2 This is a schematic diagram of the model training of the present invention. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0022] like Figure 1 to Figure 2 As shown, the embodiment of the present invention provides a method for constructing a recommendation ranking model based on prior knowledge transfer, and the specific steps of the method for constructing the recommendation ranking model include:

[0023] Step 1: Training sample production: Collect user click data, behavior data, and material data, perform data cleaning, deduplication, and preprocessing to obtain a data set for building a recommendation model;

[0024] Collect user click data, behavior data, material data, etc., perform data cleaning, deduplication, preprocessing and other operations to obtain the data set used to build the recommendation model. This step aims to ensure the accuracy and representativeness of the training samples and provide a reliable data foundation for subsequent model training.

[0025] Step 2, feature engineering: according to the target characteristics and processing requirements, feature extraction, feature selection and feature construction are performed to obtain effective features related to multiple targets;

[0026] According to the characteristics and processing requirements of each target, feature engineering is performed, including feature extraction, feature selection, feature construction, etc., to obtain effective features related to multiple targets. Feature engineering is a key link in the construction of recommendation systems. Through reasonable feature processing, the prediction ability of the model can be significantly improved.

[0027] Step 3: Model training: Train the multi-objective model and use the output of the link front-end environment as the input of the link back-end environment. The initial model l oss tends to the front-end link target with rich training data. As the number of training steps increases, the front-end link target gradually stabilizes, and the back-end data sparse environment l oss increases.

[0028] The multi-objective model is trained, and the output of the link front-end environment is used as the input of the link back-end environment. In the early stage of model training, the model l oss is biased towards the front-end link targets with rich training data, aiming to quickly stabilize these targets. As the number of training steps increases, the front-end link targets gradually stabilize, and the back-end data sparse environment l oss increases. This step gradually optimizes the performance of the model on each target through staged training, while using the prior knowledge of the front-end link to assist the training of the sparse back-end link.

[0029] Step 4: Reduce the impact on the front-end link: Before using the front-end link output as the back-end link input, prevent the reverse transmission of the front-end link output, use the Huber Loss function to process sparse targets, reduce the impact of outliers, and make the training more robust.

[0030] Before the output of the front-end link is used as the input of the back-end link, the reverse transmission of the output of the front-end link is prevented to prevent the sparse link of the back-end from biasing the trained front-end hidden layer. Sparse targets are prone to instability. For example, when encountering some incentive events or hot events, sparse targets are prone to learning bias. For this reason, the present invention adopts huber-loss functions and sets hyperparameters. When the error is large, the influence of outliers can be reduced, making the training more robust. This step effectively avoids the negative impact between the front-end and back-end links by finely controlling the training process, and improves the stability and generalization ability of the model.

[0031] Among them, the data cleaning in step one includes data collection and understanding, statistical description and visualization, missing value processing, outlier processing, data type conversion, data standardization / normalization, internal consistency check, external consistency verification, documentation, version control, and regular review.

[0032] Data collection and understanding: Obtain data from various sources (such as databases, APIs, files, etc.) and check the structure, data type, variable meaning, and missing values ​​of the data to gain a preliminary understanding of the overall situation of the data.

[0033] Statistical description and visualization: Use descriptive statistics (such as mean, median, standard deviation, etc.) to understand the distribution of data, and use charts (such as histograms, box plots, scatter plots, etc.) to visually display data characteristics and identify possible outliers and trends.

[0034] Missing value handling:

[0035] Delete missing values: Records with a large number of missing values ​​can be deleted directly, especially when the missing ratio is high.

[0036] Fill missing values: Select a filling method based on the specific situation, such as using the mean, median, or mode (applicable to numerical data), or using the most frequently occurring value (applicable to categorical data), or using more complex interpolation methods.

[0037] Mark missing values: Treat missing values ​​as a special category and create new variables to mark the missing cases.

[0038] Outlier handling: Identify and analyze outliers (e.g., through box plots) and decide whether to keep, correct, or delete them. Correction methods may include using boundary value replacement, smoothing, or filling based on model predictions.

[0039] Data type conversion: Ensure that the data type of each field is correct, such as converting a string to a date format, or converting numeric data to categorical data.

[0040] Data standardization / normalization: For numerical data, standardize (such as z-score standardization) or normalize (such as min-max scaling) as needed to make the data on the same scale.

[0041] Internal consistency check: Check whether the logical relationships between different fields in the data set are consistent, such as whether the date range is reasonable, whether the address information matches, etc.

[0042] External consistency verification: If possible, compare with external data sources to verify the accuracy and completeness of the data.

[0043] Documentation: Record each step of the cleaning process, the methods used and the reasons in detail to facilitate subsequent audits and repeat operations.

[0044] Version control: Perform version control on the cleaned data to ensure that it can be traced back to the state before cleaning.

[0045] Regular review: As new data is added, repeat the above steps regularly to ensure that data quality continues to meet requirements.

[0046] Among them, data deduplication in step one refers to: using a unique identifier or combining multiple fields to determine whether the data is duplicated; retaining one record, deleting the remaining duplicates, or merging the information of duplicate records according to business logic.

[0047] Among them, data preprocessing in step one refers to further sorting after data cleaning and deduplication to ensure that the data format is unified and meets the requirements of subsequent analysis or model training.

[0048] It may include operations such as data transformation (such as simple function transformation, normalization, etc.) and data reduction (such as attribute reduction, value reduction, etc.)

[0049] Among them, the feature extraction in step 2 includes:

[0050] Statistical methods: Calculate the chi-square value and information gain of each word to select features that have a significant impact on the target variable;

[0051] Algorithmic methods: including image feature extraction, speech feature extraction and text feature extraction.

[0052] Image feature extraction: Use SI FT, SURF, and ORB algorithms to extract key points, edges, and other features in images;

[0053] Speech feature extraction: extract Mel frequency cepstral coefficients (MFCCs) and zero crossing rate features from audio signals;

[0054] Text feature extraction: Use the bag-of-words model and TF-IDF method to extract vocabulary features from text.

[0055] Among them, the feature selection in step 2 includes: generating subsets, subset evaluation, stopping criteria and result verification; the feature selection methods include filtering methods, encapsulation methods and embedding methods.

[0056] Generate subsets: Generate a series of feature subsets from the original feature set.

[0057] Subset evaluation: Use some evaluation criteria (such as information gain, distance criterion, etc.) to evaluate the feature subset.

[0058] Stopping criteria: Set a stopping condition (such as reaching a preset evaluation value, running time or number of times, etc.) to determine when to stop generating and evaluating feature subsets.

[0059] Result verification: Verify the performance of the selected feature subset to ensure that it can meet the requirements of subsequent analysis or model training.

[0060] Commonly used feature selection methods include filtering methods (such as variance filtering, correlation filtering, etc.), encapsulation methods (such as using specific classifiers for feature selection) and embedding methods (such as feature importance evaluation in decision trees).

[0061] Among them, the feature construction in step 2 includes: based on business logic, based on data transformation, feature intersection and missing value processing; the feature construction methods include sorting features, discrete features, counting features, missing value features and intersection features.

[0062] Based on business logic: construct new features based on domain knowledge, such as calculating the user's average consumption amount, browsing time, etc.

[0063] Based on data transformation: transform the original features to construct new features, such as logarithmic transformation, square transformation, etc.

[0064] Feature intersection: Combine two or more features to construct new features, such as crossing the user's age and gender to obtain new features.

[0065] Missing value processing: treat missing values ​​as a special category or construct new features after filling missing values ​​according to certain rules.

[0066] Commonly used feature construction methods include sorting features, discrete features, counting features, missing value features, cross features, etc.

[0067] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0068] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a recommendation ranking model based on prior knowledge transfer, characterized in that: The specific steps of the method for constructing the recommendation refined ranking model include: Step 1: Training sample production: Collect user click data, behavior data, and material data, perform data cleaning, deduplication, and preprocessing to obtain a data set for building a recommendation model; Step 2, feature engineering: according to the target characteristics and processing requirements, feature extraction, feature selection and feature construction are performed to obtain effective features related to multiple targets; Step 3: Model training: Train the multi-objective model and use the output of the link front-end environment as the input of the link back-end environment. The initial model loss tends to favor the front-end link target with rich training data. As the number of training steps increases, the front-end link target gradually stabilizes, and the loss of the back-end data sparse environment increases. Step 4: Reduce the impact on the front-end link: Before using the front-end link output as the back-end link input, prevent the reverse transmission of the front-end link output, use the Huber loss function to process sparse targets, reduce the impact of outliers, and make the training more robust.

2. The method for constructing a recommendation ranking model based on prior knowledge transfer according to claim 1, characterized in that: The data cleaning in step one includes data collection and understanding, statistical description and visualization, missing value processing, outlier processing, data type conversion, data standardization / normalization, internal consistency check, external consistency verification, documentation, version control, and regular review.

3. The method for constructing a recommendation ranking model based on prior knowledge transfer according to claim 1, characterized in that: The data deduplication in step 1 refers to: using a unique identifier or combining multiple fields to determine whether the data is duplicated; retaining one record and deleting the remaining duplicates, or merging the information of duplicate records according to business logic.

4. The method for constructing a recommendation ranking model based on prior knowledge transfer according to claim 1, characterized in that: The data preprocessing in step 1 refers to further sorting of data after cleaning and deduplication to ensure that the data format is unified and meets the requirements of subsequent analysis or model training.

5. The method for constructing a recommendation ranking model based on prior knowledge transfer according to claim 1, characterized in that: The feature extraction in step 2 includes: Statistical methods: Calculate the chi-square value and information gain of each word to select features that have a significant impact on the target variable; Algorithmic methods: including image feature extraction, speech feature extraction and text feature extraction.

6. The method for constructing a recommendation ranking model based on prior knowledge transfer according to claim 1, characterized in that: The feature selection in step 2 includes: generating subsets, subset evaluation, stopping criteria and result verification; Feature selection methods include filtering methods, encapsulation methods and embedding methods.

7. The method for constructing a recommendation ranking model based on prior knowledge transfer according to claim 1, characterized in that: The feature construction in step 2 includes: based on business logic, based on data transformation, feature intersection and missing value processing; the feature construction methods include sorting features, discrete features, counting features, missing value features and intersection features.