Intelligent article selection scoring system based on fusion of rule engine and machine learning
The intelligent product selection and scoring system that integrates rule engines and machine learning solves the problems of poor adaptability, sensitivity to outliers, and difficulty in coordinating multiple targets in product evaluation, and achieves high-precision product potential evaluation and automated decision-making.
Patent Information
- Application Number
- CN202510903248.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-21
AI Technical Summary
Existing technologies in intelligent product evaluation have problems such as rigid attribute adaptation, lack of lifecycle response, cross-category feature pollution, and difficulty in multi-objective coordination, resulting in low evaluation accuracy and high dependence on manual labor.
An intelligent product selection and scoring system based on the fusion of rule engine and machine learning is adopted, including data preprocessing, feature engineering, hybrid computing and dynamic weighted fusion units. Through the dynamic weight fusion of rule engine and machine learning, the generation and scoring of product feature data sets are realized.
It realizes scientific, automated and high-precision commodity potential evaluation, improves evaluation accuracy and reduces manual intervention, adapts to the characteristic differences and outliers of different e-commerce platforms, and coordinates multi-objective optimization.
Smart Images

Figure CN120822864A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent commodity decision-making technology, and specifically to an intelligent product selection and scoring system based on the integration of a rule engine and machine learning. Background Art
[0002] The quantitative evaluation of cross-platform product attribute value includes attributes such as price range, category characteristics, and life cycle stage.
[0003] Currently, there are three major technical deficiencies in the field of intelligent commodity evaluation:
[0004] First, attribute adaptation is rigid. Traditional solutions, such as the Drools engine, rely on static rules and are unable to respond to the core attribute differences across product categories. For example, when a high-priced apparel item exceeds 5,000 yuan, the long-tail distribution compresses the score dispersion, resulting in missed potential products. Promotion-sensitive attributes for beauty and fast-moving consumer goods are not isolated, leading to a surge in prediction errors during major promotions.
[0005] Second, there is a lack of life cycle response. A single machine learning model ignores the characteristics of the product life cycle stages. For example, the weight of the communication index needs to be strengthened during the new product launch period, and the proportion of discount intensity evaluation needs to be increased during the clearance period. However, the existing technology has not built a dynamic response mechanism.
[0006] Third, cross-category feature contamination. The universal evaluation system leads to interference from features of non-related categories. For example, 3C digital indicators are misused for beauty products, and feature redundancy increases the error rate of cross-category scoring. Summary of the Invention
[0007] In response to the shortcomings of existing technologies, the present invention proposes an intelligent product selection and scoring system based on the integration of rule engines and machine learning. It solves the four core pain points of traditional apparel selection, namely poor platform adaptability, sensitivity to outliers, difficulty in multi-target coordination, and high manual dependence. It realizes scientific, automated, and high-precision product potential evaluation, providing a breakthrough technical framework for the field of e-commerce product selection.
[0008] To achieve the above objectives, the present invention designs an intelligent product selection and scoring system based on the fusion of rule engine and machine learning. The system is characterized by: including a data preprocessing unit, a feature engineering unit, a hybrid computing unit, and a dynamic weighted fusion unit;
[0009] The data preprocessing unit is used to clean the collected raw data of intelligent product selection and process missing data and abnormal data;
[0010] The feature engineering unit is used to generate a product feature dataset from the pre-processed smart product selection data according to the core value dimensions of different smart product selections. The product feature dataset includes a basic feature dataset, a statistical feature dataset, a time series feature dataset, and a composite feature dataset;
[0011] The hybrid computing unit is used to simultaneously execute preset business rule calculations and machine learning predictions; the preset business rule calculations are completed through business rule calculation functions, and the machine learning predictions are completed through a multi-dimensional linked machine learning prediction model;
[0012] The business rule calculation function is expressed by the following formula
[0013]
[0014] Where,
[0015] S rule Represents the business rule calculation function,
[0016] W i represents the dynamic weight coefficient,
[0017] Represents the coefficient of variation of each discrete feature data set, where μ represents the overall mean and σ represents the overall standard deviation.
[0018] represents the nonlinear enhancement function,
[0019] k represents the curve steepness parameter,
[0020] Z0 represents the center threshold value of the setting,
[0021] Z i Indicates standard robustness,
[0022] X i Indicates the input product characteristics,
[0023] Median(X i ) represents the median of product characteristics,
[0024] MAD(X i ) represents the median difference of product characteristics;
[0025] The dynamic weighted fusion unit is used to fuse the preset business rule score with the machine learning prediction score according to the dynamic business rule weight factor. As data accumulates and the machine learning prediction model is iterated, the business rule weight factor is automatically reduced to achieve a smooth transition from rule-driven to data-driven, and finally obtain an intelligent product selection score, and make business decisions based on the final score.
[0026] Furthermore, in the feature engineering unit, the core value dimensions include interactive performance, platform performance, and sales performance.
[0027] Furthermore, in the feature engineering unit, the basic feature data set includes directly obtained sales volume for the current period, sales revenue for the current period, total number of reviews, number of product promotion notes, product promotion interaction volume, number of product promotion influencers, number of related influencers, number of related works, and number of related live broadcasts;
[0028] The statistical feature data set includes the average order value, evaluation rate, return rate, note conversion rate, expert effectiveness, average note interaction, and interaction conversion rate obtained through formula calculation;
[0029] The time series feature dataset includes the price elasticity coefficient, sales momentum index, month-on-month sales volume, and month-on-month collection volume obtained through formula calculation;
[0030] The composite feature data set includes a social communication index, a content conversion index, and an expert effectiveness index obtained through formula calculation.
[0031] Furthermore, in the hybrid computing unit, the machine learning prediction model is an XGBoost model, and the multi-dimensional linkage in the XGBoost model includes sales dimension, social dimension, and conversion dimension.
[0032] Furthermore, in the hybrid computing unit, the training process of the XGBoost model is as follows:
[0033] S1) setting the initialization sub-model and initializing the sub-model parameters;
[0034] The initialization sub-model includes three independent sub-models, namely, a sales sub-model, a social sub-model, and a conversion sub-model;
[0035] S2) In each iteration, the residual between the predicted value of the current sub-model and the true value is calculated;
[0036] S3) constructing a decision tree in each iteration, and then constructing a new decision tree based on the residual to fit the residual;
[0037] S4) adding the newly constructed decision tree to the sub-model, adjusting the parameters of the sub-model, and updating the predicted value;
[0038] S5) Repeat steps S2) to S4) until the iteration stops.
[0039] Furthermore, in the dynamic weighted fusion unit, the intelligent product selection score is calculated by the following formula:
[0040] S=α·S rule (X)+(1-α)·f model (X)·β
[0041] Where,
[0042] S represents the intelligent product selection score,
[0043] S rule Represents the business rule calculation function,
[0044] f model represents the machine learning prediction function,
[0045] X represents the product feature vector that is uniformly input to the business rule calculation function and the machine learning prediction function.
[0046] α represents the dynamic business rule weight factor,
[0047] β represents the loss factor.
[0048] Furthermore, in the dynamic weighted fusion unit, the default initial value of the dynamic business rule weight factor α is 0.55.
[0049] The advantages of the present invention are:
[0050] 1. Traditional scoring systems rely on manual empirical rules (highly subjective) or single algorithm models (poor interpretability), and are unable to adapt to data differences across multiple platforms. This invention adopts a dual-engine dynamic fusion scoring architecture, with the rule engine and machine learning running in parallel and layered. A hybrid computing layer is designed to synchronously execute preset business rules and prediction models, integrating the results through dynamic weighting factors. Through a progressive machine learning-led mechanism, the system is initially rule-driven, and the weighting factors are automatically reduced as data accumulates and the model is iterated, achieving a smooth transition from rule-driven to data-driven, thereby improving long-term prediction accuracy.
[0051] 2. Traditional scoring systems use universal features that are unable to capture the core value dimensions of different e-commerce platforms (Taobao, Xiaohongshu, and Douyin). This invention uses a multi-platform adaptive feature engineering system to model platform-differentiated features. For example, on the Xiaohongshu platform, social interaction indicators (product interaction volume and note conversion rate) are prioritized; on the Douyin platform, communication effect indicators (number of related influencers, number of works, and number of live broadcasts) are strengthened; on the Taobao platform, time-series growth indicators (sales momentum and month-on-month sales) are focused. Through composite feature generation technology, a cross-dimensional indicator fusion formula is pioneered, such as the social communication index = (product interaction volume / number of product notes) * number of product influencers, which quantifies the breadth and efficiency of content dissemination and solves the one-sidedness of traditional single-dimensional indicators.
[0052] 3. Aiming at the statistical distortion (such as Z-score compression) caused by promotional outliers and high-priced long-tail products in traditional scoring systems, this invention adopts robust Z-score standardization. This replaces the traditional Z-score, using the median and median absolute difference (MAD) to mitigate outliers and adapt to the long-tail price distribution of the apparel industry. It also uses a nonlinear enhancement function to map the robust Z-score to the (0,1) interval, and uses the parameter k to control the steepness of the curve, improving the ability to identify high-potential products.
[0053] 4. Traditional scoring systems use a single model that makes it difficult to balance multiple optimization objectives, such as sales, social interaction, and conversion. This invention implements a customized loss function and feature weight distribution by constructing a three-channel XGBoost sub-model. By using the dynamic weight of the coefficient of variation, features with large discreteness (such as the interaction volume of popular products) are automatically given higher weights in the rule engine, thereby improving the sensitivity of the indicator.
[0054] 5. Traditional scoring systems use manual data cleaning, feature extraction, and weight adjustment, which is inefficient. The present invention uses a four-layer automated pipeline (data preprocessing → feature engineering → dual-engine parallel computing → dynamic weighted fusion) and a designed self-iteration mechanism that automatically triggers parameter updates through model performance monitoring, reducing manual intervention costs.
[0055] The present invention is based on an intelligent product selection and scoring system that integrates a rule engine and machine learning. Through a dual-engine dynamic fusion architecture, platform differentiation feature modeling, an anti-interference scoring algorithm, and a multi-dimensional machine learning linkage mechanism, it solves the four core pain points of poor platform adaptability, sensitivity to outliers, difficulty in multi-target coordination, and high manual dependence in traditional apparel selection. It achieves scientific, automated, and high-precision product potential evaluation, providing a breakthrough technical framework for the e-commerce product selection field. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 This is a framework diagram of the intelligent product selection and scoring system in the present invention;
[0057] Figure 2 This is a partial training diagram of the machine learning prediction model in the present invention;
[0058] In the figure: data preprocessing unit 1, feature engineering unit 2, hybrid computing unit 3, dynamic weighted fusion unit 4. DETAILED DESCRIPTION
[0059] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0060] like Figure 1 As shown, the present invention provides an intelligent product selection and scoring system based on the fusion of rule engine and machine learning, including a data preprocessing unit 1, a feature engineering unit 2, a hybrid computing unit 3, and a dynamic weighted fusion unit 4.
[0061] The data preprocessing unit 1 is used to clean the collected original data of intelligent product selection and process missing data and abnormal data.
[0062] The feature engineering unit 2 is used to generate a product feature data set from the preprocessed smart product selection data according to the core value dimensions of different smart product selections. The product feature data set includes a basic feature data set, a statistical feature data set, a time series feature data set and a composite feature data set.
[0063] Specifically, in feature engineering unit 2, the core value dimensions include interactive performance, platform performance, and sales performance.
[0064] Specifically, in Feature Engineering Unit 2, the basic feature dataset includes directly acquired data on sales volume for the current period, sales revenue for the current period, total reviews, number of product promotion notes, product promotion interactions, number of product promotion influencers, number of associated influencers, number of associated works, and number of associated live broadcasts. See Table 1 below for a detailed display of basic feature indicators.
[0065] Table 1 Indicators of basic characteristics
[0066]
[0067] The statistical feature dataset includes the average order value, review rate, return rate, note conversion rate, influencer effectiveness, average note interaction, and interaction conversion rate, calculated through formulas. See Table 2 below for a detailed display of these statistical feature indicators.
[0068] Table 2 Indicators display of statistical characteristics
[0069]
[0070]
[0071] The time series feature dataset includes the price elasticity coefficient, sales momentum index, month-on-month sales growth, and month-on-month collection growth, calculated using a formula. See Table 3 below for a detailed display of these time series feature indicators.
[0072] Table 3 Indicators of timing characteristics
[0073]
[0074] The composite feature dataset includes the social communication index, content conversion index, and influencer effectiveness index calculated through formulas. The composite feature indicators are shown in Table 4 below.
[0075] Table 4 Index display of composite features
[0076]
[0077] The hybrid computing unit 3 is used to synchronously execute preset business rule calculation and machine learning prediction; the preset business rule calculation is completed through the business rule calculation function, and the machine learning prediction is completed through a multi-dimensional linked machine learning prediction model.
[0078] The business rule calculation function is expressed by the following formula
[0079]
[0080] Where,
[0081] S rule Represents the business rule calculation function,
[0082] W i represents the dynamic weight coefficient,
[0083] Represents the coefficient of variation of each discrete feature data set, where μ represents the overall mean and σ represents the overall standard deviation.
[0084] represents the nonlinear enhancement function,
[0085] k represents the curve steepness parameter,
[0086] Z0 represents the center threshold value of the setting,
[0087] Z i Indicates standard robustness,
[0088] X i Indicates the input product characteristics,
[0089] Median(X i ) represents the median of product characteristics,
[0090] MAD(X i ) represents the median difference of product characteristics.
[0091] in, The robustness of the standard is improved by using the median and median absolute difference (MAD) instead of the traditional mean standard deviation to improve outlier resistance. The constant 1.4826 makes the MAD estimate consistent with the standard deviation of the normal distribution. Compared with the traditional Z-score method, the robust Z-score is more resistant to interference and is more suitable for analyzing actual business data in e-commerce scenarios where outliers are present. Furthermore, long-tail distributions (such as individual high-priced luxury items) are common in apparel data, and traditional methods will cause the Z-scores of most normal products to be compressed into a narrow range.
[0092] It represents a nonlinear enhancement function, uses the Sigmoid function to convert the robust Z score into the (0,1) interval, controls the steepness of the curve through the parameter k, and Z0 is the set central threshold.
[0093] represents the dynamic weight coefficient, where is the coefficient of variation of each discrete feature data set. In this way, the weight coefficient can be automatically adjusted according to the feature variation coefficient, and the feature data index with large discreteness will obtain a higher weight.
[0094] In hybrid computing unit 3, the machine learning prediction model is an XGBoost model. XGBoost is a machine learning algorithm based on gradient boosting decision trees. It is efficient and robust. XGBoost supports custom loss functions and can be flexibly adjusted according to different business needs. It is suitable for processing large-scale data and high-dimensional features in e-commerce business scenarios. The multi-dimensional linkage in the XGBoost model includes sales dimension, social dimension, and conversion dimension.
[0095] In hybrid computing unit 3, the training process of the XGBoost model is as follows:
[0096] S1) Set the initialization sub-model and initialize the sub-model parameters.
[0097] The initialization sub-model includes three independent sub-models, namely the sales sub-model (sales_model), the social sub-model (social_model), and the conversion sub-model (conversion_model), which respectively perform training, prediction and scoring on the sales dimension, social dimension, and conversion dimension, and initialize the three sub-model parameters, such as tree depth = 5, learning rate = 0.05, subsampling ratio and column sampling ratio, etc.
[0098] S2) In each iteration, the residual between the predicted value of the current sub-model and the true value is calculated.
[0099] The objective function is defined as the combination of prediction loss and model complexity, for the tth iteration:
[0100]
[0101] Where,
[0102] Obj (t) represents the objective function of the XGBoost model at the tth iteration,
[0103] y i is the true value of the i-th sample,
[0104] is the predicted value of the first t-1 iterations,
[0105] f t (x i ) is the prediction function of the t-th iteration,
[0106] l is the loss function, which is used to measure the difference between the predicted value and the true value.
[0107] Ω(f t ) is a regularization term used to control the complexity of the sub-model.
[0108] To optimize the objective function, XGBoost uses gradient descent. In each iteration, the gradient and second-order derivative of each sample are calculated to further determine the step size and learning rate.
[0109] S3) constructs a decision tree in each iteration, and then constructs a new decision tree based on the residual to fit the residual.
[0110] Split node selection: In each leaf node, try all possible features and split points and calculate the split gain. The calculation formula of split gain is as follows:
[0111]
[0112] Where,
[0113] Gain represents the gain value of XGBoost decision tree node splitting,
[0114] I L and I R are the sample sets of the left and right child nodes after splitting,
[0115] I represents the sample set of the original node to be split.
[0116] λ and γ are regularization parameters used to control the complexity of the model.
[0117] g i represents the gradient of the i-th sample,
[0118] h i represents the second-order derivative of the i-th sample.
[0119] Leaf node weight calculation: For each leaf node, calculate its optimal weight by the following formula:
[0120]
[0121] represents the optimal weight,
[0122] I j represents the sample set contained in the jth leaf node,
[0123] g i represents the gradient of the i-th sample,
[0124] h i represents the second-order derivative of the i-th sample,
[0125] λ represents the regularization parameter.
[0126] S4) Add the newly constructed decision tree to the sub-model, adjust the parameters of the sub-model, and update the predicted value.
[0127]
[0128] Where,
[0129] is the predicted value at the tth iteration,
[0130] is the predicted value at the tth iteration,
[0131] f t (x i ) is the prediction function for the tth iteration.
[0132] S5) Repeat steps S2) to S4) until the iteration stops.
[0133] The stopping condition is met: the maximum number of iterations = 50, at which point the acceleration of model performance improvement approaches 0. At this time, the rule score and the model prediction score are combined according to the originally preset rule weight factor α to obtain the final score, which is then uploaded to the cloud agent for the model to make relevant business decisions based on the final score.
[0134] The dynamic weighted fusion unit 4 is used to fuse the preset business rule score with the machine learning prediction score according to the dynamic business rule weight factor. As data accumulates and the machine learning prediction model is iterated, the business rule weight factor is automatically reduced to achieve a smooth transition from rule-driven to data-driven, and finally obtain an intelligent product selection score, and make business decisions based on the final score.
[0135] Specifically, the intelligent product selection score is calculated by the following formula:
[0136] S=α·S rule (X)+(1-α)·f model (X)·β
[0137] Where,
[0138] S represents the intelligent product selection score,
[0139] S rule Represents the business rule calculation function,
[0140] f model represents the machine learning prediction function,
[0141] X represents the product feature vector that is uniformly input to the business rule calculation function and the machine learning prediction function.
[0142] α represents the dynamic business rule weight factor,
[0143] β represents the loss factor.
[0144] The dynamic business rule weighting factor, α, defaults to 0.55 and is used to balance the pre-set business rules with the machine learning prediction model. As the feature database improves and the number of model iterations increases, the α factor will gradually decrease, giving the machine learning prediction score a greater weight. This will enable a smoother transition from rule-driven to data-driven approaches and improve long-term prediction accuracy.
[0145] In addition, in addition to the changes in the weight factors of dynamic business rules, a loss factor β is also manually added to ensure that all potential categories of products on major platforms are fully captured.
[0146] This example demonstrates the implementation process of the aforementioned intelligent product selection and scoring system, based on a 2024 Taobao men's clothing category analysis dataset. Through data preprocessing, feature engineering, model training, rule scoring, and score fusion, the system generates product ratings, guiding e-commerce stores in selecting high-rated products (such as polo shirts and T-shirts) for sale, thereby optimizing inventory and sales. This report, combined with specific data, demonstrates the complete process from input to output, verifying the feasibility and practicality of this invention.
[0147] The original dataset is in Excel format and contains men's apparel data for the four quarters of 2024, covering four categories: polo shirts, T-shirts, knitwear / sweaters, and windbreakers. A partial dataset is shown in Table 5 below.
[0148] Table 5. Examples of original datasets
[0149]
[0150] The DataProcessor class cleans the raw dataset to ensure it is suitable for subsequent processing. Based on the preprocessed dataset, derived features are generated to capture the product's market performance and platform characteristics. The processed dataset in this example contains the following derived features, as shown in Table 6 (example data: Polo shirts, ¥0-50, Q1 2024).
[0151] Table 6 Dataset after feature engineering (partial fields)
[0152] Category Price range Average order value Sales conversion rate Social Communication Index Polo shirt ¥0-50 24358.7 0.1104 85531.91
[0153] Use the XGBoostTrainer class to train three XGBoost sub-models to predict sales, social, and conversion performance respectively. Figure 2 The figure shows a partial training diagram of a machine learning prediction model.
[0154] Among them, the sales model: the target is the average sales volume, and the features are the average order value, average price, month-on-month sales volume, and D_promo.
[0155] Social model: The target is the average collection, and the features are social communication index, average number of reviews, and cotton percentage.
[0156] Conversion model: The target is sales conversion rate, and the features are sales conversion rate, review conversion rate, and product quantity.
[0157] Training parameters: max_depth = 5, learning_rate = 0.05, n_estimators = 50, 80 / 20 train-validation split, objective function reg:squarederror.
[0158] The training results are as follows: for the sales model, the RMSE of the validation set is approximately 150; for the social model, the RMSE of the validation set is approximately 1000; and for the conversion model, the RMSE of the validation set is approximately 0.01.
[0159] Use the RuleScorer class to calculate rule scores. The specific example results are (Polo shirt, ¥0-50, Q1 2024): Average order value Z-score: 0.72; Social communication index Z-score: 1.05; Sales conversion rate Z-score: 0.88; Comprehensive rule score: 0.6723.
[0160] Finally, we use the FusionPredictor class to combine rule scoring and model prediction. For model prediction, we load three XGBoost models to predict sales, social, and conversion scores. The loss factor β is 0.95 for polo shirts and T-shirts, and 0.90 for sweaters / cardigans and windbreakers. The dynamic weight α decays linearly from 0.55 to 0.3 with a step size of 0.01. The formula for calculating the intelligent product selection score is S = α·S. rule (X)+(1-α)·f model (X)·β, taking Polo shirts (¥0-50, first quarter of 2024) as an example, the input data is as follows:
[0161] Number of products: 44;
[0162] Average sales: 703.59;
[0163] Average price: 34.14;
[0164] Average collection: 6370.80;
[0165] Total sales: 1,071,782.64;
[0166] Average number of reviews: 591.48;
[0167] Best product 1_feature: contains "promotion", D_promo=1;
[0168] Best Product 1_Brand: Decathlon.
[0169] Run FusionPredictor.predict and the result is:
[0170] rule_score: 0.6723;
[0171] model_score: 0.7412 (sales prediction: 710, social prediction: 6200, conversion prediction: 0.11, average post-multiplication β = 0.95);
[0172] final_score: 0.7045 (α=0.55);
[0173] beta: 0.95.
[0174] The scoring results are shown in Table 7 below.
[0175] Table 7 Product rating results
[0176] Category Price range Product Name Final Rating Recommended Sales Polo shirt ¥0-50 Decathlon men's short sleeve polo shirt 0.7045 yes Polo shirt ¥50-100 WOODSOON boys' POLO shirt 0.8150 yes t-shirt ¥0-50 Uniqlo Quick-drying Crew Neck T-shirt 0.7320 yes trench coat ¥200-500 BUTTBILL Japanese retro M51 windbreaker 0.9000 yes
[0177] The scoring results show that windbreakers priced between ¥200-500 (0.9000) are suitable for priority sales due to their high unit price and seasonal demand; polo shirts and T-shirts have higher scores in the low-price range (¥0-50, ¥50-100), making them suitable for large-scale promotions.
[0178] This embodiment uses the Taobao men's clothing category dataset to generate reliable product ratings (e.g., a Polo shirt with a rating of 0.7045 on a scale of ¥0-50), providing a scientific basis for store product selection and having significant practical value.
[0179] This example also verifies the advantages of the scoring system:
[0180] First, precision: Integrating rules and model scoring, integrating sales, social engagement, and conversions to accurately reflect product potential;
[0181] Second, robustness: Promotion dummy variables and dynamic weights adapt to fluctuations during major promotions, improving scoring stability by approximately 10%.
[0182] Third, efficiency: XGBoost training time is about 10-20 seconds, which is suitable for large-scale data;
[0183] Fourth, application: After recommending high-rated products, sales are expected to increase by 15% and the turnover cycle of slow-moving products will be shortened by 20%.
[0184] The present invention is based on an intelligent product selection and scoring system that integrates a rule engine and machine learning. Through a dual-engine dynamic fusion architecture, platform differentiation feature modeling, an anti-interference scoring algorithm, and a multi-dimensional machine learning linkage mechanism, it solves the four core pain points of poor platform adaptability, sensitivity to outliers, difficulty in multi-target coordination, and high manual dependence in traditional apparel selection. It achieves scientific, automated, and high-precision product potential evaluation, providing a breakthrough technical framework for the e-commerce product selection field.
Claims
1. An intelligent product selection and scoring system based on the integration of rule engine and machine learning, characterized by: It includes a data preprocessing unit (1), a feature engineering unit (2), a hybrid computing unit (3), and a dynamic weighted fusion unit (4); The data preprocessing unit (1) is used to clean the collected original data of intelligent product selection and process missing data and abnormal data; The feature engineering unit (2) is used to generate a product feature data set from the pre-processed smart product selection data according to the core value dimensions of different smart product selections, wherein the product feature data set includes a basic feature data set, a statistical feature data set, a time series feature data set and a composite feature data set; The hybrid computing unit (3) is used to synchronously execute preset business rule calculation and machine learning prediction; The preset business rule calculation is completed through the business rule calculation function, and the machine learning prediction is completed through the multi-dimensional linked machine learning prediction model; The business rule calculation function is expressed by the following formula Where, S rule Represents the business rule calculation function, W i represents the dynamic weight coefficient, Represents the coefficient of variation of each discrete feature data set, where μ represents the overall mean and σ represents the overall standard deviation. represents the nonlinear enhancement function, k represents the curve steepness parameter, Z0 represents the center threshold value of the setting, Z i Indicates standard robustness, X i Indicates the input product characteristics, Median(X i ) represents the median of product characteristics, MAD(X i ) represents the median difference of product characteristics; The dynamic weighted fusion unit (4) is used to fuse the preset business rule score with the machine learning prediction score according to the dynamic business rule weight factor. As data accumulates and the machine learning prediction model is iterated, the business rule weight factor is automatically reduced to achieve a smooth transition from rule-driven to data-driven, and finally an intelligent product selection score is obtained, and business decisions are made based on the final score.
2. The intelligent product selection and scoring system based on the integration of rule engine and machine learning according to claim 1 is characterized by: In the feature engineering unit (2), the core value dimensions include interactive performance, platform performance, and sales performance.
3. The intelligent product selection and scoring system based on the integration of rule engine and machine learning according to claim 2 is characterized by: Feature Engineering In unit (2), the basic feature data set includes directly obtained sales volume for this period, sales revenue for this period, total number of reviews, number of product-carrying notes, product-carrying interaction volume, number of product-carrying influencers, number of related influencers, number of related works, and number of related live broadcasts; The statistical feature data set includes the average order value, evaluation rate, return rate, note conversion rate, expert effectiveness, average note interaction, and interaction conversion rate obtained through formula calculation; The time series feature dataset includes the price elasticity coefficient, sales momentum index, month-on-month sales volume, and month-on-month collection volume obtained through formula calculation; The composite feature data set includes a social communication index, a content conversion index, and an expert effectiveness index obtained through formula calculation.
4. The intelligent product selection and scoring system based on the integration of rule engine and machine learning according to claim 1 is characterized by: In the hybrid computing unit (3), the machine learning prediction model is an XGBoost model, and the multi-dimensional linkage in the XGBoost model includes a sales dimension, a social dimension, and a conversion dimension.
5. The intelligent product selection and scoring system based on the integration of rule engine and machine learning according to claim 4 is characterized by: In the hybrid computing unit (3), the training process of the XGBoost model is as follows: S1) setting the initialization sub-model and initializing the sub-model parameters; The initialization sub-model includes three independent sub-models, namely, a sales sub-model, a social sub-model, and a conversion sub-model; S2) In each iteration, the residual between the predicted value of the current sub-model and the true value is calculated; S3) constructing a decision tree in each iteration, and then constructing a new decision tree based on the residual to fit the residual; S4) adding the newly constructed decision tree to the sub-model, adjusting the parameters of the sub-model, and updating the predicted value; S5) Repeat steps S2) to S4) until the iteration stops.
6. The intelligent product selection and scoring system based on the integration of rule engine and machine learning according to claim 1 is characterized by: In the dynamic weighted fusion unit (4), the intelligent product selection score is calculated by the following formula: S=α·S rule (X)+(1-α)·f model (X)·b Where, S represents the intelligent product selection score, S rule Represents the business rule calculation function, f model represents the machine learning prediction function, X represents the product feature vector that is uniformly input to the business rule calculation function and the machine learning prediction function. α represents the dynamic business rule weight factor, β represents the loss factor.
7. The intelligent product selection and scoring system based on the integration of rule engine and machine learning according to claim 6 is characterized by: In the dynamic weighted fusion unit (4), the default initial value of the dynamic business rule weight factor α is 0.55.
Citation Information
Cited By
Intelligent contract-driven supply chain economic optimization method
CN121998204A