Demand intention degree prediction method based on relationship model between demand side and commodity

By adopting a multi-task learning method based on the relationship model of demand-side and commodity in the supply chain scheduling industry, the problems of missing data, inconsistent labeling and sparse positive samples in the demand-side intention prediction are solved, and higher prediction accuracy and recall rate are achieved.

CN116051168BActive Publication Date: 2025-06-17SHANGHAI JIAOTONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310022714.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-08
Publication Date
2025-06-17
Estimated Expiration
2043-01-08

AI Technical Summary

Technical Problem

In the supply chain scheduling industry, it is difficult for the existing technology to effectively predict the demand side's intention to high-durable items, especially in the problems of missing data, inconsistent labeling and sparse positive samples, resulting in limited model training and prediction accuracy.

Method used

A multi-task learning method based on the relationship model between the demand side and the commodity is adopted. By modeling the multi-level conversion link of the demand side and designing a multi-gated hybrid expert network, combining the aggregate link and loss functions of different weights, the input data processing and preprocessing of the model are optimized to improve the prediction accuracy of the model.

Benefits of technology

It significantly improves the accuracy and recall rate of demand-side intention prediction, especially in high-segment recalls and recall ratios, which is about 5% higher than traditional models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116051168B_ABST
    Figure CN116051168B_ABST
Patent Text Reader

Abstract

A demand intention degree prediction method based on the relationship model between the demand side and commodities can make a more accurate inference of the demand side's intention degree through modeling and prediction. Furthermore, it helps the downstream personnel of the supply chain to better focus their energy on the demand sides with higher intention degrees by using the predicted scores of the model, thereby enhancing the matching degree between the demand sides and the items and improving the overall satisfaction of the demand sides.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of neural network applications, specifically a method for predicting the Demand Intention based on the relationship model between the demand side and commodities. Background Art

[0002] In recent years, with the proposal and vigorous implementation of the digital economy, big data and artificial intelligence have also developed rapidly accordingly. These technologies have not only brought about a major revolution in the computer field, but also brought perspectives and tools for the digital transformation and innovation of the supply chain scheduling industry. With the help of new information technologies such as the Internet, it is possible to collect the data characteristics of the demand side. Through technologies such as data mining and machine learning, a better demand-side experience and more accurate marketing services can be achieved. With the accumulation and precipitation of a large amount of behavioral data, the value of informatization has gradually shifted from "business support" to "driving change"; the informatization drive has shifted from "business drive" to "data drive".

[0003] In the past, in the supply chain scheduling industry, the collection of demand-side information and the judgment of the demand-side intention degree completely relied on the downstream personnel of the supply chain. The information of the demand side was obtained through various methods such as telephone and offline, and the downstream personnel of the supply chain judged the intention degree of the demand side for the items according to the information. This method consumes a lot of time and energy of the downstream personnel of the supply chain and is likely to spend too much time on many demand sides that are not willing to purchase items. The designed data link will use both the online mini-program and the data of the downstream supply chain as the information of the demand side, and the designed machine learning model can judge the intention degree of the demand side with much higher efficiency than manual work.

[0004] Considering the characteristics of the downstream industries of the physical supply chain and the problems of the real industrial world, there are three difficulties in the model used to model the intention of the demand side: First, from the perspective of data acquisition, compared with the model training data in theory, the data generated in the real industry will have more vacant values, resulting in a lot of information loss during model training, and because the data is not labeled by professional data labelers, there are also problems of inconsistency and non-standardization in the data labels. In addition, in the process of digital transformation, the previously fully manual data and the information currently read semi-manually and semi-automatically need to be integrated and unified, and the corresponding data formats and data distributions of the two will be different. Second, some high-durability items are purchased less frequently by the demand side. Compared with the transaction volume and transaction speed of purchasing small items, the number of positive samples corresponding to high-durability items will be very sparse, which is not conducive to model training and convergence. Third, the demand side will be more rational about high-durability items than small items, and less susceptible to interference from other factors, which requires more accurate data for model training and also puts higher requirements on the prediction accuracy of the model.

[0005] In addition to these three difficulties, there is another feature that is different from other demand-side intention prediction problems. There are generally multiple indicators to show the intention of the demand side. Taking the existing demand-side intention prediction method in the recommendation field as an example, the demand side will first click, then activate, and then convert, which respectively represent the demand side clicking on the page, downloading the app or activating the product, and paying in the app or ordering the product. These three indicators progressively represent the size of the demand side's intention. The later the conversion link is, the stronger the demand side's intention. The main goal of the existing demand-side intention prediction method in the recommendation field is the final conversion, because only when the conversion occurs will the merchant get actual rewards. Therefore, the first two indicators, click and activation, are used to assist in the prediction of the final conversion. In the modeling of high-durable goods, there are also three indicators: visiting physical stores, pre-signing, and actual subscription. From the beginning to the end, the demand side becomes sparse, and the representative intention becomes larger. The indicators are similar, and the main goals are similar, all to promote more final actual subscriptions. However, the difference is that the supply chain industry has a large number of downstream supply chain personnel active in the front line of contact with customers. The modeling and prediction of customer willingness by downstream personnel at this stage steadily exceeds the prediction of the big data model. Therefore, the big data model does not need to take the final conversion as the main goal and find ways to optimize it like the recommendation system. It can use the first step of visiting the physical store as the main goal to optimize the visit prediction accuracy. In a nutshell, the model goal of matching modeling demand sides with high-durable goods is to screen high-intention demand sides for downstream supply chain personnel, not to predict the final conversion success rate. Summary of the invention

[0006] In view of the above deficiencies in the prior art, the present invention proposes a method for predicting demand intention degree based on the relationship model between the demand side and commodities, which can more accurately infer the intention degree of the demand side through modeling and prediction, and then help the downstream personnel of the supply chain to better focus their energy on the demand sides with higher intention degrees by using the predicted scores of the model, so as to enhance the matching degree between the demand side and the items and improve the overall satisfaction of the demand side.

[0007] The present invention is realized through the following technical solutions:

[0008] The present invention relates to a method for predicting demand intention degree based on the relationship model between the demand side and commodities, including:

[0009] Step 1: Model the demand side willingness as a serialized behavior model of a multi-level conversion link, that is, different behaviors of the demand side can reflect their different willingness degrees, and there is a high probability of a precondition relationship between the behaviors.

[0010] The final payment probability of the multi-level conversion link is p(b) = p(i2v) * p(v2c) * p(c2b) + p(i2v) * p(v2b) + p(i2b), where: the initial state of the demand side is init, abbreviated as i, visited is abbreviated as v, contract is abbreviated as c, and buy is abbreviated as b; the probability of the first part of the link is the probability from the initial state to the visited state * the probability from the visited state to the contract state * the probability from the contract state to the buy state, representing that this part of the demand side follows the complete conversion link in sequence. The probability of the second part is the probability from the initial state to the contract state * the probability from the contract state to the buy state. This part of the demand side does not visit the factory, directly signs the contract and then buys. The probability of the third part is the probability from the initial state directly to the buy state. This part of the demand side skips the steps of visiting and signing the contract and directly pays for the purchase. At present, the first part of the demand side accounts for the vast majority, and the second and third parts of the demand side are very few.

[0011] Step 2: Build a multi-task learning model (multi-gate mixture-of-experts) to train and predict the visited, contracted, and bought behaviors of the demand side at the same time; at the upstream of the multi-task learning model, use three expert networks, and fuse the outputs of the three expert networks through a gate as the input of the unique network.

[0012] Step 3. Based on the multi-task learning model MMOE, an additional aggregation link is designed to reduce the training difficulty caused by the sparsity of positive samples. Specifically, when the three dedicated networks respectively output A, B, and C, the product obtained by multiplying the results of A and B is used as the output of task B, and the product of A, B, and C is used as the output of task C. That is, the predicted value of task B, result(B) = output(A) * output(B|A), and the predicted value of task C, result(C) = result(B) * output(C|B). To ensure the establishment of such conditional probability, the positive samples of task B are adjusted to B or C, so that result(B) * output(C|B) must hold, simplifying the problem.

[0013] Step 4. Design loss functions with different weight sizes.

[0014] Step 5. Clean and preprocess the input data of the model.

[0015] In the training stage, the samples used include the labels of arrival, signing, and subscription. The output of the first dedicated network is arrival or signing or subscription, the output of the second dedicated network is signing or subscription, and the output of the third dedicated network is subscription; in the online stage, the output of the first dedicated network is used, and the output result is mapped from a probability of 0-1 to a score of 1-10 through a segmenter.

[0016] The present invention relates to a system for implementing the above method, including: a data processing unit, a model unit, and an output conversion unit. Among them: the data processing unit deletes invalid features and unique features according to the input data, converts time features into corresponding binary classification features, converts discrete data features into continuous variables through one-hot encoding and target encoding, and then standardizes all features to the interval of [0,1] through standardization to obtain standard data; the model unit receives the standard data submitted by the data processing unit, and outputs three-dimensional features after being processed by an expert network, a gating network, and a dedicated network; the output conversion unit receives the three-dimensional features output by the model unit, uses the first-dimensional output as the final output of the model, and converts the probability of 0-1 into a score of 1-10 through a segmenter.

[0017] Technical effects

[0018] In the present invention, through the connection path in the multi-task learning model structure, the output value of the first task is multiplied item by item with the output value of the second network, and the result obtained is used as the output value of the second task. Similarly, the output value of the second task is also multiplied item by item with the output value of the third network and used as the output value of the third task. Additionally, different loss function ratios are set for the three tasks, with a relatively high ratio set for the first task, enabling the model to be more inclined to learn the first task. Compared with the prior art, the present invention can be approximately 5% higher than other tested models in terms of the recall rate index. The tested models include logistic regression, multi-layer perceptron, and multi-gate mixture-of-experts network. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of the serialization conversion link of the behavior of the demand side;

[0020] Figure 2 Schematic diagram of the difference between the research problem of the present invention and other problems;

[0021] Figure 3 Structural diagram of the MMOE model

[0022] Figure 4 Schematic diagram of the model structure of the present invention;

[0023] Figure 5 Schematic diagram of the comparison of the present model and the old model in terms of the visit index;

[0024] Figure 6 Schematic diagram of the comparison of the present model and the old model in terms of the signing index;

[0025] Figure 7 Schematic diagram of the comparison of the present model and the old model in terms of the subscription index. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] As Figure 1 shown, the present embodiment relates to a demand intention degree prediction method based on a relationship model between the demand side and commodities, including the following steps:

[0027] First step, model the procurement willingness degree and behavior of the demand side: According to the positions of data buried points, the key behaviors of the demand side are divided into visit, signing, and payment subscription, and the degree of intention of the demand side to purchase items is represented progressively through a multi-level conversion link.

[0028] The final payment probability of the multi-level conversion link is: p(b) = p(i2v) * p(v2c) * p(c2b) + p(i2v) * p(v2b) + p(i2b), where: the initial state of the demand side is init, abbreviated as i, visited is abbreviated as v, contracted is abbreviated as c, subscribed is abbreviated as b, 2 represents to, for example, i2v represents the transition from the initial state to the visited state, and p(i2v) represents the transition probability, and the same applies hereinafter.

[0029] Step 2: Construct the model structure: While using the multi-task learning model (MMOE) to learn three output metrics simultaneously, add the concatenation of three proprietary network outputs, and change the visited metric to the OR of the three metrics, that is, instead of defining the first network as visited, define it as a metric representing the intention degree of the demand side such as visited or contracted or subscribed.

[0030] As Figure 4 As shown, the model structure includes: a data processing unit, a model unit, and an output conversion unit. Among them, the data processing unit deletes invalid features and unique features according to the input data, converts time features into corresponding binary classification features, transforms discrete data features into continuous variables through one-hot encoding and target encoding, and then standardizes all features to the interval of [0,1] through standardization, and finally obtains the standard data that can be used for training or prediction. The model unit receives the standard data submitted by the data processing unit, and after being processed by the expert network, the gating network, and the proprietary network, outputs three-dimensional features. The final output conversion unit receives the three-dimensional features output by the model unit, then takes the first-dimensional output as the final output of the model, and uses a segmenter to convert the probability of 0-1 into a score of 1-10.

[0031] When the input data is 60-dimensional, 3 expert networks are used in the shared network layer. The structure of each expert network is the same, all of which are 60*75*30 single-hidden-layer multi-layer perceptron (MLP) models. The 30-dimensional outputs of the three expert networks are linearly combined through the gate network to obtain a 30-dimensional intermediate result, which is then input into three unique networks. The structure of the unique network is designed as 30*45*2. The final output is a two-dimensional result, representing the positive sample probability and negative sample probability of an index respectively, and the sum of the two is 1. The loss function used by the network is the cross-entropy function, and the weights of the loss functions of the three sub-tasks are set to unequal values, set to 0.8, 0.1, 0.1.

[0032] The simulation experiment of this implementation example compares the performance of this model and the old model on the same batch of data for the three metrics of visited, contracted, and subscribed, and accordingly obtains the number and proportion of positive samples for a certain label in each segment from 1 to 10.

[0033] AsFigure 5 As shown in the figure, it is a comparison chart of the effects of the present model and the old model on each segment of the visit label. In the figure, the number of visitors and the visit ratio on each segment are mainly compared. It can be seen that in terms of the number of visitors, the vast majority of the visit numbers of the present model are concentrated in the relatively high score segments, and the effect is better than that of the old model which is mainly concentrated in the middle segments, that is, the high scores given by the model can better reflect the high procurement intention of the demand side. In terms of the visit ratio, it is similar to the old model, and there is a greater advantage at 10 points.

[0034] As Figure 6 、 7 shown, it is a comparison chart of the effects of the present model and the old model on each segment of the signing and subscription labels. It can be seen that in terms of the number of people signing and subscribing, the present model can assign relatively high scores to the vast majority of demand sides that have achieved procurement, and can recall high-quality demand sides better. In terms of the recall ratio in the high score segments, the present model also has a higher conversion rate than the old model.

[0035] Through the results of the above simulation experiments, the designed model does have a huge advantage compared with the existing old model in terms of the recall number and recall ratio in the high score segments.

[0036] First, using the information of the post-conversion link to help train and learn the indicators of the pre-conversion link can effectively improve the prediction stability and recall rate of the pre-conversion link indicators. Second, make full use of the differences between the large procurement scenario and other online scenarios such as advertising and recommendation, that is, a large number of personnel in the downstream of the supply chain, and the optimization goal is in the pre-conversion link, playing a role in primary screening and checking. Second, use three output indicators as training labels, but finally output one indicator as the total intention degree, which is convenient for the two usage methods of training and prediction for the model.

[0037] Those skilled in the art can make partial adjustments to the above specific implementation in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the present invention.

Claims

1. A method for predicting demand intention degree based on a relationship model between a demand side and a commodity, characterized in that, include: Step 1: Model the willingness of the demand side as a serialized behavior model of a multi-level conversion link, that is, different behaviors of the demand side can reflect their different willingness, and there is a high probability that there is a precondition relationship between the behaviors; Step 2: Build a multi-task learning model to simultaneously train and predict the visit, contract signing, and subscription behaviors of the demand side. In the upstream of the multi-task learning model, use three expert networks and use the gate to fuse the outputs of the three expert networks as the input of the unique network. Step 3. Based on the multi-task learning model MMOE, an additional aggregation link is designed to reduce the training difficulty caused by the sparse positive samples. Specifically, ABC is output in three dedicated networks respectively, and the product obtained by multiplying the results of A and B is used as the output of task B, and the product of A, B and C is used as the output of task C, that is, the predicted value of task B result(B) = output(A)*output(B|A), and the predicted value of task C result(C) = result(B)*output(C|B). To ensure the establishment of such conditional probability, the positive sample of task B is adjusted to B or C, so that result(B)*output(C|B) can definitely hold, simplifying the problem; Step 4: Design loss functions with different weights; Step 5: Perform data cleaning and data preprocessing on the model input data; Step 6: The samples used in the training phase include labels of visit, contract, and subscription, with visit or contract or subscription as the output of the first proprietary network, contract or subscription as the output of the second proprietary network, and subscription as the output of the third proprietary network; in the online phase, the output of the first proprietary network is used, and the output result is mapped from a probability of 0-1 to a score of 1-10 through a segmenter; The final payment probability of the multi-level conversion link is p(b)=p(i2v)*p(v2c)*p(c2b)+p(i2v)*p(v2b)+p(i2b), wherein: the initial state of the demander is init, abbreviated as i, the visit is visited, abbreviated as v, the contract is contract, abbreviated as c, and the subscription is buy, abbreviated as b; the probability of the first part of the link is the probability from the initial state to the visit * the probability from the visit to the contract * the probability from the contract to the subscription, which means that this part of the demander goes through the complete conversion link in sequence, the second part of the probability is the probability from the initial state to the contract * the probability from the contract to the subscription, and the third part of the probability is the probability from the initial state directly to the subscription.

2. A system for implementing the method according to claim 1, characterized in that, include: Data processing unit, model unit and output conversion unit, wherein: the data processing unit deletes invalid features and unique features according to the input data, converts the time features into corresponding binary classification features, converts the discrete data features into continuous variables through one-hot encoding and target encoding, and then standardizes all features to the interval of [0,1] to obtain standard data; the model unit receives the standard data submitted by the data processing unit, and outputs three-dimensional features after being processed by the expert network, the gated network and the proprietary network; The output conversion unit receives the three-dimensional features output by the model unit, takes the first-dimensional output as the final output of the model, and uses a segmenter to convert the probability from 0-1 to a score from 1-10.

Citation Information

Patent Citations

  • A user purchase intention prediction method based on mobile big data

    CN109741112A

  • Commodity recommendation method, device and system, and storage medium

    CN113761347A