A flow distribution method based on multi-dimensional information fitting

By establishing a traffic distribution model based on multi-dimensional information fitting, the problem of unreasonable traffic distribution in e-commerce platforms is solved, real-time monitoring and dynamic distribution of traffic are achieved, and user satisfaction and platform operation efficiency are improved.

CN119762116BActive Publication Date: 2025-10-03FOCUS TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411899667.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-10-03
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing e-commerce platforms have problems in traffic distribution, such as disproportionate investment and returns, limited product display sorting and uneven exposure, and limited traffic monitoring, which lead to low user satisfaction and unfair market competition.

Method used

By analyzing the input data and behavioral data within the service cycle, using equal-frequency binning and chi-square binning processing, a mathematical model between industry traffic and satisfaction is established, and the least squares method and third-order polynomial fitting are used to achieve real-time monitoring and dynamic allocation of traffic.

Benefits of technology

It improves the rationality and real-time performance of traffic distribution, enhances supplier satisfaction, and enhances the platform's operational efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119762116B_ABST
    Figure CN119762116B_ABST
Patent Text Reader

Abstract

This invention discloses a traffic allocation method based on multidimensional information fitting. Its characteristics are as follows: by analyzing input data and behavioral data within the service cycle, the relationship between input and traffic is explored. After processing the data into equal frequency bins and chi-square bins, the method of least squares is used to fit the data, establishing a mathematical model between industry traffic and satisfaction, thereby deriving a reasonable industry traffic allocation standard. This model can dynamically allocate traffic based on expected renewal rates and achieve real-time traffic monitoring. In this way, e-commerce platforms can more effectively manage and optimize traffic distribution, improve supplier satisfaction, and enhance the operational efficiency of the entire platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of Internet technology, and in particular relates to a traffic distribution method based on multi-dimensional information fitting. Background Art

[0002] The widespread adoption of the internet has profoundly transformed lifestyles and social structures, while also becoming a significant driver of industry diversification. The rise of e-commerce is a natural progression of this historical trend, spawning a vast array of e-commerce platforms and related service industries, which have become mainstream consumer behavior. Furthermore, the number of diverse industries is rapidly increasing, while the continued advancement of automation and artificial intelligence is further fragmenting and refining existing sectors.

[0003] As a large and comprehensive e-commerce platform, a systematic and comprehensive development of the traffic ecosystem is required. The rational distribution of traffic across different industries is particularly important. However, this is a complex technique. Currently, the most common method of allocating traffic across different industries is based on weighted distribution, which cannot effectively establish a correlation between a supplier's input-output ratio and supplier satisfaction. Furthermore, due to the significant differences between different industries, traffic distribution in light and heavy industries differs significantly. This leads to numerous related issues that need to be addressed urgently, including:

[0004] (1) The input and return are not proportional: By exploiting loopholes in the platform mechanism, the traffic entrance is controlled and traffic far exceeds the input. This leads to a siphon effect in the distribution of the introduced off-site traffic, so that the traffic is controlled by some users and cannot be distributed to other users. In this case, a small number of users who control a large amount of traffic will form an oligopoly, restrict competition and may have an adverse impact on the market.

[0005] (2) Limitations of product display sorting and uneven exposure: Products are displayed only in order of their richness and adaptability. This single sorting method can easily lead to some products occupying the main exposure positions, while other products cannot obtain effective exposure opportunities.

[0006] (3) Traffic monitoring has limitations: The platform's monitoring of industry-wide traffic only displays a numerical fluctuation. Although it can reflect the overall trend of traffic changes, it cannot specifically identify which specific industries or market segments have large traffic gaps. In addition, due to the lack of effective traffic evaluation, the platform cannot track and measure whether the traffic obtained by different users meets their expectations. If we rely on user feedback to adjust off-site traffic, since feedback is often raised after traffic problems are observed, and it takes a long time for the platform to respond and implement solutions, this delay means that the problem cannot be solved immediately, which can easily increase customer dissatisfaction and cause customer churn.

[0007] To sum up, e-commerce platforms need an effective traffic distribution technology that can not only detect traffic distribution problems in various industries early, but also regulate unreasonable traffic distribution in a timely manner to improve platform user satisfaction. Summary of the Invention

[0008] In order to solve the existing technical problems, the present invention provides a traffic allocation method based on multi-dimensional information fitting. By analyzing the input data and behavioral data within the service cycle, the relationship between input and traffic is explored, and after performing equal frequency binning and chi-square binning processing on the data, a mathematical model between industry traffic and satisfaction is established using the least squares method and third-order polynomial fitting, thereby deriving a reasonable and reliable industry traffic allocation standard. The model can dynamically allocate traffic based on the expected renewal rate and realize real-time monitoring of traffic. In this way, the e-commerce platform can more effectively manage and optimize traffic distribution, improve supplier satisfaction, and enhance the operational efficiency of the entire platform. The present invention has the characteristics of data-driven, dynamic adjustment and real-time monitoring, and is of great significance for the optimization of traffic distribution in the field of e-commerce.

[0009] The technical solution of the present invention is as follows: A flow distribution method based on multi-dimensional information fitting, specifically comprising

[0010] Step 1: Data Collection: Set up global tracking points within the e-commerce platform to obtain and record in real time the input data of supplier users and the behavioral data of buyer users within a service period; the service period refers to the contract period for supplier users to purchase platform services; the behavioral data of buyer users is data collected during the user's product search, product browsing, and product inquiries, including product ID, product industry ID, product exposure times, product visit times, and product inquiry times; the product exposure times refers to the frequency of product appearance in front of buyer users; the product visit times refers to the frequency of buyer users clicking and visiting product pages; the product inquiry times refers to the frequency of buyer users initiating product inquiries; the product supplier's input data includes renewal identifiers and daily industry input;

[0011] Step 2: Data preprocessing: mining and constructing the relationship between the behavioral data and the investment data to form an initial data set; the information in the initial data set includes industry ID, supplier ID, traffic, and renewal flag; traffic refers to the number of product exposures, visits, and product inquiries obtained per 10,000 yuan of industry investment; the renewal flag refers to a value assigned based on the supplier user's renewal status during the service period. If the supplier user renews within the service period, the renewal flag is recorded as 1, otherwise it is recorded as 0; the initial data set is divided by industry and stored in an array;

[0012] The step 2 specifically includes:

[0013] Step 2-1: Calculate the supplier's daily industry input for each product industry that the supplier inputs. Among them C i represents the supplier's daily industry input in industry i, n is the total number of product industries the supplier has invested in, C is the supplier's daily input, and the daily input = the supplier's total input / the number of days in the service cycle; the prod i is the number of valid products in industry i; valid products refer to products for which user behaviors exist, including search behaviors, visit behaviors, and inquiry behaviors;

[0014] Step 2-2: Calculate the daily industry traffic received by the supplier, and associate it with the daily industry input in step 2-1-4: Calculate the industry traffic of each industry based on the industry input of the supplier. The industry traffic refers to the sum of the number of exposures, visits, and inquiries received by products in the industry every day.

[0015] Step 2-3: Calculate the flow F corresponding to every 10,000 yuan of industry investment i =f i ×(10000 / C i ), where f i represents the industry traffic that the supplier obtains daily in industry i; Fi represents the traffic that the supplier obtains per 10,000 yuan per day in industry i;

[0016] Step 2-4: Obtain the supplier's renewal ID and associate the renewal ID with the flow F calculated in step 2-3 i , forming the initial data set;

[0017] Step 3: Process the traffic feature data using binning technology: Extract traffic and corresponding renewal identifiers from the initial data set, sort the extracted data in ascending order by traffic, and then bin the traffic using the equal-frequency binning method to evenly distribute the extracted data across different bins. Then, use the chi-square binning method to merge adjacent bins, and after merging stops, obtain the renewal rate for each bin.

[0018] In step 3, Python's pandas.qcut function is used to perform equal-frequency binning, and the initial number of bins is set to 15;

[0019] In step 3, the chi-square binning is used to merge adjacent bins, and the specific steps include:

[0020] Step 3-1: Create a contingency table, where the rows of the contingency table are the bin numbers and the columns of the contingency table are the renewal identifiers;

[0021] Step 3-2: Count the frequency of occurrence of different renewal marks in different bins and save them in a contingency table, which is recorded as the observed frequency O jm , represents the observation frequency of row j and column m;

[0022] Step 3-3: Calculate the expected frequency based on the observed frequency in the contingency table, where E km is the expected frequency of the kth row and mth column, H m is the sum of the observed frequencies in the mth column, R k is the sum of the observed frequencies in the kth row, and N is the sum of all observed frequencies in the contingency table. The calculation formula is:

[0023]

[0024] Step 3-4: Analyze and determine whether to merge adjacent boxes based on the observed frequency and expected frequency, specifically including: calculating the chi-square value of adjacent boxes respectively, screening out the minimum chi-square value of adjacent boxes, and the chi-square value of the adjacent boxes. Among them E k represents the expected frequency sum of the kth row; the minimum chi-square value is compared with a threshold value; if it is less than the threshold value, the data in the adjacent boxes are merged; if it is greater than the threshold value, the data in the adjacent boxes are stopped from being merged; the threshold value is the critical value of the chi-square distribution confirmed based on the degrees of freedom and the significance level, the significance level is 0.05, the degrees of freedom = (r-1) × (c-1), where r is the number of rows in the contingency table and c is the number of columns in the contingency table;

[0025] After stopping the merging of bins in steps 3-4, the WOE of each bin is calculated using Python's own function. d Value and IV d value, where d represents the bin number, The IV d =(Renewal rate - Non-renewal rate)*WOE d If WOE d As the bin number increases, IV d In the interval [0.02, 0.5], the binning results are robust;

[0026] In steps 3-4, the number of bins is preferably combined into 8.

[0027] Step 4: Data fitting to generate a mathematical model between traffic value and renewal rate: Obtain the binned data from step 3, select the bin number as the horizontal axis, select the renewal rate of the corresponding bin as the vertical axis, and plot the discrete points. Use the least squares method to fit the discrete points into a third-order polynomial function to generate the fitting coefficients.

[0028] In step 4, the least square method is implemented by using the polynomial function polyfit(x, y, 3) of Matlab.

[0029] Step 5: Numerical analysis to determine the standard flow rate value based on the renewal rate: Use the polynomial function fitted in step 4, substitute the specific renewal rate to generate the corresponding value, analyze the bins that the value falls into, obtain the data interval of the bins, and select the upper limit of the data interval as the standard flow rate value;

[0030] In step 5, when a specific renewal rate is substituted into the corresponding value, if the value is a decimal, the value is rounded down;

[0031] In step 5, if no corresponding value exists for the specific renewal rate, the specific renewal rate is compared with the renewal rates of all sub-boxes. If the renewal rate is greater than the renewal rates of all sub-boxes, the maximum value is selected from the minimum value and the maximum value of the function. If the polynomial function does not have a minimum value, the value corresponding to the maximum value of the function is selected.

[0032] In step 5, if there are at least two values ​​when substituting a specific renewal rate, a value is selected according to the type of monotonic interval into which the value falls; if a value falls within the monotonic increasing interval of the function, the average of these values ​​is taken as the standard flow value; if no value falls within the monotonic increasing interval of the function, the average of the values ​​falling within the monotonic decreasing interval of the function is taken as the standard flow value.

[0033] The beneficial effects achieved by the present invention are:

[0034] (1) The present invention provides a method for analyzing user input data and behavior data during the service cycle, mining the relationship between input and traffic, performing equal frequency binning and chi-square binning on the data to complete the optimal binning, and performing polynomial fitting on the scatter point relationship after the optimal binning to obtain an optimal curve, thereby generating a mathematical model that characterizes the relationship between industry traffic and renewal rate. By overlapping the curve with the predetermined renewal rate, the traffic required by the industry under the premise of meeting a certain renewal rate can be easily obtained, thereby obtaining a reasonable and reliable industry traffic allocation standard. The model can dynamically allocate traffic based on the expected renewal rate and realize real-time monitoring of traffic. In this way, the e-commerce platform can more effectively manage and optimize traffic distribution, improve supplier satisfaction, and enhance the operational efficiency of the entire platform, which is of great significance for the optimization of traffic distribution in the e-commerce field.

[0035] (2) The present invention performs optimal binning of data based on the chi-square test. This method can significantly enhance the interpretability of the mathematical model, improve the predictive ability, and capture the nonlinear relationship between traffic and renewal rate, thereby improving the model's fit. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 Flowchart of a flow distribution method based on multi-dimensional information fitting in an embodiment of the present invention;

[0037] Figure 2 This is a schematic diagram of the process of preprocessing industry traffic data in an embodiment of the present invention;

[0038] Figure 3 Schematic diagram of the process of feature data processing based on chi-square binning in an embodiment of the present invention;

[0039] Figure 4 Schematic diagram of a fitting function curve for multi-dimensional information fitting in an embodiment of the present invention;

[0040] Figure 5 FIG1 is a diagram illustrating a function curve of traffic distribution analysis based on renewal rate in an embodiment of the present invention;

[0041] Figure 6 FIG2 is a schematic diagram of a function curve of traffic distribution analysis based on renewal rate in an embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0043] Figure 1 A flow distribution method based on multi-dimensional information fitting in an embodiment of the present invention specifically includes:

[0044] Step 101: Data collection, real-time acquisition and storage of traffic data: Set up global tracking points within the e-commerce platform to acquire and record supplier user investment data and buyer user behavior data within a service period. The service period refers to the contract period for supplier users to purchase platform services. The buyer user behavior data is collected during the user's product search, product browsing, and product inquiries, including product ID, product industry ID, product exposure times, product visit times, and product inquiries. The product supplier investment data includes whether the contract has been renewed and daily industry investment.

[0045] In the embodiment of the present invention, the number of product exposures refers to the frequency with which a product appears before a buyer user; the number of product visits refers to the frequency with which a buyer user clicks on and visits a product page; the number of product inquiries refers to the frequency with which a buyer user initiates a product inquiries;

[0046] Step 102: Data preprocessing: mining the relationship characteristics between behavioral data and input data to form an initial data set: Calculate the supplier user's daily industry input. After unified data processing, generate the corresponding traffic data for each 10,000 yuan of input by the supplier in different industries. Analyzing the relationship characteristics between behavioral data and input data in step 102 is intended to more accurately mine the correlation between the supplier's input data and the buyer's behavioral data within the service cycle. On the one hand, supplier input data reflects the investment in service provision and measures supply-side capabilities. On the other hand, buyer behavioral data captures the behavioral activities of service users within the same service cycle and reflects market demand and user preferences. These two types of data essentially reflect different aspects of traffic services. Mining these correlations provides decision-making basis for subsequent traffic allocation models, helping to achieve optimal allocation.

[0047] like Figure 2 As shown, the specific steps of step 102 include:

[0048] Step 102-1: For each product industry that the supplier invests in, calculate the supplier's daily industry input, including:

[0049] Step 102-1-1: Calculate the supplier's daily investment, where the daily investment = the supplier's total investment / the number of days in the service cycle;

[0050] Step 102-1-2: Calculate daily industry proportion: Obtain all product industries provided by suppliers and count the number of valid products within each industry and across all industries on a daily basis. Valid products refer to products for which user behavior exists, including search behavior, visit behavior, and inquiry behavior. Daily industry proportion = number of valid products within a single industry / number of valid products across all industries.

[0051] Step 102-1-3: Calculate the supplier's daily industry input, where C i represents the supplier's daily industry input in industry i, n is the total number of product industries the supplier has invested in, C is the supplier's daily input, and prod i is the number of effective products in industry i;

[0052]

[0053] Step 102-2: Calculate the daily industry traffic received by the supplier, and associate it with the daily industry input in step 102-1: Calculate the industry traffic of each industry based on the industry input of the supplier. The industry traffic refers to the sum of the number of exposures, visits, and inquiries received by products in the industry every day.

[0054] Step 102-3: Calculate the flow corresponding to every 10,000 yuan of industry investment, where f i represents the industry traffic that the supplier obtains daily in industry i; Fi represents the traffic that the supplier obtains per 10,000 yuan per day in industry i;

[0055] F i =f i ×(10000 / C i ) (Formula 1)

[0056] Step 102-4: Create an initial data set, including renewal status and daily traffic per 10,000 yuan: Obtain the supplier's renewal status during the service period and record the renewal ID. If the supplier user renews during the service period, the renewal ID is recorded as 1; if the supplier user does not renew in the server, the renewal ID is recorded as 0. The supplier's renewal ID is associated with the supplier's daily traffic per 10,000 yuan in steps 2-3 to form an initial data set for recording the traffic obtained by the supplier per 10,000 yuan invested in different industries. The information in the initial data set includes industry ID, supplier ID, industry traffic, and renewal ID.

[0057] Step 102-5: Divide the initial data set by industry and save the data in an array for subsequent analysis or processing.

[0058] Step 103: Process the traffic feature data using binning technology: extract industry traffic and corresponding renewal identifiers from the initial data set, sort the extracted data in ascending order by industry traffic, bin the industry traffic using the equal-frequency binning method, and evenly distribute the extracted data in different bins. Set the initial number of bins for equal-frequency binning to 15; then use the chi-square binning method to merge adjacent bins, and obtain the renewal rate of each bin after the merging stops.

[0059] Binning is a data grouping technique used to discretize continuous variables. The principle of the equal-frequency binning method is to sort the data in ascending order and then evenly divide the data into several intervals, each containing an equal amount of data. Each interval corresponds to a bin. In this way, continuous data can be divided into multiple bins to better present the distribution of the data. In the embodiment of the present invention, industry flow is a continuous variable, so the equal-frequency binning method is used to discretize the data, and the adjacent bins are further merged through the chi-square statistic, effectively converting the continuous industry flow data into a discrete form, which facilitates subsequent data analysis and model building.

[0060] In the present invention, the initial number of boxes is 15. The purpose of box division is to perform linear fitting using the least squares method later. After testing and verification by technical personnel, for a linear fitting model, the number of boxes is at least 5 times the number of model parameters. Therefore, the initial number of boxes is set to 15.

[0061] In step 103, Python's pandas.qcut function is used to perform equal frequency binning;

[0062] like Figure 3 As shown, in step 103, the method of merging adjacent bins using chi-square binning specifically includes:

[0063] Step 103-1: Create a contingency table, where the rows of the contingency table are the box numbers and the columns of the contingency table are the renewal identifiers;

[0064] Step 103-2: Count the occurrence frequencies of different renewal indicators in different bins and save them in a contingency table, which is recorded as the observed frequency O. jm , represents the observation frequency of row j and column m;

[0065] Step 103-3: Calculate the expected frequency based on the observed frequency in the contingency table, where E km is the expected frequency of the kth row and mth column, H m is the sum of the observed frequencies in the mth column, R k is the sum of the observed frequencies in the kth row, and N is the sum of all observed frequencies in the contingency table. The calculation formula is:

[0066]

[0067] Step 103-4: Calculate the chi-square value of adjacent boxes based on the observed frequency and the expected frequency, and merge adjacent boxes based on the threshold judgment. Specifically:

[0068] Step 103-4-1: Calculate the chi-square value of adjacent boxes respectively, and filter out the minimum chi-square value of adjacent boxes. Among them E k represents the expected total frequency of the k-th row.

[0069] Step 103-4-2: Compare the minimum chi-square value to a threshold. If it is less than the threshold, merge the data in the adjacent bins. If it is greater than the threshold, stop merging the data in the adjacent bins. The threshold is the critical value of the chi-square distribution determined based on the degrees of freedom and the significance level. The significance level is 0.05. The degrees of freedom = (r-1) × (c-1), where r is the number of rows in the contingency table, c is the number of columns in the contingency table, and r is the number of rows in the contingency table. The chi-square critical value can be obtained from a chi-square distribution table or calculated using statistical software.

[0070] After the merging of bins is stopped in step 103-4, the WOE of each bin is calculated based on d Value and IV d value, where d represents the bin number, The IV d =(Renewal rate - Non-renewal rate)*WOE d .

[0071] The WOE of each bin d Value and IV d The values ​​can be obtained through Python's own functions. If WOE d As the bin number increases, IV d The range [0.02, 0.5] indicates an increase in traffic volume per 10,000 yuan and an increase in renewals, which aligns with real-world intuition and business logic, and also confirms the robustness of the binning results. Robust binning indicates an appropriate number of bins, neither too many nor too few. Excessive binning leads to overly complex models and captures noise, while too few bins leads to underfitting and poor prediction performance on unknown data. Therefore, an appropriate number of bins is crucial for building predictive models with strong generalization capabilities. In practical applications, the number of bins requires multiple iterations and optimizations. In this embodiment of the present invention, after evaluation and testing, reducing the number of bins to 8 represents a balance point. This balance point was determined through a comprehensive assessment of cross-validation and model performance metrics (such as accuracy, precision, and recall). Cross-validation demonstrated stable performance, high accuracy, precision, and recall, consistent word of essence (WOE) values ​​for the bins, and a reasonable IV value. This also aligns with business logic, ensuring that the model has both good explanatory power and high accuracy when predicting unknown data.

[0072] Step 104: Data fitting, generating a mathematical model of the relationship between traffic value and renewal rate: Obtain the data after binning in step 103, select the bin number as the horizontal axis, select the renewal rate of the corresponding bin as the vertical axis, and plot discrete points; use the least squares method to fit the discrete points into a third-order polynomial to generate fitting coefficients. The reason for choosing the third order is that the technicians found during the experiment that too high an order will lead to overfitting of the data, which is not conducive to the generalization ability of the model; overfitting causes the model to perform well on the training data, but perform poorly on the unknown test data, so it cannot predict new, unknown data well. At the same time, too high an order will also increase the computational complexity of the model and affect the operating efficiency of the model. Therefore, in order to balance the model's fit, generalization ability and computational efficiency, the technicians chose a third-order polynomial for fitting.

[0073] In step 104, the least square method is used, specifically, the polynomial function polyfit(x,y,3) of Matlab is used to implement the least square method. Table 1 shows the input data, where x is the bin number and y is the renewal rate of the corresponding bin. After fitting, the following is formed: Figure 4 The figure shows a fitting function curve diagram of multi-dimensional information fitting industry traffic distribution in an embodiment of the present invention. According to the fitting function diagram, the standard traffic value required at a specific renewal rate can be easily calculated;

[0074] Table 1

[0075] x 0 1 2 3 4 5 6 7 y 0.4285 0.5714 0.4285 0.4285 0.5714 0.8571 0.5714 0.8571

[0076] Step 105: Numerical analysis, determine the standard flow value according to the renewal rate: use the polynomial function fitted in step 104, substitute the specific renewal rate to generate the corresponding numerical value, analyze the bin into which the numerical value falls, obtain the data interval of the bin, and select the upper limit of the data interval as the standard flow value; in step 105, substitute the specific renewal rate to generate the corresponding numerical value, if the numerical value is a decimal, round it down; the bin number x represents a data interval, not a continuous value, so even if the calculation result is a decimal, in fact, only existing and complete bins can be selected, so it needs to be rounded down. Figure 4 , when the renewal rate y = 60%, the corresponding bin value x = 4.3. After rounding down the decimal, x = 4. Therefore, the upper limit of the data interval of bin 4 is taken as the standard flow value;

[0077] In step 105, if no corresponding value exists for the specific renewal rate, the specific renewal rate is compared with the renewal rates of all sub-boxes. If the renewal rate is greater than the renewal rates of all sub-boxes, the maximum value is selected from the minimum value and the maximum value of the function. If the polynomial function does not have a minimum value, the value corresponding to the maximum value of the function is selected.

[0078] like Figure 5 A diagram is provided for illustrating a function curve of traffic distribution analysis based on renewal rate in an embodiment of the present invention. The dots in the diagram are scatter points formed by the renewal rates of each bin in the coordinate system. If the renewal rate y = 70%, there is no corresponding x value, and 70% is higher than the renewal rates of all bins. The minimum values ​​of the function are calculated to be 0.59 and 4.52, and the maximum value of the function is 3.01. The maximum value of the three values ​​is x = 4.52. After rounding down the decimal, x = 4. The upper limit value of the data interval corresponding to bin No. 4 is taken as the standard traffic value.

[0079] In step 105, if there are at least two values ​​for the specific renewal rate, a value is selected according to the type of monotonic interval in which the value falls; if a value falls within the monotonic increasing interval of the function, the average of these values ​​is taken as the standard flow value; if no value falls within the monotonic increasing interval of the function, the average of the values ​​falling within the monotonic decreasing interval of the function is taken as the standard flow value; Figure 6 Figure 2 is a schematic diagram of a function curve for flow distribution analysis based on renewal rate in an embodiment of the present invention. The dots in the figure are scatter points formed by the renewal rate of each bin in the coordinate system. If the renewal rate is 70%, there are two corresponding x values, namely x1=0.2 and x2=5.8, where x1 is in the monotonically decreasing interval of the function and x2 is in the monotonically increasing interval of the function. Therefore, x=5.8 is taken, and after rounding down the decimal, x=5. The upper limit value of the data interval corresponding to bin No. 5 is taken as the standard flow value. The reason for giving priority to the monotonically increasing interval of the function and then the monotonically decreasing interval in the present invention is to follow the positive relationship between the renewal rate and the flow value. The higher the renewal rate, the more flow is allocated to the product.

[0080] The present invention may also have many other implementation methods. The above embodiments do not limit the present invention in any way. Without departing from the spirit and essence of the present invention, those skilled in the art may make various corresponding changes and modifications based on the present invention. Any other improvements and applications made to the above embodiments in an equivalent transformation manner shall fall within the scope of protection of the claims attached to the present invention.

Claims

1. A flow distribution method based on multi-dimensional information fitting, characterized in that: include: Step 1: Data Collection: Set up global tracking points within the e-commerce platform to obtain and record in real time the input data of supplier users and the behavioral data of buyer users within a service cycle; the service cycle refers to the contract period for supplier users to purchase platform services; the behavioral data of buyer users is data collected during the user's product search, product browsing, and product inquiries, including product ID, product industry ID, product exposure times, product visit times, and product inquiry times; the product exposure times refers to the frequency of product appearance in front of buyer users; the product visit times refers to the frequency of buyer users clicking and visiting product pages; the product inquiry times refers to the frequency of buyer users initiating product inquiries; the supplier user input data includes renewal identifiers and daily industry inputs; Step 2: Data preprocessing: mining and constructing the relationship between the behavioral data and the investment data to form an initial data set; the information in the initial data set includes industry ID, supplier ID, traffic, and renewal flag; traffic refers to the number of product exposures, visits, and product inquiries obtained per 10,000 yuan of industry investment; the renewal flag refers to a value assigned based on the supplier user's renewal status during the service period. If the supplier user renews within the service period, the renewal flag is recorded as 1, otherwise it is recorded as 0; the initial data set is divided by industry and stored in an array; Step 3: Process the traffic feature data using binning technology: Extract traffic and corresponding renewal identifiers from the initial data set, sort the extracted data in ascending order by traffic, and then bin the traffic using the equal-frequency binning method to evenly distribute the extracted data across different bins. Then, use the chi-square binning method to merge adjacent bins, and after merging stops, obtain the renewal rate for each bin. Step 4: Data fitting to generate a mathematical model between traffic value and renewal rate: Obtain the binned data from step 3, select the bin number as the horizontal axis, select the renewal rate of the corresponding bin as the vertical axis, and plot the discrete points. Use the least squares method to fit the discrete points into a third-order polynomial function to generate the fitting coefficients. Step 5: Numerical analysis, determine the standard flow value based on the renewal rate: Use the polynomial function fitted in step 4, substitute the specific renewal rate to generate the corresponding numerical value, analyze the bins into which the numerical value falls, obtain the data interval of the bins, and select the upper limit value of the data interval as the standard flow value.

2. A flow distribution method based on multi-dimensional information fitting according to claim 1, characterized in that: The step 2 specifically includes: Step 2-1: Calculate the supplier's daily industry input for each product industry that the supplier inputs. Among them C i represents the supplier's daily industry input in industry i, n is the total number of product industries the supplier has invested in, C is the supplier's daily input, and the daily input = the supplier's total input / the number of days in the service cycle; the prod i is the number of valid products in industry i; valid products refer to products for which user behaviors exist, including search behaviors, visit behaviors, and inquiry behaviors; Step 2-2: Calculate the supplier's daily industry traffic, linking it to the daily industry input in step 2-1. Calculate the industry traffic for each industry, using the supplier's industry input as the unit. Industry traffic refers to the sum of daily impressions, visits, and inquiries for products in that industry. Step 2-3: Calculate the flow F corresponding to every 10,000 yuan of industry investment i =f i ×(10000 / C i ), where f i represents the industry traffic that the supplier obtains daily in industry i; Fi represents the daily traffic obtained by the supplier in industry i per 10,000 yuan; Step 2-4: Obtain the supplier's renewal ID and associate the renewal ID with the flow F calculated in step 2-3 i , forming the initial data set.

3. The flow distribution method based on multi-dimensional information fitting according to claim 2, characterized in that: In step 3, Python's pandas.qcut function is used to perform equal-frequency binning, and the initial number of bins is set to 15; In step 3, the chi-square binning is used to merge adjacent bins, and the specific steps include: Step 3-1: Create a contingency table, where the rows of the contingency table are the bin numbers and the columns of the contingency table are the renewal identifiers; Step 3-2: Count the frequency of occurrence of different renewal marks in different bins and save them in a contingency table, which is recorded as the observed frequency O jm , represents the observation frequency of row j and column m; Step 3-3: Calculate the expected frequency based on the observed frequency in the contingency table, where E km is the expected frequency of the kth row and mth column, H m is the sum of the observed frequencies in the mth column, R k is the sum of the observed frequencies in the kth row, and N is the sum of all observed frequencies in the contingency table. The calculation formula is: Step 3-4: Analyze and determine whether to merge adjacent boxes based on the observed frequency and expected frequency, specifically including: calculating the chi-square value of adjacent boxes respectively, screening out the minimum chi-square value of adjacent boxes, and the chi-square value of the adjacent boxes. Among them E k represents the expected frequency sum of the kth row; the minimum chi-square value is compared with the threshold value, if it is less than the threshold value, the data in the adjacent boxes are merged; if it is greater than the threshold value, the data in the adjacent boxes are stopped from being merged; the threshold value is the critical value of the chi-square distribution confirmed according to the degrees of freedom and the significance level, the significance level is 0.05, the degrees of freedom = (r-1)×(c-1), where r is the number of rows in the contingency table, and c is the number of columns in the contingency table.

4. The flow distribution method based on multi-dimensional information fitting according to claim 3, characterized in that: After stopping the merging of bins in steps 3-4, calculate the WOE of each bin d Value and IV d value, where d represents the bin number, The IV d =(Renewal rate - Non-renewal rate)*WOE d WOE of each bin d Value and IV d The values ​​can be obtained through Python's own functions. If WOE d As the bin number increases, IV d In the interval [0.02, 0.5], the binning results are robust; In steps 3-4, the number of bins is merged to 8.

5. The flow distribution method based on multi-dimensional information fitting according to claim 4, characterized in that: In step 4, the least square method is implemented by using the polynomial function polyfit(x, y, 3) of Matlab.

6. The flow distribution method based on multi-dimensional information fitting according to claim 5, characterized in that: In step 5, when a specific renewal rate is substituted into the corresponding value, if the value is a decimal, the value is rounded down; In step 5, if no corresponding value exists for the specific renewal rate, the specific renewal rate is compared with the renewal rates of all sub-boxes. If the renewal rate is greater than the renewal rates of all sub-boxes, the maximum value is selected from the minimum value and the maximum value of the function. If the polynomial function does not have a minimum value, the value corresponding to the maximum value of the function is selected. In step 5, if there are at least two values ​​when substituting a specific renewal rate, a value is selected according to the type of monotonic interval into which the value falls; if a value falls within the monotonic increasing interval of the function, the average of these values ​​is taken as the standard flow value; if no value falls within the monotonic increasing interval of the function, the average of the values ​​falling within the monotonic decreasing interval of the function is taken as the standard flow value.

Citation Information

Patent Citations

  • Traffic analysis method, public service traffic attribution method and corresponding computer system

    CN109936512A

  • Delivery area determination method and device and storage medium

    CN115705578A