A User Mining Method and System Based on a Cross-Border E-commerce Platform

By analyzing the access data of users of cross-border e-commerce platforms, screening and classifying user behaviors, calculating the similarity of unpurchased users, identifying potential users, solving the problem of difficulty in identifying unpurchased users in the existing technology, and achieving effective expansion of sales channels.

CN120125277BActive Publication Date: 2025-07-29湖南工商大学
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510599635.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-29
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The existing technology is difficult to identify potential users from unpending users of cross-border e-commerce platforms, resulting in waste of marketing resources and limited sales channel expansion.

Method used

By analyzing the access data of users of cross-border e-commerce platforms, filter out access data with browsing time within the standard range based on the product category and browsing time of the access page, and classify them as purchased and unpurchased data, calculate the average number of visits and browsing time of purchased data as standard data, calculate the similarity of unpurchased data, and determine whether unpurchased users are marked as mining users.

Benefits of technology

It improves the accuracy of identification of unpurned users, achieves effective expansion of sales channels, and reduces the waste of marketing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125277B_ABST
    Figure CN120125277B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of user mining, and particularly to a user mining method and system based on a cross-border e-commerce platform. The present invention obtains the access data of each user in a store for a period of time from the background of the cross-border e-commerce platform, and then preprocesses and classifies the access data according to the commodity categories of the accessed pages. After that, in order to analyze the behavioral similarity between the purchased users and the non-purchased users, on the basis of classifying the preprocessed data into purchased data and non-purchased data, the data is further processed through the standard range of browsing duration, and the standard data and the comparison data are extracted. Finally, the similarity between the standard data and the comparison data is calculated through a self-designed similarity algorithm, so as to judge whether each non-purchased user is a user to be mined according to the similarity threshold, solving the technical problem that it is difficult to identify potential users from these non-purchased users by the existing methods, and achieving the purpose of expanding the sales channels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of user mining, and in particular, to a user mining method and system based on a cross-border e-commerce platform. Background Art

[0002] With the rapid development of cross-border e-commerce platforms, the mining and analysis of user behavior data have become one of the key technologies to improve the sales conversion rate. Currently, the widely used user mining method in the industry is mainly based on the RFM model (Recency, Frequency, Monetary). This model divides the purchased users into different value levels by analyzing the user's recent purchase time, purchase frequency, and consumption amount, so as to formulate differential marketing strategies.

[0003] However, the limitation of the RFM model is that its analysis object is only limited to users who have generated consumption behavior and cannot cover a large number of users who have not purchased but have potential value. According to the statistics of the 2022 Global E-commerce Industry Report, the user conversion rate of cross-border e-commerce platforms is generally lower than 5%. This means that more than 95% of the visiting users have not completed the purchase behavior. However, the existing methods are difficult to further identify potential users from these 95% of the visiting users, resulting in a waste of marketing resources and limited expansion of sales channels. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a user mining method and system based on a cross-border e-commerce platform, which solves the technical problem that it is difficult for the existing methods to identify potential users from these non-purchasing users and achieves the purpose of expanding sales channels.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: A user mining method based on a cross-border e-commerce platform, the method specifically includes the following steps:

[0006] S1. Obtain the access data of each user of any store on the cross-border e-commerce platform within any period of time;

[0007] S2. Classify the access data according to the commodity categories of the accessed pages, and screen out the access data whose browsing duration is within the standard browsing duration range to obtain a number of preprocessed data;

[0008] S3. Classify the preprocessed data into purchased data and non-purchased data, and use the average access times of each accessed page and the average browsing duration at each access time in the purchased data as standard data, and use the access times of each user to each accessed page and the browsing duration at each access time in the non-purchased data as comparison data;

[0009] S4. Calculate the similarity between the standard data and each comparison data under the corresponding commodity category in turn;

[0010] S5. Determine whether to mark the user corresponding to the un - purchased data as a user to be mined according to whether the similarity is greater than the similarity threshold;

[0011] If so, mark the user corresponding to the un - purchased data as a user to be mined and output;

[0012] If not, end.

[0013] Preferably, in step S2, it specifically includes the following steps:

[0014] S21. Obtain the commodity categories of each accessed page;

[0015] S22. According to the correspondence between each data in the access data and each accessed page, split the access data into several groups of data to be processed with the same number as the number of accessed pages;

[0016] S23. Sequentially determine whether the browsing duration of each data to be processed is within the standard range of browsing duration;

[0017] If so, mark the data to be processed as candidate data;

[0018] If not, mark the data to be processed as abnormal data;

[0019] S24. According to the commodity categories of the pages corresponding to the candidate data, divide the candidate data into several pieces of pre - processed data.

[0020] Preferably, in step S23, the method for setting the standard range of browsing duration is as follows:

[0021] S231. Obtain the accessed pages where the user has operation behaviors, mark them as the first accessed pages, and mark the remaining accessed pages as the second accessed pages;

[0022] S232. Obtain the operation interval time between any two adjacent operation behaviors of the user within the first accessed page to obtain a set of operation interval times;

[0023] S233. Obtain the browsing duration of the user within the second accessed page to obtain a set of browsing durations;

[0024] S234. Respectively obtain the numerical ranges of the user's operation interval time and the browsing duration within the second accessed page, and mark them as the operation interval and the browsing interval respectively;

[0025] S235. Determine whether there is an intersection between the operation interval and the browsing interval;

[0026] If so, enter step S236;

[0027] Otherwise, select a value between the minimum value of the operation interval and the maximum value of the browsing interval as the minimum value of the standard range of the browsing duration of the user, and use the preset value as the maximum value of the standard range of the browsing duration of the user;

[0028] S236. Calculate the first data density at any time in the set of operation interval times, and the second data density at the same time in the set of browsing durations;

[0029] S237. Establish a first curve of the first data density and a second curve of the second data density with time as the abscissa and data density as the ordinate respectively;

[0030] S238. Extract the curves where the first data density and the second data density exist simultaneously on the first curve and the second curve to obtain several segments of intersection curves;

[0031] S239. Calculate the average value of the time corresponding to each segment of the intersection curve, and establish the standard range of the browsing duration of the user with the smallest average value and the largest average value.

[0032] Preferably, in step S236, it specifically includes the following steps:

[0033] S2361. Select any time less than or equal to the minimum value of the browsing interval as the first statistical time;

[0034] S2362. Set the time gradient, and generate several statistical times according to the first statistical time and the time gradient. The largest statistical time is greater than the maximum value of the browsing interval and the operation interval; the calculation formula of the statistical time is:

[0035] t i = t1 + iΔt

[0036] In the above formula, t i and t1 respectively represent the i-th and the first statistical times, Δt represents the time gradient, and Δt > 0;

[0037] S2363. Set a time interval at any statistical time, and the median of the time interval is the corresponding statistical time;

[0038] S2364. Respectively obtain the first data density and the second data density according to the number of data in the time intervals corresponding to the statistical times in the set of operation interval times and the set of browsing durations.

[0039] Preferably, in step S2364, it specifically includes the following steps:

[0040] S23641. Obtain the first data range and the second data range corresponding to the data in the operation interval time set and the browsing duration set within the time intervals corresponding to each statistical time;

[0041] S23642. Correct the first data density and the second data density according to the first data range, the second data range and the time interval; The calculation formulas for the corrected first data density and the second data density are:

[0042]

[0043] In the above formula, ρ′ represents the corrected first data density or the second data density, ρ represents the first data density or the second data density before correction, Δt represents the time gradient, and Δt' represents the first data range corresponding to the first data density or the second data range corresponding to the second data density.

[0044] Preferably, in step S3, it specifically includes the following steps:

[0045] S31. Classify the preprocessed data into purchased data and unpurchased data;

[0046] S32. Obtain the number of access times of each user in each preprocessed data in the purchased data to obtain the access frequency, and the browsing duration of each access;

[0047] S33. Count the number of users corresponding to each access frequency in each preprocessed data, and obtain the mode of this number of users to obtain the first mode;

[0048] S34. Set several duration intervals according to the numerical distribution interval of the browsing duration;

[0049] S35. Count the number of users whose browsing duration at each access frequency is within each duration interval, and obtain the mode of this number of users to obtain the second mode;

[0050] S36. Calculate the average access frequency of the purchased data according to the first mode, and calculate the average browsing duration at each access frequency according to the second mode;

[0051] The calculation formula for the average access frequency is:

[0052]

[0053] In the above formula, represents the average access frequency, N 1 represents the number of different access frequencies corresponding to the first mode, N j represents the j-th access frequency corresponding to the first mode;

[0054] The calculation formula for the average browsing duration is as follows:

[0055]

[0056] In the above formula, represents the average browsing duration, represents the median of the k-th duration interval corresponding to the second mode, and N 2 represents the number of duration intervals corresponding to the second mode;

[0057] S37. Take the average number of visits to the purchased data and the average browsing duration at each number of visits as the standard data, and take the number of visits of each user in the non-purchased data and the browsing duration at each number of visits as the comparison data.

[0058] Preferably, in step S4, it specifically includes the following steps:

[0059] S41. Extract the standard data and the comparison data under the corresponding product category;

[0060] S42. Obtain the average number of visits in the standard data;

[0061] S43. Calculate the difference in browsing duration between the standard data and the comparison data at the corresponding number of visits according to the average number of visits; the calculation formula for the difference in browsing duration is:

[0062]

[0063] In the above formula, ΔT f represents the difference in browsing duration at the f-th number of visits, represents the average browsing duration of the standard data at the f-th number of visits, and T f represents the browsing duration of the comparison data at the f-th number of visits. The number of visits f is an integer, and its value is between 1 and the average number of visits ;

[0064] S44. Calculate the consistency of the standard data, the consistency of the comparison data, and the consistency of the difference in browsing duration respectively according to the average number of visits;

[0065] S45. Calculate the similarity between the standard data and the comparison data according to the consistency of the standard data, the consistency of the comparison data, and the consistency of the difference in browsing duration.

[0066] Preferably, the calculation formula for consistency is:

[0067]

[0068] In the above formula, Y represents the consistency of the standard data, or the consistency of the comparison data, or the consistency of the difference in browsing duration, and t fIt represents the average browsing duration of the standard data at the f - th access count, or the browsing duration of the comparison data at the f - th access count, or the difference in browsing duration at the f - th access count. It represents the average value of the average browsing durations of the standard data at each access count, or the average value of the browsing durations of the comparison data at each access count, or the average value of the differences in browsing durations at each access count.

[0069] Preferably, the calculation formula for the similarity is:

[0070]

[0071] In the above formula, S represents the similarity between the standard data and the comparison data, Y standard data and Y comparative data respectively represent the consistency of the standard data and the consistency of the comparison data, and Y Difference represents the consistency of the difference in browsing duration.

[0072] A user mining system based on a cross - border e - commerce platform, including a processor and a memory. The memory is used to store a computer program, and when the computer program is executed by the processor, it implements the user mining method based on the cross - border e - commerce platform described above.

[0073] By means of the above technical solution, the present invention provides a user mining method and system based on a cross - border e - commerce platform, having at least the following beneficial effects:

[0074] 1. The present invention obtains the access data of each user in a store within a period of time from the background of the cross - border e - commerce platform, then pre - processes and classifies the access data according to the commodity categories of the accessed pages. After that, in order to analyze the behavioral similarity between the purchased users and the non - purchased users, on the basis of classifying the pre - processed data into purchased data and non - purchased data, it further processes the data and extracts the standard data and the comparison data, and finally calculates the similarity between the standard data and the comparison data through a self - designed similarity algorithm to determine whether each non - purchased user is a user to be mined according to the similarity threshold.

[0075] 2. The present invention sets the standard range of the browsing duration by finding the time distribution law between the operation interval time and the browsing duration of the user according to the data density, so as to realize the screening of the candidate data and ensure the accuracy and representativeness of the pre - processed data selected subsequently.

[0076] 3. The present invention converts discrete operation interval times and browsing durations into a two-dimensional distribution graph of visual statistical times and data densities, enabling intuitive observation of the density conditions of operation interval times and browsing durations at each statistical time, and thus intuitively observing the distribution conditions of operation interval times and browsing durations, facilitating subsequent analysis of the standard range of browsing durations.

[0077] 4. The present invention further optimizes the calculation methods of the first data density and the second data density, further improving the accuracy of the corrected first data density and second data density, reducing the errors in the calculation results of the first data density and the second data density caused by unreasonable time gradient settings, and improving the accuracy of the data.

[0078] 5. The present invention obtains the average access times and average browsing durations by extracting the mode, which can reduce the influence of a large amount of abnormal discrete data on the calculation results, ensuring that the calculated average access times and average browsing durations are more representative. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0080] Figure 1 is a flowchart of the user mining method based on the cross-border e-commerce platform of the present invention;

[0081] Figure 2 is a schematic diagram of the first curve and the second curve of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0082] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments. Thereby, a full understanding of how the present application uses technical means to solve technical problems and achieve technical effects can be obtained and implemented accordingly.

[0083] To solve the technical problem that it is difficult to identify potential users from these non-purchasing users in the existing methods, the present invention provides a user mining method based on a cross-border e-commerce platform, aiming to efficiently identify potential users to be mined by analyzing the behavioral similarities between non-purchasing users and purchasing users, thereby breaking through the limitations of the existing technology and providing a new growth path for the sales volume of the cross-border e-commerce platform. As Figure 1 shown, the method specifically includes the following steps:

[0084] S1. Obtain the access data of each user of any store on the cross-border e-commerce platform within any period of time. The cross-border e-commerce platform usually provides official channels for sellers to obtain access data. In this embodiment, the access data to be obtained includes user ID, user access time, accessed page for each visit, product category of the accessed page, and browsing duration of a single page. The browsing duration of a single page can be calculated by subtracting the time when the user enters the page from the time when the user exits the page. Of course, the store background can also directly obtain this data;

[0085] S2. Classify the access data according to the product category of the accessed page, and screen out the access data whose browsing duration is within the standard browsing duration range to obtain a number of preprocessed data. When users purchase products, they usually browse relevant pages of a certain type of product. For example, when searching for an electric rice cooker, the user will selectively enter the product pages related to the electric rice cooker. The number of visits to such pages and the browsing duration of a single page by the user have certain regularities. And according to the standard browsing duration range, some access data of users who accidentally enter this page can be excluded, so as to ensure the accuracy and representativeness of the data. Generally, the lower limit of the standard browsing duration range is 3 seconds, that is, the time when the user directly exits after accidentally entering a certain product page. The upper limit can be set to 20 minutes or 30 minutes, etc. The specific method for obtaining the preprocessed data includes the following steps:

[0086] S21. Obtain the product category of each accessed page. For example, an electric rice cooker can be used as a product category. Further, each brand of electric rice cooker can also be used as a product category respectively, or electric rice cookers of different sizes can be used as a category. The specific classification can be carried out according to the product types of the merchant;

[0087] S22. According to the corresponding relationship between each data in the access data and each accessed page, split the access data into several groups of data to be processed with the same number as the number of accessed pages, that is, take the user ID, user access time, accessed page for each visit, product category of the accessed page, and browsing duration of a single page within each accessed page as a group of data to be processed;

[0088] S23. Judge in turn whether the browsing duration of each group of data to be processed is within the standard browsing duration range;

[0089] If so, mark this group of data to be processed as data to be selected;

[0090] If not, mark this group of data to be processed as abnormal data;

[0091] S24. According to the product category of the page corresponding to the data to be selected, divide the data to be selected into several preprocessed data. For example, all those with the product category belonging to the electric rice cooker are used as a product category to obtain a preprocessed data;

[0092] The standard range of browsing duration is to better distinguish abnormal data. If the user's browsing duration is less than the lower limit of the standard range of browsing duration, the reason may be that the user has accidentally entered this page. If the user's browsing duration is greater than the upper limit of the standard range of browsing duration, it may be that the user has left due to other things during the browsing process. Moreover, for different products, there are also differences in the amount of information of different products. Therefore, there are also obvious differences in browsing duration. Therefore, reasonably setting the standard range of browsing duration is beneficial to ensuring the representativeness of the preprocessed data used for subsequent calculations. The method for setting the standard range of browsing duration is as follows:

[0093] S231. Obtain the access page where the user has an operation behavior and mark it as the first access page. The operation behavior refers to the operation instruction sent by the user to the access page other than entering the page and exiting the page. Mark the remaining access pages as the second access pages;

[0094] S232. Obtain the operation interval time between any two adjacent operation behaviors of the user within the first access page, so as to obtain a data set, that is, obtain the operation interval time set;

[0095] S233. Obtain the browsing duration of the user within the second access page, so as to obtain a data set, that is, obtain the browsing duration set;

[0096] S234. Respectively obtain the numerical ranges of the user's operation interval time and the browsing duration within the second access page, and mark them as the operation interval and the browsing interval respectively;

[0097] S235. Determine whether there is an intersection between the operation interval and the browsing interval;

[0098] If so, go to step S236;

[0099] If not, select a value between the minimum value of the operation interval and the maximum value of the browsing interval as the minimum value of the standard range of the user's browsing duration, and use a preset value as the maximum value of the standard range of the user's browsing duration. The preset value can be 20 minutes or 30 minutes, etc., which is determined according to the amount of information contained in the product page;

[0100] S236. Calculate the first data density of any time within the operation interval time set, and the second data density of the same time within the browsing duration set. The first density and the second density are used to reflect the density of the data, identify outliers through the density, so as to better set the lower limit value and the upper limit value of the standard range of browsing duration. The specific steps are as follows:

[0101] S2361. Select any time less than or equal to the minimum value of the browsing interval as the first statistical time. To reduce the computational load, the minimum value of the browsing interval can be directly selected as the first statistical time;

[0102] S2362. Set the time gradient. Generally, it can be set to 0.5 seconds, and generate several statistical times based on the first statistical time and the time gradient. The maximum value of the statistical times is greater than the maximum value of the browsing interval and the operation interval; the calculation formula for the statistical time is:

[0103] t i = t1 + iΔt

[0104] In the above formula, t i and t1 represent the i-th and the first statistical times respectively, Δt represents the time gradient, Δt > 0. For example, if the first statistical time is 3 s and the time gradient is 0.5 s, then 3.5 s, 4 s, and 4.5 s are the second, third, and fourth statistical times in sequence, and so on.

[0105] S2363. Set the time interval at any statistical time. The size of the time interval is generally not less than the size of the time gradient to prevent data omission. The median of the time interval is the corresponding statistical time;

[0106] S2364. Obtain the first data density and the second data density respectively according to the number of data in the time intervals corresponding to each statistical time in the operation interval time set and the browsing duration set. For some data, if the range of an interval is [3.25 s, 3.75 s], and the data in it are 3.26, 3.49, 3.41, and 3.45, then the data density can be 4. Of course, a certain formula can also be substituted for further calculation. Here, 4 is used as an example. If the data in it are 3.27, 3.45, 3.57, and 3.65, then the data density is also 4. However, it can be clearly seen that the former data is more concentrated. Therefore, the calculation method of the data density is further optimized. In step S2364, it specifically includes the following steps:

[0107] S23641. Obtain the first data range and the second data range of the data in the operation interval time set and the browsing duration set in the time intervals corresponding to each statistical time. The first data range and the second data range are obtained by subtracting the minimum value from the maximum value of the data in the operation interval time set and the browsing duration set. For example, for the above-mentioned distance with a data density of 4, the data range of the first group of data is 3.49 - 3.26 = 0.23, and the data range of the first group of data is 3.65 - 3.27 = 0.38.

[0108] S23642. Correct the first data density and the second data density according to the first data range, the second data range, and the time interval. The calculation formulas for the corrected first data density and the second data density are as follows:

[0109]

[0110] In the above formula, ρ′ represents the corrected first data density or the second data density, ρ represents the first data density or the second data density before correction, Δt represents the time gradient, and Δt′ represents the first data range corresponding to the first data density or the second data range corresponding to the second data density. It can be seen that in the example where the data densities of the above two groups are both 4, the time gradients Δt of the two groups of data are the same, but the data range of the first group of data is smaller. Therefore, the corrected data density calculated according to the formula is larger, which also conforms to the actual characteristics of the data.

[0111] S237. With time as the abscissa and data density as the ordinate, establish the first curve and the second curve of the first data density and the second data density respectively. The abscissa of the curve is the statistical time, and the ordinate is the data density. As Figure 2 shown, when there is data density at adjacent times, connect them. The black line in the figure is the second curve, and the red line is the first curve.

[0112] S238. Extract the curves on the first curve and the second curve where the first data density and the second data density exist simultaneously, that is, the statistical time corresponding to this section of the curve has corresponding first data density and second data density on the first curve and the second curve respectively, that is Figure 2 the partial curves on the first curve and the second curve corresponding to 2s, 2.5s, 4.5s, and 5s in, to obtain several sections of intersection curves. This section of the curve is the area where the two intersect. Therefore, it is necessary to select the lower limit and the upper limit of the standard range of browsing duration from this area.

[0113] S239. Calculate the average value of the time corresponding to each section of the intersection curve, and establish the standard range of the browsing duration of this user with the smallest average value and the largest average value. In practice, there may be intersection curves corresponding to multiple sections of statistical time. At this time, establish the standard range of the browsing duration of this user with the smallest average value and the largest average value.

[0114] S3. Classify the preprocessed data into purchased data and unpurchased data, and use the average number of visits to each accessed page and the average browsing duration at each number of visits in the purchased data as the standard data, and use the number of visits of each user to each accessed page and the browsing duration at each number of visits in the unpurchased data as the comparison data. To further illustrate the classification method of the purchased data and the unpurchased data, as well as the screening method of the comparison data and the standard data, the specific steps are as follows:

[0115] S31. Classify the preprocessed data into purchased data and unpurchased data;

[0116] S32. Obtain the number of access times for each user in each preprocessed data in the purchased data to obtain the access frequency, and the browsing duration of each access, that is, respectively count the access frequency and the browsing duration of each access for each user entering the same product type page;

[0117] S33. Count the number of users corresponding to each access frequency in each preprocessed data, that is, the number of users corresponding to each access frequency entering the same product type page, and obtain the mode of this number of users to obtain the first mode;

[0118] S34. Set several duration intervals according to the numerical distribution interval of the browsing duration;

[0119] S35. Count the number of users whose browsing duration corresponding to the corresponding access frequency is within each duration interval, and obtain the mode of this number of users to obtain the second mode;

[0120] S36. Calculate the average access frequency of the purchased data according to the first mode, and calculate the average browsing duration of each access frequency according to the second mode. By calculating the average access frequency and the average browsing duration through the mode, some abnormal outlier data can be excluded. The different access frequencies corresponding to the first mode may be one or more, and the different browsing durations corresponding to the second mode may also be one or more.

[0121] The calculation formula for the average access frequency is:

[0122]

[0123] In the above formula, represents the average access frequency, N 1 represents the number of different access frequencies corresponding to the first mode, N j represents the j-th access frequency corresponding to the first mode;

[0124] The calculation formula for the average browsing duration is:

[0125]

[0126] In the above formula, represents the average browsing duration, represents the median of the k-th duration interval corresponding to the second mode, N 2 represents the number of duration intervals corresponding to the second mode;

[0127] S37. Take the average number of accesses to the purchased data and the average browsing duration at each access count as the standard data, and take the number of accesses of each user in the unpurchased data and the browsing duration at each access count as the comparison data.

[0128] S4. Calculate the similarity between the standard data and each piece of comparison data under the corresponding product category in turn. To achieve the similarity degree between the standard data and the comparison data, since there are certain differences in the browsing duration among different people, a new algorithm is developed to calculate the similarity of the numerical change trends between the two, which specifically includes the following steps:

[0129] S41. Extract the standard data and the comparison data under the corresponding product category.

[0130] S42. Obtain the average number of accesses in the standard data.

[0131] S43. Calculate the difference in browsing duration between the standard data and the comparison data at the corresponding access count according to the average number of accesses; the calculation formula for the difference in browsing duration is:

[0132]

[0133] In the above formula, ΔT f represents the difference in browsing duration at the f-th access count, represents the average browsing duration of the standard data at the f-th access count, T f represents the browsing duration of the comparison data at the f-th access count, and the value is 0 when not accessed at the corresponding access count. The access count f is an integer, and the numerical value is between 1 and the average number of accesses ;

[0134] S44. Calculate the consistency of the standard data, the consistency of the comparison data, and the consistency of the difference in browsing duration respectively according to the average number of accesses; the calculation formula for consistency is:

[0135]

[0136] In the above formula, Y represents the consistency of the standard data, or the consistency of the comparison data, or the consistency of the difference in browsing duration, t f represents the average browsing duration of the standard data at the f-th access count, or the browsing duration of the comparison data at the f-th access count, or the difference in browsing duration at the f-th access count, represents the average value of the average browsing durations of the standard data at each access count, or the average value of the browsing durations of the comparison data at each access count, or the average value of the differences in browsing duration at each access count.

[0137] S45. Calculate the similarity between the standard data and the comparison data based on the consistency of the standard data, the consistency of the comparison data, and the consistency of the browsing duration difference. The formula for calculating the similarity is as follows:

[0138]

[0139] In the above formula, S represents the similarity between the standard data and the comparison data, and Y standard data and Y comparative data represent the consistency of the standard data and the consistency of the comparison data respectively, and Y Difference represents the consistency of the browsing duration difference.

[0140] S5. Determine whether to mark the users corresponding to the non-purchase curve as users to be mined according to whether the similarity is greater than the similarity threshold; if so, mark the users corresponding to the non-purchase curve as users to be mined; if not, end.

[0141] In the present invention, by obtaining the access data of each user of a store within a period of time from the background of the cross-border e-commerce platform, then preprocessing and classifying the access data according to the commodity categories of the accessed pages, and then, in order to analyze the behavioral similarity between the purchased users and the non-purchased users, on the basis of classifying the preprocessed data into purchased data and non-purchased data, further data processing is performed on them and standard data and comparison data are extracted, and finally, the similarity between the standard data and the comparison data is calculated through a self-designed similarity algorithm to determine whether each non-purchased user is a user to be mined according to the similarity threshold.

[0142] The present invention also provides a user mining system based on a cross-border e-commerce platform, including a processor and a memory. The memory is used to store a computer program, and when the computer program is executed by the processor, the user mining method based on the cross-border e-commerce platform is implemented.

[0143] Those of ordinary skill in the art can understand that all or part of the steps of implementing the method in the above embodiments can be completed by instructing relevant hardware through a program. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0144] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts between each embodiment can be referred to each other. For the above embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0145] The above embodiments have introduced the present invention in detail. Specific examples are used herein to elaborate on the principle and implementation of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A user mining method based on a cross-border e-commerce platform, characterized in that, The method specifically includes the following steps: S1. Obtain the access data of each user of any store on a cross-border e-commerce platform within any period of time; S2. Classify the access data according to the product categories of the accessed pages, and filter out the access data with the browsing duration within the standard browsing duration range to obtain several preprocessed data; Among them, the method for setting the standard browsing duration range is as follows: S231. Obtain the accessed pages where the user has operation behaviors, and mark them as the first accessed pages, and mark the remaining accessed pages as the second accessed pages; S232. Obtain the operation interval time between any two adjacent operation behaviors of the user within the first accessed page to obtain an operation interval time set; S233. Obtain the browsing duration of the user within the second accessed page to obtain a browsing duration set; S234. Respectively obtain the numerical ranges of the operation interval time of the user and the browsing duration within the second accessed page, and mark them as the operation interval and the browsing interval respectively; S235. Determine whether there is an intersection between the operation interval and the browsing interval; If so, go to step S236; If not, select a value between the minimum value of the operation interval and the maximum value of the browsing interval as the minimum value of the browsing duration standard range of the user, and use the preset value as the maximum value of the browsing duration standard range of the user; S236. Calculate the first data density of any time within the operation interval time set and the second data density of the same time within the browsing duration set; S237. Respectively establish a first curve of the first data density and a second curve of the second data density with time as the abscissa and data density as the ordinate; S238. Extract the curves where the first data density and the second data density exist simultaneously on the first curve and the second curve to obtain several segments of intersection curves; S239. Calculate the average value of the time corresponding to each segment of the intersection curve, and establish the browsing duration standard range of the user with the smallest average value and the largest average value; S3. Classify the preprocessed data into purchased data and unpurchased data, and use the average access times of each accessed page and the average browsing duration under each access time in the purchased data as standard data, and use the access times of each user to each accessed page and the browsing duration under each access time in the unpurchased data as comparison data; S4. Calculate the similarity between the standard data and each comparison data under the corresponding product category in turn; S5. Determine whether to mark the user corresponding to the unpurchased data as a user to be mined according to whether the similarity is greater than the similarity threshold; If so, mark the user corresponding to the unpurchased data as a user to be mined and output; If not, end.

2. The user mining method according to claim 1, wherein In step S2, it specifically includes the following steps: S21. Obtain the product categories of each accessed page; S22. According to the corresponding relationship between each data in the access data and each accessed page, split the access data into several groups of data to be processed with the same number as the number of accessed pages; S23. Determine in turn whether the browsing duration of each data to be processed is within the standard browsing duration range; If so, mark the data to be processed as data to be selected; If not, mark the data to be processed as abnormal data; S24. Divide the candidate data into several preprocessed data according to the product category of the page corresponding to the candidate data.

3. The user mining method according to claim 1, wherein In step S236, it specifically includes the following steps: S2361. Select any time less than or equal to the minimum value of the browsing interval as the first statistical time; S2362. Set the time gradient, and generate several statistical times according to the first statistical time and the time gradient, where the largest statistical time is greater than the maximum value of the browsing interval and the operation interval; the calculation formula for the statistical time is: t i = t1 + iΔt In the above formula, t i and t1 respectively represent the i-th and the 1st statistical time, Δt represents the time gradient, and Δt > 0; S2363. Set the time interval at any statistical time, and the median of the time interval is the corresponding statistical time; S2364. Obtain the first data density and the second data density respectively according to the number of data in the operation interval time set and the browsing duration set within the time intervals corresponding to each statistical time.

4. The user mining method according to claim 3, characterized in that In step S2364, it specifically includes the following steps: S23641. Obtain the first data range and the second data range of the data in the operation interval time set and the browsing duration set within the time intervals corresponding to each statistical time; S23642. Correct the first data density and the second data density according to the first data range, the second data range and the time interval; the calculation formulas for the corrected first data density and the second data density are: In the above formula, ρ ′ represents the corrected first data density or second data density, ρ represents the first data density or second data density before correction, Δt represents the time gradient, and Δt ' represents the first data range corresponding to the first data density or the second data range corresponding to the second data density.

5. The user mining method according to claim 1, characterized in that, In step S3, it specifically includes the following steps: S31. Classify the preprocessed data into purchased data and unpurchased data; S32. Obtain the number of access times of each user in each preprocessed data in the purchased data to obtain the access times, and the browsing duration of each access; S33. Count the number of users corresponding to each access time in each preprocessed data, and obtain the mode of the number of users to obtain the first mode; S34. Set several duration intervals according to the numerical distribution interval of the browsing duration; S35. Count the number of users whose browsing duration at the corresponding number of times is within each duration interval for each number of access times, and obtain the mode of the number of users to obtain the second mode; S36. Calculate the average access times of the purchased data according to the first mode, and calculate the average browsing duration at each number of access times according to the second mode; The calculation formula for the average access times is: In the above formula, represents the average number of accesses, N 1 represents the number of different access times corresponding to the first mode, N j represents the j-th access time corresponding to the first mode; The calculation formula for the average browsing duration is: In the above formula, represents the average browsing duration, represents the median of the k-th duration interval corresponding to the second mode, and N 2 represents the number of duration intervals corresponding to the second mode; S37. Take the average access times of the purchased data and the average browsing duration at each number of access times as the standard data, and take the access times of each user in the unpurchased data and the browsing duration at each number of access times as the comparison data.

6. The user mining method according to claim 1, wherein In step S4, it specifically includes the following steps: S41. Extract the standard data and the comparison data under the corresponding product category; S42. Obtain the average access times in the standard data; S43. Calculate the browsing duration difference between the standard data and the comparison data at the corresponding number of access times according to the average access times; The calculation formula for the browsing duration difference is: In the above formula, ΔT f represents the difference in browsing duration at the f-th visit count, represents the average browsing duration of the standard data at the f-th visit count, and T f represents the browsing duration of the comparison data at the f-th visit count. The visit count f is an integer, and its value is between 1 and the average visit count inclusive; S44. Calculate the consistency of the standard data, the consistency of the comparison data and the consistency of the browsing duration difference respectively according to the average access times; S45. Calculate the similarity between the standard data and the comparison data according to the consistency of the standard data, the consistency of the comparison data and the consistency of the browsing duration difference.

7. The user mining method according to claim 6, wherein The calculation formula for the consistency is: In the above formula, Y represents the consistency of the standard data, or the consistency of the comparison data, or the consistency of the difference in browsing duration, t f represents the average browsing duration of the standard data at the f-th access count, or the browsing duration of the comparison data at the f-th access count, or the difference in browsing duration at the f-th access count, represents the average value of the average browsing durations of the standard data at each access count, or the average value of the browsing durations of the comparison data at each access count, or the average value of the differences in browsing duration at each access count.

8. The user mining method according to claim 6, characterized in that, The calculation formula for similarity is as follows: In the above formula, S represents the similarity between the standard data and the comparison data, Y standarddata and Y comparativedata respectively represent the consistency of the standard data and the consistency of the comparison data, Y Difference represents the consistency of the difference in browsing duration.

9. A system for implementing the user mining method according to any one of claims 1-8, characterized in that, It includes a processor and a memory. The memory is used to store a computer program, and when the computer program is executed by the processor, it implements the user mining method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Method and apparatus for judging user interest degree

    CN105469284A

  • Potential customer identification method, device and system based on user behaviors and medium

    CN115063170A