A commercial data mining method, device, electronic device and product
By acquiring and analyzing user's personal information and historical behavior data, performing preference feature mining and user data mining, the problem of traditional methods relying on display feedback data is solved, and the efficiency of personalized product recommendations is achieved.
Patent Information
- Application Number
- CN202411334909.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-09-24
AI Technical Summary
Traditional user data mining methods based on collaborative filtering rely on user ratings and other display feedback data, resulting in the inability to effectively mine user data when the display feedback data is missing, and thus unable to make effective recommendations.
By obtaining the target user's personal information and historical behavior data, determining the purchase data of the product, and mining preference characteristics, the user's preference product characteristics are obtained. Then, based on these characteristics, first user data mining and second user data mining are carried out to generate optimal commercial mining data to realize personalized product recommendations.
There is no need to rely on user ratings and other display feedback data, which can effectively mine user data, improve the effectiveness of data mining, and realize personalized recommendations of products.
Smart Images

Figure CN119273378B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data mining, and particularly relates to a commercial data mining method, apparatus, electronic device and product. Background Art
[0002] With the continuous development of the Internet, online sales have become one of the main sales methods of many enterprises. When users browse Web sites online, product recommendation has become a main promotion method for enterprises. Among them, the more widely used personalized recommendation algorithms mainly include: content-based personalized recommendation, collaborative filtering personalized recommendation, and Web log-based personalized recommendation. The above-mentioned recommendation methods can provide personalized recommendation services for customers, which not only improve the sales performance of enterprises, but also greatly improve the user experience.
[0003] Currently, the recommendation method based on collaborative filtering (CF) is one of the most widely used recommendation methods. It mainly uses explicit feedback such as rating data to mine user commercial data, and obtains user preferences, interests, purchase tendencies, etc. So far, many researchers have conducted in-depth research on the collaborative filtering recommendation method based on user data mining from various aspects and have also achieved some remarkable results. However, the traditional user data mining method based on collaborative filtering has the following deficiencies: it needs to rely on explicit feedback data such as user ratings to mine user data, but this explicit feedback data is more often unclear. Therefore, in the case of missing explicit feedback data, user data cannot be effectively mined, and thus effective recommendation cannot be made. Therefore, based on the above deficiencies, how to provide a data mining method that does not rely on explicit feedback data and can effectively mine user commercial data has become an urgent problem to be solved. Summary of the Invention
[0004] The purpose of the present invention is to provide a commercial data mining method, apparatus, electronic device and product, which are used to solve the problem that in the case of missing explicit feedback data in the prior art, user data cannot be effectively mined, and thus effective recommendation cannot be made.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] In a first aspect, a commercial data mining method is provided, including:
[0007] Obtain the personal information of the target user and the historical behavior data of the target user on the e-commerce platform;
[0008] Based on the historical behavior data, determine the purchased product data of the target user;
[0009] Based on the purchase product data of the target user, perform preference feature mining processing on the target user, so as to obtain the preference product features of the target user after the preference feature mining processing;
[0010] Utilize the preference product features of the target user to perform the first user data mining processing on the target user, and obtain the first commercial mining data of the target user, wherein the first commercial mining data includes a number of first similar users having similar preference product features to the target user and the first user data corresponding to each first similar user;
[0011] Based on the historical behavior data and the personal information, perform the second user data mining processing on the target user, and obtain the second commercial mining data of the target user, wherein the second commercial mining data includes a number of second similar users having similar personal information and behavior data to the target user and the second user data corresponding to each second similar user;
[0012] Generate the optimal commercial mining data of the target user according to the first commercial mining data and the second commercial mining data, so as to utilize the optimal commercial mining data to perform personalized product recommendation for the target user.
[0013] Based on the above-disclosed content, the present invention first obtains the personal information of the target user and its historical behavior data on the e-commerce platform; then, according to the historical behavior data, determines the purchase product data of the target user, and based on this, performs user preference feature mining processing to obtain the preference product features of the target user; then, according to the preference product features, performs the first data mining processing on the target user to obtain the first commercial mining data including the first similar users having similar preference product features to the target user and the corresponding user data; then, utilizes the historical behavior data and the personal information to perform the second data mining processing to obtain the second commercial mining data including the second similar users having similar personal information and behavior data to the target user and the corresponding user data; finally, based on the foregoing first and second commercial mining data, the optimal commercial mining data of the target user can be obtained; thus, the personalized product recommendation for the target user can be completed by utilizing the optimal commercial mining data.
[0014] Through the above design, when the present invention performs user data mining, it is based on the attributes of the products themselves to mine and process the user preference features, so as to obtain the preferred product features of the user; then, based on the preferred product features, the first user data mining process is performed to obtain the first commercial mining data related to the user preference features; then, combined with the personal information and historical behavior data of the user himself, the second user data mining process is performed to obtain the second commercial mining data related to the user personal information and historical behavior; then, based on the aforementioned first commercial mining data and second commercial mining data, the optimal commercial mining data of the user can be obtained; finally, based on the optimal commercial mining data, the personalized recommendation service for the user can be completed; thus, the present invention can effectively mine the user data based on the products purchased by the user, personal information and historical behavior, without relying on explicit feedback data such as user ratings. Therefore, compared with the traditional technology, the effectiveness of data mining can be greatly improved, so that effective personalized recommendation of products can be realized. Therefore, it is very suitable for large-scale application and promotion.
[0015] In a possible design, according to the purchase product data of the target user, the preference feature mining process is performed on the target user to obtain the preferred product features of the target user after the preference feature mining process, including:
[0016] Based on the purchase product data of the target user, a product attribute set corresponding to the purchase product data is generated;
[0017] For the i-th product attribute in the product attribute set, the total number of products in the purchase product data that contain the i-th product attribute is counted as the attribute influence value corresponding to the i-th product attribute;
[0018] According to the attribute influence value corresponding to the i-th product attribute and according to the following formula (1), the importance of the i-th product attribute to the target user is calculated;
[0019]
[0020] In the above formula (1), P i represents the importance of the i-th product attribute to the target user, Y i represents the attribute influence value of the i-th product attribute, Y m represents the attribute influence value of the m-th product attribute, and n represents the total number of product attributes in the product attribute set;
[0021] Increment i by 1 and re-count the number of products in the purchase product data that contain the i-th product attribute until i is equal to n to obtain the importance of each product attribute to the target user, where the initial value of i is 1;
[0022] Filter out several product attributes from the set of product attributes according to the importance of each product attribute to the target user, so as to serve as the preferred product features of the target user.
[0023] In a possible design, the number of preferred product features of the target user is multiple. Among them, using the preferred product features of the target user, perform the first user data mining process on the target user to obtain the first commercial mining data of the target user, including:
[0024] Filter out several initial similar users corresponding to the target user from the user database of the e-commerce platform;
[0025] Determine the feature weights of each preferred product feature in the preferred product features of the target user according to the several initial similar users;
[0026] Calculate the first similarity between the target user and each user in the user database by using the multiple preferred product features of the target user and the feature weights of each preferred product feature;
[0027] Determine several first similar users corresponding to the target user from the user database according to the first similarity between the target user and each user;
[0028] Obtain the first user data of each first similar user, and use each first similar user and the first user data of each first similar user to form the first commercial mining data.
[0029] In a possible design, determining the feature weights of each preferred product feature in the preferred product features of the target user according to the several initial similar users includes:
[0030] Determine any product from the historical purchased products corresponding to each initial similar user as the test product;
[0031] Perform multiple initialization processes on the feature weights of each preferred product feature, so that after each initialization process, use the initialized feature weights of each preferred product feature to form an initial weight vector;
[0032] Generate multiple particle individuals by using the initial weight vector obtained from each initialization process. Among them, each particle individual corresponds to an initial position vector and an initial velocity vector, and the initial position vector of any particle individual corresponds to an initial weight vector;
[0033] Initialize the iteration count \(t\) to 1, and obtain the position vectors and velocity vectors of each particle individual at the \(t\)-th iteration. Among them, when \(t = 1\), the position vector and velocity vector of any particle individual at the \(t\)-th iteration are the initial position vector and initial velocity vector of the particle individual;
[0034] Calculate the fitness of each particle individual at the \(t\)-th iteration according to a number of initial similar users and the position vectors of each particle individual at the \(t\)-th iteration. Among them, the fitness of any particle individual at the \(t\)-th iteration is used to characterize the error of the purchase probability of the target user for the test product calculated using the position vector of the particle individual at the \(t\)-th iteration;
[0035] Update the global optimal position and the individual optimal positions of each particle individual based on the fitness of each particle individual at the \(t\)-th iteration. Among them, the global optimal position is the position vector corresponding to the particle individual with the smallest fitness during the \(1\)-st to \(t\)-th iterations;
[0036] Determine whether the iteration stop condition is satisfied. Among them, the iteration stop condition is that the fitness of the particle individual corresponding to the updated global optimal position is less than the fitness threshold, or \(t\) reaches the maximum number of iterations;
[0037] If not, then use the updated global optimal position and the updated individual optimal positions of each particle individual to update the velocity vector of each particle individual at the \(t\)-th iteration, obtain the updated velocity vectors corresponding to each particle individual, and use the updated velocity vectors corresponding to each particle individual to update the position vector of each particle individual, obtain the updated position vectors corresponding to each particle individual;
[0038] Increment \(t\) by 1, and replace the position vector and velocity vector of each particle individual at the \(t\)-th iteration with the corresponding updated position vector and updated velocity vector, and recalculate the fitness of each particle individual at the \(t\)-th iteration according to a number of initial similar users and the position vectors of each particle individual at the \(t\)-th iteration until the iteration stop condition is satisfied, so as to determine the feature weights of each preference product feature according to the global optimal position that satisfies the iteration stop condition.
[0039] In a possible design, calculating the fitness of each particle individual at the \(t\)-th iteration according to a number of initial similar users and the position vectors of each particle individual at the \(t\)-th iteration includes:
[0040] For any particle individual, determine the weight vector of the particle individual at the \(t\)-th iteration according to the position vector of the particle individual at the \(t\)-th iteration;
[0041] According to the characteristic of each preferred commodity of the target user and the weight vector of any particle individual at the t-th iteration, and according to the following formula (2), calculate the true similarity between the target user and each initial similar user;
[0042]
[0043] In the above formula (2), S(K,a) represents the true similarity between the target user K and the a-th initial similar user among several initial similar users, is the feature weight of the j-th preferred commodity feature in the weight vector of any particle individual at the t-th iteration, represents the importance of the j-th preferred commodity feature to the target user, represents the importance of the j-th preferred commodity feature to the a-th initial similar user, R represents the total number of preferred commodity features, where a = 1, 2,..., A, and A represents the total number of initial similar users;
[0044] Based on the true similarity between the target user and each initial similar user, and according to the following formula (3), calculate the purchase probability of the target user for the test commodity when using the weight vector corresponding to any particle individual as the feature weight of the preferred commodity feature;
[0045]
[0046] In the above formula (3), G represents the purchase probability of the target user for the test commodity, U(a) represents whether the a-th initial similar user has purchased the test commodity, where if the a-th initial similar user has purchased the test commodity, U(a) is 1, otherwise, U(a) is -1;
[0047] According to the purchase probability of the target user for the test commodity, and according to the following formula (4), calculate the fitness of any particle individual at the t-th iteration;
[0048] D t =|G - U(K)| (4)
[0049] In the above formula (4), D t represents the fitness of any particle individual at the t-th iteration, U(K) represents whether the target user has purchased the test commodity, where if the target user has purchased the test commodity, then U(K) is 1, otherwise, U(K) is -1.
[0050] In a possible design, based on the historical behavior data and the personal information, perform a second user data mining process on the target user to obtain the second commercial mining data of the target user, including:
[0051] Obtain the personal information and historical behavior data of each user in the user database of the e-commerce platform;
[0052] According to the personal information of the target user and the personal information of each user, calculate the personal information similarity between the target user and each user in the user database;
[0053] Based on the historical behavior data of the target user and the historical behavior data of each user, calculate the behavior similarity between the target user and each user in the user database;
[0054] According to the personal information similarity and behavior similarity between the target user and each user in the user database, calculate the second similarity between the target user and each user in the user database;
[0055] According to the second similarity between the target user and each user, determine several second similar users corresponding to the target user from the user database;
[0056] Obtain the second user data of each second similar user, and use each second similar user and the second user data of each second similar user to form the second business mining data.
[0057] In a possible design, based on the historical behavior data of the target user and the historical behavior data of each user, calculating the behavior similarity between the target user and each user in the user database includes:
[0058] Based on the historical behavior data of the target user, determine the set of behavior types of the target user;
[0059] For any user in the user database, count the number of behaviors of each type of behavior in the set of behavior types generated by the any user on the e-commerce platform;
[0060] Based on the number of behaviors of each type of behavior in the set of behavior types generated by the any user on the e-commerce platform, and according to the following formula (5), calculate the behavior similarity between the target user and the any user;
[0061]
[0062] In the above formula (5), Q(K,b) represents the behavior similarity between the target user and the any user, represents the number of behaviors of the x-th type of behavior generated by the target user on the e-commerce platform, represents the number of behaviors of the x-th type of behavior generated by the any user on the e-commerce platform, and X represents the total number of types of behaviors in the set of behavior types.
[0063] In a second aspect, a commercial data mining device is provided, including:
[0064] An acquisition unit, configured to acquire personal information of a target user and historical behavior data of the target user on an e-commerce platform;
[0065] The acquisition unit is further configured to determine purchase commodity data of the target user based on the historical behavior data;
[0066] A commodity feature mining unit, configured to perform preference feature mining processing on the target user according to the purchase commodity data of the target user, so as to obtain preference commodity features of the target user after the preference feature mining processing;
[0067] A commercial data mining unit, configured to perform first user data mining processing on the target user by using the preference commodity features of the target user to obtain first commercial mining data of the target user, where the first commercial mining data includes a plurality of first similar users having similar preference commodity features to the target user and first user data corresponding to each first similar user;
[0068] The commercial data mining unit is configured to perform second user data mining processing on the target user based on the historical behavior data and the personal information to obtain second commercial mining data of the target user, where the second commercial mining data includes a plurality of second similar users having similar personal information and behavior data to the target user and second user data corresponding to each second similar user;
[0069] The commercial data mining unit is further configured to generate optimal commercial mining data of the target user according to the first commercial mining data and the second commercial mining data, so as to perform personalized commodity recommendation for the target user by using the optimal commercial mining data.
[0070] In a third aspect, another commercial data mining device is provided. Taking the device as an electronic device as an example, it includes a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the commercial data mining method as described in the first aspect or any possible design in the first aspect.
[0071] In a fourth aspect, a storage medium is provided, on which instructions are stored. When the instructions run on a computer, the commercial data mining method as described in the first aspect or any possible design in the first aspect is executed.
[0072] Fifth aspect, there is provided a computer program product including instructions, which when running on a computer, cause the computer to execute the commercial data mining method as described in the first aspect or any possible design in the first aspect.
[0073] Advantageous effects:
[0074] (1) When performing user data mining, the present invention is based on the attributes of the products themselves to mine and process user preference characteristics, thereby obtaining the preferred product characteristics of the users; then, based on the preferred product characteristics, the first user data mining process is performed to obtain the first commercial mining data related to the user preference characteristics; then, combined with the personal information and historical behavior data of the users themselves, the second user data mining process is performed to obtain the second commercial mining data related to the user personal information and historical behavior; then, based on the foregoing first commercial mining data and second commercial mining data, the optimal commercial mining data of the users can be obtained; finally, based on the optimal commercial mining data, the personalized recommendation service for the users can be completed; thus, the present invention can effectively mine user data based on the products purchased by the users, personal information, and historical behavior, without relying on explicit feedback data such as user ratings. Therefore, compared with the traditional technology, the effectiveness of data mining can be greatly improved, and thus effective personalized recommendation of products can be achieved. Therefore, it is very suitable for large-scale application and promotion. Description of the drawings
[0075] Figure 1 It is a schematic flowchart of the steps of the commercial data mining method provided by the embodiment of the present invention;
[0076] Figure 2 It is a schematic structural diagram of the commercial data mining device provided by the embodiment of the present invention;
[0077] Figure 3 It is a schematic structural diagram of the electronic device provided by the embodiment of the present invention. Detailed implementation manners
[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the drawings and the descriptions of the embodiments or the prior art. Obviously, the following descriptions of the structures of the drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. It should be noted here that the descriptions of these embodiments are used to help understand the present invention, but do not constitute a limitation to the present invention.
[0079] It should be understood that although terms such as first and second may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, the first unit may be referred to as the second unit, and similarly, the second unit may be referred to as the first unit, without departing from the scope of the exemplary embodiments of the present invention.
[0080] It should be understood that for the term "and / or" that may appear herein, it is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and both A and B exist simultaneously; for the term " / and" that may appear herein, it is a description of another association object relationship, indicating that two relationships may exist. For example, A / and B may represent: A exists alone, and both A and B exist; in addition, for the character " / " that may appear herein, generally it represents that the associated objects before and after are in an "or" relationship.
[0081] Embodiment:
[0082] See Figure 1 As shown, the commercial data mining method provided in this embodiment mines and processes user data based on the attributes of the product itself (such as color, brand, price, type, etc.), user personal information (such as age, education level, gender, etc.), and the historical behavior data of the user (such as collection, purchase, browsing, etc.), so as to obtain user data with similar preferences, similar personal information, and behaviors; then, based on this, personalized recommendation services for users can be carried out; in this way, compared with traditional technologies, when performing data mining, this method does not need to rely on explicit feedback data such as user ratings. Therefore, the problem of missing explicit feedback data can be solved, thereby greatly improving the effectiveness of data mining, and then achieving a more effective product personalized recommendation effect. Therefore, it is very suitable for large-scale application and promotion; among them, for example, this method can but is not limited to running on the data mining side. Optionally, the data mining side can but is not limited to a personal computer (PC), a tablet computer, or a smart phone. It can be understood that the foregoing execution subject does not constitute a limitation to the embodiments of the present application. Correspondingly, the running steps of this method can but are not limited to the following steps S1 to S6.
[0083] S1. Obtain the personal information of the target user and the historical behavior data of the target user on the e-commerce platform; in this embodiment, for example, the personal information may include, but is not limited to, basic information such as the age, gender, and education level of the target user, and the historical behavior data is various behaviors of the target user on the e-commerce platform, such as browsing behavior, product collection behavior, purchase behavior, adding to the shopping cart behavior, and so on; among them, the foregoing data can be obtained from the operation log of the e-commerce platform, and the data of the previous month can be obtained to obtain the foregoing historical behavior data.
[0084] After obtaining the personal information and historical behavior data of the target user, user data mining can be carried out based on this; among them, this embodiment conducts data mining processing from two aspects. On the one hand, based on the attributes of the product itself, the preferred product characteristics of the user are mined, and based on this, data mining is carried out; on the other hand, data mining is carried out based on the historical behavior data and personal information; finally, the optimal commercial mining data of the target user is generated by integrating the mining data of the foregoing two aspects; optionally, the mining process of the user's preferred product characteristics can be, but is not limited to, as shown in the following steps S2 and S3.
[0085] S2. Based on the historical behavior data, determine the purchased product data of the target user; in this embodiment, it has been described above that the historical behavior data includes purchase behavior. Therefore, the purchase behavior data can be screened out from the historical behavior data, and then based on the purchase behavior data, the purchased product data of the target user can be obtained; then, based on the purchased product data, the mining of the user's preference characteristics can be carried out, and the process can be, but is not limited to, as shown in the following step S3.
[0086] S3. According to the purchased product data of the target user, perform preference characteristic mining processing on the target user to obtain the preferred product characteristics of the target user after the preference characteristic mining processing; in specific implementation, for example, but not limited to, the following steps S31 to S35 can be used to mine the preferred product characteristics of the user.
[0087] S31. Based on the purchased product data of the target user, generate a product attribute set corresponding to the purchased product data; in this embodiment, each product has corresponding attributes, such as type, brand, price, color, and so on. Therefore, in this embodiment, all the attributes in the products purchased by the target user are counted to obtain the product attribute set.
[0088] After obtaining the product attribute set, the influence value of each product attribute can be determined based on this, so that the importance of each product attribute to the target user can be calculated based on the influence value of the product attribute; among them, the determination process of the influence value of the product attribute can be, but is not limited to, as shown in the following step S32.
[0089] S32. For the i-th product attribute in the set of product attributes, count the total number of products containing the i-th product attribute in the purchased product data as the attribute influence value corresponding to the i-th product attribute. In this embodiment, if the i-th product attribute is the red color, then count the total number of red products among all the products purchased. Similarly, if it is the green color, count the total number of green products among all the products purchased, and so on. For another example, if the i-th product attribute is the product category, which includes mobile phones, computers, etc., then count the total number of mobile phones and computers among all the products purchased. Of course, the determination process of the attribute influence values of the remaining product attributes is the same as the foregoing examples and will not be elaborated herein.
[0090] After obtaining the attribute influence value corresponding to the i-th product attribute, based on this, the importance of the i-th product attribute to the target user can be calculated. The calculation process can be but is not limited to the steps shown in the following S33.
[0091] S33. According to the attribute influence value corresponding to the i-th product attribute and in accordance with the following formula (1), calculate the importance of the i-th product attribute to the target user.
[0092]
[0093] In the above formula (1), P i represents the importance of the i-th product attribute to the target user, Y i represents the attribute influence value of the i-th product attribute, Y m represents the attribute influence value of the m-th product attribute, and n represents the total number of product attributes in the set of product attributes.
[0094] It can be seen from the above formula (1) that the importance of the i-th product attribute to the target user is essentially the ratio of a certain attribute in the customer's purchased products to the total attributes included in all the purchased products. Therefore, the larger the proportion, the more important the attribute is to the customer, and the more the customer attaches importance to the attribute.
[0095] Thus, after calculating the importance of the i-th product attribute to the target user through the foregoing formula (1), the importance of the remaining product attributes to the target user can be calculated in the same way. Among them, the loop calculation process can be but is not limited to the steps shown in the following S34.
[0096] S34. Increment i by 1 and re-count the number of products containing the i-th product attribute in the purchased product data until i is equal to n, and obtain the importance of each product attribute to the target user, where the initial value of i is 1.
[0097] After calculating the importance of each product attribute in the product attribute set for the target user based on the foregoing steps S31 to S34, the preferred product features of the target user can be determined based on the importance; among them, the process of determining the preferred product features can be but is not limited to the following steps shown in S35.
[0098] S35. According to the importance of each product attribute to the target user, several product attributes are screened out from the product attribute set as the preferred product features of the target user; in this embodiment, for example, but not limited to, the product attributes are sorted in descending order of importance, and then the top H are taken as the preferred product features of the target user. In this way, the number of preferred product features of the target user is multiple; in addition, the value of H can be specifically set according to actual use and is not specifically limited here.
[0099] In this way, through the foregoing steps S31 to S35, the preferred product features of the user can be mined based on the attributes of the product itself; then, based on this, the first user data mining of the target user can be performed; among them, the first user data mining process can be but is not limited to the following steps shown in S4.
[0100] S4. Using the preferred product features of the target user, the first user data mining process is performed on the target user to obtain the first commercial mining data of the target user, where the first commercial mining data includes several first similar users with similar preferred product features to the target user and the first user data corresponding to each first similar user; in specific applications, for example, but not limited to, the following steps S41 to S45 are used to complete the first user data mining process of the target user.
[0101] S41. From the user database of the e-commerce platform, several initial similar users corresponding to the target user are screened out; in this embodiment, traditional similar user recommendation methods (such as collaborative filtering, association rules, etc.) can be used to obtain similar users of the target user as the initial similar users; of course, the foregoing traditional similar user recommendation is a common technology for content recommendation, and its principle will not be elaborated.
[0102] After determining several initial similar users corresponding to the target user, the feature weights of each preferred product feature can be calculated based on this, so as to calculate the similarity between users based on this in the follow-up, thereby completing the mining process of the first user data; among them, the process of determining the feature weights of each preferred product feature can be but is not limited to the following steps shown in S42.
[0103] S42. Determine the feature weights of each preferred product feature in the preferred product features of the target user based on the several initial similar users. In this embodiment, a genetic algorithm is used to optimize and obtain the weights of each preferred product feature. The optimization process can be but is not limited to the steps shown in S42a - S42i below.
[0104] S42a. Select any product from the historical purchased products corresponding to each initial similar user as a test product. In this embodiment, the selected test product is used to calculate the probability that the target user purchases the test product under different feature weights, and the error between the purchase probability and whether the target user actually purchases the product is used as the fitness function for optimization. The specific calculation process of the fitness function is described below.
[0105] After determining the test product, the feature weights can be initialized to form several particle individuals using the initialized feature weights. The initialization process can be but is not limited to the steps shown in S42b below.
[0106] S42b. Perform multiple initialization processes on the feature weights of each preferred product feature. After each initialization process, use the initialized feature weights of each preferred product feature to form an initial weight vector. In this embodiment, it is equivalent to generating an initial weight vector in one initialization process. The initial weight vector contains the feature weights of each preferred product feature during this initialization. In this way, after multiple initialization processes, multiple initial weight vectors can be obtained. Then, based on this, particle individuals can be generated, and the process can be but is not limited to the steps shown in S42c below.
[0107] S42c. Use the initial weight vector obtained from each initialization process to generate multiple particle individuals. Each particle individual corresponds to an initial position vector and an initial velocity vector, and the initial position vector of any particle individual corresponds to an initial weight vector. In this embodiment, the number of particle individuals is equal to the number of initial weight vectors, and the initial position vector of the particle individual is actually the initial weight vector. In this way, during the optimization process, continuously updating the position of the particle individual is actually updating the initial weight vector corresponding to the particle individual. Based on this, when the iteration stop condition is met, the optimal weight vector can be obtained, and the weights in the optimal weight vector are the feature weights of each preferred product feature.
[0108] Optionally, the optimization process of the weight vector can be but is not limited to the steps shown in S42d - S42i below.
[0109] S42d. Initialize the iteration count t to 1, and obtain the position vector and velocity vector of each particle individual at the t-th iteration. Among them, when t is 1, the position vector and velocity vector of any particle individual at the t-th iteration are the initial position vector and initial velocity vector of that particle individual. In this embodiment, after obtaining the position vector of each particle individual at the t-th iteration, based on this, the fitness of each particle individual can be calculated, so as to update the global optimal position and individual optimal position based on the fitness subsequently. Among them, the calculation process of the fitness can be but is not limited to the steps shown in the following step S42e.
[0110] S42e. Calculate the fitness of each particle individual at the t-th iteration according to a number of initial similar users and the position vector of each particle individual at the t-th iteration. Among them, the fitness of any particle individual at the t-th iteration is used to represent the error of the purchase probability of the target user for the test commodity calculated using the position vector of that particle individual at the t-th iteration. In specific implementation, taking any particle individual as an example for specific elaboration. Optionally, the calculation process of its fitness can be but is not limited to the steps shown in the following steps S42e1 to S42e4.
[0111] S42e1. For any particle individual, determine the weight vector of that particle individual at the t-th iteration according to the position vector of that particle individual at the t-th iteration. In this embodiment, it has been described above that the initial position vector of each particle individual corresponds to an initial weight vector. Therefore, the position vector at the t-th iteration is also the weight vector at the t-th iteration. Based on this, the position vector of any particle individual at the t-th iteration is its corresponding weight vector, and each element inside it is the feature weight of each preferred commodity feature at the t-th iteration.
[0112] In this way, after obtaining the weight vector of the above-mentioned any particle individual at the t-th iteration, the similarity can be calculated, so as to calculate the probability of the target user purchasing the test commodity under this weight vector based on the true similarity between the target user and each initial similar user subsequently. Among them, the calculation process of the true similarity can be but is not limited to the steps shown in the following step S42e2.
[0113] S42e2. According to each preferred commodity feature of the target user and the weight vector of any particle individual at the t-th iteration, and according to the following formula (2), calculate the true similarity between the target user and each initial similar user.
[0114]
[0115] In the above formula (2), S(K,a) represents the true similarity between the target user K and the a-th initial similar user among a number of initial similar users. is the feature weight of the j-th preferred commodity feature in the weight vector of any particle individual at the t-th iteration. represents the importance of the j-th preferred commodity feature to the target user. represents the importance of the j-th preferred commodity feature to the a-th initial similar user. R represents the total number of preferred commodity features. Here, a = 1, 2,..., A, and A represents the total number of initial similar users. In this embodiment, the calculation method of the importance of the i-th preferred commodity feature to the a-th initial similar user can be referred to the foregoing step S33, and its principle will not be elaborated.
[0116] Thus, after calculating the true similarity between the target user and each initial similar user based on the foregoing formula (2), the probability that the target user purchases the test commodity when the feature weights are based on the weight vector corresponding to any particle individual can be calculated according to each true similarity. The calculation process of the purchase probability is as shown in the following step S42e3.
[0117] S42e3. Based on the true similarity between the target user and each initial similar user, and according to the following formula (3), calculate the purchase probability of the target user for the test commodity when the feature weights are based on the weight vector corresponding to any particle individual.
[0118]
[0119] In the above formula (3), G represents the purchase probability of the target user for the test commodity, and U(a) represents whether the a-th initial similar user has purchased the test commodity. Here, if the a-th initial similar user has purchased the test commodity, U(a) is 1, otherwise, U(a) is -1.
[0120] Therefore, based on the foregoing formula (3), the purchase probability of the target user for the test commodity can be calculated when the weight vector of any particle individual at the t-th iteration is used as the feature weights of the preferred commodity features. Then, based on this purchase probability, the fitness of any particle individual at the t-th time can be calculated. In this embodiment, the error between the predicted purchase probability and the result of the target user's actual purchase of the test commodity is used as the fitness function. Optionally, its calculation process can be but is not limited to as shown in the following step S42e4.
[0121] S42e4. According to the purchase probability of the target user for the test commodity, and according to the following formula (4), calculate the fitness of any particle individual at the t-th iteration.
[0122] D t = |G - U(K)| (4)
[0123] In the above formula (4), D t represents the fitness of any one of the particle individuals at the t-th iteration, and U(K) represents whether the target user has purchased the test product. Among them, if the target user has purchased the test product, then U(K) is 1; otherwise, U(K) is -1.
[0124] It can be seen from the foregoing formula (4) that G represents the purchase probability of the test product when the target user takes the weight vector of any one of the particle individuals at the t-th iteration as the characteristic weight of the preferred product features, while U(K) represents the representation value of the true purchase result of the target user for the test product; therefore, the smaller the absolute value of the error between the two, the higher the accuracy of the characteristic weight. Based on this, the purpose of optimization is to find the position vector with a fitness less than the fitness threshold (i.e., the error threshold), so as to obtain its optimal characteristic weight.
[0125] Thus, through the foregoing steps S42e1 to S42e4, the fitness of each particle individual at the t-th iteration can be calculated. Then, based on each fitness, the global optimal position and the individual optimal positions of each particle individual can be updated. The process can be but is not limited to the following steps shown in S42f.
[0126] S42f. Based on the fitness of each particle individual at the t-th iteration, update the global optimal position and the individual optimal positions of each particle individual. Among them, the global optimal position is the position vector corresponding to the particle individual with the smallest fitness during the 1st to t-th iterations; in this embodiment, if the smallest fitness at the t-th iteration is less than the fitness of the particle individual corresponding to the global optimal position at the (t - 1)-th iteration, then the position vector of the particle individual corresponding to the smallest fitness at the t-th iteration is used as the global optimal position; otherwise, no update is performed. Similarly, the individual optimal position of a particle individual is carried out in units of the particle individual, that is, the individual optimal position of each particle individual is the position vector corresponding to the smallest fitness during the first to t-th iterations; of course, the same is true for the global optimal position, which will not be elaborated here.
[0127] After completing the update of the global optimal position and the individual optimal positions of each particle individual, it can be determined whether to stop the iteration; among them, the determination process can be but is not limited to the following steps shown in S42g.
[0128] S42g. Determine whether the iteration stop condition is satisfied. The iteration stop condition is that the fitness of the particle individual corresponding to the updated global optimal position is less than the fitness threshold, or t reaches the maximum number of iterations. In this embodiment, if the iteration stop condition is satisfied, then each element in the weight vector corresponding to the global optimal position at this time is used as the feature weight of each optimal preference commodity feature, that is, based on the global optimal position when the iteration stop condition is satisfied, the feature weights of each preference commodity feature are determined. Otherwise, the iteration needs to continue, that is, the positions and velocities of each particle individual are updated. The position and velocity update process can be but is not limited to the steps shown in S42h below.
[0129] S42h. If not, then use the updated global optimal position and the updated individual optimal positions of each particle individual to update the velocity vector of each particle individual at the t-th iteration, obtaining the updated velocity vector corresponding to each particle individual, and use the updated velocity vector corresponding to each particle individual to update the position vector of each particle individual, obtaining the updated position vector corresponding to each particle individual. In specific applications, in this embodiment, the position update formula and velocity update formula of the traditional PSO (Particle Swarm Optimization) algorithm are used to update the position vector and velocity vector of each particle. Of course, the aforementioned PSO algorithm's position update and velocity update formulas are common formulas for optimization algorithms, and their principles will not be elaborated.
[0130] After the positions and velocities of each particle individual are updated, the updated positions and velocities can be used as the initial positions and velocities for the next iteration. Thus, the above process is continuously repeated until the iteration condition is satisfied, and then the feature weights of each preference commodity feature can be obtained. The cyclic optimization process can be but is not limited to the steps shown in S42i below.
[0131] S42i. Increment t by 1, and replace the position vector and velocity vector of each particle individual at the t-th iteration with the corresponding updated position vector and updated velocity vector, and recalculate the fitness of each particle individual at the t-th iteration based on a number of initial similar users and the position vector of each particle individual at the t-th iteration until the iteration stop condition is satisfied, so as to determine the feature weights of each preference commodity feature based on the global optimal position that satisfies the iteration stop condition. In this embodiment, the global optimal position when the iteration stop condition is satisfied is used as the optimal weight vector, and based on the optimal weight vector, the feature weights of each preference commodity feature can be obtained.
[0132] Thus, based on the foregoing steps S42a - S42i, the feature weights of each preferred product feature can be obtained through optimization; then, based on the feature weights of each preferred product feature, the user similarity can be recalculated to obtain multiple first similar users whose preferred product features are similar to those of the target user; among them, the calculation process of user similarity can be but is not limited to the following steps shown in S43.
[0133] S43. Use multiple preferred product features of the target user and the feature weights of each preferred product feature to calculate the first similarity between the target user and each user in the user database; in this embodiment, the foregoing formula (2) can be used to calculate the first similarity between the target user and each user in the user database, and its calculation principle will not be elaborated; thus, after obtaining the first similarity between the target user and each user, the first similar users can be determined based on this; among them, the determination process of the first similar users can be but is not limited to the following steps shown in S44.
[0134] S44. According to the first similarity between the target user and each user, determine several first similar users corresponding to the target user from the user database; in this embodiment, the users are sorted in descending order according to the first similarity, and then the top 5 or 10 users in the ranking are used as the first similar users; then, the user data (i.e., their historical purchase data) of the first similar users can be obtained to form the first commercial mining data; among them, the generation process of the first commercial mining data can be but is not limited to the following steps shown in S45.
[0135] S45. Obtain the first user data of each first similar user, and use each first similar user and the first user data of each first similar user to form the first commercial mining data.
[0136] Thus, through the foregoing steps S41 - S45, the first data mining process based on the user's preferred product features can be completed; then, the second data mining process can be performed according to the historical behavior data and personal information, and its process can be but is not limited to the following steps shown in S5.
[0137] S5. Based on the historical behavior data and the personal information, perform a second user data mining process on the target user to obtain the second commercial mining data of the target user, where the second commercial mining data includes several second similar users who have similar personal information and behavior data to the target user and the second user data corresponding to each second similar user; in specific implementation, for example, the following steps S51 - S56 can be but are not limited to be used to generate the second commercial mining data.
[0138] S51. Obtain the personal information and historical behavior data of each user in the user database of the e-commerce platform; in this embodiment, the process of obtaining the personal information and historical behavior data of each user in the user database is the same as that of the aforementioned target user, and will not be elaborated here.
[0139] After obtaining the personal information and historical behavior data of each user in the user database, the second data mining process can be carried out from two aspects of personal information and behavior data; among them, the data mining process based on personal information can be but is not limited to the following steps shown in S52.
[0140] S52. According to the personal information of the target user and the personal information of each user, calculate the personal information similarity between the target user and each user in the user database; in this embodiment, for the target user and any user, if they have the same personal attributes (such as male gender, or both have a bachelor's degree or above), then the similarity of this personal attribute is 1, otherwise it is 0; at the same time, for personal attributes such as age, they are divided according to a certain range and compared in the same way. If they belong to the same age range, then the similarity of the personal attribute of age between the two is 1; in this way, by obtaining the weights of different personal attributes in the personal information, and then multiplying the weights of each personal attribute by the similarity of the personal attribute and summing them up, the personal information similarity between the target user and the any user can be obtained; based on this, in the same way as above, the personal information similarity between the target user and each user can be obtained.
[0141] After completing the calculation of the similarity based on personal information, the calculation of behavior similarity can be carried out based on historical behavior data, and its calculation process can be but is not limited to the following steps shown in S53.
[0142] S53. Based on the historical behavior data of the target user and the historical behavior data of each user, calculate the behavior similarity between the target user and each user in the user database; in specific implementation, taking any user as an example, to specifically elaborate, the calculation process of its behavior similarity with the target user can be but is not limited to the following steps S53a to S53c.
[0143] S53a. Based on the historical behavior data of the target user, determine the set of behavior types of the target user; in this embodiment, an example is used to elaborate this step. Suppose there are browsing behavior, purchase behavior, collection behavior, and adding to cart behavior in the historical behavior data of the target user, then the set of behavior types includes browsing behavior, purchase behavior, collection behavior, and adding to cart behavior; in this way, based on this set of behavior types, the calculation of behavior similarity can be carried out, and its process can be but is not limited to the following steps S53b to S53c.
[0144] S53b. For any user in the user database, count the number of behaviors of each type in the set of behavior types generated by the any user on the e-commerce platform; in this embodiment, on the basis of the above example, it is elaborated that the number of browsing behaviors, the number of favorite behaviors, the number of adding-to-cart behaviors, and the number of purchase behaviors of the any user on the e-commerce platform in the same time period (such as within the past month) are counted; thus, after the counting is completed, the behavior similarity between the target user and the any user can be calculated according to the number of behaviors of each type counted, and the calculation process is as shown in the following step S53c.
[0145] S53c. Based on the number of behaviors of each type in the set of behavior types generated by the any user on the e-commerce platform, and according to the following formula (5), calculate the behavior similarity between the target user and the any user.
[0146]
[0147] In the above formula (5), Q(K,b) represents the behavior similarity between the target user and the any user. represents the number of behaviors of the target user generating the x-th type of behavior on the e-commerce platform. represents the number of behaviors of the any user generating the x-th type of behavior on the e-commerce platform, and X represents the total number of types of behaviors in the set of behavior types.
[0148] Thus, through the foregoing steps S53a to S53c, the behavior similarity between the target user and the any user can be calculated. Then, in the same way as above, the behavior similarity between the target user and each of the remaining users can be calculated; then, the second similarity between the target user and each user in the user database can be obtained by combining the personal information similarity; among them, the calculation process of the second similarity can be but is not limited to the following step S54.
[0149] S54. Calculate the second similarity between the target user and each user in the user database based on the personal information similarity and behavior similarity between the target user and each user in the user database. In this embodiment, any user is still taken as an example for illustration. For example, the total number of various behaviors in the historical behavior dataset of the target user can be calculated to obtain the first behavior number, and the total number of various behaviors in the historical behavior dataset of any user can be calculated to obtain the second behavior number. Then, sum the first behavior number and the second behavior number and take the average value to calculate the behavior control factor based on the average value. Finally, according to this behavior control factor, calculate the second similarity between the target user and any user based on the personal information similarity and behavior similarity between the target user and any user.
[0150] Optionally, for example, but not limited to, the following formula (6) can be used to calculate the second similarity between the target user and any user.
[0151]
[0152] In the above formula (6), Q′(K, b) represents the second similarity between the target user and any user, and s′(K, b) represents the personal information similarity between the target user and any user. represents the behavior control factor, and ζ represents the aforementioned average value.
[0153] In this way, after calculating the second similarity between the target user and each user through the aforementioned formula (6), several second similar users corresponding to the target user can be determined based on the second similarity. The determination process can be, but not limited to, as shown in the following step S55.
[0154] S55. Determine several second similar users corresponding to the target user from the user database according to the second similarity between the target user and each user. In this embodiment, the users are also sorted in descending order according to the second similarity, and then the top 5 or 10 users are selected as the second similar users. Then, the user data of the second similar users can be obtained, and thus the second commercial mining data can be formed. The process can be, but not limited to, as shown in the following step S56.
[0155] S56. Obtain the second user data of each second similar user, and use each second similar user and the second user data of each second similar user to form the second commercial mining data.
[0156] In this way, through the foregoing steps S51 to S56, the second user data mining process based on personal information and historical behavior data can be completed. Then, the optimal commercial mining data of the target user can be generated by combining the foregoing first commercial mining data. The process can be but is not limited to the following steps shown in S6.
[0157] S6. Generate the optimal commercial mining data of the target user based on the first commercial mining data and the second commercial mining data, so as to use the optimal commercial mining data for personalized product recommendation for the target user; in this embodiment, taking the union of the first commercial mining data and the second commercial mining data can obtain the optimal commercial mining data; among them, if there are the same users among the first similar users and the second similar users, only one is taken, and the similarity between the similar user and the target user in the optimal commercial mining data is the average of the corresponding first similarity and the second similarity; thus, after obtaining the optimal commercial mining data, personalized product recommendation for the target user can be performed based on this.
[0158] In this embodiment, one of the recommended methods is disclosed below, and the process is as shown below:
[0159] The first step: According to the product category currently browsed by the target user and based on the product features preferred by the user, screen out the products to be recommended from the products of the same category as this product category on the e-commerce platform; in this embodiment, the products that are the same as the product category browsed by the user and whose product attributes conform to the product features preferred by the user are used as the products to be recommended.
[0160] The second step: Calculate the recommendation scores of each product to be recommended according to the optimal commercial mining data; in this embodiment, for example, but not limited to, the following formula (7) can be used to calculate the recommendation scores of each product to be recommended.
[0161]
[0162] In the above formula (7), T(K,δ) represents the recommendation score of the δth product to be recommended for the target user, represents the similarity (the first similarity, the second similarity or the average of the two) between the rth similar user (the first similar user or the second similar user) in the optimal commercial mining data and the target user, pc(r,j) represents whether the rth similar user has purchased the δth product to be recommended, where if so, pc(r,j) takes the value of 1, otherwise, it takes the value of 0, and δ = 1, 2,..., μ, and μ represents the total number of products to be recommended.
[0163] Thus, through the foregoing formula (7), the recommendation scores of each product to be recommended for the target user can be calculated. Then, based on the foregoing recommendation scores, personalized product recommendations can be made, and the process is as shown in the third step below.
[0164] Third step: Sort the products to be recommended in descending order according to the recommendation scores to obtain a sorted sequence, and use the top 5 or 10 products to be recommended in the sorted sequence as the recommended products for the target user.
[0165] Based on the foregoing steps, personalized product recommendations for the target user can be completed.
[0166] Thus, through the commercial data mining method described in detail in the foregoing steps S1 to S6, the present invention mines and processes user data based on the attributes of the products themselves, the personal information of the users, and the historical behavior data of the users, so as to obtain user data with similar preferences, similar personal information, and behaviors; then, personalized recommendation services for users can be based on this; thus, compared with the traditional technology, when the present invention performs data mining, it does not need to rely on explicit feedback data such as user ratings. Therefore, the problem of missing explicit feedback data can be solved, thereby greatly improving the effectiveness of data mining, and further achieving a more effective personalized product recommendation effect. Therefore, it is very suitable for large-scale application and promotion.
[0167] As Figure 2 shown, in the second aspect of this embodiment, a hardware device for implementing the commercial data mining method described in the first aspect of the embodiment is provided, including:
[0168] An acquisition unit, configured to acquire the personal information of the target user and the historical behavior data of the target user on the e-commerce platform.
[0169] The acquisition unit is further configured to determine the purchased product data of the target user based on the historical behavior data.
[0170] A product feature mining unit, configured to perform preference feature mining processing on the target user according to the purchased product data of the target user, so as to obtain the preference product features of the target user after the preference feature mining processing.
[0171] A commercial data mining unit, configured to perform first user data mining processing on the target user by using the preference product features of the target user to obtain the first commercial mining data of the target user, where the first commercial mining data includes a plurality of first similar users having similar preference product features to the target user and the first user data corresponding to each first similar user.
[0172] A business data mining unit, configured to perform a second user data mining process on the target user based on the historical behavior data and the personal information, so as to obtain second business mining data of the target user, where the second business mining data includes a plurality of second similar users having similar personal information and behavior data to the target user, and second user data corresponding to each second similar user.
[0173] The business data mining unit is further configured to generate optimal business mining data of the target user according to the first business mining data and the second business mining data, so as to perform personalized commodity recommendation for the target user by using the optimal business mining data.
[0174] For the working process, working details and technical effects of the device provided in this embodiment, reference may be made to the first aspect of the embodiment, which will not be elaborated herein.
[0175] As Figure 3 shown, in the third aspect of this embodiment, another business data mining device is provided. Taking the device as an electronic device as an example, it includes: a memory, a processor, and a transceiver that are communicatively connected in sequence, where the memory is configured to store a computer program, the transceiver is configured to send and receive messages, and the processor is configured to read the computer program and execute the business data mining method described in the first aspect of the embodiment.
[0176] Specifically, the memory may include, but is not limited to, a random access memory (RAM), a read only memory (ROM), a flash memory, a first input first output (FIFO) memory, and / or a first in last out (FILO) memory, etc.; specifically, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor may be implemented in at least one hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). At the same time, the processor may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state.
[0177] In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. For example, the processor may not be limited to using a microprocessor of the STM32F105 series, a reduced instruction set computer (RISC) microprocessor, an X86 architecture processor, or a processor integrated with an embedded neural-network processing unit (NPU); the transceiver may be, but is not limited to, a Wi-Fi wireless transceiver, a Bluetooth wireless transceiver, a General Packet Radio Service (GPRS) wireless transceiver, a ZigBee (low-power local area network protocol based on the IEEE 802.15.4 standard) wireless transceiver, a 3G transceiver, a 4G transceiver, and / or a 5G transceiver, etc. In addition, the device may also include, but is not limited to, a power module, a display screen, and other necessary components.
[0178] For the working process, working details, and technical effects of the electronic device provided in this embodiment, reference may be made to the first aspect of the embodiment, which will not be elaborated herein.
[0179] The fourth aspect of this embodiment provides a storage medium storing instructions including the commercial data mining method described in the first aspect of the embodiment, that is, instructions are stored on the storage medium, and when the instructions are run on a computer, the commercial data mining method described in the first aspect of the embodiment is executed.
[0180] Among them, the storage medium refers to a carrier for storing data, and may include, but is not limited to, a floppy disk, an optical disc, a hard disk, a flash memory, a USB flash drive, and / or a Memory Stick, etc. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices.
[0181] For the working process, working details, and technical effects of the storage medium provided in this embodiment, reference may be made to the first aspect of the embodiment, which will not be elaborated herein.
[0182] The fifth aspect of this embodiment provides a computer program product including instructions, and when the instructions are run on a computer, the computer is made to execute the commercial data mining method described in the first aspect of the embodiment, where the computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices.
[0183] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A commercial data mining method, characterized in that: include: Obtain the target user’s personal information and historical behavior data on the e-commerce platform; Based on the historical behavior data, determining the purchase commodity data of the target user; According to the purchase commodity data of the target user, a preference feature mining process is performed on the target user, so as to obtain the preferred commodity features of the target user after the preference feature mining process; Using the preferred commodity characteristics of the target user, first user data mining processing is performed on the target user to obtain first commercial mining data of the target user, wherein the first commercial mining data includes a plurality of first similar users having similar preferred commodity characteristics as the target user and first user data corresponding to each first similar user; Based on the historical behavior data and the personal information, performing second user data mining processing on the target user to obtain second business mining data of the target user, wherein the second business mining data includes a number of second similar users having similar personal information and behavior data to the target user and second user data corresponding to each second similar user; Generate optimal business mining data for the target user based on the first business mining data and the second business mining data, so as to make personalized product recommendations for the target user by using the optimal business mining data; The number of the target user's preferred product features is multiple, wherein the target user's preferred product features are used to perform first user data mining processing on the target user to obtain first commercial mining data of the target user, including: Filtering a number of initial similar users corresponding to the target user from a user database of the e-commerce platform; Determining, based on the plurality of initial similar users, a feature weight of each of the preferred product features of the target user; Calculate a first similarity between the target user and each user in the user database by using a plurality of preferred product features of the target user and a feature weight of each preferred product feature; According to the first similarity between the target user and each user, determining a plurality of first similar users corresponding to the target user from the user database; Acquire the first user data of each first similar user, and use each first similar user and the first user data of each first similar user to form the first business mining data; Determining the feature weight of each of the preferred product features of the target user according to the plurality of initial similar users includes: Determine any product from the historical purchase products corresponding to each initial similar user as a test product; Performing multiple initialization processing on the feature weights of each preferred product feature, so that after each initialization processing, an initial weight vector is formed using the initialized feature weights of each preferred product feature; Using the initial weight vector obtained in each initialization process, multiple individual particles are generated, wherein each individual particle corresponds to an initial position vector and an initial velocity vector, and the initial position vector of any individual particle corresponds to an initial weight vector; Initialize the number of iterations t to 1, and obtain the position vector and velocity vector of each particle individual at the t-th iteration, wherein when t is 1, the position vector and velocity vector of any particle individual at the t-th iteration are the initial position vector and initial velocity vector of any particle individual; According to a number of initial similar users and the position vectors of each particle individual at the t-th iteration, the fitness of each particle individual at the t-th iteration is calculated, wherein the fitness of any particle individual at the t-th iteration is used to characterize the error of the target user's purchase probability of the test product calculated using the position vector of any particle individual at the t-th iteration; Based on the fitness of each particle individual at the tth iteration, the global optimal position and the individual optimal position of each particle individual are updated, wherein the global optimal position is the position vector corresponding to the particle individual with the smallest fitness during the 1st to tth iterations; Determine whether an iteration stop condition is met, wherein the iteration stop condition is that the fitness of the particle individual corresponding to the updated global optimal position is less than the fitness threshold, or t reaches the maximum number of iterations; If not, then use the updated global optimal position and the updated individual optimal position of each individual particle to update the velocity vector of each individual particle at the tth iteration to obtain the updated velocity vector corresponding to each individual particle, and use the updated velocity vector corresponding to each individual particle to update the position vector of each individual particle to obtain the updated position vector corresponding to each individual particle; Add 1 to t, replace the position vector and velocity vector of each particle individual at the tth iteration with the corresponding updated position vector and updated velocity vector, and recalculate the fitness of each particle individual at the tth iteration based on several initial similar users and the position vector of each particle individual at the tth iteration until the iteration stop condition is met, so as to determine the feature weights of each preferred product feature according to the global optimal position that meets the iteration stop condition.
2. The method according to claim 1, characterized in that According to the purchase commodity data of the target user, a preference feature mining process is performed on the target user, so as to obtain the preferred commodity features of the target user after the preference feature mining process, including: Based on the purchase commodity data of the target user, generating a commodity attribute set corresponding to the purchase commodity data; For the i-th commodity attribute in the commodity attribute set, the total number of commodities in the purchased commodity data containing the i-th commodity attribute is counted as the attribute influence value corresponding to the i-th commodity attribute; According to the attribute influence value corresponding to the i-th commodity attribute and in accordance with the following formula (1), the importance of the i-th commodity attribute to the target user is calculated; (1) In the above formula (1), represents the importance of the i-th product attribute to the target user, represents the attribute impact value of the i-th commodity attribute, represents the attribute impact value of the mth product attribute, Indicates the total number of product attributes in the product attribute set; i is incremented by 1, and the number of commodities in the purchased commodity data that contain the i-th commodity attribute is recounted until i is equal to n, and the importance of each commodity attribute to the target user is obtained, wherein the initial value of i is 1; According to the importance of each commodity attribute to the target user, several commodity attributes are selected from the commodity attribute set to serve as the preferred commodity features of the target user.
3. The method according to claim 1, characterized in that According to several initial similar users and the position vectors of each particle individual at the tth iteration, the fitness of each particle individual at the tth iteration is calculated, including: For any individual particle, according to the position vector of the individual particle at the t-th iteration, determine the weight vector of the individual particle at the t-th iteration; According to the preferred product features of the target user and the weight vector of any particle individual at the tth iteration, the real similarity between the target user and each initial similar user is calculated according to the following formula (2); (2) In the above formula (2), Indicates the target user The first of several initial similar users The true similarity of the initial similar users, is the weight vector of any particle individual at the tth iteration. The feature weights of the preferred product features, represents the importance of the jth preferred product feature to the target user, Indicates the jth preferred product feature for the The importance of initial similar users, Represents the total number of preferred product features, where ,and represents the total number of initial similar users; Based on the real similarity between the target user and each initial similar user, and according to the following formula (3), the probability of the target user purchasing the test product is calculated when the weight vector corresponding to any particle individual is used as the feature weight of the preferred product feature; (3) In the above formula (3), represents the purchase probability of the target user for the test product, Indicates Whether the initial similar users purchased the test product, Initial similar users purchased the test product. is 1, otherwise, is -1; According to the purchase probability of the target user for the test product, the fitness of any particle individual at the tth iteration is calculated according to the following formula (4); (4) In the above formula (4), represents the fitness of any particle individual at the tth iteration, Indicates whether the target user has purchased the test product, wherein if the target user has purchased the test product, then is 1, otherwise, is -1.
4. The method according to claim 1, characterized in that Based on the historical behavior data and the personal information, a second user data mining process is performed on the target user to obtain second commercial mining data of the target user, including: Obtain personal information and historical behavior data of each user in the user database of the e-commerce platform; Calculate the similarity between the personal information of the target user and the personal information of each user in the user database according to the personal information of the target user and the personal information of each user; Based on the historical behavior data of the target user and the historical behavior data of each user, calculating the behavior similarity between the target user and each user in the user database; Calculate the second similarity between the target user and each user in the user database based on the similarity of personal information and behavior between the target user and each user in the user database; Determining a plurality of second similar users corresponding to the target user from the user database according to the second similarity between the target user and each user; The second user data of each second similar user is obtained, and the second business mining data is formed by using each second similar user and the second user data of each second similar user.
5. The method according to claim 4, characterized in that Based on the historical behavior data of the target user and the historical behavior data of each user, the behavior similarity between the target user and each user in the user database is calculated, including: Determining a behavior category set of the target user based on the historical behavior data of the target user; For any user in the user database, counting the number of behaviors of each type of behaviors in the behavior category set generated by the any user on the e-commerce platform; Based on the number of behaviors of each type of behaviors in the behavior type set generated by any user on the e-commerce platform, the behavior similarity between the target user and any user is calculated according to the following formula (5); (5) In the above formula (5), represents the behavior similarity between the target user and any of the users, Indicates that the target user generates the first The number of behaviors of the class behavior, Indicates that any user generates the first The number of behaviors of the class behavior, Indicates the total number of behavior categories in the behavior category set.
6. A commercial data mining device for executing the commercial data mining method according to any one of claims 1 to 5, characterized in that: include: An acquisition unit, used to acquire the target user's personal information and the target user's historical behavior data on the e-commerce platform; The acquisition unit is further used to determine the purchase commodity data of the target user based on the historical behavior data; A commodity feature mining unit, configured to perform a preference feature mining process on the target user according to the commodity purchase data of the target user, so as to obtain the preferred commodity features of the target user after the preference feature mining process; A commercial data mining unit, configured to perform first user data mining processing on the target user by using the preferred commodity characteristics of the target user to obtain first commercial mining data of the target user, wherein the first commercial mining data includes a plurality of first similar users having similar preferred commodity characteristics as the target user and first user data corresponding to each of the first similar users; A business data mining unit, configured to perform a second user data mining process on the target user based on the historical behavior data and the personal information to obtain second business mining data of the target user, wherein the second business mining data includes a plurality of second similar users having similar personal information and behavior data to the target user and second user data corresponding to each second similar user; The business data mining unit is further used to generate optimal business mining data for the target user based on the first business mining data and the second business mining data, so as to use the optimal business mining data to make personalized product recommendations for the target user.
7. An electronic device, characterized in that: include: A memory, a processor and a transceiver which are sequentially communicatively connected, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program to execute the commercial data mining method as described in any one of claims 1 to 5.
8. A computer program product comprising instructions, characterized in that When the instructions are executed on a computer, the computer is caused to execute the commercial data mining method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Web service recommendation method based on user preference feature modeling
CN103544623A
Preference-based game commodity recommending method and device
CN108261766A
Commodity recommendation method and system based on big data analysis
CN116402565A