A method, device and medium for processing user behavior characteristics
By obtaining the target item category set of the specified item category and its associated item categories, screening the user identification set associated with it, and calculating the similarity based on the user behavior feature vector, the user identification most similar to the specified user behavior feature vector is selected, which solves the problem of insufficient scale and accuracy of information push and achieves higher matching degree and push scale.
Patent Information
- Application Number
- CN202411969394.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-12-30
AI Technical Summary
In the existing technology, the scale and accuracy of information push related to specific item categories are insufficient, making it difficult to accurately select more users.
By obtaining the target item category set of the specified item category and its associated item categories, filtering the user identification set associated with it, and calculating the similarity based on the user behavior feature vector, the user identification most similar to the specified user behavior feature vector is selected, and the target user identification set is constructed for information push.
Improved the scale and accuracy of information push related to specific item categories, ensuring a high degree of match between push users and item categories.
Smart Images

Figure CN119863293B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a method, device and medium for processing user behavior characteristics. Background Art
[0002] User characteristics can be determined based on their behavior. For example, their consumption preferences can be determined based on their consumption behavior. Once the user's consumption preferences are determined, matching item information can be pushed to the user, that is, item information that matches the user's consumption preferences can be pushed to the user. A user's consumption behavior can be represented by the names of the items they purchased. Each item has its own item category. Based on their consumption behavior, users can be associated with item categories to form a user list that stores these associations. In scenarios where information is pushed for a specific item category, the number of users in the user list associated with that specific item category may be relatively small. Therefore, how to accurately select more users and improve the scale and accuracy of information push related to a specific item category is an urgent problem that needs to be solved. Summary of the Invention
[0003] The present invention aims to provide a method, device and medium for processing user behavior characteristics to accurately select more users and improve the scale and accuracy of information push related to specific item categories.
[0004] According to a first aspect of the present invention, a method for processing user behavior characteristics is provided, the method comprising the following steps:
[0005] A target item category set is obtained according to a specified item category and a list of associated item categories; the target item category set includes the specified item category and item categories associated with the specified item category, and the list of associated item categories includes a plurality of records, each record including an item category and item categories associated with the item category.
[0006] A designated user identification set is obtained based on a first user list and a target item category set; the first user list includes a plurality of records, each record including a user identification and an item category set corresponding to the user identification; the item category set corresponding to each user identification includes a plurality of item categories; the designated user identification set includes a plurality of designated user identifications, and the item category set corresponding to each designated user identification includes at least one item category in the target item category set.
[0007] A designated user behavior feature vector set corresponding to the designated user identification set is obtained; the designated user behavior feature vector set includes a designated user behavior feature vector corresponding to each designated user identification.
[0008] Obtain a preset first number of global user behavior feature vectors in the global user behavior feature vector set that have the greatest similarity to each designated user behavior feature vector, and append the user identifiers and similarity values corresponding to the global user behavior feature vectors to the target user identifier set; the designated user behavior feature vector set is a subset of the global user behavior feature vector set, and the product of the preset first number and the number of designated user identifiers is greater than a preset number of target users.
[0009] The user identification corresponding to the largest similarity value of the preset number of target users in the target user identification set is determined as the target user identification corresponding to the specified item category.
[0010] According to a second aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for processing user behavior characteristics when executing the computer program.
[0011] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for processing user behavior characteristics is implemented.
[0012] Compared with the prior art, the present invention has at least the following beneficial effects:
[0013] The present invention obtains other item categories associated with a designated item category based on the designated item category, thereby obtaining a target item category set, wherein any item category in the target item category set is relatively similar to the designated item category. The present invention also filters user identifiers associated with the target item category set from a first user list based on the target item category set, thereby obtaining a designated user identifier set, wherein any designated user in the designated user identifier set has a behavior related to at least one item category in the target item category set. The present invention further obtains a designated user behavior feature vector corresponding to each user identifier in the designated user identifier set, and the designated user behavior feature vector can be used to characterize the behavior characteristics of the corresponding user. By selecting the first first preset number of user behavior feature vectors that are most similar to each designated user behavior feature vector from the global user behavior feature vectors, the present invention obtains user behavior feature vectors that exceed a preset target number of users and are similar to the designated user behavior feature vectors in the designated user behavior feature vector set, and further obtains a preset target number of user identifiers that best match the item categories in the target item category set by comparing the overall similarities. Thus, the present invention not only achieves the purpose of obtaining the preset target number of user identifiers, but also has a high degree of match between the obtained preset target number of user identifiers and the designated item category, thereby improving the scale and accuracy of information push related to the designated item category. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0015] Figure 1 This is a flowchart of a method for processing user behavior characteristics provided in Example 1 of the present invention;
[0016] Figure 2 A flowchart of the steps of obtaining a designated user identification set based on a first user list and a target item category set provided in the first embodiment of the present invention;
[0017] Figure 3 Flowchart of the process of obtaining a global user behavior feature vector set provided in the first embodiment of the present invention;
[0018] Figure 4 A flowchart of the steps of obtaining a global user behavior feature vector corresponding to a user identifier included in a first record provided in the first embodiment of the present invention;
[0019] Figure 5 This is a flowchart of the steps of obtaining a first number of global user behavior feature vectors having the greatest similarity to each designated user behavior feature vector in a global user behavior feature vector set provided in the first embodiment of the present invention. DETAILED DESCRIPTION
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0021] Example 1:
[0022] According to this embodiment, Figure 1 As shown, a method for processing user behavior characteristics is provided, the method comprising the following steps:
[0023] S100, obtaining a target item category set based on a designated item category and a list of associated item categories; the target item category set includes the designated item category and item categories associated with the designated item category, the list of associated item categories includes a plurality of records, each record including an item category and item categories associated with the item category.
[0024] In this embodiment, the designated item category is the item category for which information is to be pushed.
[0025] In this embodiment, the associated item category list is a pre-established list. For example, the designated item category is chewing gum, and the item categories associated with it include bubble gum, QQ candy, chewable tablets, and mouth fresheners.
[0026] S200, obtaining a designated user identification set based on a first user list and a target item category set; the first user list includes a plurality of records, each record including a user identification and an item category set corresponding to the user identification; the item category set corresponding to each user identification includes a plurality of item categories; the designated user identification set includes a plurality of designated user identifications, and the item category set corresponding to each designated user identification includes at least one item category in the target item category set.
[0027] In this embodiment, the first user list is a known list and can be directly obtained through a third party. The third party can know the consumption behavior of a large number of users, that is, know the names and purchase times of items purchased by a large number of users.
[0028] In this embodiment, the user identifier is a user ID. Different records in the first user list contain different user identifiers. The item category set included in any record in the first user list is the set of item categories to which items purchased by the user identifier included in the record during a certain time period belong. It should be noted that the designated user identifier is also a user identifier, where the designated user identifier specifically refers to a user identifier in the first user list whose corresponding item category set includes at least one item category in the target item category set.
[0029] As a specific embodiment, Figure 2 As shown, obtaining the designated user identification set according to the first user list and the target item category set includes:
[0030] S210, traversing the first user list, if a record included in the first user list includes at least one item category in the target item category set, appending the user identifier included in the record to a preset user identifier set; the preset user identifier set is initialized to an empty set.
[0031] S220: After the traversal is completed, the preset user identification set is determined as the designated user identification set.
[0032] Based on S210 - S220 , the specified user identification set can be accurately obtained.
[0033] S300 , obtaining a designated user behavior feature vector set corresponding to a designated user identification set; the designated user behavior feature vector set includes a designated user behavior feature vector corresponding to each designated user identification.
[0034] In this embodiment, user behavior refers to a user's consumption behavior, and a user behavior feature vector is used to represent the corresponding user's consumption preferences (i.e., the user's habit of purchasing certain categories of goods). Two user behavior feature vectors with high similarity correspond to two users with similar consumption preferences, while two user behavior feature vectors with low similarity correspond to two users with significantly different consumption preferences. It should be noted that the designated user behavior feature vector in this embodiment is also a user behavior feature vector, where designated refers specifically to the user behavior feature vector corresponding to the designated user identifier.
[0035] In this embodiment, a global user behavior feature vector set is pre-established. Any global user behavior feature vector in this global user behavior feature vector set corresponds to a single user. The number of global user behavior feature vectors included in this global user behavior feature vector set is equal to the number of user identifiers appearing in all user lists. It should be noted that the global user behavior feature vector in this embodiment is also a user behavior feature vector, where "global" is used to represent all, and any user identifier has a corresponding global user behavior feature vector. It should be understood that the specified user behavior feature vector and the global user behavior feature vector for the same user identifier are the same user behavior feature vector.
[0036] In this embodiment, the designated user behavior feature vector set is a subset of the global user behavior feature vector set. It is known which user identifier each global user behavior feature vector in the global user behavior feature vector set corresponds to. When the designated user identifier set is known, the designated user behavior feature vector set corresponding to the designated user identifier set can be filtered out from the global user behavior feature vector set.
[0037] S400, obtaining a preset first number of global user behavior feature vectors in the global user behavior feature vector set that have the greatest similarity to each specified user behavior feature vector, and appending the user identifiers and similarity values corresponding to the global user behavior feature vectors to the target user identifier set; the specified user behavior feature vector set is a subset of the global user behavior feature vector set, and the product of the preset first number and the number of specified user identifiers is greater than a preset number of target users.
[0038] In this embodiment, the target user identification set is initialized to an empty set, and the preset first number and the preset number of target users are empirical values; preferably, the preset first number is the product of the ratio of the preset number of target users to the number of designated user identifications and a preset target multiple, where the preset target multiple is a multiple of 10. As a specific embodiment, the number of designated user identifications is 1 million, the preset number of target users is 10 million, and the preset first number is 100. Then, for each designated user behavior feature vector, the top 100 global user behavior feature vectors with the greatest similarity to it are selected from the global user behavior feature vector set. As a result, the target user identification set ultimately includes 100 million target identifications. In this embodiment, the product of the preset first number and the number of designated user identifications is greater than the preset number of target users, which is conducive to improving the matching degree between the target user identification and the designated item category in the final target user identification set.
[0039] In this embodiment, when the user identification and the similarity value are added to the target user identification set, any user identification and the corresponding similarity value are constructed into a tuple including two elements.
[0040] Those skilled in the art know that any method for obtaining the similarity between two vectors in the prior art falls within the protection scope of the present invention.
[0041] As a preferred embodiment, any record included in the first user list also includes an item name set, and the item name set includes several item names; Figure 3 As shown, the process of obtaining the global user behavior feature vector set includes:
[0042] S401, input each item name in the item name set included in the first record into the trained neural network model, and obtain a sub-feature vector corresponding to each item name in the item name set included in the first record; the first record is any record in the first user list; the trained neural network model is used to obtain the sub-feature vector corresponding to the input item name.
[0043] In this embodiment, the trained neural network model is a CoRom model, which is capable of inferring a sub-feature vector corresponding to an input item name. This sub-feature vector can represent information about the input item name and has 768 dimensions. Those skilled in the art will appreciate that any prior art method for training a neural network model falls within the scope of protection of the present invention.
[0044] It should be understood that if the number of item names in the item name set included in the first record is 20, then it needs to be input 20 times, and the number of sub-feature vectors obtained is also 20.
[0045] S402: Obtain the item category corresponding to each item name in the item name set included in the first record.
[0046] In this embodiment, there is a correspondence between the item category set and the item name set included in any record in the first user list, that is, the item category to which any item name in the item name set included in any record in the first user list belongs is an item category in the item category set included in the record; therefore, the item category corresponding to each item name in the item name set can be selected from the item category set included in the first record.
[0047] S403: Obtain the frequency of occurrence of the item category corresponding to each item name in the item name set included in the first record in the first user list.
[0048] In this embodiment, the same item category may exist in the item category sets included in different records in the first user list. The more frequently a certain item category appears in the first user list, the more likely it is that most users need to purchase items in this category, and this item category has relatively low user diversity. The less frequently a certain item category appears in the first user list, the more likely it is that only a small number of users need to purchase items in this category, and this item category has relatively high user diversity.
[0049] S404, obtaining a weight corresponding to each item name in the item name set included in the first record according to the frequency; the weight corresponding to any item name in the item name set included in the first record is negatively correlated with the frequency of the item name appearing in the first user list.
[0050] As a preferred embodiment, the total number of item names in the item name set included in the first record is n, and the frequency of the item category corresponding to the i-th item name in the item name set included in the first record in the first user list is f. i , then the weight corresponding to the i-th item name in the item name set included in the first record is w i , w i =(1-f i / ∑ n i=1 f i ) / (n-1). Thus, the greater the frequency of an item name in the item name set included in the first record appearing in the first user list, the smaller the weight corresponding to the item name.
[0051] S405 : Obtain a global user behavior feature vector corresponding to the user identifier included in the first record according to the weight and sub-feature vector corresponding to each item name in the item name set included in the first record.
[0052] As a preferred embodiment, a weighted summation formula is used to obtain the global user behavior feature vector corresponding to the user identifier included in the first record. If the sub-feature vector corresponding to the i-th item name in the item name set included in the first record is H i , then the global user behavior feature vector corresponding to the user identifier included in the first record is G, G = ∑ n i=1 (H i ×w i ). Thus, the global user behavior feature vector corresponding to the user identifier included in the first record obtained in this embodiment can more accurately represent the behavior feature corresponding to the user identifier.
[0053] As a more preferred embodiment, any record included in the first user list also includes time information; Figure 4 As shown, obtaining the global user behavior feature vector corresponding to the user identifier included in the first record according to the weight and sub-feature vector corresponding to each item name in the item name set included in the first record includes:
[0054] S4051: If there are other user lists other than the first user list, determine whether there is a record including a first user identifier in the other user list; the first user identifier is a user identifier included in the first record.
[0055] In this embodiment, if there are no other user lists except the first user list, S4052-S4053 are not executed. Instead, the weighted summation formula is used to obtain the global user behavior feature vector corresponding to the user identifier included in the first record. If the sub-feature vector corresponding to the i-th item name in the item name set included in the first record is H i , then the global user behavior feature vector corresponding to the user identifier included in the first record is G, G = ∑ n i=1 (H i ×w i ); Thus, the global user behavior feature vector corresponding to the user identifier included in the first record obtained in this embodiment can more accurately represent the behavior feature corresponding to the user identifier.
[0056] S4052. If the other user list includes a record of the first user identifier, a first behavior feature vector corresponding to the first user identifier is obtained based on the weight and sub-feature vector corresponding to each item name in the item name set included in the first record, and other behavior feature vectors corresponding to the first user identifier are obtained based on the other user list; the number of other behavior feature vectors is equal to the number of user lists in the other user list that include records of the first user identifier; any record included in the other user list includes time information.
[0057] In this embodiment, if there are user lists in the user table other than the first user list, it is possible that the user identifiers included in some records in the first user list are the same as the user identifiers included in some records in other user lists. In this case, not only the global user behavior feature vector corresponding to the user identifier must be determined based on the first user list, but also the global user behavior feature vector corresponding to the user identifier must be determined in combination with the other user lists.
[0058] In this embodiment, if there is a user table user list other than the first user list, the weighted sum formula is used to obtain the first behavior feature vector. If the sub-feature vector corresponding to the i-th item name in the item name set included in the first record is H i , then the first behavior feature vector corresponding to the user identifier included in the first record is P, P = ∑ n i=1 (H i ×w i ).
[0059] In this embodiment, the process of obtaining the behavior feature vector corresponding to the first user identifier based on any user list is similar to the process of obtaining the first behavior feature vector corresponding to the first user identifier described above, and will not be repeated here. It should be understood that if the number of user lists in which records corresponding to the first user identifier exist in other user lists is m, then the number of other behavior feature vectors obtained is also m. It should be noted that the process of obtaining the first behavior feature vector and other behavior feature vectors is similar, where the first and other are only used to distinguish whether the behavior feature vector is obtained based on the first user list or based on other lists other than the first user list.
[0060] In this embodiment, if there is no record corresponding to the first user identifier in the list of other users, the first behavior feature vector corresponding to the first user identifier is determined as the global user behavior feature vector corresponding to the first user identifier, and the step of obtaining other behavior feature vectors corresponding to the first user identifier and S4053 are no longer performed.
[0061] S4053, obtaining a global user behavior feature vector corresponding to the first user identifier based on the first behavior feature vector corresponding to the first user identifier, other behavior feature vectors, time information included in the first record, and time information included in the target record in the other user lists; the target record is a record including the first user identifier.
[0062] As a preferred specific implementation, the weight of the first behavior feature vector and the weights of other behavior feature vectors are determined based on the time information included in the first record and the time information included in the target record in other user lists, wherein the later the time corresponding to the included time information, the greater the weight of the behavior feature vector.
[0063] As a preferred embodiment, a weighted summation formula is used to obtain the global user behavior feature vector G corresponding to the first user identifier, G = P × c + ∑ z j=1 (D j ×b j ), c is the weight of the first line feature vector, b j is the weight of the jth other behavior feature vector, D j is the jth other behavior feature vector, the value range of j is 1 to z, z is the number of other user lists except the first user list, c+∑ z j=1 b j =1, c and b j are all greater than 0. It should be understood that D j It is obtained based on the jth user list except the first user list.
[0064] As a preferred embodiment, Figure 5 As shown, obtaining a first preset number of global user behavior feature vectors in the global user behavior feature vector set having the greatest similarity to each designated user behavior feature vector includes:
[0065] S410: Clustering the global user behavior feature vector set using a clustering algorithm to obtain a number of clusters.
[0066] Optionally, use the k-means clustering algorithm for clustering.
[0067] S420: Obtain the similarity between the first designated user behavior feature vector and the center vector of each cluster. If the similarity between the first designated user behavior feature vector and the center vector of a cluster is greater than or equal to a preset similarity threshold, determine the cluster as a cluster associated with the first designated user behavior feature vector.
[0068] In this embodiment, the center vector of any cluster is the average of the global user behavior feature vectors included in the cluster.
[0069] In this embodiment, the preset similarity threshold is an empirical value. If the similarity between two user behavior feature vectors is greater than or equal to the preset similarity threshold, it means that the two user behavior feature vectors are relatively similar and the corresponding user behavior features are relatively similar.
[0070] S430: Perform dimensionality reduction processing on the first designated user behavior feature vector to obtain a first dimensionality-reduced user behavior feature vector.
[0071] In this embodiment, the product quantization (PQ) method is used to achieve dimensionality reduction. Those skilled in the art will appreciate that the PQ method is a prior art process and will not be further described here. As a specific embodiment, the first designated user behavior feature vector is 768-dimensional, and the first reduced user behavior feature vector is 16-dimensional.
[0072] S440 : Perform dimensionality reduction processing on the global user behavior feature vectors included in all clusters associated with the first designated user behavior feature vector to obtain a set of reduced-dimensional user behavior feature vectors.
[0073] S450, obtaining a preset first number of reduced dimensionality user behavior feature vectors in the reduced dimensionality user behavior feature vector set that have the greatest similarity to the first reduced dimensionality user behavior feature vector, and determining the global user behavior feature vectors corresponding to the preset first number of reduced dimensionality user behavior feature vectors as the preset first number of global user behavior feature vectors that have the greatest similarity to the first specified user behavior feature vector.
[0074] This embodiment obtains a first preset number of global user behavior feature vectors with the greatest similarity to the first specified user behavior feature vector based on the first reduced-dimensionality user behavior feature vector and the reduced-dimensionality user behavior feature vector set, and can quickly obtain a first preset number of global user behavior feature vectors with the greatest similarity to the first specified user behavior feature vector.
[0075] S500: Determine the user identifier corresponding to the largest similarity value of the preset number of target users in the target user identifier set as the target user identifier corresponding to the designated item category.
[0076] In this embodiment, the target user identifier corresponding to the designated item category is the object to which information related to the designated item category is to be pushed.
[0077] This embodiment obtains other item categories associated with the designated item category based on the designated item category, thereby obtaining a target item category set, wherein any item category in the target item category set is relatively similar to the designated item category. This embodiment also filters user identifiers associated with the target item category set from the first user list based on the target item category set, thereby obtaining a designated user identifier set, wherein any designated user in the designated user identifier set has a behavior related to at least one item category in the target item category set. This embodiment further obtains a designated user behavior feature vector corresponding to each user identifier in the designated user identifier set, which can be used to characterize the behavior characteristics of the corresponding user. By selecting the first preset number of user behavior feature vectors that are most similar to each designated user behavior feature vector from the global user behavior feature vectors, this embodiment obtains user behavior feature vectors that are similar to the designated user behavior feature vectors in the designated user behavior feature vector set, exceeding a preset target number of users, and further obtains a preset target number of user identifiers that best match the item categories in the target item category set by comparing the overall similarities. Thus, this embodiment not only achieves the purpose of obtaining the preset target number of user identifiers, but also has a high degree of match between the obtained preset target number of user identifiers and the designated item category, thereby improving the scale and accuracy of information push related to the designated item category.
[0078] Example 2:
[0079] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0080] A target item category set is obtained according to a specified item category and a list of associated item categories; the target item category set includes the specified item category and item categories associated with the specified item category, and the list of associated item categories includes a plurality of records, each record including an item category and item categories associated with the item category.
[0081] A designated user identification set is obtained based on a first user list and a target item category set; the first user list includes a plurality of records, each record including a user identification and an item category set corresponding to the user identification; the item category set corresponding to each user identification includes a plurality of item categories; the designated user identification set includes a plurality of designated user identifications, and the item category set corresponding to each designated user identification includes at least one item category in the target item category set.
[0082] A designated user behavior feature vector set corresponding to the designated user identification set is obtained; the designated user behavior feature vector set includes a designated user behavior feature vector corresponding to each designated user identification.
[0083] Obtain a preset first number of global user behavior feature vectors in the global user behavior feature vector set that have the greatest similarity to each designated user behavior feature vector, and append the user identifiers and similarity values corresponding to the global user behavior feature vectors to the target user identifier set; the designated user behavior feature vector set is a subset of the global user behavior feature vector set, and the product of the preset first number and the number of designated user identifiers is greater than a preset number of target users.
[0084] The user identification corresponding to the largest similarity value of the preset number of target users in the target user identification set is determined as the target user identification corresponding to the specified item category.
[0085] Example 3:
[0086] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the following steps are implemented:
[0087] A target item category set is obtained according to a specified item category and a list of associated item categories; the target item category set includes the specified item category and item categories associated with the specified item category, and the list of associated item categories includes a plurality of records, each record including an item category and item categories associated with the item category.
[0088] A designated user identification set is obtained based on a first user list and a target item category set; the first user list includes a plurality of records, each record including a user identification and an item category set corresponding to the user identification; the item category set corresponding to each user identification includes a plurality of item categories; the designated user identification set includes a plurality of designated user identifications, and the item category set corresponding to each designated user identification includes at least one item category in the target item category set.
[0089] A designated user behavior feature vector set corresponding to the designated user identification set is obtained; the designated user behavior feature vector set includes a designated user behavior feature vector corresponding to each designated user identification.
[0090] Obtain a preset first number of global user behavior feature vectors in the global user behavior feature vector set that have the greatest similarity to each designated user behavior feature vector, and append the user identifiers and similarity values corresponding to the global user behavior feature vectors to the target user identifier set; the designated user behavior feature vector set is a subset of the global user behavior feature vector set, and the product of the preset first number and the number of designated user identifiers is greater than a preset number of target users.
[0091] The user identification corresponding to the largest similarity value of the preset number of target users in the target user identification set is determined as the target user identification corresponding to the specified item category.
[0092] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0093] Although some specific embodiments of the present invention have been described in detail by way of example, it should be understood by those skilled in the art that the above examples are for illustration only and are not intended to limit the scope of the present invention. It should also be understood by those skilled in the art that various modifications may be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A method for processing user behavior characteristics, characterized in that: The processing method comprises the following steps: Obtaining a target item category set based on a specified item category and a list of associated item categories; the target item category set includes the specified item category and item categories associated with the specified item category, the list of associated item categories includes a plurality of records, each record including an item category and item categories associated with the item category; Obtaining a designated user identification set based on a first user list and a target item category set; the first user list includes a plurality of records, each record including a user identification and an item category set corresponding to the user identification; the item category set corresponding to each user identification includes a plurality of item categories; the designated user identification set includes a plurality of designated user identifications, and the item category set corresponding to each designated user identification includes at least one item category in the target item category set; Obtaining a designated user behavior feature vector set corresponding to a designated user identification set; the designated user behavior feature vector set includes a designated user behavior feature vector corresponding to each designated user identification; Obtaining a first preset number of global user behavior feature vectors from the global user behavior feature vector set that have the greatest similarity to each designated user behavior feature vector, and appending the user identifiers and similarity values corresponding to the global user behavior feature vectors to the target user identifier set; the designated user behavior feature vector set is a subset of the global user behavior feature vector set, and the product of the first preset number and the number of designated user identifiers is greater than a preset number of target users; The user identification corresponding to the largest similarity value of the preset number of target users in the target user identification set is determined as the target user identification corresponding to the specified item category.
2. The method for processing user behavior characteristics according to claim 1, characterized in that: Any record included in the first user list also includes an item name set, and the item name set includes a plurality of item names; and the process of obtaining the global user behavior feature vector set includes: Inputting each item name in the item name set included in the first record into a trained neural network model, respectively, to obtain a sub-feature vector corresponding to each item name in the item name set included in the first record; the first record is any record in the first user list; the trained neural network model is used to obtain the sub-feature vector corresponding to the input item name; Obtain the item category corresponding to each item name in the item name set included in the first record; Obtaining a frequency of occurrence of an item category corresponding to each item name in the item name set included in the first record in the first user list; Obtaining a weight corresponding to each item name in the item name set included in the first record according to the frequency; the weight corresponding to any item name in the item name set included in the first record is negatively correlated with the frequency of the item name appearing in the first user list; A global user behavior feature vector corresponding to the user identifier included in the first record is obtained according to the weight and sub-feature vector corresponding to each item name in the item name set included in the first record.
3. The method for processing user behavior characteristics according to claim 1, characterized in that: Any record included in the first user list also includes time information; obtaining a global user behavior feature vector corresponding to the user identifier included in the first record based on the weight and sub-feature vector corresponding to each item name in the item name set included in the first record includes: If there are other user lists other than the first user list, determining whether there is a record including the first user identifier in the other user list; the first user identifier is the user identifier included in the first record; If the other user list includes a record for the first user identifier, obtaining a first behavior feature vector corresponding to the first user identifier based on the weight and sub-feature vector corresponding to each item name in the item name set included in the first record, and obtaining other behavior feature vectors corresponding to the first user identifier based on the other user list; the number of other behavior feature vectors is equal to the number of user lists in the other user list that include records for the first user identifier; and any record included in the other user list includes time information; The global user behavior feature vector corresponding to the first user identifier is obtained based on the first behavior feature vector corresponding to the first user identifier, other behavior feature vectors, time information included in the first record, and time information included in the target record in other user lists; the target record is a record including the first user identifier.
4. The method for processing user behavior characteristics according to claim 1, characterized in that: Obtaining a first preset number of global user behavior feature vectors having the greatest similarity to each designated user behavior feature vector in the global user behavior feature vector set includes: Use clustering algorithm to cluster the global user behavior feature vector set to obtain several clusters; Obtaining a similarity between the first designated user behavior feature vector and the center vector of each cluster; if the similarity between the first designated user behavior feature vector and the center vector of a cluster is greater than or equal to a preset similarity threshold, determining the cluster as a cluster associated with the first designated user behavior feature vector; the first designated user behavior feature vector is any designated user behavior feature vector; Performing dimensionality reduction processing on the first specified user behavior feature vector to obtain a first dimensionality-reduced user behavior feature vector; Performing dimensionality reduction processing on the global user behavior feature vectors included in all clusters associated with the first specified user behavior feature vector to obtain a set of reduced-dimensionality user behavior feature vectors; A first preset number of reduced dimensionality user behavior feature vectors having the greatest similarity to the first reduced dimensionality user behavior feature vector in the reduced dimensionality user behavior feature vector set are obtained, and global user behavior feature vectors corresponding to the first preset number of reduced dimensionality user behavior feature vectors are determined as the first preset number of global user behavior feature vectors having the greatest similarity to the first specified user behavior feature vector.
5. The method for processing user behavior characteristics according to claim 1, characterized in that: The preset first number is the product of the ratio of the preset target number of users to the number of designated user identifiers and a preset target multiple, where the preset target multiple is a multiple of 10.
6. The method for processing user behavior characteristics according to claim 3, characterized in that: Obtaining the global user behavior feature vector corresponding to the user identifier included in the first record based on the weight and sub-feature vector corresponding to each item name in the item name set included in the first record also includes: if there are no other user lists except the first user list, determining the first behavior feature vector corresponding to the first user identifier as the global user behavior feature vector corresponding to the first user identifier.
7. The method for processing user behavior characteristics according to claim 3, characterized in that: Obtaining the global user behavior feature vector corresponding to the user identifier included in the first record based on the weight and sub-feature vector corresponding to each item name in the item name set included in the first record also includes: if there is no record corresponding to the first user identifier in the other user list, determining the first behavior feature vector corresponding to the first user identifier as the global user behavior feature vector corresponding to the first user identifier.
8. The method for processing user behavior characteristics according to claim 1, characterized in that: Acquiring a specified user identification set according to the first user list and the target item category set includes: Traversing the first user list, if a record included in the first user list includes at least one item category in the target item category set, appending the user identifier included in the record to a preset user identifier set; the preset user identifier set is initialized to an empty set; After the traversal is completed, the preset user identification set is determined as the designated user identification set.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for processing user behavior characteristics according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for processing user behavior characteristics according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Product information recommendation method and device
CN111861569A
Consumer purchase intention prediction method and device, electronic equipment and storage medium
CN115456656A