Relationship analysis method, apparatus, computer device, and storage medium
By acquiring and analyzing relationship characteristics and behavioral characteristics in social networks, and using the target community discovery algorithm, the problem of insufficient accuracy of social network division is solved, achieving more efficient user relationship determination.
Patent Information
- Application Number
- CN202110543065.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-05-18
AI Technical Summary
The existing social network construction methods cannot obtain sufficient user characteristic attributes, resulting in insufficient accuracy of social network division.
By obtaining the reference features of M relationship characteristics and N reference user groups, combining the target behavior characteristics of the target user group, the target characteristics of the target user group are constructed, and the target community discovery algorithm is used for relationship analysis to determine the target user relationship between users.
It improves the accuracy of social network division and enhances the generalization ability and accuracy of computer equipment when determining user relationships.
Smart Images

Figure CN114676294B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a relationship analysis method, apparatus, computer device, and storage medium. Background Art
[0002] With the continuous in-depth development of computer technology, constructing and mining social networks using computer methods is of great significance in real production and life. Existing methods for constructing social networks can cluster based on the characteristic attributes of each user, and by referring to the characteristic attributes, the construction of social networks can be achieved. However, when a computer device obtains the characteristic attributes of each user, since it is unable to obtain a large number of characteristic attributes, the method of constructing a social network based on clustering of characteristic attributes cannot accurately divide the social network. Therefore, how to improve the accuracy of social network division has become a current research hotspot. Summary of the Invention
[0003] Embodiments of the present invention provide a relationship analysis method, apparatus, computer device, and storage medium, which can improve the accuracy of social network division.
[0004] On the one hand, embodiments of the present invention provide a relationship analysis method, including:
[0005] Obtain M relationship features and reference features of N reference user groups. One relationship feature is used to represent a type of user relationship; one reference user group corresponds to one reference feature; the reference feature of any one of the N reference user groups is used to indicate: the interaction behavior order among users in the any one of the reference user groups, where N and M are both integers greater than 1;
[0006] Obtain the target behavior feature of the target user group to be processed, where the target behavior feature is used to indicate the interaction behavior order among users in the target user group;
[0007] Determine a target relationship feature that matches the target behavior feature from the M relationship features, and construct a target feature of the target user group according to the target relationship feature and the target behavior feature;
[0008] According to the target feature and the N reference features, perform relationship analysis on the target user group and the N reference user groups, determine the target community to which the target user group belongs, and determine the target user relationship among users in the target user group according to the user relationships among users included in each reference user group in the target community.
[0009] On the other hand, an embodiment of the present invention provides a relationship analysis device, including:
[0010] An acquisition unit, configured to acquire M relationship features and reference features of N reference user groups. One relationship feature is used to represent a type of user relationship; one reference user group corresponds to one reference feature; the reference feature of any one of the N reference user groups is used to indicate: the interaction behavior order among the users in the any one of the reference user groups. Both N and M are integers greater than 1.
[0011] The acquisition unit is further configured to acquire target behavior features of a target user group to be processed, where the target behavior features are used to indicate the interaction behavior order among the users in the target user group.
[0012] A determination unit, configured to determine a target relationship feature that matches the target behavior features from the M relationship features, and construct target features of the target user group according to the target relationship feature and the target behavior features.
[0013] A division unit, configured to perform relationship analysis on the target user group and the N reference user groups according to the target features and the N reference features, and determine a target community to which the target user group belongs.
[0014] The determination unit is further configured to determine target user relationships among the users in the target user group according to the user relationships among the users included in each reference user group in the target community.
[0015] In an embodiment, both the target behavior features and the relationship features are represented by an identification sequence. The target behavior identification sequence representing the target behavior features and the relationship identification sequence representing the relationship features both include at least one behavior identification. One behavior identification corresponds to one interaction behavior. The arrangement order of the at least one behavior identification in the corresponding identification sequence is used to indicate the execution order of the corresponding interaction behavior. The determination unit 602 is specifically configured to:
[0016] Acquire the relationship identification sequence corresponding to each relationship feature among the M relationship features, and the target behavior identification sequence corresponding to the target behavior features.
[0017] From the relationship identification sequences respectively corresponding to the M relationship features, find a sequence in which at least two target behavior identifications are the same as those in the target behavior identification sequence and the corresponding arrangement order is the same, and use the relationship feature corresponding to the found sequence as the target relationship feature that matches the target behavior features.
[0018] In an embodiment, the determination unit is specifically configured to:
[0019] Obtain the target portraits of each user included in the target user group, and perform feature mapping on the target portraits of each user included in the target user group respectively to obtain the portrait features of each user included in the target user group;
[0020] Perform feature splicing on the portrait features, the target behavior features, and the target relationship features of each user in the target user group to obtain the target features of the target user group.
[0021] In one embodiment, the dividing unit is specifically configured to:
[0022] Construct a graph model according to the target features and each reference feature. The graph model includes a target node for indicating the target user group and multiple other nodes. One other node is used to indicate one reference user group, and the weight of the edge existing between two nodes is determined according to the similarity between the features of the corresponding user groups;
[0023] Adopt a target community discovery algorithm to perform node clustering operation on the nodes in the graph model, and determine the target node set where the target node is located. The user groups corresponding to the nodes included in the target node set constitute the target community to which the target user group belongs.
[0024] In one embodiment, each node in the graph model is associated with the features of the corresponding user group; the determining unit is specifically configured to:
[0025] Based on the random walk algorithm, perform random walk in the graph model, and sort the nodes in the graph model according to the order corresponding to the nodes passed by the random walk in the graph model to obtain a node sequence, where the node sequence includes at least one sequence class;
[0026] Perform hierarchical encoding on the node sequence, where one sequence class included in the node sequence corresponds to one category encoding, and the nodes in one sequence class correspond to one node encoding;
[0027] According to the hierarchical encoding result, determine the total average encoding length of the node sequence, and when the total average encoding length obtains the minimum value, obtain the target sequence class where the target node is located, and use the nodes included in the target sequence class as the nodes in the target node set where the target node is located.
[0028] In one embodiment, the obtaining unit is further configured to obtain the occurrence probability corresponding to each node included in the node sequence, and the weight of the edge corresponding to the two nodes with an edge existing between them;
[0029] The determining unit is further configured to calculate a class transition probability according to the occurrence probability corresponding to each node and the weight of the edge connecting two nodes with an edge;
[0030] The partitioning unit is further configured to partition the node sequence into at least one sequence class according to the class transition probability, and each sequence class includes one or more nodes.
[0031] In one embodiment, the determining unit is specifically configured to:
[0032] Perform class encoding on the sequence classes included in the node sequence, where the class encodings of different sequence classes are different;
[0033] Determine the sequence class to which the corresponding node belongs, and perform node encoding on the corresponding node according to the probability corresponding to each node in the sequence class, where the length of the node encoding corresponding to each node is negatively correlated with the probability that the corresponding node is passed through during random walk;
[0034] Use the node encoding of each node and the class encoding of the corresponding sequence class as the encoding of each node.
[0035] In one embodiment, the determining unit is specifically configured to:
[0036] Obtain the average encoding length corresponding to each sequence class in the node sequence and the average encoding length of each node in each sequence class according to the hierarchical encoding result;
[0037] Perform weighted averaging on the average encoding lengths corresponding to each class and the average encoding lengths of the nodes in each sequence class, and use the weighted average encoding length as the total average encoding length of the node sequence.
[0038] In one embodiment, the determining unit is specifically configured to:
[0039] Determine the type of user relationship among the users included in each reference user group in the target community, and determine the proportion of the number of reference user groups corresponding to the user relationship of the same type;
[0040] Use the user relationship corresponding to the reference user group with the largest proportion in the target community as the target user relationship among the users in the target user group.
[0041] In one embodiment, the obtaining unit is specifically configured to:
[0042] Obtain the behavioral characteristics of each reference user group under the user relationship of the target type, and use the sequence of behavior identifiers corresponding to the behavioral characteristics corresponding to each reference user group as the target sequence set, and determine the number of sequences in the target sequence set, the included behavior identifiers, and the occurrence times corresponding to each behavior identifier;
[0043] According to the occurrence times of each behavior identifier in the target sequence set, select multiple one-item prefixes from the target sequence set, and each one-item prefix includes: a behavior identifier whose occurrence times in the target sequence set are greater than the quantity threshold;
[0044] Construct sequence patterns using each one-item prefix respectively, and obtain the projection data sets of the respective one-item prefixes. The projection data sets contain the suffixes of the corresponding one-item prefixes in the corresponding interaction behavior sequences, and the suffixes include the behavior identifiers in the corresponding interaction behavior sequences that are located after the prefix;
[0045] Perform recursive mining on the projection data sets of the respective one-item prefixes to obtain K-item prefixes, and use the K-item prefixes to construct sequences respectively to obtain multiple behavior sequence patterns. The obtained behavior sequence patterns are used to represent the target relationship characteristics corresponding to the user relationship of the target type, and K is an integer greater than 1.
[0046] In one embodiment, the device further includes a sending unit.
[0047] The obtaining unit is further configured to obtain the information to be recommended, and the historical user groups interested in the information to be recommended, and determine the communities to which the historical user groups belong;
[0048] The sending unit is configured to send the information to be recommended to each user in the communities to which the historical user groups belong.
[0049] On the other hand, an embodiment of the present invention provides a computer device, including a processor, an input device, an output device, and a memory. The processor, the input device, the output device, and the memory are interconnected. Among them, the memory is used to store a computer program that supports the computer device to execute the above method. The computer program includes program instructions, and the processor is configured to call the program instructions to execute the following steps:
[0050] Obtain M relationship characteristics and the reference characteristics of N reference user groups. One relationship characteristic is used to represent a type of user relationship; one reference user group corresponds to one reference characteristic; the reference characteristic of any one of the N reference user groups is used to indicate: the interaction behavior order among the users in the any one of the reference user groups. Both N and M are integers greater than 1;
[0051] Obtain the target behavior characteristics of the target user group to be processed, where the target behavior characteristics are used to indicate the interaction behavior order among users in the target user group;
[0052] Determine the target relationship characteristics that match the target behavior characteristics from the M relationship characteristics, and construct the target characteristics of the target user group according to the target relationship characteristics and the target behavior characteristics;
[0053] According to the target characteristics and the N reference characteristics, perform relationship analysis on the target user group and the N reference user groups, determine the target community to which the target user group belongs, and determine the target user relationship among the users in the target user group according to the user relationships among the users included in each reference user group in the target community.
[0054] In another aspect, an embodiment of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions are executed by a processor, the program instructions are used to execute the relationship analysis method as described in the first aspect.
[0055] In the embodiment of the present invention, the computer device can obtain multiple relationship characteristics based on the behavior characteristics corresponding to each reference user group in the N reference user groups and by mining the behavior characteristics of the reference user groups with the same user relationship. Then, the computer device can obtain the reference characteristics of each reference user group based on the obtained relationship characteristics and the behavior characteristics of each reference user group. Moreover, the computer device can also match the target behavior characteristics of the target user group with unknown social types with the obtained multiple relationship characteristics, and then match the target relationship characteristics, so that the computer device can construct the target characteristics of the target user group based on the matched target relationship characteristics and target behavior characteristics, enabling the computer device to realize the construction of multi-dimensional characteristics of the user group, thereby improving the accuracy of the characteristics corresponding to the user group obtained by the computer device. After the computer device obtains the reference characteristics corresponding to each reference user group and the target characteristics corresponding to the target user group, it can perform relationship analysis on the target user group and the reference user groups based on the target characteristics and the reference characteristics, so that the computer device can determine the target user relationship among the users in the target user group based on the user relationships corresponding to the other reference user groups in the target community to which the target user group is assigned. The generalization ability of the method for determining the user relationship based on relationship analysis is strong, and since the accuracy of the characteristics corresponding to the user group obtained by the computer device is relatively high, the method for the computer device to obtain the user relationship among users based on relationship analysis is also helpful for improving the accuracy of the computer device in determining the user relationship among users. Description of the Drawings
[0056] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 It is a schematic diagram of a relationship analysis method provided by an embodiment of the present invention;
[0058] Figure 2 It is a schematic flowchart of a relationship analysis method provided by an embodiment of the present invention;
[0059] Figure 3 It is a schematic flowchart of a relationship analysis method provided by an embodiment of the present invention;
[0060] Figure 4 It is a schematic diagram of obtaining a node sequence provided by an embodiment of the present invention;
[0061] Figure 5 It is a schematic diagram of determining a target community provided by an embodiment of the present invention;
[0062] Figure 6 It is a schematic block diagram of a relationship analysis device provided by an embodiment of the present invention;
[0063] Figure 7 It is a schematic block diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners
[0064] An embodiment of the present invention provides a relationship analysis method, enabling a computer device to determine relationship features corresponding to each type of user relationship according to the interaction behaviors among users in a reference user group under each known type of user relationship. The relationship features can be used to reflect the constraint relationships that the interaction behaviors performed by the users in the reference user group under a type of user relationship need to satisfy. Then, further, after obtaining the target behavior features of the target user group to be processed, the computer device can determine the relationship features that match the target behavior features based on the comparison between the target behavior features and the relationship features corresponding to each type of user relationship that have been determined. Thus, the computer device can further construct the features of each user group based on the behavior features among users in each reference user group (and the target user group) and the feature data such as the relationship features that match the behavior features. Then, the computer device can perform relationship analysis on the reference user group and the target user group based on the features of each user group obtained. Thus, the computer device can determine the user relationship type corresponding to the target user group according to the result of the relationship analysis. Based on the relationship analysis, the generalization ability of the method for determining the relationship type of the target user group to be processed is strong, and the determination efficiency of the computer device in determining the user relationship type corresponding to the target user group to be processed can be improved. In a specific implementation, as Figure 1 shown, the computer device can first obtain the reference behavior features of each reference user group in the reference user group with known user relationship types. Among them, one type of user relationship corresponds to one or more reference user groups, and the number of users included in each reference user group is at least two. And in the embodiment of the present invention, the case where each reference user group includes 2 users is mainly described in detail. When the reference user group includes more than two users, the embodiment of the present invention can also be referred to.
[0065] The reference behavior characteristics corresponding to each reference user group obtained by the computer device are represented in the form of an identification sequence. That is to say, after the computer device obtains the behavior identification sequence corresponding to each reference user group, it can use the behavior time series pattern mining algorithm to perform sequence pattern mining on the behavior identification sequences corresponding to the reference user groups under the same relationship type, so as to mine the frequent sequence patterns under the corresponding relationship type. Then, it can be understood that the mined frequent sequence patterns can be used to represent the relationship characteristics corresponding to the corresponding relationship type. It can be understood that since a frequent sequence pattern is mined by the computer device based on the behavior identification sequence of the reference user group corresponding to a specific user relationship, the frequent sequence pattern can be used to indicate: the interaction behaviors that the users included in the user group under the corresponding type of user relationship should perform, and the order that the interaction behaviors need to satisfy. Among them, the mined frequent sequence patterns under the corresponding relationship type can also be represented by the corresponding identification sequences. Then, after the computer device obtains the relationship characteristics corresponding to multiple types of user relationships, it can construct user characteristics when it is necessary to determine the type of user relationship between the users in the target user group. In a specific implementation, the computer device can first construct the reference characteristics (i.e., user characteristics) of each reference user group based on the relationship characteristics and the reference behavior characteristics corresponding to the interaction behaviors performed by the users in the reference user group. Among them, when constructing the reference characteristics of the reference user group, the computer device can use the method of feature splicing. For example, the user relationship between user 1 and user 2 included in a reference user group is a couple relationship, and the interaction behaviors performed by user 1 and user 2 are: sequentially performing interaction behavior 1, interaction behavior 2, and interaction behavior 3. Then, when constructing the reference characteristics of this reference user group, the computer device can first map the couple relationship to space to obtain the corresponding relationship characteristics, assuming that feature a is obtained, and can map the interaction behaviors performed by user 1 and user 2 and the corresponding execution order to space to obtain the corresponding reference behavior characteristics, assuming that feature b is obtained. In addition, the computer device can also map the user portraits of user 1 and user 2 in this reference user group to space, assuming that feature c is obtained. Then, the reference characteristics of a reference user group obtained by the computer device can be represented as <a, b, c>. Among them, when the computer device splices the characteristics mapped to space to obtain the reference characteristics, the splicing order is not limited. For the above reference user group, the reference characteristics of the reference user group spliced by the computer device can also be <a, c, b>.
[0066] In addition, the computer device will also construct user characteristics of the target user group based on relationship characteristics. In a specific implementation, the computer device may first match the target behavior characteristics of the target user group with each relationship characteristic, so as to select the target relationship characteristics that match the target behavior characteristics successfully. Then, the computer device can construct the target characteristics of the target user group based on the target relationship characteristics, the behavior characteristics of the target user group, and other characteristics related to the target user group. In one embodiment, after the computer device determines the reference characteristics of each reference user group and the target characteristics of the target user group, the computer device can perform relationship analysis on the reference user group and the target user group based on the reference characteristics and the target characteristics, so as to divide the user groups with high similarity between characteristics into the same community, and divide the user groups corresponding to the characteristics with low similarity into different communities. Furthermore, the computer device can determine the target user relationship between the users in the target user group based on the target community to which the target user group belongs.
[0067] In one embodiment, the relationship analysis performed by the computer device on the reference user group and the target user group based on user characteristics can result in one or more communities being divided. Then, the other user groups included in the target community including the target user group must be reference user groups with known user relationships. Further, the computer device can determine the target user relationship corresponding to the users in the target user group by calculating the proportion of the number of reference user groups with known user relationships in the target user group. Among them, the computer device can use the user relationship with a corresponding proportion greater than or equal to the threshold as the target user relationship corresponding to each user in the target user group, or the computer device can also use the user relationship when the corresponding proportion in the target community reaches the maximum value as the target user relationship corresponding to each user in the target user group, so that the computer device can construct user characteristics and combine relationship analysis to obtain the target user relationship between the users in the target user group with unknown user relationship types. The generalization ability of this user relationship determination method is strong, and it can effectively improve the efficiency and accuracy of the computer device in determining the user relationship between the users in the user group with unknown user relationship types.
[0068] Please refer to Figure 2 , which is a schematic flowchart of a relationship analysis method proposed by an embodiment of the present invention. This relationship analysis method can be executed by the above-mentioned computer device. As Figure 2 shown, this method may include:
[0069] S201. Obtain M relationship features and reference features of N reference user groups. One relationship feature is used to represent one type of user relationship; one reference user group corresponds to one reference feature. The reference feature of any one of the N reference user groups is used to indicate the interaction behavior order among the users in any one of the reference user groups. Both N and M are integers greater than 1.
[0070] In one embodiment, both the relationship features and the reference features are determined by the computer device based on the obtained behavior features of the reference user groups. That is, before obtaining the relationship features and the reference features, the computer device needs to first obtain the reference user groups of different types of user relationships and the behavior data of each reference user group. The user relationships may include social relationships, such as couple relationships, father-son relationships, etc., or the user relationships may also include family relationships, such as mother-son relationships, colleague relationships, etc. In addition, the user relationships may also include other relationships between users, etc. One user relationship is one type of user relationship. That is, a couple relationship is one type of user relationship, and a father-son relationship is another type of user relationship. When the computer device obtains the reference user groups, it can obtain one or more reference user groups under each type of user relationship and add a category label corresponding to the user relationship to each reference user group obtained under each user relationship by using the category identifier (such as the category identity identifier) corresponding to the user relationship. The reference user groups obtained by the computer device may be as shown in Table 1:
[0071] Table 1
[0072]
[0073]
[0074] Among them, the above 0, 1, …, n, etc. are the category labels corresponding to the user relationships added by the computer device to each reference user group, so as to be used to refer to the type of user relationship among the users in the corresponding reference user group. For example, the above 0 can be used to represent the reference user group with the corresponding user relationship being a couple relationship, and 1 is used to represent the reference user group with the corresponding user relationship being a colleague relationship. The manner in which the computer device obtains multiple reference user groups and the manner of adding a category label corresponding to the user relationship to the users in each reference user group are not limited.
[0075] In one embodiment, if the number of reference user groups obtained by a computer device is N, after the computer device obtains N reference user groups and the category labels of the corresponding user relationships, it can also construct the behavior characteristics corresponding to the interaction behaviors performed by the users included in the reference user groups of each type of user relationship. Among them, when the computer device obtains the behavior characteristics corresponding to each reference user group, it needs to first obtain the interaction behaviors performed by the users in each reference user group, and when the computer device obtains the interaction behaviors performed by the users in each reference user group, it can obtain the interaction behaviors performed by the users in the corresponding reference user group within the key time range associated with the user relationship based on the user relationship corresponding to the users in each reference user group. It can be understood that the user group of a certain user relationship will perform different interaction behaviors within a specific time range (i.e., the key time range) from those within other time ranges. For example, the user group in a couple relationship will transfer electronic resources of a specific amount on Valentine's Day or the Qixi Festival (i.e., the key time range), such as the users included in the user group transfer electronic resources of 520 or 999 to each other. And within other time ranges, the user group in a couple relationship rarely forwards the above-mentioned electronic resources of a specific amount. That is to say, within this key time range, the interaction behaviors performed by the users in the user group can, to a certain extent, reflect the social meaning of this key time range. Then, by obtaining the interaction behaviors within the key time range associated with the user relationship, the computer device can make the obtained interaction behaviors best reflect the user relationship between user groups. Further, the computer device can perform feature mapping on the interaction behaviors performed by the users in the corresponding reference user group within the key time range, so as to obtain the behavior characteristics of the corresponding reference user group. In one embodiment, when the computer device performs feature mapping on the interaction behaviors performed by the users in the corresponding reference user group to obtain the behavior characteristics corresponding to the corresponding reference user group, the computer device can first map each interaction behavior into a corresponding behavior identifier, and then arrange the corresponding behavior identifiers in ascending (or descending) order in sequence according to the order in which each interaction behavior is performed within this key time range, so as to use the obtained sequence of behavior identifiers as the behavior characteristics of the corresponding reference user group.
[0076] In one embodiment, if the user relationship among users in the reference user group 1 is a couple, the computer device may take Valentine's Day as the key time point and the time range of x days before and after Valentine's Day as the key time range. Then, the computer device can obtain the interaction behaviors performed by the users in the reference user group 1 within x days before and after Valentine's Day. If the interaction behaviors performed by the users in the reference user group 1 include user A transferring 520 yuan to user B and B receiving it, the computer device can first mark this interaction behavior as A-transfer 520-B and B-receive transfer 520-A. And the interaction behaviors performed by the users in the reference user group 1 also include that user A and user B posted moments in sequence or simultaneously within the key time node, and user A invited user B to follow a public account of a specific theme (such as wedding photography, etc.). The computer device can first mark the interaction behaviors performed by the users in the reference user group 1. Further, the computer device can determine the behavior identifier corresponding to each interaction behavior according to this mark. Specifically, when the computer device maps each interaction behavior into the corresponding behavior identifier, it maps the same type of behaviors into the same behavior identifier. For example, it maps the transfer behavior into one behavior identifier and the behavior of following a public account into another behavior identifier. Then, in order to avoid different transfer amounts or different public accounts followed, and mapping the same type of interaction behaviors into different behavior identifiers, the computer device can first perform the same-type identification processing, that is, unify the transfer amounts such as 520, 5.20, 52.0, 99, 1314, 13.14, 131.4, 999, etc. under the transfer behavior and mark them as the transfer behavior. Thus, when performing the behavior identifier conversion, the same type of transfer amounts can be represented by the same identifier. That is, when the computer device maps the interaction behaviors of transferring 520 and transferring 5.2, they are both mapped to the behavior identifier a.
[0077] By using the above method for determining behavioral characteristics, the computer device can determine the behavioral characteristics of each reference user group. After obtaining the behavioral characteristics of each reference user group, the computer device can mine the behavioral characteristics of reference user pairs under the same type of user relationship, so as to obtain the relationship characteristics corresponding to each user relationship. In one embodiment, since the behavioral characteristics of each reference user group obtained by the computer device are a sequence of behavior identifiers composed of one or more behavior identifiers, when the computer device determines the relationship characteristics corresponding to each user relationship, it can perform frequent time series pattern mining on the sequence of behavior identifiers corresponding to the behavioral characteristics of the reference user group under each user relationship, so as to obtain multiple behavior sequence patterns, and then determine the sequence pattern corresponding to each user relationship, and determine the relationship characteristics corresponding to the corresponding user relationship. In addition, the computer device can also obtain the relationship characteristics corresponding to some user relationships based on prior knowledge. Then, after the computer device obtains M relationship characteristics and the behavioral characteristics corresponding to N reference user groups, it can construct the reference characteristics of each reference user group, so as to obtain the reference characteristics of each reference user group among the N reference user groups. After the computer device determines the sequence pattern corresponding to each user relationship, it can first send the sequence pattern corresponding to each user relationship to the blockchain network, and when user characteristics need to be constructed, obtain the corresponding sequence pattern from the blockchain network for user characteristic construction.
[0078] In one embodiment, when the computer device determines the reference characteristics of each reference user in the N reference user groups, it can only refer to the behavioral characteristics of each reference user group and the relationship characteristics corresponding to the corresponding user relationship; or, in another implementation manner, the computer device can further obtain other characteristics of each reference user group, such as the user portraits of the users in the reference user group, etc., to construct the reference characteristics of each reference user in the N reference user groups. In one embodiment, a user portrait is virtual representative data of an actual user. The user portrait is jointly transformed based on data such as the basic user situation of the actual user (such as age, occupation, etc.), social preferences (such as wealth, whether married, whether having children, etc.), and consumption preferences (such as consumption preference merchants, consumption preference amounts). Based on the construction of the user portrait, each user can be effectively distinguished from other users, and based on the differences in user portraits, the differences between each user and other users can be directly understood.
[0079] S202, obtain the target behavioral characteristics of the target user group to be processed, where the target behavioral characteristics are used to indicate the interaction behavior order among the users in the target user group.
[0080] S203. Determine a target relationship feature that matches the target behavior feature from the M relationship features, and construct the target feature of the target user group based on the target relationship feature and the target behavior feature.
[0081] In steps S202 and S203, when the computer device needs to determine the user relationship type between users in the target user group, it can first obtain the target behavior feature of the target user group to be processed. Similarly, when the computer device obtains the target behavior feature of the target user group, it can also use the above method of obtaining the behavior features corresponding to the users in the reference user group. Then, after the computer device obtains the target behavior feature corresponding to the target user group, the computer device can match the target behavior feature with the M relationship features respectively, so as to determine the target relationship feature that matches the target behavior feature from the M relationship features. In one embodiment, both the relationship feature and the behavior feature are identification sequences composed of one or more behavior identifiers. Then, in the process of the computer device matching the target behavior feature of the target user group with the M relationship features respectively, it is to match whether the behavior identifiers included in the target behavior identifier sequence corresponding to the target behavior feature and the relationship identifier sequence corresponding to each relationship feature are the same, and whether the arrangement order of the same behavior identifiers in the target behavior feature sequence and the relationship identifier sequence is the same. The target relationship feature obtained by the computer that matches the target behavior feature is the target relationship identifier sequence corresponding to the target relationship identifier sequence and the target behavior identifier sequence corresponding to the target behavior feature, that is, there are at least two identical behavior identifiers, and the same behavior identifiers are in the same order in the target behavior identifier sequence and the target relationship identifier sequence respectively.
[0082] In one embodiment, if the computer device determines that the users in the target user group have successively performed interaction behaviors including: the behavior of transferring electronic resources, and the behavior of following the same official account, and first perform the behavior of transferring electronic resources, and then perform the behavior of following the same official account, then, after determining the interaction behaviors performed by the users in the target user group, the computer device can convert the interaction behaviors into corresponding behavior identifiers. For example, it can convert the behavior of performing electronic resource conversion into behavior identifier a, and convert the behavior of following the same official account into behavior identifier b. Furthermore, based on the execution order of the interaction behaviors, the computer device constructs a target behavior identifier sequence corresponding to the target user group, which is ac. If the relationship identifier sequences corresponding to the relationship features obtained by the computer device include abc and dca, then the computer device can determine that in each obtained relationship identifier sequence, there are behavior identifiers that are the same as the behavior identifier a in the target behavior identifier sequence, and behavior identifiers that are the same as the behavior identifier c in the target behavior identifier sequence. However, based on the arrangement order of the behavior identifier a and the behavior identifier c in the target behavior identifier sequence ac, it can be seen that the users in the target user group first perform the interaction behavior corresponding to the behavior identifier a, and then perform the interaction behavior corresponding to the behavior identifier c. According to the relationship identifier sequence abc obtained by the computer device, the interaction behavior corresponding to the behavior identifier a will be performed first, and then the interaction behavior corresponding to the behavior identifier c will be performed. According to the relationship identifier sequence dca, the interaction behavior corresponding to the behavior identifier c will be performed first, and then the interaction behavior corresponding to the behavior identifier a will be performed. Thus, the computer device can determine that the relationship identifier sequence abc matches the target behavior identifier sequence ac successfully, and use the relationship feature corresponding to the relationship identifier sequence abc as the target relationship feature that matches the target behavior feature.
[0083] After the computer device determines the target relationship feature that matches the target behavior feature, the computer device can construct the target feature of the target user group according to the target relationship feature and the target behavior feature. In the same way as constructing the reference feature of the reference user group, the computer device can construct the target feature of the target user group only based on the target behavior feature and the target relationship feature of the target user group. Or, similarly, the computer device can also obtain other features related to the users in the target user group, such as the user portraits of the users in the target user group, etc., and construct the user feature of the target user group. After the computer device obtains the reference features of each of the N reference user groups and the target feature of the target user group, it can perform relationship analysis on the target user group and each reference user group based on the target feature and the reference feature, and based on the result of the relationship analysis, determine the user relationship among the users in the target user group, that is, then execute step S204.
[0084] S204. Perform relationship analysis on the target user group and N reference user groups according to the target feature and the N reference features, determine the target community to which the target user group belongs, and determine the target user relationship among the users in the target user group according to the user relationships among the users included in each reference user group in the target community.
[0085] When a computer device performs relationship analysis on the target user group and the reference user groups based on the target feature and the reference feature of each reference user, the computer device may first construct a graph network model based on the features of each user group (including the target user group and the reference user groups). Then, the computer device can perform relationship analysis on each user group based on the graph network model and determine the target community to which the target user group belongs (i.e., the set of target user groups to which the target user group belongs). In one embodiment, when the computer device constructs a graph network model according to the features of the user groups, the computer device may first map each user group to a node in space and determine whether there is an edge between different nodes based on the similarity between the corresponding features of each user group. It can be understood that the computer device can determine that there is an edge between two nodes when the similarity between the features corresponding to the user groups of the two nodes is greater than or equal to the similarity threshold, and when the computer device determines that the similarity between the features corresponding to the user groups of the two nodes is less than the similarity threshold, it determines that there is no edge between the two nodes. Then, based on the nodes in space and the edges between the nodes, the computer device can construct a graph model. Among them, when the computer device constructs a graph model, it can be constructed based on artificial intelligence (AI) technology. Artificial intelligence is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0086] After the computer device constructs a graph model, the computer device can further optimize the constructed graph model based on a target community discovery algorithm. It can be understood that the process of optimizing the graph model is the process of re-dividing different nodes in the graph network model into different node sets. That is to say, after the computer device obtains the optimal graph network model, it can realize dividing each node into different node sets. Further, the computer device can determine the target user relationship among the users in the target user group based on the target node set where the target node corresponding to the target user group is located. In one embodiment, when the computer device optimizes the constructed graph model based on the target community discovery algorithm, since the nodes in the graph network model represent a user group and the edges represent the similarity degree between the characteristics of the corresponding user groups, and the greater the similarity degree, the greater the relevance (or intimacy degree) between the corresponding user groups. Then, the greater the weight value assigned by the computer device to the corresponding edge. Further, the computer device can perform coding clustering based on the initial network model and the similarity degree between the edges corresponding to each node in the initial network model, so as to continuously reduce the coding length corresponding to the node set obtained by clustering. When the coding length of the node set obtains the minimum value, the node set where the target node corresponding to the target user group is located is used as the target node set, and the target user relationship among the users in the target user group is determined according to the user relationships of the user groups corresponding to the other nodes in the target node set.
[0087] In one embodiment, after the computer device determines the target user relationship among the users in the target user group, the computer device not only realizes the determination of the user relationship among the users in the target user group with unknown user relationships, but also can perform precise recommendation marketing based on the optimized graph network model. For example, if the computer device has information to be recommended, the computer device can obtain the historical user group interested in the information to be recommended, and then can determine the community to which the historical user group belongs based on the optimized graph network model. Then the computer device can think that the users in the community to which the historical user group belongs are more likely to be interested in the information to be recommended. Further, the computer device can send the information to be recommended to each user in the community to which the historical user group belongs, instead of sending it to the users in other communities, so as to realize the precise recommendation of the information to be recommended.
[0088] In an embodiment of the present invention, a computer device can obtain multiple relationship features by mining the behavioral features corresponding to each of the N reference user groups and mining the behavioral features of the reference user groups with the same user relationship. Then, based on the obtained relationship features and the behavioral features of each reference user group, the computer device can obtain the reference features of each reference user group. Moreover, the computer device can also match the target behavioral features of the target user group with an unknown social type with the obtained multiple relationship features, and then match the target relationship features, so that the computer device can construct the target features of the target user group based on the matched target relationship features and target behavioral features, enabling the computer device to realize the construction of multi-dimensional features of the user group, thereby improving the accuracy of the features corresponding to the user group obtained by the computer device. After the computer device obtains the reference features corresponding to each reference user group and the target features corresponding to the target user group, it can perform relationship analysis on the target user group and the reference user groups based on the target features and reference features, so that the computer device can determine the target user relationships among the users in the target user group based on the user relationships corresponding to the other reference user groups in the target community to which the target user group belongs. The generalization ability of the method for determining user relationships based on relationship analysis is strong, and since the accuracy of the features corresponding to the user group obtained by the computer device is relatively high, the method for the computer device to obtain the user relationships between users based on relationship analysis also helps to improve the accuracy of the computer device in determining the user relationships between users.
[0089] Please refer to Figure 3 , which is a schematic flowchart of a relationship analysis method provided by an embodiment of the present invention. The relationship analysis method can be executed by the above computer device, as Figure 3 shown. The method may include:
[0090] S301, obtain M relationship features and the reference features of N reference user groups. One relationship feature is used to represent a type of user relationship; one reference user group corresponds to one reference feature; the reference feature of any one of the N reference user groups is used to indicate: the type of user relationship and the interaction behavior order among the users in any one reference user group. Both N and M are integers greater than 1.
[0091] In one embodiment, the relationship features obtained by the computer device are mined from the behavioral features corresponding to each reference user group in the reference user group of the same user relationship. If the relationship features corresponding to the user relationship of the target type are included in the M relationship features obtained by the computer device, then when the computer device obtains the relationship features corresponding to the user relationship of the target type, it can first obtain the behavioral features of each reference user group under the user relationship of the target type, and use the behavioral identification sequence corresponding to the behavioral features of each reference user group as the target sequence set. Further, the computer device can determine the number of sequences in the target sequence set, the included behavioral identifications, and the occurrence times corresponding to each behavioral identification. Furthermore, the computer device can perform frequent sequence pattern mining (prefixspan) based on the number of sequences in the target sequence set, the included behavioral identifications, and the occurrence times corresponding to each behavioral identification to obtain the target relationship features corresponding to the target relationship features. In one embodiment, when the computer device performs frequent sequence pattern mining, the computer device can first select multiple one-item prefixes from the target sequence set according to the occurrence times of each behavioral identification in the target sequence set. Each one-item prefix includes: the behavioral identifications whose occurrence times in the target sequence set are greater than the quantity threshold, and further, it can use each one-item prefix to construct sequence patterns respectively, and obtain the projection data sets of each one-item prefix. The projection data set contains the suffixes of the corresponding one-item prefix in the corresponding interaction behavior sequence, and the suffix includes the behavioral identifications in the corresponding interaction behavior sequence that are located after the prefix; then it can perform recursive mining on the projection data sets of each one-item prefix to obtain K-item prefixes, and use the K-item prefixes to construct sequences respectively to obtain multiple behavioral sequence patterns. The obtained behavioral sequence patterns are used to represent the target relationship features corresponding to the user relationship of the target type, and K is an integer greater than 1.
[0092] In one embodiment, when the computer device mines the behavioral sequence patterns between each user group, it can use the frequent sequence mining algorithm to mine the temporal behavioral sequence patterns of the feature relationship type. Then the computer device can use the temporal behavioral sequences of each reference user group (that is, the behavioral identification sequences composed of the interaction behaviors performed by the users in each reference user group) as the target sequence set to be mined. It can be understood that each behavioral identification sequence included in the target sequence set is based on the mining object to be processed by the mining algorithm. Moreover, when the computer device uses this algorithm, it will also use the minimum support strategy for mining, and the minimum support strategy can be shown as formula (1):
[0093] min_sup=a*n Formula (1)
[0094] Among them, n is the total number of reference user groups under the same user relationship type, and a is the minimum support rate. The minimum support rate parameter is adjusted according to the reference user groups under the same user relationship type. The specific operation steps of the algorithm include the following ① to ③, where:
[0095] ① Determine, for each behavior identifier sequence in each target sequence set, the behavior identifier sequence in which each sequence element (i.e., sequence identifier) is located, as well as the corresponding sequence prefix and projection data set;
[0096] ② Statistically analyze the behavior identifiers in the target sequence set for which the corresponding occurrence times of the time behavior sequence elements are greater than the quantity threshold. That is, if the occurrence frequency of the behavior identifier with a corresponding occurrence times greater than the quantity threshold as the prefix is higher than the minimum support threshold, then it can be understood that the behavior identifier with a corresponding occurrence times greater than the quantity threshold is a prefix item;
[0097] ③ Recursively mine all prefixes with length i that meet the minimum support requirement.
[0098] In one embodiment, when the computer device recursively mines all prefixes with length i that meet the minimum support requirement, the following steps (1) - (3) can be specifically executed
[0099] (1) Mine the prefix to obtain the projection data set of the prefix. If the projection data set is empty, return the recursion; and
[0100] (2) If the obtained projection data set of the prefix is not empty, then statistically analyze the minimum support of each item in the corresponding projection data set, and merge each single item that meets the minimum support with the current prefix to obtain a new prefix. If the support requirement is not met, return the recursion;
[0101] (3) Let i = i + 1, so that the prefix is each new prefix after merging the single items, and recursively execute step (3).
[0102] After the computer device executes the above steps, it can return all the sequence patterns of each sequence identifier in the target sequence set. For example, if the user relationship of the target type obtained by the computer is a couple relationship, the computer device can map the behavior identifiers of the interaction behaviors generated by each user in the reference user group with the couple relationship. If the computer device can map the interaction behaviors generated by each user in a certain reference user group as follows: mapping the transfer or red envelope of the same type of red envelope transfer amount to the behavior identifier a, posting a moment within the key time node to the behavior identifier b, and following the same public account to the behavior identifier c, etc., then the computer device can map the interaction behaviors generated by each user in the certain reference user group to the behavior identifier sequence abc. Suppose the reference user groups obtained by the computer device only include reference user group 1 and reference user group 2, and the behavior identifier sequence of reference user group 1 is bcagh, and the behavior identifier sequence of reference user group B is bcdaghf. Then when the computer device mines the sequence patterns contained in the behavior sequences of each user in the reference user group based on the Prefixspan algorithm, it can assume that the set minimum support threshold is 0.5. Then the computer device can determine that the minimum number of times that the mined prefix needs to meet should be greater than 0.5 * 2 = 1, that is, the mined prefix appears at least twice in the target sequence set composed of the behavior identifier sequence bcdaghf and the behavior identifier sequence bcagh. Then the computer device can determine a prefix that meets the threshold and the corresponding suffix (i.e., the corresponding projected data set) as shown in Table 2:
[0103] Table 2
[0104]
[0105] Then, the computer device can similarly determine the two prefixes that meet the minimum support threshold and the corresponding suffixes, as shown in Table 3 specifically:
[0106] Table 3
[0107]
[0108] In addition, the computer device can use the same method to determine the three prefixes that meet the minimum support threshold and the corresponding suffixes, as shown in Table 4 specifically:
[0109] Table 4
[0110]
[0111] The computer device can similarly determine the four prefixes that meet the minimum support threshold and the corresponding suffixes, as shown in Table 5 specifically:
[0112] Table 5
[0113]
[0114] Finally, the computer device can determine the five prefixes that meet the minimum support threshold and their corresponding suffixes, as shown in Table 6 specifically:
[0115] Table 6
[0116] Five prefixes Corresponding suffix bcagh f
[0117] For the sequence pattern mining of Tables 2 to 6 as described above, the computer device can use the above method to mine the behavior identification sequences corresponding to the reference user groups under each user relationship type, so as to obtain the frequent sequence model corresponding to each user relationship type, that is, the relationship behavior sequence pattern (i.e., the relationship identification sequence). For example, when mining the behavior identification sequences of the two reference user groups, namely reference user group 1 and reference user group 2, the relationship behavior sequences obtained both include the sequence pattern bcagh. Then the computer device can determine the interaction behaviors indicated by each behavior identification in the sequence pattern bcagh, as well as the execution order of the corresponding interaction behaviors identified by the corresponding arrangement order of each behavior identification, so as to obtain the relationship characteristics corresponding to the user relationship (i.e., lovers) under the target type.
[0118] In addition, based on the above mining steps, the computer device can also determine the support of the mined sequence pattern through the calculation method such as formula (2), and obtain the frequency of the occurrence of the sequence pattern in this type of user relationship. Specifically,
[0119]
[0120] Then, for example, for the mined sequence pattern bcagh, the computer device knows that the support corresponding to this sequence pattern is 1.
[0121] After the computer device separately mines the relationship features corresponding to the user relationships under each type, the computer device can also obtain the relationship features corresponding to the user relationships under other types that have not been mined based on prior knowledge, and then can obtain multiple relationship features and construct the reference features for each reference user group. For example, if the computer device has not mined the relationship features corresponding to the user relationships of the friend type, and based on prior knowledge, it is known that for the reference user group under the user relationship of the friend type, the interactive behaviors that will be sequentially executed necessarily include interactive behavior X and interactive behavior Y, that is, first execute interactive behavior X and then execute interactive behavior Y. If the behavior identifier mapped by this interactive behavior X is x, and the behavior identifier mapped by interactive behavior Y is y, then based on this prior knowledge, the frequent sequence pattern corresponding to the reference user group under the friend type can be obtained as xy. Then it can be understood that the interactive behaviors indicated by the behavior identifiers in the frequent sequence pattern xy, and the execution order of the interactive behaviors indicated by the arrangement order of the corresponding identifiers can be used to represent the relationship features corresponding to the user relationships of the friend type.
[0122] S302. Obtain the target behavior features of the target user group to be processed, where the target behavior features are used to indicate the interactive behavior order among the users in the target user group.
[0123] S303. Determine the target relationship feature that matches the target behavior feature from the M relationship features, and construct the target feature of the target user group according to the target relationship feature and the target behavior feature.
[0124] In step S302 and step S303, after the computer device obtains the target behavior features of the target user group to be processed, since both the target behavior features and the relationship features are represented by an identifier sequence, and the target behavior identifier sequence representing the target behavior features and the relationship identifier sequence representing the relationship features both include at least one behavior identifier, and one behavior identifier corresponds to one interactive behavior, the arrangement order of the at least one behavior identifier in the corresponding identifier sequence can be used to indicate the execution order of the corresponding interactive behavior; then, when the computer device determines the target relationship feature that matches the target behavior feature from the M relationship features, it can first obtain the relationship identifier sequence corresponding to each relationship feature among the M relationship features, and the target behavior identifier sequence corresponding to the target behavior features; furthermore, the computer device can find out the sequences in the relationship identifier sequences respectively corresponding to the M relationship features that are the same as at least two target behavior identifiers in the target behavior identifier sequence and have the same corresponding arrangement order, and use the relationship feature corresponding to the found sequence as the target relationship feature that matches the target behavior feature.
[0125] When the computer device constructs the target features of the target user group based on the target relationship features and target behavior features, in order to further improve the accuracy of the target features of the constructed target user group when describing each user in the target user group, the computer device can obtain the target portraits of each user included in the target user group, and perform feature mapping on the target portraits of each user included in the target user group respectively to obtain the portrait features of each user included in the target user group. Furthermore, the computer device can perform feature splicing on the portrait features, target behavior features, and target relationship features of each user in the target user group to obtain the target features of the target user group. In a specific implementation, the portrait features obtained by performing feature mapping based on the user portrait may include: basic portrait features, wealth features, life stage features, consumption preference features, etc.; and the basic portrait features include features such as age, gender, permanent residence, occupation, etc., the wealth features include features such as asset score, consumption ability, risk tolerance, etc., the life stage features include features such as whether married, whether having children, whether having a house, etc., and the consumption preference features include features such as consumption preference merchants, consumption preference amounts, consumption preference times, etc.
[0126] In one embodiment, after the computer device obtains the above portrait features, it will also perform preprocessing and feature construction on the portrait features, and use the processed features as the portrait features of the user group. Among them, when the computer device performs preprocessing and feature construction on the above portrait features, it will execute one or more of the following steps ① to ⑤.
[0127] ① Discard features with too many missing values. Among them, the computer device can set a filtering threshold using a formula such as formula (3), so that by calculating the discard values corresponding to each portrait feature, the features with corresponding discard values reaching the filtering threshold can be deleted.
[0128] Filtering threshold = sample data volume * n Formula (3)
[0129] Among them, the general value of n is 0.4. However, the computer device can also adjust the value of n according to specific application scenarios.
[0130] ② Process outliers. The computer can, according to the feature distributions corresponding to various portrait features in the obtained portrait features, regard the features with corresponding feature values being too large and ranked in the top 1 / m as outlier features, and discard the determined outlier features, where m can be set to 10000 and can also be flexibly adjusted according to specific application scenarios.
[0131] ③ Missing value processing. For the portrait features with missing feature values, the computer can fill the missing values of continuous features with the mean value, and fill the discrete features with constants and regard them as separate categories.
[0132] ④ Feature derivation. The computer device can combine and derive features through feature transformation, feature squaring, and feature addition and subtraction.
[0133] ⑤ Feature processing. The computer device will discretize continuous features and perform feature encoding on discrete features. The computer device can use one-hot encoding for feature encoding.
[0134] After the computer device obtains the portrait features of each user included in the target user group, when the computer device splices features based on the portrait features, target behavior features, and target relationship features of each user in the target user group to obtain the target features of the target user group, the computer device can first perform feature encoding on the portrait features, target behavior features, and target relationship features of each user in the target user respectively, and splice the obtained feature encodings, so as to use the spliced feature encodings as the target features of the target user group. It can be understood that there must be certain associations and causal relationships among the interaction behaviors performed by users in a user group. If the target user group is a user group in a couple relationship, and the user group in a couple relationship includes user a1 and user a2, and the corresponding interaction behaviors include: a1 reposted a2's Moments within a critical time range, a2 liked a1's Moments, a2 forwarded a certain shopping link to a1 and a1 made a purchase through this shopping link, etc. It can be seen that this series of behaviors actually reflects the correlation between the user behaviors in the target object group. Then, after the computer device obtains the target behavior features of the target user group, when encoding the target behavior features, it can perform one-hot encoding on different behaviors and sort them based on the order of occurrence of each behavior feature within the critical time range, so as to obtain the encoding corresponding to the target behavior feature. For example, the computer device can encode the behavior feature of reposting Moments as [1,0,0,0,0,0], the behavior feature of liking Moments as [0,1,0,0,0,0], the behavior of forwarding a shopping link as [0,0,1,0,0,0], and the behavior of clicking on this shopping link to make a purchase as [0,0,0,1,0,0].
[0135] Using the same encoding method, the computer device can respectively perform feature encoding on the target relationship features and target behavior features corresponding to the target user group, and combine the portrait features of each user in the target user group to obtain the target features corresponding to the target user group. In one embodiment, the computer device can simultaneously perform feature encoding on the above-mentioned target relationship features, target behavior features, and the portrait features of each user in the target user group to obtain the target features corresponding to the target user group. However, in the actual application process, the computer device generally may first obtain the portrait features of each user in the target user group and the target behavior features. Therefore, the computer device can first perform feature encoding on the portrait features of each user in the target user group and the target behavior features. For example, when the target user group includes user a1 and user a2, if the encoded portrait features of user a1 are [0.2, -0.4, 0.09, 0.54, -2.5], the encoded portrait features of user a2 are [0.7, -0.01, 0.3, 0.4, 9], and the encoded target behavior features are [1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0], then the computer device can first splice the encoded portrait features of user a1, the encoded portrait features of user a2, and the encoded portrait features of the target behavior features. The spliced encoded features can be: [0.2, -0.4, 0.09, 0.54, -2.5, 0.7, -0.01, 0.3, 0.4, 9, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0]. Further, after the computer device determines the target relationship features corresponding to the target user group in the subsequent process, it can perform encoding processing on the target relationship features and further splice the encoded target relationship features with the spliced encoded features to obtain the target features corresponding to the target user group. In one embodiment, when performing feature encoding, the computer device can directly perform encoding, or, it can also refer to the support degree corresponding to the relationship features included in the features of each user group, and perform weighted encoding on the relationship features based on the support degree when performing feature encoding on the relationship features.
[0136] S304, according to the target features and N reference features, perform relationship analysis on the target user group and N reference user groups to determine the target community to which the target user group belongs.
[0137] After the computer device obtains the target features corresponding to the target user group and the reference features corresponding to N reference user groups, it can perform relationship analysis on the target user group and the N reference user groups based on the target features and the N reference features, and then determine the target community to which the target user group belongs. In a specific implementation, the computer device can first construct a graph model according to the target features and each reference feature. The graph model includes a target node for indicating the target user group and multiple other nodes. One other node is used to indicate one reference user group. The weight of the edge between two nodes is determined according to the similarity between the features of the corresponding user groups. Then, the computer device can use the target community discovery algorithm to perform node clustering operations on the nodes in the graph model to determine the target node set where the target node is located. The user groups corresponding to the nodes included in the target node set form the target community to which the target user group belongs. In one embodiment, when the computer device determines whether there is an edge between two corresponding nodes according to the similarity between the features of the two corresponding user groups, the computer device can first determine the feature codes corresponding to the features of the two corresponding user groups, so as to determine the similarity corresponding to the features of the two corresponding user groups based on the feature codes. Subsequently, the computer device can be based on the set similarity threshold S (assumed to be 0.7), and when the similarity corresponding to the features of the two corresponding user groups is higher than the similarity threshold S, add the corresponding edge to these two nodes, and after adding the corresponding edges to each of the mapped nodes, normalize the weight of the corresponding edge (i.e., similarity), and use the normalized weight as the transition probability p α→β 。
[0138] In one embodiment, each node in the graph model is associated with the characteristics of the corresponding user group. When the computer device uses the target community discovery algorithm to perform node clustering operations on the nodes in the graph model and determine the target node set where the target node is located, it can first perform random walks in the graph model based on the random walk algorithm, and sort the nodes in the graph model according to the order corresponding to the nodes passed through in the graph model during the random walk, obtaining a node sequence. Among them, the node sequence includes at least one sequence class. Further, the computer device can perform hierarchical encoding on the node sequence. Among them, one sequence class included in the node sequence corresponds to one category encoding, and the nodes in one sequence class correspond to one node encoding. Then, according to the hierarchical encoding result, the computer device can determine the total average encoding length of the node sequence, and when the total average encoding length obtains the minimum value, obtain the target sequence class where the target node is located, and use the nodes included in the target sequence class as the nodes in the target node set where the target node is located. In one embodiment, when the computer device determines the sequence classes included in the obtained node sequence based on the obtained node sequence, the computer device can first obtain the occurrence probability corresponding to each node included in the node sequence, and the weight of the edge connecting two nodes with an edge. Further, the computer device can calculate the category transition probability according to the occurrence probability corresponding to each node and the weight of the edge connecting two nodes with an edge. It can be understood that the computer device can divide the node sequence into at least one sequence class according to the category transition probability, and each sequence class includes one or more nodes.
[0139] In one embodiment, if the graph model constructed by the computer device is the graph model marked by 40 in Figure 4 then, the node sequence determined by the computer device after random walk can be as shown in the sequence marked by 41 in Figure 4 and after the computer device divides the node sequence into at least one sequence class based on the category transition probability, it can be as shown in the sequence marked by 42 in Figure 4 Among them, in the sequence marked by 42 in Figure 4 one filled color is used to indicate one sequence class, and the nodes corresponding to different indicated colors belong to different sequence classes.
[0140] The Figure 4 In the graph model marked by 40 in, when the user relationship type of the user group corresponding to the nodes is the same relationship type, the filled colors marked are the same. Then, the node sequence obtained by the computer device sorting the nodes in the graph model can be as shown in Figure 4As shown by the sequence marked by 41, when the computer device sorts the nodes based on the relationship type to obtain the corresponding node sequence, it will not consider the arrangement order of the nodes of the user groups corresponding to different types of user relationships, nor the arrangement order of the nodes corresponding to the user groups under the user relationships of the same type.
[0141] In one embodiment, when the computer device calculates the category transition probability based on the occurrence probability corresponding to the node and the weight of the edge connecting the two nodes with an edge, assuming that the node is among the two nodes (α, β) of the edge, the occurrence probability corresponding to node α is p α , and the occurrence probability corresponding to node β is p β , then since both of the two nodes (α, β) with an edge have an edge β→α, and the weight of this edge is the corresponding transition probability p β→α , the computer determines that there must be formula (4).
[0142] p β =∑ α p α *p α→β Formula (4)
[0143] If there are some isolated regions in the graph model constructed by the computer, then after the computer device randomly walks to these isolated regions, it cannot jump out, resulting in the probability that the nodes outside this region are passed through being 0. To avoid this situation, the computer device can perform random walks according to the probability of 1-τ and in combination with the probability of p β→α . Among them, τ is the crossing probability, which is a hyperparameter introduced externally, and its value range is 0 to 1. Then formula (4) can be rewritten as formula (5).
[0144]
[0145] Among them, n is the total number of nodes included in the i-th sequence class. Then, from formulas (4) and (5), it can be seen that the relationship between the transition probability corresponding to the node is as shown in formula (6).
[0146] p α→β =(1-τ)p β→α +τ / n Formula (6)
[0147] In one embodiment, to determine p α and p β , the computer device can further assume that the sequence class where node α is located is i, and then can use to represent the occurrence probability of sequence class i. Since the occurrence of each sequence class must end with the termination mark of this sequence class, so is also the occurrence probability of the termination marker of category i, then there exists the above probability, and there exists the relationship shown in Equation (7).
[0148]
[0149] And appearance means that the sequence class where the next node is located is not the i-th class. Therefore which is also the probability that the event of leaving the i-th class occurs. Therefore, it can be denoted as Equation (8).
[0150]
[0151] Then, based on Equations (4) to (8), the computer device can calculate p α 、p β and It can be understood that according to the above steps, the computer device can obtain the occurrence probability corresponding to each node in the node sequence and calculate the category transition probability, so that the computer device can determine the total average coding length corresponding to the node sequence based on the result of the hierarchical coding of each sequence class in the node sequence. In one embodiment, when the computer device performs hierarchical coding on the node sequence, it can first perform category coding on the sequence classes included in the node sequence, where the category coding of different sequence classes is different; and determine the sequence class to which the corresponding node belongs, and perform node coding on the corresponding node according to the probability corresponding to each node in the sequence class, where the length of the node coding corresponding to each node is negatively correlated with the probability of being passed by during the random walk; then, the computer device can use the node coding of each node and the category coding of the corresponding sequence class as the coding of each node.
[0152] When a computer device performs a random walk, it can start from a certain point j, jump to the next point i with reference to the probability p(i|j), and then start jumping sequentially from i. By repeating this process, the computer device can determine the number of times each node is visited, and then statistically obtain the probability of each node being visited. Thus, Huffman coding can be performed based on the probability of each node being visited during the random walk, such that the length of the node code corresponding to each node is negatively correlated with the probability of the corresponding node being visited during the random walk. When the computer device performs hierarchical coding on the node sequence, a class label can be inserted before the relationship pairs in the same sequence class, and a termination label can be inserted at the end of the class. The class label uses a separate set of codes, such as 000, 001, 002 to represent, and the relationship pairs within the class and the termination label are represented by another set of codes. Since class labels are considered, the relationship pairs of different user relationship classes also use the same set of codes, such as 000, 001, 010, 011, 100 to represent, enabling the computer device to cluster object groups by minimizing the total average coding length.
[0153] In one embodiment, when the computer device determines the total average coding length of the node sequence according to the hierarchical coding result, it can obtain the average coding length corresponding to each sequence class in the node sequence and the average coding length of each node in each sequence class according to the hierarchical coding result. Thus, the average coding lengths corresponding to each class and the average coding lengths of the nodes in each sequence class can be weighted and averaged, and the weighted average coding length is used as the total average coding length of the node sequence.
[0154] Since the computer device uses two different coding methods for sequence classes and the objects (i.e., nodes) within the sequence classes, the shortest average coding lengths of both can be calculated separately. Among them, the shortest average coding length of the sequence class (i.e., the class average coding length) can be as shown in Equation (9).
[0155]
[0156] Among them, After normalizing the probabilities of the sequence classes among different sequence classes, the theoretically shortest average coding length of the sequence classes can be obtained. Similarly, the shortest average coding length of the objects within each sequence class i can be as shown in Equation (10).
[0157]
[0158] Among them, Therefore, after the encoding adopted by the objects in each sequence class, together with the termination marking probability of each sequence class, is normalized within each sequence, the average encoding length of the intra-class objects in each sequence class can be obtained. Furthermore, the computer device can perform a weighted average on the average encoding lengths obtained from the above formulas (9) and (10) to obtain the total average encoding length, and the obtained total average encoding length can be as shown in formula (11).
[0159]
[0160] Correspondingly, after the computer device obtains the node sequence (or object sequence), it can search for a target hierarchical encoding scheme so that formula (11) can achieve the minimum value, thereby realizing the clustering of nodes. It can be understood that the process of minimizing formula (11) is the process of continuously optimizing the sequence classes divided in the node sequence, that is, the process of optimizing the clustering of different object groups. Then, after the computer device completes the above optimization clustering process, the computer device can determine the target sequence class where the target node is located, and then use the user group corresponding to the nodes included in the target sequence class as the user group included in the target community. As Figure 5 shown, if the computer device determines that the target node corresponding to the target user group is the node marked by 50 as shown in Figure 5 , then based on the target sequence group where the target node is located, it can be determined that the target community where the target user group is located can be the set marked by 501 as shown in Figure 5 .
[0161] S305. Determine the types of user relationships among the users included in each reference user group in the target community, and determine the proportion of the number of reference user groups corresponding to the user relationships of the same type.
[0162] S306. Use the user relationship corresponding to the reference user group with the largest proportion in the target community as the target user relationship among the users in the target user group.
[0163] In steps S305 and S306, after the computer device determines the target community where the target user group is located, it can determine the number of reference user groups corresponding to the user relationships of the same type based on the types of user relationships among the users included in each reference user group in the target community. Furthermore, it can determine the proportion of the number of reference user groups corresponding to the user relationships of each type. If it is assumed that in the target community determined by the computer device, the proportion of the number of reference user groups corresponding to the a user relationship in the target community is 53%, the proportion of the number of reference user groups corresponding to the b user relationship in the target community is 37%, and the proportion of the number of reference user groups corresponding to other user relationships in the target community is 10%, then the computer device can use the user relationship corresponding to the reference user group with the largest proportion as the user relationship among the users in the target user group, and thus determine that the target user relationship among the users in the target user group is the a user relationship.
[0164] In one embodiment, after the computer device constructs the above-mentioned graph model, when it is necessary to add the nodes corresponding to the new user group to the graph model to update the graph model, the computer device can adjust the similarity threshold of the edges between the nodes corresponding to the user groups based on the step adjustment relationship, and gradually adjust the similarity threshold S mentioned in step S304 by ±0.05, so as to reclassify the high target value area in the initial target value. Furthermore, based on the new clustering result and according to the proportion of the number of user groups with known user relationships in each community obtained by clustering, it determines the types of users in the user group with unknown user relationships. Based on the determined target community to which each user group belongs, the computer device can apply it to information recommendation. In a specific implementation, the computer device can first obtain the information to be recommended and the historical user groups interested in the information to be recommended, and determine the community to which the historical user groups belong; then, further, the computer device can send the information to be recommended to each user in the community to which the historical user groups belong, so as to achieve accurate recommendation of the information to be recommended.
[0165] In an embodiment of the present invention, after a computer device obtains a plurality of relationship features and reference features of a reference user group, when it is necessary to determine the user relationship between users in a target user group to be processed, based on the target behavior feature corresponding to the user group, it searches for a target relationship feature that matches the target behavior feature, and then constructs a target feature corresponding to the target user group. Then, the computer device can perform a clustering operation based on the target feature and a plurality of reference features, so as to obtain a target community to which the target user group belongs, and use the user relationship corresponding to the reference user group with the largest proportion in the target community as the target user relationship between users in the target user group. This enables the computer device to construct the behavior features and relationship features of the user group for building user relationships, construct transition probabilities, generate sequences based on random walks on the graph model, construct the association relationship between the features of the user group, and predict the relationship type of the target user group based on relationship analysis, thereby improving the generalization ability of the computer device when determining the user relationship in the target user group, and having wide application value and reference significance in the field of social network construction.
[0166] Based on the description of the embodiments of the above relationship analysis method, an embodiment of the present invention also proposes a relationship analysis device, which can be a computer program (including program code) running in the above computer device. The relationship analysis device can be used to execute as Figure 2 and Figure 3 the relationship analysis method described, please refer to Figure 6 , and the relationship analysis device includes: an acquisition unit 601, a determination unit 602, and a division unit 603.
[0167] The acquisition unit 601 is used to acquire M relationship features and reference features of N reference user groups. One relationship feature is used to represent a type of user relationship; one reference user group corresponds to one reference feature; the reference feature of any one of the N reference user groups is used to indicate: the interaction behavior order between users in the any one of the reference user groups, and both N and M are integers greater than 1;
[0168] The acquisition unit 601 is further used to acquire the target behavior feature of the target user group to be processed, and the target behavior feature is used to indicate the interaction behavior order between users in the target user group;
[0169] The determination unit 602 is used to determine a target relationship feature that matches the target behavior feature from the M relationship features, and construct a target feature of the target user group according to the target relationship feature and the target behavior feature;
[0170] A dividing unit 603, configured to perform relationship analysis on the target user group and the N reference user groups according to the target feature and the N reference features, and determine a target community to which the target user group belongs;
[0171] The determining unit 602 is further configured to determine a target user relationship among users in the target user group according to user relationships among users included in each reference user group in the target community.
[0172] In one embodiment, both the target behavior feature and the relationship feature are represented by an identification sequence. The target behavior identification sequence representing the target behavior feature and the relationship identification sequence representing the relationship feature both include at least one behavior identification. One behavior identification corresponds to one interaction behavior, and the arrangement order of the at least one behavior identification in the corresponding identification sequence is used to indicate the execution order of the corresponding interaction behavior. The determining unit 602 is specifically configured to:
[0173] Obtain a relationship identification sequence corresponding to each relationship feature among the M relationship features, and a target behavior identification sequence corresponding to the target behavior feature;
[0174] From the relationship identification sequences respectively corresponding to the M relationship features, find sequences that are the same as at least two target behavior identifications in the target behavior identification sequence and have the same corresponding arrangement order, and use the relationship features corresponding to the found sequences as target relationship features that match the target behavior feature.
[0175] In one embodiment, the determining unit 602 is specifically configured to:
[0176] Obtain a target portrait of each user included in the target user group, and perform feature mapping on the target portrait of each user included in the target user group respectively to obtain portrait features of each user included in the target user group;
[0177] Perform feature splicing on the portrait features of each user in the target user group, the target behavior feature, and the target relationship feature to obtain a target feature of the target user group.
[0178] In one embodiment, the dividing unit 603 is specifically configured to:
[0179] Construct a graph model according to the target feature and each reference feature. The graph model includes a target node for indicating the target user group and multiple other nodes. One other node is used to indicate one reference user group, and the weight of the edge existing between two nodes is determined according to the similarity between the features of the corresponding user groups;
[0180] Using the target community discovery algorithm, perform node clustering operations on the nodes in the graph model to determine the target node set where the target node is located. The user groups corresponding to the nodes included in the target node set constitute the target community to which the target user group belongs.
[0181] In one embodiment, each node in the graph model is associated with the characteristics of the corresponding user group; the determining unit 602 is specifically configured to:
[0182] Based on the random walk algorithm, perform random walks in the graph model, and sort the nodes in the graph model according to the order of the nodes passed through during the random walk in the graph model to obtain a node sequence, where the node sequence includes at least one sequence class;
[0183] Perform hierarchical encoding on the node sequence, where one sequence class included in the node sequence corresponds to one category encoding, and the nodes in one sequence class correspond to one node encoding;
[0184] According to the hierarchical encoding result, determine the total average encoding length of the node sequence, and when the total average encoding length obtains the minimum value, obtain the target sequence class where the target node is located, and use the nodes included in the target sequence class as the nodes in the target node set where the target node is located.
[0185] In one embodiment, the obtaining unit 601 is further configured to obtain the occurrence probability corresponding to each node included in the node sequence, and the weight of the edge connecting two nodes with an edge;
[0186] The determining unit 602 is further configured to calculate the category transition probability according to the occurrence probability corresponding to each node and the weight of the edge connecting two nodes with an edge;
[0187] The partitioning unit 603 is further configured to partition the node sequence into at least one sequence class according to the category transition probability, and each sequence class includes one or more nodes.
[0188] In one embodiment, the determining unit 602 is specifically configured to:
[0189] Perform category encoding on the sequence classes included in the node sequence, where the category encodings of different sequence classes are different;
[0190] Determine the sequence class to which the corresponding node belongs, and perform node encoding on the corresponding node according to the probability corresponding to each node in the sequence class, where the length of the node encoding corresponding to each node is negatively correlated with the probability that the corresponding node is passed through during the random walk;
[0191] Use the node encoding of each node and the category encoding of the corresponding sequence class as the encoding of each node.
[0192] In one embodiment, the determining unit 602 is specifically configured to:
[0193] According to the hierarchical encoding result, obtain the average encoding length corresponding to each sequence class in the node sequence, and the average encoding length of each node in each sequence class;
[0194] Perform a weighted average on the average encoding lengths corresponding to each class and the average encoding lengths corresponding to each node in each sequence class, and use the weighted average encoding length as the total average encoding length of the node sequence.
[0195] In one embodiment, the determining unit 602 is specifically configured to:
[0196] Determine the type of user relationship among the users included in each reference user group in the target community, and determine the proportion of the number of reference user groups corresponding to the user relationship of the same type;
[0197] Use the user relationship corresponding to the reference user group with the largest proportion in the target community as the target user relationship among the users in the target user group.
[0198] In one embodiment, the obtaining unit 601 is specifically configured to:
[0199] Obtain the behavior characteristics of each reference user group under the user relationship of the target type, and use the behavior identification sequence corresponding to the behavior characteristics corresponding to each reference user group as the target sequence set, and determine the number of sequences in the target sequence set, the included behavior identifications, and the occurrence times corresponding to each behavior identification;
[0200] According to the occurrence times of each behavior identification in the target sequence set, select multiple one-item prefixes from the target sequence set, and each one-item prefix includes: the behavior identification with an occurrence time greater than the quantity threshold in the target sequence set;
[0201] Construct sequence patterns using each one-item prefix respectively, and obtain the projection data set of each one-item prefix, where the projection data set contains the suffixes of the corresponding one-item prefix in the corresponding interaction behavior sequence, and the suffixes include the behavior identifications in the corresponding interaction behavior sequence that are after the prefix;
[0202] Perform recursive mining on the projection data sets of each one-item prefix to obtain K-item prefixes, and use the K-item prefixes to construct sequences respectively to obtain multiple behavior sequence patterns, and the obtained behavior sequence patterns are used to represent the target relationship characteristics corresponding to the user relationship of the target type, where K is an integer greater than 1.
[0203] In one embodiment, the device further includes a sending unit 604.
[0204] The obtaining unit 601 is further configured to obtain information to be recommended, and a historical user group interested in the information to be recommended, and determine the community to which the historical user group belongs;
[0205] The sending unit 604 is configured to send the information to be recommended to each user in the community to which the historical user group belongs.
[0206] In the embodiment of the present invention, after the obtaining unit 601 obtains the behavior characteristics corresponding to each reference user group in the N reference user groups, a plurality of relationship characteristics can be obtained by mining the behavior characteristics of the reference user groups with the same user relationship. Then, the obtaining unit 601 can obtain the reference characteristics of each reference user group based on the obtained relationship characteristics and the behavior characteristics of each reference user group. Moreover, the obtaining unit 601 can also match the target behavior characteristics of the target user group with the unknown social type with the obtained plurality of relationship characteristics, and then match the target relationship characteristics, so that the determining unit 602 can construct the target characteristics of the target user group based on the matched target relationship characteristics and target behavior characteristics, so that the multi-dimensional characteristics of the user group can be constructed, and the accuracy of the characteristics corresponding to the obtained user group is improved. After the obtaining unit 601 obtains the reference characteristics corresponding to each reference user group and the target characteristics corresponding to the target user group, the relationship between the target user group and the reference user group can be analyzed based on the target characteristics and the reference characteristics, so that the dividing unit 603 can determine the target user relationship between the users in the target user group based on the user relationships corresponding to the other reference user groups in the target community to which the target user group is assigned. The generalization ability of the method for determining the user relationship based on the relationship analysis is strong, and since the accuracy of the characteristics corresponding to the obtained user group is relatively high, the method for analyzing the user relationship between users by the relationship analysis is also helpful to improve the accuracy when determining the user relationship between users.
[0207] Please refer to Figure 7, which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present invention. Among them, the computer device can be a server, and the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; or, the computer device can also be a terminal device, and the terminal device can be, for example, a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. As Figure 7 shown, the computer device in this embodiment may include: one or more processors 701; one or more input devices 702, one or more output devices 703, and a memory 704. The above-mentioned processors 701, input devices 702, output devices 703, and memory 704 are connected through a bus 705. The memory 704 is used to store a computer program, and the computer program includes program instructions. The processor 701 is used to execute the program instructions stored in the memory 704.
[0208] The memory 704 may include a volatile memory, such as a random-access memory (RAM); the memory 704 may also include a non-volatile memory, such as a flash memory, a solid-state drive (SSD), etc.; the memory 704 may further include a combination of the above types of memories.
[0209] The processor 701 may be a central processing unit (CPU). The processor 701 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), etc. The PLD may be a field-programmable gate array (FPGA), a generic array logic (GAL), etc. The processor 701 may also be a combination of the above structures.
[0210] In an embodiment of the present invention, the memory 704 is used to store a computer program, and the computer program includes program instructions. The processor 701 is used to execute the program instructions stored in the memory 704 to implement the steps of the corresponding method as described above Figure 2 and Figure 3 in the corresponding method.
[0211] In one embodiment, the processor 701 is configured to call the program instructions for execution:
[0212] Obtain M relationship features and reference features of N reference user groups. One relationship feature is used to represent a type of user relationship; one reference user group corresponds to one reference feature; the reference feature of any one of the N reference user groups is used to indicate: the interaction behavior order among users in the any one of the reference user groups. Both N and M are integers greater than 1;
[0213] Obtain the target behavior feature of the target user group to be processed, where the target behavior feature is used to indicate the interaction behavior order among users in the target user group;
[0214] Determine a target relationship feature that matches the target behavior feature from the M relationship features, and construct a target feature of the target user group according to the target relationship feature and the target behavior feature;
[0215] Perform relationship analysis on the target user group and the N reference user groups according to the target feature and the N reference features, determine the target community to which the target user group belongs, and determine the target user relationship among users in the target user group according to the user relationships among users included in each reference user group in the target community.
[0216] In one embodiment, both the target behavior feature and the relationship feature are represented by an identification sequence. The target behavior identification sequence representing the target behavior feature and the relationship identification sequence representing the relationship feature both include at least one behavior identification. One behavior identification corresponds to one interaction behavior. The arrangement order of the at least one behavior identification in the corresponding identification sequence is used to indicate the execution order of the corresponding interaction behavior; the processor 701 is configured to call the program instructions for execution:
[0217] Obtain the relationship identification sequence corresponding to each relationship feature among the M relationship features, and the target behavior identification sequence corresponding to the target behavior feature;
[0218] From the relationship identification sequences respectively corresponding to the M relationship features, find the sequences that are the same as at least two target behavior identifications in the target behavior identification sequence and have the same corresponding arrangement order, and use the relationship features corresponding to the found sequences as the target relationship features that match the target behavior features.
[0219] In one embodiment, the processor 701 is configured to call the program instructions for performing:
[0220] Obtain the target portraits of each user included in the target user group, and perform feature mapping on the target portraits of each user included in the target user group respectively to obtain the portrait features of each user included in the target user group;
[0221] Perform feature splicing on the portrait features of each user in the target user group, the target behavior features, and the target relationship features to obtain the target features of the target user group.
[0222] In one embodiment, the processor 701 is configured to call the program instructions for performing:
[0223] Construct a graph model according to the target features and each reference feature, where the graph model includes a target node for indicating the target user group and a plurality of other nodes, and the weight of the edge existing between two nodes is determined according to the similarity between the features of the corresponding user groups;
[0224] Adopt a target community discovery algorithm to perform node clustering operation on the nodes in the graph model, and determine the target node set where the target node is located. The user groups corresponding to the nodes included in the target node set constitute the target community to which the target user group belongs.
[0225] In one embodiment, each node in the graph model is associated with the features of the corresponding user group; the processor 701 is configured to call the program instructions for performing:
[0226] Based on the random walk algorithm, perform random walk in the graph model, and sort the nodes in the graph model according to the order corresponding to the nodes passed by the random walk in the graph model to obtain a node sequence, where the node sequence includes at least one sequence class;
[0227] Perform hierarchical encoding on the node sequence, where one sequence class included in the node sequence corresponds to one category encoding, and the nodes in one sequence class correspond to one node encoding;
[0228] Determine the total average coding length of the node sequence according to the hierarchical coding result. When the total average coding length reaches the minimum value, obtain the target sequence class where the target node is located, and use the nodes included in the target sequence class as the nodes in the target node set where the target node is located.
[0229] In one embodiment, the processor 701 is configured to call the program instructions to perform:
[0230] Obtain the occurrence probability corresponding to each node included in the node sequence, and the weight of the edge connecting two nodes with an edge;
[0231] Calculate the class transition probability according to the occurrence probability corresponding to each node and the weight of the edge connecting two nodes with an edge;
[0232] Divide the node sequence into at least one sequence class according to the class transition probability, and each sequence class includes one or more nodes.
[0233] In one embodiment, the processor 701 is configured to call the program instructions to perform:
[0234] Perform class coding on the sequence classes included in the node sequence, where the class codings of different sequence classes are different;
[0235] Determine the sequence class to which the corresponding node belongs, and perform node coding on the corresponding node according to the probability corresponding to each node in the sequence class, where the length of the node coding corresponding to each node is negatively correlated with the probability of being passed through during random walk;
[0236] Use the node coding of each node and the class coding of the corresponding sequence class as the coding of each node.
[0237] In one embodiment, the processor 701 is configured to call the program instructions to perform:
[0238] According to the hierarchical coding result, obtain the average coding length corresponding to each sequence class in the node sequence, and the average coding length of each node in each sequence class;
[0239] Perform weighted averaging on the average coding lengths corresponding to each class and the average coding lengths of the nodes in each sequence class, and use the weighted average coding length as the total average coding length of the node sequence.
[0240] In one embodiment, the processor 701 is configured to call the program instructions to perform:
[0241] Determine the types of user relationships among the users included in each reference user group in the target community, and determine the proportion of the number of reference user groups corresponding to the user relationships of the same type;
[0242] In the target community, use the user relationship corresponding to the reference user group with the largest proportion of the number as the target user relationship among the users in the target user group.
[0243] In one embodiment, the processor 701 is configured to call the program instructions to perform:
[0244] Obtain the behavior characteristics of each reference user group under the user relationship of the target type, and use the behavior identification sequence corresponding to the behavior characteristics corresponding to each reference user group as the target sequence set, and determine the number of sequences in the target sequence set, the included behavior identifications, and the occurrence times corresponding to each behavior identification;
[0245] According to the occurrence times of each behavior identification in the target sequence set, select multiple one-item prefixes from the target sequence set, and each one-item prefix includes: the behavior identification with the occurrence times greater than the quantity threshold in the target sequence set;
[0246] Use each one-item prefix to construct sequence patterns respectively, and obtain the projection data set of each one-item prefix. The projection data set contains the suffixes of the corresponding one-item prefix in the corresponding interaction behavior sequence, and the suffixes include the behavior identifications after the prefix in the corresponding interaction behavior sequence;
[0247] Perform recursive mining on the projection data sets of each one-item prefix to obtain K-item prefixes, and use the K-item prefixes to construct sequences respectively to obtain multiple behavior sequence patterns. The obtained behavior sequence patterns are used to represent the target relationship characteristics corresponding to the user relationship of the target type, where K is an integer greater than 1.
[0248] In one embodiment, the processor 701 is configured to call the program instructions to perform:
[0249] Obtain the information to be recommended and the historical user groups interested in the information to be recommended, and determine the community to which the historical user groups belong;
[0250] Send the information to be recommended to each user in the community to which the historical user groups belong.
[0251] An embodiment of the present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above asFigure 2 or Figure 3 The method embodiments shown. Among them, the computer-readable storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), etc.
[0252] The above-disclosed are only partial embodiments of the present invention. Of course, the scope of rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand the entire or partial processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A relationship analysis method, characterized in that, Including: Obtain M relationship features and reference features of N reference user groups, where one relationship feature is used to represent one type of user relationship; One reference user group corresponds to one reference feature; the reference feature of any one of the N reference user groups is used to indicate: the interaction behavior order among users in the any one of the reference user groups, and both N and M are integers greater than 1; the reference feature is constructed from reference behavior features and relationship features matching the reference behavior features; Obtain the target behavior features of the target user group to be processed, where the target behavior features are used to indicate the interaction behavior order among users in the target user group; Determine a target relationship feature matching the target behavior features from the M relationship features, and construct the target features of the target user group according to the target relationship feature and the target behavior features; According to the target features and the N reference features, perform relationship analysis on the target user group and the N reference user groups, determine the target community to which the target user group belongs, and determine the target user relationships among users in the target user group according to the user relationships among users included in each reference user group in the target community.
2. The method according to claim 1, characterized in that Both the target behavior features and the relationship features are represented by identification sequences. The target behavior identification sequence representing the target behavior features and the relationship identification sequence representing the relationship features both include at least one behavior identification. One behavior identification corresponds to one interaction behavior, and the arrangement order of the at least one behavior identification in the corresponding identification sequence is used to indicate the execution order of the corresponding interaction behavior; The determining a target relationship feature matching the target behavior features from the M relationship features includes: Obtain the relationship identification sequence corresponding to each relationship feature among the M relationship features, and the target behavior identification sequence corresponding to the target behavior features; From the relationship identification sequences respectively corresponding to the M relationship features, find a sequence with at least two target behavior identifications the same as those in the target behavior identification sequence and the corresponding arrangement order the same, and use the relationship feature corresponding to the found sequence as the target relationship feature matching the target behavior features.
3. The method according to claim 1, wherein The constructing the target features of the target user group according to the target relationship feature and the target behavior features includes: Obtain the target portraits of each user included in the target user group, and perform feature mapping on the target portraits of each user included in the target user group respectively to obtain the portrait features of each user included in the target user group; Perform feature splicing on the portrait features of each user in the target user group, the target behavior features, and the target relationship features to obtain the target features of the target user group.
4. The method according to claim 1, characterized in that, The performing relationship analysis on the target user group and the N reference user groups according to the target features and the N reference features to determine the target community to which the target user group belongs includes: Construct a graph model according to the target feature and each reference feature. The graph model includes a target node for indicating a target user group and a plurality of other nodes. One other node is used to indicate a reference user group. The weight of the edge existing between two nodes is determined according to the similarity between the features of the corresponding user groups. Adopt a target community discovery algorithm to perform node clustering operation on the nodes in the graph model, and determine a target node set where the target node is located. The user groups corresponding to the nodes included in the target node set constitute the target community to which the target user group belongs.
5. The method according to claim 4, characterized in that, Each node in the graph model is associated with the feature of the corresponding user group. The adopting a target community discovery algorithm to perform node clustering operation on the nodes in the graph model and determine the target node set where the target node is located includes: Based on the random walk algorithm, perform random walk in the graph model, and sort the nodes in the graph model according to the order corresponding to the nodes passed by during the random walk in the graph model to obtain a node sequence. Among them, the node sequence includes at least one sequence class. Perform hierarchical encoding on the node sequence. Among them, one sequence class included in the node sequence corresponds to a category encoding, and the nodes in one sequence class correspond to a node encoding. According to the hierarchical encoding result, determine the total average encoding length of the node sequence. When the total average encoding length obtains the minimum value, obtain the target sequence class where the target node is located, and use the nodes included in the target sequence class as the nodes in the target node set where the target node is located.
6. The method according to claim 5, wherein The method further includes: Obtain the occurrence probability corresponding to each node included in the node sequence, and the weight of the edge corresponding to the two nodes with an edge. Calculate the category transition probability according to the occurrence probability corresponding to each node and the weight of the edge corresponding to the two nodes with an edge. Divide the node sequence into at least one sequence class according to the category transition probability. Each sequence class includes one or more nodes.
7. The method according to claim 5, wherein The performing hierarchical encoding on the node sequence includes: Perform category encoding on the sequence classes included in the node sequence. Among them, the category encodings of different sequence classes are different. Determine the sequence class to which the corresponding node belongs, and perform node encoding on the corresponding node according to the probability corresponding to each node in the sequence class. Among them, the length of the node encoding corresponding to each node is negatively correlated with the probability of being passed by during the random walk of the corresponding node. Use the node encoding of each node and the category encoding of the corresponding sequence class as the encoding of each node.
8. The method according to claim 5, wherein The determining the total average encoding length of the node sequence according to the hierarchical encoding result includes: According to the hierarchical encoding result, obtain the average encoding length corresponding to each sequence class in the node sequence, and the average encoding length of each node in each sequence class. Perform weighted average on the average encoding lengths corresponding to each class and the average encoding lengths of the nodes in each sequence class, and use the weighted average encoding length as the total average encoding length of the node sequence.
9. The method according to claim 1, wherein Determining the target user relationships among the users in the target user group according to the user relationships among the users included in each reference user group in the target community includes: Determining the types of user relationships among the users included in each reference user group in the target community, and determining the proportion of the number of reference user groups corresponding to the user relationships of the same type; Taking the user relationship corresponding to the reference user group with the largest proportion of the number in the target community as the target user relationship among the users in the target user group.
10. The method according to claim 1, characterized in that, The M relationship features include the relationship features corresponding to the user relationships of the target type. The method for obtaining the relationship features corresponding to the user relationships of the target type includes: Obtaining the behavior features of each reference user group under the user relationship of the target type, and taking the behavior identification sequence corresponding to the behavior feature corresponding to each reference user group as the target sequence set, and determining the number of sequences in the target sequence set, the behavior identifications included, and the occurrence times corresponding to each behavior identification; Selecting a plurality of one-item prefixes from the target sequence set according to the occurrence times of the behavior identifications in the target sequence set. Each one-item prefix includes: the behavior identifications with occurrence times greater than the quantity threshold in the target sequence set; Constructing sequence patterns using each one-item prefix respectively, and obtaining the projection data set of each one-item prefix. The projection data set contains the suffixes of the corresponding one-item prefix in the corresponding interaction behavior sequence, and the suffixes include the behavior identifications located after the prefix in the corresponding interaction behavior sequence; Performing recursive mining on the projection data sets of the one-item prefixes respectively to obtain K-item prefixes, and constructing sequences using the K-item prefixes respectively to obtain a plurality of behavior sequence patterns. The obtained behavior sequence patterns are used to represent the target relationship features corresponding to the user relationships of the target type, where K is an integer greater than 1.
11. The method according to claim 1, characterized in that, The method further includes: Obtaining the information to be recommended, and the historical user group interested in the information to be recommended, and determining the community to which the historical user group belongs; Sending the information to be recommended to each user in the community to which the historical user group belongs.
12. A relationship analysis device, characterized in that, Including: An obtaining unit, configured to obtain M relationship features and the reference features of N reference user groups. One relationship feature is used to represent one type of user relationship; One reference user group corresponds to one reference feature; the reference feature of any one of the N reference user groups is used to indicate: the interaction behavior order among the users in the any one of the reference user groups. Both N and M are integers greater than 1; the reference feature is constructed from the reference behavior feature and the relationship feature matching the reference behavior feature; The obtaining unit is further configured to obtain the target behavior feature of the target user group to be processed, and the target behavior feature is used to indicate the interaction behavior order among the users in the target user group; A determining unit, configured to determine the target relationship feature matching the target behavior feature from the M relationship features, and construct the target feature of the target user group according to the target relationship feature and the target behavior feature; A division unit, configured to perform relationship analysis on the target user group and the N reference user groups according to the target feature and the N reference features, and determine a target community to which the target user group belongs; The determination unit is further configured to determine a target user relationship among users in the target user group according to user relationships among users included in each reference user group in the target community.
13. A computer device, characterized in that, It includes a processor, an input device, an output device, and a memory, and the processor, the input device, the output device, and the memory are interconnected. Among them, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 11.
15. A computer program product, characterized in that, The computer program product includes a computer program, the computer program includes program instructions, and when the program instructions are called by a processor, the processor is caused to execute the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Invitation behavior prediction method and device
CN108875993A
Network community management method, terminal equipment and storage medium
CN111966975A