Target user identification method and apparatus, and electronic device

By grouping and similarity calculations of user behavior data, identifying target users is solved, and the problem of accurately identifying users in network information security is improved, and identification efficiency and security are improved.

WO2025139642A1PCT designated stage expired Publication Date: 2025-07-03CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/136464
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-26
Filing Date
2024-12-03
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

How to accurately identify target users to improve network information security and prevent bad organizations or individuals from exploiting website vulnerabilities to engage in destructive activities.

Method used

By grouping user behavior data based on preset time periods and device information, identifying similar behavior information pairs, calculating user behavior similarity, and establishing a similarity correlation link to identify the highest level user as the target user.

Benefits of technology

It improves the accuracy and efficiency of target user identification, enhances network information security, supports fast and efficient identification of target users, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136464_03072025_PF_FP_ABST
    Figure CN2024136464_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers, in particular to a target user identification method and apparatus, and an electronic device. The method comprises: on the basis of a preset first time period and device information, grouping a plurality of pieces of user behavior data of a plurality of different devices in a set second time period according to a preset rule; on the basis of the user behavior data, respectively determining a plurality of similar behavior information pairs in the plurality of data groups; on the basis of the number of similar behavior information pairs between two users and the volume of user behavior data, calculating a user behavior similarity, and associating users having user behavior similarities not less than a set similarity threshold value to obtain similarity association links; and comparing the number of users in the similarity association links with a preset rank threshold value to determine the ranks of the similarity association links, and determining users in a similarity association link having the highest rank as target users. The solution can identify target users accurately, thus improving the security of network information.
Need to check novelty before this filing date? Find Prior Art

Description

Target user identification method, device and electronic equipment

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 26, 2023, with application number 202311802065.3 and application name "A target user identification method, device and electronic device", the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of computer technology, and in particular to a target user identification method, device, and electronic device. Background Art

[0004] With the rapid development of internet technology, the importance of network information security has become increasingly prominent. Organizations or individuals with malicious intent often exploit website vulnerabilities to disrupt internet order and seek illicit gains. This poses a significant threat to network data and property security.

[0005] How to accurately identify target users and effectively improve network information security has become a question worth discussing. Summary of the Invention

[0006] The embodiments of the present application provide a target user identification method, device, and electronic device for accurately identifying target users and effectively improving network information security.

[0007] In a first aspect, an embodiment of the present application provides a target user identification method, comprising:

[0008] Based on a preset first time period and device information, multiple pieces of user behavior data from multiple different devices during a preset second time period are grouped according to preset rules to generate multiple data groups. The user behavior data includes behavior object information, behavior type information, and behavior time information, and the duration of the first time period is shorter than the duration of the second time period. Multiple similar behavior information pairs are determined in the multiple data groups based on the user behavior data. Each similar behavior information pair includes two pieces of user behavior data with the same behavior object information and the same behavior type information, and a data time span that is less than or equal to a preset time span. The data time span represents the time span between the behavior time information of the two pieces of user behavior data. A user behavior similarity, representing the correlation between user behaviors, is calculated based on the number of similar behavior information pairs between the two users and the amount of user behavior data for the two users. Users whose user behavior similarity is not less than a preset similarity threshold are then associated to generate a similarity association link. The similarity association link includes multiple nodes, which are connected based on the order of the behavior time information, and each node represents a user. The number of users in the similarity association link is compared with a preset level threshold to determine the level of the similarity association link, and the user in the similarity association link with the highest level is determined as the target user.

[0009] The above method, based on calculating user behavior similarity, can accurately determine the correlation between different user behaviors, thereby improving the accuracy and efficiency of identifying target users and enhancing network information security. It supports calculating user behavior similarity based on multiple user behavior data. This allows for timely identification of target users based on user behavior similarity at any point in time, quickly and efficiently meeting business needs and improving the user experience.

[0010] Optionally, the above-mentioned method of grouping multiple pieces of user behavior data of multiple devices in a set second time period according to preset rules based on the preset first time period and device information to obtain multiple data groups specifically includes:

[0011] Sort the multiple user behavior data in chronological order, and group them based on a preset first time period to obtain multiple time window data segments, each time window data segment including multiple user behavior data;

[0012] Multiple pieces of user behavior data with the same device information in the time window data segment are divided into one group to obtain multiple data groups.

[0013] In the above method, multiple pieces of user behavior data are sorted chronologically and grouped based on preset time periods to obtain multiple time window data segments, thereby enabling initial correlation of the multiple pieces of user behavior data in chronological order. Furthermore, by grouping multiple pieces of user behavior data with identical device information within the time window data segments into a single group to obtain multiple data groups, multiple pieces of user behavior data initiated based on the same device within the preset time period can be further identified, thereby enhancing the correlation of the multiple pieces of user behavior data within the data group and facilitating subsequent calculation of user behavior similarity.

[0014] Optionally, the user behavior data further includes user information. Before sorting the plurality of user behavior data in chronological order, the method further includes:

[0015] Data cleaning is performed on multiple pieces of user behavior data to obtain multiple pieces of cleaned user behavior data, wherein any two pieces of user behavior data in the multiple pieces of cleaned user behavior data are different, and any piece of user behavior data includes behavior object information, behavior type information, behavior time information, user information, and device information.

[0016] In the above method, by performing data cleaning on multiple user behavior data to obtain multiple cleaned user behavior data, it is convenient to save computing resources, prevent waste of computing resources, and improve computing efficiency when subsequently calculating user behavior similarity based on multiple user behavior data.

[0017] Optionally, after grouping the plurality of pieces of user behavior data of the plurality of different devices in the set second time period according to a preset rule to obtain a plurality of data groups based on the preset first time period and the device information, the method further includes:

[0018] Determine the number of users for each data group;

[0019] When the number of users is less than a preset effective threshold, the data group with the number of users less than the effective threshold is deleted.

[0020] In the above method, by deleting data groups with fewer than the valid threshold number of users, the remaining data groups can be made valid. This can save computing resources, prevent waste of computing resources, and improve computing efficiency when subsequently calculating user behavior similarity based on multiple user behavior data in the data group.

[0021] Optionally, before determining a plurality of similar behavior information pairs in a plurality of data groups based on the user behavior data, the method further includes:

[0022] Multiple pieces of user behavior data with the same user information, behavior type information, and behavior object information in the same data group are merged into one piece of user behavior data according to preset rules.

[0023] In the above method, by combining multiple pieces of user behavior data with the same user information, behavior type information, and behavior object information in the same data group into one piece of user behavior data according to preset rules, it is possible to eliminate duplicates of the same user's identical behaviors at different times, thereby improving the accuracy of subsequent calculations of user behavior similarity. This facilitates saving computing resources when subsequently calculating user behavior similarity based on multiple pieces of user behavior data in the data group.

[0024] Optionally, comparing the number of users in the similarity association link with a preset level threshold to determine the level of the similarity association link specifically includes:

[0025] Determine the number of users of each similarity association link;

[0026] If the number of users of the similarity-associated link is less than or equal to the preset first sub-level threshold, the level of the similarity-associated link is the lowest level;

[0027] If the number of users of the similarity-associated link is greater than the first sub-level threshold and less than or equal to the preset second sub-level threshold, the level of the similarity-associated link is the second lowest level;

[0028] If the number of users of the similarity-associated link is greater than the second sub-level threshold and less than or equal to the preset third sub-level threshold, the level of the similarity-associated link is the second highest level;

[0029] If the number of users of the similarity-related link is greater than the third sub-level threshold, the level of the similarity-related link is the highest level.

[0030] In the above method, by comparing the number of users in the similarity association link with a plurality of preset sub-level thresholds, the level of the similarity association link can be determined more accurately.

[0031] Optionally, the above preset duration is determined based on behavior type information.

[0032] In the above method, the method of determining the preset time length based on the behavior type information can divide different conditions for determining similar behavior information pairs for different behavior types, so that the subsequent calculation of user behavior similarity based on similar behavior information pairs is more accurate.

[0033] In a second aspect, an embodiment of the present application provides a target user identification device, comprising:

[0034] A processing module is used to group multiple pieces of user behavior data according to preset rules based on a preset time period and device information to obtain multiple data groups, where the user behavior data includes behavior object information, behavior type information, and behavior time information;

[0035] The processing module is further configured to treat any two pieces of user behavior data having the same behavior object information and behavior type information and a data time span less than or equal to a preset time length as a similar behavior information pair, where the data time span represents the time length between the behavior time information of any two pieces of user behavior data;

[0036] a calculation module for calculating user behavior similarity, used to characterize the correlation between user behaviors, based on the number of similar behavior information pairs between two users and the amount of user behavior data of the two users, and associating users whose user behavior similarity is not less than a set similarity threshold to obtain a similarity association link, where users in the similarity association link have the same type of behavior on the same behavior object within a preset time period;

[0037] The determination module is used to compare the number of users in the similarity association link with a preset level threshold, determine the level of the similarity association link, and determine the user in the similarity association link with the highest level as the target user.

[0038] In a third aspect, an embodiment of the present application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the processor implements any one of the target user identification methods described in the first aspect above.

[0039] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, any one of the target user identification methods in the first aspect is implemented.

[0040] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program, which is executed by a processor to implement any target user identification method as described in the first aspect above.

[0041] The technical effects brought about by any implementation method in the second to fifth aspects can refer to the technical effects brought about by the corresponding implementation method in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] FIG1 is a schematic diagram of an application scenario of a target user identification method provided by an embodiment of the present application;

[0043] FIG2 is a flow chart of a target user identification method provided in an embodiment of the present application;

[0044] FIG3 is a schematic diagram of a user behavior database provided in an embodiment of the present application;

[0045] FIG4 is a schematic diagram of another user behavior database provided in an embodiment of the present application;

[0046] FIG5 is a schematic diagram of another user behavior database provided in an embodiment of the present application;

[0047] FIG6 is a flow chart of a method for determining a data group provided in an embodiment of the present application;

[0048] FIG7 is a schematic diagram of another user behavior database provided in an embodiment of the present application;

[0049] FIG8 is a schematic diagram of another user behavior database provided in an embodiment of the present application;

[0050] FIG9 is a schematic diagram of another user behavior database provided in an embodiment of the present application;

[0051] FIG10 is a schematic diagram of user behavior similarity matching provided by an embodiment of the present application;

[0052] FIG11 is a schematic diagram of a process for establishing a similarity association link according to an embodiment of the present application;

[0053] FIG12 is a schematic diagram of a similarity association link provided in an embodiment of the present application;

[0054] FIG13 is an exemplary flowchart of a target user identification method provided in an embodiment of the present application;

[0055] FIG14 is a schematic diagram of a device for identifying a target user provided in an embodiment of the present application;

[0056] FIG15 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] To make the objectives, technical solutions, and advantages of this application more clear, this application will be further described in detail below with reference to the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0058] The application scenarios described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Persons skilled in the art will appreciate that, as new application scenarios emerge, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems. In the description of this application, unless otherwise specified, "multiple" means two or more.

[0059] With the rapid development of internet technology, the importance of network information security has become increasingly prominent. Organizations or individuals with malicious intent often exploit website vulnerabilities to disrupt internet order and seek illicit gains. This poses a significant threat to network data and property security.

[0060] How to accurately identify target users and effectively improve network information security has become a question worth discussing.

[0061] To address the above-mentioned issues, embodiments of the present application provide a target user identification method, apparatus, and electronic device. For example, based on a preset first time period and device information, multiple pieces of user behavior data from multiple different devices in a set second time period are grouped according to preset rules to obtain multiple data groups. The user behavior data includes behavior object information, behavior type information, and behavior time information, and the duration of the first time period is shorter than the duration of the second time period. Based on the user behavior data, multiple similar behavior information pairs are determined in the multiple data groups. Each similar behavior information pair includes two user behavior data with the same behavior object information, the same behavior type information, and a data time span that is less than or equal to a preset time span. The data time span represents the duration between the behavior time information of the two user behavior data. Based on the number of similar behavior information pairs between the two users and the amount of user behavior data of the two users, a user behavior similarity is calculated to characterize the correlation between the user behaviors, and users whose user behavior similarity is not less than a set similarity threshold are associated to obtain a similarity association link. The similarity association link includes multiple nodes, which are connected based on the order of the behavior time information, and each node represents a user. The number of users in the similarity association link is compared with a preset level threshold to determine the level of the similarity association link, and the user in the similarity association link with the highest level is determined as the target user.

[0062] This calculation of user behavior similarity accurately identifies the correlation between different user behaviors, improving the accuracy and efficiency of identifying target users and enhancing network information security. By calculating user behavior similarity based on multiple pieces of user behavior data, target users can be identified based on similarity at any time, quickly and efficiently meeting business needs and improving the user experience.

[0063] As shown in Figure 1, an application scenario diagram of an optional target user identification method in an embodiment of the present application includes a server 100 and a terminal 101. The server 100 and the terminal 101 can be communicatively connected through a network to implement the target user identification method of the present application.

[0064] The user can use the server 100 to interact with the terminal 101 via the network, such as receiving or sending messages, etc. Various client applications can be installed on the terminal 101, such as programming applications, web browser applications, search applications, etc.

[0065] It is understood that in the embodiment of the present application, the server 100 can be implemented as an independent server or a server cluster composed of multiple servers. The terminal 101 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, desktop computers, etc.

[0066] As shown in FIG2 , a flowchart of a target user identification method provided in an embodiment of the present application may specifically include the following steps.

[0067] S201: Based on a preset first time period and device information, a plurality of user behavior data of a plurality of different devices in a set second time period are grouped according to a preset rule to obtain a plurality of data groups.

[0068] The user behavior data includes behavior object information, behavior type information, and behavior time information, and the duration of the first time period is shorter than the duration of the second time period.

[0069] For example, behavior type information includes other types of behavior, such as fake orders, likes, forwarding, and comments. The behavior time information is 12:28 on December 19, 2023. Device information includes the device identifier, such as a digital sequence.

[0070] Action target information is used to indicate the target object corresponding to the user's action. For example, if user A likes user N, user N is the action target. It is understood that the action target information can be a specific user, a specific video, etc. This application does not make specific restrictions on this.

[0071] It is understood that the first time period and the second time period in this application may be preset by those skilled in the art. They may also be changed according to the application scenario. For example, the first time period may be 1 hour, and the second time period may be 24 hours. For another example, the first time period may be 30 minutes, and the second time period may be 12 hours. This application does not specifically limit this.

[0072] Optionally, user behavior data also includes user information. User information includes a user ID. The user ID is used to identify the user in a specific region. For example, the user ID can be the characters "Zhang San" or an identity identifier. The user ID is information used to identify users within the enterprise, such as a numerical sequence.

[0073] In a possible embodiment, the server may obtain a user behavior database from the platform, and then obtain multiple pieces of user behavior data for a set second time period from the user behavior database, wherein the user behavior database includes multiple pieces of user behavior data.

[0074] Optionally, the plurality of user behavior data may be collected internally by the enterprise or provided by other enterprise organizations. This application does not limit the method of collecting the plurality of user behavior data.

[0075] As shown in Figure 3, the present application provides a schematic diagram of a user behavior database. In Figure 3, each line of user behavior data is a piece of user behavior data. For example, (user A, device 1, moment 1, object 1, behavior 1) is T1 data in the user behavior database shown in Figure 1. The user information of this T1 data is user A, the device information is device 1, the behavior time information is moment 1, the behavior object information is object 1, and the behavior type is behavior 1. For another example, (user B, device 2, moment 2, object 2, behavior 1) is T3 data in the user behavior database shown in Figure 1. The user information of this T3 data is user B, the device information is device 2, the behavior time information is moment 2, the behavior object information is object 2, and the behavior type is behavior 1.

[0076] Optionally, after obtaining the multiple pieces of user behavior data, the server may perform data cleaning on the multiple pieces of user behavior data to obtain multiple pieces of cleaned user behavior data.

[0077] Among them, any two pieces of user behavior data among the multiple pieces of cleaned user behavior data are different, and any piece of user behavior data includes behavior object information, behavior type information, behavior time information, user information and device information.

[0078] Specifically, the server can determine whether the data elements of each piece of user behavior data are complete. Data with incomplete data elements will be deleted. If there are multiple pieces of data with identical data elements, one duplicate piece will be retained and the remaining duplicates will be deleted. The data elements include user information, information about the user's behavior target, information about the behavior type, and information about the time of the behavior.

[0079] In the above method, by performing data cleaning on multiple user behavior data to obtain multiple cleaned user behavior data, it is convenient to save computing resources, prevent waste of computing resources, and improve computing efficiency when subsequently calculating user behavior similarity based on multiple user behavior data.

[0080] For example, as shown in Figure 4, the present application provides a schematic diagram of another user behavior database. In Figure 4, each line of the user behavior data is a piece of user behavior data. For example, (User A, Device 1, Time 1, Object 1, Behavior 1) is T1 data in the user behavior database shown in Figure 4. (User A, Device 1, Object 1, Behavior 1) is T2 data in the user behavior database shown in Figure 4. The data element of the T2 data lacks behavior time information, and the data element is incomplete, so the T2 data is deleted.

[0081] For another example, as shown in Figure 5, the present application provides a schematic diagram of another user behavior database. Each line of the user behavior data in Figure 5 is a piece of user behavior data. For example, (user A, device 1, time 1, object 1, behavior 1) is T1 data in the user behavior database shown in Figure 5. (User A, device 1, time 1, object 1, behavior 1) is T4 data in the user behavior database shown in Figure 5. T1 data and T4 data have completely repeated data elements of T4 data. Then retain T1 data and delete T4 data. Or, retain T4 data and delete T1 data.

[0082] In an optional embodiment, after obtaining multiple pieces of user behavior data, the server may process the multiple pieces of user behavior data to obtain multiple data groups.

[0083] Specifically, as shown in FIG6 , the present application provides a flow chart of a method for determining a data group.

[0084] Step S601: sorting multiple pieces of user behavior data in chronological order.

[0085] Step S602: grouping based on a preset first time period to obtain a plurality of time window data segments.

[0086] Each time window data segment includes multiple pieces of user behavior data, and the time span of behavior time information between any two pieces of user behavior data in the time window data segment is less than or equal to a preset time period.

[0087] For example, suppose the total time span of the behavior time information of multiple pieces of user behavior data is 24 hours. If the preset first time period, i.e., the time window length, is determined to be 1 hour, the multiple pieces of user behavior data can be divided into 24 time window data segments. These are the first time window data segment, the second time window data segment, ... the 24th time window data segment. The maximum time span of each time window data segment after segmentation is 1 hour.

[0088] For example, as shown in Figure 7, this application provides a schematic diagram of another user behavior database. Referring to Figure 7, after segmenting the user behavior database using time windows, the user behavior database is segmented into multiple time window data segments, where the first time window data segment includes user behavior data at T1, T2, T3, T4, and T5. The second time window data segment includes user behavior data at T7 and T8.

[0089] Step S603: Divide multiple pieces of user behavior data with the same device information in the time window data segment into one group to obtain multiple data groups.

[0090] For example, as shown in Figure 8, this application provides a schematic diagram of another user behavior database. Referring to Figure 8, the first time window data segment is grouped to form a first data group, a second data group, and a third data group. The device information for the first data group is the same, all being device 1; the device information for the first data group is device 2; and the device information for the third data group is device 5.

[0091] In the above method, multiple pieces of user behavior data are sorted chronologically and grouped based on preset time periods to obtain multiple time window data segments, thereby enabling initial correlation of the multiple pieces of user behavior data in chronological order. Furthermore, by grouping multiple pieces of user behavior data with identical device information within the time window data segments into a single group to obtain multiple data groups, multiple pieces of user behavior data initiated based on the same device within the preset time period can be further identified, thereby enhancing the correlation of the multiple pieces of user behavior data within the data group and facilitating subsequent calculation of user behavior similarity.

[0092] After determining the multiple data groups, the server may calculate the user behavior similarity between any two users in each data group of the current time window data segment.

[0093] For example, taking the first time window data segment shown in Figure 8 as the current time window data segment, the user information in the first data group includes user A, user C, user D and user F, the user information in the second data group includes user B, and the user information in the third data group includes user E.

[0094] For the first data set: the similarity between user A and user C in the first data set can be calculated. The similarity between user A and user D can also be calculated. The similarity between user A and user F can also be calculated. The similarity between user C and user D can also be calculated. The similarity between user C and user F can also be calculated. The similarity between user D and user F can also be calculated.

[0095] Optionally, after determining the multiple data groups and before determining the similar behavior information pairs, the server may further determine whether the data groups are valid data groups. Specifically, the server may determine the number of users in each data group. If the number of users is less than a preset valid threshold, the server may delete the data groups with fewer users than the valid threshold.

[0096] That is, when a data group is a valid data group, step S202 is executed for the data group; otherwise, the data group is skipped or deleted.

[0097] It is understood that the above-mentioned effective threshold can be preset by those skilled in the art. It can also be changed according to the application scenario. For example, the effective threshold is 2. This application does not make specific limitations on this.

[0098] In the above method, by deleting data groups with fewer than the valid threshold number of users, the remaining data groups can be made valid. This can save computing resources, prevent waste of computing resources, and improve computing efficiency when subsequently calculating user behavior similarity based on multiple user behavior data in the data group.

[0099] For example, assuming the validity threshold is 2, a determination is made as to whether the number of users in the data group is greater than or equal to 2. If the number of users is greater than or equal to 2, the data group is determined to be a valid data group, and step S202 is executed. If the number of users is less than 2, the data group is determined to be an invalid data group and is deleted. Among the multiple data groups shown in FIG8 , since the number of users in the second and third data groups is less than 2, the second and third data groups can be deleted before executing step S202. Alternatively, the second and third data groups can be skipped during step S202.

[0100] Step S202: determining a plurality of similar behavior information pairs in a plurality of data groups based on the user behavior data.

[0101] Among them, each similar behavior information pair includes two user behavior data with the same behavior object information, the same behavior type information, and a data time span less than or equal to a preset duration. The data time span represents the duration between the behavior time information of the two user behavior data.

[0102] Optionally, before executing step S202 , the server may further merge multiple pieces of user behavior data with the same user information, behavior type information, and behavior object information in the same data group into one piece of user behavior data according to preset rules.

[0103] For example, as shown in Figure 9, the present application provides a schematic diagram of another user behavior database. Based on the user behavior database in Figure 8, the T5 data item in the first data group {user F, device 1, time 6, object 1, behavior 1} and the T6 data item in the first data group {user F, device 1, time 7, object 1, behavior 1} have the same user information, and the behavior object information and behavior type information in the behavior information are the same. Therefore, the T5 data item and the T6 data item in the figure can be merged into the T5,6 data item {user F, device 1, [time 6, time 7], object 1, behavior 1}, resulting in Figure 9.

[0104] In the above method, by combining multiple pieces of user behavior data with the same user information, behavior type information, and behavior object information in the same data group into one piece of user behavior data according to preset rules, it is possible to eliminate duplicates of the same user's identical behaviors at different times, thereby improving the accuracy of subsequent calculations of user behavior similarity. This facilitates saving computing resources when subsequently calculating user behavior similarity based on multiple pieces of user behavior data in the data group.

[0105] In an optional embodiment, if the behavior object information and behavior type information of two pieces of user behavior data are consistent, and the time span between the two pieces of user behavior data is less than or equal to a preset time length, then the two pieces of user behavior data are determined to be similar, and the two similar pieces of data are a similar behavior information pair. Otherwise, they are not similar.

[0106] Optionally, the preset duration may be determined based on behavior type information.

[0107] In the above method, the method of determining the preset time length based on the behavior type information can divide different conditions for determining similar behavior information pairs for different behavior types, so that the subsequent calculation of user behavior similarity based on similar behavior information pairs is more accurate.

[0108] For example, assume that behavior type 1 may be set with a first preset duration, which is 30 minutes, and behavior type 2 may be set with a second preset duration, which is 60 minutes.

[0109] For another example, take the first data group in the first time window data segment of Figure 8 as an example. T2 data is not similar to T1 data, T3 data, T4 data, T5, and T6 data (the behavior type information is different, and the behavior object information is different). If the time span between moment 1 and moment 3 in the first data group is less than or equal to the preset duration, then T1 data is similar to T3 data, otherwise they are not similar. If the time span between moment 1 and moment 4 in the first data group is less than or equal to the preset duration, then T1 data is similar to T4 data, otherwise they are not similar. If the time span between moment 1 and moment 6 in the first data group is less than or equal to the preset duration, then T1 data is similar to T5 data, otherwise they are not similar. If the time span between moment 1 and moment 7 in the first data group is less than or equal to the preset duration, then T1 data is similar to T6 data, otherwise they are not similar.

[0110] Step S203: Calculate the user behavior similarity used to characterize the user behavior correlation based on the number of similar behavior information pairs between the two users and the amount of user behavior data of the two users, and associate users whose user behavior similarity is not less than a set similarity threshold to obtain a similarity association link.

[0111] The similarity association link includes multiple nodes. The multiple nodes are connected based on the order of behavior time information. Each node represents a user.

[0112] In an optional embodiment, the server may determine the number of similar behavior information pairs between any two different user combinations in each data group in the current time window data segment, and the amount of user behavior data of each user.

[0113] Optionally, the following formula may be used to calculate user behavior similarity.

[0114] Where J(X,Y) represents user behavior similarity. X represents user X. Y represents user Y. |X∩Y| represents the number of similar behavior information pairs between user X and user Y. |X∪Y| represents the sum of the number of user behavior data for user X and the number of user behavior data for user Y.

[0115] For example, as shown in Figure 10, the present application provides a schematic diagram of user behavior similarity matching. Specifically, Figure 10 is a similarity matching diagram between any two different user combinations determined based on the first data group of the first time window data segment in Figure 8.

[0116] As shown in Figure 8, there are four different users in the first data group of the first time window data segment: User A (two pieces of behavior information), User C (one piece of behavior information), User D (one piece of behavior information), and User F (one piece of behavior information). When performing step S203 on the first time window data segment, the users can be grouped into the following five different combinations to obtain the user behavior similarity between different users.

[0117] The first type: X = user A, Y = user C, |X∩Y| = 1, |X∪Y| = 3, J(X,Y) = 0.5;

[0118] The second type: X = user A, Y = user D, |X∩Y| = 1, |X∪Y| = 3, J(X,Y) = 0.5;

[0119] The third type: X = user A, Y = user F, |X∩Y| = 1, |X∪Y| = 3, J(X,Y) = 0.5;

[0120] The fourth type: X = user C, Y = user D, |X∩Y| = 1, |X∪Y| = 2, J(X,Y) = 1;

[0121] The fifth type: X = user D, Y = user F, |X∩Y| = 1, |X∪Y| = 2, J(X,Y) = 1.

[0122] It is understandable that when calculating user behavior similarity, it is necessary to traverse all time window data segments, calculate the user behavior similarity based on the number of similar behavior information pairs of any two users, and the amount of user behavior data of any two users, until the similarity between all any two users in each data group of all time window data segments is determined.

[0123] For example, after calculating and determining the similarity between all arbitrary two users in each data group in the first time window data segment, the similarity between all arbitrary two users in each data group in the second time window data segment is calculated, and then the similarity between all arbitrary two users in each data group in the third time window data segment is calculated... until all arbitrary two users in each data group in all time window data segments are traversed.

[0124] In an optional embodiment, after determining the user behavior similarity of all two users in each data group of all time window data segments, the server may establish a similarity association link based on the user behavior similarity and a similarity threshold.

[0125] Specifically, as shown in FIG11 , the present application provides a flowchart of establishing a similarity association link.

[0126] Step S1101: comparing user behavior similarity with a similarity threshold;

[0127] It is understandable that the similarity threshold in the embodiment of the present application can be pre-set by those skilled in the art. The similarity threshold can also be set according to needs. For example, the similarity threshold in this embodiment can be set to 0.5.

[0128] Step S1102: obtaining similarity association links for users whose user behavior similarity is not less than a set similarity threshold.

[0129] For example, Figure 12 is a schematic diagram of a similarity association link provided by an embodiment of the present application. Assume that in the first time window data segment, user A is associated with user C, user A is associated with user D, user A is associated with user F, user C is associated with user D, and user D is associated with user F; in other time window data segments, user A is associated with user B, and user B is associated with user E. Thus, a similarity association link as shown in Figure 12 can be formed. It can be seen from Figure 12 that the similarity association link includes user A, user B, user C, user D, user E, and user F.

[0130] Step S204: Compare the number of users in the similarity association link with a preset level threshold, determine the level of the similarity association link, and determine the user in the similarity association link with the highest level as the target user.

[0131] Specifically, the server determines the number of users of each similarity-associated link. If the number of users of the similarity-associated link is less than or equal to a preset first sub-level threshold, the level of the similarity-associated link is the lowest level. If the number of users of the similarity-associated link is greater than the first sub-level threshold and less than or equal to a preset second sub-level threshold, the level of the similarity-associated link is the second lowest level. If the number of users of the similarity-associated link is greater than the second sub-level threshold and less than or equal to a preset third sub-level threshold, the danger level of the similarity-associated link is the second highest level. If the number of users of the similarity-associated link is greater than the third sub-level threshold, the level of the similarity-associated link is the highest level.

[0132] It is understood that this application does not specifically limit the number and values ​​of sub-level thresholds for determining the level of similarity-related links. For example, the sub-level thresholds may include a first sub-level threshold, a second sub-level threshold, a third sub-level threshold, a fourth sub-level threshold, and so on. Those skilled in the art may modify or pre-set these thresholds based on specific application scenarios.

[0133] In the above method, by comparing the number of users in the similarity association link with a plurality of preset sub-level thresholds, the level of the similarity association link can be determined more accurately.

[0134] For example, the number of users of the similarity association link shown in FIG12 is 6. Assume that the first sub-level threshold is 3, the second sub-level threshold is 6, and the third sub-level threshold is 12. Then, the level of the similarity association link is determined to be intermediate.

[0135] For another example, suppose the similarity association link represents the risk level of a dangerous user. The risk levels, from low to high, may include: lowest risk level, second lowest risk level, second lowest risk level, and second highest risk level. Assume the first sub-level threshold is 3, the second sub-level threshold is 6, and the third sub-level threshold is 12.

[0136] If the number of users on a similarity association link is less than or equal to 3, the risk level of the similarity association link is determined to be the lowest risk level. If the number of users on a similarity association link is greater than 3 and less than or equal to 6, the risk level of the similarity association link is determined to be the second lowest risk level.

[0137] If the number of users in a similarity association link is greater than 6 and less than or equal to 12, the risk level of the similarity association link is determined to be the second highest risk level. If the number of users in a similarity association link is greater than 12, the risk level of the similarity association link is determined to be the highest risk level. After determining the risk levels of all similarity association links, all users in the similarity association link with the highest risk level can be determined as target users.

[0138] The embodiments of the present application can improve computing efficiency while improving the accuracy and efficiency of target user identification by cleaning data in a user behavior database, dividing data segments by time windows, grouping the same data segment by device, calculating the similarity between all any two users in each data group, and determining similarity association links based on the similarity.

[0139] As shown in FIG13 , an embodiment of the present application provides an exemplary flow chart of a target user identification method, which may include the following steps:

[0140] S1301. Acquire multiple pieces of user behavior data of multiple different devices in a set second time period from a user behavior database;

[0141] S1302. Cleaning multiple pieces of user behavior data.

[0142] S1303: Sort the multiple pieces of user behavior data in chronological order, and group them based on a preset first time period to obtain multiple time window data segments;

[0143] S1304: Group multiple pieces of user behavior data with the same device information in the time window data segment into one group to obtain multiple data groups;

[0144] S1305: Determine the number of users in each data group;

[0145] S1306: If the number of users is less than the preset effective threshold, delete the data group whose number of users is less than the effective threshold;

[0146] S1307: Merge multiple pieces of user behavior data with the same user information, behavior type information, and behavior object information in the same data group into one piece of user behavior data according to preset rules;

[0147] S1308: any two user behavior data with the same behavior object information and behavior type information and a data time span less than or equal to a preset time length are regarded as a similar behavior information pair;

[0148] S1309: Calculating a user behavior similarity for characterizing user behavior relevance based on the number of similar behavior information pairs between the two users and the amount of user behavior data of the two users;

[0149] S1310: Determine whether all time window data segments have been traversed to calculate user behavior similarity. If so, execute S1311; if not, execute S1304.

[0150] S1311: Associating users whose user behavior similarity is not less than a set similarity threshold, obtaining similarity association links;

[0151] S1312: Compare the number of users in the similarity association link with a preset level threshold to determine the level of the similarity association link;

[0152] S1313: Determine the user in the highest-level similarity association link as the target user.

[0153] FIG14 is a schematic structural diagram of a target user identification device provided in an embodiment of the present application. As shown in FIG14 , the device includes: a processing module 1401 , a calculation module 1402 , and a determination module 1403 .

[0154] Processing module 1401 is configured to group, based on a preset first time period and device information, a plurality of pieces of user behavior data from a plurality of different devices in a set second time period according to a preset rule to obtain a plurality of data groups, wherein the user behavior data includes behavior object information, behavior type information, and behavior time information, and the duration of the first time period is shorter than the duration of the second time period;

[0155] The processing module 1401 is further configured to determine, based on the user behavior data, a plurality of similar behavior information pairs in the plurality of data groups, each similar behavior information pair including two user behavior data having the same behavior object information, the same behavior type information, and a data time span that is less than or equal to a preset time length, where the data time span represents the time length between the behavior time information of the two user behavior data;

[0156] Calculation module 1402, configured to calculate user behavior similarity, used to characterize user behavior correlation, based on the number of similar behavior information pairs between two users and the amount of user behavior data of the two users, and to associate users whose user behavior similarity is not less than a set similarity threshold to obtain a similarity association link, wherein the similarity association link includes multiple nodes, the multiple nodes being connected based on the order of behavior time information, and each node representing a user;

[0157] The determination module 1403 is configured to compare the number of users in the similarity association link with a preset level threshold, determine the level of the similarity association link, and determine the user in the similarity association link with the highest level as the target user.

[0158] Optionally, based on the preset first time period and device information, multiple pieces of user behavior data of multiple different devices in the set second time period are grouped according to preset rules to obtain multiple data groups. The processing module 1401 is specifically configured to:

[0159] Sort the multiple user behavior data in chronological order, and group them based on a preset first time period to obtain multiple time window data segments, each time window data segment including multiple user behavior data;

[0160] Multiple pieces of user behavior data with the same device information in the time window data segment are divided into one group to obtain multiple data groups.

[0161] Optionally, the user behavior data further includes user information. Before sorting the plurality of user behavior data in chronological order, the processing module 1401 is further configured to:

[0162] Data cleaning is performed on multiple pieces of user behavior data to obtain multiple pieces of cleaned user behavior data, wherein any two pieces of user behavior data in the multiple pieces of cleaned user behavior data are different, and any piece of user behavior data includes behavior object information, behavior type information, behavior time information, user information, and device information.

[0163] Optionally, after grouping the plurality of user behavior data of the plurality of different devices in the set second time period according to a preset rule to obtain a plurality of data groups based on the preset first time period and the device information, the processing module 1401 is further configured to:

[0164] Determine the number of users for each data group;

[0165] When the number of users is less than a preset effective threshold, the data group with the number of users less than the effective threshold is deleted.

[0166] Optionally, before determining the plurality of similar behavior information pairs in the plurality of data groups based on the user behavior data, the processing module 1401 is further configured to:

[0167] Multiple pieces of user behavior data with the same user information, behavior type information, and behavior object information in the same data group are merged into one piece of user behavior data according to preset rules.

[0168] Optionally, the above-mentioned method compares the number of users in the similarity association link with a preset level threshold to determine the level of the similarity association link. The determination module 1403 is specifically used to:

[0169] Determine the number of users of each similarity association link;

[0170] If the number of users of the similarity-associated link is less than or equal to the preset first sub-level threshold, the level of the similarity-associated link is the lowest level;

[0171] If the number of users of the similarity-associated link is greater than the first sub-level threshold and less than or equal to the preset second sub-level threshold, the level of the similarity-associated link is the second lowest level;

[0172] If the number of users of the similarity-associated link is greater than the second sub-level threshold and less than or equal to the preset third sub-level threshold, the level of the similarity-associated link is the second highest level;

[0173] If the number of users of the similarity-related link is greater than the third sub-level threshold, the level of the similarity-related link is the highest level.

[0174] Based on the same technical concept, an electronic device is also provided in an embodiment of the present application, which can realize the functions of the aforementioned target user identification device.

[0175] FIG15 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0176] At least one processor 1501, and a memory 1502 connected to at least one processor 1501. In the embodiments of the present application, the specific connection medium between the processor 1501 and the memory 1502 is not limited. Figure 15 takes the connection between the processor 1501 and the memory 1502 via the bus 1500 as an example. The bus 1500 is represented by a bold line in Figure 15, and the connection method between other components is only for schematic illustration and is not intended to be limiting. The bus 1500 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one bold line is used in Figure 15, but this does not mean that there is only one bus or one type of bus. Alternatively, the processor 1501 can also be called a controller, and there is no limitation on the name.

[0177] In this embodiment of the present application, memory 1502 stores instructions executable by at least one processor 1501. At least one processor 1501 can perform a target user identification method discussed above by executing the instructions stored in memory 1502. Processor 1501 can implement the functions of each module in the apparatus shown in FIG14.

[0178] Among them, processor 1501 is the control center of the device, which can use various interfaces and lines to connect the various parts of the entire control device, and monitor the device as a whole by running or executing instructions stored in memory 1502 and calling data stored in memory 1502, various functions of the device and processing data.

[0179] In one possible design, processor 1501 may include one or more processing units. Processor 1501 may integrate an application processor and a modem processor. The application processor primarily processes the operating system, driver interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 1501. In some embodiments, processor 1501 and memory 1502 may be implemented on the same chip. In some embodiments, they may also be implemented on separate chips.

[0180] The processor 1501 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of a target user identification method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.

[0181] Memory 1502 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. Memory 1502 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. Memory 1502 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 1502 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0182] By designing and programming the processor 1501, the code corresponding to the target user identification method described in the aforementioned embodiment can be embedded in the chip, thereby enabling the chip to execute the target user identification method of the embodiment shown in FIG2 during operation. Designing and programming the processor 1501 is well known to those skilled in the art and will not be further described here.

[0183] It should be noted here that the above-mentioned electronic device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned method embodiment and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those in the method embodiment will not be described in detail here.

[0184] An embodiment of the present application further provides a computer-readable storage medium, which stores computer-executable instructions. The computer-executable instructions are used to enable a computer to execute a target user identification method in the above embodiment.

[0185] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0186] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0187] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0188] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0189] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for identifying target users, characterized in that, The method includes: Based on a preset first time period and device information, grouping multiple pieces of user behavior data of multiple different devices in a set second time period according to preset rules to obtain multiple data groups. The user behavior data includes behavior object information, behavior type information, and behavior time information, and the duration of the first time period is less than the duration of the second time period; Based on the user behavior data, respectively determining multiple pairs of similar behavior information in the multiple data groups. Each pair of similar behavior information includes two pieces of user behavior data with the same behavior object information, the same behavior type information, and a data time span less than or equal to a preset duration. The data time span represents the duration between the behavior time information of the two pieces of user behavior data; Based on the number of pairs of similar behavior information between two users and the quantity of user behavior data of the two users, calculating a user behavior similarity for characterizing the relevance of user behavior, and associating users with a user behavior similarity not less than a set similarity threshold to obtain a similarity association link. The similarity association link includes multiple nodes, and the multiple nodes are connected in the order of the sequence of the behavior time information. Each node represents a user; Comparing the number of users in the similarity association link with a preset level threshold to determine the level of the similarity association link, and determining the users in the similarity association link with the highest level as target users.

2. The method according to claim 1, wherein The step of grouping multiple pieces of user behavior data of multiple different devices in a set second time period according to preset rules based on a preset first time period and device information to obtain multiple data groups specifically includes: Sorting the multiple pieces of user behavior data in chronological order, and grouping them based on the preset first time period to obtain multiple time window data segments. Each time window data segment includes multiple pieces of user behavior data; Dividing multiple pieces of user behavior data with the same device information in the time window data segment into one group to obtain the multiple data groups.

3. The method according to claim 2, characterized in that, The user behavior data further includes user information. Before sorting the multiple pieces of user behavior data in chronological order, the method further includes: Performing data cleaning on the multiple pieces of user behavior data to obtain multiple pieces of cleaned user behavior data. Any two pieces of the cleaned user behavior data are different, and any piece of user behavior data includes the behavior object information, the behavior type information, the behavior time information, the user information, and the device information.

4. The method according to claim 1, characterized in that, After grouping multiple pieces of user behavior data of multiple different devices in a set second time period according to preset rules based on a preset first time period and device information to obtain multiple data groups, the method further includes: Determining the number of users in each data group; In the case where the number of users is less than a preset effective threshold, deleting the data group with the number of users less than the effective threshold.

5. The method according to claim 1 or 3, characterized in that, Before respectively determining multiple pairs of similar behavior information in the multiple data groups based on the user behavior data, the method further includes: Multiple pieces of user behavior data with the same user information, behavior type information, and behavior object information in the same data group are merged into one piece of user behavior data according to a preset rule.

6. The method according to claim 1, wherein Comparing the number of users in the similarity association link with a preset level threshold to determine the level of the similarity association link specifically includes: Determining the number of users in each similarity association link; If the number of users in the similarity association link is less than or equal to a preset first sub-level threshold, the level of the similarity association link is the lowest level; If the number of users in the similarity association link is greater than the first sub-level threshold and less than or equal to a preset second sub-level threshold, the level of the similarity association link is the second lowest level; If the number of users in the similarity association link is greater than the second sub-level threshold and less than or equal to a preset third sub-level threshold, the level of the similarity association link is the second highest level; If the number of users in the similarity association link is greater than the third sub-level threshold, the level of the similarity association link is the highest level.

7. The method according to any one of claims 1 to 4, characterized in that The preset duration is determined based on the behavior type information.

8. A target user identification device, characterized in that, Including: A processing module, configured to group multiple pieces of user behavior data of multiple different devices in a set second time period according to a preset rule based on a preset first time period and device information. The user behavior data includes behavior object information, behavior type information, and behavior time information, and the duration of the first time period is less than the duration of the second time period; The processing module is further configured to respectively determine multiple pairs of similar behavior information in the multiple data groups based on the user behavior data. Each pair of similar behavior information includes two pieces of user behavior data with the same behavior object information, the same behavior type information, and a data time span less than or equal to a preset duration. The data time span represents the duration between the behavior time information of the two pieces of user behavior data; A calculation module, configured to calculate a user behavior similarity for characterizing the relevance of user behavior based on the number of pairs of similar behavior information between two users and the number of pieces of user behavior data of the two users, and associate users with a user behavior similarity not less than a set similarity threshold to obtain a similarity association link. The similarity association link includes multiple nodes, and the multiple nodes are connected in the order of the sequence of the behavior time information. Each node represents a user; A determination module, configured to compare the number of users in the similarity association link with a preset level threshold, determine the level of the similarity association link, and determine the users in the highest-level similarity association link as target users.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.

11. A computer program product, characterized in that, When the computer program product is called by a computer, it causes the computer to execute the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device for identifying abnormal object

    CN110390585A

  • Malicious account identification method, malicious account identification device, medium and electronic equipment

    CN111371767A

  • Target user identification method and device, electronic equipment and storage medium

    CN114257427A

  • Abnormal group identification method and device, equipment, readable storage medium and product

    CN114676290A

  • Target user identification method and device and electronic equipment

    CN117834231A