Target account identification method and device, electronic equipment and storage medium

By recording and analyzing the posting behavior of candidate accounts in real time, traffic time-series features and correlation graphs are generated. Combined with the K-clique algorithm, black and gray market accounts and teams are accurately identified, solving the problem of insufficient identification accuracy in existing technologies and improving community security.

CN120893025APending Publication Date: 2025-11-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510804910.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

On forums and social media platforms, black market accounts engage in malicious activities through methods such as automated posting, leading to community disorder and reduced security. Existing technologies struggle to accurately identify these accounts.

Method used

By recording the posting behavior of candidate accounts in real time, traffic time-series features are generated. Abnormal posting behavior is identified using sequence variation coefficient and correlation graph. The target team is identified by combining the K-clique algorithm.

Benefits of technology

It has improved the accuracy of identifying black and gray market accounts, reduced missed and false positives, and maintained community order and cybersecurity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120893025A_ABST
    Figure CN120893025A_ABST
Patent Text Reader

Abstract

The invention provides a target account identification method and device, electronic equipment and a storage medium, and relates to the field of artificial intelligence such as network security and knowledge maps. The method comprises the following steps: recording posting behaviors of candidate accounts in a target application in real time; in response to determining that the first time window is passed, determining the posting amount of each candidate account in the first time window according to the recorded content; in response to the fact that a second time window is determined to pass, the second time window comprises N continuous first time windows, N is a positive integer larger than 1, corresponding target sequences are generated for all the candidate accounts, and the target sequences comprise the posting amount of the corresponding candidate accounts in the N first time windows, and the posting amount of the corresponding candidate accounts in the N first time windows is larger than the posting amount of the corresponding candidate accounts in the N first time windows; and according to the target sequence, determining traffic time sequence characteristics of the candidate accounts, and according to the traffic time sequence characteristics, determining a target account with an abnormal posting behavior from the candidate accounts.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the fields of network security and knowledge graph, and more particularly to a target account identification method and device, an electronic device and a storage medium. BACKGROUND

[0002] In community applications such as forums and social media platforms, black and gray production accounts (accounts participating in black or gray industries) can perform malicious behaviors through machine posting and other means, seriously disrupting community order, affecting the information acquisition efficiency of normal accounts, and reducing the security of the community, etc. SUMMARY

[0003] The present disclosure provides a target account identification method and device, an electronic device and a storage medium.

[0004] A target account identification method includes:

[0005] Real-time recording of posting behaviors of each candidate account in a target application;

[0006] In response to determining that a first time window has elapsed, the posting quantity of each candidate account within the first time window is determined according to the recorded content;

[0007] In response to determining that a second time window has elapsed, the second time window includes consecutive N first time windows, and N is a positive integer greater than 1, for each candidate account, a corresponding target sequence is generated, the target sequence includes the posting quantity of the corresponding candidate account within the N first time windows, the traffic time sequence feature of each candidate account is determined according to the target sequence, and the target account with abnormal posting behavior is determined from each candidate account according to the traffic time sequence feature.

[0008] A target account identification device includes a recording module, a statistical module and an identification module.

[0009] The recording module is configured to record the posting behaviors of each candidate account in a target application in real time;

[0010] The statistical module is configured to determine the posting quantity of each candidate account within the first time window according to the recorded content in response to determining that a first time window has elapsed;

[0011] The identification module is configured to, in response to determining that a second time window is passed, the second time window including consecutive N first time windows, N being a positive integer greater than 1, generate, for each candidate account, a corresponding target sequence, the target sequence including post quantities of the corresponding candidate account in the N first time windows, determine, according to the target sequence, a traffic time sequence feature of each candidate account, and determine, according to the traffic time sequence feature, a target account having an abnormal posting behavior from the candidate accounts.

[0012] An electronic device comprising:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein

[0015] the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described above.

[0016] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method as described above.

[0017] A computer program product comprising computer programs / instructions that, when executed by a processor, implement the method as described above.

[0018] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the disclosure, nor is it used to limit the scope of the disclosure. Other features of the disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings are used to better understand the present scheme and do not limit the disclosure. Among them:

[0020] Figure 1 a flowchart of a first embodiment of the target account identification method according to the disclosure;

[0021] Figure 2 a schematic diagram of the target sequence according to the disclosure;

[0022] Figure 3 a flowchart of a second embodiment of the target account identification method according to the disclosure;

[0023] Figure 4 a schematic diagram of the composition structure of the target account identification device embodiment 400 according to the disclosure;

[0024] Figure 5A schematic block diagram of an electronic device 500 that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0025] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are incorporated in, and constitute a part of, this specification. Various details of the embodiments of the present disclosure are described herein in order to provide a thorough understanding of the embodiments of the present disclosure. It will be understood by those of ordinary skill in the art that the embodiments described herein can be practiced without these details, and that numerous implementations can be derived from the embodiments described herein without departing from the scope of the present disclosure. Similarly, some devices are not shown so as not to obscure the disclosure. Also, well-known functions or constructions are not described in detail so as not to obscure the disclosure.

[0026] In addition, it should be understood that the term "and / or" as used herein merely describes associated objects, and can exist in three forms, for example, A and / or B can mean that A exists alone, A and B exist together, or B exists alone. In addition, the character " / " in this paper generally represents an "or" relationship between the front and rear associated objects.

[0027] Figure 1 The flowchart of the first embodiment of the target account identification method described in the present disclosure is shown. As shown in Figure 1 the following specific implementation is included.

[0028] In step 101, the posting behavior of each candidate account in the target application is recorded in real time.

[0029] In step 102, in response to determining that the first time window has elapsed, the posting quantity of each candidate account in the first time window is determined according to the recorded content.

[0030] In step 103, in response to determining that the second time window has elapsed, the second time window includes N consecutive first time windows, and N is a positive integer greater than 1, for each candidate account, a corresponding target sequence is generated, the target sequence includes the posting quantity of the corresponding candidate account in the N first time windows, the traffic time sequence feature of each candidate account is determined according to the target sequence, and the target account with abnormal posting behavior is determined from each candidate account according to the traffic time sequence feature.

[0031] In the conventional way, when identifying black and gray production accounts, the posting quantity and other indicators of the account within a certain time period are usually counted. When these indicators exceed the preset threshold, the account is determined to be a black and gray production account. However, the accuracy of this method is usually poor. For example, black and gray production accounts may control the posting time interval to make the posting quantity appear normal in statistics, so that some black and gray production accounts cannot be discovered in time, and in addition, some normal accounts may be misjudged as black and gray production accounts.

[0032] According to the above method and the scheme, for any candidate account, after obtaining the post quantity in the continuous multiple first time windows, the corresponding target sequence can be generated, and the traffic time sequence feature of the candidate account can be determined according to the target sequence. The traffic time sequence feature can help more accurately identify the candidate account with the machine post feature. Accordingly, the traffic time sequence feature is used to identify the target account with abnormal post behavior from the candidate accounts, which can improve the accuracy of the identification result and reduce the missed judgment and misjudgment.

[0033] The candidate account can be any account that posts in the target application, and the target application usually refers to a community application such as a forum, a social media platform, etc.

[0034] In some embodiments of the present disclosure, the manner of recording the post behavior of each candidate account in the target application in real time can include: in response to determining that any candidate account has a post behavior, recording the following post data corresponding to the post behavior: the identity identifier (UID, User IDentifier) of the candidate account, the Internet protocol address (IP, Internet Protocol) used when posting, and the post time, and determining the UID or IP as the account identifier corresponding to the candidate account.

[0035] That is, for each post behavior of each candidate account, the corresponding post data can be recorded respectively, which can include the UID of the candidate account, the IP used when posting, and the post time, etc. How to obtain the post data is not limited. In addition, the UID or IP can be used as the identity identifier of the candidate account. For different candidate accounts, either the UID is used as the identity identifier, or the IP is used as the identity identifier.

[0036] Through the above real-time recording manner, the omission of post data can be avoided, thereby laying a good data foundation for subsequent processing.

[0037] The post data and the like in the embodiments of the present disclosure are not for a specific account and are not used to reflect the personal information of a specific account. In addition, the execution subject of the method of the present disclosure can obtain the post data through various public, legal and compliant manners, such as obtaining the post data in the case of authorization of the post account, etc. In the technical scheme of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of personal information are in line with the relevant legal regulations and do not violate public order and good customs.

[0038] Accordingly, after each first time window is passed, the post quantity of each candidate account in the first time window can be determined according to the recorded post data.

[0039] In some embodiments of the present disclosure, the posting quantity of each account identifier in the first time window can be determined according to the recorded posting data. That is, the posting quantity can be counted by time according to the UID dimension or the IP dimension. For example, if the length of the first time window is 1 hour, the posting quantity of each UID or IP in each hour can be counted respectively.

[0040] The specific length of the first time window can be determined according to actual needs, which is not limited to 1 hour, and can be adjusted at any time according to actual needs.

[0041] In addition, after the second time window, a corresponding target sequence can be generated for each candidate account. The specific length of the second time window can also be determined according to actual needs, such as one day.

[0042] For any candidate account, the corresponding target sequence can include the posting quantity of the candidate account in the N first time windows respectively. Figure 2 A schematic diagram of the target sequence of the present disclosure is shown in FIG. 1. As shown in FIG. 1, assuming that the length of the first time window is 1 hour and the length of the second time window is one day, the target sequence can include posting quantity 1, posting quantity 2, posting quantity 3,..., posting quantity 23 and posting quantity 24, wherein posting quantity 1 refers to the posting quantity of the candidate account between 0 and 1 o'clock, posting quantity 2 refers to the posting quantity of the candidate account between 1 and 2 o'clock, and the others are sequentially similar. Figure 2

[0043] According to the corresponding target sequence, the traffic time sequence feature of each candidate account can be determined. In some embodiments of the present disclosure, for any candidate account, the mean and the standard deviation of the N posting quantities in the corresponding target sequence can be determined, and then the sequence coefficient of variation (CV) can be determined according to the mean and the standard deviation, and the sequence coefficient of variation can be determined as the traffic time sequence feature of the candidate account.

[0044] The mean and the standard deviation are determined as follows.

[0045] The standard deviation is determined as follows.

[0046] x i represents the posting quantity in the i-th first time window, that is, the i-th posting quantity in the target sequence.

[0047] After obtaining the mean and the standard deviation, the sequence coefficient of variation can be determined in combination with the mean and the standard deviation. For example, the ratio of the standard deviation to the mean can be determined as the sequence coefficient of variation.

[0048] ​It can be seen that the sequence variation coefficient required can be obtained simply and quickly in the above manner. The sequence variation coefficient can be used to evaluate the dispersion degree of the post quantity of the candidate account in the time sequence, that is, to evaluate the time sequence dispersion degree of the post behavior of the candidate account, so that the candidate account with the machine post feature can be more accurately identified.

[0049] Correspondingly, the target account with abnormal post behavior can be determined from the candidate accounts according to the sequence variation coefficient.

[0050] In some embodiments of the present disclosure, the candidate account that meets the following condition can be screened from the candidate accounts: the corresponding sequence variation coefficient is less than a predetermined threshold. Then, the screened candidate account can be determined as the target account, or the screened candidate account can be determined as a suspicious account, and the target account can be determined from the suspicious account according to the post data of the suspicious account recorded in the second time window.

[0051] That is, for each candidate account, the following processing can be performed respectively: comparing the sequence variation coefficient of the candidate account with a preset threshold, and if it is determined that the sequence variation coefficient of the candidate account is less than the threshold, the candidate account can be taken as a screened candidate account. The specific value of the threshold can be determined according to actual needs, such as 0.5, and can be adjusted at any time according to actual needs.

[0052] Suppose that 20 candidate accounts are screened, then the 20 candidate accounts can be directly determined as target accounts, or the 20 candidate accounts can be determined as suspicious accounts, and then the target account can be determined from the 20 suspicious accounts according to the post data of the 20 suspicious accounts recorded in the second time window. The specific way can be determined according to actual needs, which is very flexible and convenient. Preferably, the latter way is adopted to further improve the accuracy of the identified target account.

[0053] In some embodiments of the present disclosure, when the target account is determined from the suspicious account, the post data of the suspicious account recorded in the second time window can be first determined as target data, and then the association relationship graph can be generated according to the target data, and the target account can be determined according to the association relationship graph.

[0054] Specifically, in some embodiments of the present disclosure, the post data can further include a post identifier (TID) of the posted post, i.e., the post data can include a UID, an IP, a TID, and a post time, etc. Accordingly, the UID, the IP, and the TID can be extracted from each target data, and the extraction results can be de-duplicated. Then, the remaining UID, IP, and TID after de-duplication can be determined as vertices in the association relationship graph, and for any two vertices, the two vertices can be connected by an edge in response to determining that the two vertices appear in the same target data.

[0055] As all the target data can be traversed, the UID, the IP, and the TID can be extracted therefrom, and the association relationship graph can be constructed by an adjacency list after de-duplication. Each vertex in the association relationship graph is a UID, an IP, or a TID, and the edges between the vertices are determined according to the association relationship between the UID, the IP, and the TID. For example, for a UID vertex and an IP vertex, if it is determined that the UID and the IP appear in the same target data, the two vertices can be connected by an edge. For example, for a UID vertex and a TID vertex, if it is determined that the UID and the TID appear in the same target data, the two vertices can be connected by an edge. For example, for an IP vertex and a TID vertex, if it is determined that the IP and the TID appear in the same target data, the two vertices can be connected by an edge.

[0056] The association relationship graph can be a weighted graph or an unweighted graph. If it is a weighted graph, the number of common occurrences can be used as the weight. For example, for a UID vertex and a TID vertex, if it is determined that the UID and the TID appear in 10 target data, 10 can be determined as the weight of the edge between the two vertices.

[0057] Through the association relationship graph, the association relationship between the UID, the IP, and the TID can be accurately reflected, thereby laying a good foundation for subsequent processing. In addition, in the scheme described in the present disclosure, the association relationship graph is taken as an example. If other graph construction methods are used, as long as the association relationship between the UID, the IP, and the TID can be accurately reflected, they are also acceptable.

[0058] Based on the association relationship graph, the target account can be determined. In some embodiments of the present disclosure, the UID in the association relationship graph can be clustered according to a K-clique algorithm, and then the target account can be determined according to the UID set obtained by clustering.

[0059] The specific value of K can be determined according to actual needs, such as 2. That is, the K-clique algorithm (K = 2) can be used to perform clique clustering on the UID-IP relationship. In specific implementation, all edges in the association relationship graph can be traversed to find a UID set that meets the K-clique condition. The number of UID sets can be one, multiple, or zero. The K-clique algorithm is a mature algorithm, and accordingly, using the algorithm to cluster the UIDs in the association relationship graph can improve the processing efficiency and the accuracy of the processing result.

[0060] In some embodiments of the present disclosure, for any UID set, a UID-TID relationship subgraph corresponding to the UID set can be extracted from the association relationship graph. Then, in response to determining that the UID-TID relationship subgraph is a fully connected relationship subgraph, the suspicious accounts corresponding to each UID in the UID set can be determined as target accounts.

[0061] In some embodiments of the present disclosure, for any UID set, a UID-TID relationship subgraph corresponding to the UID set can be extracted from the association relationship graph. Then, in response to determining that the UID-TID relationship subgraph is a fully connected relationship subgraph, the suspicious accounts corresponding to each UID in the UID set can be determined as target accounts.

[0062] In addition, in some embodiments of the present disclosure, for any UID-TID relationship subgraph corresponding to a UID set, in response to determining that the UID-TID relationship subgraph is a fully connected relationship subgraph, a team composed of suspicious accounts corresponding to each UID in the UID set can be determined as a target team identified.

[0063] The target team refers to an identified black production team, or black production clique. Black and gray production accounts usually have team behavior, that is, multiple black and gray production accounts form a black production team to post maliciously.

[0064] In the traditional way, when identifying a target team, only the association relationship between accounts and IPs is considered, and the association relationship between accounts and posts is ignored, resulting in poor accuracy of the identification result, such as splitting or missing the target team. However, in fact, the target team usually has a clear posting behavior pattern and clear division of labor. By analyzing the association relationship between accounts and posts, the target team can be more accurately identified.

[0065] In addition, the black and gray production account can use a public IP export. According to the traditional target team identification method, only the association between the account and the IP is considered, so that the normal account using the public IP export can be misjudged as a black and gray production account, thereby causing unnecessary disturbance to the normal account. In the scheme, the UID-TID relationship subgraph is combined to determine the target team, thereby reducing the misjudgment of the normal account, and further improving the accuracy of the identification result of the target team.

[0066] In actual application, the scheme can be continuously executed, such as performing target account and target team identification once every second time window.

[0067] In addition, the determined target account can be fed back to the audit system corresponding to the target application, and the target account can be processed by the audit system, such as being banned, limited to post, etc. For the target team, the countermeasures can be determined by analyzing the organization structure and activity rules, so as to improve the security of the target application.

[0068] In combination with the above introduction, Figure 3 The flowchart of the second embodiment of the target account identification method of the present disclosure is shown in FIG. 3. Figure 3 As shown in FIG. 3, the following specific implementation modes are included.

[0069] In step 301, in response to determining that any candidate account has a post behavior in the target application, the following post data corresponding to the post behavior is recorded: the UID of the candidate account, the IP used for posting, the UID of the post published, and the posting time.

[0070] In step 302, in response to determining that a first time window has passed, the posting quantity of each candidate account in the first time window is determined according to the recorded post data.

[0071] In step 303, in response to determining that a second time window has passed, the second time window includes N consecutive first time windows, and N is a positive integer greater than 1. For each candidate account, steps 304-305 are performed.

[0072] In step 304, a target sequence corresponding to the candidate account is generated, and the target sequence includes the posting quantity of the candidate account in N first time windows.

[0073] In step 305, the mean and standard deviation of the N posting quantities in the target sequence are determined, and the sequence coefficient of variation of the candidate account is determined according to the mean and standard deviation.

[0074] In step 306, the candidate account that meets the following condition is screened out from each candidate account: the corresponding sequence variation coefficient is less than a predetermined threshold, and the screened-out candidate account is determined as a suspicious account.

[0075] In step 307, the posting data of the suspicious account in the recorded second time window is determined as target data, the UID, IP and TID are extracted from each target data, and the extracted result is de-duplicated, the remaining UID, IP and TID after de-duplication are determined as vertices in the association relationship graph, and for any two vertices, in response to determining that the two vertices appear in the same target data, the two vertices are connected through an edge.

[0076] In step 308, according to the K-clique algorithm, the UIDs in the association relationship graph are clustered.

[0077] In step 309, for any UID set obtained by clustering, the corresponding UID-TID relationship subgraph is extracted from the association relationship graph, in response to determining that the relationship subgraph is a fully connected relationship subgraph, the suspicious accounts corresponding to each UID in the UID set are all determined as target accounts, and the team composed of the suspicious accounts corresponding to each UID in the UID set is determined as a target team.

[0078] For the identified target accounts and target teams, corresponding processing measures can be taken subsequently, and the first time window, the second time window and the threshold can be adjusted according to the processing effect, so as to better adapt to the actual demand and improve the accuracy and processing efficiency of the processing result of subsequent processing.

[0079] It should be noted that for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited by the order of the described actions, because according to the present disclosure, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present disclosure. In addition, the parts not described in detail in a certain embodiment can refer to the related description in other embodiments.

[0080] The above is the introduction of the method embodiment, and the following will further illustrate the scheme of the present disclosure through the device embodiment.

[0081] Figure 4 The constituent structure schematic diagram of the target account identification device embodiment 400 of the present disclosure is shown in FIG. 4. Figure 4 As shown in FIG. 4, it includes a recording module 401, a statistical module 402 and an identification module 403.

[0082] The recording module 401 is configured to record, in real time, posting behaviors of each candidate account in the target application.

[0083] The statistical module 402 is configured to, in response to determining that the first time window has elapsed, determine, according to the recorded content, posting quantities of each candidate account in the first time window, respectively.

[0084] The identification module 403 is configured to, in response to determining that the second time window has elapsed, the second time window including consecutive N first time windows, N being a positive integer greater than 1, generate, for each candidate account, a corresponding target sequence, the target sequence including posting quantities of the corresponding candidate account in the N first time windows, respectively, determine, according to the target sequence, a traffic time sequence feature of each candidate account, and determine, according to the traffic time sequence feature, a target account having abnormal posting behaviors from the candidate accounts.

[0085] In some embodiments of the present disclosure, the recording module 401 can record, in real time, the posting behaviors of each candidate account in the target application in the following manner: in response to determining that any candidate account has a posting behavior, recording posting data corresponding to the posting behavior, the posting data including a UID of the candidate account, an IP used for posting, and a posting time, and determining the UID or the IP as an account identifier corresponding to the candidate account. Accordingly, when determining, according to the recorded content, the posting quantities of each candidate account in the first time window, respectively, the statistical module 402 can determine, according to the posting data, the posting quantities of each account identifier in the first time window, respectively.

[0086] The identification module 403 can generate, for each candidate account, a corresponding target sequence after the second time window has elapsed, wherein for any candidate account, the corresponding target sequence can include, respectively, posting quantities of the candidate account in the N first time windows. According to the corresponding target sequence, the identification module 403 can also determine a traffic time sequence feature of each candidate account, respectively.

[0087] In some embodiments of the present disclosure, for any candidate account, the identification module 403 can determine, respectively, a mean value and a standard deviation of the N posting quantities in the corresponding target sequence, then determine a sequence coefficient of variation according to the mean value and the standard deviation, and further determine the sequence coefficient of variation as the traffic time sequence feature of the candidate account.

[0088] In some embodiments of the present disclosure, the identification module 403 can further determine the target account with abnormal posting behavior from the candidate accounts according to the sequence variation coefficient. Specifically, the identification module 403 can filter the candidate accounts that meet the following condition from the candidate accounts: the corresponding sequence variation coefficient is less than a predetermined threshold, and then determine the filtered candidate accounts as the target account, or determine the filtered candidate accounts as suspicious accounts, and determine the target account from the suspicious accounts according to the recorded posting data of the suspicious accounts in the second time window.

[0089] In some embodiments of the present disclosure, when determining the target account from the suspicious accounts, the identification module 403 can first determine the recorded posting data of the suspicious accounts in the second time window as target data, and then generate an association graph according to the target data, and further determine the target account according to the association graph.

[0090] In some embodiments of the present disclosure, the posting data can further include the TID of the published post, that is, the posting data can include UID, IP, TID and posting time, and the like. Accordingly, the identification module 403 can extract UID, IP and TID from each target data, and perform deduplication processing on the extraction result, and then determine the remaining UID, IP and TID after deduplication as the vertices in the association graph, and connect any two vertices through an edge in response to determining that the two vertices appear in the same target data.

[0091] Based on the association graph, the target account can be determined. In some embodiments of the present disclosure, the identification module 403 can cluster the UIDs in the association graph according to the K-clique algorithm, and then determine the target account according to the UID set obtained by clustering.

[0092] In some embodiments of the present disclosure, for any UID set, the identification module 403 can extract the UID-TID relationship subgraph corresponding to the UID set from the association graph, and then determine the suspicious accounts corresponding to each UID in the UID set as the target account in response to determining that the UID-TID relationship subgraph is a fully connected relationship subgraph.

[0093] In addition, in some embodiments of the present disclosure, for any UID-TID relationship subgraph corresponding to a UID set, the identification module 403 can further determine a team composed of the suspicious accounts corresponding to each UID in the UID set as a target team identified in response to determining that the UID-TID relationship subgraph is a fully connected relationship subgraph.

[0094] The specific working process of the above-mentioned device embodiments can refer to the related description in the foregoing method embodiments, which will not be described here.

[0095] In summary, by using the scheme provided in the disclosure, the target account and the target team can be identified in a timely manner through real-time monitoring and analysis of the posting behavior of the candidate account, and the accuracy of the identification result can be improved based on the traffic time sequence characteristics and the correlation relationship graph, and accordingly, the professionalism and order of the forum can be maintained, the social media environment can be purified, and the network security can be improved.

[0096] The scheme provided in the disclosure can be applied to the field of artificial intelligence, and particularly relates to the fields of network security and knowledge graph. Artificial intelligence is a discipline that studies enabling a computer to simulate some thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) of a person, and includes both hardware technologies and software technologies. The artificial intelligence hardware technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, and big data processing, and the artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0097] According to the embodiments of the disclosure, the disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0098] Figure 5 A schematic block diagram of an electronic device 500 that can be used to implement embodiments of the disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, servers, servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0099] As shown in Figure 5 The electronic device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded into a random access memory (RAM) 503 from a storage unit 508. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0100] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0101] The computing unit 501 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphic processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded onto the RAM 503 and executed by the computing unit 501, one or more steps of the methods described in the present disclosure can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the methods described in the present disclosure by any other appropriate means, such as by means of firmware.

[0102] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0103] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as part of a separate software package, and partially on a remote machine or server.

[0104] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a linearly-programmed electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0105] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0106] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0107] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0108] It should be understood that the various forms of flow shown above can be re-ordered, added to, or have steps deleted, for example. The steps recited in the present disclosure can be performed in parallel, in series, or in different orders, as long as the desired results of the technology disclosed in the present disclosure are achieved, and the present disclosure is not limited herein.

[0109] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for identifying target accounts, comprising: Record the posting behavior of each candidate account in the target application in real time; In response to the determination that the first time window has passed, the number of posts made by each candidate account within the first time window is determined based on the recorded content. In response to determining that a second time window has passed, the second time window includes N consecutive first time windows, where N is a positive integer greater than 1. For each candidate account, a corresponding target sequence is generated. The target sequence includes the number of posts made by the corresponding candidate account in the N first time windows. The traffic time sequence characteristics of each candidate account are determined according to the target sequence. Based on the traffic time sequence characteristics, the target account with abnormal posting behavior is determined from each candidate account.

2. The method according to claim 1, wherein, The real-time recording of posting behavior of each candidate account in the target application includes: in response to determining that any candidate account has posted, recording the following posting data corresponding to the posting behavior: the candidate account's identity identifier UID, the Internet Protocol address (IP) used when posting, and the posting time, and determining the UID or the IP as the account identifier corresponding to the candidate account; The step of determining the number of posts by each candidate account within the first time window based on the recorded content includes: determining the number of posts by each account identifier within the first time window based on the posting data.

3. The method according to claim 2, wherein, The step of determining the traffic time-series characteristics of each candidate account based on the target sequence includes: For any candidate account, determine the mean and standard deviation of the number of posts in the corresponding target sequence for each of the N posts; The coefficient of variation is determined based on the mean and the standard deviation, and the coefficient of variation is used as the traffic time-series characteristic of the candidate account.

4. The method according to claim 3, wherein, The target accounts identified from each candidate account based on the traffic time-series characteristics include: Candidate accounts that meet the following criteria are selected from all candidate accounts: the corresponding sequence variation coefficient is less than a predetermined threshold; The selected candidate accounts are determined as the target accounts, or the selected candidate accounts are determined as suspicious accounts, and the target accounts are determined from the suspicious accounts based on the posting data of the suspicious accounts within the recorded second time window.

5. The method according to claim 4, wherein, The step of identifying the target account from the suspicious accounts includes: The posting data of the suspicious account within the recorded second time window is identified as target data, and a relationship graph is generated based on the target data; The target account is determined based on the relationship diagram.

6. The method according to claim 5, wherein, The posting data also includes: the post identifier (TID) of the posted post; The step of generating the relationship diagram based on the target data includes: Extract the UID, IP, and TID from each target data, and perform deduplication on the extraction results; The remaining UID, IP, and TID after deduplication are determined as vertices in the association graph, and for any two vertices, in response to determining that the two vertices appear in the same target data, the two vertices are connected by an edge.

7. The method according to claim 6, wherein, The step of determining the target account based on the relationship diagram includes: The UIDs in the association graph are clustered according to the K-cluster clustering algorithm; The target account is determined based on the set of UIDs obtained from clustering.

8. The method according to claim 7, wherein, The determination of the target account based on the UID set obtained from clustering includes: For any set of UIDs, extract the corresponding UID-TID relationship subgraph from the relationship graph; In response to determining that the UID-TID relationship subgraph is a fully connected relationship subgraph, all suspicious accounts corresponding to each UID in the UID set are identified as the target accounts.

9. The method according to claim 8, further comprising: In response to determining that the UID-TID relationship subgraph is a fully connected relationship subgraph, the team composed of suspicious accounts corresponding to each UID in the UID set is identified as the target team.

10. A target account identification device, comprising: The module includes a recording module, a statistics module, and a recognition module. The recording module is used to record the posting behavior of each candidate account in the target application in real time; The statistics module is used to determine the number of posts made by each candidate account within the first time window based on the recorded content in response to the determination that the first time window has passed. The identification module is configured to, in response to determining that a second time window has passed, the second time window includes N consecutive first time windows, where N is a positive integer greater than 1, generate a corresponding target sequence for each candidate account, the target sequence including: the number of posts made by the corresponding candidate account in the N first time windows, determine the traffic time sequence characteristics of each candidate account according to the target sequence, and determine the target account with abnormal posting behavior from each candidate account according to the traffic time sequence characteristics.

11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-9.

13. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the method of any one of claims 1-9.