Account identification method and device, computer device, and storage medium
By analyzing the content posted by target accounts and combining features at the levels of content similarity and frequency patterns, a neural network model is used to automatically identify accounts that repost content. This solves the problems of low efficiency and low accuracy of manual review in existing technologies, and achieves efficient and accurate account identification.
Patent Information
- Application Number
- CN202111033044.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-03
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-09-03
AI Technical Summary
Current technologies rely on manual review to identify accounts that have been copied, which is inefficient and inaccurate, making it difficult to effectively identify whether an account is a copy account.
The first type of plagiarism features is statistically analyzed based on the content posted by the target account. The second type of plagiarism features are constructed by combining the frequent itemsets in the sample account set. A predetermined neural network model is used for automatic identification, and account identification is performed by combining features at the content similarity level and the frequent pattern level.
It improves the efficiency and accuracy of identifying accounts that have been illegally accessing content, avoids the inefficiency and insufficient recall of manual review, and achieves more accurate account identification.
Smart Images

Figure CN115757766B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to an account identification method and device, computer equipment and storage medium. BACKGROUND
[0002] In the information flow platform, there are a large number of carrying accounts. The carrying account refers to an account that generates and publishes content by relying on carrying, copying others' creation or performing some low-cost creation, such as a video account, a public account, a social account and the like. Obviously, the carrying account is not conducive to the reasonable development of the original account and the ecological health of the platform.
[0003] In the related art, the identification of the carrying account usually relies on the manual auditing mode to determine whether the account carries the content of others. The efficiency of manual auditing is very low, and it is inaccurate to directly conclude whether it is a carrying account according to a small amount of published content. SUMMARY
[0004] Therefore, it is necessary to provide an account identification method, device, computer equipment and storage medium capable of improving the identification efficiency and accuracy of the carrying account, and to provide a processing method and device of an account identification model, computer equipment and storage medium.
[0005] An account identification method, the method comprising:
[0006] Based on the published content corresponding to the target account, the first type of carrying feature of the target account in the content similarity layer is counted.
[0007] Based on the published content corresponding to each sample account in the sample account set, a target attribute frequent item set set about the carrying account is determined, and the sample account set includes the carrying account and the reference account.
[0008] According to the matching degree of the published content corresponding to the target account and the target attribute frequent item set set, the second type of carrying feature of the target account in the frequent pattern layer is obtained.
[0009] Based on the first type of carrying feature and the second type of carrying feature, the target account is identified to obtain the identification result of whether the target account is a carrying account.
[0010] An account identification device, the device comprising:
[0011] A first feature counting module is configured to count, based on the published content corresponding to the target account, the first type of carrying feature of the target account in the content similarity layer.
[0012] The frequent item set determination module is configured to determine a target attribute frequent item set set of the forwarding account based on the publishing content corresponding to each sample account in a sample account set, the sample account set including the forwarding account and the reference account;
[0013] The second feature statistics module is configured to obtain a second type of forwarding feature of the target account at a frequent pattern level according to a matching degree of the publishing content corresponding to the target account and the target attribute frequent item set set;
[0014] The identification module is configured to identify the target account based on the first type of forwarding feature and the second type of forwarding feature, and obtain an identification result of whether the target account is a forwarding account.
[0015] A computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the account identification method when executing the computer program
[0016] A computer readable storage medium stores a computer program, and the computer program implements the steps of the account identification method when executed by a processor.
[0017] A computer program includes computer instructions stored in a computer readable storage medium, a processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the steps of the account identification method.
[0018] The account identification method, device, computer device, and storage medium described above, for a target account, on the one hand, based on the publishing content corresponding thereto, aggregates the similarity relationship between the contents to the account dimension, mines the first type of forwarding feature of the target account at the content similarity level, and on the other hand, by focusing on analyzing the publishing content of the forwarding account and the reference account in the sample account set, constructs a target attribute frequent item set set of the forwarding account, according to the matching degree of the target attribute of the publishing content corresponding to the target account and the target attribute frequent item set set constructed, mines the frequent pattern of the target attribute of the target account as the second type of forwarding feature of the target account at the frequent pattern level. Then, the first type of forwarding feature and the second type of forwarding feature can be combined to automatically identify the target account and obtain an identification result of whether the target account is a forwarding account. Not only can the problem of low efficiency and low accuracy of manual review be avoided, but also by using all the publishing content of the target account to mine the frequent pattern, a content set is identified, and when the target account has certain forwarding behavior characteristics, it can be identified, which can avoid the problem of insufficient recall rate caused by identification from the content dimension.
[0019] A processing method of an account identification model, the method comprising:
[0020] obtaining unidentified accounts and reference accounts;
[0021] screening, from the unidentified accounts, accounts having an account similarity relationship with any of the reference accounts, and based on published content corresponding to the screened accounts, counting a first type of carrying probability of the screened accounts at a content similarity level, and determining a first carrying account from the screened accounts according to the first type of carrying probability;
[0022] based on the published content corresponding to the first carrying account and the reference accounts, determining a target attribute frequent item set sample set about carrying accounts;
[0023] obtaining, according to a matching degree of the published content corresponding to the unidentified accounts and the target attribute frequent item set, a second type of carrying probability of the unidentified accounts at a frequent pattern level, and determining a second carrying account from the unidentified accounts according to the second type of carrying probability;
[0024] based on a sample account set composed of the first carrying account, the second carrying account, and the reference accounts, determining a target attribute frequent item set set about carrying accounts and a target attribute frequent item set set about reference accounts;
[0025] using the published content corresponding to each sample account in the sample account set and the target attribute frequent item set set about carrying accounts and the target attribute frequent item set set about reference accounts, performing model training on a predetermined neural network model to obtain an account identification model for identifying carrying accounts.
[0026] A processing device of an account identification model, the device comprising:
[0027] an account obtaining module configured to obtain unidentified accounts and reference accounts;
[0028] a first carrying account mining module configured to screen, from the unidentified accounts, accounts having an account similarity relationship with any of the reference accounts, and based on published content corresponding to the screened accounts, count a first type of carrying probability of the screened accounts at a content similarity level, and determine a first carrying account from the screened accounts according to the first type of carrying probability;
[0029] a frequent item set determining module configured to determine, based on the published content corresponding to the first carrying account and the reference accounts, a target attribute frequent item set sample set about carrying accounts;
[0030] The second carrying account mining module is configured to obtain a second type of carrying probability of the unidentified account at a frequent pattern level according to a matching degree of the publishing content corresponding to the unidentified account and the target attribute frequent item set collection, and determine a second carrying account from the unidentified account according to the second type of carrying probability.
[0031] The frequent item set determination module is further configured to determine a target attribute frequent item set collection about a carrying account and a target attribute frequent item set collection about a reference account based on a sample account set composed of the first carrying account, the second carrying account, and the reference account.
[0032] The training module is configured to perform model training on a predetermined neural network model using publishing content corresponding to each sample account in the sample account set and the target attribute frequent item set collection about the carrying account and the target attribute frequent item set collection about the reference account, to obtain an account identification model for identifying a carrying account.
[0033] A computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the steps of the processing method of the account identification model when executing the computer program.
[0034] A computer readable storage medium stores a computer program, and the computer program implements the steps of the processing method of the account identification model when executed by a processor.
[0035] A computer program includes computer instructions stored in a computer readable storage medium, a processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the steps of the processing method of the account identification model.
[0036] The processing method and device of the account identification model, the computer device, and the storage medium divide the training process of the account identification model into three stages. In the first stage, after screening out accounts having a similar relationship with the reference account, the first type of transfer probability in the content similarity level is obtained through analysis and calculation of content similarity relationship and other characteristics, and the transfer account in the first stage is identified. In the second stage, based on the reference account and the transfer account obtained in the previous stage, a target attribute frequent item set sample set about the transfer account is mined, and the corresponding second type of transfer probability is obtained according to the matching degree of the content published by the unidentified account and the frequent item set, and the transfer account in the second stage is identified. In the third stage, after the target attribute frequent item set set about the transfer account and the target attribute frequent item set set about the reference account are re-determined using the reference account and the transfer account identified in the previous two stages, the transfer account identified in the previous two stages and the applied characteristics are used to train the predetermined neural network model, so that the predetermined neural network model learns the artificial experience and summarizes the high-order characteristics possessed by the transfer account, and the account identification model is obtained, thereby identifying more transfer accounts. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 An application environment diagram of an account identification method in an embodiment;
[0038] Figure 2 A flowchart of an account identification method in an embodiment;
[0039] Figure 3 A processing framework diagram of an account identification method in an embodiment;
[0040] Figure 4 A flowchart of obtaining the first type of transfer feature in an embodiment;
[0041] Figure 5 A framework diagram of obtaining the first type of transfer feature corresponding to the target account in an embodiment;
[0042] Figure 6 A flowchart of obtaining the target attribute frequent item set set in an embodiment;
[0043] Figure 7 A flowchart of obtaining the frequent item set set about the theme and the category in an embodiment;
[0044] Figure 8 A flowchart of obtaining the second type of transfer feature in an embodiment;
[0045] Figure 9 A framework diagram of obtaining the second type of transfer feature corresponding to the target account in an embodiment;
[0046] Figure 10 a schematic diagram of a training framework of a three-stage account identification model in one embodiment;
[0047] Figure 11 a schematic diagram of an identification process of a first transfer account in one embodiment;
[0048] Figure 12 a schematic diagram of a process of determining whether two accounts have a similar relationship in one embodiment;
[0049] Figure 13 a schematic diagram of a framework of identifying a first transfer account in one embodiment;
[0050] Figure 14 a schematic diagram of an identification process of a second transfer account in one embodiment;
[0051] Figure 15 a schematic diagram of a framework of identifying a second transfer account in one embodiment;
[0052] Figure 16 a schematic diagram of a training process of an account identification model in one specific embodiment;
[0053] Figure 17 a schematic diagram of a processing method of an account identification model in one embodiment;
[0054] Figure 18 a structural block diagram of an account identification device in one embodiment;
[0055] Figure 19 a structural block diagram of a processing device of an account identification model in one embodiment;
[0056] Figure 20 an internal structural diagram of a computer device in one embodiment. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0058] The method provided in the application relates to artificial intelligence (AI) technology. Artificial intelligence technology is a theory, method, technology and application system for using a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive an environment, acquire knowledge and use the knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning and decision-making.
[0059] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning, etc.
[0060] Machine learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. This subject is dedicated to studying how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to enabling computers to have intelligence. It is applied in various fields of artificial intelligence. Artificial neural network is an important machine learning technology, which has a wide application prospect in system identification, pattern recognition, intelligent control, etc. For example, in the present application, each sample account in the sample account set can be used to train a predetermined neural network model to obtain an account identification model for identifying a transfer account. The sample account includes a transfer account and a reference account. The reference account can be an official account or an original account.
[0061] The method provided in the application can also involve natural language processing (NLP) technology. Natural language processing is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science and mathematics. For example, in the present application, natural language processing technology can be used to identify the categories of the title and the body of the published content to obtain the corresponding categories.
[0062] The method provided in the application can also use computer vision (CV) technology, which simulates biological vision using computers and related devices. Its main task is to obtain three-dimensional information of the corresponding scene by processing the collected pictures or videos, just like what humans and many other living beings do every day. For example, in the present application, computer vision technology can be used to mine the similarity between the cover pictures of videos or texts. Computer vision technology can also be used to identify the category of a video.
[0063] The method provided in the embodiments of the present application also relates to some terms and terminologies, which are explained and described as follows:
[0064] Terminal: The terminal can be, but is not limited to, various personal computers, notebook computers, smart televisions, smart homes, smart phones, tablet computers, vehicle-mounted terminals and portable wearable devices.
[0065] Server: It can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.
[0066] User: A person who uses a terminal to use computer services or Internet services. Usually, a user needs to register a corresponding user account on an information flow platform to publish generated publication content to the network.
[0067] Carrying account: An account that generates publication content by relying on carrying, copying others' creations or performing some low-cost creations, such as video accounts, public accounts, social accounts, etc.
[0068] Account owner: An institution or individual who publishes content through an account in a content generation platform (also known as an information flow platform). The publication content, videos or articles created or carried by the account owner can be browsed by users and generate corresponding user behaviors.
[0069] Terminal program: Also known as a client or an application program, it is an application program with specific functions, such as QQ, WeChat, Today's Headlines, etc.
[0070] Information flow product: A product form of a terminal program, through which various video and article information can be obtained, and video and article information can be recommended to users.
[0071] User behavior: The collective term for the operations of browsing, clicking, commenting, liking, following, etc. in the information flow product.
[0072] Number owner: The number owner is the person who owns the account on the information flow platform and has the right to publish, delete, change, and comment on the content.
[0073] Content category: The content category is used to describe the general characteristics and similarities of the published content, such as movies, variety shows, technology, sports, education, and society.
[0074] Content tag: The content tag is used to represent the concepts and content actually contained in the published content, including detailed information about the characters, locations, and scenes involved in the published content.
[0075] p-norm: The p-norm refers to the p-norm in the real number field, which calculates the sum of the p-th power of each value in a vector, and then takes the 1 / p-th power of the result, which is the length or size of a vector in a certain vector space. When p is equal to 0, the resulting value is the number of non-zero values in the vector, when p is equal to 1, the resulting value is the sum of the absolute values of all elements in the vector, when p is equal to 2, the resulting value is the square root of the sum of the squares of all elements in the vector, and when p is equal to positive infinity, the resulting value is the absolute value of the maximum value in the vector. Therefore, the larger the value of p, the more attention the p-norm pays to the maximum value in the vector, and the smaller the value of p, the more attention the resulting value pays to the number of non-zero values in the vector.
[0076] p-norm aggregation: Calculate the sum of the p-th power of each value in a vector, and then take the 1 / p-th power of the result, which is the aggregation result. For example, assume that p = 3, vector s = (s1, s2,..., s 12 ), then the aggregation result is
[0077] Video classification: Video classification refers to identifying the category of content contained in a given video clip.
[0078] Wilson score: The Wilson score is the lower bound of the Wilson confidence interval for a binomial distribution sample. Assuming a sample of samples, each sample takes the value 0 or 1, and the proportion of 1 is calculated. Since the true proportion is often unknown, it is assumed that the sample follows a Bernoulli distribution, and the estimate of this proportion is calculated based on the central limit theorem.
[0079] The formula for calculating the Wilson score is as follows:
[0080] n = u + v;
[0081] p = u / n;
[0082]
[0083] Wherein, u represents the number of samples with value A, v represents the number of samples with value B, n represents the number of samples, p represents the proportion of samples with value A, and z α is the quantile of normal distribution, and S represents the Wilson score. z α Generally, the value is 2, that is, the confidence level is 95%. It can be understood that A and B are only used for illustration, and u and v are not respectively equal to A and B. For example, A is 1, and B is 0. For another example, A represents a hit, and B represents a miss.
[0084] Frequent pattern: a frequent pattern is a set of items, subsequence or substructure that appears in a dataset with a frequency not less than a specified threshold. For example, all the goods in a supermarket are a dataset, and diapers and beer can constitute a frequent item set because they account for a relatively high proportion among all the people who buy them. For example, a computer is purchased first, then a digital camera, and then a memory card. If this purchase order often appears in a shopping history database, it is a frequent sequential pattern, that is, a subsequence. In different structural forms, such as subgraph, subtree or subgrid, these structural forms can be combined with item sets or subsequences. If a substructure often appears in a graph database, it is a frequent substructure.
[0085] Expert model: an expert model is an organic combination of rules and models. What the expert model does is to artificially encode the business experience of an expert into the structure of a machine learning model, and can realize the reuse and rapid deployment of the same business scenario.
[0086] The account identification method provided in the application can be applied to an application environment as shown in Figure 1 The terminal 102 communicates with the server 104 through a network. In an embodiment, the server 104 can count a first type of transfer feature of a target account on a content similarity level based on the publishing content corresponding to the target account, determine a target attribute frequent item set set about a transfer account based on the publishing content corresponding to each sample account in a sample account set, the sample account set including the transfer account and a reference account, obtain a second type of transfer feature of the target account on a frequent pattern level according to the matching degree of the publishing content corresponding to the target account and the target attribute frequent item set set, and identify the target account based on the first type of transfer feature and the second type of transfer feature to obtain an identification result of whether the target account is a transfer account.
[0087] The processing method of the account identification model provided in the embodiment of the application can be applied to Figure 1In the illustrated application environment. For example, the server 104 can obtain the unidentified account and the reference account; filter the accounts having an account similarity relationship between any reference account from the unidentified account, based on the publishing content corresponding to the filtered accounts, count the first type of carrying probability of the filtered accounts in the content similarity layer, determine the first carrying account from the filtered accounts according to the first type of carrying probability; based on the publishing content corresponding to the first carrying account and the reference account, determine the target attribute frequent item set sample set of the carrying account; according to the matching degree of the publishing content corresponding to the unidentified account and the target attribute frequent item set set, obtain the second type of carrying probability of the unidentified account in the frequent mode layer, determine the second carrying account from the unidentified account according to the second type of carrying probability; based on the sample account set composed of the first carrying account, the second carrying account and the reference account, determine the target attribute frequent item set set of the carrying account and the target attribute frequent item set set of the reference account; using the publishing content corresponding to each sample account in the sample account set and the target attribute frequent item set set of the carrying account and the target attribute frequent item set set of the reference account, the predetermined neural network model is trained to obtain an account identification model for identifying carrying accounts.
[0088] In other embodiments, the above-mentioned account identification method and the processing method of the account identification model can be implemented by the terminal 102 or the server 104 alone, or by the terminal 102 and the server 104 together, for example, the terminal 102 can upload the publishing content corresponding to the account to the server 104, the server 104 can respond to the carrying account identification request of the terminal 102, or the server 104 can periodically trigger the identification of the carrying account, or the server 104 can randomly trigger the identification of the carrying account. That is, based on the publishing content corresponding to the target account, count the first type of carrying feature of the target account in the content similarity layer and the second type of carrying feature in the frequent mode layer, identify the target account based on the first type of carrying feature and the second type of carrying feature, and return the obtained identification result to the terminal 102.
[0089] The account identification method provided by the embodiments of the present application can, for a target account, on the one hand, aggregate the similarity relationship between contents to the account dimension based on the corresponding publishing contents of the target account, and mine the first type of carrying characteristics of the target account in the content similarity layer; on the other hand, construct a target attribute frequent item set set about the carrying account by focusing on analyzing the publishing contents of the carrying account and the reference account in the sample account set, mine the frequent pattern about the target attribute of the target account as the second type of carrying characteristics of the target account in the frequent pattern layer according to the target attribute of the publishing contents corresponding to the target account and the matching degree of the constructed target attribute frequent item set set. Then, the target account can be automatically identified by combining the first type of carrying characteristics and the second type of carrying characteristics, and the identification result of whether the target account is a carrying account can be obtained. Not only the problems of low efficiency and low accuracy of manual review can be avoided, but also the frequent pattern of the target account can be mined by using all the publishing contents of the target account, which is equivalent to identifying a content set, so that the target account can be identified when the target account has certain carrying behavior characteristics, and the problem of insufficient recall rate caused by identification in the content dimension can be avoided.
[0090] In one embodiment, as shown in Figure 2 , an account identification method is provided. The method is applied to a computer device (terminal 102 or server 104) in Figure 1 for example, and includes the following steps:
[0091] Step 202, based on the publishing contents corresponding to the target account, the first type of carrying characteristics of the target account in the content similarity layer is counted.
[0092] In the present application, the account is an account that has the function of publishing content in the information flow platform, and the published content is referred to as publishing content. The account can be a media account or a self-media account, more specifically, a video account, a public account, a social account, etc. The account is used to uniquely identify a user. The target account is an account to be identified as a carrying account. In order to ensure the reasonable development of the account in the information flow platform and the health of the platform ecology, the computer device can identify the carrying account of the account in the platform.
[0093] The first type of carrying characteristics is the characteristics of the account in the content similarity layer. The first type of carrying characteristics is obtained by analyzing and extracting all the publishing contents corresponding to the target account based on expert experience. The computer device can periodically, periodically, or under a specified trigger condition, identify the carrying account of the account, that is, obtain all the publishing contents corresponding to the target account, count the first type of carrying characteristics of the target account in the content similarity layer based on the publishing contents corresponding to the target account. It can be understood that, with the development of time, the publishing contents corresponding to the target account can be updated due to the modification operation, deletion operation or publishing new content operation of the user.
[0094] Specifically, the computer device extracts the published content corresponding to the target account when it needs to identify the account carrying account of the target account, pre-processes the published content based on expert experience, extracts all features of the target account in the content similarity layer, and fuses all features in the content similarity layer to obtain the first type of carrying features of the target account in the content similarity layer. The features in the content similarity layer can at least include the reference vocabulary hit degree, the field distribution concentration degree and the content carrying degree corresponding to the target account.
[0095] Step 204, based on the published content corresponding to each sample account in the sample account set, determine the target attribute frequent item set set of the carrying account, the sample account set includes the carrying account and the reference account.
[0096] Among them, the sample account set is a set of sample accounts used in the training phase of the account identification model, and the sample account set includes the carrying account and the reference account. The carrying account can be used as a positive sample in model training, and the reference account can be used as a negative sample in model training. The reference account can be some official account or original account. It should be noted that the carrying account in the sample account set is identified based on expert experience, which will be introduced and explained in the training process of the account identification model in the following.
[0097] After the computer device obtains the sample account set, the sample accounts in the sample account set represent expert experience, so the computer device can focus on analyzing the published content corresponding to the carrying account and the reference account in the sample account set, respectively, and mine the frequent pattern features of the carrying account and the reference account, so as to extract the features of the target account from the frequent pattern level.
[0098] The published content as a kind of media product has attributes, including the theme of the published content, the category of the published content, the label of the published content, and the title, the body, the cover picture, the label, the video OCR text, etc. of the published content. Since the carrying account and the reference account have differences in the frequent patterns of these attributes, the computer device can mine the frequent item sets of the published content corresponding to the carrying account and the reference account in the target attribute, and extract the target attribute frequent item set set of the carrying account and the frequent item set set of the reference account from these frequent item sets according to expert experience. The target attribute can be one or more, for example, the target attribute can be the theme, or the category.
[0099] In one embodiment, the computer device can extract target attributes of the publishing content corresponding to the carrying account, and use a frequent item set mining algorithm, such as an FP-Growth algorithm, an Apriori algorithm, etc., to mine 1 frequent item set, 2 frequent item set, …, N frequent item set of the target attributes of the carrying account. Similarly, 1 frequent item set, 2 frequent item set, …, N frequent item set of the target attributes of the reference account are mined, where N represents the number of items in each frequent item set. Based on expert experience, a set of frequent item sets of the target attributes of the carrying account is extracted from the frequent item sets, and a set of frequent item sets of the reference account can also be extracted. The carrying account is illustrated as follows based on the target attributes as the theme:
[0100] The theme words of the first publishing content are w1, w2, w3, w4, and w5;
[0101] The theme words of the second publishing content are w2, w3, w4, w5, and w6;
[0102] The theme words of the third publishing content are w3, w4, w5, w6, and w7;
[0103] The theme words of the fourth publishing content are w11, w2, w14, w15, and w16;
[0104] Mining 1 frequent item set: There are a total of 4 publishing contents,
w2
w2
[0105] Mining 2 frequent item set: There are a total of 4 publishing contents,
w3, w4
w3, w4
[0106] Step 206: According to the matching degree of the publishing content corresponding to the target account and the set of frequent item sets of the target attributes, a second type of carrying feature of the target account at the frequent pattern level is obtained.
[0107] The second type of carrying feature is a feature of the target account at a frequent pattern level. Specifically, after obtaining the target attribute frequent item set set of the carrying account and the target attribute frequent item set set of the reference account, the computer device can calculate, for the target account, a matching degree of the published content of the target account and the target attribute frequent item set set based on the target attribute of the published content corresponding to the target account, as the second type of carrying feature of the target account. The second type of carrying feature is a feature of the account publishing behavior pattern, which, in combination with the first type of carrying feature of the account publishing at a content similarity level, can multi-dimensionally describe the features of the target account, thereby improving the accuracy of the carrying account identification.
[0108] At step 208, the target account is identified based on the first type of carrying feature and the second type of carrying feature, to obtain an identification result of whether the target account is a carrying account.
[0109] After the first type of carrying feature and the second type of carrying feature of the target account are extracted through the first two steps, the computer device can identify the target account as a carrying account based on the features at the two levels, to obtain an identification result.
[0110] In an embodiment, the computer device can input the first type of carrying feature and the second type of carrying feature into a trained account identification model, mine low-order and high-order cross features hidden in the features through a feature fusion layer of the account identification model, and obtain an identification result based on the mined cross features through a classification prediction layer of the account identification model. The account identification model can be a deepFM model, a PNN model, or a Wide&Deep model, etc.
[0111] In an embodiment, step 208 includes: performing feature fusion on the first type of carrying feature and the second type of carrying feature through a feature fusion layer of the account identification model, to obtain a fusion feature corresponding to the target account; and performing classification based on the fusion feature through a classification prediction layer of the account identification model, to obtain an identification result of whether the target account is a carrying account.
[0112] Specifically, the computer device can process the carrying feature corresponding to the target account through a trained account identification model, to identify whether the account is a carrying account. The account identification model is a neural network model based on machine learning, which can be learned through training samples, thereby having a specific ability. In this embodiment, the account identification model is a pre-trained model having a carrying account identification ability.
[0113] In an embodiment, the computer device can construct a predetermined neural network model based on the initial model parameters, train the predetermined neural network model through each account sample in the account set, and obtain trained model parameters when a training stop condition is met. When identifying the target account, the model parameters can be obtained and imported into the predetermined neural network model to obtain an account identification model that can identify the target account. When identifying the account, the carrying features corresponding to the target account need to be extracted as input features of the account identification model in the manner described in the foregoing embodiments.
[0114] In an embodiment, the computer device can further obtain a third type of carrying feature of the target account according to the number of times that the target attributes in the published content of the target account hit the target attribute frequent item set set of the carrying account and the target attribute frequent item set set of the reference account, respectively. The third type of carrying feature is also a feature of the target account at the frequent pattern level.
[0115] In an embodiment, the target attributes include the theme of the published content and the category of the published content, the target attribute frequent item set set of the carrying account includes a first theme frequent item set set of the carrying account and a first category frequent item set set of the carrying account, and the target attribute frequent item set set of the reference account includes a second theme frequent item set set of the reference account and a second category frequent item set set of the reference account. The account identification method further includes obtaining a third type of carrying feature of the target account according to the number of times that the theme in the published content of the target account hits each frequent item set in the first theme frequent item set set and the second theme frequent item set set, respectively, and the number of times that the category hits each frequent item set in the first category frequent item set set and the second category frequent item set set, respectively.
[0116] Specifically, after obtaining the theme frequent item set set and the category frequent item set set of the carrying account and the reference account, respectively, the computer device determines the theme and the category of all published content of the target account, counts the number of times that each frequent item set in the first theme frequent item set set, the second theme frequent item set set, the first category frequent item set set, and the second category frequent item set set is hit, respectively, and obtains the third type of carrying feature of the target account based on the number of times.
[0117] In an embodiment, the computer device can generate a frequent item set bag of words about the published content according to the first topic frequent item set set, the second topic frequent item set set, the first category frequent item set set and the second category frequent item set set; obtain the topic vocabulary and the category of each piece of published content corresponding to the target account; determine the hit times of each frequent item set in the frequent item set bag of words corresponding to the target account according to the topic vocabulary and the category of each piece of published content; and generate a bag of words vector corresponding to the target account as the third type of carrying feature of the target account according to the hit times of each frequent item set corresponding to the target account.
[0118] Specifically, the computer device can generate a frequent item set bag of words about the frequent patterns involved in the published content of the account according to all the frequent item sets in the four sets, obtain the topic vocabulary and the category of the corresponding published content for the target account, and count the times of the topic vocabulary and the category hitting each frequent item set in the bag of words, or the times of the frequent items hitting each frequent item set in the bag of words, to obtain a corresponding bag of words vector. The bag of words vector is used as the third type of carrying feature of the target account, and is used together with the first type of carrying feature and the second type of carrying feature to identify whether the account is a carrying account.
[0119] For example, the bag of words formed by the four sets includes d frequent item sets in total, involving m frequent items: the first frequent item set
w1 w2 w3
w2 w3
w4
w5 w6
wm
[0120] For the first frequent item set
w1 w2 w3
[0121] In an embodiment, the computer device obtains the topic vocabulary {w1 w2 w3} of the published content A and the topic vocabulary {w1 w2 w3 w5} of the published content B, and can consider that both hit the first frequent item set
w1 w2 w3
[0122] In an embodiment, the computer device can also count the frequent item set of the topic vocabulary and the frequent item set of the category of all published content, respectively. If the counted frequent item set includes the three frequent item sets
w1 w2 w3
[0123] Accordingly, step 208 includes: fusing the first type of reposting features, the second type of reposting features, and the third type of reposting features to obtain a fused feature; and identifying the target account based on the fused feature to obtain an identification result of whether the target account is a reposting account.
[0124] In this embodiment, in addition to using the first type of copying features and the second type of copying features as input features of the account recognition model, the bag-of-words vector corresponding to the target account is also used as input features of the account recognition model. By inputting all features into the account recognition model, the recognition result of the target account is more accurate because it takes into account the frequent patterns of copying accounts and non-copying accounts.
[0125] like Figure 3 The diagram shown illustrates the processing framework of an account recognition method in one embodiment. (Refer to...) Figure 3 Before identification, a sample account set is obtained, including accounts that copy content and reference accounts. The sample account set is used to mine frequent itemsets related to the topic and category of accounts that copy content, and to mine frequent itemsets related to the topic and category of accounts that reference accounts. Each frequent itemset in these four sets is combined to obtain a bag-of-frequent-itemsets. Using each sample account and its corresponding published content, the four sets, and the bag-of-frequent-itemsets, a predetermined neural network model based on machine learning is trained to obtain an account identification model. When identifying a target account as a copying account, the corresponding first-type, second-type, and third-type copying features are extracted using the target account's published content, the four sets, and the bag-of-frequent-itemsets. These features are then used as input features to the trained account identification model to obtain the identification result.
[0126] The account identification method described above, for a target account, on the one hand, based on the corresponding published content, aggregates the similarity relationship between the content to the account dimension, mines the first type of carrying characteristics of the account in the content similarity layer, on the other hand, by focusing on the published content of the carrying account and the reference account in the sample account set, a target attribute frequent item set set about the carrying account is constructed, according to the matching degree of the target attribute of the published content corresponding to the target account and the target attribute frequent item set set constructed, the frequent pattern about the target attribute is mined as the second type of carrying characteristics of the target account in the frequent pattern layer. Then, by combining the first type of carrying characteristics and the second type of carrying characteristics, the target account can be automatically identified, and the identification result of whether the target account is a carrying account is obtained. Not only can the problem of low efficiency and low accuracy of manual review be avoided, but also by using all the published content of the target account to mine the frequent pattern, a content set is identified, and when the target account has certain carrying behavior characteristics, it can be identified, which can avoid the problem of insufficient recall rate caused by identifying from the content dimension.
[0127] In one embodiment, as shown in FIG. 4, step 202 includes steps 402-406: Figure 4
[0128] Step 402, based on the published content corresponding to the target account, the reference vocabulary hit degree, the field distribution concentration degree and the content carrying degree corresponding to the target account are counted.
[0129] Among them, the reference vocabulary hit degree is used to represent the degree of account plagiarism of the reference account. The more times the reference vocabulary hits in the published content corresponding to an account, the more likely the account is a carrying account. The reference vocabulary is one or more exclusive vocabularies constructed for the reference account, also known as original vocabulary. The reference vocabulary can be a vocabulary or a sentence. The reference vocabulary can be derived from the name of the reference account. The reference vocabulary can also be extracted from the published content corresponding to the reference account, for example, the tags of the published content can be identified, and the words or short sentences with a tf-idf value greater than a threshold value are identified from the identified tags as reference vocabularies. The reference vocabulary can also be reported by artificial means, and the present application does not limit this. For example, the reference vocabulary of the reference account "Wang Xiaoming" is "Wang Xiaoming", and the reference vocabulary of the reference account "Beijing Six O'Clock" is "Beijing Six O'Clock".
[0130] In an embodiment, the statistical step of the reference vocabulary hit degree corresponding to the target account comprises: obtaining a reference vocabulary generated according to the published content of a reference account; obtaining attribute information of each published content of the target account, obtaining the number of times of hitting the reference vocabulary according to the attribute information, and obtaining the number of times of hitting the reference vocabulary for each published content according to the attribute information; and calculating the reference vocabulary hit degree corresponding to the target account according to the total number of times of hitting the reference vocabulary of the published content corresponding to the target account and the number of published contents of the target account.
[0131] Specifically, the computer device can obtain a set of reference vocabularies constructed for all reference accounts, obtain all published contents corresponding to the target account for the target account, traverse the attribute information of each published content, and count the number of times of hitting the reference vocabulary according to the attribute information. The attribute information can include the title, cover picture, text, label, and video OCR text obtained by character recognition of the video subtitle of the published content, and the like. If the reference vocabulary of the reference account is traversed from the attribute information, it is determined that the target account hits the reference vocabulary.
[0132] More specifically, different hit scores can be set for different attribute information based on expert experience. For example, for each published content, if the label hits the reference vocabulary, the number of times of hitting the reference vocabulary of the published content is directly recorded as 1, if the label does not hit but the title hits the reference vocabulary, the number of times of hitting the reference vocabulary of the published content is directly recorded as 0.8, if the label does not hit and the title does not hit, the text or the video OCR text hits the reference vocabulary, the number of times of hitting is recorded as 0.5, otherwise the number of times is 0.
[0133] It should be noted that the embodiments of the present application do not limit the way of extracting the attribute information of the published content. For example, the computer device can use an image processing model to extract the video OCR text, and can use a text classification model to identify the label and category.
[0134] In an embodiment, after the computer device obtains the total number of times of hitting the reference vocabulary of the published content corresponding to the target account, the Wilson score s is calculated according to the total number of times u and the number of published contents n of the target account, and the score s is taken as the reference vocabulary hit degree corresponding to the target account.
[0135] The field distribution concentration degree is used to represent the concentration degree of the fields to which the published contents corresponding to the account belong. If the fields to which the published contents corresponding to an account belong are uneven, such as food, technology, and parenting, the more dispersed and larger the span of the fields, the more likely it is that the published contents corresponding to the account are obtained by transferring other people's works.
[0136] In an embodiment, the statistical step of the field distribution concentration degree corresponding to the target account includes: performing category detection on the publishing content corresponding to the target account to obtain the category of each publishing content; and taking the ratio of the number of publishing content belonging to the target category to the total number of publishing content of the target account as the field distribution concentration degree corresponding to the target account, where the number of categories in the publishing content corresponding to the target account belonging to the target category is the largest.
[0137] Specifically, the computer device can use a text classification model to obtain the category corresponding to the publishing content. More specifically, the computer device can input the title, cover picture, video OCR text, and other content of the publishing content into the text classification model to obtain the category of the publishing content. For a target account, the target category with the largest number in the publishing content is determined, and the proportion of the number belonging to the target category to the total number of the publishing content is taken as the field distribution concentration degree.
[0138] The content carrying degree represents the proportion of the publishing content corresponding to the account that has a similar relationship with the publishing content corresponding to other accounts. If a part of the publishing content of an account is carried from other accounts, the account has the possibility of being a carrying account, and other publishing content of the account can be carrying content. The larger the proportion of the publishing content carried from other accounts of the account, the greater the possibility of the account being a carrying account.
[0139] In an embodiment, the statistical step of the content carrying degree of the target account includes: respectively obtaining the attribute feature vectors of each attribute of the publishing content of the target account and the reference account; calculating the vector similarity between the attribute feature vectors corresponding to the same attribute for two publishing contents, and obtaining the content similarity between the two publishing contents according to the vector similarity of multiple attributes; wherein the two publishing contents are respectively from the target account and the reference account, and the publishing time of the publishing content from the reference account is earlier than that of the publishing content from the target account; counting the number of publishing content in the target account with a content similarity higher than a threshold; and calculating the content carrying degree corresponding to the target account according to the number and the total number of publishing content of the target account.
[0140] The attributes of the published content can include cover pictures, titles, content, texts, and the like. For the cover pictures, the computer device can input the cover pictures of the two published content into a siamese network model to directly calculate the similarity, or can use an image processing model to extract cover picture vectors and then calculate the similarity. For the titles, the computer device can input the titles of the two published content into a BERT model to extract title vectors and then calculate the similarity. For the content, the computer device can input the content (including cover pictures, titles, videos, audios, texts, and the like) of the two published content into a multi-modal feature extraction model to extract content vectors and then calculate the similarity. For the texts, mainly the text words and pictures in the texts, the computer device can also input the texts into a multi-modal feature extraction model to extract text vectors and then calculate the similarity.
[0141] After obtaining the vector similarity between the attribute feature vectors corresponding to the same attribute of the two published content, the computer device obtains the content similarity between the two published content according to the vector similarities corresponding to multiple attributes. For example, the similarity between the title vectors is s1, the similarity between the cover picture vectors is s2, the similarity between the content vectors is s3, and the similarity between the text vectors is s4. Then, the average of the four numbers or the weighted average is taken as the content similarity between the two published content.
[0142] Among the two published content, one published content is from the target account, and the other published content is from the reference account, that is, the content similarity between any two published content of the target account and the reference account. The computer device screens the content similarity whose publishing time is earlier than the publishing time of the corresponding reference account, and counts the number of published content whose content similarity is higher than a threshold, and calculates the content carrying degree corresponding to the target account according to the number and the total number of published content of the target account.
[0143] For example, the reference account A publishes content a1, a2, and a3, and each content has a publishing time. The target account B publishes content b1, b2, b3, and b4, and each content has a publishing time. Then, 12 content similarities between A and B will be calculated, and the computer device will screen according to the publishing time. Assuming that the publishing time of b1 is earlier than that of a1 and a2, the content similarity between b1 and a1 and the content similarity between b1 and a2 will be filtered out, leaving 10 content similarities. Among the 10 content similarities, 5 are greater than the threshold, and all of them are corresponding to the published content b2 and b3 of the target account. Therefore, the number of published content whose content similarity is higher than the threshold is 2, the total number of published content of the target account B is 4, and the content carrying degree corresponding to the target account is calculated as 2 / 4.
[0144] In one embodiment, after obtaining the content similarity between any two posts from two accounts, the computer device will have multiple content similarities between different posts from any two accounts. The computer device can then aggregate these multiple content similarities to obtain the account similarity between the two accounts. For example, the computer device can use the P-norm to aggregate the multiple content similarities between any two accounts.
[0145] Step 404: The target account's reference keyword hit rate, domain distribution concentration, and content plagiarism rate are weighted and summed to obtain the first type of plagiarism probability.
[0146] Specifically, the computer device performs a weighted summation of three features of the target account at the content similarity level to obtain the first type of plagiarism probability of the target account. The weight of each feature can be customized based on expert experience.
[0147] Step 406: The degree of reference word hit, the degree of concentration of domain distribution, the degree of content plagiarism, and the probability of first-type plagiarism are used as the first-type plagiarism features of the target account in terms of content similarity.
[0148] It can be assumed that the first type of plagiarism features include four parts: the degree of reference word hit, the degree of domain distribution concentration, the degree of content plagiarism, and the probability of first-type plagiarism. Computer devices can use these four parts as the first type of plagiarism features of the target account at the level of content similarity, and use them to identify the target account as a plagiarism account in the future.
[0149] In one embodiment, the computer device may also use only the degree of reference word hit, the degree of domain distribution concentration, and the degree of content copying as the first type of copying feature of the target account at the content similarity level.
[0150] like Figure 5 The diagram shown illustrates a framework for obtaining the first type of data transfer characteristics corresponding to a target account in one embodiment. (Refer to...) Figure 5The required data includes reference accounts, their corresponding published content, and reference keywords. For the target account, obtain all its published content. For each published content, obtain the corresponding cover image, title, body text, and video OCR text. Compare these attribute information from all published content of the target account with the reference keywords to obtain the keyword hit rate for the target account. Using these attribute information from all published content of the target account, perform category detection to obtain the domain distribution concentration for the target account. Use these attribute information from all published content of the target account to perform a similarity comparison with the attribute information from all published content of the reference account to obtain the content plagiarism rate for the target account. Finally, calculate the first-type plagiarism probability based on the keyword hit rate, domain distribution concentration, and content plagiarism rate, and use these factors as the first-type plagiarism feature for the target account.
[0151] In the above embodiments, based on the content published by the account, the similarity relationship between the content is aggregated to the account dimension, and the first type of copying feature at the content similarity level is mined, which can improve the accuracy of copying account identification. As long as an account's published content is similar to the published content of the reference account, or an account copies the recently published content of the reference account, or an account has copying features at the content similarity level, it will be identified.
[0152] In one embodiment, such as Figure 6 As shown, step 204 includes the following steps 602-604:
[0153] Step 602: Based on the topics of the published content corresponding to each sample account in the sample account set, determine the first topic frequent itemset set for the reposting account and the second topic frequent itemset set for the reference account.
[0154] Step 604: Based on the category of the published content corresponding to each sample account in the sample account set, determine the first category frequent itemset set for the reposting account and the second category frequent itemset set for the reference account.
[0155] In this embodiment, taking the target attributes as theme and category as an example, based on the theme and category of each published content, we mine frequent itemsets and frequent itemsets of the theme and category of the reposting account, and mine frequent itemsets and frequent itemsets of the theme and category of the reference account. Then, we extract the first theme frequent itemset set and the second theme frequent itemset set of the reference account, as well as the first category frequent itemset set of the reposting account and the second category frequent itemset set of the reference account from these frequent itemsets.
[0156] In one embodiment, such asFigure 7 As shown, step 204 includes steps 702-708.
[0157] At step 702, based on the publishing content corresponding to the carrying account in the sample account set, the first topic frequent item set about the carrying account and the first category frequent item set about the carrying account are extracted.
[0158] In an embodiment, step 702 includes: after extracting the attribute information of the publishing content corresponding to the carrying account, performing word segmentation on the attribute information and counting the importance of each word, determining the topic word of the publishing content according to the importance, and performing category detection on the publishing content to obtain the category of the publishing content; according to the topic word of each piece of publishing content corresponding to the carrying account, the first topic frequent item set about the carrying account is extracted, and according to the category of each piece of publishing content corresponding to the carrying account, the first category frequent item set about the carrying account is extracted.
[0159] Specifically, for each sample account in the sample account set, the attribute information such as the title, label, body, and ocr text data of the publishing content corresponding to the carrying account and the reference account is obtained, and the topic word and the category are determined according to the attribute information. The computer device can use word segmentation technology to perform word segmentation on the title, body, label, and ocr text data, and then use tf-idf to determine important words therefrom as the topic of a piece of publishing content. The computer device can use a text classification model to perform category detection on the attribute information to obtain the category of the publishing content. For each detected category of the publishing content, the computer device also performs some category merging, merging similar or similar categories into the same category, for example, technology and science, education and parenting, etc.
[0160] After the computer device obtains the topic word of all publishing content corresponding to the carrying account, the first topic frequent item set about the carrying account is extracted. For example, the computer device can use a frequent item set mining algorithm such as the FP-Growth algorithm, the Apriori algorithm, etc., to mine the 1 frequent item set, the 2 frequent item set, …, the N frequent item set about the topic of the carrying account according to the topic word of all publishing content of the carrying account. The value of N can be customized, for example, the value is 5. Similarly, after the computer device obtains the category of all publishing content corresponding to the carrying account, the first category frequent item set about the carrying account is extracted.
[0161] At step 704, based on the publishing content corresponding to the reference account in the sample account set, the second topic frequent item set about the reference account and the second category frequent item set about the reference account are extracted.
[0162] In one embodiment, step 704 comprises: extracting attribute information of the publishing content corresponding to the reference account, performing word segmentation on the attribute information and counting the importance of each word, determining the theme word of the publishing content according to the importance, and performing category detection on the publishing content to obtain the category of the publishing content; extracting the second theme frequent item set about the reference account according to the theme word of each publishing content corresponding to the reference account, and extracting the second category frequent item set about the reference account according to the category of each publishing content corresponding to the reference account.
[0163] Similarly, after the computer device obtains the theme word of all publishing content corresponding to the reference account, the second theme frequent item set about the reference account is extracted. Similarly, after the computer device obtains the category of all publishing content corresponding to the reference account, the second category frequent item set about the reference account is extracted.
[0164] Step 706, a set of frequent item sets belonging to the first theme frequent item set and not belonging to the second theme frequent item set is taken as the first theme frequent item set set about the transfer account, and a set of frequent item sets belonging to the second theme frequent item set and not belonging to the first theme frequent item set is taken as the second theme frequent item set set about the reference account.
[0165] Specifically, the computer device can take a set of frequent item sets that have appeared in the first theme frequent item set of the transfer account but have not appeared in the second theme frequent item set of the reference account as the first theme frequent item set set about the transfer account. Similarly, a set of frequent item sets that have appeared in the second theme frequent item set of the reference account but have not appeared in the first theme frequent item set of the transfer account is taken as the second theme frequent item set set about the reference account.
[0166] In one embodiment, the first theme frequent item set set about the transfer account also has another source. For a frequent item set A that appears in both the first theme frequent item set of the transfer account and the second theme frequent item set of the reference account, if the support degree of the frequent item set A from the transfer account is p1 and the support degree of the frequent item set A from the reference account is p2, the computer device can add the frequent item set with a support degree ratio p1 / p2 greater than a preset threshold to the first theme frequent item set set about the transfer account.
[0167] Similarly, regarding the second topic frequent itemset set of the reference account, the computer device can determine the frequent itemset B that appears simultaneously in the first topic frequent itemset set of the copy account and the second topic frequent itemset set of the reference account. If the support of the frequent itemset B from the copy account is p1 and the support of the frequent itemset B from the reference account is p2, the computer device can add the frequent itemsets with a support ratio p2 / p1 greater than a preset threshold to the second topic frequent itemset set of the reference account.
[0168] Step 708: The set of frequent itemsets belonging to the first category but not to the second category is taken as the first category frequent itemset set for the purpose of transferring accounts, and the set of frequent itemsets belonging to the second category but not to the first category is taken as the second category frequent itemset set for the purpose of the purpose of referencing accounts.
[0169] The method for determining the first category of frequent itemsets for reposting account categories and the second category of frequent itemsets for reference account categories is the same as that for the first topic of frequent itemsets for reposting account topics and the second topic of frequent itemsets for reference account topics, and will not be repeated here.
[0170] In this embodiment, by mining the frequent pattern features of the reposting account and the reference account regarding themes and categories, we can focus on analyzing the differences in the behavioral patterns of the reposting account and the reference account, which facilitates the subsequent extraction of the target account's features from the frequent pattern level.
[0171] In one embodiment, such as Figure 8 As shown, step 206 includes the following steps 802 to 808:
[0172] Step 802: Based on the matching degree between the target account and the frequent itemset set of the first topic, obtain the topic hit rate of the reposting account corresponding to the target account.
[0173] In one embodiment, step 802 includes: extracting the topic keywords of the published content based on the attribute information of the published content corresponding to the target account; counting the first number of published content whose topic keywords belong to the first topic frequent itemset set; and calculating the topic hit rate of the reposting account corresponding to the target account based on the first number and the total number of published content corresponding to the target account.
[0174] Specifically, the computer device obtains the topic keywords of all the content published by the target account, and counts the number u1 of the topic keywords that match the first frequent itemset set of the reposting account. Based on this number u1 and the total number n of the content published by the target account, the Wilson score is calculated, and this score is used as the topic hit degree of the reposting account corresponding to the target account.
[0175] At step 804, the carrying account category hit degree corresponding to the target account is obtained according to the matching degree of the target account and the first category frequent item set set.
[0176] In an embodiment, step 804 comprises: performing category detection on the publishing content according to the attribute information of the publishing content corresponding to the target account, to obtain the category of the publishing content; counting the second quantity of the publishing content whose category belongs to the first category frequent item set set, and calculating the carrying account category hit degree corresponding to the target account according to the second quantity and the total quantity of the publishing content corresponding to the target account.
[0177] Specifically, the computer device obtains the category of all publishing content corresponding to the target account, and counts the quantity u2 of the categories in the publishing content that hit the first category frequent item set set of the carrying account, and calculates the Wilson score according to the quantity u2 and the total quantity n of the publishing content corresponding to the target account, and takes the score as the carrying account category hit degree corresponding to the target account.
[0178] At step 806, the carrying account theme hit degree and the carrying account category hit degree are weighted and summed to obtain the second type of carrying probability.
[0179] Specifically, the computer device weights and sums the two features of the target account at the frequent pattern level to obtain the second type of carrying probability of the target account. The weight of each feature can be customized according to expert experience.
[0180] At step 808, the carrying account theme hit degree, the carrying account category hit degree, and the second type of carrying probability are taken as the second type of carrying feature of the target account.
[0181] It can be considered that the second type of carrying feature includes the carrying account theme hit degree, the carrying account category hit degree, and the second type of carrying probability, and the computer device can take the three parts as the second type of carrying feature of the target account at the frequent pattern level, which is used for subsequent identification of the target account as a carrying account.
[0182] In an embodiment, the computer device can also only take the carrying account theme hit degree and the carrying account category hit degree as the second type of carrying feature of the target account at the frequent pattern level.
[0183] As shown in Figure 9 , it is a framework schematic diagram for obtaining the second type of carrying feature of the target account in an embodiment. Referring to Figure 9 , the data to be prepared includes a sample account set and the publishing content corresponding to each sample account in the set.
[0184] For the sample account set, the corresponding attribute information of the transfer account and the reference account is obtained, such as title, label, text, video OCR text and other data, the topic vocabulary and category of each published content are mined, and then the frequent item set mining algorithm is used to mine the frequent item set of the topic and category of the transfer account and the reference account, and the first topic frequent item set and the first category frequent item set of the transfer account are extracted.
[0185] For the target account, all published contents corresponding to the target account are obtained. For each published content, the corresponding attribute information is obtained, and the topic vocabulary and category of each published content are extracted. According to the number of hits of the first topic frequent item set and the first category frequent item set, the transfer account topic hit degree and the transfer account category hit degree corresponding to the target account are obtained. Finally, the second transfer probability is calculated according to the transfer account topic hit degree and the transfer account category hit degree, and the transfer account topic hit degree, the transfer account category hit degree and the second transfer probability are used as the second transfer feature of the target account.
[0186] In the above embodiment, for the obtained sample account set, the frequent item set set of the transfer account topic and category and the frequent item set set of the reference account topic and category are extracted based on the published content corresponding to the sample account in the sample account set, which is helpful to mine the second transfer feature and the third transfer feature of the target account at the frequent pattern level. As long as an account has the behavior characteristics of the transfer account at the frequent pattern level, it will be identified.
[0187] In one embodiment, the account identification model is obtained by using the transfer account and the reference account in the sample account set to train a predetermined neural network model. The transfer account includes a first transfer account and a second transfer account. The first transfer account is an account with a first transfer probability greater than a first threshold value. The second transfer account is an account with a second transfer probability greater than a second threshold value. The second transfer probability is calculated based on the target attribute frequent item set sample set of the transfer account determined based on the first transfer account and the reference account.
[0188] In the embodiments of the present application, a three-stage expert model paradigm is proposed, which combines expert experience and the fitting ability of supervised machine learning model to train an expert model with certain interpretability and generalization ability. The first stage is expert experience analysis and identification, which performs the first round of identification. The second stage focuses on analysis and identification, analyzes the identification results of the previous stage, has certain data selection and bias, relies on the completeness of expert experience, and performs the second round of identification. The third stage is generalization, which trains a complex model to learn high-order features based on the identification results of the previous two stages to achieve the purpose of generalization. The three stages are progressive, step by step increasing the number of identifications, starting from expert experience and generalizing to machine learning model. This paradigm can not only be used in the identification of account transfer, but also be reused in other problems in similar scenarios without labels or with few labels. In particular, due to the integration of expert experience, this paradigm also has certain interpretability.
[0189] Referring to Figure 10 Applying the three-stage expert model paradigm to the identification of account transfer, the specific scheme is: in the first stage, for the un-identified account and the reference account, the attribute information of the corresponding published content on the information stream platform is obtained, including the title, text, video, cover picture, video OCR text, etc. of the published content. These contents are obtained through different machine learning models to obtain the content similarity between the two published contents. Through the analysis and calculation of the content similarity between any two published contents of the un-identified account and the reference account, the first account transfer account is obtained in the first stage. In the second stage, based on the reference account and the first account transfer account obtained in the first stage, the attributes of the published content are analyzed and focused, and the features of the account transfer account and the reference account on the frequent pattern level are obtained. According to the analyzed features, the second account transfer account in the second stage is obtained from the un-identified account. In the third stage, the reference account is used as a negative sample, and the account transfer accounts obtained in the previous two stages are used as positive samples to form a sample account set. The account transfer features extracted in the previous two stages are applied to train a more complex predetermined neural network model, so that the model can capture the hidden low-order and high-order cross features between the account transfer features, thereby having strong generalization ability. After training, the account identification model is obtained, and the trained account identification model can be used to automatically identify more account transfer accounts.
[0190] In one embodiment, as shown in Figure 11 the identification steps of the first account transfer account include:
[0191] Step 1102, obtaining an un-identified account, and screening an account having an account similarity relationship with any reference account from the un-identified account.
[0192] The unidentified account is an account of which the carrying type is unknown. The computer device needs to extract corresponding carrying features from the batch of accounts according to expert experience, identify whether the accounts are carrying accounts, take the identified carrying accounts as training samples, and train the predetermined neural network model to obtain an account identification model.
[0193] Specifically, as shown in FIG. 11, in one embodiment, before step 1102, the above method further includes the step of determining whether there is an account similarity relationship between two accounts: Figure 12
[0194] Step 1202, respectively, obtaining attribute feature vectors of attributes of respective publishing contents of the unidentified account and the reference account.
[0195] Step 1204, for two publishing contents, calculating vector similarities between attribute feature vectors corresponding to the same attribute, and obtaining a content similarity between the two publishing contents according to the vector similarities corresponding to multiple attributes; wherein the two publishing contents are respectively from the unidentified account and the reference account.
[0196] The attributes of the publishing contents can include cover pictures, titles, contents, and texts. For the cover pictures, the computer device can input the cover pictures of the two publishing contents into the siamese network model to directly calculate the similarity, or can extract corresponding cover picture vectors using an image processing model and then calculate the similarity. For the titles, the computer device can input the titles of the two publishing contents into the bert model to extract corresponding title vectors and then calculate the similarity. For the contents, the computer device can input the contents (including cover pictures, titles, videos, audios, texts, etc.) of the two publishing contents into a multi-modal feature extraction model to extract corresponding content vectors and then calculate the similarity. For the texts, mainly the text words and pictures in the texts, the computer device can also input the texts into the multi-modal feature extraction model to extract corresponding text vectors and then calculate the similarity.
[0197] After obtaining the vector similarities between the attribute feature vectors corresponding to the same attribute of the two publishing contents, the computer device obtains the content similarity between the two publishing contents according to the vector similarities corresponding to multiple attributes. For example, the similarity between the title vectors is s1, the similarity between the cover picture vectors is s2, the similarity between the content vectors is s3, and the similarity between the text vectors is s4. Then, the average of the four numbers or the weighted average is taken as the content similarity between the two publishing contents.
[0198] In step 1206, the content similarities between the two published contents are aggregated to obtain the account similarity between the unidentified account and the reference account, and the unidentified account with an account similarity greater than a preset threshold is regarded as an account having an account similarity relationship with the reference account.
[0199] After the computer device obtains the content similarity between any two published contents of two accounts, different published contents between any two accounts have multiple content similarities. One of the any two published contents is from an unidentified account, and the other is from a reference account. Then, the computer device can aggregate the multiple content similarities between any unidentified account and reference account to obtain the account similarity between the two accounts. For example, the computer device can use the P-norm to aggregate the multiple content similarities between any two accounts. The computer device obtains a threshold set according to expert experience, and two accounts having a content similarity greater than the threshold have an account similarity relationship.
[0200] Through the above account similarity calculation, the computer device screens a batch of accounts having a similarity relationship with the reference account from the unidentified accounts, and continues the subsequent identification.
[0201] In step 1104, based on the published content corresponding to the screened account, the reference vocabulary hit degree, the field distribution concentration degree, and the content carrying degree are counted, the reference vocabulary hit degree, the field distribution concentration degree, and the content carrying degree corresponding to the screened account are weighted and summed to obtain the first type of carrying probability of the screened account.
[0202] The reference vocabulary hit degree is used to represent the degree of account plagiarism of the reference account. The more times the reference vocabulary hits in the published content corresponding to an account, the greater the possibility that the account is a carrying account. Specifically, the computer device can obtain a set of reference vocabularies constructed for all reference accounts. For an unidentified account, all published contents corresponding to the unidentified account are obtained. For each published content, the attribute information is traversed, and the number of times of hitting the reference vocabulary is counted. The attribute information can include the title, cover picture, text, label of the published content, and the text obtained by character recognition of the video internal subtitles (referred to as video OCR text), etc. If the reference vocabulary of the reference account is traversed from the attribute information, it is determined that the unidentified account hits the reference vocabulary.
[0203] More specifically, different hit scores can also be set for different attribute information based on expert experience. For example, for each published content, if the label hits the reference vocabulary, the number of times the published content hits the reference vocabulary is directly recorded as 1, if the label does not hit but the title hits the reference vocabulary, the number of times the published content hits the reference vocabulary is directly recorded as 0.8, if the label does not hit, the title does not hit, and the text or video OCR text hits the reference vocabulary, the number of hits is recorded as 0.5, otherwise the number is 0.
[0204] The field distribution concentration degree is used to represent the concentration degree of the fields to which the published content corresponding to the account belongs. If the fields to which the multiple published contents corresponding to an account belong are uneven, such as food, technology, and parenting, the more dispersed and larger the span of the fields, the more likely it is that the published content corresponding to the account is obtained by carrying other people's works. The computer device can obtain the categories of the published content using a text classification model. More specifically, the computer device can input the title, cover picture, video OCR text, and the like of the published content into the text classification model to obtain the category of the published content. For an unidentified account, the target category with the largest number of published contents is determined, and the proportion of the number of categories belonging to the target category to the total number of published contents is taken as the field distribution concentration degree.
[0205] The content carrying degree represents the proportion of the published content corresponding to the account that has a similar relationship with the published content corresponding to other accounts. If a part of the published content of an account is carried from other accounts, the account has a possibility of being a carrying account, and other published content of the account can be carrying content. The larger the proportion of the published content of the account carried from other accounts, the more likely it is that the account is a carrying account.
[0206] For the samples having a similar relationship with the reference account, the computer device screens the content similarity whose publishing time is earlier than the publishing time of the corresponding reference account, and counts the number of published contents whose content similarity is higher than a threshold value, and calculates the content carrying degree corresponding to the unidentified account according to the number and the total number of published contents of the unidentified account.
[0207] Then, the computer device performs weighted summation on the three features of the unidentified account in the content similarity level to obtain the first type of carrying probability of the unidentified account. The weight of each feature can be customized according to expert experience.
[0208] Step 1106, the samples in the screened account whose first type of carrying probability is greater than a first threshold value are taken as the first carrying account.
[0209] The computer device can obtain a first threshold value designed according to expert experience, and the un-identified account with a first type of carrying probability exceeding the first threshold value is considered as a first carrying account identified in the first stage. In addition, in order to ensure the high accuracy of the carrying account, the computer device can also obtain a pre-set account whitelist, which includes official accounts, government accounts, operation accounts, etc. For the identified first carrying account, the account in the account whitelist needs to be filtered out, and finally the first batch of high-accuracy carrying accounts is obtained.
[0210] As shown in Figure 13 , it is a framework diagram for identifying the first carrying account in an embodiment. Referring to Figure 13 , in the first step, the problem is analyzed and defined based on expert experience, that is, what is a carrying account. Generally speaking, a carrying account refers to an account that publishes original content of others without any secondary creation or with little secondary creation. In the second step, the content similarity between any two published contents is calculated, and the similar relationship is aggregated to the account dimension, so as to identify a batch of accounts having an account similar relationship with the reference account from the un-identified accounts. In the third step, for the batch of accounts, three features are calculated: reference vocabulary hit degree, field distribution concentration degree and content carrying degree. The first type of carrying probability is obtained by weighted summation of the three features. According to the first threshold value designed according to expert experience, the account exceeding the first threshold value is the first carrying account identified in the first stage. Further filtering is performed using the account whitelist, and finally the first batch of high-accuracy carrying accounts is obtained.
[0211] In the above embodiment, based on the published content corresponding to the account, the similar relationship between the contents is aggregated to the account dimension, and the first carrying account with high accuracy is mined from the content similarity level.
[0212] After the first carrying account is mined in the first stage, the attribute of the published content of the reference account and the first carrying account obtained in the first stage is further analyzed in the second stage, and the features of the carrying account and the reference account in the frequent pattern level are obtained. According to the analyzed features, the second carrying account in the second stage is obtained from the un-identified accounts.
[0213] In an embodiment, as shown in Figure 14 , the identification steps of the second carrying account include:
[0214] Step 1402, based on the theme and category of the published content corresponding to each first carrying account and reference account, a theme frequent item set sample set about the carrying account and a category frequent item set sample set about the carrying account are determined.
[0215] In an embodiment, the computer device extracts, based on the publishing content corresponding to the first re-posting account, a first topic frequent item set sample about the re-posting account and a first category frequent item set sample about the re-posting account; extracts, based on the publishing content corresponding to the reference account, a second topic frequent item set sample about the reference account and a second category frequent item set sample about the reference account. A set of each frequent item set sample belonging to the first topic frequent item set sample and not belonging to the second topic frequent item set sample is taken as a topic frequent item set sample set about the re-posting account, and a set of each frequent item set sample belonging to the first category frequent item set sample and not belonging to the second category frequent item set sample is taken as a category frequent item set sample set about the re-posting account category.
[0216] Specifically, for the reference account and the first re-posting account, the computer device respectively acquires corresponding attribute information such as title, label, body, and ocr text data according to the corresponding publishing content, and determines the topic vocabulary and the category according to the attribute information. The computer device can use word segmentation technology to perform word segmentation on the title, the body, the label, and the ocr text data, and then determine important words from the title, the body, the label, and the ocr text data using tf-idf as the topic of a publishing content. The computer device can use a text classification model to perform category detection on the attribute information to obtain the category of the publishing content. For each detected category of the publishing content, the computer device also performs some category merging, and merges similar or similar categories into the same category, for example, science and technology and science, education and parenting, and the like.
[0217] After the computer device acquires the topic vocabulary of all the publishing content corresponding to the first re-posting account, the computer device extracts the first topic frequent item set sample about the re-posting account. After the computer device acquires the category of all the publishing content corresponding to the first re-posting account, the computer device extracts the first category frequent item set sample about the re-posting account.
[0218] Similarly, after the computer device acquires the topic vocabulary of all the publishing content corresponding to the reference account, the computer device extracts the second topic frequent item set sample about the reference account. After the computer device acquires the category of all the publishing content corresponding to the reference account, the computer device extracts the second category frequent item set sample about the reference account.
[0219] Further, the computer device can take a set of frequent item set samples that appear in the first topic frequent item set sample of the re-posting account but do not appear in the second topic frequent item set sample of the reference account as a topic frequent item set sample set about the re-posting account. Similarly, the computer device can take a set of frequent item set samples that appear in the first category frequent item set sample of the re-posting account but do not appear in the second category frequent item set sample of the reference account as a category frequent item set sample set about the re-posting account.
[0220] In one embodiment, the frequent itemset sample set for the topic of a reposting account also has another source. For a frequent itemset sample A that appears simultaneously in the first frequent itemset sample set of the reposting account and the second frequent itemset sample set of the reference account, if the support of the frequent itemset sample A from the first reposting account is p1 and the support of the frequent itemset sample A from the reference account is p2, the computer device can add frequent itemset samples with a support ratio p1 / p2 greater than a preset threshold to the first frequent itemset sample set for the reposting account. Similarly, the extraction method for the category frequent itemset sample set for the reposting account is similar and will not be repeated here.
[0221] Step 1404: Based on the matching degree between the unidentified account and the frequent itemset sample set of the topic, obtain the topic hit degree of the reposting account corresponding to the unidentified account.
[0222] In one embodiment, the computer device extracts the topic keywords of the published content based on the attribute information of the published content corresponding to the unidentified account; counts the first number of published content whose topic keywords belong to the topic frequent itemset sample set; and calculates the topic hit rate of the reposting account corresponding to the unidentified account based on the first number and the total number of published content corresponding to the unidentified account.
[0223] Specifically, the computer device obtains the topic keywords of all published content corresponding to the unidentified account, and counts the number u1 of the number of frequent itemsets of the first topic related to the reposting account in the published content. Based on this number u1 and the total number n of published content corresponding to the unidentified account, the Wilson score is calculated, and this score is used as the topic hit degree of the reposting account corresponding to the unidentified account.
[0224] Step 1406: Based on the matching degree between the unidentified account and the frequent itemset sample set of the category, obtain the hit degree of the reposting account category corresponding to the unidentified account.
[0225] In one embodiment, the computer device performs category detection on the published content based on the attribute information of the published content corresponding to the unidentified account to obtain the category of the published content; counts the second number of published content whose category belongs to the category frequent itemset sample set; and calculates the category hit degree of the reposting account corresponding to the unidentified account based on the second number and the total number of published content corresponding to the unidentified account.
[0226] Specifically, the computer device obtains the categories of all published content corresponding to the unidentified account and counts the number u2 of frequent itemsets of categories that match the reposting account in the published content. Based on this number u2 and the total number n of published content corresponding to the unidentified account, the Wilson score is calculated, and this score is used as the category hit degree of the reposting account corresponding to the unidentified account.
[0227] Step 1410, the carrying account theme hit degree and the carrying account category hit degree are weighted and summed to obtain a second type carrying probability corresponding to the un-identified account.
[0228] Specifically, the computer device weights and sums two features of the un-identified account at the frequent pattern level to obtain the second type carrying probability of the un-identified account. The weight of each feature can be customized according to expert experience.
[0229] Step 1412, samples with the second type carrying probability greater than a second threshold value in the un-identified account are taken as second carrying accounts.
[0230] The computer device can obtain a second threshold value set according to expert experience, and the un-identified account with the second type carrying probability greater than the second threshold value is considered as a carrying account identified in the second stage.
[0231] As shown in Figure 15 , it is a framework diagram for identifying the second carrying account in an embodiment. Referring to Figure 15 , first, for the reference account and the first carrying account identified in the first stage, data preprocessing is performed: extracting title, label, category, video OCR text data, etc. For the title, label and video OCR text data, all the texts are segmented using the word segmentation technology, and the important words are calculated as the theme words of the published content using tf-idf. For the category, similar or similar categories are combined into the same one. Second, focus on analyzing the first carrying account, and use the frequent pattern mining method to construct multiple frequent item sets and multiple features: for the theme words and categories of the identified reference account and the first carrying account, 1 to N frequent item set samples about the theme and category of the carrying account, and 1 to N frequent item set samples about the theme and category of the reference account are mined, respectively. Then, according to the mined frequent item set samples, the theme frequent item set sample set and the category frequent item set sample set about the carrying account are extracted. Third, the number of contents hitting the two sets and the total number of published contents of the un-identified account are counted, and the Wilson scores are calculated respectively to obtain the score of the theme frequent item set sample set hitting the carrying account and the score of the category frequent item set sample set hitting the carrying account. The final second type carrying probability is obtained by using the weighted sum method, and the second carrying account is selected from the un-identified account.
[0232] With the carrying accounts identified in the first two stages based on expert experience, the samples are sufficient. In the third stage, these carrying accounts are used to train a complex predetermined neural network model, so that the model learns the artificial experience and summarizes the low-order and high-order cross features that cannot be understood by artificial, so that the model has generalization ability and identifies more carrying accounts.
[0233] AsFigure 16 FIG. 1 is a schematic diagram of a training process of an account identification model in a specific embodiment, as shown. Referring to FIG. 1, the training process of the account identification model includes the following steps: Figure 16 On the one hand, the account identification model is trained by using the accounts identified in the first and second stages as positive samples and using all the reference accounts as negative samples to construct a sample account set. On the other hand, the first and second types of carrying features extracted in the first two stages are used, and the first topic frequent item set collection about the carrying accounts and the first category frequent item set collection about the carrying accounts, the second topic frequent item set collection about the reference accounts and the second category frequent item set collection about the reference accounts are extracted according to all the reference accounts and the carrying accounts identified in the first and second stages. The extraction methods of the four collections are consistent with the extraction method of the topic frequent item set sample collection about the carrying accounts mentioned above, which will not be repeated here. According to the frequent item sets in the four collections, the frequent item set bag about the published content is generated, the number of hits of each frequent item set in the bag in the published content of the sample account is counted, and the corresponding bag vector is generated. This bag vector is input into the predetermined neural network model together with the first and second types of carrying features, so that the model can learn the high-order cross features hidden in the features. This implicit feature enables the model to have generalization ability. After training, the trained account identification model can be used to automatically identify more carrying accounts.
[0234] As shown in FIG. 2, it is a flowchart of a processing method of an account identification model in an embodiment. Referring to FIG. 2, the processing method of the account identification model includes the following steps: Figure 17 Figure 17 The following steps are included:
[0235] Step 1702, obtaining an unidentified account and a reference account.
[0236] Step 1704, screening the account that has an account similarity relationship with any reference account from the unidentified account, and counting the first type of carrying probability of the screened account in the content similarity layer based on the published content corresponding to the screened account. The first carrying account is determined from the screened account according to the first type of carrying probability.
[0237] Step 1706, determining a target attribute frequent item set sample collection about the carrying account based on the published content corresponding to the first carrying account and the reference account;
[0238] Step 1708, obtaining the second type of carrying probability of the unidentified account in the frequent pattern layer according to the matching degree of the published content corresponding to the unidentified account and the target attribute frequent item set collection, and determining the second carrying account from the unidentified account according to the second type of carrying probability;
[0239] Step 1710, based on the first carrying account, the second carrying account and the sample account set composed of the reference account, determine the target attribute frequent item set set about the carrying account and the target attribute frequent item set set about the reference account;
[0240] Step 1712, using the publishing content corresponding to each sample account in the sample account set and the target attribute frequent item set set about the carrying account and the target attribute frequent item set set about the reference account, model training is performed on the predetermined neural network model to obtain an account identification model for identifying the carrying account.
[0241] The processing method of the above account identification model divides the training process of the account identification model into three stages: in the first stage, after screening out the accounts having similar relationship with the reference account, the first type of carrying probability in the content similar level is obtained through analysis and calculation of the content similar relationship and other characteristics, and the carrying accounts in the first stage are identified. In the second stage, based on the reference account and the carrying account obtained in the previous stage, the target attribute frequent item set sample set about the carrying account is mined, and the corresponding second type of carrying probability is obtained according to the matching degree of the publishing content of the un-identified account and the frequent item set, and the carrying accounts in the second stage are identified. In the third stage, after using the reference account and the carrying accounts identified in the previous two stages to re-determine the target attribute frequent item set set about the carrying account and the target attribute frequent item set set about the reference account, using the carrying accounts identified in the previous two stages and the applied features, the predetermined neural network model is trained, so that the predetermined neural network model learns the artificial experience and summarizes the high-order features possessed by the carrying account, and the account identification model is obtained, thereby identifying more carrying accounts.
[0242] In one embodiment, the method further comprises: when the identification result indicates that the target account is a carrying account, generating a carrying mark for the publishing content corresponding to the target account; and when pushing the publishing content, filtering the publishing content to be pushed according to the carrying mark.
[0243] Optionally, when the account is identified as a carrying account according to the account identification method provided in the foregoing, a corresponding carrying mark can be generated for the carrying account and the publishing content generated based on the carrying account. When performing content recommendation, the marked publishing content is subjected to batch filtering and weight reduction of the content, so as to achieve the purpose of filtering the carrying content and make the information flow ecology healthier.
[0244] In one specific embodiment, the account identification method is roughly divided into two parts, one part is the training process of the account identification model, and the other part is the account identification process.
[0245] The training process of the account identification model includes three stages, and the data required for preparation includes reference accounts, published content corresponding to the reference accounts, reference vocabulary, and unidentified accounts.
[0246] In the first stage, according to the published content corresponding to the reference accounts and the unidentified accounts, the accounts with account similarity relationship with the reference accounts are screened out from the unidentified accounts, and then the first type of carrying probability is extracted according to the reference vocabulary, category and content similarity, and the first carrying account is further screened out.
[0247] In the second stage, according to the published content corresponding to the reference accounts and the first carrying account obtained in the previous stage, the frequent item set sample set about the theme and classification of the carrying account is mined, and the corresponding second type of carrying probability is obtained according to the number of the two sets hit in the published content of the unidentified account, and the second carrying account is identified from the unidentified account.
[0248] In the third stage, the reference accounts and the carrying accounts identified in the previous two stages are used as sample accounts. According to these sample accounts, the frequent item set set about the theme and classification of the carrying account and the frequent item set set about the theme and classification of the reference account are re-determined, and then the frequent item set bag is generated, and the bag-of-words vector of each sample account is counted. The bag-of-words vector combines the carrying accounts identified in the previous two stages and the first type of carrying feature and the second type of carrying feature applied, and is input into the predetermined neural network model together to perform model training, so that the predetermined neural network model learns artificial experience and summarizes the high-order features possessed by the carrying account, and obtains the trained account identification model.
[0249] The account identification process is consistent with the third stage of the account identification model training process. That is, the processing flow for each sample account during model training is consistent with the processing flow for the target account during account identification. Specifically, based on reference vocabulary, the category of the target account's posted content, and the content similarity between the target account's and reference account's posted content, the first type of plagiarism features for the target account are extracted. Based on all reference accounts, plagiarism accounts identified in the first and second stages, the first topic frequent itemset set and the first category frequent itemset set for plagiarism accounts, as well as the second topic frequent itemset set and the second category frequent itemset set for reference accounts, are extracted. The corresponding second type of plagiarism features are obtained based on the number of times the target account's posted content matches the first topic frequent itemset set and the first category frequent itemset set for plagiarism accounts. Based on the frequent itemsets in these four sets, a bag-of-words for frequent itemsets related to the posted content is generated. The number of times the target account's posted content matches each frequent itemset in this bag-of-words is counted, generating the corresponding bag-of-words vector as the third type of plagiarism feature. These three types of features are input into the trained account recognition model to obtain the recognition result of whether the target account is a copy account.
[0250] It should be understood that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0251] In one embodiment, such as Figure 18 As shown, an account recognition device 1800 is provided. This device can be a software module, a hardware module, or a combination of both as part of a computer device. Specifically, the device includes: a first feature statistics module 1802, a frequent itemset determination module 1804, a second feature statistics module 1806, and a recognition module 1808, wherein:
[0252] The first feature statistics module 1802 is used to count the first type of plagiarism features of the target account at the content similarity level based on the published content of the target account.
[0253] The frequent item set determination module 1804 is configured to determine a target attribute frequent item set set of the carrier account based on the publishing content corresponding to each sample account in a sample account set, the sample account set including the carrier account and the reference account.
[0254] The second feature statistics module 1806 is configured to obtain a second type of carrier feature of the target account at a frequent mode level according to a matching degree of the publishing content corresponding to the target account and the target attribute frequent item set set.
[0255] The identification module 1808 is configured to identify the target account based on the first type of carrier feature and the second type of carrier feature, and obtain an identification result of whether the target account is a carrier account.
[0256] In an embodiment, the first feature statistics module 1802 is further configured to: determine a reference vocabulary hit degree, a field distribution concentration degree, and a content carrier degree of the target account based on the publishing content corresponding to the target account; perform weighted summation on the reference vocabulary hit degree, the field distribution concentration degree, and the content carrier degree of the target account to obtain a first type of carrier probability of the target account; and take the reference vocabulary hit degree, the field distribution concentration degree, the content carrier degree, and the first type of carrier probability as the first type of carrier feature of the target account at a content similarity level.
[0257] In an embodiment, the first feature statistics module 1802 is further configured to: obtain a reference vocabulary, the reference vocabulary being generated according to the publishing content of the reference account; obtain attribute information of each publishing content of the target account, and obtain a number of times of hitting the reference vocabulary for each publishing content according to a number of times of hitting the reference vocabulary according to the attribute information; and calculate the reference vocabulary hit degree of the target account according to a total number of times of hitting the reference vocabulary for the publishing content corresponding to the target account and a number of publishing times of the target account.
[0258] In an embodiment, the first feature statistics module 1802 is further configured to: perform category detection on the publishing content corresponding to the target account to obtain a category of each publishing content; and take a ratio of a number of publishing contents belonging to a target category to a total number of publishing contents of the target account as the field distribution concentration degree of the target account, the number of categories belonging to the target category in the publishing content corresponding to the target account being the largest.
[0259] In an embodiment, the first feature statistics module 1802 is further configured to respectively perform feature extraction on attributes of the respective published content of the target account and the reference account, to obtain attribute feature vectors of each attribute of the respective published content; for the two published content, calculate vector similarity between the attribute feature vectors corresponding to the same attribute, and obtain content similarity between the two published content according to the vector similarity corresponding to multiple attributes; wherein the two published content are respectively from the target account and the reference account, and the published time of the published content from the reference account is earlier than that of the published content from the target account; count the number of published content in the target account with content similarity higher than a threshold; and calculate the content carrying degree corresponding to the target account according to the number and the total number of published content of the target account.
[0260] In an embodiment, the frequent item set determination module 1804 is further configured to determine a first topic frequent item set set about the carrying account and a second topic frequent item set set about the reference account based on the topics of the published content corresponding to each sample account in the sample account set; and determine a first category frequent item set set about the carrying account and a second category frequent item set set about the reference account based on the categories of the published content corresponding to each sample account in the sample account set.
[0261] In an embodiment, the frequent item set determination module 1804 is further configured to extract the first topic frequent item set about the carrying account and the first category frequent item set about the carrying account based on the published content corresponding to the carrying account in the sample account set; extract the second topic frequent item set about the reference account and the second category frequent item set about the reference account based on the published content corresponding to the reference account in the sample account set; form a set of each frequent item set belonging to the first topic frequent item set and not belonging to the second topic frequent item set as the first topic frequent item set set about the carrying account, and form a set of each frequent item set belonging to the second topic frequent item set and not belonging to the first topic frequent item set as the second topic frequent item set set about the reference account; form a set of each frequent item set belonging to the first category frequent item set and not belonging to the second category frequent item set as the first category frequent item set set about the carrying account category, and form a set of each frequent item set belonging to the second category frequent item set and not belonging to the first category frequent item set as the second category frequent item set set about the reference account category.
[0262] In an embodiment, the frequent item set determination module 1804 is further configured to, after extracting attribute information of the published content, perform word segmentation on the attribute information and count the importance of each word, determine a theme word of the published content according to the importance, and perform category detection on the published content to obtain a category of the published content; extract a first theme frequent item set about the carrier account according to the theme word of each piece of published content corresponding to the carrier account, and extract a first category frequent item set about the carrier account according to the category of each piece of published content corresponding to the carrier account.
[0263] In an embodiment, the second feature statistics module 1806 is further configured to obtain a carrier account theme hit degree corresponding to the target account according to a matching degree of the target account and the first theme frequent item set collection; obtain a carrier account category hit degree corresponding to the target account according to a matching degree of the target account and the first category frequent item set collection; and obtain a second carrier probability by weighted summation of the carrier account theme hit degree and the carrier account category hit degree. The carrier account theme hit degree, the carrier account category hit degree, and the second carrier probability are taken as the second carrier feature of the target account.
[0264] In an embodiment, the second feature statistics module 1806 is further configured to extract a theme word of the published content according to attribute information of the published content corresponding to the target account, and perform category detection on the published content to obtain a category of the published content; count a first number of published contents in which the theme word belongs to the first theme frequent item set collection, and calculate a carrier account theme hit degree corresponding to the target account according to the first number and a total number of published contents corresponding to the target account; count a second number of published contents in which the category belongs to the first category frequent item set collection, and calculate a carrier account category hit degree corresponding to the target account according to the second number and the total number of published contents corresponding to the target account.
[0265] In an embodiment, the apparatus further includes a third feature statistics module configured to obtain a third carrier feature of the target account according to a number of themes that hit each frequent item set in the first theme frequent item set collection and the second theme frequent item set collection, and a number of categories that hit each frequent item set in the first category frequent item set collection and the second category frequent item set collection in the published content corresponding to the target account. The identification module 1808 is further configured to fuse the first carrier feature, the second carrier feature, and the third carrier feature to obtain a fusion feature, and identify the target account based on the fusion feature to obtain an identification result of whether the target account is a carrier account.
[0266] In an embodiment, the third feature statistics module is further configured to generate a frequent item set bag of words about the published content according to the first topic frequent item set collection, the second topic frequent item set collection, the first category frequent item set collection, and the second category frequent item set collection; obtain the topic vocabulary and the category of each piece of published content corresponding to the target account; determine the hit times of each frequent item set in the frequent item set bag of words corresponding to the target account according to the topic vocabulary and the category of each piece of published content; and generate a bag of words vector corresponding to the target account as the third type of carrying feature of the target account according to the hit times of each frequent item set corresponding to the target account.
[0267] In an embodiment, the identification module 1808 is further configured to perform feature fusion on the first type of carrying feature and the second type of carrying feature through a feature fusion layer of the account identification model to obtain the fusion feature corresponding to the target account, and perform classification based on the fusion feature through a classification prediction layer of the account identification model to obtain the identification result of whether the target account is a carrying account.
[0268] In an embodiment, the account identification model is obtained by using the carrying accounts and the reference accounts in the sample account collection to perform model training on a predetermined neural network model, the carrying accounts include the first carrying accounts and the second carrying accounts, the first carrying accounts are accounts with a corresponding first type of carrying probability greater than a first threshold, the second carrying accounts are accounts with a corresponding second type of carrying probability greater than a second threshold, and the second type of carrying probability is calculated according to a target attribute frequent item set sample collection about the carrying accounts determined based on the first carrying accounts and the reference accounts.
[0269] In an embodiment, the device further includes a training module, and the training module includes a first carrying account identification unit configured to obtain unidentified accounts, and filter accounts having an account similarity relationship with any reference account from the unidentified accounts; based on the published content corresponding to the filtered accounts, statistically determine the reference vocabulary hit degree, the field distribution concentration degree, and the content carrying degree, weight and sum the reference vocabulary hit degree, the field distribution concentration degree, and the content carrying degree corresponding to the filtered accounts to obtain the first type of carrying probability of the filtered accounts; and take samples with the first type of carrying probability greater than a first threshold in the filtered accounts as the first carrying accounts.
[0270] In an embodiment, the first account identification unit is further configured to respectively extract features of attributes of respective published content of the unidentified account and the reference account, and obtain attribute feature vectors of each attribute of the respective published content; for the two published content, calculate vector similarity between the attribute feature vectors of the same attribute, and obtain content similarity between the two published content according to the vector similarity of the multiple attributes; wherein the two published content are respectively from the unidentified account and the reference account; aggregate the content similarity between the two published content to obtain the account similarity between the unidentified account and the reference account, and identify the unidentified account as the account having the account similarity relationship with the reference account when the account similarity is greater than a preset threshold.
[0271] In an embodiment, the account identification device 1800 further comprises a training module, and the training module comprises a second account identification unit configured to determine a topic frequent item set sample set of the account and a category frequent item set sample set of the account based on the topics and categories of the published content corresponding to the respective first account and the reference account; obtain a topic hit degree of the account corresponding to the unidentified account according to the matching degree of the unidentified account and the topic frequent item set sample set; obtain a category hit degree of the account corresponding to the unidentified account according to the matching degree of the unidentified account and the category frequent item set sample set; and obtain a second type of account probability corresponding to the unidentified account by weighted sum of the topic hit degree and the category hit degree of the account; and identify the sample with the second type of account probability greater than a second threshold value in the unidentified account as the second account.
[0272] In an embodiment, the account identification device 1800 further comprises a carrying mark module and a pushing module, configured to generate a carrying mark for the published content corresponding to the target account when the identification result indicates that the target account is a carrying account; and the pushing module is configured to filter the published content to be pushed according to the carrying mark when pushing the published content.
[0273] The account identification device 1800, for a target account, on the one hand, based on the corresponding publishing content, aggregates the similarity relationship between the contents to the account dimension, mines the first type of carrying characteristics of the target account in the content similarity layer, and on the other hand, by focusing on the publishing content of the carrying account and the reference account in the sample account set, a target attribute frequent item set set about the carrying account is constructed, and according to the matching degree of the target attribute of the corresponding publishing content of the target account and the constructed target attribute frequent item set set, the frequent pattern about the target attribute is mined as the second type of carrying characteristics of the target account in the frequent pattern layer. Then, by combining the first type of carrying characteristics and the second type of carrying characteristics, the target account can be automatically identified, and the identification result of whether the target account is a carrying account can be obtained. Not only can the problems of low efficiency and low accuracy of manual review be avoided, but also by using all the publishing content of the target account to mine the frequent pattern, a content set is identified, and when the target account has certain carrying behavior characteristics, it can be identified, which can avoid the problem of insufficient recall rate caused by identifying from the content dimension.
[0274] In one embodiment, as shown in Figure 19 A processing device 1900 of an account identification model is provided, which can be a part of a computer device in the form of a software module or a hardware module, or a combination of the two. The device specifically includes an account obtaining module 1902, a first carrying account mining module 1904, a frequent item set determination module 1906, a second carrying account mining module 1908, and a training module 1910, wherein:
[0275] The account obtaining module 1902 is configured to obtain an unidentified account and a reference account.
[0276] The first carrying account mining module 1904 is configured to filter, from the unidentified account, an account having an account similarity relationship with any reference account, and based on the publishing content corresponding to the filtered account, to count a first type of carrying probability of the filtered account in the content similarity layer, and to determine a first carrying account from the filtered account according to the first type of carrying probability.
[0277] The frequent item set determination module 1906 is configured to determine, based on the publishing content corresponding to the first carrying account and the reference account, a target attribute frequent item set sample set about the carrying account.
[0278] The second carrying account mining module 1908 is configured to obtain, according to the matching degree of the publishing content corresponding to the unidentified account and the target attribute frequent item set set, a second type of carrying probability of the unidentified account in the frequent pattern layer, and to determine a second carrying account from the unidentified account according to the second type of carrying probability.
[0279] The frequent itemset determination module 1906 is also used to determine the frequent itemset set of the target attribute of the transfer account and the frequent itemset set of the target attribute of the reference account based on the sample account set consisting of the first transfer account, the second transfer account and the reference account.
[0280] Training module 1910 is used to train a predetermined neural network model using the published content corresponding to each sample account in the sample account set, as well as the frequent itemset set of target attributes for the plagiarizing account and the frequent itemset set of target attributes for the reference account, to obtain an account recognition model for identifying plagiarizing accounts.
[0281] The processing device 1900 of the aforementioned account recognition model divides the training process of the account recognition model into three stages: In the first stage, after filtering out accounts with similar relationships to the reference account, the first-type plagiarism probability at the content similarity level is obtained through analysis and calculation of features such as content similarity, and the plagiarism accounts of the first stage are identified. In the second stage, based on the reference account and the plagiarism accounts obtained in the previous stage, a frequent itemset sample set of target attributes for the plagiarism accounts is mined, and the corresponding second-type plagiarism probability is obtained based on the matching degree between the content published by the unidentified account and these frequent itemsets, thus identifying the plagiarism accounts of the second stage. In the third stage, after redefining the frequent itemset set of target attributes for the plagiarism accounts and the frequent itemset set of target attributes for the reference account using the reference account and the plagiarism accounts identified in the first two stages, a predetermined neural network model is trained using the plagiarism accounts identified in the first two stages and the applied features. This allows the predetermined neural network model to learn from human experience and summarize the high-order features possessed by plagiarism accounts, resulting in an account recognition model that identifies more plagiarism accounts.
[0282] Specific limitations regarding the account recognition device 1800 and the account recognition model processing device 1900 can be found in the above descriptions of the limitations on the account recognition method and the account recognition model processing method, and will not be repeated here. Each module in the aforementioned account recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. Similarly, each module in the aforementioned account recognition model processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0283] In one embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows: Figure 20As shown in the figure. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is used to communicate with the external terminal through the network connection. The computer program is executed by the processor to implement an account identification method and / or a processing method of an account identification model.
[0284] When the computer device is a terminal, its internal structure can also include a display screen and an input device. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.
[0285] Those skilled in the art can understand that, Figure 20 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0286] In one embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in the above method embodiments.
[0287] In one embodiment, a computer readable storage medium is provided, storing a computer program, which is executed by a processor to implement the steps in the above method embodiments.
[0288] In one embodiment, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to make the computer device execute the steps in the above method embodiments.
[0289] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0290] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, but as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0291] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for those skilled in the art, without departing from the concept of the present application, some modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. An account identification method, characterized in that, The method includes: Based on the content published by the target account, the hit rate of reference words, the concentration of domain distribution, and the degree of content plagiarism of the target account are statistically analyzed. Based on the hit rate of reference words, the concentration of domain distribution, and the degree of content plagiarism of the target account, the first type of plagiarism feature of the target account at the content similarity level is obtained. Obtain a sample account set, which includes reposting accounts and reference accounts. Based on the topics and categories of the content published by each reposting account and each reference account, determine the topic frequent itemset set and category frequent itemset set for the reposting accounts, and the topic frequent itemset set and category frequent itemset set for the reference accounts. Each frequent itemset in the topic frequent itemset set originates from a topic word that appears more frequently than a specified threshold in the topic of the corresponding account's published content, and each frequent itemset in the category frequent itemset set originates from a category word that appears more frequently than a specified threshold in the category of the corresponding account's published content. Based on the degree of matching between the topic of the content published by the target account and the frequent itemset set of the topic, and the degree of matching between the category of the content published by the target account and the frequent itemset set of the category, the second type of copying feature of the target account at the frequent pattern level is obtained. Based on the first type of reposting features and the second type of reposting features, the target account is identified to obtain the identification result of whether the target account is a reposting account.
2. The method according to claim 1, characterized in that, The method of statistically analyzing the first type of plagiarism characteristics of the target account based on its published content includes: Based on the content published by the target account, the hit rate of reference keywords, the concentration of domain distribution, and the degree of content plagiarism for the target account are statistically analyzed. The first type of plagiarism probability of the target account is obtained by weighted summing of the hit rate of reference keywords, the concentration of domain distribution, and the degree of content plagiarism corresponding to the target account. The degree of hit of the reference words, the degree of concentration of domain distribution, the degree of content plagiarism, and the first type of plagiarism probability are used as the first type of plagiarism feature of the target account at the content similarity level.
3. The method according to claim 2, characterized in that, The statistical steps for determining the hit rate of the reference vocabulary corresponding to the target account include: Obtain reference vocabulary, which is generated based on the content published by the reference account; Obtain the attribute information of each piece of content published by the target account, and obtain the number of times each piece of content hits the reference words based on the number of times the attribute information hits the reference words; The hit rate of the reference keywords for the target account is calculated based on the total number of times the content published by the target account matches the reference keywords and the number of posts published by the target account.
4. The method according to claim 2, characterized in that, The statistical steps for determining the concentration of domain distribution corresponding to the target account include: The categories of the content published by the target account are determined by category detection. The ratio of the number of posts belonging to the target category to the total number of posts by the target account is used as the degree of concentration of the domain distribution corresponding to the target account, where the number of posts belonging to the target category is the largest among the posts by the target account.
5. The method according to claim 2, characterized in that, The statistical steps for determining the extent of content plagiarism by the target account include: The attributes of the content published by the target account and the reference account are obtained respectively, and the attribute feature vectors of each attribute are obtained for each published content. For two published content items, after calculating the vector similarity between the attribute feature vectors corresponding to the same attribute, the content similarity between the two published content items is obtained based on the vector similarity of multiple corresponding attributes; wherein, the two published content items originate from the target account and the reference account respectively, and the publication time of the published content originating from the reference account is earlier than that of the published content originating from the target account; Count the number of posts published by the target account whose content similarity exceeds a threshold; The degree of content plagiarism corresponding to the target account is calculated based on the stated quantity and the total number of posts published by the target account.
6. The method according to claim 1, characterized in that, The process of determining the frequent itemset set and frequent itemset set for themes and categories of the content posted by each reposting account in the sample account set, and the frequent itemset set for themes and frequent itemset set for the content posted by each reference account, based on the themes and categories of the content posted by each reposting account in the sample account set, includes: Based on the topic of the published content corresponding to each sample account in the sample account set, determine the first topic frequent itemset set for reposting accounts and the second topic frequent itemset set for reference accounts; Based on the category of the content published by each sample account in the sample account set, determine the first category frequent itemset set for reposting accounts and the second category frequent itemset set for reference accounts.
7. The method according to claim 6, characterized in that, The process involves determining a first set of frequent itemsets for reposting accounts and a second set of frequent itemsets for reference accounts based on the topics of the published content corresponding to each sample account in the sample account set; and determining a first set of frequent itemsets for reposting accounts and a second set of frequent itemsets for reference accounts based on the categories of the published content corresponding to each sample account in the sample account set, including: Based on the content posted by the reposting accounts in the sample account set, extract the first topic frequent itemset and the first category frequent itemset of the reposting accounts; Based on the published content corresponding to the reference account in the sample account set, extract the second topic frequent itemset and the second category frequent itemset of the reference account; The set of frequent itemsets that belong to the first topic frequent itemset but not to the second topic frequent itemset is taken as the first topic frequent itemset set for the reposting account, and the set of frequent itemsets that belong to the second topic frequent itemset but not to the first topic frequent itemset is taken as the second topic frequent itemset set for the reference account. The set of frequent itemsets belonging to the first category but not to the second category is taken as the first category frequent itemset set for the purpose of transferring accounts, and the set of frequent itemsets belonging to the second category but not to the first category is taken as the second category frequent itemset set for the purpose of the purpose of referencing accounts.
8. The method according to claim 7, characterized in that, Based on the content posted by the accounts that reposted content in the sample account set, the first topic frequent itemset and the first category frequent itemset for the reposting accounts are extracted, including: After extracting the attribute information of the published content, the attribute information is segmented into words and the importance of each word is counted. The topic words of the published content are determined based on the importance, and the category of the published content is obtained by category detection. Based on the keywords of each post from the reposting account, extract the first frequent itemset of the reposting account by topic, and based on the category of each post from the reposting account, extract the first frequent itemset of the reposting account by category.
9. The method according to claim 6, characterized in that, The step of obtaining the second type of copying feature of the target account at the frequent pattern level based on the matching degree between the topic of the content published by the target account and the frequent itemset set of the topic, and the matching degree between the category of the content published by the target account and the frequent itemset set of the category, includes: Based on the degree of matching between the target account and the first set of frequent itemsets of the topic, the topic hit rate of the reposting account corresponding to the target account is obtained; Based on the degree of matching between the target account and the first category of frequent itemsets, the hit rate of the reposting account category corresponding to the target account is obtained; The second type of reposting probability is obtained by weighting and summing the hit rate of the reposting account topic and the hit rate of the reposting account category; The degree of topic hit of the reposting account, the degree of category hit of the reposting account, and the second type of reposting probability are used as the second type of reposting feature of the target account.
10. The method according to claim 6, characterized in that, The method further includes: Based on the number of frequent itemsets in the first and second frequent itemsets sets of the topic and the number of frequent itemsets in the first and second frequent itemsets sets of the category in the content published by the target account, the third type of copying feature of the target account is obtained. The step of identifying the target account based on the first type of reposting features and the second type of reposting features, and obtaining an identification result as to whether the target account is a reposting account, includes: The first type of reposting feature, the second type of reposting feature, and the third type of reposting feature are fused to obtain a fused feature; the target account is identified based on the fused feature to obtain an identification result as to whether the target account is a reposting account.
11. The method according to claim 1, characterized in that, The step of identifying the target account based on the first type of reposting features and the second type of reposting features, and obtaining an identification result as to whether the target account is a reposting account, includes: The feature fusion layer of the account recognition model is used to fuse the first type of transfer features and the second type of transfer features to obtain the fused features corresponding to the target account. The classification prediction layer of the account recognition model is used to classify the target account based on the fused features, thereby obtaining the identification result of whether the target account is a copycat account.
12. The method according to claim 11, characterized in that, The account recognition model is obtained by training a predetermined neural network model using the copy accounts and reference accounts in the sample account set. The copy accounts include a first copy account and a second copy account. The first copy account is an account with a first type of copy probability greater than a first threshold, and the second copy account is an account with a second type of copy probability greater than a second threshold. The second type of copy probability is calculated based on the frequent itemset sample set of the target attribute of the copy account determined based on the first copy account and the reference account.
13. The method according to claim 12, characterized in that, The identification steps for the first transfer account include: Unidentified accounts are obtained, and accounts that have an account similarity relationship with any of the reference accounts are filtered from the unidentified accounts. Based on the published content of the filtered accounts, the corresponding reference keyword hit rate, domain distribution concentration, and content plagiarism degree are statistically analyzed. The reference keyword hit rate, domain distribution concentration, and content plagiarism degree of the filtered accounts are weighted and summed to obtain the first type of plagiarism probability of the filtered accounts. Samples of the filtered accounts whose first type of plagiarism probability is greater than a first threshold are taken as the first plagiarism accounts. Alternatively, the identification steps for the second transfer account include: Based on the topics and categories of the content published by each of the first reposting accounts and reference accounts, a frequent itemset sample set for topics and a frequent itemset sample set for categories of the reposting accounts are determined. The topic hit rate of the reposting account corresponding to the unidentified account is obtained based on the matching degree between the unidentified account and the frequent itemset sample set for topics. The category hit rate of the reposting account corresponding to the unidentified account is obtained based on the matching degree between the unidentified account and the frequent itemset sample set for categories. The topic hit rate and the category hit rate of the reposting account are weighted and summed to obtain the second type of reposting probability corresponding to the unidentified account. Samples among the unidentified accounts whose second type of reposting probability is greater than a second threshold are designated as second reposting accounts.
14. The method according to any one of claims 1 to 13, characterized in that, The method further includes: When the identification result indicates that the target account is a reselling account, then Generate a reposting tag for the content published by the target account; When pushing out content, the content to be pushed is filtered according to the aforementioned reposting tags.
15. A method for processing an account recognition model, characterized in that, The method includes: Retrieve unidentified accounts and reference accounts; From the unidentified accounts, select accounts that have an account similarity relationship with any of the reference accounts. Based on the published content of the selected accounts, calculate the first type of plagiarism probability of the selected accounts at the content similarity level. Determine the first plagiarism account from the selected accounts according to the first type of plagiarism probability. Based on the themes and categories of the content published by the first reposting account and the reference account, determine the frequent itemset set of themes and the frequent itemset set of categories for the reposting account; Based on the degree of matching between the topic of the published content corresponding to the unidentified account and the frequent itemset set of the topic, and the degree of matching between the category of the published content corresponding to the unidentified account and the frequent itemset set of the category, the second type of copying probability of the unidentified account at the frequent pattern level is obtained, and the second copying account is determined from the unidentified account according to the second type of copying probability. Based on the sample account set consisting of the first reposting account, the second reposting account, and the reference account, determine the topic frequent itemset set and category frequent itemset set for the reposting account, and the topic frequent itemset set and category frequent itemset set for the reference account; wherein, any frequent itemset in the topic frequent itemset set originates from topic words that appear more frequently than a specified threshold in the topic of the content published by the corresponding account, and any frequent itemset in the category frequent itemset set originates from category words that appear more frequently than a specified threshold in the category of the content published by the corresponding account; Using the topics and categories of the content published by each sample account in the sample account set, the frequent itemsets of topics and categories of the reposting accounts, and the frequent itemsets of topics and categories of the reference accounts, a predetermined neural network model is trained to obtain an account recognition model for identifying reposting accounts.
16. An account recognition device, characterized in that, The device includes: The first feature statistics module is used to calculate the hit rate of reference words, the concentration of domain distribution, and the degree of content plagiarism of the target account based on the published content of the target account, and to obtain the first type of plagiarism feature of the target account at the content similarity level based on the hit rate of reference words, the concentration of domain distribution, and the degree of content plagiarism of the target account. The frequent itemset determination module is used to obtain a sample account set, which includes reposting accounts and reference accounts. Based on the topic and category of the published content corresponding to each reposting account and the topic and category of the published content corresponding to each reference account in the sample account set, the module determines the topic frequent itemset set and category frequent itemset set for the reposting accounts, as well as the topic frequent itemset set and category frequent itemset set for the reference accounts. Specifically, any frequent itemset in the topic frequent itemset set originates from a topic word that appears more frequently than a specified threshold in the topic of the corresponding account's published content, and any frequent itemset in the category frequent itemset set originates from a category word that appears more frequently than a specified threshold in the category of the corresponding account's published content. The second feature statistics module is used to obtain the second type of copying feature of the target account at the frequent pattern level based on the degree of matching between the topic of the content published by the target account and the frequent itemset of the topic, and the degree of matching between the category of the content published by the target account and the frequent itemset of the category. The identification module is used to identify the target account based on the first type of reposting features and the second type of reposting features, and to obtain the identification result of whether the target account is a reposting account.
17. A processing device for an account recognition model, characterized in that, The device includes: The account acquisition module is used to acquire unidentified accounts and reference accounts; The first reposting account mining module is used to filter accounts that have an account similarity relationship with any of the reference accounts from the unidentified accounts, and based on the published content corresponding to the filtered accounts, to calculate the first type of reposting probability of the filtered accounts at the content similarity level, and to determine the first reposting account from the filtered accounts according to the first type of reposting probability. The frequent itemset determination module is used to determine the topic frequent itemset set and category frequent itemset set for the reposting account based on the topic and category of the published content corresponding to the first reposting account and the reference account. The second reposting account mining module is used to obtain the second type of reposting probability of the unidentified account at the frequent pattern level based on the matching degree between the topic of the published content corresponding to the unidentified account and the matching degree between the category of the published content corresponding to the unidentified account and the category frequent itemset set, and to determine the second reposting account from the unidentified account according to the second type of reposting probability. The frequent itemset determination module is further configured to determine, based on the sample account set consisting of the first reposting account, the second reposting account, and the reference account, a set of frequent itemsets for topics and categories for reposting accounts, and a set of frequent itemsets for topics and categories for reference accounts; wherein, any frequent itemset in the set of frequent itemsets for topics originates from topic words that appear more frequently than a specified threshold in the topics of the content published by the corresponding account, and any frequent itemset in the set of frequent itemsets for categories originates from category words that appear more frequently than a specified threshold in the categories of the content published by the corresponding account; The training module is used to train a predetermined neural network model using the topics and categories of the content published by each sample account in the sample account set, the frequent itemsets of topics and categories of the reposting accounts, and the frequent itemsets of topics and categories of the reference accounts, to obtain an account recognition model for identifying reposting accounts.
18. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 15.
19. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 15.
20. A computer program comprising computer instructions stored in a computer-readable storage medium, wherein a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the steps of the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Carrying account identification method, device and equipment and computer readable storage medium
CN112989167A