Detection Method, Device, Electronic Device and Storage Medium for Junk Accounts
By generating user matrix and text vectors, using deep learning models for classification, and finally determining whether the account is a spam account through the weighted average value, the problem of low detection accuracy of spam account in the existing technology is solved, and higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202210923309.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-08-02
AI Technical Summary
The prior art has low accuracy when detecting spam accounts, traditional classifiers are insufficient, and multi-layer reverse neural networks have low detection accuracy due to insufficient depth and activation function selection.
By generating an adjacency list of user information in the target account, a user matrix is generated, and node aggregation is performed to obtain the target user matrix. Then, the vector is updated using the preset update function and user classification is performed. At the same time, the text in the target account is transformed vectorically, and the text classification model is used for classification, and finally determine whether the account is a spam account through the weighted average.
The accuracy of spam account detection is improved, and the analysis of user attributes and text attributes by using deep learning models separately is reduced, the matrix dimension is reduced, the calculation amount is reduced, and the detection accuracy is improved.
Smart Images

Figure CN115238041B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a junk account detection method, device, electronic device and computer-readable storage medium. Background Art
[0002] With the rapid development of computer and Internet technology, people are increasingly accustomed to using the Internet to handle various work and life matters. To handle these matters, users generally register accounts in the business systems that provide corresponding services, such as e-commerce platform accounts, third-party payment platform accounts, forum platform accounts, etc., and then use the accounts as representatives of their identities to run related business logic. The Internet has two sides. While it greatly facilitates users, it also inevitably brings some security risks. Some users or organizations will register a large number of accounts for bad purposes and use these accounts to perform some abnormal operations, such as spreading messages, promoting false advertisements, and brushing orders, which not only wastes the resources of the business system, but also damages the interests of legitimate users.
[0003] Nowadays, traditional classifiers including Randforest, SVM, NB, and KNN are used to detect spam. However, due to their performance shortcomings and the need to manually extract the attributes of user information and spam information, the detection accuracy is not high. The use of multi-layer inverse neural networks to detect spam may result in a detection accuracy lower than that of some traditional classifiers due to insufficient network depth, incorrect activation function selection, parameter settings, etc. Therefore, how to improve the detection accuracy of spam accounts has become an urgent problem to be solved. Summary of the invention
[0004] The present invention provides a junk account detection method, device and computer-readable storage medium, the main purpose of which is to solve the problem of low accuracy in junk account detection.
[0005] To achieve the above object, the present invention provides a method for detecting junk accounts, comprising:
[0006] Generate an adjacency list of user information in a target account, and generate a user matrix of the user information using the adjacency list;
[0007] Performing node aggregation on the user matrix to obtain a target user matrix;
[0008] Updating the vector in the target user matrix according to a preset update function to obtain an updated vector;
[0009] Classifying the update vector using a preset user classification model to obtain a user classification result of the update vector;
[0010] Vectorize the text in the target account to obtain a text vector;
[0011] Input the text vector into a preset text classification model to obtain the text classification result of the text vector;
[0012] Calculate the weighted average of the user classification result and the text classification result, compare the weighted average with the threshold. When the weighted average is greater than the threshold, determine that the target account is a spam account.
[0013] Optionally, generating an adjacency list of user information in the target account includes:
[0014] Perform binary classification on the user information in the target account according to a preset index to obtain vertex information and edge information;
[0015] Initialize a preset vertex table, and write the vertex information into the initialized vertex table to obtain a target vertex table;
[0016] Write the edge information into a preset edge table in sequence according to the connection relationship to obtain a target edge table;
[0017] Generate an adjacency list based on the target vertex table and the target edge table.
[0018] Optionally, generating a user matrix of the user information using the adjacency list includes:
[0019] Generate a vertex array based on the vertex information in the adjacency list, and generate an edge array based on the edge information in the adjacency list;
[0020] Determine the total number of vertices according to the vertex array, and determine the total number of edges according to the edge array;
[0021] Construct an adjacency matrix using the total number of vertices and the total number of edges, and initialize the adjacency matrix;
[0022] Fill the initialized adjacency matrix according to the adjacency list to obtain the user matrix of the user information.
[0023] Optionally, performing node aggregation on the user matrix to obtain a target user matrix includes:
[0024] Partition the user matrix according to a preset node partitioning method to obtain local matrices;
[0025] Perform average pooling on all local matrices to obtain the eigenvalues of the local matrices;
[0026] Concatenate the eigenvalues to obtain a target user matrix.
[0027] Optionally, updating the vector in the target user matrix according to a preset update function to obtain an updated vector includes:
[0028] Updating the vector in the target user matrix by using the following update function:
[0029]
[0030] where, is the feature vector of the node u in the (k + 1)-th layer, represents the feature vector of the node v in the k-th layer, N(u) represents the set of neighbor nodes of the node u, UPDATE (k) represents updating at the nodes in the k-th layer, AGGREGATE (k) represents summing up the nodes in the k-th layer, means that the node v can take any value in N(u).
[0031] Optionally, converting the text in the target account into a vector to obtain a text vector includes:
[0032] Performing word segmentation on the text in the target account to obtain text word segments;
[0033] Obtaining the word vectors of the text word segments, and performing clustering processing on the word vectors to obtain the clustering categories of each feature word;
[0034] Calculating the weight of each feature word in the text based on a weight algorithm;
[0035] Generating a text vector according to the clustering categories and the weights.
[0036] Optionally, converting the text in the target account into a vector to obtain a text vector includes:
[0037] Performing word segmentation on the text in the target account to obtain text word segments;
[0038] Marking the text word segments to obtain marked word segments;
[0039] Performing marked padding on the marked word segments according to a preset sentence length to obtain standard word segments;
[0040] Performing attention masking on the marked word segments to obtain attention word segments;
[0041] Mapping the elements in the attention word segments to obtain the unique IDs of all the attention word segments;
[0042] Inputting the attention word segments and the unique IDs into a pre-trained language representation model to obtain the text vectors of each text word segment.
[0043] To solve the above problems, the present invention also provides a detection device for junk accounts, the device comprising:
[0044] A user matrix module, configured to generate an adjacency list of user information in a target account, and generate a user matrix of the user information by using the adjacency list;
[0045] A node aggregation module, configured to perform node aggregation on the user matrix to obtain a target user matrix;
[0046] A vector update module, configured to update vectors in the target user matrix according to a preset update function to obtain updated vectors;
[0047] A user classification module, configured to classify the updated vectors by using a preset user classification model to obtain a user classification result of the updated vectors;
[0048] A vectorization module, configured to perform vectorization conversion on text in the target account to obtain text vectors;
[0049] A text classification module, configured to input the text vectors into a preset text classification model to obtain a text classification result of the text vectors;
[0050] A weighted average module, configured to calculate a weighted average of the user classification result and the text classification result, compare the weighted average with a threshold, and when the weighted average is greater than the threshold, determine that the target account is a junk account.
[0051] To solve the above problems, the present invention also provides an electronic device, the electronic device comprising:
[0052] At least one processor; and,
[0053] A memory communicatively connected to the at least one processor; wherein,
[0054] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned detection method for junk accounts.
[0055] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned detection method for junk accounts.
[0056] In the embodiments of the present invention, different features of user attributes and text attributes are separately analyzed and processed using deep learning models. By generating an adjacency list of user information in the target account, a user matrix is generated, thereby realizing a simple and easy-to-understand representation of the user information. By partitioning and aggregating the user matrix, a target user matrix is obtained, reducing the matrix dimension and the subsequent computational amount. The vectors of the target user matrix are updated, and the updated vectors are classified to obtain a user classification result. A vectorization model is used to perform vectorization conversion on the text in the target account to obtain a text vector, and the text vector is classified to obtain a text classification result. The two classification results of the user classification result and the text classification result are used to determine whether the target account is a spam account, improving the detection accuracy. Therefore, the present invention proposes a detection method, device, electronic device, and computer-readable storage medium for spam accounts, which can solve the problem of low detection accuracy of spam accounts. Description of the Drawings
[0057] Figure 1 It is a flowchart of the detection method for spam accounts provided by an embodiment of the present invention;
[0058] Figure 2 It is a flowchart of generating a user matrix provided by an embodiment of the present invention;
[0059] Figure 3 It is a flowchart of text vectorization provided by an embodiment of the present invention;
[0060] Figure 4 It is a functional module diagram of the detection device for spam accounts provided by an embodiment of the present invention;
[0061] Figure 5 It is a structural diagram of an electronic device for implementing the detection method for spam accounts provided by an embodiment of the present invention.
[0062] The realization, functional characteristics, and advantages of the objectives of the present invention will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Embodiments
[0063] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0064] An embodiment of the present application provides a method for detecting spam accounts. The execution subject of the spam account detection method includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiment of the present application. In other words, the spam account detection method can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0065] Referring to Figure 1 As shown, it is a flowchart of the spam account detection method provided by an embodiment of the present invention. In this embodiment, the spam account detection method includes:
[0066] S1. Generate an adjacency list of user information in the target account, and generate a user matrix of the user information by using the adjacency list.
[0067] In the embodiment of the present invention, the adjacency list is a chained storage structure for storing a graph, which is composed of a vertex table and an edge table. The vertex table is composed of a vertex field and a pointer pointing to the first adjacent edge. The edge table node is composed of an adjacent point field and a pointer field pointing to the next adjacent edge. Using the adjacency list can reduce the storage space.
[0068] Specifically, the user matrix is a matrix representing the adjacent relationship between vertices. Data can be obtained intuitively and simply according to the user matrix, and it is easy to check whether there is an edge between any pair of vertices and all "adjacent points" of any vertex.
[0069] In the embodiment of the present invention, the generation of the adjacency list of user information in the target account includes:
[0070] Perform binary classification on the user information in the target account according to a preset index to obtain vertex information and edge information; initialize a preset vertex table, write the vertex information into the initialized vertex table to obtain a target vertex table; write the edge information into the preset edge table in sequence according to the connection relationship to obtain a target edge table; generate an adjacency list according to the target vertex table and the target edge table.
[0071] Specifically, for the binary classification of the user information in the target account according to a preset index, vertex information can be determined using the vertex numbers. After the vertex information screening of the user information in the target account is completed, the edge information can also be directly determined using the numbers and weights of the two vertices corresponding to each edge.
[0072] In an embodiment of the present invention, the generation of the user matrix of the user information using the adjacency list includes:
[0073] S21. Generate a vertex array according to the vertex information in the adjacency list, and generate an edge array according to the edge information in the adjacency list;
[0074] S22. Determine the total number of vertices according to the vertex array, and determine the total number of edges according to the edge array;
[0075] S23. Use the total number of vertices and the total number of edges to construct an adjacency matrix, and initialize the adjacency matrix;
[0076] S24. Fill the initialized adjacency matrix according to the adjacency list to obtain the user matrix of the user information.
[0077] Specifically, the generation of the vertex array according to the vertex information in the adjacency list is arranged according to the vertex identifier; the determination of the total number of vertices according to the vertex array can be as follows: from the vertex array (1, 2, 3, 4, 5, 6), it can be seen that the total number of vertices is 6.
[0078] Further, the initialization of the adjacency matrix can set the weight of the edge to infinity, or the weight of the edge can be set to 0 to obtain the initialized adjacency matrix.
[0079] S2. Perform node aggregation on the user matrix to obtain a target user matrix.
[0080] In an embodiment of the present invention, the dimensionality reduction processing of the user matrix is realized by performing node aggregation on the user matrix, reducing the subsequent calculation amount.
[0081] In an embodiment of the present invention, the generation of the target user matrix by performing node aggregation on the user matrix includes: partitioning the user matrix according to a preset node partitioning method to obtain local matrices; performing average pooling on all local matrices to obtain the eigenvalues of the local matrices; splicing the eigenvalues to obtain a target user matrix.
[0082] Specifically, the average pooling takes the average value of the feature points in each region of the local matrix, which can well retain the background.
[0083] S3. Update the vectors in the target user matrix according to a preset update function to obtain updated vectors.
[0084] In the embodiment of the present invention, the step of updating the vectors in the target user matrix according to a preset update function to obtain updated vectors includes:
[0085] Update the vectors in the target user matrix by using the following update function:
[0086]
[0087] Where, is the feature vector of the node u at the k + 1 layer, represents the feature vector of the node v at the k layer, N(u) represents the set of neighbor nodes of the node u, UPDATE (k) represents an update at the k layer node, AGGREGATE (k) represents a summation of the k layer nodes, means that the node v can take any value in N(u).
[0088] Specifically, updating the vectors in the user matrix by using the update function, the obtained updated vectors realize new representations from node to node, from edge to edge, from node to edge, and from edge to node.
[0089] Further, the updated vectors include updated node vectors, updated edge vectors, and updated global vectors. Multilayer perceptrons are respectively constructed for the updated node vectors, updated edge vectors, and updated global vectors, and the features of the feature vectors, node sets, and edge sets are further learned according to the multilayer perceptrons; the multilayer perceptron is also called an artificial neural network. Except for the input and output layers, it can have multiple hidden layers in the middle. The simplest multilayer perceptron only contains one hidden layer, that is, a three-layer structure. The layers of the multilayer perceptron are fully connected. The bottom layer of the multilayer perceptron is the input layer, the middle is the hidden layer, and the last is the output layer; the neurons in the hidden layer are fully connected to the input layer. Assuming that the input layer is represented by the vector X, the output of the hidden layer is f(W 1 *X + b 1 ), W 1 is the weight (also called the connection coefficient), b 1 is the bias, and the function f can be a common sigmoid function or tanh function. By using the activation function, non-linear factors can be introduced into the neurons, enabling the neural network to approximate any non-linear function arbitrarily, so that the neural network can be applied to more non-linear models.
[0090] S4. Classify the updated vector using a preset user classification model to obtain the user classification result of the updated vector.
[0091] In an embodiment of the present invention, the preset user classification model includes a preset SoftMax function, and the Sigmoid function can be used to complete the classification of the updated vector.
[0092] Specifically, for the multi-label classification problem, the Sigmoid function outputs multiple correct answers. When the preset user classification model solves the problem of having multiple correct answers, the Sigmoid function processes each original output value separately. The output of the Sigmoid function is between (0, 1), and the output range is limited, and the optimization is stable, so it can be used as the output layer.
[0093] S5. Perform vectorization conversion on the text in the target account to obtain a text vector.
[0094] In an embodiment of the present invention, the performing vectorization conversion on the text in the target account to obtain a text vector includes:
[0095] S31. Perform word segmentation on the text in the target account to obtain text word segments;
[0096] S32. Obtain the word vectors of the text word segments, perform clustering processing on the word vectors to obtain the clustering categories of each feature word;
[0097] S33. Calculate the weight of each feature word in the text based on a weight algorithm;
[0098] S34. Generate a text vector according to the clustering category and the weight.
[0099] Specifically, the word2vec model can be used to obtain the word vectors of the text word segments. The word2vec model is a simple neural network, which consists of the following several layers: 1 input layer, and the input uses one-hot encoding, that is: assuming there are n words, then each word can be represented by an n-dimensional vector, and only one position in this n-dimensional vector is 1, and the rest of the positions are 0. 1 hidden layer, and the number of neurons in the hidden layer represents the dimension size of each word represented by the vector. For the weight matrix between the input layer and the hidden layer, its shape should be a matrix of [vocab_size, hidden_size]. 1 output layer, and the output layer is a vector of size [vocab_size], and each value represents the probability of outputting a word; in the Word2Vec model, there are mainly two training models, Skip-Gram and CBOW. Intuitively understood, Skip-Gram is to predict the context given the current value, while CBOW is to predict the current value given the context.
[0100] Specifically, the weight algorithm can be the TF-IDF algorithm, which is a commonly used weighting technique for information retrieval and text mining. The importance of a word increases in direct proportion to the number of times it appears in a document, but at the same time decreases in inverse proportion to the frequency of its appearance in the corpus. If a certain word appears frequently (TF) in an article and rarely appears in other articles, it is considered that this word or phrase has good category discrimination ability and is suitable for classification.
[0101] In the embodiment of the present invention, the vectorization conversion of the text in the target account to obtain a text vector includes: performing word segmentation on the text in the target account to obtain text word segments; marking the text word segments to obtain marked word segments; performing marked filling on the marked word segments according to a preset sentence length to obtain standard word segments; performing attention masking on the marked word segments to obtain attention word segments; mapping the elements in the attention word segments to obtain the unique ID of all attention word segments; and inputting the attention word segments and the unique ID into a pre-trained language representation model to obtain the text vector of each text word segment.
[0102] Specifically, the language representation model can be the Bert model. Only by inserting the input and output of a specific task into the Bert model and then using the powerful attention mechanism of Transformer can many downstream tasks be simulated. The downstream tasks include: sentence pair relationship judgment, single text topic classification, question answering task, single sentence tagging. The Bert model has powerful language representation ability and feature extraction ability.
[0103] S6. Input the text vector into a preset text classification model to obtain the text classification result of the text vector.
[0104] In the embodiment of the present invention, a fully connected layer and an activation function are added to the text vector to complete the construction of the migration model, and the migration model is used to classify the text vector. Among them, the text vector is subjected to fully connected processing to obtain a fully connected vector, and the activation function is used to calculate the category prediction value of the fully connected vector as the target category, and the text classification result is obtained according to the category prediction value.
[0105] Specifically, it plays the role of a classifier in the entire convolutional neural network. If operations such as convolutional layers, pooling layers, and activation functions map the original data to the hidden layer feature space, the fully connected layer plays the role of mapping the learned "distributed feature representation" to the sample label space. The role of the fully connected layer is to achieve classification through feature extraction.
[0106] S7. Calculate the weighted average of the user classification result and the text classification result, compare the size of the weighted average and the threshold, and when the weighted average is greater than the threshold, determine that the target account is a spam account.
[0107] In the embodiment of the present invention, calculating the weighted average of the user classification result and the text classification result is obtained according to the user category prediction value for generating the user classification result and the text category prediction value for generating the text classification result.
[0108] Further, when the user category prediction value is equal to 0.4, the weight ratio of the user category prediction value is 0.5, the text category prediction value is equal to 0.6, and the weight ratio of the text category prediction value is 0.5, then the weighted average is 0.5.
[0109] Specifically, when the threshold is 0.4 and the weighted average is 0.5, it can be determined that the target account is a spam account.
[0110] In the embodiment of the present invention, for different characteristics of user attributes and text attributes, deep learning models are separately used for analysis and processing. By generating an adjacency list of user information in the target account, a user matrix is generated, realizing a simple and easy-to-understand representation of the user information. By partitioning and aggregating the user matrix, a target user matrix is obtained, reducing the matrix dimension and the subsequent calculation amount. The vector of the target user matrix is updated, and the updated vector is classified to obtain a user classification result. A vectorization model is used to perform vectorization conversion on the text in the target account to obtain a text vector, and the text vector is classified to obtain a text classification result. Using the two classification results of the user classification result and the text classification result to determine whether the target account is a spam account improves the detection accuracy. Therefore, the present invention proposes a detection method for spam accounts, which can solve the problem of low detection accuracy of spam accounts.
[0111] As Figure 4 shown, it is a functional module diagram of a detection device for spam accounts provided by an embodiment of the present invention.
[0112] The detection device 100 for spam accounts of the present invention can be installed in an electronic device. According to the realized functions, the detection device 100 for spam accounts can include a user matrix module 101, a node aggregation module 102, a vector update module 103, a user classification module 104, a vectorization module 105, a text classification module 106, and a weighted average module 107. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0113] In this embodiment, the functions of each module / unit are as follows:
[0114] The user matrix module is used to generate an adjacency list of user information in the target account and generate a user matrix of the user information by using the adjacency list;
[0115] The node aggregation module is used to perform node aggregation on the user matrix to obtain a target user matrix;
[0116] The vector update module is used to update the vectors in the target user matrix according to a preset update function to obtain updated vectors;
[0117] The user classification module is used to classify the updated vectors by using a preset user classification model to obtain a user classification result of the updated vectors;
[0118] The vectorization module is used to perform vectorization conversion on the text in the target account to obtain a text vector;
[0119] The text classification module is used to input the text vector into a preset text classification model to obtain a text classification result of the text vector;
[0120] The weighted average module is used to calculate a weighted average of the user classification result and the text classification result, compare the size of the weighted average with a threshold, and when the weighted average is greater than the threshold, determine that the target account is a spam account.
[0121] As Figure 5 shown, it is a schematic structural diagram of an electronic device for implementing a method for detecting spam accounts provided in an embodiment of the present invention.
[0122] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as a spam account detection program.
[0123] Among them, in some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. By running or executing programs or modules stored in the memory 11 (such as executing the detection program of junk accounts, etc.), and calling the data stored in the memory 11, it can perform various functions of the electronic device and process data.
[0124] The memory 11 includes at least one type of readable storage medium. The readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In some other embodiments, the memory 11 may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 11 may also include both the internal storage unit and the external storage device of the electronic device. The memory 11 can not only be used to store application software installed on the electronic device and various types of data, such as the code of the detection program of junk accounts, etc., but also be used to temporarily store data that has been output or will be output.
[0125] The communication bus 12 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.
[0126] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.
[0127] Only the electronic device with components is shown in the figure. Those skilled in the art can understand that the structure shown in the figure does not constitute a limitation on the electronic device, and it may include fewer or more components than shown in the figure, or combine certain components, or have different component arrangements.
[0128] For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0129] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0130] The detection program of the junk account stored in the memory 11 in the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:
[0131] Generate an adjacency list of user information in the target account, and generate a user matrix of the user information by using the adjacency list;
[0132] Perform node aggregation on the user matrix to obtain a target user matrix;
[0133] Update the vectors in the target user matrix according to a preset update function to obtain updated vectors;
[0134] Classify the updated vector using a preset user classification model to obtain the user classification result of the updated vector;
[0135] Perform vectorization conversion on the text in the target account to obtain a text vector;
[0136] Input the text vector into a preset text classification model to obtain the text classification result of the text vector;
[0137] Calculate the weighted average of the user classification result and the text classification result, compare the size of the weighted average and the threshold, and when the weighted average is greater than the threshold, determine that the target account is a spam account.
[0138] Specifically, for the specific implementation method of the above instructions by the processor 10, reference can be made to the description of the relevant steps in the corresponding embodiments of the accompanying drawings, which will not be elaborated here.
[0139] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM, Read-Only Memory).
[0140] The present invention also provides a computer-readable storage medium, where the readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, it can implement:
[0141] Generate an adjacency list of user information in the target account, and use the adjacency list to generate a user matrix of the user information;
[0142] Perform node aggregation on the user matrix to obtain a target user matrix;
[0143] Update the vectors in the target user matrix according to a preset update function to obtain an updated vector;
[0144] Classify the updated vector using a preset user classification model to obtain the user classification result of the updated vector;
[0145] Perform vectorization conversion on the text in the target account to obtain a text vector;
[0146] Input the text vector into a preset text classification model to obtain the text classification result of the text vector;
[0147] Calculate the weighted average of the user classification result and the text classification result, compare the size of the weighted average and the threshold, and when the weighted average is greater than the threshold, determine that the target account is a spam account.
[0148] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0149] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0150] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0151] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0152] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.
[0153] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.
[0154] Embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0155] In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as first and second are used to represent names and do not indicate any specific order.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for detecting spam accounts, characterized in that, the method includes: generating an adjacency list of user information in the target account, and generating a user matrix of the user information by using the adjacency list; performing node aggregation on the user matrix to obtain a target user matrix; updating the vectors in the target user matrix according to a preset update function to obtain updated vectors; classifying the updated vectors by using a preset user classification model to obtain the user classification result of the updated vectors; performing vectorization conversion on the text in the target account to obtain a text vector; inputting the text vector into a preset text classification model to obtain the text classification result of the text vector; calculating the weighted average of the user classification result and the text classification result, comparing the size of the weighted average and a threshold, and when the weighted average is greater than the threshold, determining that the target account is a spam account; wherein, generating the adjacency list of user information in the target account includes: performing binary classification on the user information in the target account according to a preset index to obtain vertex information and edge information; initializing a preset vertex table, writing the vertex information into the initialized vertex table to obtain a target vertex table; writing the edge information into a preset edge table in sequence according to the connection relationship to obtain a target edge table; generating an adjacency list according to the target vertex table and the target edge table; using the adjacency list to generate the user matrix of the user information includes: generating a vertex array according to the vertex information in the adjacency list, and generating an edge array according to the edge information in the adjacency list; determining the total number of vertices according to the vertex array, and determining the total number of edges according to the edge array; constructing an adjacency matrix by using the total number of vertices and the total number of edges, and initializing the adjacency matrix; filling the initialized adjacency matrix according to the adjacency list to obtain the user matrix of the user information.
2. The method for detecting spam accounts according to claim 1, characterized in that, performing node aggregation on the user matrix to obtain a target user matrix includes: partitioning the user matrix according to a preset node partitioning method to obtain local matrices; performing average pooling on all local matrices to obtain the eigenvalues of the local matrices; concatenating the eigenvalues to obtain a target user matrix.
3. The method for detecting spam accounts according to claim 1, characterized in that, updating the vectors in the target user matrix according to a preset update function to obtain updated vectors includes: updating the vectors in the target user matrix by using the following update function: Among them, is the feature vector of the layer node, represents the feature vector of the layer node, represents the set of neighbor nodes of node , represents an update at the layer node, represents a summation of the layer node, represents that node takes any value in .
4. The method for detecting spam accounts according to claim 1, characterized in that, performing vectorization conversion on the text in the target account to obtain a text vector includes: performing word segmentation on the text in the target account to obtain text word segments; obtaining the word vectors of the text word segments, performing clustering processing on the word vectors to obtain the clustering categories of each feature word; calculating the weight of each feature word in the text based on a weight algorithm; generating a text vector according to the clustering categories and the weights.
5. The detection method of a junk account according to any one of claims 1 to 4, characterized in that, the vectorization conversion of the text in the target account to obtain a text vector includes: performing word segmentation on the text in the target account to obtain text word segments; marking the text word segments to obtain marked word segments; performing marked padding on the marked word segments according to a preset sentence length to obtain standard word segments; performing attention masking on the marked word segments to obtain attention word segments; mapping the elements in the attention word segments to obtain the unique ID of all attention word segments; inputting the attention word segments and the unique ID into a pre-trained language representation model to obtain the text vector of each text word segment.
6. A detection device for a junk account, which is used to implement the detection method of a junk account according to any one of claims 1 to 5, characterized in that, the device includes: a user matrix module, which is used to generate an adjacency list of user information in the target account, and generate a user matrix of the user information by using the adjacency list; a node aggregation module, which is used to perform node aggregation on the user matrix to obtain a target user matrix; a vector update module, which is used to update the vectors in the target user matrix according to a preset update function to obtain updated vectors; a user classification module, which is used to classify the updated vectors by using a preset user classification model to obtain the user classification result of the updated vectors; a vectorization module, which is used to perform vectorization conversion on the text in the target account to obtain a text vector; a text classification module, which is used to input the text vector into a preset text classification model to obtain the text classification result of the text vector; a weighted average module, which is used to calculate the weighted average of the user classification result and the text classification result, compare the size of the weighted average and a threshold, and when the weighted average is greater than the threshold, determine that the target account is a junk account.
7. An electronic device, characterized in that, the electronic device includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the detection method of a junk account according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the detection method of a junk account according to any one of claims 1 to 5.
Citation Information
Patent Citations
Account risk model training method and risk user group determination method
CN114187112A
Method and electronic device for video content recommendation
US20170188102A1