Data processing method, device, medium and program product

By building a user interaction data graph structure on a social networking platform and conducting cluster analysis, we can identify and process abnormal data, solve the problem of identifying black market information, and improve the security of the platform.

CN118981719BActive Publication Date: 2025-09-30SHANGHAI BILIBILI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410979613.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2025-09-30
Estimated Expiration
2044-07-19

AI Technical Summary

Technical Problem

On Internet social networking platforms, illegal activities threaten user safety by publishing obscure and negative information, which is difficult to effectively identify and prevent with existing technologies.

Method used

By obtaining user interaction data, building a graph structure and performing cluster analysis, we can identify abnormal communities and detect black market-related content and accounts based on preset rules.

Benefits of technology

It achieves efficient identification of black market-related content and accounts, reduces potential dangers on the platform, and protects user safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118981719B_ABST
    Figure CN118981719B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, medium and program product for data processing. The method according to the present application includes: obtaining target text data to be processed, the target text data being obtained based on user interaction data; performing graph processing based on the similarity information between the various text data contained in the target text data set to obtain a corresponding graph structure; clustering the nodes in the graph structure to divide the nodes contained in the graph structure into multiple communities; analyzing the interaction data of each community based on predetermined anomaly analysis rules to determine whether there are abnormal numbers. The present application obtains a graph structure by performing similarity analysis on text data obtained based on the interaction data of a large number of users, and obtains multiple communities by clustering the graph structure, thereby mining data and corresponding accounts that may have abnormalities in each community based on preset conditions, thereby realizing the identification of black market-related content and accounts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method, device, computer-readable medium, and computer program product. Background Art

[0002] With the development of internet technology, the dangers posed by cybercrime on social networking platforms, such as video sites, can endanger ordinary users due to their massive user base. For example, within the communities of social networking platforms, cybercrime can spread harmful information, harming the community environment. The content posted by these actors is often highly cryptic, attempting to circumvent sensitive keywords and other policies set by the platforms. Summary of the Invention

[0003] Multiple aspects of the present application provide a data processing method, an apparatus, a computer-readable medium, and a computer program product.

[0004] In one aspect of the present application, a data processing method is provided, wherein the method comprises:

[0005] Acquiring target text data to be processed, wherein the target text data is obtained based on user interaction data;

[0006] Based on the similarity information between the various text data contained in the target text dataset, the graph is processed to obtain the corresponding graph structure;

[0007] Clustering the nodes in the graph structure to divide the nodes into multiple communities;

[0008] Analyze the interaction data of each community based on the predetermined abnormal analysis rules to determine whether there is abnormal data.

[0009] In one aspect of the present application, a device for data processing is provided, wherein the device includes:

[0010] means for acquiring target text data to be processed, wherein the target text data is obtained based on user interaction data;

[0011] A device for performing graph composition processing based on similarity information between various text data contained in a target text dataset to obtain a corresponding graph structure;

[0012] means for dividing the nodes in the graph structure into a plurality of communities by clustering the nodes in the graph structure;

[0013] A device for analyzing the interaction data of each community based on predetermined abnormality analysis rules to determine whether there is abnormal data.

[0014] Another aspect of the present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of an embodiment of the present application.

[0015] In another aspect of the present application, a computer-readable storage medium is provided, on which computer program instructions are stored. The computer program instructions can be executed by a processor to implement the method of the embodiment of the present application.

[0016] In another aspect of the present application, a computer program product is provided, including a computer program, which implements the method of the embodiment of the present application when executed by a processor.

[0017] In the solution provided in the embodiment of the present application, a graph structure is obtained by performing similarity analysis on text data obtained based on interactive data of a large number of users, and multiple communities are obtained by clustering the graph structure, thereby mining possible abnormal data and corresponding accounts in each community based on preset conditions, thereby realizing the identification of black market-related content and accounts. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, a brief introduction will be given below to the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Other features, objects and advantages of the present application will become more apparent upon reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0020] Figure 1 A schematic diagram showing a flow chart of a data processing method provided in an embodiment of the present application is shown;

[0021] Figure 2 A schematic structural diagram of a device for data processing provided in an embodiment of the present application is shown;

[0022] Figure 3 A structural diagram of a device suitable for implementing the solution in the embodiments of the present application is shown.

[0023] The same or similar reference numerals in the drawings represent the same or similar components. DETAILED DESCRIPTION

[0024] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] In a typical configuration of the present application, the terminal and the equipment of the service network each include one or more processors (CPUs), input / output interfaces, network interfaces and memories.

[0026] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0027] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer program instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc-read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0028] Figure 1 The flowchart of a method provided in an embodiment of the present application is shown. The method comprises at least step S101, step S102, step S103 and step S104.

[0029] In practical scenarios, the execution subject of this method can be a network device or an application running on a network device. The network device includes, but is not limited to, a network host, a single network server, a set of multiple network servers, or a collection of computers based on cloud computing, and can be used to implement some of the processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a group of loosely coupled computers forming a virtual computer.

[0030] The concepts involved in the embodiments of this application are explained below.

[0031] Black industry refers to illegal activities that use the Internet as a medium and network technology as the main means to pose potential threats to the security of computer information systems and the management order of cyberspace, and even national security and social and political stability.

[0032] In the context of social networks such as video websites, the method can be executed by a server used to identify illegal activities. After a user, acting as a content publisher, uploads a video clip, other users can interact with it by liking, commenting on, or forwarding it. The server exploits the fact that users associated with illegal activities tend to post similar content. By collecting large amounts of user interaction data and performing text similarity analysis on this interaction data, the server can identify potentially anomalous data and corresponding accounts, thereby identifying illegal content.

[0033] Reference Figure 1 In step S101, target text data to be processed is obtained, and the target text data is obtained based on user interaction data.

[0034] The user interaction data includes various information related to user interaction behaviors, including but not limited to users liking a video, posting comments, sending comments, or @ing other users.

[0035] The user interaction data includes but is not limited to at least one of the following:

[0036] 1) Text information corresponding to interactive behaviors; for example, dynamic text content posted by users, titles, text or introductions of video manuscripts posted by users, text content of barrages or comments posted by users, etc.;

[0037] 1) Account-related information of the user who made the interactive behavior; for example, the user's account ID, the identification information of the client device corresponding to the user, etc.;

[0038] 2) Identification information of the object of the user's interactive behavior; for example, the video ID, comment ID, and activity ID of the video being liked, commented on, or forwarded;

[0039] 3) Identification information of the user targeted by the user interaction and the corresponding client device of the user; for example, the user ID of the uploader who liked, commented on, or forwarded the post and the client device ID of the uploader;

[0040] 4) Time information of user interaction behavior.

[0041] Optionally, the method directly obtains text information corresponding to the user interaction data as the target text data.

[0042] According to one embodiment, the method integrates the text information corresponding to the user interaction data, and uses the text data obtained by the integration as the target text data.

[0043] For example, all text information corresponding to the video manuscript information is integrated according to title, partition, label and description.

[0044] Optionally, the method converts text information into word vectors of fixed dimensions through a pre-trained model, thereby using the corresponding word vector data as target text data.

[0045] Optionally, after obtaining a large amount of user interaction data, the method trains the pre-trained model based on the user interaction data, and the training task is to determine whether two interactive contents are in the same user account.

[0046] Specifically, for the task of determining whether two interactive contents are in the same user account, data is collected according to the account dimension during the data collection process of the pre-training task. Two texts posted by the same account in a short period of time are marked as 1, and the contents posted by different accounts are marked as 0. The data set is divided and then the classification task is performed. The task itself is to distinguish whether the two input texts are in the same account, thereby making the model parameters sensitive to the language of different accounts.

[0047] For example, in order to make the pre-trained model more adaptable to the characteristics of user interactive content on video websites, further pre-training is performed on the basis of Sim-BERT parameters using tens of millions of interactive data. The training task is to determine whether two interactive contents come from the same account and perform the corresponding classification task.

[0048] According to one embodiment, before step S101 , the method includes steps S105 and S106 .

[0049] In step S105 , user interaction data within a predetermined time period in the past is obtained.

[0050] In step S106, the interactive data is filtered based on a predetermined data filtering rule, thereby obtaining target text data based on the filtered interactive data.

[0051] For example, by analyzing multi-dimensional features such as user account level, behavior tags, interaction history, and user profiles, a pre-screening mechanism can be established to initially identify potential black market accounts. For example, in the comment scenario, user accounts are pre-screened using the following screening rules: the account has fewer than 1,000 followers, the commenter is not an Up Master, the comment can be seen by other users, the Chinese text length is greater than 12, and the comment rate within an hour exceeds 10. Accounts that meet all the screening rules are identified as suspected black market accounts, and all their comments and other information visible to other users in the past two hours are obtained as target text data for further processing, thereby reducing unnecessary subsequent computing costs.

[0052] In step S102 , graph composition processing is performed based on the similarity information between the various text data included in the target text dataset to obtain a corresponding graph structure.

[0053] The graph structure includes multiple nodes and edges, the nodes correspond to text content, and the edges correspond to similarity distances between text contents.

[0054] The similarity information includes various information that can indicate the similarity between two texts.

[0055] Specifically, step S102 further includes step S1021 and step S1022.

[0056] In step S1021 , the text similarities between the various text data included in the target text dataset are calculated based on a predetermined similarity algorithm.

[0057] The similarity algorithms include but are not limited to cosine similarity, Pearson correlation coefficient, vector inner product, etc.

[0058] Optionally, the text similarity information is the distance between two texts, and the distance includes but is not limited to cosine distance, Chebyshev distance, etc.

[0059] It should be noted that those skilled in the art should be familiar with the fact that a variety of similarity algorithms can be used to calculate the text similarity between text data, and are not limited to the algorithms mentioned above. Those skilled in the art can select a suitable algorithm to calculate the text similarity between text data based on actual needs.

[0060] In step S1022 , a graph structure of the target text dataset is constructed based on the calculated text similarity.

[0061] According to one embodiment, after obtaining the graph structure, the method executes step S107.

[0062] In step S107 , adjustment processing is performed based on the obtained graph structure.

[0063] The adjustment process includes but is not limited to at least one of the following processes:

[0064] 1) Adjust the weights of the edges of the graph structure;

[0065] For example, in order to highlight the similarity differences between different texts and reduce the impact of uneven data distribution density, the weight of the edge will be scaled. The calculation formula of the weight is: weight = -log(1-cosine similarity, e);

[0066] 2) Pruning the graph structure;

[0067] Because all nodes in a graph have edges between them, subsequent algorithms require significant computational resources and require more hardware resources to support the computation. To reduce computational resources and time, the graph structure is pruned. Specifically, if the weights between nodes fall below a preset threshold, the corresponding weights are removed. If, after removing the weights, the number of edges between a node and other nodes falls below a preset threshold, the node is directly deleted.

[0068] In step S103 , clustering is performed on the nodes in the graph structure to divide the nodes in the graph structure into multiple communities.

[0069] The method divides communities based on text similarity, so that the similarity of text contents belonging to a community is high.

[0070] According to one embodiment, step S103 includes step S1031.

[0071] In step S1031 , based on a predetermined community identification algorithm, the nodes in the graph structure are divided into multiple communities.

[0072] The community identification algorithm includes various algorithms that can be used to divide nodes in a graph structure into multiple communities based on the similarity between nodes, such as a label propagation algorithm, a GN algorithm, a random walk algorithm, etc.

[0073] According to one embodiment, the Louvain community discovery algorithm is used as the community identification algorithm.

[0074] The core goal of the Louvain community discovery algorithm is to identify community structures by maximizing the network modularity value.

[0075] Specifically, the Louvain community discovery algorithm clusters nodes in a graph structure. If the textual similarity between the textual content of certain nodes meets a predetermined requirement, these nodes are assigned to the same community. After algorithm iteration, the Louvain algorithm can divide the nodes in the graph structure into multiple communities, where the textual content belonging to the same community has a high degree of similarity.

[0076] According to one embodiment, in the scenario of video manuscripts, the method uses a density-based clustering algorithm to cluster multiple video manuscripts.

[0077] The density-based clustering algorithm includes, but is not limited to, Density-Based Spatial Clustering of Applications with Noise (DBSCAN). DBSCAN defines a cluster as the largest set of density-connected points, can partition areas with sufficiently high density into clusters, and can discover clusters of any shape in a spatial database of noise.

[0078] For example, since video manuscript information related to the black industry generally has a high similarity, the video manuscripts are clustered and the video manuscripts that may be related to the black industry are determined based on the density of the obtained clusters. The greater the density, the greater the probability of being related to the black industry.

[0079] Specifically, multiple video clips are encoded according to their titles, partitions, tags, and descriptions to generate corresponding feature vectors. These feature vectors are then clustered using the DBSCAN algorithm, and the density of each cluster is calculated. Clusters with a density greater than a preset threshold are then identified as potentially related to illegal activities.

[0080] In step S104, the interaction data of each community is analyzed based on predetermined abnormality analysis rules to determine whether there is abnormal data.

[0081] Optionally, the anomaly analysis rule is set based on at least any one of the following information:

[0082] 1) Distribution of published content; the distribution includes various information indicating the coverage of the corresponding published content, for example, the number of partitions of the video website where texts from the same community are distributed.

[0083] 2) The account level of the user who engaged in the interactive behavior; for example, the account level of the user who posted the video manuscript, barrage, or comment;

[0084] 3) Frequency information of interactive behaviors; for example, frequency information of posting the same or similar video manuscripts from the same community.

[0085] For example, since accounts related to the black industry usually roam around in the partitions of various communities to post video manuscripts and have relatively low account levels, for each community, texts or accounts that meet at least one of the following conditions will be regarded as texts or accounts related to the black industry: the number of partitions of the video website where texts from the same community are distributed is greater than or equal to 3; the minimum account level is less than or equal to 2 and the number of partitions of the video website where texts from the account are distributed is greater than or equal to 2; the minimum account level is less than or equal to 2 and the account posts content under more than 10 different video manuscripts.

[0086] According to one embodiment, the method further includes step S108.

[0087] In step S108, if it is determined that abnormal data exists in the community, the corresponding abnormal interaction data and abnormal accounts are determined, and the determined abnormal interaction data and abnormal accounts are processed.

[0088] For example, deleting abnormal text content, banning abnormal accounts, etc.

[0089] Optionally, the method uses keyword or account matching to remove texts and accounts that are easily accidentally damaged from the mining results, providing a backup measure for accidentally damaged texts and accounts.

[0090] According to the method of the embodiment of the present application, a graph structure is obtained by performing similarity analysis on text data obtained based on interactive data of a large number of users, and multiple communities are obtained by clustering the graph structure, thereby mining possible abnormal data and corresponding accounts in each community based on preset conditions, thereby realizing the identification of black market-related content and accounts.

[0091] Figure 2 A schematic structural diagram of a device for data processing provided in an embodiment of the present application is shown.

[0092] The device shown in the figure includes: a device for acquiring target text data to be processed (hereinafter referred to as "data acquisition device 101"), a device for performing composition processing based on the similarity information between the various text data contained in the target text data set to obtain a corresponding graph structure (hereinafter referred to as "composition processing device 102"), a device for clustering the nodes in the graph structure to divide the nodes contained in the graph structure into multiple communities (hereinafter referred to as "community division device 103") and a device for analyzing the interaction data of each community based on predetermined anomaly analysis rules to determine whether there is abnormal data (hereinafter referred to as "abnormal analysis device 104").

[0093] Reference Figure 2The data acquisition device 101 acquires target text data to be processed, and the target text data is obtained based on user interaction data.

[0094] The user interaction data includes various information related to user interaction behaviors, including but not limited to users liking a video, posting comments, sending comments, or @ing other users.

[0095] The user interaction data includes but is not limited to at least one of the following:

[0096] 1) Text information corresponding to interactive behaviors; for example, dynamic text content posted by users, titles, text or introductions of video manuscripts posted by users, text content of barrages or comments posted by users, etc.;

[0097] 1) Account-related information of the user who made the interactive behavior; for example, the user's account ID, the identification information of the client device corresponding to the user, etc.;

[0098] 2) Identification information of the object of the user's interactive behavior; for example, the video ID, comment ID, and activity ID of the video being liked, commented on, or forwarded;

[0099] 3) Identification information of the user targeted by the user interaction and the corresponding client device of the user; for example, the user ID of the uploader who liked, commented on, or forwarded the post and the client device ID of the uploader;

[0100] 4) Time information of user interaction behavior.

[0101] Optionally, the data acquisition device 101 directly acquires text information corresponding to the user interaction data as target text data.

[0102] According to one embodiment, the data acquisition device 101 integrates the text information corresponding to the user interaction data, and uses the text data obtained by the integration as the target text data.

[0103] For example, all text information corresponding to the video manuscript information is integrated according to title, partition, label and description.

[0104] Optionally, the data acquisition device 101 converts text information into word vectors of fixed dimensions through a pre-trained model, thereby using the corresponding word vector data as target text data.

[0105] Optionally, after acquiring a large amount of user interaction data, the device trains the pre-training model based on the user interaction data, and the training task is to determine whether two interactive contents are in the same user account.

[0106] Specifically, for the task of determining whether two interactive contents are in the same user account, data is collected according to the account dimension during the data collection process of the pre-training task. Two texts posted by the same account in a short period of time are marked as 1, and the contents posted by different accounts are marked as 0. The data set is divided and then the classification task is performed. The task itself is to distinguish whether the two input texts are in the same account, thereby making the model parameters sensitive to the language of different accounts.

[0107] For example, in order to make the pre-trained model more adaptable to the characteristics of user interactive content on video websites, further pre-training is performed on the basis of Sim-BERT parameters using tens of millions of interactive data. The training task is to determine whether two interactive contents come from the same account and perform the corresponding classification task.

[0108] According to one embodiment, the apparatus includes an interactive data acquisition device and a data screening device, and the operations of the interactive data acquisition device and the data screening device are performed before the operation of the data acquisition device 101 .

[0109] The interactive data acquisition device acquires user interactive data within a predetermined time period in the past.

[0110] The data screening device performs screening processing on the interactive data based on a predetermined data screening rule, thereby obtaining target text data based on the screened interactive data.

[0111] For example, the data screening device establishes a pre-screening mechanism to preliminarily identify potential black market accounts by analyzing multi-dimensional features such as the user's account level, behavior tags, interaction history, and user portrait. For example, in the comment scenario, user accounts are pre-screened using the following screening rules: the account has fewer than 1,000 followers, the commenter is not an Up master, the comment has not been viewed by the user, the Chinese text length is greater than 12, and the comment rate within one hour exceeds 10. Accounts that meet all the screening rules are identified as suspected black market accounts, and all their comment texts and other information that can be seen by other users in the past two hours are obtained as target text data for the next step of processing, thereby reducing unnecessary subsequent computing costs.

[0112] The composition processing device 102 performs composition processing based on the similarity information between the various text data included in the target text data set to obtain a corresponding graph structure.

[0113] The graph structure includes multiple nodes and edges, the nodes correspond to text content, and the edges correspond to similarity distances between text contents.

[0114] The similarity information includes various information that can indicate the similarity between two texts.

[0115] Specifically, the composition processing device 102 further includes a similarity calculation device and a sub-composition processing device.

[0116] The similarity calculation device calculates the text similarity between each text data included in the target text data set based on a predetermined similarity algorithm.

[0117] The similarity algorithms include but are not limited to cosine similarity, Pearson correlation coefficient, vector inner product, etc.

[0118] Optionally, the text similarity information is the distance between two texts, and the distance includes but is not limited to cosine distance, Chebyshev distance, etc.

[0119] It should be noted that those skilled in the art should be familiar with the fact that a variety of similarity algorithms can be used to calculate the text similarity between text data, and are not limited to the algorithms mentioned above. Those skilled in the art can select a suitable algorithm to calculate the text similarity between text data based on actual needs.

[0120] The sub-graph processing device constructs a graph structure of the target text data set based on the calculated text similarity.

[0121] According to one embodiment, the apparatus comprises a graph structure adjustment device, and the operation of the graph structure adjustment device is performed after the graph structure is obtained.

[0122] The graph structure adjustment device performs adjustment processing based on the obtained graph structure.

[0123] The adjustment process includes but is not limited to at least one of the following processes:

[0124] 1) Adjust the weights of the edges of the graph structure;

[0125] For example, in order to highlight the similarity differences between different texts and reduce the impact of uneven data distribution density, the weight of the edge will be scaled. The calculation formula of the weight is: weight = -log(1-cosine similarity, e);

[0126] 2) Pruning the graph structure;

[0127] Because all nodes in a graph have edges between them, subsequent algorithms require significant computational resources and require more hardware resources to support the computation. To reduce computational resources and time, the graph structure is pruned. Specifically, if the weights between nodes fall below a preset threshold, the corresponding weights are removed. If, after removing the weights, the number of edges between a node and other nodes falls below a preset threshold, the node is directly deleted.

[0128] Continue to Figure 2 To illustrate, the community division device 103 performs clustering processing on the nodes in the graph structure to divide the nodes included in the graph structure into multiple communities.

[0129] In the method, the community division device 103 divides communities based on text similarity, so that the similarity of text contents belonging to a community is high.

[0130] According to one embodiment, the community division device 103 divides the nodes in the graph structure into multiple communities based on a predetermined community identification algorithm.

[0131] The community identification algorithm includes various algorithms that can be used to divide nodes in a graph structure into multiple communities based on the similarity between nodes, such as a label propagation algorithm, a GN algorithm, a random walk algorithm, etc.

[0132] According to one embodiment, the community division device 103 uses the Louvain community discovery algorithm as the community identification algorithm.

[0133] The core goal of the Louvain community discovery algorithm is to identify community structures by maximizing the network modularity value.

[0134] Specifically, the Louvain community discovery algorithm clusters nodes in a graph structure. If the textual similarity between the textual content of certain nodes meets a predetermined requirement, these nodes are assigned to the same community. After algorithm iteration, the Louvain algorithm can divide the nodes in the graph structure into multiple communities, where the textual content belonging to the same community has a high degree of similarity.

[0135] According to one embodiment, in the scenario of video manuscripts, the device uses a density-based clustering algorithm to cluster multiple video manuscripts.

[0136] The density-based clustering algorithm includes, but is not limited to, Density-Based Spatial Clustering of Applications with Noise (DBSCAN). DBSCAN defines a cluster as the largest set of density-connected points, can partition areas with sufficiently high density into clusters, and can discover clusters of any shape in a spatial database of noise.

[0137] For example, since video manuscripts related to the black market generally have a high similarity, the video manuscripts are clustered and the density of the obtained clusters is used to determine the video manuscripts that may be related to the black market. The greater the density, the greater the probability of being related to the black market. Specifically, by encoding multiple video manuscripts according to the title, partition, tag, and description, the corresponding feature vectors are obtained. Then, these feature vectors are clustered based on the DBSCAN algorithm, and the density of each cluster is calculated, so that the clusters with a density greater than the preset threshold are regarded as clusters that may be related to the black market.

[0138] The abnormality analysis device 104 analyzes the interaction data of each community based on predetermined abnormality analysis rules to determine whether there is abnormal data.

[0139] Optionally, the anomaly analysis rule is set based on at least any one of the following information:

[0140] 1) Distribution of published content; the distribution includes various information indicating the coverage of the corresponding published content, for example, the number of partitions of the video website where texts from the same community are distributed.

[0141] 2) The account level of the user who engaged in the interactive behavior; for example, the account level of the user who posted the video manuscript, barrage, or comment;

[0142] 3) Frequency information of interactive behaviors; for example, frequency information of posting the same or similar video manuscripts from the same community.

[0143] For example, since accounts related to the black industry usually roam around in the partitions of various communities to post video manuscripts and have relatively low account levels, for each community, texts or accounts that meet at least one of the following conditions will be regarded as texts or accounts related to the black industry: the number of partitions of the video website where texts from the same community are distributed is greater than or equal to 3; the minimum account level is less than or equal to 2 and the number of partitions of the video website where texts from the account are distributed is greater than or equal to 2; the minimum account level is less than or equal to 2 and the account posts content under more than 10 different video manuscripts.

[0144] According to one embodiment, the device further comprises an exception handling device.

[0145] If it is determined that abnormal data exists in the community, the abnormality processing device determines the corresponding abnormal interaction data and abnormal accounts, and then processes the determined abnormal interaction data and abnormal accounts.

[0146] For example, deleting abnormal text content, banning abnormal accounts, etc.

[0147] According to one embodiment, the device uses keyword or account matching to remove texts and accounts that are easily accidentally damaged from the mining results, providing a backup measure for accidentally damaged texts and accounts.

[0148] According to the device of the embodiment of the present application, a graph structure is obtained by performing similarity analysis on text data obtained based on interactive data of a large number of users, and multiple communities are obtained by clustering the graph structure, thereby mining possible abnormal data and corresponding accounts in each community based on preset conditions, thereby realizing the identification of black market-related content and accounts.

[0149] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application. The method corresponding to the electronic device may be the method for data processing in the aforementioned embodiment, and its principle of solving the problem is similar to that of the method. The electronic device provided in an embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods and / or technical solutions of the aforementioned multiple embodiments of the present application.

[0150] The electronic device may be a user device, or a device formed by integrating a user device and a network device via a network, or an application running on the above device. The user device includes but is not limited to various terminal devices such as computers, mobile phones, tablets, smart watches, and bracelets. The network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer collections, and can be used to implement some of the processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing (Cloud Computing), where cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computers.

[0151] Figure 3The structure of a device suitable for implementing the method and / or technical solution in the embodiment of the present application is shown. The device 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage part 1208 into the random access memory (RAM) 1203. Various programs and data required for system operation are also stored in RAM 1203. CPU 1201, ROM 1202 and RAM 1203 are connected to each other through a bus 1204. Input / output (I / O) interface 1205 is also connected to bus 1204.

[0152] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, and the like; an output section 1207 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, and a speaker; a storage section 1208 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, and a semiconductor memory; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication section 1209 performs communication processing via a network such as the Internet.

[0153] In particular, the methods and / or embodiments of the present application can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 1201, the above-mentioned functions defined in the method of the present application are performed.

[0154] Another embodiment of the present application further provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of the present application.

[0155] Specifically, the present embodiment can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.

[0156] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0157] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0158] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0159] The flow chart or block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the equipment, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code include one or more executable instructions for realizing the logical function of the specification. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated system for hardware that performs the function or operation of the specification, or can be implemented with a combination of dedicated hardware and computer instructions.

[0160] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0161] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0162] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0163] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0164] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code.

[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

[0166] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. Terms such as "first" and "second" are used to indicate names and do not imply any particular order.

Claims

1. A method for data processing, wherein: The method comprises: Acquire target text data to be processed, where the target text data is obtained based on user interaction data; Performing graph processing based on similarity information between various text data contained in the target text dataset to obtain a corresponding graph structure, wherein the graph structure includes a plurality of nodes and edges, wherein the nodes correspond to text content, and the edges correspond to similarity distances between the text content; Clustering the nodes in the graph structure to divide the nodes into a plurality of communities, wherein the communities are divided based on text similarity so that the text contents belonging to a community have a high degree of similarity; Analyze the interaction data of each community based on the predetermined abnormal analysis rules to determine whether there is abnormal data; The method further comprises: If it is determined that there is abnormal data in the community, the corresponding abnormal interaction data and abnormal accounts are determined, and the determined abnormal interaction data and abnormal accounts are processed.

2. The method according to claim 1, wherein The clustering process of the nodes in the graph structure to divide the nodes in the graph structure into a plurality of communities includes: Based on a predetermined community identification algorithm, the nodes in the graph structure are divided into multiple communities.

3. The method according to claim 1, wherein In the video manuscript scenario, the method further includes: A density-based clustering algorithm is used to cluster multiple video manuscripts.

4. The method according to claim 1, wherein The graph structure obtained by performing graph composition processing based on the similarity information between the various text data contained in the target text dataset includes: Calculating text similarity between each text data included in the target text data set based on a predetermined similarity algorithm; Based on the calculated text similarity, a graph structure of the target text dataset is constructed.

5. The method according to claim 4, wherein The method further comprises: An adjustment process is performed based on the obtained graph structure, wherein the adjustment process includes at least any one of the following processes: Adjust the weights of the edges of the graph structure; Prune the graph structure.

6. The method according to claim 1, wherein The method further comprises: Obtain user interaction data within a predetermined time period in the past; The data is filtered based on predetermined data filtering rules, thereby obtaining target text data based on the filtered interactive data.

7. A device for data processing, wherein: The device comprises: means for acquiring target text data to be processed, wherein the target text data is obtained based on user interaction data; A device for performing graph processing based on similarity information between various text data contained in a target text dataset to obtain a corresponding graph structure, wherein the graph structure includes a plurality of nodes and edges, wherein the nodes correspond to text content and the edges correspond to similarity distances between the text content; means for clustering the nodes in the graph structure to divide the nodes into a plurality of communities, wherein the communities are divided based on text similarity such that the text contents belonging to a community have a high degree of similarity; A device for analyzing the interaction data of each community based on predetermined abnormality analysis rules to determine whether there is abnormal data; The device further includes an exception handling device: If it is determined that abnormal data exists in the community, the abnormality processing device determines the corresponding abnormal interaction data and abnormal accounts, and then processes the determined abnormal interaction data and abnormal accounts.

8. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A computer-readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Account data intelligent processing method and device

    CN112084422A

  • Information mining method and device, electronic equipment and storage medium

    CN117093627A