Data filtering method, related device, equipment and storage medium

By building a correlation map and similarity fusion of social networks, the problems of large and low quality of conversation data in social networks are solved, and the accuracy and real-time improvement of data filtering are achieved.

CN120541284APending Publication Date: 2025-08-26HEFEI IFLY DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410210397.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The large amount of conversation data in social networks is high and low quality, resulting in low data analysis efficiency and inability to respond in real time. Accurate filtering methods are required to improve the quality of data analysis.

Method used

By obtaining the historical session data of the target account in the target social network, using the personalized graph neural network to construct the association map of topics, accounts and social groups, combined with graph vectorization and similarity fusion, determine whether to filter the current session data.

Benefits of technology

Improves the accuracy of data filtering, and can measure the similarity between accounts and topics from macro and micro dimensions, ensuring the accuracy and real-timeness of data filtering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541284A_ABST
    Figure CN120541284A_ABST
Patent Text Reader

Abstract

The invention discloses a data filtering method, a related device, equipment and a storage medium, and the method comprises the steps: obtaining historical session data of a plurality of target accounts in a target social network; based on the historical session data, predicting first similarities between the plurality of target accounts and the target topics, and based on newly filtered target session data in the historical session data of the target accounts, determining second similarities between the corresponding target accounts and the target topics; performing fusion based on the first similarity and the second similarity to obtain fusion similarity between each target account and each target theme; and based on the fusion similarity between each target account and each target theme, determining whether to filter the current session data belonging to any target account. According to the scheme, the data filtering accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data filtering method and related devices, equipment, and storage media. Background Art

[0002] With the rapid development of electronic information technology, social networks, especially those with instant messaging functions, have been widely used in daily life, business negotiations, office entertainment and other scenarios. This has generated massive amounts of conversation data, posing huge challenges to data analysis.

[0003] Generally speaking, data analysis relies on data quality; low-quality data leads to low-quality analytical results. Furthermore, due to hardware resource limitations, analyzing the entire volume of social network conversation data would significantly reduce efficiency, making real-time operation and response impossible. Therefore, precise filtering of social data is essential. Summary of the Invention

[0004] The main technical problem solved by this application is to provide a data filtering method and related devices, equipment and storage media, which can improve the accuracy of data filtering.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a data filtering method, including: obtaining historical conversation data of several target accounts in a target social network; based on the historical conversation data, predicting the first similarities between the several target accounts and each target topic, and based on the newly filtered target conversation data in the historical conversation data of the target account, determining the second similarities between the corresponding target account and each target topic; fusing the first similarity and the second similarity to obtain the fused similarity between each target account and each target topic; and determining whether to filter the current conversation data belonging to any target account based on the fused similarity between each target account and each target topic.

[0006] In order to solve the above technical problems, the second aspect of the present application provides a data filtering device, including: a data acquisition module, a similarity measurement module, a similarity fusion module and a filtering analysis module, the data acquisition module is used to obtain historical session data of several target accounts in the target social network; the similarity measurement module is used to predict the first similarities between several target accounts and each target topic based on the historical session data, and determine the second similarities between the corresponding target accounts and each target topic based on the newly filtered target session data in the historical session data of the target accounts; the similarity fusion module is used to fuse the first similarity and the second similarity to obtain the fused similarity between each target account and each target topic; the filtering analysis module is used to determine whether to filter the current session data belonging to any target account based on the fused similarity between each target account and each target topic.

[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the data filtering method in the above first aspect.

[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be executed by a processor, and the program instructions are used to implement the data filtering method of the first aspect.

[0009] The above scheme obtains historical conversation data for several target accounts in a target social network, predicts first similarities between each of the target accounts and each target topic based on the historical conversation data, and determines second similarities between each of the target accounts and each target topic based on newly filtered target conversation data from the target accounts' historical conversation data. Based on this, the first and second similarities are fused to obtain a fused similarity between each target account and each target topic. Based on this fused similarity, a determination is then made as to whether to filter the current conversation data belonging to any target account. This allows for both macro-level predictions of the similarity between the target account and the target topic based on the historical conversation data, and micro-level predictions of the similarity between the target account and the target topic based on the newly filtered target conversation data from the historical conversation data. Combining these predicted macro- and micro-level similarities allows for the most accurate measurement of the similarity between the target account and the target topic, which is then used to determine whether to filter the current conversation data, thereby improving data filtering accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a flow chart of an embodiment of the data filtering method of the present application;

[0011] Figure 2a This is a process diagram of an embodiment of the data filtering method of the present application;

[0012] Figure 2b is a schematic diagram of the first atlas;

[0013] Figure 2c is a schematic diagram of the second atlas;

[0014] Figure 2d is a schematic diagram of the third atlas;

[0015] Figure 3 This is a schematic diagram of the framework of an embodiment of the data filtering device of the present application;

[0016] Figure 4 This is a schematic diagram of the framework of an embodiment of the electronic device of the present application;

[0017] Figure 5 It is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0018] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0019] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0020] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" is simply a description of an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the fragment " / " generally indicates that the related objects are in an "or" relationship. Furthermore, "multiple" in this document refers to two or more than two.

[0021] See also Figure 1 , Figure 1 It is a flow chart of an embodiment of the data filtering method of the present application.

[0022] Specifically, the following steps may be included:

[0023] Step S11: Obtain historical session data of several target accounts in the target social network.

[0024] In one implementation scenario, the target social network can be set based on actual application needs. For example, if data filtering is required for social network A, the target social network can be set to social network A; or, if data filtering is required for social network B, the target social network can be set to social network B, and so on. Examples are not given here one by one.

[0025] In one implementation scenario, the target account may be all accounts in the target social network, that is, all accounts in the target social network may be selected as target accounts; or, the target account may be some accounts in the target social network, that is, some accounts in the target social network may be selected as target accounts separately, without limitation here.

[0026] In one implementation scenario, the historical session data may be the full amount of session data that has been generated by the target account, that is, all session data of the target account since its opening can be selected as historical session data; or, the historical session data may also be the session data generated by the target account in a recent preset time period, that is, the historical session data can be selected from the session data generated by the target account in a preset time period (such as one day, one week, etc.) from now on, without limitation. It should be noted that the historical session data may come from the social group where the target account is located. For example, if target account A and target account C are both in social group 1, the session data generated by target account A and target account C in social group 1 can be collected as their respective historical session data. Of course, the historical session data can be marked with the target account to which it belongs and the social group to which it belongs.

[0027] Step S12: Based on the historical conversation data, predict the first similarities between several target accounts and each target topic, and based on the newly filtered target conversation data in the historical conversation data of the target account, determine the second similarities between the corresponding target account and each target topic.

[0028] In one implementation scenario, target topics can be set based on specific data filtering requirements. For example, based on data filtering requirements, if conversation data involving topics such as advertising, entertainment, food, and games is considered low-quality data and needs to be filtered, then topics such as advertising, entertainment, food, and games can be set as target topics. Of course, other topics can be defined based on specific requirements in a similar manner, and no further examples are given here.

[0029] In an implementation scenario, please refer to Figure 2a , Figure 2a This is a process diagram of an embodiment of the data filtering method of this application. Figure 2aAs shown, after obtaining historical conversation data, a personalized graph neural network can be trained based on the historical conversation data, and then predictions can be made based on the personalized graph neural network to obtain a first graph containing the first correlation between each target topic, a second graph containing the second correlation between each target account, and a third graph containing the third correlation between each social group. Therefore, based on the first graph and the second graph, the first topic feature of the target topic can be extracted, and based on the second graph and the third graph, the first account feature of the target account can be extracted. Then, based on the first account feature of the target account and the first topic feature of the target topic, the first similarity between the target account and the target topic can be measured. The above method predicts the first graph about the topic, the second graph about the account, and the third graph about the social group, so as to extract the topic feature based on the first two and the account feature based on the latter two, and measure the first similarity. Therefore, the similarity between the account feature and the topic feature can be measured through multi-source graphs, which helps to improve the accuracy of the similarity measurement.

[0030] In a specific implementation scenario, the personalized graph neural network can include, but is not limited to, being constructed through PPNP. It should be noted that PPNP can directly simulate the complete propagation process of features without the need for multiple convolutional layers to simulate multiple propagations, and there is a probability of jumping back to the initial node during the propagation process, thereby balancing neighborhood features of different sizes. In addition, the specific process of obtaining a personalized graph neural network through training of historical session data can be referred to the technical details of PPNP, which will not be repeated here.

[0031] In a specific implementation scenario, the personalized graph neural network can predict the first correlation between each target topic, and thus construct a first graph. For example, when the first correlation between two target topics is greater than a preset threshold, the two target topics can be connected. Conversely, when the first correlation between two target topics is not greater than the preset threshold, the two target topics can be disconnected. Please refer to 2b, Figure 2b is a schematic diagram of the first spectrum. Figure 2b As shown, taking the target topics including advertising, entertainment, food and games as an example, the first correlation between the above target topics can be predicted by the personalized graph neural network (not shown in the figure). According to the first correlation, the target topics "advertisement" and "game", "advertisement" and "food", "advertisement" and "entertainment" can be connected, and the target topics "entertainment" and "food" can be connected in the form of Figure 2b The first spectrum shown. Of course, Figure 2b What is shown is only a possible example of the first graph, and does not limit the specific target topics contained in the first graph and the specific connections between the target topics.

[0032] In a specific implementation scenario, the personalized graph neural network can predict the second correlation between each target account, and thus construct a second graph. For example, when the second correlation between two target accounts is greater than a preset threshold, the two target accounts can be connected. Conversely, when the first correlation between the two target accounts is not greater than the preset threshold, the two target accounts can be disconnected. Figure 2c , Figure 2c is a schematic diagram of the second spectrum. Figure 2c As shown, taking the target accounts including Account A to Account D as an example, the second correlation degree between the above target accounts can be predicted separately through the personalized graph neural network (not shown in the figure). According to the second correlation degree, the target accounts "Account A" and "Account B", the target accounts "Account A" and "Account C", the target accounts "Account A" and "Account D", the target accounts "Account B" and "Account D", and the target accounts "Account C" and "Account D" can be connected. Of course, Figure 2c What is shown is only a possible example of the second graph, and does not limit the specific target accounts contained in the second graph and the specific connections between the target accounts.

[0033] In a specific implementation scenario, the personalized graph neural network can predict the third degree of connection between various social groups, and thus construct a third graph. For example, when the third degree of connection between two social groups is greater than a preset threshold, the two social groups can be connected. Conversely, when the third degree of connection between two social groups is not greater than the preset threshold, the two social groups can be disconnected. Figure 2d , Figure 2d is a schematic diagram of the third spectrum. Figure 2d As shown, taking the social groups including social groups 1 to 3 as an example, the third correlation degree between the above social groups can be predicted by the personalized graph neural network (not shown in the figure). According to the third correlation degree, the social groups "social group 1" and "social group 2" and the social groups "social group 1" and "social group 3" can be connected. Of course, Figure 2d The diagram is merely a possible example of the third graph, and does not limit the specific social groups contained in the third graph and the specific connections between the social groups.

[0034] In a specific implementation scenario, after obtaining the first graph and the second graph, a first heterogeneous graph can be constructed based on the first graph and the second graph, and the first heterogeneous graph includes a first connection weight between the target topic and the target account, and the first connection weight is obtained based on the first correlation metric. On this basis, graph quantization can be performed based on the first heterogeneous graph to obtain the first topic features of each target topic in the first heterogeneous graph. For example, for the target topic Ti, the first graph related to the target topic T iThe adjacent topic set can be recorded as C, and the first correlation degree between the target topics connected to the target topic Ti can be recorded as In addition, the topic set adjacent to the above topic set C can be recorded as D, then the target topic T i The first connection weight with the target account in the second graph can be expressed as:

[0035]

[0036] In the above formula (1), represents the first connection weight, α k Represents a learnable hyperparameter. At this point, the first heterogeneous graph can be obtained. Based on this, the graph network can be represented as F, L is the number of network layers, and the first topic feature T of the target topic can be obtained through F. emb :

[0037]

[0038] In the above formula (2), F L represents the Lth layer of the graph network, F L-1 Represents the L-1 layer of the graph network. The above method constructs a first heterogeneous graph based on the first and second graphs. The first heterogeneous graph includes a first connection weight between the target topic and the target account, and the first connection weight is obtained based on a first correlation metric. Graph quantization is then performed based on the first heterogeneous graph to obtain first topic features for each target topic in the first heterogeneous graph. This method can measure topic features using the heterogeneous graph, helping to improve the accuracy of topic features.

[0039] In a specific implementation scenario, after obtaining the second graph and the third graph, a second heterogeneous graph can be constructed based on the second graph and the third graph, and the second heterogeneous graph includes a second connection weight between the target account and the social group, and the second connection weight is obtained based on the third correlation metric. On this basis, graph quantization is performed based on the second heterogeneous graph to obtain the first account features of each target account in the second heterogeneous graph. For example, for the social group G i , the third correlation degree between the social groups connected to the social group in the third graph can be recorded as The set of groups adjacent to the social group in the third graph can be denoted as P, then the social group G i The second connection weight with the target account in the second graph can be expressed as:

[0040]

[0041] In the above formula (3), represents the second connection weight, β kRepresents a learnable hyperparameter. At this point, the second heterogeneous graph is obtained. Based on this, the graph network can be represented as F, where L is the number of network layers. Through F, the first account feature A of the target account can be obtained. emb :

[0042]

[0043] In the above formula (4), F L represents the Lth layer of the graph network, F L-1 Represents the L-1 layer of the graph network. The above method constructs a second heterogeneous graph based on the second and third graphs; the second heterogeneous graph includes second connection weights between target accounts and social groups, and the second connection weights are measured based on the third association degree. Graph quantization is performed based on the second heterogeneous graph to obtain first account features for each target account in the second heterogeneous graph. Account features can be measured using the heterogeneous graph, which helps improve the accuracy of account features.

[0044] In an implementation scenario, please continue to refer to Figure 2a After obtaining the historical conversation data, the second account feature of the target account can be updated based on each target conversation data in the target account, and vectorized based on the target topic to obtain the second topic feature of the target topic. On this basis, the second similarity between the target account and the target topic can be measured based on the second account feature of the target account and the second topic feature of the target topic. The above method obtains the second account feature of the target account by updating the target conversation data in the target account, and obtains the second similarity based on this and the topic feature of the target topic. It can synchronize the feature representation of the target account with the conversation update of the target social network, which helps to improve the accuracy of the second similarity.

[0045] In a specific implementation scenario, the session vector features of each target session data in the target account can be fused to obtain the second account feature of the target account, and the session vector features of each target session data in the target account can be fused to obtain the second account feature of the target account. For example, taking the newly filtered target session data of a target account including: session data 1, session data 2, ..., session data n as an example, the above session data can be vectorized to obtain the session vector features, which can be recorded as: {v1, v2, ...v n}, the session vector features can be fused by averaging, weighting, etc. to obtain the second account feature of the target account. Taking averaging as an example, the second account feature can be expressed as:

[0046]

[0047] In the above formula (5), u represents the second account feature, v i represents the session vector feature of the i-th target session data, and n represents the total number of target session data in the target account. This approach vectorizes each target session data in the target account to obtain a session vector feature for each target session data. The session vector features of each target session data in the target account are then fused to obtain the second account feature of the target account. This allows the feature representation of the target account to be updated synchronously with the session updates of the target account on the target social network, helping to improve the accuracy of the second similarity measure.

[0048] In a specific implementation scenario, as a possible implementation method, the above vectorization can be achieved through a network model such as the Deep Structured Semantic Model (DSSM). It should be noted that the core idea of ​​the Twin Towers Model is to map topics and accounts into a semantic space of common dimensions, and to train a latent semantic model by maximizing the cosine similarity between the semantic vectors of topics and accounts. Assuming that x represents the input topic feature and y represents the vector representation after DNN, the calculation is as follows:

[0049] l1=W1x……(6)

[0050] l i =f(W i l i-1 +b i ),i=2,…,N-1……(7)

[0051] y=f(W N l N-1 +b N )……(8)

[0052] In the above formulas (6) to (8), W,b are the parameters of the model, and f is the activation function, where f is the tanh activation function. Based on this, the similarity can be calculated as:

[0053]

[0054] In the above formula (10), y Q and y D are the semantic vectors of the topic and account. As previously mentioned, during training, the cosine similarity between the topic and account semantic vectors is maximized to train a latent semantic model, enabling vectorization using the Twin Towers model. It should be noted that the above is only one possible example of vectorization using the Twin Towers model and does not limit the steps involved in implementing vectorization in practice.

[0055] Step S13: Fusion is performed based on the first similarity and the second similarity to obtain the fusion similarity between each target account and each target topic.

[0056] Specifically, for a target topic and a target account, the first similarity and the second similarity can be fused by averaging, summing, weighting, etc. to obtain the fused similarity between the target topic and the target account. Of course, for ease of processing, the fused similarity between each target topic and each target account can be represented by a similarity matrix. Please continue to refer to Figure 2b and 2c , taking the target topics including advertising, games, entertainment, and food as an example, and the target accounts including accounts A to D, the similarity matrix can be expressed as:

[0057] Account A Account B Account C Account D game 0.1 0.32 0.57 0.2 advertise 0.19 0.08 0.23 0.1 entertainment 0.43 0.27 0.14 0.3 gourmet food 0.2 0.3 0.05 0.4

[0058] Of course, the above example is only a possible example of the similarity matrix in actual application, and does not limit the specific scale and content of the similarity matrix.

[0059] Step S14: Based on the fusion similarity between each target account and each target topic, determine whether to filter the current session data belonging to any target account.

[0060] Specifically, as mentioned above, a similarity matrix can be constructed based on the fused similarity between each target account and each target topic. On this basis, decomposition can be performed based on the similarity matrix to obtain the account latent features of each target account and the topic latent features of each target topic. Then, based on the account latent features of each target account and the topic latent features of each target topic, it is determined whether to filter the current session data belonging to any target account. The above method, by decomposing the similarity matrix to obtain the topic latent features and account latent features, and performing data filtering based on them, can achieve data filtering through semantic information, help to improve the refinement of data filtering, and thus improve the accuracy of data filtering.

[0061] In one implementation scenario, after obtaining the similarity matrix, SVD (Singular Value Decomposition) can be used to decompose the similarity matrix to obtain the topic latent features of each target topic and the account latent features of each target account. It should be noted that the specific process of decomposing the similarity matrix can be referred to the technical details of decomposition techniques such as SVD, and will not be repeated here.

[0062] In one implementation scenario, after obtaining the account latent features of each target account and the topic latent features of each target topic, for the current session data belonging to any target account, the similarity between the account latent features of the target account to which the current session data belongs and the topic latent features of each target topic can be obtained. If there is at least one target topic whose topic latent features meet a preset condition and the account latent features of the target account to which the current session data belongs, it can be determined to filter the current session data. It should be noted that the preset condition can be set to a similarity greater than a preset threshold. Furthermore, the specific value of the preset threshold can be set according to actual application needs. For example, when the tolerance for data filtering is relatively loose, the preset threshold can be set to a relatively small value. Alternatively, when the tolerance for data filtering is relatively high, the preset threshold can be set to a relatively large value. In this manner, determining to filter the current session data in response to the similarity between the account latent features of the target account to which the current session data belongs and the topic latent features of at least one target topic meeting the preset condition can minimize the omission of session data that should be filtered, thereby helping to improve the accuracy of data filtering.

[0063] In one implementation scenario, as described above, after obtaining the account latent features of each target account and the subject latent features of each target topic, for the current session data belonging to any target account, the similarity between the account latent features of the target account to which the current session data belongs and the subject latent features of each target topic can be obtained. If the similarity between the subject latent features of each target topic and the account latent features of the target account to which the current session data belongs does not meet the preset conditions, it can be determined to retain the current session data. It should be noted that the specific setting of the preset conditions can refer to the aforementioned related descriptions and will not be repeated here. In the above method, in response to the fact that the similarity between the account latent features of the target account to which the current session data belongs and the subject latent features of each target topic does not meet the preset conditions, it is determined to retain the current session data, so that it can avoid the misfiltering of session data that should not be filtered as much as possible, which helps to improve the accuracy of data filtering.

[0064] In one implementation scenario, if the current session data is determined to be filtered, the newly filtered target session data in the historical session data of the target account to which the current session data belongs can be updated, and the process returns to step S12 to determine whether to filter new current session data belonging to any target account. It should be noted that if the current session data is determined to be filtered, as previously described, the newly filtered target session data in the target account to which the current session data belongs will change, thereby also changing the second similarity. Furthermore, since the historical session data is constantly updated, the first similarity and the second similarity can be re-updated to determine whether to perform data filtering on the new current session data. In the above method, in response to the determination to filter the current session data, the newly filtered target session data in the historical session data of the target account to which the current session data belongs is updated, and the process returns to the step of predicting the first similarity between several target accounts and each target topic based on the historical session data to determine whether to filter new current session data belonging to any target account. Therefore, the first similarity and the second similarity can be automatically updated as the session data continuously changes, thereby determining whether to filter the new current session data, thereby further improving the accuracy of data filtering.

[0065] The above scheme obtains historical conversation data for several target accounts in a target social network, predicts first similarities between each of the target accounts and each target topic based on the historical conversation data, and determines second similarities between each of the target accounts and each target topic based on newly filtered target conversation data from the target accounts' historical conversation data. Based on this, the first and second similarities are fused to obtain a fused similarity between each target account and each target topic. Based on this fused similarity, a determination is then made as to whether to filter the current conversation data belonging to any target account. This allows for both macro-level predictions of the similarity between the target account and the target topic based on the historical conversation data, and micro-level predictions of the similarity between the target account and the target topic based on the newly filtered target conversation data from the historical conversation data. Combining these predicted macro- and micro-level similarities allows for the most accurate measurement of the similarity between the target account and the target topic, which is then used to determine whether to filter the current conversation data, thereby improving data filtering accuracy.

[0066] See also Figure 3 , Figure 3Schematic diagram of a framework of an embodiment of a data filtering device 30 of the present application. The data filtering device 30 includes: a data acquisition module 31, a similarity measurement module 32, a similarity fusion module 33, and a filtering analysis module 34. The data acquisition module 31 is used to acquire historical conversation data of several target accounts in a target social network; the similarity measurement module 32 is used to predict the first similarities between several target accounts and each target topic based on the historical conversation data, and to determine the second similarities between the corresponding target accounts and each target topic based on the newly filtered target conversation data in the historical conversation data of the target accounts; the similarity fusion module 33 is used to fuse the first similarity and the second similarity to obtain the fused similarity between each target account and each target topic; and the filtering analysis module 34 is used to determine whether to filter the current conversation data belonging to any target account based on the fused similarity between each target account and each target topic.

[0067] In the above scheme, the data filtering device 30 obtains historical conversation data of several target accounts in the target social network, predicts first similarities between each of the target accounts and each target topic based on the historical conversation data, and determines second similarities between each of the target accounts and each target topic based on the newly filtered target conversation data from the target accounts' historical conversation data. Based on this, the first and second similarities are fused to obtain a fused similarity between each target account and each target topic. Based on this fused similarity, the device then determines whether to filter the current conversation data belonging to any target account. This allows the device to predict the similarity between the target account and the target topic from a macro perspective based on the historical conversation data, and to predict the similarity between the target account and the target topic from a micro perspective based on the newly filtered target conversation data from the historical conversation data. Combining these predicted similarities from the macro and micro perspectives allows the device to measure the similarity between the target account and the target topic as accurately as possible, which is then used to determine whether to filter the current conversation data, thereby improving the accuracy of data filtering.

[0068] In some disclosed embodiments, the similarity measurement module 32 includes a network training submodule for training a personalized graph neural network based on historical session data; the similarity measurement module 32 includes a graph prediction submodule for performing prediction based on the personalized graph neural network to obtain a first graph containing a first correlation between each target topic, a second graph containing a second correlation between each target account, and a third graph containing a third correlation between each social group; the similarity measurement module 32 includes a first extraction submodule for extracting a first topic feature of the target topic based on the first graph and the second graph; the similarity measurement module 32 includes a second extraction submodule for extracting a first account feature of the target account based on the second graph and the third graph; the similarity measurement module 32 includes a first similarity submodule for measuring a first similarity between the target account and the target topic based on the first account feature of the target account and the first topic feature of the target topic.

[0069] In some disclosed embodiments, the first extraction submodule includes a first construction unit for constructing a first heterogeneous graph based on the first graph and the second graph; wherein the first heterogeneous graph includes a first connection weight between the target topic and the target account, and the first connection weight is obtained based on a first association measure; the first extraction submodule includes a first vector unit for performing graph vectorization based on the first heterogeneous graph to obtain first topic features of each target topic in the first heterogeneous graph.

[0070] In some disclosed embodiments, the second extraction submodule includes a second construction unit for constructing a second heterogeneous graph based on the second graph and the third graph; wherein the second heterogeneous graph includes a second connection weight between the target account and the social group, and the second connection weight is obtained based on the third association measure; the second extraction submodule includes a second extraction unit for performing graph quantization based on the second heterogeneous graph to obtain the first account features of each target account in the second heterogeneous graph.

[0071] In some disclosed embodiments, the similarity measurement module 32 includes a feature update submodule for updating the second account feature of the target account based on each target session data in the target account; the similarity measurement module 32 includes a topic vector submodule for performing vectorization based on the target topic to obtain the second topic feature of the target topic; the similarity measurement module 32 includes a second similarity submodule for measuring the second similarity between the target account and the target topic based on the second account feature of the target account and the second topic feature of the target topic.

[0072] In some disclosed embodiments, the feature update submodule includes a session vectorization unit, which is used to perform vectorization based on each target session data in the target account to obtain the session vector features of each target session data; the feature update submodule includes a vector fusion unit, which is used to perform fusion based on the session vector features of each target session data in the target account to obtain the second account feature of the target account.

[0073] In some disclosed embodiments, the filtering analysis module 34 includes a matrix construction submodule for constructing a similarity matrix based on the fused similarity between each target account and each target topic; the filtering analysis module 34 includes a matrix decomposition submodule for performing decomposition based on the similarity matrix to obtain the account latent features of each target account and the topic latent features of each target topic; the filtering analysis module 34 includes a filtering determination submodule for determining whether to filter the current session data belonging to any target account based on the account latent features of each target account and the topic latent features of each target topic.

[0074] In some disclosed embodiments, the filtering determination submodule includes a first response unit for determining to filter the current session data in response to the similarity between the account latent features of the target account to which the current session data belongs and the subject latent features of at least one target subject satisfying a preset condition; the filtering determination submodule includes a second response unit for determining to retain the current session data in response to the similarity between the account latent features of the target account to which the current session data belongs and the subject latent features of each target subject not satisfying the preset condition.

[0075] In some disclosed embodiments, the data filtering device 30 also includes a data update module for updating the newly filtered target session data in the historical session data of the target account to which the current session data belongs in response to determining to filter the current session data; the data filtering device 30 also includes a loop execution module for returning to the step of predicting the first similarity between several target accounts and each target topic based on the historical session data to determine whether to filter the new current session data belonging to any target account.

[0076] See also Figure 4 , Figure 4 This is a schematic diagram of an embodiment of an electronic device 40 of the present application. Electronic device 40 includes a memory 41 and a processor 42. Memory 41 stores program instructions, and processor 42 is configured to execute the program instructions to implement the steps in any of the aforementioned data filtering method embodiments. For details, please refer to the previously disclosed embodiments and will not be further described here. Electronic device 40 may specifically include, but is not limited to, a server, etc., and is not limited here.

[0077] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above-mentioned data filtering method embodiments. The processor 42 can also be called a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 42 can be implemented by an integrated circuit chip.

[0078] In the above scheme, electronic device 40 obtains historical conversation data for several target accounts in the target social network, predicts first similarities between each of the target accounts and each target topic based on the historical conversation data, and determines second similarities between each of the target accounts and each target topic based on the newly filtered target conversation data from the target accounts' historical conversation data. Based on this, electronic device 40 fuses the first and second similarities to obtain a fused similarity between each target account and each target topic. Based on this fused similarity, electronic device 40 determines whether to filter the current conversation data belonging to any target account. This allows for both macro-level predictions of similarity between the target account and the target topic based on the historical conversation data, and micro-level predictions of similarity between the target account and the target topic based on the newly filtered target conversation data from the historical conversation data. Combining these predicted macro- and micro-level similarities allows for the most accurate measurement of the similarity between the target account and the target topic, which is then used to determine whether to filter the current conversation data, thereby improving data filtering accuracy.

[0079] See also Figure 5 , Figure 5 1 is a schematic diagram of a framework of an embodiment of a computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be executed by a processor, and the program instructions 51 are used to implement the steps of any of the above data filtering method embodiments.

[0080] In the above scheme, the computer-readable storage medium 50 obtains historical conversation data of several target accounts in the target social network, predicts first similarities between each of the target accounts and each target topic based on the historical conversation data, and determines second similarities between each of the target accounts and each target topic based on newly filtered target conversation data from the target accounts' historical conversation data. Based on this, the first and second similarities are fused to obtain a fused similarity between each target account and each target topic. Based on this fused similarity, a determination is then made as to whether to filter the current conversation data belonging to any target account. This allows, on the one hand, to predict the similarity between the target account and the target topic from a macro perspective based on the historical conversation data, and, on the other hand, to predict the similarity between the target account and the target topic from a micro perspective based on newly filtered target conversation data from the historical conversation data. Combining these predicted similarities from the macro and micro perspectives allows the similarity between the target account and the target topic to be measured as accurately as possible, which is then used to determine whether to filter the current conversation data, thereby improving the accuracy of data filtering.

[0081] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0082] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0083] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0084] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0085] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0086] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0087] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

Claims

1. A data filtering method, characterized in that: include: Obtain historical conversation data of several target accounts in the target social network; Predicting first similarities between each of the target accounts and each target topic based on the historical conversation data, and determining second similarities between each of the target accounts and each target topic based on newly filtered target conversation data from the historical conversation data of the target accounts; Fusion is performed based on the first similarity and the second similarity to obtain fused similarities between each target account and each target topic; Based on the fusion similarity between each target account and each target topic, it is determined whether to filter the current session data belonging to any target account.

2. The method according to claim 1, characterized in that The predicting, based on the historical conversation data, first similarities between the plurality of target accounts and each target topic, includes: Based on the historical session data, a personalized graph neural network is trained; Predicting based on the personalized graph neural network, obtaining a first graph including a first correlation between each of the target topics, a second graph including a second correlation between each of the target accounts, and a third graph including a third correlation between each of the social groups; Extracting a first topic feature of the target topic based on the first graph and the second graph, and extracting a first account feature of the target account based on the second graph and the third graph; Based on the first account feature of the target account and the first topic feature of the target topic, a first similarity between the target account and the target topic is measured.

3. The method according to claim 2, characterized in that The extracting a first topic feature of the target topic based on the first graph and the second graph includes: Constructing a first heterogeneous graph based on the first graph and the second graph; wherein the first heterogeneous graph includes a first connection weight between the target topic and the target account, and the first connection weight is obtained based on the first association metric; Graph quantization is performed based on the first heterogeneous graph to obtain first topic features of each target topic in the first heterogeneous graph.

4. The method according to claim 2, characterized in that The extracting the first account feature of the target account based on the second graph and the third graph includes: Constructing a second heterogeneous graph based on the second graph and the third graph; wherein the second heterogeneous graph includes a second connection weight between the target account and the social group, and the second connection weight is obtained based on the third association metric; Graph quantization is performed based on the second heterogeneous graph to obtain the first account feature of each of the target accounts in the second heterogeneous graph.

5. The method according to claim 1, wherein The determining, based on the newly filtered target conversation data in the historical conversation data of the target account, the second similarity between the corresponding target account and each of the target topics, includes: Based on each target session data in the target account, a second account feature of the target account is updated, and vectorization is performed based on the target topic to obtain a second topic feature of the target topic; Based on the second account feature of the target account and the second topic feature of the target topic, a second similarity between the target account and the target topic is measured.

6. The method according to claim 5, characterized in that The updating and obtaining of the second account feature of the target account based on each target session data in the target account includes: Vectorizing each target session data in the target account to obtain a session vector feature of each target session data; The second account feature of the target account is obtained by fusing the session vector features of each target session data in the target account.

7. The method according to claim 1, characterized in that The determining whether to filter the current session data belonging to any of the target accounts based on the fusion similarity between each of the target accounts and each of the target topics includes: Constructing a similarity matrix based on the fusion similarities between each of the target accounts and each of the target topics; Decomposing the similarity matrix to obtain the account latent features of each target account and the topic latent features of each target topic; Based on the account implicit features of each of the target accounts and the topic implicit features of each of the target topics, it is determined whether to filter the current session data belonging to any of the target accounts.

8. The method according to claim 7, characterized in that The determining whether to filter the current session data belonging to any of the target accounts based on the account implicit features of each of the target accounts and the topic implicit features of each of the target topics includes at least one of the following: In response to a similarity between an account latent feature of the target account to which the current session data belongs and a topic latent feature of at least one target topic meeting a preset condition, determining to filter the current session data; In response to the similarity between the account latent feature of the target account to which the current session data belongs and the topic latent features of each of the target topics not satisfying a preset condition, it is determined to retain the current session data.

9. The method according to any one of claims 1 to 8, characterized in that After determining whether to filter the current session data belonging to any of the target accounts based on the fusion similarities between each of the target accounts and each of the target topics, the method further includes: In response to determining to filter the current session data, updating the historical session data of the target account to which the current session data belongs with newly filtered target session data; Return to the step of predicting the first similarities between the target accounts and the target topics based on the historical conversation data to determine whether to filter new current conversation data belonging to any of the target accounts.

10. A data filtering device, characterized in that: include: A data acquisition module is used to obtain historical session data of several target accounts in the target social network; a similarity measurement module, configured to predict, based on the historical conversation data, a first similarity between each of the plurality of target accounts and each target topic, and determine, based on newly filtered target conversation data from the historical conversation data of the target accounts, a second similarity between each of the target accounts and each of the target topics; a similarity fusion module, configured to fuse the first similarity and the second similarity to obtain a fused similarity between each target account and each target topic; The filtering analysis module is used to determine whether to filter the current session data belonging to any of the target accounts based on the fusion similarity between each of the target accounts and each of the target topics.

11. An electronic device, characterized in that: The method comprises at least a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the data filtering method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the data filtering method according to any one of claims 1 to 9.