User generation content judgment method based on user historical behaviors

By obtaining the topic tags and historical application behavior of the generated content to be determined by the target user and calculating the similarity, the problem of difficulty for users to judge the authenticity of the content is solved, and the effective identification of the content automatically generated by AI is achieved.

CN120145033APending Publication Date: 2025-06-13ZHEJIANG MEIRI HUDONG NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510241496.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

During the process of information dissemination and sharing, it is difficult for users to effectively determine whether the articles shared by the target user are created by the target user themselves or are automatically generated through AI.

Method used

By obtaining the collection of topic tags for the content to be determined by the target user, filtering the application set in the historical behavior of the target user based on these tags, obtaining the collection of portrait tags for the target user, calculating the similarity between the target user and the content to be determined, and determining the authenticity of the content.

Benefits of technology

It realizes reliable judgment on the authenticity of the content generated by the target user, and improves the user's recognition efficiency of automatically generated content of AI.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145033A_ABST
    Figure CN120145033A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric digital data processing, in particular to a user generation content judgment method based on user historical behaviors. The method comprises the following steps: acquiring a theme label set B of to-be-determined generation content of a target user; acquiring an application set APP used by the target user in a historical time period; screening the APPs according to the B to obtain a screened application set APP '; obtaining a portrait label set D of the target user according to the APP '; obtaining the similarity sim between the target user and the to-be-determined generation content of the target user according to B and D; if sim is greater than or equal to sim0, judging that the to-be-judged generation content of the target user is the content generated by the target user; and otherwise, judging that the to-be-judged generation content of the target user is not the content generated by the target user. According to the invention, whether the content generated by the target user is true or false can be judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electrical digital data processing, and particularly to a method for judging user-generated content based on user historical behavior. Background Art

[0002] With the rapid development of artificial intelligence technology, AI automatic generation technology has become increasingly popular, and various types of content generated by users with the help of this technology are increasing. For example, articles can be automatically generated by AI. This makes other users fall into a dilemma when seeing an article shared by a certain user (denoted as the target user) during the process of information dissemination and sharing: it is very difficult for other users to effectively distinguish whether the article they see is actually created by the target user based on the knowledge they have, or is automatically generated by the target user with the help of AI automatic generation technology and other means. In other words, for the article shared by the target user, whether it is created by the target user based on knowledge (that is, the user-generated content is true), or is automatically generated by the user using means such as AI (that is, the user-generated content is false), other users lack a reliable basis for judgment. How to judge the authenticity of the content generated by the target user is an urgent problem to be solved. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for judging user-generated content based on user historical behavior to judge the authenticity of the content generated by the target user.

[0004] According to the present invention, there is provided a method for judging user-generated content based on user historical behavior, and the method includes the following steps:

[0005] S100, obtaining a set B of topic tags of the to-be-determined generated content of the target user; B = {b 1 , b 2 , …, b i , …, b n}, where b i is the i-th topic tag of the to-be-determined generated content of the target user, and the value range of i is from 1 to n, and n is the number of topic tags of the to-be-determined generated content of the target user.

[0006] S200, obtaining a set APP of application programs used by the target user within the historical time period; APP = {app 1 , app 2 , …, app j , …, app m}, where app j is the j-th application program used by the target user within the historical time period, and the value range of j is from 1 to m, and m is the number of application programs used by the target user within the historical time period.

[0007] S300, Screen the APP according to B to obtain the screened application set APP'; APP' = {app' 1 , app' 2 , …, app' k , …, app' u}, where app' k is the k-th application relevant to B used by the target user within the historical time period, and the value range of k is from 1 to u, where u is the number of applications relevant to B used by the target user within the historical time period.

[0008] S400, Obtain the portrait label set D of the target user according to APP'; D = {d 1 , d 2 , …, d t , …, d T}, where d t is the t-th portrait label of the target user, and the value range of t is from 1 to T, where T is the number of portrait labels of the target user.

[0009] S500, Obtain the similarity sim between the target user and the content to be determined for the target user according to B and D.

[0010] S600, If sim ≥ sim 0 , then determine that the content to be determined for the target user is the content generated by the target user; otherwise, determine that the content to be determined for the target user is not the content generated by the target user; sim 0 is the preset similarity threshold.

[0011] The present invention has at least the following beneficial effects compared with the prior art:

[0012] The present invention first obtains the set of topic labels of the content to be determined for the target user, and each topic label in this set is a possible topic label of the content to be determined for the target user; then, according to the applications used by the target user within the historical time period, the present invention obtains the set of portrait labels of the target user. By calculating the similarity between the topic labels in the set of topic labels and the portrait labels in the set of portrait labels of the target user, the similarity between the target user and the content to be determined for the target user is obtained. If this similarity is high, it is determined that the content to be determined for the target user is the content generated by the target user, that is, the content personally created by the target user based on the knowledge they have mastered; otherwise, it is determined that the content to be determined for the target user is not the content generated by the target user, that is, the content automatically generated by means such as AI. Thus, the present invention realizes the true or false judgment of the content to be determined for the target user based on the topic labels of the content to be determined for the target user and the portrait labels of the target user.

[0013] Moreover, the portrait tags of the present invention are not obtained based on all apps used by the target user within a historical time period. Instead, first, the above-mentioned all apps are screened according to the theme tags of the content to be determined generated by the target user, and only the apps relevant to the theme tags of the content to be determined generated by the target user are retained. Since multiple portrait tags of the target user may be obtained based on the behavior of the target user using one app, therefore, the method of the present invention for screening apps first and then obtaining the portrait tags of the target user can greatly reduce the number of obtained target portrait tags on the premise of ensuring the accuracy of the final judgment result, and greatly reduce the workload of obtaining the similarity between the target user and the content to be determined generated by the target user, thereby improving the judgment efficiency of the authenticity of the content to be determined generated by the target user as a whole. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0015] Figure 1 It is a flowchart of a method for judging user-generated content based on user historical behavior provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0017] According to this embodiment, as Figure 1 shown, a method for judging user-generated content based on user historical behavior is provided. The method includes the following steps:

[0018] S100, obtaining a set B of theme tags of the content to be determined generated by the target user; B = {b 1 , b 2 , …, b i , …, b n}, where b i is the i-th theme tag of the content to be determined generated by the target user, and the value range of i is from 1 to n, and n is the number of theme tags of the content to be determined generated by the target user.

[0019] In this embodiment, the topic tags of the content to be determined for generation are used to characterize the topic of the content to be determined for generation. For example, if the content to be determined for generation is an article, the set of topic tags of this article is {Artificial Intelligence, Healthcare}.

[0020] As a preferred specific implementation manner, S100 includes:

[0021] S110, input the text in the content to be determined for generation of the target user into the trained first neural network model, and obtain the output result E of the trained first neural network model, E = ((e 1, x 1 ),(e 2, x 2 ),…,(e r, x r ),…,(e R, x R ))), where e r is the r-th topic tag output by the trained first neural network model, and x r is the probability of the r-th topic tag output by the trained first neural network model. The value range of r is from 1 to R, and R is the number of topic tags output by the trained first neural network model; the trained first neural network model has the function of inferring the topic tags corresponding to the input text and the probabilities of the topic tags according to the input text.

[0022] Those skilled in the art know that any training method of a neural network model in the prior art falls within the protection scope of the present invention; optionally, obtain a set of sample generated contents, which includes a large number of sample generated contents, and any sample generated content is an article; use the text in each sample generated content as the training sample of the first neural network model, and use the set of topics of each sample generated content as the label corresponding to the training sample to train the first neural network model to obtain the trained first neural network model. Optionally, the first neural network model is a convolutional neural network. The first neural network model includes a text representation layer for converting text into a word vector matrix representation, a convolutional layer for extracting features, a pooling layer for pooling the output of the convolutional layer, a fully connected layer for mapping the features extracted by the pooling layer to the topic tag space, and an output layer for converting the output of the fully connected layer into the probability distribution of each topic tag.

[0023] S120, input the image in the content to be determined for generation of the target user into the trained second neural network model, and obtain the output result F of the trained second neural network model, F = ((f 1, y 1 ),(f 2, y 2 ),…,(f q,y q ),…,(f Q, y Q )),f q is the q-th topic label output by the trained second neural network model, and y q is the probability of the q-th topic label output by the trained second neural network model. The value range of q is from 1 to Q, and Q is the number of topic labels output by the trained second neural network model; the trained second neural network model has the function of inferring the topic label corresponding to the input image and the probability of the topic label according to the input image.

[0024] Those skilled in the art know that the training methods of any neural network models in the prior art fall within the protection scope of the present invention; optionally, a sample generation content set is obtained, and the sample generation content set includes a large number of sample generation contents, and any sample generation content is an article; the image set in each sample generation content is used as the training sample of the second neural network model, and the topic set of each sample generation content is used as the label of the corresponding training sample, and the second neural network model is trained to obtain the trained second neural network model. Optionally, the second neural network model is a multi-input convolutional neural network, which includes an image input layer with multiple input branches, a convolutional and pooling layer for extracting each image feature and downsampling the feature map, a feature fusion layer for fusing the features extracted by each input branch, a fully connected layer, and an output layer for outputting the probabilities of each topic label.

[0025] S130. Obtain the initial label set G of the to-be-determined generated content of the target user according to E and F; G = {g 1 , g 2 , …, g p , …, g z}, where g p is the p-th topic label appearing in E and F, and the value range of p is from 1 to z, and z is the number of topic labels appearing in E and F; wherein, when g p appears in both E and F at the same time, the probability of g p is w 1 ×h 1,p +(1 - w 1 )×h 2,p , and h 1,p and h 2,p are the probabilities of g p output by the trained first neural network model and the trained second neural network model respectively, and w 1 is the preset text weight; when g p appears in E but does not appear in F, the probability of g p is h 1,p ; when gp When it appears in F and does not appear in E, g p has a probability of h 2,p .

[0026] In this embodiment, during the training of the first neural network model and the second neural network model, the expressions of the topic tags corresponding to the same topic are consistent. Therefore, if a certain topic tag output by the first neural network model is the same as a certain topic tag output by the second neural network model, then the probability that this topic tag is the topic tag of the content to be determined and generated by the target user is greater.

[0027] In this embodiment, w 1 is an empirical value, w 1 > 0.5. Optionally, w 1 = 0.7 or 0.8.

[0028] S140. Construct B according to the top n topic tags with the largest probability values in G.

[0029] In this embodiment, B is composed of the top n topic tags with the largest probability values in G; optionally, n is an empirical value. For example, n is 3.

[0030] Based on S110 - S140, this embodiment can accurately and comprehensively obtain the set of topic tags of the content to be determined and generated by the target user.

[0031] S200. Obtain the set of application programs APP used by the target user within the historical time period; APP = {app 1 , app 2 , …, app j , …, app m}, where app j is the jth application program used by the target user within the historical time period, and the value range of j is from 1 to m, where m is the number of application programs used by the target user within the historical time period.

[0032] In this embodiment, the duration corresponding to the historical time period is an empirical value. Optionally, the historical time period is the 1 - month time period closest to the current moment.

[0033] S300. Screen APP according to B to obtain the screened set of application programs APP'; APP' = {app' 1 , app' 2 , …, app' k , …, app' u}, where app' kIt is the k-th application related to B used by the target user within the historical time period, where the value range of k is from 1 to u, and u is the number of applications related to B used by the target user within the historical time period.

[0034] As a preferred specific implementation manner, S300 includes:

[0035] S310, traverse the APPs. If the description text of an app j is less than or equal to the preset relevance value threshold for any b i in B, then delete the app j in the APP; otherwise, retain the app j in the APP.

[0036] Optionally, the preset relevance value threshold is an empirical value, and the preset relevance value threshold is greater than 0 and less than 1.

[0037] As a preferred specific implementation manner, S310 includes:

[0038] S311, obtain the description text of the app j and extract the keywords of the description text of the app j

[0039] In this embodiment, the description text of the app j includes the name and core function of the app j and the core function can characterize the type of the app j , such as social, shopping, travel, food, medical, learning, etc.

[0040] Those skilled in the art know that any keyword extraction method in the prior art falls within the protection scope of the present invention.

[0041] S312, input each keyword of the description text of the app j into the trained word vector conversion model to obtain the vector corresponding to each keyword of the description text of the app j ; the trained word vector conversion model has the function of converting the input word into a vector representation.

[0042] Those skilled in the art know that any word vector conversion model in the prior art falls within the protection scope of the present invention. Optionally, the word vector conversion model is word2vec.

[0043] S313, input b i into the trained word vector conversion model to obtain the vector corresponding to b i .

[0044] ​S314, Obtain the app j The similarity between the vector corresponding to each keyword in the description text of the app and the vector corresponding to b i is determined, and the maximum similarity is determined as the correlation value between the description text of the app j and b i .

[0045] In this embodiment, if the correlation value between the description text of the app j and b i is less than or equal to the preset correlation value threshold, it is determined that the app j has no correlation with b i ; otherwise, it is determined that the app j has a correlation with b i .

[0046] In this embodiment, if the app j has no correlation with any b in B i , it is determined that the app j has no correlation with B; otherwise, it is determined that the app j has a correlation with B

[0047] S320, Determine the updated APP as APP'.

[0048] It should be understood that after executing S310, the updated APP only includes the application programs that have a correlation with B

[0049] Based on S310 - S320, APP' can be accurately obtained. The portrait tags in this embodiment are not obtained based on all apps used by the target user in the historical time period, but first, all the above apps are screened according to the theme tags of the content to be determined generated by the target user, and only the apps that have a correlation with the theme tags of the content to be determined generated by the target user are retained; since multiple portrait tags of the target user may be obtained according to the behavior of the target user using one app, therefore, the method of screening the app first and then obtaining the portrait tags of the target user in this embodiment can greatly reduce the number of subsequent obtained target portrait tags on the premise of ensuring the accuracy of the final judgment result, and greatly reduce the workload of subsequently obtaining the similarity between the target user and the content to be determined generated by the target user, thereby improving the judgment efficiency of the authenticity of the content to be determined generated by the target user as a whole

[0050] S400, Obtain the portrait tag set D of the target user according to APP'; D = {d 1 , d 2 , …, d t , …, d T}, d tIt is the t-th portrait label of the target user, where the value range of t is from 1 to T, and T is the number of portrait labels of the target user.

[0051] Those skilled in the art know that any method in the prior art for obtaining the portrait labels of a user based on the behavior of the application programs used by the user falls within the protection scope of the present invention. As a preferred specific embodiment, S400 includes:

[0052] S410, traverse APP’, and according to app’ k Obtain the k-th portrait label subset of the target user.

[0053] As a preferred specific embodiment, S410 includes:

[0054] S411, obtain the usage frequency, usage duration, and category of the content browsed by the target user using app’ k of.

[0055] S412, obtain the k-th portrait label subset of the target user according to the usage frequency, usage duration, and category of the content browsed by the target user using app’ k of.

[0056] In this embodiment, the category of the content frequently and for a long time browsed by the target user is used as the basis for obtaining the portrait labels of the target user. For example, if the target user uses an information app with a high frequency and for a long time, and the main category of the content browsed by the target user in the information app is technology news, then "technology information follower" can be used as the portrait label of the target user. Another example, if the target user uses a reading app with a high frequency and for a long time, and the target user often reads articles on food making in the reading app, then "food preference person" can be used as the portrait label of the target user.

[0057] Based on S411 - S412, partial portrait labels of the target user can be obtained based on app’ k of.

[0058] S420, perform a union operation on all portrait label subsets of the target user to obtain D.

[0059] Based on S410 - S420, the portrait label set of the target user can be accurately obtained.

[0060] S500, obtain the similarity sim between the target user and the content to be determined and generated by the target user according to B and D.

[0061] In this embodiment, 0 < sim < 1. As a preferred specific embodiment, S500 includes:

[0062] S510, traverse B, and use bi Input the trained word vector conversion model to obtain b i The corresponding vector Vb i ; The trained word vector conversion model has the function of converting the input word into a vector representation.

[0063] S520, Traverse D, and input d t Input the trained word vector conversion model to obtain d t The corresponding vector Vd t .

[0064] S530, Obtain the similarity sim i between each Vb t and each Vd i,t .

[0065] It should be understood that the value range of i is from 1 to n, and the value range of t is from 1 to T. Then, n×T similarities can be obtained.

[0066] Those skilled in the art know that any method for obtaining the similarity between vectors in the prior art falls within the protection scope of the present invention. Optionally, the cosine similarity between Vb i and Vd t is determined as sim i,t .

[0067] S540, Determine the maximum sim i,t as sim.

[0068] Based on S510 - S540, the similarity between the target user and the to-be-determined generated content of the target user can be obtained. This similarity can represent the matching degree between the target user and the to-be-determined generated content of the target user. The greater the similarity, the higher the matching degree.

[0069] S600, If sim≥sim 0 , then determine that the to-be-determined generated content of the target user is the content generated by the target user; otherwise, determine that the to-be-determined generated content of the target user is not the content generated by the target user; sim 0 is the preset similarity threshold.

[0070] Optionally, the preset similarity threshold is an empirical value.

[0071] In this embodiment, the set of topic tags of the target user's content to be determined is first obtained, and each topic tag in this set is a possible topic tag of the target user's content to be determined. Then, according to the application programs used by the target user in the historical time period, the set of portrait tags of the target user is obtained. By calculating the similarity between the topic tags in the set of topic tags and the portrait tags in the set of portrait tags of the target user, the similarity between the target user and the target user's content to be determined is obtained. If the similarity is high, it is determined that the target user's content to be determined is the content generated by the target user, that is, the content personally created by the target user based on the knowledge he / she has mastered. Otherwise, it is determined that the target user's content to be determined is not the content generated by the target user, that is, the content automatically generated by means of AI or the like. Thus, in this embodiment, the authenticity of the target user's content to be determined is judged based on the topic tags of the target user's content to be determined and the portrait tags of the target user.

[0072] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.

Claims

1. A method for determining user-generated content based on user historical behavior, characterized in that: The method comprises the following steps: S100, obtaining a subject tag set B of the target user's generated content to be determined; B = {b1, b2, ..., b i ,…,b n }, b i The i-th topic label of the target user's to-be-determined generated content, where i ranges from 1 to n, and n is the number of topic labels of the target user's to-be-determined generated content; S200, obtaining the application set APP used by the target user in the historical time period; APP = {app1, app2, ..., app j ,…,app m }, app j is the jth application used by the target user in the historical time period, where j ranges from 1 to m, and m is the number of applications used by the target user in the historical time period; S300, filter the APP according to B to obtain a filtered application set APP'; APP' = {app'1, app'2, ..., app' k ,…,app' u }, app' k is the kth application related to B used by the target user in the historical time period, the value range of k is 1 to u, and u is the number of applications related to B used by the target user in the historical time period; S400, obtaining a target user's portrait tag set D according to APP'; D = {d1, d2, ..., d t ,…,d T }, d t is the t-th portrait label of the target user, the value range of t is 1 to T, and T is the number of portrait labels of the target user; S500, obtaining the similarity sim between the target user and the target user's generated content to be determined according to B and D; S600, if sim≥sim0, determine that the target user's generated content to be determined is the content generated by the target user; otherwise, determine that the target user's generated content to be determined is not the content generated by the target user; sim0 is a preset similarity threshold.

2. The method for determining user-generated content based on user historical behavior according to claim 1, characterized in that: S100 includes: S110, inputting the text of the target user's generated content to be determined into the trained first neural network model, and obtaining the output result E of the trained first neural network model, where E=((e 1, x1),(e 2, x2),…,(e r, x r ),…,(e R, x R )), e r is the rth topic label output by the first trained neural network model, x r is the probability of the rth topic label output by the trained first neural network model, where r ranges from 1 to R, and R is the number of topic labels output by the trained first neural network model; the trained first neural network model has the function of inferring the topic labels corresponding to the input text and the probability of the topic labels based on the input text; S120, inputting the image of the target user's to-be-determined generated content into the trained second neural network model, and obtaining the output result F of the trained second neural network model, where F=((f 1, y1),(f 2, y2),…,(f q, y q ),…,(f Q, y Q )),f q is the qth topic label output by the trained second neural network model, y q is the probability of the qth topic label output by the trained second neural network model, where the value of q ranges from 1 to Q, and Q is the number of topic labels output by the trained second neural network model; the trained second neural network model has the function of inferring the topic labels corresponding to the input image and the probability of the topic labels according to the input image; S130, obtaining the initial tag set G of the target user's to-be-determined generated content according to E and F; G={g1,g2,…,g p ,…,g z }, g p is the pth topic label that appears in E and F, the value of p ranges from 1 to z, and z is the number of topic labels that appear in E and F; p When it appears in both E and F, g p The probability is w1×h 1,p +(1-w1)×h 2,p ,h 1,p and h 2,p are the g output by the trained first neural network model and the trained second neural network model respectively. p The probability of w1 is the preset text weight; when g p When it appears in E and not in F, g p The probability is h 1,p When g p When it appears in F and not in E, g p The probability is h 2,p ; S140, construct B according to the first n topic labels with the largest probability values ​​in G.

3. The method for determining user-generated content based on user historical behavior according to claim 1, characterized in that: S500 includes: S510, traverse B, and b i Input the trained word vector conversion model to obtain b i The corresponding vector Vb i ; The trained word vector conversion model has the function of converting input words into vector representations; S520, traverse D, and convert d t Input the trained word vector conversion model to obtain d t The corresponding vector Vd t ; S530, obtain each Vb i With each Vd t Similarity sim i,t ; S540, the largest sim i,t Confirmed as sim.

4. The method for determining user-generated content based on user historical behavior according to claim 1, characterized in that: S300 includes: S310, traverse APP, if app j The description text of B is the same as any b in B i If the correlation values ​​of are less than or equal to the preset correlation value threshold, the app in the APP j Delete; otherwise, the app in the APP j reserve; S320, determining the updated APP as APP'.

5. The method for determining user-generated content based on user historical behavior according to claim 4, characterized in that: S310 includes: S311, get the app j Description text of the app and extract it j Keywords of the description text; S312, app j Each keyword of the description text is input into the trained word vector conversion model to obtain app j The trained word vector conversion model has the function of converting the input words into vector representations; S313, b i Input the trained word vector conversion model to obtain b i The corresponding vector; S314, get app j The vector corresponding to each keyword in the description text is i The similarity of the corresponding vectors, and the largest similarity is determined as app j Description text and b i The relevant value of .

6. The method for determining user-generated content based on user historical behavior according to claim 1, characterized in that: S400 includes: S410, traverse APP', according to app' k Get the kth portrait tag subset of the target user; S420, performing a union process on all subsets of the target user's portrait labels to obtain D.

7. The method for determining user-generated content based on user historical behavior according to claim 6, characterized in that: S410 includes: S411, obtain the target user's app' k Frequency of use, duration of use, and categories of content viewed; S412, use app' according to target users k The k-th portrait label subset of the target user is obtained based on the usage frequency, usage duration and category of browsed content.