Information processing method and device and electronic equipment
By extracting the fine-grained semantic features of social media text data and combining user characteristics and emotional categories, the problem of low accuracy in false information detection is solved, and more efficient identification of false information is achieved.
Patent Information
- Application Number
- CN202510329781.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has low accuracy in false information detection, which may lead to missed identification or misidentification, and does not consider the association between user emotions and false information.
By obtaining the text data to be identified and its corresponding user characteristics, combining the preset bidirectional long and short-term memory network Bi-LSTM and convolutional neural network, the fine-grained semantic features of the text data are extracted, and the probability value of the text data being false information is calculated based on user characteristics and emotional categories.
The accuracy of false information detection is improved, and by combining user emotions and characteristics, it can effectively curb the spread of false information on social networks.
Smart Images

Figure CN120217157A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular, to an information processing method, apparatus, and electronic device. Background Art
[0002] With the rapid development of social media networks, the behavior of users sharing and reading various posts on social media networks has increased significantly. However, due to the low threshold and fast speed of spreading content, and the lack of effective supervision measures, there are many false information on social media networks, which has a serious impact on users.
[0003] In related technologies, false information detection methods can be implemented based on deep learning technologies. These methods use different features related to context content, network behavior, and user patterns to distinguish false information. However, these methods do not consider the association between user emotions and false information. In actual scenarios, there is a close connection between false information and user emotions. False information can cause emotional reactions of users, which in turn lead to specific behaviors, such as liking and forwarding false information.
[0004] Currently, in the field of false information detection, if only deep learning technology is used to extract features of false information and then detect false information based on the extracted features, the accuracy of detecting false information will be low due to the single extracted features, and there may be cases of missed recognition or misrecognition of false information. Therefore, how to improve the accuracy of false information detection is a technical problem that urgently needs to be solved currently. Summary of the Invention
[0005] Embodiments of the present application provide an information processing method, apparatus, and electronic device, which are used to combine user emotion analysis with user features to detect false information, thereby improving the accuracy of false information detection and effectively curbing the spread of false information on social networks.
[0006] In a first aspect, an embodiment of the present application provides an information processing method, and the method includes:
[0007] Obtain text data to be recognized, and obtain user features of the publishing account corresponding to the text data; wherein, the user features at least include the influence value of the publishing account and the registration years;
[0008] Obtain a first feature corresponding to the text data by using a preset bidirectional long short-term memory network Bi-LSTM, and extract a second feature corresponding to the first feature by using a preset convolutional neural network;
[0009] Obtain an emotion category corresponding to the text data based on the second feature;
[0010] Obtain the probability value that the text data is false information based on user characteristics and emotion categories.
[0011] Through the above method, based on the Bi-LSTM network and the preset convolutional neural network, fine-grained semantic features of the text data to be recognized can be extracted, so as to better realize the prediction of the user's emotion category, and then apply the emotion category to the recognition of false information, thereby improving the accuracy of false information recognition.
[0012] In an alternative embodiment, obtaining the user characteristics of the publishing account corresponding to the text data includes:
[0013] Obtain the registration time and the current time of the publishing account, and obtain the number of the first people subscribing to the publishing account and the number of the second people subscribed by the publishing account;
[0014] Based on the registration time and the current time, calculate the registration years, and based on the number of the first people and the number of the second people, calculate the influence value.
[0015] Through the above method, the user characteristics corresponding to the account publishing the text data to be recognized can be accurately obtained, so as to use the user characteristics in the detection of false information.
[0016] In an alternative embodiment, obtaining the emotion category corresponding to the text data based on the second feature includes:
[0017] Input the second feature into the target classifier to obtain the emotion category corresponding to the text data; wherein, the target classifier is constructed based on the softmax classification model.
[0018] Through the above method, the emotion category prediction can be performed on the text data to be recognized, and then the emotion category is applied to the recognition of false information.
[0019] In an alternative embodiment, obtaining the probability value that the text data is false information based on user characteristics and emotion categories includes:
[0020] Use the preset membership function to calculate the membership degrees corresponding to the user characteristics and the emotion categories respectively; wherein, the membership degree represents the degree to which the user characteristics and the emotion categories belong to their respective fuzzy sets;
[0021] Based on the membership degrees corresponding to the user characteristics and the emotion categories respectively, determine the mapping levels corresponding to the user characteristics and the emotion categories respectively;
[0022] Based on the mapping levels corresponding to the user characteristics and the emotion categories respectively, match the probability value that the text data is false information according to the preset rules.
[0023] Through the above method, a fuzzy inference mechanism is adopted to detect false information for text data. By combining user emotions with user characteristics, false information can be effectively detected, thereby improving the accuracy of false information detection.
[0024] In an alternative embodiment, the preset convolutional neural network includes N convolutional neural network groups; where each convolutional neural network group includes: a one-dimensional convolutional layer, a hyperbolic linear unit activation function, and an average pooling layer, and N is a positive integer greater than or equal to 1.
[0025] Through the above method, by adopting a structure that combines a hybrid one-dimensional convolutional layer, a hyperbolic linear unit activation function, and an average pooling layer, the convolutional process can be carried out more smoothly, reducing feature conflicts and extracting fine-grained semantic features corresponding to the text data to be recognized.
[0026] In a second aspect, the present application provides an information processing device, and the device includes:
[0027] An acquisition module, configured to acquire the text data to be recognized and acquire the user characteristics of the publishing account corresponding to the text data; where the user characteristics at least include the influence value of the publishing account and the registration years.
[0028] A processing module, configured to obtain a first feature corresponding to the text data by using a preset bidirectional long short-term memory network Bi-LSTM, and extract a second feature corresponding to the first feature by using a preset convolutional neural network.
[0029] A classification module, configured to obtain the emotion category corresponding to the text data based on the second feature.
[0030] An identification module, configured to obtain the probability value that the text data is false information based on the user characteristics and the emotion category.
[0031] In an alternative embodiment, when acquiring the user characteristics of the publishing account corresponding to the text data, the acquisition module is specifically configured to:
[0032] Acquire the registration time and the current time of the publishing account, and acquire the first number of people subscribing to the publishing account and the second number of people subscribed to by the publishing account.
[0033] Based on the registration time and the current time, calculate the registration years, and based on the first number of people and the second number of people, calculate the influence value.
[0034] In an alternative embodiment, when obtaining the emotion category corresponding to the text data based on the second feature, the classification module is specifically configured to:
[0035] Input the second feature into the target classifier to obtain the emotion category corresponding to the text data; wherein, the target classifier is constructed based on the softmax classification model.
[0036] In an alternative embodiment, when obtaining the probability value that the text data is false information based on the user feature and the emotion category, the recognition module is specifically configured to:
[0037] Use a preset membership function to calculate the membership degrees corresponding to the user feature and the emotion category respectively; wherein, the membership degree represents the degree to which the user feature and the emotion category belong to their respective fuzzy sets.
[0038] Based on the membership degrees corresponding to the user feature and the emotion category respectively, determine the mapping levels corresponding to the user feature and the emotion category respectively.
[0039] Based on the mapping levels corresponding to the user feature and the emotion category, match the probability value that the text data is false information according to a preset rule.
[0040] In an alternative embodiment, the preset convolutional neural network includes N convolutional neural network groups; wherein, each convolutional neural network group includes: a one-dimensional convolutional layer, a hyperbolic linear unit activation function, and an average pooling layer, and N is a positive integer greater than or equal to 1.
[0041] In a third aspect, the present application provides an electronic device, which includes a processor and a memory. Wherein, the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the information processing method described in the first aspect above.
[0042] In a fourth aspect, the present application provides a computer-readable storage medium, which includes program code, and when the program code runs on an electronic device, the program code is used to cause the electronic device to execute the steps of the information processing method described in the first aspect above.
[0043] In a fifth aspect, the present application provides a computer program product, and when the computer program product is called by a computer, the computer is caused to execute the steps of the information processing method described in the first aspect.
[0044] In addition, other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0046] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0047] Figure 2 It is a schematic diagram of the implementation process of an information processing method provided by an embodiment of the present application;
[0048] Figure 3 It is a schematic diagram of the emotion analysis of true and false information provided by an embodiment of the present application;
[0049] Figure 4 It is a schematic diagram of the structure of an information processing device provided by an embodiment of the present application;
[0050] Figure 5 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the technical solutions of the present application. Based on the embodiments described in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the technical solutions of the present application.
[0052] It should be noted that in the description of the present application, "a plurality of" is understood as "at least two". "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The connection between A and B may represent: A is directly connected to B and A is connected to B through C. In addition, in the description of the present application, words such as "first" and "second" are only used for the purpose of distinguishing descriptions and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.
[0053] In addition, in the technical solutions of the present application, the collection, dissemination, use, etc. of data all comply with the requirements of relevant national laws and regulations.
[0054] The following briefly introduces the design concept of the embodiments of the present application:
[0055] With the rapid development of social media networks, the behavior of users sharing and reading various posts on social media networks is increasing day by day. However, due to the low threshold and fast speed of spreading content, and the lack of effective supervision measures, there are many false information on social media networks, which has a serious impact on users.
[0056] In related technologies, the detection methods of false information can be implemented based on deep learning technologies. These methods use different features related to context content, network behavior, and user models to distinguish false information. However, these methods do not consider the association between user emotions and false information. In actual scenarios, there is a close connection between false information and user emotions. False information can cause emotional reactions of users, which in turn lead to specific behaviors, such as liking and forwarding false information.
[0057] Currently, in the field of false information detection, if only deep learning technology is used to extract features of false information and then detect false information based on the extracted features, due to the single extracted features, the accuracy of detecting false information is relatively low, and there may be situations of missed recognition or misrecognition of false information.
[0058] In view of this, the embodiments of the present application provide an information processing method, which includes: First, obtain the text data to be recognized, and obtain the user features of the publishing account corresponding to the text data; Then, use a preset bidirectional long short-term memory network Bi-LSTM to obtain the first feature corresponding to the text data, and use a preset convolutional neural network to extract the second feature corresponding to the first feature; Further, obtain the emotion category corresponding to the text data based on the second feature; Finally, obtain the probability value that the text data is false information based on the user features and the emotion category. Through the above method, when identifying false information, the real emotions of users can be combined, and when extracting features, deep features of the text data to be recognized can be extracted, enhancing the understanding of deep semantics, thereby improving the accuracy of false information recognition.
[0059] Next, the information processing method provided by the exemplary embodiment of the present application will be described with reference to the accompanying drawings.
[0060] Refer to Figure 1 As shown, it is a schematic diagram of a possible application scenario in the embodiments of the present application. In this application scenario diagram, it includes a server 11 and terminal devices 12 (including terminal devices 121, 122... 12n).
[0061] The server 11 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device 12 and the server 11 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.
[0062] The terminal device 12 includes, but is not limited to, devices such as mobile phones, tablet computers, laptop computers, desktop computers, e-readers, intelligent voice interaction devices, smart home appliances, in-vehicle terminals, etc.; various software can be installed on the terminal device, such as application programs and applets, etc.
[0063] It should be noted that Figure 1 The above are only examples. In fact, the number of terminal devices 12 and servers 11 is not limited and is not specifically defined in the embodiments of this application.
[0064] Exemplarily, the terminal device 12 can access the network through cellular mobile communication technology, thereby communicating with the server 11, and then sending the text data to be recognized to the server. Among them, the flavor mobile communication technology, for example, includes the 5th Generation Mobile Networks (5G) technology.
[0065] Optionally, the terminal 12 can access the network through short-range wireless communication methods, thereby communicating with the server 11, and then sending the text data to be recognized to the server. Among them, the short-range wireless communication methods, for example, include Wireless Fidelity (Wi-Fi) technology.
[0066] Referring to Figure 2 shown, which is a schematic diagram of the implementation process of an information processing method provided by an embodiment of this application. The specific implementation process of this method is as follows:
[0067] S1: Obtain the text data to be recognized and obtain the user characteristics of the publishing account corresponding to the text data.
[0068] In actual application scenarios, users can generate their own opinions or views on the information browsed on social media, comment on their views on the browsed information, or users can also directly post their own remarks on social media.
[0069] Exemplarily, a user can browse information published by other users on different social media platforms, and then post their own comments or opinions on the browsed information, or directly post their own remarks on the social media. For example, post remarks about the dining experience for a restaurant, etc.
[0070] However, due to the openness of social media, the remarks posted by users may be false information and lack authenticity. Therefore, in the embodiments of the present application, text data to be recognized can be obtained from social media, that is, the remarks posted by users are obtained, and then the obtained text data to be recognized is recognized to determine whether the text data to be recognized is false information.
[0071] In the embodiments of the present application, user characteristics of the publishing account corresponding to the text data to be recognized also need to be obtained. Specifically, in the embodiments of the present application, the user characteristics corresponding to the publishing account at least include: the influence value corresponding to the publishing account and the registration years.
[0072] The publishing account is the account used by a user to post remarks or information on social media. For example, user A uses account A registered with their real-name information to post remarks or information. Accordingly, it can be understood that the registration years are the total time from the registration of the publishing account to the current time, and the influence value represents the degree of influence of an account on other users. According to the size of the influence value, the publishing accounts corresponding to the text data to be recognized can be classified into ordinary accounts, influential accounts, highly influential accounts, accounts of well-known personalities / celebrities, or accounts of well-known institutions, etc.
[0073] In an alternative embodiment, the following steps can be adopted to obtain the user characteristics of the publishing account corresponding to the text data to be recognized, that is, to obtain the influence value and the registration years of the publishing account:
[0074] First, the registration time and the current time of the publishing account can be obtained. Specifically, when obtaining the registration time and the current time of the publishing account, the registration time corresponding to the publishing account can be directly extracted from the account information corresponding to the publishing account, and the current time can be directly obtained. Then, based on the registration time and the current time of the publishing account, the registration years of the publishing account are calculated.
[0075] In the embodiments of the present application, the registration years of the publishing account can be calculated by the following formula:
[0076] A age =A cur -A reg ;
[0077] Wherein, A age represents the registration years of the publishing account, A cur represents the current time, Areg Indicates the registration time of the publishing account.
[0078] Further, when obtaining the influence value corresponding to the publishing account, first, the first number of people subscribing to the publishing account and the second number of people subscribed to by the publishing account can be obtained.
[0079] It can be understood that in the actual application scenario, subscribing to a publishing account usually means that on social media, other users use their own accounts to actively choose to receive content updates or information pushes from the publishing account. Through the act of "subscribing" or "subscribing", it expresses the interest or expectation of other users in the content of the publishing account, and the latest information can be obtained at the first time when the publishing account publishes a message. Similarly, the publishing account in the embodiment of the present application can also "subscribe" to other accounts to obtain the latest news corresponding to other accounts.
[0080] Based on this, in the embodiment of the present application, the first number of people subscribing to the publishing account and the second number of people subscribed to by the publishing account can be obtained from the account information corresponding to the publishing account.
[0081] Then, based on the first number and the second number, the influence value corresponding to the publishing account is calculated. That is: the ratio of the first number to the second number is used as the influence value corresponding to the publishing account.
[0082] Specifically, in the embodiment of the present application, the influence value of the publishing account can be calculated by the following formula:
[0083]
[0084] where, w inf represents the influence value of the publishing account, A followers represents the first number of people subscribing to the publishing account, A following represents the second number of people subscribed to by the publishing account.
[0085] Through the above method, the registration age and the influence value corresponding to the publishing account can be calculated, that is, the user characteristics corresponding to the publishing account are obtained.
[0086] S2: Obtain the first feature corresponding to the text data by using a preset bidirectional long short-term memory network Bi-LSTM, and extract the second feature corresponding to the first feature by using a preset convolutional neural network.
[0087] In the embodiment of the present application, in order to better capture the long-term dependence relationship in the text data to be recognized, a preset bidirectional long short-term memory network (Bidirectional Long Short Term Memory, Bi-LSTM) can be used to extract the first feature corresponding to the text data to be recognized.
[0088] Specifically, the Bi-LSTM network consists of a forward long short-term memory network and a backward long short-term memory network. Therefore, when the text data to be recognized is input into the preset Bi-LSTM network, the forward hidden state corresponding to the text data to be recognized can be obtained through the forward long short-term memory network of the Bi-LSTM, and the backward hidden state corresponding to the text data to be recognized can be obtained through the backward long short-term memory network of the Bi-LSTM network. Then, the obtained forward hidden state and backward hidden state are concatenated to obtain the first feature in the embodiments of the present application. The first feature can be expressed as: where H t represents the first feature, represents the forward hidden state output by the forward long short-term memory network, represents the backward hidden state output by the backward long short-term memory network.
[0089] Since the Bi-LSTM network has memory units and gating mechanisms inside, it can learn the long-term dependencies in the text data and can better understand the context of the text data.
[0090] Furthermore, it is not enough to only extract the utterance representation corresponding to the text data to be recognized through the Bi-LSTM network. Therefore, in the embodiments of the present application, a preset convolutional neural network is also provided, which is used to perform fine-grained feature extraction based on the first feature extracted by the Bi-LSTM network, that is, to obtain the second feature, so as to better obtain the complete information of the text data to be recognized.
[0091] In an alternative embodiment, in the embodiments of the present application, the preset convolutional neural network includes N convolutional neural network groups.
[0092] It should be noted that in the embodiments of the present application, N is a positive integer greater than or equal to 1. For the number of convolutional neural network groups, the embodiments of the present application do not make specific limitations, and can be selected according to the actual application scenario. Preferably, in the embodiments of the present application, the number of convolutional neural network groups can be set to 3. Wherein, each convolutional neural network group includes: a one-dimensional convolutional layer, a hyperbolic linear unit activation function, and an average pooling layer.
[0093] In the embodiments of the present application, the first feature extracted by the Bi-LSTM network is input into the preset convolutional neural network. The one-dimensional convolutional layer in the preset convolutional neural network can process the first feature and extract the corresponding features. The one-dimensional convolution has a more stable effect compared with the high-dimensional convolution and will not cause too much perturbation to the first feature extracted by the Bi-LSTM network.
[0094] The output after the one-dimensional convolutional layer extracts features from the first feature can be expressed as: H c = Conv1D(H t ).
[0095] Furthermore, using an average pooling layer to process the output of the one-dimensional convolutional layer can perform dimensionality reduction on the output of the one-dimensional convolutional layer through pooling operations and reduce noise interference, effectively retaining the feature information of the data.
[0096] The output after the average pooling layer processes the output of the one-dimensional convolutional layer can be expressed as: H p = aveage - pooling(H c ).
[0097] Furthermore, using a hyperbolic linear unit activation function to process the output of the average pooling layer, the average output value of the hyperbolic linear unit activation function approaches 0, which can reduce the deviation between the natural gradient and the standard gradient, make the convolutional process smoother, and enhance the non-linear expression ability of the convolutional neural network.
[0098] The output after the hyperbolic linear unit activation function processes the output of the average pooling layer can be expressed as: H f = HLU(H p ).
[0099] Inputting the first feature output by the Bi-LSTM network into the preset convolutional neural network in the embodiment of the present application, and processing it successively through a one-dimensional convolutional layer, an average pooling layer, and a hyperbolic linear unit activation function, the second feature in the embodiment of the present application can be obtained, that is, the fine-grained feature that can represent the complete information of the text data to be recognized. Moreover, stacking multiple convolutional neural network groups helps prevent overfitting and enhances the generalization ability.
[0100] S3: Obtain the emotion category corresponding to the text data based on the second feature.
[0101] In the embodiment of the present application, after obtaining the second feature corresponding to the text data to be recognized through the above steps, the emotion category of the text data to be recognized can be identified based on the second feature.
[0102] In an alternative embodiment, input the second feature into a target classifier to implement emotion category classification of the text data to be recognized, thereby obtaining the emotion category corresponding to the text data to be recognized.
[0103] It should be noted that the target classifier in the embodiments of the present application is constructed based on a softmax classification model. When training the target classifier, first, various text data are collected from social media to form a data set, and corresponding category labels are assigned to each sample in the data set. For example, the label corresponding to sample 1 is anger, the label corresponding to sample 2 is happiness, the label corresponding to sample 3 is sadness, and so on. Then, the data set is divided into a training set, a validation set, and a test set. The training set is input into a pre-constructed classifier, the loss function value is calculated through forward propagation, the gradient value is calculated through backward propagation, and the parameters of the pre-constructed classifier are updated through a suitable optimization algorithm, such as the stochastic gradient descent algorithm, to minimize the loss function value. During the training process, the validation set is used to evaluate the performance of the classifier, and the corresponding training parameters or model structure are adjusted according to actual needs. Finally, the trained classifier is tested using the test set, and the target classifier is obtained when the evaluation metrics (such as accuracy, F1 value, etc.) corresponding to the classifier meet the requirements.
[0104] Furthermore, the target classifier can be applied to obtain the emotion category corresponding to the text data to be recognized based on the second feature. In the embodiments of the present application, predicting the emotion category of the text data to be recognized can be expressed as: Y i = softmax(Wh i + b).
[0105] Wherein, Y i represents the weight corresponding to the emotion category, softmax represents the activation function, W represents the weight matrix, h i represents the second feature, and b represents the bias term.
[0106] Exemplarily, the result output by the target classifier is that the weight corresponding to happiness is 0.8, the weight corresponding to anger is 0.1, and the weight corresponding to sadness is 0.1. Then, it can be determined that the emotion category with the highest final corresponding weight value is the target emotion category, that is, the obtained emotion category is happiness.
[0107] S4: Obtain the probability value that the text data is false information based on the user feature and the emotion category.
[0108] In the embodiments of the present application, based on the user feature and the emotion category obtained in the above steps, it is possible to determine whether the text data to be recognized is false information.
[0109] Specifically, referring to Figure 3 as shown, compared with true information, the emotional intensity brought by false information is usually higher. This indicates that there is a significant correlation between emotion and false information. Applying the emotion category to the detection of false information can further improve the detection accuracy.
[0110] In an alternative embodiment, when identifying false information based on user characteristics and emotion categories, a preset membership function can be used to calculate the membership degrees corresponding to the user characteristics and emotion categories respectively. That is, calculate the membership degrees corresponding to the influence value of the publishing account, the registration duration, and the emotion category respectively.
[0111] It should be noted that the membership degree in the embodiments of the present application represents the degree to which the user characteristics and emotion category belong to their respective fuzzy sets. Moreover, the embodiments of the present application do not specifically limit the membership function used, and it can be selected according to requirements in actual applications.
[0112] Then, based on the membership degrees corresponding to the user characteristics and emotion categories respectively, determine the mapping levels corresponding to the user characteristics and emotion categories respectively.
[0113] Exemplarily, for example, there is currently a fuzzy set "influence" used to describe the influence value corresponding to the publishing account. To quantify whether a publishing account has influence, a membership function can be defined to describe the relationship between the influence value and the fuzzy probability of "influence". Then, the corresponding mapping rule can be set as follows: when the influence value is greater than or equal to 0 and less than or equal to 0.3, the mapping level corresponding to this influence value is "low"; when the influence value is greater than 0.3 and less than or equal to 0.6, the mapping level corresponding to this influence value is "medium"; when the influence value is greater than 0.6, the mapping level corresponding to this influence value is "high". For example, if the calculated influence value of the publishing account is 0.7, then the mapping level corresponding to this publishing account is "high". Similarly, the mapping levels corresponding to the registration duration and emotion category can also be determined, and the embodiments of the present application will not elaborate on them one by one here.
[0114] Furthermore, based on the mapping levels corresponding to the user characteristics and emotion categories respectively, the probability value that the text data to be identified is false information can be matched according to a preset rule.
[0115] Referring to Table 1 shown, the embodiments of the present application have designed corresponding fuzzy rules, and based on the mapping levels corresponding to the user characteristics and emotion categories respectively, the probability value that the text data to be identified is false information can be determined.
[0116] Table 1
[0117]
[0118]
[0119] As can be seen from the above table, the probability value that the text data to be recognized is false information can be determined according to the mapping levels of the registration years, influence values, and emotion categories respectively. For example, if the mapping level of the registration years is high, the mapping level of the influence value is high, and the mapping level of the emotion category is high, then the probability value that the corresponding text data to be recognized is false information is high; for another example, if the mapping level of the registration years is high, the mapping level of the influence value is medium, and the mapping level of the emotion category is medium, then the probability value that the corresponding text data to be recognized is false information is medium.
[0120] Through the above method, using the Bi-LSTM network and the preset convolutional neural network to extract features from the text data to be recognized, the fine-grained semantic features of the text data to be recognized can be obtained, and the emotion category prediction can be performed more accurately. Further applying the emotion category to the recognition of false information, the recognition ability of false information can be further improved according to the correlation between the emotion category and false information.
[0121] Furthermore, based on the same technical concept, the embodiment of the present application provides an information processing device, and this information processing device is used to implement the above method flow of the embodiment of the present application. Refer to Figure 4 As shown, the device includes: an acquisition module 401, a processing module 402, a classification module 403, and an identification module 404, where
[0122] The acquisition module 401 is used to acquire the text data to be recognized and acquire the user features of the publishing account corresponding to the text data; wherein, the user features at least include the influence value and registration years of the publishing account;
[0123] The processing module 402 is used to obtain the first feature corresponding to the text data by using the preset bidirectional long short-term memory network Bi-LSTM, and extract the second feature corresponding to the first feature by using the preset convolutional neural network;
[0124] The classification module 403 is used to obtain the emotion category corresponding to the text data based on the second feature;
[0125] The identification module 404 is used to obtain the probability value that the text data is false information based on the user features and the emotion category.
[0126] In an alternative embodiment, when acquiring the user features of the publishing account corresponding to the text data, the acquisition module 401 specifically is used to:
[0127] Acquire the registration time of the publishing account and the current time, and acquire the first number of people subscribing to the publishing account and the second number of people subscribed by the publishing account;
[0128] Based on the registration time and the current time, the registration years are calculated, and based on the first number and the second number, the influence value is calculated.
[0129] In an alternative embodiment, when obtaining the emotion category corresponding to the text data based on the second feature, the classification module 403 is specifically configured to:
[0130] Input the second feature into the target classifier to obtain the emotion category corresponding to the text data; wherein, the target classifier is constructed based on the softmax classification model.
[0131] In an alternative embodiment, when obtaining the probability value that the text data is false information based on the user feature and the emotion category, the recognition module 404 is specifically configured to:
[0132] Use a preset membership function to calculate the membership degrees corresponding to the user feature and the emotion category respectively; wherein, the membership degree represents the degree to which the user feature and the emotion category belong to their respective fuzzy sets;
[0133] Based on the membership degrees corresponding to the user feature and the emotion category respectively, determine the mapping levels corresponding to the user feature and the emotion category respectively;
[0134] Based on the mapping levels corresponding to the user feature and the emotion category, match the probability value that the text data is false information according to a preset rule.
[0135] In an alternative embodiment, the preset convolutional neural network includes N convolutional neural network groups; wherein, each convolutional neural network group includes: a one-dimensional convolutional layer, a hyperbolic linear unit activation function, and an average pooling layer, and N is a positive integer greater than or equal to 1.
[0136] Based on the same technical concept, the embodiments of the present application further provide an electronic device, and this electronic device can implement the information processing method flow provided in the above embodiments of the present application. In one embodiment, this electronic device can be a server, or a terminal device or other electronic devices. Refer to Figure 5 As shown, this electronic device may include:
[0137] At least one processor 501, and a memory 502 connected to at least one processor 501. In the embodiments of the present application, the specific connection medium between the processor 501 and the memory 502 is not limited. Figure 5 Here, it is taken as an example that the processor 501 and the memory 502 are connected through a bus 500. The bus 500 is Figure 5 shown by a thick line. The connection manners between other components are only for illustrative purposes and are not to be taken as limitations. The bus 500 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation,Figure 5 It is only represented by a thick line, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 501 may also be referred to as a controller, and there is no limitation on the name.
[0138] In the embodiment of the present application, the memory 502 stores instructions that can be executed by at least one processor 501. By executing the instructions stored in the memory 502, the at least one processor 501 can execute an information processing method described above. The processor 501 can implement Figure 4 the functions of each module in the device shown.
[0139] Among them, the processor 501 is the control center of the device. It can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 502 and calling the data stored in the memory 502, various functions of the device and process data, so as to monitor the device as a whole.
[0140] In a possible design, the processor 501 may include one or more processing units. The processor 501 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 501. In some embodiments, the processor 501 and the memory 502 may be implemented on the same chip. In some embodiments, they may also be implemented separately on independent chips.
[0141] The processor 501 may be a general-purpose processor, such as a CPU, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of an information processing method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0142] The memory 502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 502 can include at least one type of storage medium. For example, it can include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical discs, and so on. The memory 502 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 502 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0143] By programming the design of the processor 501, the code corresponding to the information processing method introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute Figure 1 the steps of an information processing method of the illustrated embodiment. How to program the design of the processor 501 is a well-known technology to those skilled in the art and will not be elaborated here.
[0144] Based on the same inventive concept, the embodiments of the present application also provide a storage medium that stores computer instructions. When the computer instructions run on a computer, the computer is caused to execute an information processing method discussed above.
[0145] In some possible implementation manners, the present application also provides that various aspects of an information processing method can also be implemented in the form of a program product, which includes program code. When the program product runs on a device, the program code is used to cause the control device to execute the steps in an information processing method according to various exemplary embodiments of the present application described above in this specification.
[0146] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0147] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0148] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0149] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a server, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0150] The program code for performing the operations of the present application can be written using any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0151] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for the functions specified in one block or a plurality of blocks.
[0152] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to cover these changes and modifications.
Claims
1. An information processing method, characterized in that: The method comprises Acquire text data to be identified, and acquire user characteristics of a publishing account corresponding to the text data; wherein the user characteristics at least include the influence value and registration years of the publishing account; A preset bidirectional long short-term memory network Bi-LSTM is used to obtain a first feature corresponding to the text data, and a preset convolutional neural network is used to extract a second feature corresponding to the first feature; Obtaining the emotion category corresponding to the text data based on the second feature; Based on the user characteristics and the emotion category, a probability value that the text data is false information is obtained.
2. The method according to claim 1, characterized in that The obtaining of user characteristics of a publishing account corresponding to the text data includes: Acquire the registration time and current time of the publishing account, and acquire the first number of people subscribed to the publishing account and the second number of people subscribed to the publishing account; The registration years are calculated based on the registration time and the current time, and the influence value is calculated based on the first number of people and the second number of people.
3. The method according to claim 1, characterized in that The obtaining the emotion category corresponding to the text data based on the second feature includes: The second feature is input into a target classifier to obtain the emotion category corresponding to the text data; wherein the target classifier is constructed based on a softmax classification model.
4. The method according to claim 1, characterized in that The obtaining, based on the user characteristics and the emotion category, a probability value of the text data being false information includes: The preset membership function is used to calculate the membership degree corresponding to the user feature and the emotion category respectively; wherein the membership degree represents the degree to which the user feature and the emotion category belong to the corresponding fuzzy set respectively; Based on the membership degrees corresponding to the user features and the emotion categories, respectively, determining the mapping levels corresponding to the user features and the emotion categories; Based on the mapping levels corresponding to the user features and the emotion categories, respectively, a probability value of the text data being false information is matched according to a preset rule.
5. The method according to claim 1, characterized in that The preset convolutional neural network includes N convolutional neural network groups; wherein each convolutional neural network group includes: a one-dimensional convolution layer, a hyperbolic unit activation function and an average pooling layer, and N is a positive integer greater than or equal to 1.
6. An information processing device, characterized in that: The device comprises: An acquisition module, used to acquire text data to be identified, and acquire user characteristics of a publishing account corresponding to the text data; wherein the user characteristics at least include the influence value and registration years of the publishing account; A processing module, used to obtain a first feature corresponding to the text data by using a preset bidirectional long short-term memory network Bi-LSTM, and to extract a second feature corresponding to the first feature by using a preset convolutional neural network; A classification model, used for obtaining the emotion category corresponding to the text data based on the second feature; The recognition module is used to obtain a probability value that the text data is false information based on the user characteristics and the emotion category.
7. The device according to claim 6, characterized in that When obtaining the user characteristics of the publishing account corresponding to the text data, the obtaining module is specifically used to: Acquire the registration time and current time of the publishing account, and acquire the first number of people subscribed to the publishing account and the second number of people subscribed to the publishing account; The registration years are calculated based on the registration time and the current time, and the influence value is calculated based on the first number of people and the second number of people.
8. The device according to claim 6, characterized in that When obtaining the emotion category corresponding to the text data based on the second feature, the classification module is specifically used to: The second feature is input into a target classifier to obtain the emotion category corresponding to the text data; wherein the target classifier is constructed based on a softmax classification model.
9. An electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.