A Model for Evaluating the Social Responsibility of Online Speech and Its Creation Method
By constructing a social responsibility evaluation model based on LSTM neural network, the problem of difficulty in effectively evaluating individual speech social responsibility in the existing technology is solved, and rapid and accurate quantitative evaluation is achieved, which improves the objectivity and efficiency of evaluation.
Patent Information
- Application Number
- CN202010933551.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2040-09-08
AI Technical Summary
The existing technology is difficult to effectively evaluate the social responsibility of individual speech, lacks quantitative research methods and mature corpus, manual judgment is time-consuming and labor-intensive and lacks fixed standards.
The double screening method is used to obtain positive and negative samples, and a mathematical model is constructed through LSTM neural network training. The method of combining a fully connected neural network and an LSTM neural network is used to construct an initial and final social responsibility evaluation model.
A fast and accurate social responsibility evaluation of a large number of remarks is achieved, and a feasible technical method is provided to build a social responsibility evaluation model, which improves the objectivity and efficiency of evaluation.
Smart Images

Figure CN112052677B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data analysis, and particularly to a model for evaluating the social responsibility of online speech based on big data analysis and a method for creating the same. Background Art
[0002] The two-step flow theory proposed by Lazarsfeld et al. in their works emphasizes that the information of the mass media is often more persuasive and acceptable and believable to ordinary people after being reprocessed by opinion leaders. Some social events are usually more influential after being reprocessed by Internet users, especially Internet opinion leaders. Therefore, effectively evaluating the social responsibility of online speech has important guiding significance and practical value for constructing a positive and harmonious online environment.
[0003] Currently, in the research on social responsibility, the focus of many scholars is mainly on the research of corporate social responsibility. In terms of evaluating the social responsibility of individual speech, scholars are less involved, and mostly mainly focus on how to cultivate and improve students' social responsibility. Moreover, the research is mostly qualitative research, and there is no quantitative statistical method based on big data to study and evaluate individual social responsibility. Scholars mainly focus on the research of corporate social responsibility in the study of social responsibility, and there is less research on individual social responsibility and it is qualitative research for a certain subdivision group. In the quantitative research on the social responsibility of individual speech, the relevant achievements are still blank. There is less research on individual social responsibility, its evaluation criteria are vague, subjective, and there is a lack of a mature corpus and annotation set in the relevant field for constructing a social responsibility evaluation model. In addition, there are also individual cases using the method of manual judgment. Since it is time-consuming and laborious to manually judge the social responsibility of some speech, and there are large differences among individuals in the cognitive social responsibility of speech, there is no fixed standard. Summary of the Invention
[0004] The invention purpose of the present invention is to create a set of feasible technical methods to construct a social responsibility evaluation model, and through the model, it can more accurately, quantitatively and quickly complete the evaluation of the social responsibility of a large number of speech.
[0005] To achieve the above invention purpose, the technical solution adopted by the present invention is: a model for evaluating the social responsibility of online speech, and the model is a mathematical model constructed by using a large number of positive and negative samples obtained by a double screening method to participate in the training of an LSTM neural network.
[0006] The positive and negative samples include positive texts and negative texts; the positive texts and negative texts are non-artificially marked samples, and are samples obtained by respectively capturing speeches with strong social responsibility and speeches without obvious social responsibility through the public network and then performing secondary screening.
[0007] The positive and negative texts include the initial positive and negative texts and the final positive and negative texts; the initial positive text and the initial negative text are trained by a fully connected neural network to obtain an initial evaluation model. After judgment according to the initial evaluation model, the first 50% of the initial positive text and the last 50% of the initial negative text are selected to form the final positive text and the final negative text, and then used as the final samples for training by the LSTM neural network to obtain the final social responsibility evaluation model.
[0008] A method for creating a model to evaluate the social responsibility of online speech, comprising the following steps:
[0009] S1. Respectively select the same number of positive and negative text data as the initial positive and negative texts, and use a fully connected neural network to construct an initial evaluation model;
[0010] S2. Apply the obtained initial evaluation model to the initial positive and negative texts for predicting the social responsibility index. The higher the predicted index, the stronger the social responsibility;
[0011] S3. According to the prediction results, respectively select the first 50% of the speeches of the initial positive text and the last 50% of the speeches of the initial negative text as the final positive and negative texts;
[0012] S4. Use the obtained final positive and negative texts to train with the LSTM neural network; construct the final model.
[0013] In step S1: After selecting the initial positive and negative texts, the initial positive and negative text data are segmented using the Jieba segmentation system in the accurate mode, and the segmented results are matched with the pre-trained dense word vector set to obtain the sample data word vector matrix;
[0014] Before training the initial evaluation model, the sample data word vector matrix is imported as the input layer into a two-layer fully connected neural network, using the ReLU activation function, and the number of neurons is 256 and 32 respectively; a Dropout layer with a value of 0.3 is added after each layer.
[0015] In step S2: The trained initial evaluation model is respectively applied to the initial positive and negative texts for prediction. During the prediction process, the method of batch reading texts from the MySQL database using Python is used. The texts are input into the initial evaluation model to obtain the prediction results, and then the prediction results are written into the database in batches.
[0016] In step S4: During the process of constructing the final model, before starting the training, operations such as segmenting and matching word vectors are first performed on the final positive and negative texts;
[0017] The final model is trained using an LSTM neural network. It has two hidden layers, namely an LSTM neural network layer with 256 neurons and a fully connected neural network layer with 64 neurons. The output layer consists of two neurons.
[0018] During training, the Adam optimizer is selected, and the ReLU activation function is used.
[0019] The model accuracy evaluation is carried out using a validation set of manually judged social responsibilities to evaluate the model accuracy, and then the model iteration is completed to optimize the model evaluation accuracy.
[0020] The above-mentioned model accuracy evaluation is to verify the accuracy of the test set of manually judged social responsibilities. The specific method is as follows:
[0021] Randomly select 1000 speech data from the microblog speeches of a number of online opinion leader groups as the validation set. The validation set is evenly distributed to a number of subjects. Subjectively judge the distributed speeches according to the discriminant basis of having social responsibility and having no obvious social responsibility. Mark the speeches with social responsibility as "1" and the speeches with no obvious social responsibility as "0". Apply the final model to the validation set for accuracy verification;
[0022] The sample predicted values in the validation set are mainly distributed in the low value interval of 0 - 0.05, accounting for 63% of the total samples, indicating that most of the validation set selected in the experiment is speeches without obvious social responsibility, which matches the statistical results of the validation set;
[0023] When the value for dividing having and not having social responsibility is 0.025, the model accuracy rate is 66.1%, and the accuracy rate gradually increases as the value increases;
[0024] It reaches the peak value when the value is 0.775, and the model prediction accuracy rate is 90.7%. After that, the model accuracy rate decreases;
[0025] The experiment uses 0.775 as the cut-off value for dividing having and not having social responsibility, that is, the interval [0, 0.775) is content without social responsibility, and [0.775, 1] is content with social responsibility.
[0026] Compared with the prior art, the social responsibility rating model and its construction method disclosed in the present invention are convenient to apply. The model obtains the social responsibility index of the speech by inputting the speech, and judges the intensity of social responsibility through the index; this model can quickly complete a relatively accurate social responsibility evaluation of a large number of speeches, and has strong practicability in applications such as monitoring network public opinion and network speech management. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a flow chart of the model training of the present invention.
[0028] Figure 2 This is the distribution diagram of the sample quantity corresponding to each segment of the social responsibility index and the model accuracy rate in the present invention.
[0029] Figure 3 This is the diagram of the proportion of socially responsible remarks over the years.
[0030] Figure 4 This is the relationship diagram between the segments of the social responsibility index of the ordinary user group and the average number of forwards and likes.
[0031] Figure 5 This is the relationship diagram between the segments of the social responsibility index of all online opinion leaders and the average number of forwards and likes. Specific implementation manner
[0032] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0033] For the technical solution disclosed by the present invention, since it is time-consuming and laborious to manually judge the social responsibility of certain remarks, and there are significant differences among individuals in recognizing the social responsibility of remarks without a fixed standard. The present invention mainly solves the problem of how to quickly evaluate the social responsibility of remarks using computer algorithms.
[0034] The present invention uses the LSTM neural network implemented with the Keras API in the TensorFlow open source framework using the Python language to construct a social responsibility evaluation model, and the construction process is as Figure 1 shown.
[0035] The positive and negative texts include initial positive and negative texts and final positive and negative texts; the positive and negative texts are remarks with strong social responsibility and no obvious social responsibility respectively. As an embodiment of the present invention, a total of 104,758 pieces of content data of the official microblog posts of People's Daily from July 2012 to April 2020 are selected, and used as the initial positive texts for constructing the social responsibility evaluation model to participate in model training, and the remarks published by ordinary users are used as the initial negative texts.
[0036] The construction method includes the following steps:
[0037] S1. Respectively select the same quantity of positive and negative text data as the initial positive and negative texts, and use a fully connected neural network to construct an initial evaluation model;
[0038] S2. Apply the obtained initial evaluation model to the initial positive and negative texts for predicting the social responsibility index, and the higher the predicted index, the stronger the social responsibility;
[0039] S3. To obtain more significant sample differences, eliminate the positive content in the negative texts, and separately select the first 50% of the statements in the initial positive texts and the last 50% of the statements in the initial negative texts as the final positive and negative texts.
[0040] S4. Use the obtained final positive and negative texts to train with an LSTM neural network; construct the final model.
[0041] In step S1, after selecting the initial positive and negative texts, use the Jieba word segmentation system to perform word segmentation on the initial positive and negative text data in the accurate mode, match the segmented results with the pre-trained dense word vector set, and obtain the sample data word vector matrix. In this embodiment, the selected dense word vector set contains a total of 195,203 Chinese words and the corresponding 300-dimensional word vectors. This set is pre-trained using the Skip-Gram with Negative Sampling (SGNS) model, which converts each word into a 300-dimensional word vector. This model predicts the context words based on the center word and can take into account the context relationship between words, improving the accuracy of the word vectors.
[0042] Before training the initial evaluation model, import the sample data word vector matrix as the input layer into a two-layer fully connected neural network, use the ReLU activation function, and the number of neurons is 256 and 32 respectively. To prevent overfitting during training, add a Dropout layer with a value of 0.3 after each layer. Since the experiment sets the label values of the initial positive and negative texts to 1 and 0 respectively, the output layer of the training consists of two neurons.
[0043] In step S2, apply the trained initial evaluation model to the initial positive and negative texts for prediction respectively. During the prediction process, use the method of batch reading texts from the MySQL database in Python, input the texts into the initial evaluation model to obtain the prediction results, and then batch write the prediction results into the database. Through experiments, using this method can effectively improve the model prediction speed. The prediction result is a decimal between 0 and 1. The closer the result is to 1, the stronger the social responsibility of the text, and vice versa.
[0044] In step S4, during the process of constructing the final model, before starting the training, perform operations such as word segmentation and matching word vectors on the final positive and negative texts (using the method in step S1); the final model is trained using an LSTM neural network, with two hidden layers, namely an LSTM neural network layer with 256 neurons and a fully connected neural network layer with 64 neurons, and the output layer consists of two neurons. During the training, the optimizer selects the Adam optimizer proposed by two scholars, Kingma and Lei Ba, which has excellent performance and strong adaptability, and the activation function uses the commonly used ReLU.
[0045] To improve accuracy, model accuracy evaluation is also included throughout the training process. This model accuracy evaluation uses a validation set of manually judged social responsibilities to evaluate the model accuracy, and then completes model iteration to optimize the model evaluation accuracy.
[0046] The specific method is as follows;
[0047] In an embodiment of the present invention, accuracy verification is performed on the test set of manually judged social responsibilities. 1000 speech data are randomly selected from 566817 microblog speeches of the network opinion leader group as the validation set, and the validation set is evenly distributed to 57 subjects. The subjects are undergraduate students aged between 20 and 22 years old. According to the discriminant basis of having social responsibility and having no obvious social responsibility, the assigned speeches are subjectively judged. Speeches with social responsibility are labeled as "1", and speeches with no obvious social responsibility are labeled as "0". The final model is applied to the validation set for accuracy verification, and the verification results are as Figure 2 shown.
[0048] The predicted values of the samples in the validation set are mainly distributed in the low value interval of 0 - 0.05, accounting for 63% of the total samples, indicating that most of the validation set selected in the experiment are speeches without obvious social responsibility, which matches the statistical results of the validation set; when the value for dividing having and not having social responsibility is 0.025, the model accuracy rate is 66.1%, and the accuracy rate gradually increases as the value increases; it reaches the peak when the value is 0.775, and the model prediction accuracy rate is 90.7%. After that, the model accuracy rate decreases; therefore, the experiment uses 0.775 as the segmentation value for dividing having and not having social responsibility, that is, the interval [0, 0.775) is content without social responsibility, and [0.775, 1] is content with social responsibility.
[0049] The following is a detailed description with specific embodiments
[0050] In this embodiment, the model is respectively applied to the network opinion leader group and the ordinary user group on Sina Weibo for calculation and analysis, and good application effects are obtained. Among them, for the network opinion leader group, the top 350 users in the average KOL index (a comprehensive index reflecting the activity, user coverage, and influence of a Weibo account) in the Lingku big data platform in three months are selected. For the ordinary Weibo user group, the random selection principle is adopted. Using the "Find People" function provided by the Weibo platform, searches are performed under the "Ordinary Users" category with "a", "b", "c", "d", "e" as keywords respectively, and a total of 1480 ordinary Weibo accounts are obtained as the ordinary Weibo user group.
[0051] The application effects are as follows:
[0052] (1) Compared with the ordinary user group, the network opinion leader group has a stronger sense of social responsibility
[0053] When the social responsibility index is 0.775, the evaluation accuracy of the test set is the highest. Therefore, this study selects the social responsibility prediction index of 0.775 as the standard for dividing obvious social responsibility. The study divides the remarks with a social responsibility index < 0.775 into remarks without obvious social responsibility; those with an index >= 0.775 are divided into responsible remarks. By statistically calculating the proportion of responsible remarks of different groups in the overall remarks annually, as Figure 3 shown, generally speaking, the social responsibility of all research objects shows an increasing trend on the time scale, and the social responsibility of Chinese Internet users is gradually increasing. Since the data of this batch of online opinion leaders was selected from November 2019 to January 2020, and the attention and activity of Weibo users vary greatly in different time periods, the performance of online opinion leaders in recent years should be mainly observed. The study found that in recent years, the social responsibility of online opinion leaders is significantly higher than that of ordinary user groups, which can reflect the traditional perception that "the greater the ability, the greater the responsibility".
[0054] (2) Responsible remarks are more recognizable and contagious
[0055] This study segments the social responsibility index and statistically calculates the average number of forwards and likes of the remarks of the research objects within the segmented intervals, as Figure 4 、 Figure 5 shown.
[0056] From Figure 4 it can be seen that the average number of forwards, the average number of likes and the social responsibility index of the ordinary user group show a strong positive correlation, that is, the higher the social responsibility of the remarks, the more the number of forwards and likes.
[0057] The correlation coefficients calculated according to Formula 1 (x = lower bound of the segmented interval, y = number of times) are 0.954 and 0.980 respectively.
[0058] Formula 1 Correl correlation coefficient calculation formula
[0059] From Figure 5 it can be seen that the number of forwards of the online opinion leader group also shows a strong positive correlation with the social responsibility index, and the correlation coefficient is 0.947. However, there is no obvious correlation in the number of likes of this group, and the correlation coefficient is -0.336.
[0060] The number of likes and reposts of online remarks is an important indicator of the sense of identity, influence, and dissemination power of such remarks. Fans express their sense of identity with the blogger's remarks through likes and reposts and even participate in the further dissemination of such remarks. The number of likes on Weibo mainly measures the broad sense of identity of remarks. Remarks with a high number of likes have a stronger social sense of identity. On the basis of measuring the broad sense of identity, the repost volume index can better express fans' willingness to spread remarks, and it is the remarks that generate greater dissemination power and influence. It can be seen that as the social responsibility of all users' remarks increases, the sense of identity, influence, and dissemination power of their remarks are also gradually increasing, and the above conclusion is more obvious in groups with relatively weak influence.
[0061] Thus, it can be seen that the research object responds promptly in major public events and can demonstrate a strong sense of social responsibility. The online opinion leaders show a more rapid and intense response, but with insufficient persistence. The group of online opinion leaders has extensive social attention and the power to spread remarks. They can demonstrate a sense of responsibility corresponding to their influence in the initial stage of the epidemic and play an important positive role in responding to the national call and publicizing epidemic prevention and control. However, compared with the ordinary user group, their subsequent persistence is insufficient. If the subsequent persistence of the group of online opinion leaders can be improved, better results can be achieved in the publicity and prevention of major public events in the future.
Claims
1. A method for creating a model to evaluate the social responsibility of online speech. The model is a mathematical model constructed by training an LSTM neural network with a large number of positive and negative samples obtained by a double screening method. The positive and negative samples include positive texts and negative texts. The positive and negative texts are non-artificially labeled texts, which are samples obtained by respectively scraping speeches with strong social responsibility and speeches without obvious social responsibility from the public network and then performing secondary screening. The positive and negative texts include initial positive and negative texts and final positive and negative texts. The initial positive text and the initial negative text are trained by a fully connected neural network to obtain an initial evaluation model. After judgment according to the initial evaluation model, the first 50% of the initial positive text and the last 50% of the initial negative text are selected to form the final positive text and the final negative text, and then used as the final samples for training the LSTM neural network to obtain the final social responsibility evaluation model. Characterized in that, The construction method includes the following steps: S1. Respectively select the same number of positive and negative text data as the initial positive and negative texts, and use a fully connected neural network to construct an initial evaluation model. S2. Apply the obtained initial evaluation model to the initial positive and negative texts for predicting the social responsibility index. The higher the predicted index, the stronger the social responsibility. The input of the initial evaluation model is to perform word segmentation on the initial positive and negative text data using the Jieba word segmentation system in the accurate mode, and match the segmented results with the pre-trained dense word vector set to obtain the sample data word vector matrix. The input layer of the initial evaluation model includes two layers of fully connected neural networks, using the ReLU activation function, with the number of neurons being 256 and 32 respectively. A Dropout layer with a value of 0.3 is added after each layer. The label values of the initial positive and negative texts are set to 1 and 0 respectively. The output layer consists of two neurons. The output of the initial evaluation model is the social responsibility index, specifically a decimal between 0 and 1. S3. According to the prediction results, respectively select the first 50% of the speeches in the initial positive text and the last 50% of the speeches in the initial negative text as the final positive and negative texts. S4. Use the LSTM neural network to train the obtained final positive and negative texts; construct the final model. The input of the final model is the word vector matrix after performing word segmentation and matching word vectors on the final positive and negative texts. The final model is trained using the LSTM neural network. The hidden layer has two layers, namely an LSTM neural network layer with 256 neurons and a fully connected neural network layer with 64 neurons. The output layer consists of two neurons.
2. The model creation method according to claim 1, Characterized in that, In step S2: Apply the trained initial evaluation model to the initial positive and negative texts for prediction respectively. In the prediction process, use the method of batch reading texts from the MySQL database in Python. Input the texts into the initial evaluation model to obtain the prediction results, and then batch write the prediction results into the database.
3. The model creation method according to claim 1, Characterized in that, During training, the Adam optimizer is selected, and the ReLU activation function is used.
4. The model creation method according to claim 1, wherein, it further includes model accuracy evaluation; the model accuracy evaluation is to use a validation set for manually judging social responsibility to evaluate the model accuracy, and then complete model iteration to optimize the model evaluation accuracy.
5. The model creation method according to claim 4, wherein, the model accuracy evaluation is to perform accuracy verification on a test set for manually judging social responsibility, and the specific method is as follows: Input the validation set into the final model to obtain predicted values, and use 0.775 as the segmentation value for dividing whether there is social responsibility, that is, the interval [0, 0.775) is content without social responsibility, and [0.775, 1] is content with social responsibility; after dividing the predicted values of the final model, compare them with the results of manual judgment to perform accuracy verification.
Citation Information
Patent Citations
A method for constructing emotion recognition model of Chinese social text based on deep fusion neural network
CN109299253A
Drunk driving detection method and system based on sensor and machine vision
CN110070078A