Data detection method, device and storage medium
By combining multidimensional data to cyclically train the training text data set, the current text recognition model is constructed, which solves the problem of low ironic text recognition efficiency in group conversations, and achieves more efficient text risk control and intelligent recognition.
Patent Information
- Application Number
- CN202110896864.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-05
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-08-05
AI Technical Summary
The prior art is difficult to effectively identify ironic texts in group conversations, resulting in low recognition efficiency and high manslaughter rate in text risk control, which affects the construction of intelligent capabilities.
By obtaining the training text data set, combining multi-dimensional data to extract multiple ironic text information, the current text recognition model is trained cyclically, and the current text recognition model is constructed to improve the recognition accuracy using part-of-speech, emotional information, portrait and evaluation classification information.
It improves the recognition efficiency of ironic texts in group conversations, reduces the rate of manslaughter, and enhances the accuracy and intelligence of text risk control.
Smart Images

Figure CN113609854B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of Internet risk control technology, and in particular to a data detection method, device, and storage medium. Background Art
[0002] Currently, most companies face fierce industry competition and the various risks posed by the shady industry. Text-based risk control is a crucial risk management tool on the Chinese internet. Sensitive content such as pornography, harassment, insults, political content, and fraud is prevalent in forums, business consulting systems, online reviews, personal information websites, Weibo, and other communities and business domains. This negative content is then exposed in various public domains. Furthermore, this negative content can impact the development of intelligent capabilities such as input association and auto-answering robots, amplifying negative impacts.
[0003] Currently, among all types of negative text, ironic text in group conversations is the most challenging to identify. Ironic text refers to users expressing dissatisfaction through irony or sarcasm. Irony involves using words that contradict the intended meaning to convey a negative, sarcastic, or mocking message. Irony is a rhetorical device with strong emotional overtones, such as "I just love this lousy place." Sarcasm uses metaphors and exaggeration to expose, criticize, or ridicule people or events, such as "XX Express, a fighter among pheasants." These ironic texts often express dissatisfaction with positive sentiment, making it difficult to accurately identify the actual sentiment of ironic text using traditional sentiment analysis. Irony and other rhetorical devices significantly impact text recognition accuracy, making the identification of ironic text crucial for text risk control.
[0004] Currently, ironic text detection technology primarily utilizes methods such as keyword spotting, sentiment analysis, and supervised learning, each with its own advantages and disadvantages. Keyword spotting typically involves building a manually collected corpus. Then, during text filtering, ironic text is identified based on whether the corpus identifies the keywords in the text. This approach is inefficient for complex ironic text. Sentiment analysis uses sentiment as a primary criterion, scoring the text using a sentiment recognition model. However, this approach can lead to false negatives and is inefficient for identifying ironic text. The third approach, supervised learning, requires a large ironic corpus to identify ironic text. This approach is also less timely. After a long period of corpus collection and training, users' ironic expressions often change, leading to model failure and, consequently, low ironic text recognition efficiency. Summary of the Invention
[0005] The embodiments of the present invention provide a data detection method, device, and storage medium, which can improve the recognition efficiency of ironic text in group conversations.
[0006] The technical solution of the present invention is achieved as follows:
[0007] An embodiment of the present invention provides a data detection method, including:
[0008] Get the current text information to be checked;
[0009] Detecting the current text information to be detected using the current text recognition model to determine a detection result of the current text information to be detected;
[0010] The current text recognition model is obtained by obtaining a training text data set, extracting multiple ironic text information from the training text data set in combination with multidimensional data, and cyclically training the previous text recognition model using the multiple ironic text information; the multidimensional data includes: at least two of: part of speech class, sentiment information class, portrait class and evaluation classification information.
[0011] In the above scheme, the current text recognition model is based on the portrait of the target object corresponding to the training text data set containing the context, and the evaluation classification information of the conversation text information in the training text data set, and the multiple ironic text information is determined in the training text data set, and the previous text recognition model is cyclically trained using the multiple ironic text information.
[0012] In the above solution, before detecting the current text information to be detected using the current text recognition model and determining the detection result of the current text information to be detected, the method further includes:
[0013] Acquire a training text data set, and extract multiple ironic text information from the training text data set by combining the multidimensional data; the training text data set includes conversation text information and related information of multiple target objects, wherein each text information represents a single scene conversation text information of the corresponding target object;
[0014] The previous text recognition model is trained based on the multiple ironic text information until the current text recognition model is obtained.
[0015] In the above solution, the method further includes:
[0016] extracting again a plurality of supplementary ironic text information from the training text data set using the current text recognition model;
[0017] Combining the plurality of supplementary ironic text information with the plurality of ironic text information, constructing a plurality of portraits of a plurality of target objects corresponding to the training text data set; each of the plurality of portraits represents ironic text score information corresponding to the corresponding target object;
[0018] Based on the multiple portraits, multiple ironic text information is extracted from the training text data set to be obtained next time, so as to train the current text recognition model to obtain a next text recognition model.
[0019] In the above solution, the step of obtaining a training text data set and extracting a plurality of ironic text information from the training text data set in combination with multidimensional data includes:
[0020] Acquire the training text data set, segment the text in the training text set to obtain multiple text information;
[0021] Combining at least two of the part-of-speech class, the sentiment information class, the portrait class, and the evaluation classification information, performing text recognition on the plurality of text information to obtain a plurality of scores corresponding to the plurality of text information;
[0022] Based on the plurality of scores, a plurality of ironic text information is determined among the plurality of text information.
[0023] In the above solution, the multiple scores include: multiple four-level scores;
[0024] The combining of at least two of the part-of-speech class, the sentiment information class, the portrait class, and the evaluation classification information to perform text recognition on the multiple text information to obtain multiple scores corresponding to the multiple text information includes:
[0025] Performing word segmentation on multiple pieces of text information to be selected from the multiple pieces of text information to obtain multiple keywords corresponding to the multiple pieces of text information; the multiple pieces of text information to be selected are the text information remaining after the multiple pieces of text information are filtered by the previous text recognition model;
[0026] If the multidimensional data includes four of the part-of-speech class, the sentiment information class, the portrait class, and the evaluation classification information, then based on the parts of speech of the multiple keywords, a plurality of first-level ironic text information is determined from the multiple candidate text information, and scores are added to the multiple first-level ironic text information to obtain a plurality of first-level scores corresponding to the multiple candidate text information;
[0027] Based on the sentiment information of the plurality of to-be-selected text information and the sentiment information of the corresponding context information, determining a plurality of secondary ironic text information from the plurality of to-be-selected text information, and adding the primary scores corresponding to the plurality of secondary ironic text information to obtain a plurality of secondary scores corresponding to the plurality of to-be-selected text information;
[0028] Based on the multiple associated portraits corresponding to the target objects of the multiple text information to be selected, determining multiple third-level ironic text information from the multiple text information to be selected, and adding the second-level scores corresponding to the multiple third-level ironic text information to obtain multiple third-level scores corresponding to the multiple text information to be selected;
[0029] Based on the multiple evaluation scores corresponding to the multiple text information to be selected and the emotional information of the multiple text information to be selected, multiple four-level ironic text information are determined from the multiple text information to be selected, and the three-level scores corresponding to the multiple four-level ironic text information are added to obtain multiple four-level scores corresponding to the multiple text information to be selected.
[0030] In the above solution, based on the parts of speech of the multiple keywords, multiple first-level ironic text information is determined from the multiple text information to be selected, and the multiple first-level ironic text information is scored to obtain multiple first-level scores corresponding to the multiple text information to be selected, including:
[0031] Finding the parts of speech of the multiple keywords in the sentiment word list; the sentiment word list is a preset list of parts of speech including all keywords;
[0032] Determining, from the plurality of text information to be selected, a plurality of first-level ironic text information including keywords with opposite parts of speech;
[0033] The plurality of first-level ironic text information are scored based on the initial scores corresponding to the plurality of text information to be selected, to obtain a plurality of first-level scores corresponding to the plurality of text information to be selected.
[0034] In the above solution, the method of determining a plurality of secondary ironic text information from the plurality of text information to be selected based on the emotional information of the plurality of text information to be selected and the emotional information of the corresponding context information, and adding the primary scores corresponding to the plurality of secondary ironic text information to obtain a plurality of secondary scores corresponding to the plurality of text information to be selected, includes:
[0035] Extracting multiple context information corresponding to the multiple text information to be selected from the training text data set;
[0036] Detecting and obtaining emotional information of the plurality of contextual information and emotional information of the plurality of selected text information;
[0037] Determining, from the plurality of text information to be selected, a plurality of secondary ironic text information whose sentiment information is positive and whose corresponding contextual information has negative sentiment information;
[0038] The plurality of secondary ironic text information are added to the primary scores corresponding to the plurality of text information to be selected, to obtain a plurality of secondary scores corresponding to the plurality of text information to be selected.
[0039] In the above solution, the method of determining a plurality of third-level ironic text information from the plurality of text information to be selected based on the plurality of associated portraits corresponding to the target objects of the plurality of text information to be selected, and adding the second-level scores corresponding to the plurality of third-level ironic text information to obtain a plurality of third-level scores corresponding to the plurality of text information to be selected includes:
[0040] Extracting a plurality of associated portraits corresponding to the target objects corresponding to the plurality of text information to be selected from the plurality of portraits;
[0041] Determining, from the plurality of text information to be selected, a plurality of third-level ironic text information whose target objects represented by the corresponding associated portraits are users prone to irony;
[0042] The plurality of third-level ironic text information are added to the second-level scores corresponding to the plurality of text information to be selected, to obtain a plurality of third-level scores corresponding to the plurality of text information to be selected.
[0043] In the above solution, the method of determining a plurality of four-level ironic text information from the plurality of text information to be selected based on the plurality of evaluation scores corresponding to the plurality of text information to be selected and the sentiment information of the plurality of text information to be selected, and adding the three-level scores corresponding to the plurality of four-level ironic text information to obtain the plurality of four-level scores corresponding to the plurality of text information to be selected includes:
[0044] Extracting multiple evaluation scores corresponding to the multiple pieces of text information to be selected from the training text data set;
[0045] Determining, from the plurality of text messages to be selected, a plurality of level four ironic text messages having positive sentiment information and corresponding evaluation scores lower than a first threshold;
[0046] The plurality of fourth-level ironic text information are added to the third-level scores corresponding to the plurality of text information to be selected, so as to obtain a plurality of fourth-level scores corresponding to the plurality of text information to be selected.
[0047] In the above solution, the detecting and obtaining the emotional information of the plurality of contextual information and the emotional information of the plurality of selected text information includes one of the following:
[0048] Inputting the plurality of context information and the plurality of to-be-selected text information into an emotion recognition model to obtain emotion information of the plurality of context information and emotion information of the plurality of to-be-selected text information;
[0049] Segmenting the multiple context information into multiple upper and lower keywords;
[0050] Find the parts of speech of the multiple upper and lower keywords in the sentiment word list, and find the parts of speech of the multiple keywords in the sentiment word list;
[0051] The sentiment information of the plurality of contextual information and the sentiment information of the plurality of candidate text information are determined based on the parts of speech of the plurality of contextual keywords and the parts of speech of the plurality of keywords.
[0052] In the above solution, before performing word segmentation processing on the plurality of candidate text information of the plurality of text information to obtain a plurality of keywords corresponding to the plurality of candidate text information, the method further includes:
[0053] The plurality of text information are input into the previous text recognition model to obtain a plurality of five-level ironic text information from the plurality of text information.
[0054] In the above solution, determining a plurality of ironic text messages from the plurality of text messages based on the plurality of scores includes:
[0055] Determining, from the plurality of text messages to be selected, a plurality of level 6 ironic text messages whose corresponding level 4 scores are higher than a second threshold;
[0056] The plurality of fifth-level ironic text messages and the plurality of sixth-level ironic text messages are combined to form the plurality of ironic text messages.
[0057] In the above solution, the step of combining the plurality of supplementary ironic text information with the plurality of ironic text information to construct a plurality of portraits of the plurality of target objects corresponding to the training text data set includes:
[0058] Inputting multiple text information in the training text data set into the emotion recognition model to obtain emotional text information corresponding to multiple target objects corresponding to the training text data set;
[0059] Extracting the number of evaluations and the number of negative reviews corresponding to the multiple target objects from the training text data set;
[0060] The multiple portraits are constructed based on the multiple supplementary ironic text information, the multiple ironic text information, the emotional information and the negative review quantity information.
[0061] In the above solution, the constructing of the multiple portraits based on the multiple supplementary ironic text information, the multiple ironic text information, the emotional information, and the negative review quantity information includes:
[0062] Comparing the sum of the supplementary ironic text information and the ironic text information corresponding to the plurality of target objects with the total number of corresponding text information, obtaining irony scores corresponding to the plurality of target objects respectively;
[0063] Comparing the emotional text information corresponding to the multiple target objects with the total number of corresponding text information to obtain emotional scores corresponding to the multiple target objects;
[0064] Divide the number of negative reviews corresponding to the multiple target objects by the number of corresponding reviews to obtain negative impact scores corresponding to the multiple target objects;
[0065] The multiple portraits are calculated based on the irony scores, emotion scores, and negative impact scores respectively corresponding to the multiple target objects.
[0066] In the above solution, the calculation of the multiple portraits based on the irony scores, emotion scores, and negative impact scores corresponding to the multiple target objects includes:
[0067] Adding the product of the irony score and the irony weight corresponding to each of the multiple target objects, the product of the emotion score and the emotion weight, and the product of the negative influence score and the negative influence weight, to obtain multiple comprehensive scores corresponding to the multiple target objects;
[0068] A correspondence between the multiple comprehensive scores and the multiple target objects is formed, thereby forming the multiple portraits.
[0069] In the above solution, after forming the correspondence between the multiple comprehensive scores and the multiple goals, and then forming the multiple portraits, the method further includes:
[0070] Determine that the target objects corresponding to the first proportion with the largest comprehensive score among the plurality of target objects are easy-to-ironic objects;
[0071] The target objects corresponding to the first second proportion having the smallest comprehensive score among the plurality of target objects are determined to be non-ironic objects.
[0072] An embodiment of the present invention further provides a data detection device, comprising:
[0073] An acquisition unit, used to acquire the current text information to be inspected;
[0074] A detection unit, configured to detect the current text information to be detected using a current text recognition model, and determine a detection result of the current text information to be detected;
[0075] The current text recognition model is obtained by obtaining a training text data set, extracting multiple ironic text information from the training text data set in combination with multidimensional data, and cyclically training the previous text recognition model with the multiple ironic text information; the multidimensional data represents at least two of the classification information of part of speech, emotional information, portrait, and evaluation.
[0076] An embodiment of the present invention further provides a data detection device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps in the above method when executing the program.
[0077] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above method when executed by a processor.
[0078] In an embodiment of the present invention, current text information to be inspected is obtained; the current text information to be inspected is inspected using a current text recognition model to determine a detection result for the current text information to be inspected; wherein the current text recognition model is obtained by obtaining a training text data set, extracting multiple ironic text information from the training text data set in combination with multidimensional data, and cyclically training a previous text recognition model using the multiple ironic text information; the multidimensional data represents at least two of the following classification information: part of speech, sentiment information, image, and evaluation. Because the multiple ironic text information is obtained by extracting the multiple ironic text information from the training text data set in combination with the multidimensional data, and because the multidimensional data considers multiple factors in the ironic text information, the previous text recognition model can learn a large amount of comprehensive ironic text information. Furthermore, the current text recognition model, trained with this large amount of comprehensive ironic text information, can accurately identify ironic text information in group conversations, thereby improving the efficiency of ironic text recognition in group conversations. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0080] Figure 2 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0081] Figure 3 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0082] Figure 4 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0083] Figure 5 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0084] Figure 6 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0085] Figure 7 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0086] Figure 8 An optional effect diagram of the data detection method provided by an embodiment of the present invention;
[0087] Figure 9 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0088] Figure 10 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0089] Figure 11 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0090] Figure 12 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0091] Figure 13 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0092] Figure 14 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0093] Figure 15 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0094] Figure 16 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0095] Figure 17 An optional effect diagram of the data detection method provided by an embodiment of the present invention;
[0096] Figure 18 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0097] Figure 19 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0098] Figure 20 An optional flowchart of a data detection method provided in an embodiment of the present invention;
[0099] Figure 21 A schematic structural diagram of a data detection device provided in an embodiment of the present invention;
[0100] Figure 22 A schematic diagram of a hardware entity of a data detection device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0101] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting the present invention. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0102] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0103] If similar descriptions of "first / second" appear in the invention document, the following explanation is added. In the following description, the terms "first\second\third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with the specific order or sequence where permitted, so that the embodiments of the invention described herein can be implemented in an order other than that illustrated or described herein.
[0104] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention pertains. The terms used herein are for the purpose of describing embodiments of the present invention only and are not intended to limit the present invention.
[0105] Figure 1 An optional flow chart of the data detection method provided in the embodiment of the present invention is combined with Figure 1 The steps shown are explained.
[0106] S101: Obtain the current text information to be inspected.
[0107] In the embodiment of the present invention, the server obtains the current to-be-checked text information at the current moment in the multi-object conversation text information in the client through a communication line pre-established with the client.
[0108] In the embodiment of the present invention, the server obtains the current to-be-checked text information within a period of time in the multi-object conversation text information of a certain application in the client through a communication line pre-established with the client.
[0109] In the embodiment of the present invention, the server obtains, through a communication line pre-established with the client, current to-be-checked text information corresponding to multiple objects in a multi-object session of a certain application in the client within a period of time.
[0110] In the embodiment of the present invention, the current text information to be checked is a text sentence transmitted by the corresponding object in the client. For example, the current text information to be checked can be a text conversation transmitted by the corresponding object in the client: "It will definitely rain tomorrow."
[0111] S102: Detect the current text information to be detected using the current text recognition model to determine a detection result of the current text information to be detected.
[0112] In the embodiment of the present invention, the server inputs the current text information to be checked into the current text recognition model, and the current text recognition model outputs a detection result corresponding to the current text information to be checked, that is, a result indicating whether the current text information to be checked is ironic text.
[0113] In an embodiment of the present invention, the current text recognition model includes a detection model based on a convolutional neural network model. The current text recognition model is a previous text recognition model obtained by obtaining a training text data set, extracting multiple ironic text information from the training text data set in combination with multidimensional data, and cyclically training the previous text recognition model using the multiple ironic text information; the multidimensional data includes at least two of the following: part-of-speech, sentiment information, portrait, and evaluation classification information. Because the ironic text information in the training text data set is fully extracted during the training process, and the ironic text library is updated in real time through cyclic training, the number and quality of learning samples of the previous text recognition model are improved. As a result, the current text recognition model obtained after training can efficiently recognize ironic text information in multi-object conversations.
[0114] In an embodiment of the present invention, the current text recognition model may include: an optical character recognition model (OCR) or a text recognition model (Convolutional Recurrent Neural Network, CRNN). Among them, the network structure of the current text recognition model may include: an input layer, several intermediate layers and an output layer. The current text recognition model obtains the current text information to be inspected through the acquisition unit. The current text recognition model inputs the current text information to be inspected into the input layer. After the intermediate layer of the current text recognition model processes the current text information to be inspected, the output layer of the previous text recognition model outputs the confidence of the current text information to be inspected. The previous text recognition model judges the size of the confidence. When the confidence is greater than a certain value, the current text recognition model outputs the corresponding current text information to be inspected as ironic text information.
[0115] In an embodiment of the present invention, the current text recognition model is based on a portrait of a target object corresponding to a training text data set containing context, and classification information obtained by evaluating conversational text information in the training text data set. A plurality of ironic text information is determined in the training text data set, and the previous text recognition model is cyclically trained using the plurality of ironic text information.
[0116] In an embodiment of the present invention, current text information to be inspected is obtained; the current text information to be inspected is inspected using a current text recognition model to determine a detection result of the current text information to be inspected; wherein the current text recognition model is obtained by obtaining a training text data set, extracting multiple ironic text information from the training text data set in combination with multidimensional data, and cyclically training a previous text recognition model using the multiple ironic text information; the multidimensional data represents the classification data of part of speech, sentiment information, image, and evaluation. Because the multiple ironic text information is obtained by extracting the training text data set in combination with the multidimensional data, and because the multidimensional data includes multiple factors in the ironic text information, the previous text recognition model can learn a large amount of comprehensive ironic text information. The current text recognition model, which is trained with a large amount of comprehensive ironic text information, can accurately identify ironic text information in group conversations, thereby improving the recognition efficiency of ironic text in group conversations.
[0117] In some embodiments, see Figure 2 , Figure 2 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 1 The steps before S102 shown in the figure also include S103 to S104, which will be described in conjunction with each step.
[0118] S103: Acquire a training text data set, and extract multiple ironic text information from the training text data set by combining the multidimensional data.
[0119] In an embodiment of the present invention, a server obtains a training text dataset corresponding to multiple clients via pre-established communication lines with the multiple clients. The training text dataset may include multiple text information about multiple target objects. The server extracts multiple ironic text information from the training text dataset by combining multiple portraits corresponding to the multiple target objects, the parts of speech of keywords in the training text dataset, and the sentiment and evaluation information of the keywords in the training text dataset.
[0120] The training text dataset includes conversational text information and related information for multiple target subjects, wherein each conversational text information represents a single scene of conversational text information for the corresponding target subject. Each of the multiple portraits represents ironic text score information corresponding to the corresponding target subject. The related information may be evaluation or scoring information for the corresponding text information.
[0121] In the embodiment of the present invention, the server may extract multiple ironic text information from multiple text information in the training text data set using the previous text recognition model.
[0122] In the embodiment of the present invention, the server may determine a plurality of ironic text messages from a plurality of text messages through evaluation information or scoring information corresponding to the plurality of text messages.
[0123] In an embodiment of the present invention, the server can determine ironic text information among multiple ironic text information based on the sentiment information of the subject words in the multiple text information. For example, when the sentiment information of the subject word of a text information is positive, and the evaluation or score information corresponding to the text information is negative or low, the text information is ironic text information.
[0124] For example, combined Figure 3 The server can extract consultation texts, evaluation texts, and user ratings corresponding to multiple target objects from a locally stored original database. The service combines the consultation texts, evaluation texts, and user ratings to form a training text data set. The server performs ironic text recognition in step S201. New corpus mining is performed on the training text data set, that is, multiple ironic text information is extracted therefrom. The server extracts multiple ironic text information from the training text data set through ironic text recognition.
[0125] S104: Training the previous text recognition model based on the multiple ironic text information until a current text recognition model is obtained.
[0126] In the embodiment of the present invention, the server sequentially inputs multiple ironic text information into the previous text recognition model for training until the function of the previous text recognition model converges or reaches a certain number of training times, thereby obtaining the current text recognition model.
[0127] In an embodiment of the present invention, the training of a previous text recognition model can be represented by the following stages: a forward propagation stage, a backward propagation stage, and a weight update stage. The forward propagation stage is the backward transmission of text information from the input layer to the output layer. The backward propagation stage is the forward transmission from the output layer to the input layer. In the data detection method proposed in an embodiment of the present invention, during the forward propagation stage, text information is input into the network structure of the previous text recognition model to be trained. The network structure of the previous text recognition model calculates the loss corresponding to the first negative sample using a loss function based on the text information.
[0128] In an embodiment of the present invention, the network structure of the previous text recognition model calculates the loss of the corresponding text information based on the loss function. If the loss is greater than the loss threshold, the network structure of the previous text model will be back-propagated layer by layer through the output layer to the intermediate layer and the input layer based on the loss, and the weights of each layer will be corrected in a gradient descent manner. After the weights of each layer of the network structure of the previous text recognition model are corrected, the network structure of the previous text recognition model will continue to train the newly acquired text information. The process of training the previous text recognition model to obtain the current text recognition model continues until the current loss calculated by the previous text recognition model is no greater than the loss threshold, or the number of times the previous text recognition model is trained reaches a preset number of training times, and the current text recognition model is obtained.
[0129] For example, combined Figure 3 . The server inputs the extracted multiple ironic text information into the previous text recognition model, and the server performs step S202 and self-learning model. Then the current text recognition model is formed. The server then performs step S203 and secondary content screening on the training text data set through the current text recognition model. Then the supplementary ironic text information is extracted. The server constructs user portraits corresponding to multiple target objects based on multiple ironic text information and supplementary text information. The server extracts multiple ironic text information in the next training text data set through multiple user portraits and the current text recognition model. The server forms the current text recognition model by training the previous text model, and then forms the user portrait, which is a process of model iteration. After the server obtains the current text recognition model, it can perform step S204 and automatic deployment of the new version of the model on the current text recognition model, and then recognize the current text information to be inspected to complete the model application.
[0130] In some embodiments, see Figure 4, Figure 4 An optional flow chart of the data detection method provided by the embodiment of the present invention also includes S105 to S107, which will be described in conjunction with each step.
[0131] S105 , using the current text recognition model to extract multiple supplementary ironic text information from the training text data set again.
[0132] In the embodiment of the present invention, the server inputs a plurality of text information in the training text data set into the current text recognition model, and obtains a plurality of supplementary ironic text information in the training text data set through the current text recognition model.
[0133] In this embodiment of the present invention, since the multiple ironic text information in the training text dataset was partially acquired by the previous text recognition model, the multiple ironic text information in the training text dataset does not represent all the ironic text information in the training text dataset. However, the current text recognition model, which is fully trained based on the training text recognition model, has higher recognition efficiency and accuracy than the previous text recognition model. Therefore, the server uses the current text recognition model to extract multiple supplementary ironic text information from the training text dataset, thus compensating for the deficiencies in the previous text recognition model's ironic text information extraction.
[0134] S106. Combining the plurality of supplementary ironic text information and the plurality of ironic text information, construct a plurality of portraits of the plurality of target objects corresponding to the training text data set.
[0135] In an embodiment of the present invention, the server combines the plurality of supplementary ironic text information and the individual ironic text information to construct a plurality of portraits of the plurality of target objects corresponding to the training text data set.
[0136] In an embodiment of the present invention, the server can extract the ironic text information and supplementary ironic text information corresponding to each target object from the plurality of ironic text information and the plurality of supplementary ironic text information. The server calculates a score for each target object based on the scores of the ironic text information and supplementary ironic text information, and then labels each target object based on the score information. This results in a profile for each target object. Exemplarily, the server can establish a correspondence between the identification information of each target object and the corresponding score, thereby constructing a profile corresponding to each target object.
[0137] S107. Based on the multiple portraits, extract multiple ironic text information from the training text data set to be obtained next time, so as to train the current text recognition model to obtain the next text recognition model.
[0138] In an embodiment of the present invention, the server obtains multiple portraits of multiple target objects. When the clients of the multiple target objects transmit new text information to the server, the server can form a new training text data set, i.e., the next training text data set obtained. The server can extract multiple ironic text information from the new training text data set using the current text recognition model. The server trains the current text recognition model using the multiple ironic text information obtained this time to obtain the next text recognition model.
[0139] In some embodiments, see Figure 5 , Figure 5 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 2 The illustrated S103 can also be implemented through S108 to S110 , which will be described in conjunction with each step.
[0140] S108: Acquire a training text data set, segment the text in the training text set, and obtain multiple text information.
[0141] In an embodiment of the present invention, the training text data set includes conversation texts of multiple target objects. The conversation texts may be incoherent or semantically discontinuous. The server segments the conversation texts using a segmentation model to obtain multiple text information.
[0142] S109 , combining at least two of the classified information of part of speech, sentiment information, portrait, and evaluation, and performing text recognition on the plurality of text information to obtain a plurality of scores corresponding to the plurality of text information.
[0143] In an embodiment of the present invention, the server may first identify some ironic text information from the multiple text messages using a previous text recognition model. The server then combines multiple profiles of the multiple target objects and the parts of speech of the keywords in the candidate text information to determine multiple scores for the candidate text information corresponding to the multiple target objects. The multiple candidate text information is the text information remaining after the multiple text messages have been filtered by the previous text recognition model.
[0144] In an embodiment of the present invention, the server may further combine multiple portraits of multiple target objects and sentiment information of keywords in the text information to be selected to determine multiple scores in the text information to be selected corresponding to the multiple target objects.
[0145] In the embodiment of the present invention, the server may also determine multiple scores of the text information to be selected corresponding to multiple target objects by combining the part of speech and sentiment information of the keywords in the text information to be selected.
[0146] In the embodiment of the present invention, the server may further determine multiple scores of the text information to be selected corresponding to multiple target objects by combining the parts of speech of the keywords in the text information to be selected and the evaluation information of the text information to be selected.
[0147] In an embodiment of the present invention, a server identifies multiple first-level ironic text messages having opposite keyword polarities in multiple text messages. The server adds points to the multiple first-level ironic text messages based on the initial scores of the multiple text messages to be selected, thereby obtaining multiple first-level scores corresponding to the multiple text messages to be selected. The server then identifies multiple second-level ironic text messages having positive sentiment information and corresponding contextual information having negative sentiment information from the multiple text messages to be selected. The server adds points to the multiple second-level ironic text messages based on the first-level scores of the multiple text messages to obtain multiple scores corresponding to the multiple text messages to be selected.
[0148] In an embodiment of the present invention, if a target object's profile indicates that the target object is prone to irony, and the text information corresponding to the target object is sentimentally positive and has a low evaluation, the server may assign a confidence score to the text information corresponding to the target object, thereby obtaining a score for the text information corresponding to the target object.
[0149] In an embodiment of the present invention, the server can also determine the intermediate scores of multiple text messages based on the emotional information of the keywords in the multiple text messages. The server adds the intermediate scores of the multiple text messages to the scores of the text information obtained by the server by combining the portraits to obtain the final scores of the multiple text messages.
[0150] S110 : Determine a plurality of ironic text information from a plurality of text information based on a plurality of scores.
[0151] In the embodiment of the present invention, the server determines a plurality of ironic text messages having scores higher than a threshold among the plurality of text messages.
[0152] In some embodiments, see Figure 6 , Figure 6 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 5 The illustrated S109 can also be implemented through S111 to S115 , which will be described in conjunction with each step.
[0153] S111 , performing word segmentation processing on a plurality of text information to be selected from a plurality of text information to obtain a plurality of keywords corresponding to the plurality of text information to be selected.
[0154] In the embodiment of the present invention, the server performs word segmentation processing on a plurality of candidate text information among a plurality of text information through a word segmentation model, and obtains a plurality of keywords corresponding to the plurality of candidate text information respectively.
[0155] In the embodiment of the present invention, the plurality of text information to be selected is the text information remaining after the plurality of text information is filtered by the previous text recognition model. Figure 7 , each step will be explained.
[0156] S205: Input text recognition model.
[0157] In the embodiment of the present invention, the server inputs multiple text information into the previous text recognition model.
[0158] S206: Whether it has been identified as an ironic sentence.
[0159] In this embodiment of the present invention, the server uses a previous text recognition model to determine whether each text message in the plurality of text messages is an ironic sentence. If the previous text recognition model identifies n text messages in the plurality of text messages as ironic sentences, the server classifies the n ironic text messages into the plurality of ironic text messages. The remaining text messages in the plurality of text messages are the plurality of candidate text messages.
[0160] In an embodiment of the present invention, the server may segment multiple text messages using a mechanical segmentation algorithm to obtain multiple keywords corresponding to each text message. The server may also segment each text message using a Markov model segmentation algorithm to obtain multiple keywords corresponding to each text message. In other embodiments, the server may also use other segmentation algorithms to segment the identification information into corresponding keywords, which is not limited in the embodiment of the present invention.
[0161] Among them, multiple keywords may include: nouns, verbs and adjectives.
[0162] S112. If the multidimensional data includes four of the following: part-of-speech, sentiment information, portrait, and evaluation classification information, then based on the parts of speech of the multiple keywords, multiple first-level ironic text information is determined from the multiple text information to be selected, and the multiple first-level ironic text information is scored to obtain multiple first-level scores corresponding to the multiple text information to be selected.
[0163] In an embodiment of the present invention, if the multidimensional data includes four of the following: part-of-speech category, sentiment information category, image category, and evaluation category information, the server determines multiple first-level ironic text information with opposite keyword polarity in the multiple text information. The server adds points to the multiple first-level ironic text information based on the initial scores of the multiple candidate text information, thereby obtaining multiple first-level scores corresponding to the multiple candidate text information.
[0164] The opposite keyword polarity indicates that at least one pair of keywords appear in the same text information, one of which has a positive polarity and the other has a negative polarity.
[0165] For example, the initial scores of the plurality of text information to be selected may be 0 points.
[0166] For example, combined Figure 7 , which will be explained in combination with the steps.
[0167] S207, word segmentation + sentiment part-of-speech recognition.
[0168] In the embodiment of the present invention, the server first performs word segmentation on the plurality of candidate text information to obtain a plurality of keywords, and then performs sentiment part-of-speech recognition on the plurality of keywords to obtain the parts of speech of the plurality of keywords.
[0169] S208: Whether the word pair contains emotional conflict.
[0170] In the embodiment of the present invention, the server further identifies whether the keywords corresponding to the plurality of candidate text information include word pairs corresponding to sentiment conflicts. The server performs step S209 and confidence scoring on the word pairs with sentiment part-of-speech conflicts to obtain a plurality of first-level scores corresponding to the plurality of candidate text information.
[0171] S113. Based on the emotional information of the multiple text information to be selected and the emotional information of the corresponding context information, determine multiple secondary ironic text information from the multiple text information to be selected, and add the primary scores corresponding to the multiple secondary ironic text information to obtain multiple secondary scores corresponding to the multiple text information to be selected.
[0172] In an embodiment of the present invention, a server identifies multiple secondary ironic text messages from a plurality of candidate text messages, each having positive sentiment information and corresponding contextual information having negative sentiment information. The server then adds points to the multiple secondary ironic text messages based on the primary scores of the multiple candidate text messages, thereby obtaining multiple secondary scores corresponding to the multiple candidate text messages.
[0173] The context information is the text information adjacent to the selected text information in the conversation text of the target object.
[0174] In an embodiment of the present invention, the server may input the candidate text information into an emotion recognition model to obtain emotion information corresponding to the plurality of candidate text information.
[0175] Combine Figure 7 , each step will be explained.
[0176] S210, context sentence emotion recognition.
[0177] In an embodiment of the present invention, the server performs context emotion recognition on multiple candidate text information to determine whether the contexts corresponding to the multiple candidate text information are negative emotions and whether the corresponding candidate text information are positive emotions.
[0178] In the embodiment of the present invention, S211: whether the context has a negative sentiment.
[0179] The server performs step S212 and confidence scoring on multiple secondary ironic text messages whose corresponding emotional information is positive and whose corresponding context is negative in the multiple text messages to be selected, and obtains multiple third-level scores corresponding to the multiple text messages to be selected.
[0180] S114. Based on the multiple associated portraits corresponding to the target objects of the multiple text information to be selected, determine multiple third-level ironic text information from the multiple text information to be selected, and add the second-level scores corresponding to the multiple third-level ironic text information to obtain multiple third-level scores corresponding to the multiple text information to be selected.
[0181] In an embodiment of the present invention, a server determines associated profiles of target objects corresponding to a plurality of candidate text messages. The server determines that the associated profiles represent a plurality of third-level ironic text messages corresponding to a user prone to irony. The server then adds points to the plurality of third-level ironic text messages based on the plurality of second-level scores of the plurality of candidate text messages, thereby obtaining a plurality of third-level scores corresponding to the plurality of candidate text messages.
[0182] Combine Figure 7 , which will be explained in combination with the steps.
[0183] S213. Obtain user portrait.
[0184] S214: Whether the user is prone to irony.
[0185] In an embodiment of the present invention, the server identifies whether multiple target objects are prone to irony based on the associated profiles corresponding to multiple targets corresponding to the multiple candidate text information. The server then determines multiple third-level ironic text information corresponding to the prone to irony among the multiple target objects. The server then performs S215 confidence scoring on the multiple third-level ironic text information based on the multiple second-level scores of the multiple candidate text information, thereby obtaining multiple third-level scores corresponding to the multiple candidate text information.
[0186] S115. Based on the multiple evaluation scores corresponding to the multiple text information to be selected and the emotional information of the multiple text information to be selected, determine multiple fourth-level ironic text information from the multiple text information to be selected, and add the third-level scores corresponding to the multiple fourth-level ironic text information to obtain multiple fourth-level scores corresponding to the multiple text information to be selected.
[0187] In an embodiment of the present invention, a server identifies, from among a plurality of candidate text messages, a plurality of fourth-level ironic text messages having corresponding low evaluation scores and positive sentiment information. The server then adds points to the plurality of fourth-level ironic text messages based on the plurality of third-level scores corresponding to the plurality of candidate text messages, thereby obtaining a plurality of fourth-level scores corresponding to the plurality of candidate text messages.
[0188] Combine Figure 7 , which will be explained in combination with the steps.
[0189] S216. Obtain evaluation scores.
[0190] S217: Is the review negative and the sentiment positive?
[0191] In this embodiment of the present invention, a server obtains evaluation scores corresponding to multiple candidate text messages. The server identifies whether the multiple candidate text messages are level 4 ironic text messages with negative reviews and positive sentiment. The server performs S218 confidence scoring on the multiple level 4 ironic text messages based on the multiple level 3 scores corresponding to the multiple candidate text messages, thereby obtaining multiple level 4 scores corresponding to the multiple candidate text messages.
[0192] S219. Comprehensive judgment score.
[0193] S220: Is the score higher than the threshold?
[0194] S221, multiple ironic text messages
[0195] In an embodiment of the present invention, a server obtains comprehensive judgment scores corresponding to a plurality of candidate text messages, i.e., a plurality of four-level scores. The server determines whether the plurality of four-level scores are above a threshold. If so, the server classifies the candidate text messages corresponding to the scores above the threshold into a plurality of ironic text messages.
[0196] In the embodiment of the present invention, combined with Figure 8 The server determines the ironic text information among the multiple candidate text information based on the intra-sentence emotional conflict information, context emotional conflict information, user language structure habit information and text score conflict information corresponding to the multiple candidate text information.
[0197] In some embodiments, see Figure 9 , Figure 9 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 6 The illustrated S112 can also be implemented through S116 to S118 , which will be described in conjunction with each step.
[0198] S116. Find the parts of speech of multiple keywords in the sentiment word list.
[0199] In the embodiment of the present invention, the server searches for parts of speech corresponding to multiple keywords in the sentiment word list, wherein the parts of speech can be positive or negative.
[0200] Among them, the sentiment word list is a preset list of parts of speech that includes all keywords.
[0201] S117 . Determine, from the plurality of text information to be selected, a plurality of first-level ironic text information including keywords with opposite parts of speech.
[0202] In the embodiment of the present invention, the server determines a plurality of first-level ironic text information having opposite parts of speech among the corresponding plurality of keywords.
[0203] For example, the multiple keywords may include: "like", "only", and "damn". Among them, "like" and "damn" have opposite polarities, and the server may determine that the candidate text information corresponding to the multiple keywords is first-level ironic text information.
[0204] S118 , adding points to the multiple first-level ironic text information based on the initial scores corresponding to the multiple text information to be selected, to obtain multiple first-level scores corresponding to the multiple text information to be selected.
[0205] In the embodiment of the present invention, the plurality of text messages to be selected have a plurality of initial scores corresponding to them. The server adds points to the plurality of first-level ironic text messages in the plurality of text messages to be selected based on the initial scores to obtain a plurality of first-level scores corresponding to the plurality of text messages to be selected.
[0206] The additional points may be 1 point or 10 points, which is not limited here.
[0207] In some embodiments, see Figure 10 , Figure 10 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 6 The illustrated S113 can also be implemented through S119 to S122 , which will be described in conjunction with each step.
[0208] S119: Extracting multiple context information corresponding to multiple text information to be selected from the training text data set.
[0209] In the embodiment of the present invention, each text information to be selected includes adjacent text information, that is, context information. The server extracts multiple pieces of contextual text information corresponding to the multiple text information to be selected from the training text data set.
[0210] S120: Detect and obtain emotional information of multiple context information and emotional information of multiple candidate text information.
[0211] In the embodiment of the present invention, the server may obtain the emotion information of the plurality of photogenic text information and the emotion information of the plurality of candidate text information through emotion recognition model processing.
[0212] In the embodiment of the present invention, the server can search for the emotional information of multiple context information and the emotional information of multiple candidate text information through the emotional word list.
[0213] S121. Determine, from among the plurality of text information to be selected, a plurality of secondary ironic text information whose emotional information is positive and whose corresponding contextual information has negative emotional information.
[0214] In an embodiment of the present invention, the server determines, among a plurality of candidate text messages, a plurality of secondary ironic text messages whose sentiment information is positive and whose corresponding context information is negative.
[0215] S122 , adding points to the multiple secondary ironic text information on the primary scores corresponding to the multiple text information to be selected, to obtain multiple secondary scores corresponding to the multiple text information to be selected.
[0216] In the embodiment of the present invention, the server adds points to the first-level scores of the second-level ironic text information in the plurality of text information to be selected, and obtains a plurality of second-level scores corresponding to the plurality of text information to be selected.
[0217] In some embodiments, see Figure 11 , Figure 11 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 6 The illustrated S114 can also be implemented through S123 to S125 , which will be described in conjunction with each step.
[0218] S123. Extract multiple associated portraits corresponding to target objects corresponding to multiple pieces of text information to be selected from the multiple portraits.
[0219] In the embodiment of the present invention, the server determines multiple associated portraits of target objects corresponding to multiple pieces of text information to be selected from the multiple portraits.
[0220] S124 , determining, from among the multiple text information to be selected, multiple third-level ironic text information whose target objects represented by the corresponding associated portraits are users prone to irony.
[0221] In the embodiment of the present invention, since each associated profile indicates whether the corresponding target object is a user prone to irony, the server determines a plurality of third-level ironic text messages from the plurality of candidate text messages. The target objects corresponding to the plurality of third-level ironic text messages are users prone to irony.
[0222] S125 , adding points to the multiple third-level ironic text information on the second-level scores corresponding to the multiple text information to be selected, to obtain multiple third-level scores corresponding to the multiple text information to be selected.
[0223] In the embodiment of the present invention, the server adds points to the second-level scores of the third-level ironic text messages in the plurality of text messages to be selected, and obtains a plurality of third-level scores corresponding to the plurality of text messages to be selected.
[0224] In some embodiments, see Figure 12 , Figure 12 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 6 The illustrated S115 can also be implemented through S126 to S128 , which will be described in conjunction with each step.
[0225] S126 . Extracting multiple evaluation scores corresponding to multiple pieces of text information to be selected from the training text data set.
[0226] In the embodiment of the present invention, the server extracts multiple evaluation scores corresponding to multiple pieces of text information to be selected from the training text data set.
[0227] S127. Determine, from among the plurality of text information to be selected, a plurality of level four ironic text information whose sentiment information is positive and whose corresponding evaluation scores are lower than a first threshold.
[0228] In the embodiment of the present invention, the server determines, from a plurality of candidate text messages, a plurality of fourth-level ironic text messages whose sentiment information is positive and whose corresponding evaluation scores are lower than a first threshold.
[0229] S128 , adding points to the multiple fourth-level ironic text information based on the third-level scores corresponding to the multiple text information to be selected, to obtain multiple fourth-level scores corresponding to the multiple text information to be selected.
[0230] In the embodiment of the present invention, the server adds points to the third-level scores for the fourth-level ironic text messages among the plurality of text messages to be selected, and obtains a plurality of fourth-level scores corresponding to the plurality of text messages to be selected.
[0231] In some embodiments, see Figure 13 , Figure 13 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 10 The illustrated S120 can also be implemented through S129 , which will be described in conjunction with each step.
[0232] S129: Input the multiple context information and the multiple text information to be selected into the emotion recognition model to obtain the emotion information of the multiple context information and the emotion information of the multiple text information to be selected.
[0233] In the embodiment of the present invention, the server inputs multiple context information and multiple text information to be selected into the emotion recognition model respectively, and can obtain the emotion information of the multiple context information and the emotion information of the multiple text information to be selected.
[0234] In some embodiments, see Figure 14 , Figure 14 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 10The illustrated S120 may also be implemented through S130 to S132 , which will be described in conjunction with each step.
[0235] S130: Segment the multiple context information into multiple upper and lower keywords.
[0236] S131. Find the parts of speech of multiple upper and lower keywords in the sentiment word list, and find the parts of speech of multiple keywords in the sentiment word list.
[0237] S132. Determine the sentiment information of the multiple contextual information and the sentiment information of the multiple candidate text information based on the parts of speech of the multiple upper and lower keywords and the parts of speech of the multiple keywords.
[0238] In the embodiment of the present invention, the server determines the part of speech of each keyword in the multiple context keywords, and if the corresponding context information includes a keyword with a positive part of speech, then the context information is positive in sentiment information. Accordingly, the server can determine the sentiment information of the corresponding selected text information.
[0239] In some embodiments, see Figure 15 , Figure 15 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 6 The step S111 shown includes S133 before the step S133 , which will be described in conjunction with each step.
[0240] S133: Input the multiple text information into the previous text recognition model to obtain multiple five-level ironic text information from the multiple text information.
[0241] In the embodiment of the present invention, the server inputs a plurality of candidate text information into the previous text recognition model, and obtains a plurality of five-level ironic text information from the plurality of text information.
[0242] In some embodiments, see Figure 15 , Figure 15 An optional flow chart of a data detection method provided in an embodiment of the present invention is provided. Figure 6 The illustrated S110 can be implemented through S134 to S135 , which will be described in conjunction with each step.
[0243] S134 . Determine, from among the plurality of text messages to be selected, a plurality of level 6 ironic text messages whose corresponding level 4 scores are higher than a second threshold.
[0244] In the embodiment of the present invention, the server determines, from a plurality of candidate text messages, a plurality of level 6 ironic text messages whose corresponding level 4 scores are higher than a second threshold.
[0245] In the embodiment of the present invention, the server may further determine, from among the plurality of text messages to be selected, the top m corresponding fourth-level scores as a plurality of sixth-level ironic text messages.
[0246] S135 . Combining the plurality of level five ironic text information and the plurality of level six ironic text information to form a plurality of ironic text information.
[0247] In the embodiment of the present invention, the server combines a plurality of level five ironic text messages and a plurality of level six ironic text messages into a plurality of ironic text messages.
[0248] In some embodiments, see Figure 16 , Figure 16 An optional flow chart of a data detection method provided by an embodiment of the present invention, Figure 4 The illustrated S106 can be implemented through S136 to S138 , which will be described in conjunction with each step.
[0249] S136 , inputting multiple text information in the training text data set into the emotion recognition model to obtain emotional text information corresponding to multiple target objects corresponding to the training text data set.
[0250] In an embodiment of the present invention, a server extracts multiple text messages corresponding to multiple target objects from a training text dataset. The server inputs the multiple text messages corresponding to the multiple target objects into an emotion recognition model. Thus, emotional text messages corresponding to the multiple target objects are obtained.
[0251] In the embodiment of the present invention, combined with Figure 17 The server needs to obtain data such as user sarcasm, emotional sentences, and conversation service evaluations. Sarcasm is simply multiple pieces of ironic text. Using a sentiment recognition model, the server can obtain emotional sentences from multiple target users, and then extract conversation service evaluation information for each of these sarcasm and emotional sentences.
[0252] S137. Extract the number of evaluations and the number of negative reviews corresponding to the multiple target objects in the training text data set.
[0253] In the embodiment of the present invention, the server extracts the number of positive reviews and the number of negative reviews corresponding to a plurality of target objects from the training text data set.
[0254] S138. Construct multiple portraits based on multiple supplementary ironic text information, multiple ironic text information, emotional information, and negative review quantity information.
[0255] In an embodiment of the present invention, the server constructs multiple portraits based on multiple supplementary ironic text information, multiple ironic text information, emotional information, and information on the number of negative reviews.
[0256] In the embodiment of the present invention, combined with Figure 17 The server obtains the sarcasm score, sentiment score, and negative impact score for each of the multiple target objects. Based on the sarcasm score, sentiment score, and negative impact score, the server generates a comprehensive score for the multiple target objects. The server then performs a weighted ranking of the multiple comprehensive scores to identify users prone to sarcasm among the multiple target objects. The sarcasm score includes the percentage of ironic conversations. The sentiment score includes the percentage of complaining conversations and the percentage of angry conversations. The negative impact score includes the percentage of conversations with negative reviews and the percentage of negative service reviews.
[0257] In some embodiments, see Figure 18 , Figure 18 An optional flow chart of a data detection method provided by an embodiment of the present invention, Figure 16 The illustrated S138 can be implemented through S139 to S142 , which will be described in conjunction with each step.
[0258] S139: Compare the sum of the supplementary ironic text information and the ironic text information corresponding to the multiple target objects with the total number of corresponding text information to obtain irony scores corresponding to the multiple target objects.
[0259] In the embodiment of the present invention, the server divides the sum of the supplementary ironic text information and the ironic text information corresponding to the plurality of target objects by the total amount of text information of the target object to obtain the irony score of the corresponding target object.
[0260] S140 : Compare the emotional text information corresponding to the multiple target objects with the total number of corresponding text information to obtain emotional scores corresponding to the multiple target objects.
[0261] In the embodiment of the present invention, the server compares the emotional text information corresponding to the multiple target objects with the total number of corresponding text information to obtain the emotional scores corresponding to the multiple target objects.
[0262] S141. Divide the number of negative reviews corresponding to the multiple target objects by the number of corresponding reviews to obtain negative impact scores corresponding to the multiple target objects.
[0263] In the embodiment of the present invention, the server divides the number of negative reviews corresponding to the multiple target objects by the number of corresponding evaluations to obtain negative impact scores corresponding to the multiple target objects.
[0264] S142. Calculate multiple portraits based on the irony scores, emotion scores, and negative impact scores corresponding to the multiple target objects.
[0265] In some embodiments, see Figure 19 , Figure 19 An optional flow chart of a data detection method provided by an embodiment of the present invention, Figure 18The illustrated S142 can be implemented through S143 to S144 , which will be described in conjunction with each step.
[0266] S143. Add the product of the irony scores and the irony weights corresponding to the multiple target objects, the product of the emotion scores and the emotion weights, and the product of the negative impact scores and the negative impact weights to obtain multiple comprehensive scores corresponding to the multiple target objects.
[0267] In an embodiment of the present invention, the server adds the product of the irony scores and the irony weights corresponding to the multiple target objects, the product of the emotion scores and the emotion weights, and the product of the negative influence scores and the negative influence weights to obtain multiple comprehensive scores corresponding to the multiple target objects.
[0268] Among them, the irony weight can be any value, the emotion weight can be any value, and the negative impact weight can be any value.
[0269] S144. Form a correspondence between multiple comprehensive scores and multiple target objects, and then form multiple portraits.
[0270] In the embodiment of the present invention, the server establishes a correspondence between multiple target objects and corresponding comprehensive scores, and then determines multiple portraits corresponding to the multiple target objects.
[0271] In some embodiments, see Figure 20 , Figure 20 An optional flow chart of a data detection method provided by an embodiment of the present invention, Figure 19 S144 shown includes S145 to S146, which will be explained in conjunction with each step.
[0272] S145 , determining that the target objects corresponding to the first proportion with the largest comprehensive score among the multiple target objects are easy-to-ironic objects.
[0273] In the embodiment of the present invention, the server determines that the target objects corresponding to the first proportion with the largest comprehensive scores among the plurality of target objects are ironic objects.
[0274] The first ratio may be 10 percent.
[0275] S146 , determining that the target objects corresponding to the first second proportion with the smallest comprehensive score among the multiple target objects are non-ironic objects.
[0276] In the embodiment of the present invention, the server determines that the target objects corresponding to the first second proportions of the target objects with the smallest comprehensive scores are non-ironic objects.
[0277] The second ratio may also be 10 percent.
[0278] See also Figure 21, Figure 21 A schematic structural diagram of a data detection device provided by an embodiment of the present invention.
[0279] The embodiment of the present invention further provides a data detection device 800 , including: an acquisition unit 803 and a detection unit 804 .
[0280] The acquisition unit 803 is used to acquire the current text information to be inspected;
[0281] The detection unit 804 is configured to detect the current text information to be detected using the current text recognition model and determine a detection result of the current text information to be detected;
[0282] The current text recognition model is obtained by obtaining a training text data set, extracting multiple ironic text information from the training text data set in combination with multidimensional data, and cyclically training the previous text recognition model using the multiple ironic text information; the multidimensional data includes: at least two of: part of speech class, sentiment information class, portrait class and evaluation classification information.
[0283] In an embodiment of the present invention, the current text recognition model is based on the portrait of the target object corresponding to the training text data set containing the context, and the evaluation classification information of the conversation text information in the training text data set, and the multiple ironic text information is determined in the training text data set, and the previous text recognition model is cyclically trained using the multiple ironic text information.
[0284] In an embodiment of the present invention, the acquisition unit 803 in the data detection device 800 is used to acquire a training text data set, and extract multiple ironic text information from the training text data set in combination with multidimensional data; the training text data set includes conversation text information and related information of multiple target objects, wherein each text information represents a single-scene conversation text information of the corresponding target object; each portrait of the multiple portraits represents the ironic text score information corresponding to the corresponding target object; and the previous text recognition model is trained based on the multiple ironic text information until the current text recognition model is obtained.
[0285] In the embodiment of the present invention, the data detection device 800 is used to extract a plurality of supplementary ironic text information from the training text data set again using the current text recognition model;
[0286] Combining the plurality of supplementary ironic text information and the plurality of ironic text information, constructing a plurality of portraits of a plurality of target objects corresponding to the training text data set;
[0287] Based on the multiple portraits, multiple ironic text information is extracted from the training text data set to be obtained next time, so as to train the current text recognition model to obtain a next text recognition model.
[0288] In an embodiment of the present invention, the acquisition unit 803 in the data detection device 800 is used to obtain a training text data set, segment the text in the training text set to obtain multiple text information; combine at least two of the part-of-speech class, emotional information class, portrait class and evaluation classification information to perform text recognition on the multiple text information to obtain multiple scores corresponding to the multiple text information; based on the multiple scores, determine multiple ironic text information in the multiple text information.
[0289] In an embodiment of the present invention, the data detection device 800 is used to perform word segmentation processing on multiple text information to be selected from the multiple text information to obtain multiple keywords corresponding to the multiple text information to be selected; the multiple text information to be selected is the text information remaining after the multiple text information is filtered by the previous text recognition model; if the multidimensional data includes: four of the part-of-speech class, the emotional information class, the portrait class and the evaluation classification information, then based on the part-of-speech of the multiple keywords, multiple first-level ironic text information is determined in the multiple text information to be selected, and the multiple first-level ironic text information is scored to obtain multiple first-level scores corresponding to the multiple text information to be selected; based on the emotional information of the multiple text information to be selected and the emotional information of the corresponding context information, the multiple first-level ironic text information is scored in the multiple text information to be selected A plurality of second-level ironic text information is determined in this information, and the first-level scores corresponding to the plurality of second-level ironic text information are added to obtain a plurality of second-level scores corresponding to the plurality of text information to be selected; based on a plurality of associated portraits corresponding to the target objects of the plurality of text information to be selected, a plurality of third-level ironic text information is determined in the plurality of text information to be selected, and the second-level scores corresponding to the plurality of third-level ironic text information are added to obtain a plurality of third-level scores corresponding to the plurality of text information to be selected; based on a plurality of evaluation scores corresponding to the plurality of text information to be selected and the emotional information of the plurality of text information to be selected, a plurality of fourth-level ironic text information is determined in the plurality of text information to be selected, and the third-level scores corresponding to the plurality of fourth-level ironic text information are added to obtain a plurality of fourth-level scores corresponding to the plurality of text information to be selected.
[0290] In an embodiment of the present invention, the data detection device 800 is used to find the parts of speech of the multiple keywords in the sentiment word list; the sentiment word list is a preset list of parts of speech including all keywords; determine multiple first-level ironic text information with opposite parts of speech of the keywords in the multiple text information to be selected; add points to the multiple first-level ironic text information on the initial scores corresponding to the multiple text information to be selected, and obtain multiple first-level scores corresponding to the multiple text information to be selected.
[0291] In an embodiment of the present invention, the data detection device 800 is used to extract multiple contextual information corresponding to the multiple text information to be selected from the training text data set; detect and obtain emotional information of the multiple contextual information and emotional information of the multiple text information to be selected; determine, among the multiple text information to be selected, multiple secondary ironic text information whose emotional information is positive and whose corresponding contextual information is negative; add points to the multiple secondary ironic text information on the first-level scores corresponding to the multiple text information to obtain multiple secondary scores corresponding to the multiple text information to be selected.
[0292] In an embodiment of the present invention, the data detection device 800 is used to extract multiple associated portraits corresponding to the target objects corresponding to the multiple text information to be selected from the multiple portraits; determine that the target objects represented by the corresponding associated portraits are multiple third-level ironic text information of users who are prone to irony in the multiple text information to be selected; add points to the multiple third-level ironic text information on the secondary scores corresponding to the multiple text information to be selected, and obtain multiple third-level scores corresponding to the multiple text information to be selected.
[0293] In an embodiment of the present invention, the data detection device 800 is used to extract multiple evaluation scores corresponding to the multiple text information to be selected from the training text data set; determine multiple fourth-level ironic text information whose emotional information is positive and whose corresponding evaluation scores are lower than a first threshold value among the multiple text information to be selected; and add points to the multiple fourth-level ironic text information on the third-level scores corresponding to the multiple text information to obtain multiple fourth-level scores corresponding to the multiple text information to be selected.
[0294] In an embodiment of the present invention, the data detection device 800 is used to input the multiple contextual information and the multiple text information to be selected into a sentiment recognition model to obtain the sentiment information of the multiple contextual information and the sentiment information of the multiple text information to be selected; segment the multiple contextual information into multiple contextual keywords; find the parts of speech of the multiple contextual keywords in the sentiment word list, and find the parts of speech of the multiple keywords in the sentiment word list; determine the sentiment information of the multiple contextual information and the sentiment information of the multiple text information to be selected based on the parts of speech of the multiple contextual keywords and the parts of speech of the multiple keywords.
[0295] In the embodiment of the present invention, the data detection device 80 is configured to input the plurality of text information into the previous text recognition model to obtain a plurality of five-level ironic text information from the plurality of text information.
[0296] In an embodiment of the present invention, the data detection device 800 is configured to determine, from the plurality of candidate text information, a plurality of level 6 ironic text information whose corresponding level 4 scores are higher than a second threshold; and combine the plurality of level 5 ironic text information with the plurality of level 6 ironic text information to form the plurality of ironic text information.
[0297] In an embodiment of the present invention, the data detection device 800 is used to input multiple text information in the training text data set into an emotion recognition model to obtain emotional text information corresponding to multiple target objects corresponding to the training text data set; extract the number of evaluations and the number of negative reviews corresponding to the multiple target objects in the training text data set; and construct the multiple portraits based on the multiple supplementary ironic text information, the multiple ironic text information, the emotional information and the number of negative reviews.
[0298] In an embodiment of the present invention, the data detection device 800 is used to divide the sum of the supplementary ironic text information and the ironic text information corresponding to the multiple target objects by the total number of corresponding text information to obtain the irony scores corresponding to the multiple target objects respectively; divide the emotional text information corresponding to the multiple target objects by the total number of corresponding text information to obtain the emotion scores corresponding to the multiple target objects respectively; divide the number of negative reviews corresponding to the multiple target objects by the corresponding number of evaluations to obtain the negative impact scores corresponding to the multiple target objects respectively; and calculate the multiple portraits based on the irony scores, emotion scores, and negative impact scores corresponding to the multiple target objects respectively.
[0299] In an embodiment of the present invention, the data detection device 800 is used to add the product of the irony scores and the irony weights corresponding to the multiple target objects, the product of the emotion scores and the emotion weights, and the product of the negative influence scores and the negative influence weights, to obtain multiple comprehensive scores corresponding to the multiple target objects; form a corresponding relationship between the multiple comprehensive scores and the multiple target objects, and then form the multiple portraits.
[0300] In an embodiment of the present invention, the data detection device 800 is used to determine that the target objects corresponding to the first proportion with the largest comprehensive scores of the multiple target objects are easy to be ironic objects; and determine that the target objects corresponding to the second proportion with the smallest comprehensive scores of the multiple target objects are non-ironic objects.
[0301] In an embodiment of the present invention, the current text information to be inspected is obtained by an acquisition unit 803; the detection unit 804 detects the current text information to be inspected using a current text recognition model to determine a detection result of the current text information to be inspected; wherein the current text recognition model is obtained by obtaining a training text data set, extracting multiple ironic text information from the training text data set in combination with multidimensional data, and cyclically training a previous text recognition model using the multiple ironic text information; the multidimensional data represents at least two of the classification information of part of speech, sentiment information, portrait, and evaluation. Because the multiple ironic text information is obtained by extracting the training text data set in combination with the multidimensional data, and because the multidimensional data takes into account multiple factors in the ironic text information, the previous text recognition model can learn a large amount of comprehensive ironic text information. Then, the current text recognition model, which is trained with a large amount of comprehensive ironic text information, can accurately identify ironic text information in group conversations, thereby improving the recognition efficiency of ironic text in group conversations.
[0302] It should be noted that, in the embodiment of the present invention, if the above-mentioned data detection method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a data detection device (which can be a personal computer, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present invention is not limited to any specific combination of hardware and software.
[0303] Correspondingly, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above method when executed by a processor.
[0304] Correspondingly, an embodiment of the present invention provides a data detection device 800, including a memory 802 and a processor 801, wherein the memory 802 stores a computer program that can be run on the processor 801, and the processor 801 implements the steps in the above method when executing the program.
[0305] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of the present invention, please refer to the description of the method embodiments of the present invention for understanding.
[0306] It should be noted that Figure 22 A schematic diagram of a hardware entity of a data detection device provided by an embodiment of the present invention, such as Figure 22 As shown, the hardware entity of the data detection device 800 includes: a processor 801 and a memory 802, wherein;
[0307] The processor 801 generally controls the overall operation of the data detection apparatus 800 .
[0308] The memory 802 is configured to store instructions and applications executable by the processor 801, and can also cache data to be processed or processed by the processor 801 and each module in the data detection device 800 (for example, image data, audio data, voice communication data and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0309] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention. The serial numbers of the above-mentioned embodiments of the present invention are for description only and do not represent the advantages and disadvantages of the embodiments.
[0310] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0311] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0312] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0313] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0314] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.
[0315] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0316] The above description is merely an embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A data detection method, characterized in that: include: Get the current text information to be checked; Acquire a training text data set, segment the text in the training text data set to obtain a plurality of text information; the training text data set includes conversation text information and related information of a plurality of target objects, wherein each text information represents a single scene conversation text information of a corresponding target object; the target object is a user; Combining at least two of the classified information of part of speech, sentiment information, portrait, and evaluation, text recognition is performed on the plurality of text information to obtain a plurality of scores corresponding to the plurality of text information; the portrait information includes a portrait of the target object, and the portrait of the target object represents the ironic text score information corresponding to the target object; determining, based on the plurality of scores, a plurality of ironic text messages having scores higher than a threshold value among the plurality of text messages; cyclically training the previous text recognition model based on the plurality of ironic text information until a current text recognition model is obtained; The current text information to be detected is detected using the current text recognition model to determine a detection result of the current text information to be detected.
2. The data detection method according to claim 1, characterized in that: The current text recognition model is based on the portrait of the target object corresponding to the training text data set containing the context, and the evaluation classification information of the conversation text information in the training text data set, and is obtained by determining the multiple ironic text information in the training text data set and cyclically training the previous text recognition model through the multiple ironic text information.
3. The data detection method according to any one of claims 1 or 2, characterized in that: The method further comprises: extracting again a plurality of supplementary ironic text information from the training text data set using the current text recognition model; Combining the plurality of supplementary ironic text information with the plurality of ironic text information, constructing a plurality of portraits of a plurality of target objects corresponding to the training text data set; each of the plurality of portraits represents ironic text score information corresponding to the corresponding target object; Based on the multiple portraits, multiple ironic text information is extracted from the training text data set to be obtained next time, so as to train the current text recognition model to obtain a next text recognition model.
4. The data detection method according to claim 1, characterized in that: The plurality of scores include: a plurality of four-level scores; The combining of at least two of the part-of-speech class, the sentiment information class, the portrait class, and the evaluation classification information to perform text recognition on the plurality of text information to obtain a plurality of scores corresponding to the plurality of text information includes: Performing word segmentation on multiple pieces of text information to be selected from the multiple pieces of text information to obtain multiple keywords corresponding to the multiple pieces of text information; the multiple pieces of text information to be selected are the text information remaining after the multiple pieces of text information are filtered by the previous text recognition model; If the multidimensional data includes four of the part-of-speech class, the sentiment information class, the portrait class, and the evaluation classification information, then based on the parts of speech of the multiple keywords, a plurality of first-level ironic text information is determined from the multiple candidate text information, and scores are added to the multiple first-level ironic text information to obtain a plurality of first-level scores corresponding to the multiple candidate text information; Based on the sentiment information of the plurality of to-be-selected text information and the sentiment information of the corresponding context information, determining a plurality of secondary ironic text information from the plurality of to-be-selected text information, and adding the primary scores corresponding to the plurality of secondary ironic text information to obtain a plurality of secondary scores corresponding to the plurality of to-be-selected text information; Based on the multiple associated portraits corresponding to the target objects of the multiple text information to be selected, determining multiple third-level ironic text information from the multiple text information to be selected, and adding the second-level scores corresponding to the multiple third-level ironic text information to obtain multiple third-level scores corresponding to the multiple text information to be selected; Based on the multiple evaluation scores corresponding to the multiple text information to be selected and the emotional information of the multiple text information to be selected, multiple four-level ironic text information are determined from the multiple text information to be selected, and the three-level scores corresponding to the multiple four-level ironic text information are added to obtain multiple four-level scores corresponding to the multiple text information to be selected.
5. The data detection method according to claim 4, characterized in that: The step of determining a plurality of first-level ironic text information from the plurality of text information to be selected based on the parts of speech of the plurality of keywords, and adding points to the plurality of first-level ironic text information to obtain a plurality of first-level scores corresponding to the plurality of text information to be selected includes: Finding the parts of speech of the multiple keywords in the sentiment word list; the sentiment word list is a preset list of parts of speech including all keywords; Determining, from the plurality of text information to be selected, a plurality of first-level ironic text information including keywords with opposite parts of speech; The plurality of first-level ironic text information are scored based on the initial scores corresponding to the plurality of text information to be selected, to obtain a plurality of first-level scores corresponding to the plurality of text information to be selected.
6. The data detection method according to claim 4 or 5, characterized in that: The method of determining a plurality of secondary ironic text information from the plurality of text information to be selected based on the emotional information of the plurality of text information to be selected and the emotional information of the corresponding context information, and adding the primary scores corresponding to the plurality of secondary ironic text information to obtain a plurality of secondary scores corresponding to the plurality of text information to be selected, includes: Extracting multiple context information corresponding to the multiple text information to be selected from the training text data set; Detecting and obtaining emotional information of the plurality of contextual information and emotional information of the plurality of selected text information; Determining, from the plurality of text information to be selected, a plurality of secondary ironic text information whose sentiment information is positive and whose corresponding contextual information has negative sentiment information; The plurality of secondary ironic text information are added to the primary scores corresponding to the plurality of text information to be selected, to obtain a plurality of secondary scores corresponding to the plurality of text information to be selected.
7. The data detection method according to claim 4 or 5, characterized in that: The method of determining a plurality of third-level ironic text information from the plurality of text information to be selected based on the plurality of associated portraits corresponding to the target objects of the plurality of text information to be selected, and adding the second-level scores corresponding to the plurality of third-level ironic text information to obtain a plurality of third-level scores corresponding to the plurality of text information to be selected includes: Extracting a plurality of associated portraits corresponding to the target objects corresponding to the plurality of text information to be selected from the plurality of portraits; Determining, from the plurality of text information to be selected, a plurality of third-level ironic text information whose target objects represented by the corresponding associated portraits are users prone to irony; The plurality of third-level ironic text information are added to the second-level scores corresponding to the plurality of text information to be selected, to obtain a plurality of third-level scores corresponding to the plurality of text information to be selected.
8. The data detection method according to claim 6, characterized in that: The method of determining a plurality of four-level ironic text information from the plurality of text information to be selected based on the plurality of evaluation scores corresponding to the plurality of text information to be selected and the sentiment information of the plurality of text information to be selected, and adding the three-level scores corresponding to the plurality of four-level ironic text information to obtain the plurality of four-level scores corresponding to the plurality of text information to be selected, includes: Extracting multiple evaluation scores corresponding to the multiple pieces of text information to be selected from the training text data set; Determining, from the plurality of text messages to be selected, a plurality of level four ironic text messages having positive sentiment information and corresponding evaluation scores lower than a first threshold; The plurality of fourth-level ironic text information are added to the third-level scores corresponding to the plurality of text information to be selected, so as to obtain a plurality of fourth-level scores corresponding to the plurality of text information to be selected.
9. The data detection method according to claim 6, characterized in that: The detecting and obtaining the emotional information of the plurality of contextual information and the emotional information of the plurality of selected text information includes one of the following: Inputting the plurality of context information and the plurality of to-be-selected text information into an emotion recognition model to obtain emotion information of the plurality of context information and emotion information of the plurality of to-be-selected text information; Segmenting the multiple context information into multiple upper and lower keywords; Find the parts of speech of the multiple upper and lower keywords in the sentiment word list, and find the parts of speech of the multiple keywords in the sentiment word list; The sentiment information of the plurality of contextual text information and the sentiment information of the plurality of candidate text information are determined based on the parts of speech of the plurality of contextual keywords and the parts of speech of the plurality of keywords.
10. The data detection method according to any one of claims 4-5, 8-9, characterized in that: Before performing word segmentation processing on the plurality of candidate text information of the plurality of text information to obtain a plurality of keywords respectively corresponding to the plurality of candidate text information, the method further includes: The plurality of text information are input into the previous text recognition model to obtain a plurality of five-level ironic text information from the plurality of text information.
11. The data detection method according to claim 10, characterized in that: The step of determining a plurality of ironic text messages from the plurality of text messages based on the plurality of scores includes: Determining, from the plurality of text messages to be selected, a plurality of level 6 ironic text messages whose corresponding level 4 scores are higher than a second threshold; The plurality of fifth-level ironic text messages and the plurality of sixth-level ironic text messages are combined to form the plurality of ironic text messages.
12. The data detection method according to claim 3, characterized in that: The step of combining the plurality of supplementary ironic text information and the plurality of ironic text information to construct a plurality of portraits of the plurality of target objects corresponding to the training text data set includes: Inputting multiple text information in the training text data set into the emotion recognition model to obtain emotional text information corresponding to multiple target objects corresponding to the training text data set; Extracting the number of evaluations and the number of negative reviews corresponding to the multiple target objects from the training text data set; The multiple portraits are constructed based on the multiple supplementary ironic text information, the multiple ironic text information, the emotional text information and the negative review quantity information.
13. The data detection method according to claim 12, characterized in that: The constructing of the multiple portraits based on the multiple supplementary ironic text information, the multiple ironic text information, the emotional text information, and the negative review quantity information includes: Comparing the sum of the supplementary ironic text information and the ironic text information corresponding to the plurality of target objects with the total number of corresponding text information, obtaining irony scores corresponding to the plurality of target objects respectively; Comparing the emotional text information corresponding to the multiple target objects with the total number of corresponding text information to obtain emotional scores corresponding to the multiple target objects; Divide the number of negative reviews corresponding to the multiple target objects by the number of corresponding reviews to obtain negative impact scores corresponding to the multiple target objects; The multiple portraits are calculated based on the irony scores, emotion scores, and negative impact scores respectively corresponding to the multiple target objects.
14. The data detection method according to claim 13, characterized in that: The calculating the multiple portraits based on the irony scores, emotion scores, and negative impact scores respectively corresponding to the multiple target objects includes: Adding the product of the irony score and the irony weight corresponding to each of the multiple target objects, the product of the emotion score and the emotion weight, and the product of the negative influence score and the negative influence weight, to obtain multiple comprehensive scores corresponding to the multiple target objects; A correspondence between the multiple comprehensive scores and the multiple target objects is formed, thereby forming the multiple portraits.
15. The data detection method according to claim 14, characterized in that: After forming the correspondence between the multiple comprehensive scores and the multiple target objects, and then forming the multiple portraits, the method further includes: Determine that the target objects corresponding to the first proportion with the largest comprehensive score among the plurality of target objects are easy-to-ironic objects; The target objects corresponding to the first second proportion having the smallest comprehensive score among the plurality of target objects are determined to be non-ironic objects.
16. A data detection device, characterized in that: include: An acquisition unit, used to acquire the current text information to be inspected; It is also used to obtain a training text data set, segment the text in the training text data set to obtain multiple text information; combine at least two of the part-of-speech class, sentiment information class, portrait class and evaluation classification information to perform text recognition on the multiple text information to obtain multiple scores corresponding to the multiple text information; based on the multiple scores, determine multiple ironic text information with scores higher than a threshold in the multiple text information; based on the multiple ironic text information, cyclically train the previous text recognition model until the current text recognition model is obtained; the training text data set includes conversation text information and related information of multiple target objects, wherein each text information represents a single-scene conversation text information corresponding to the target object; the target object is a user; the portrait information includes a portrait of the target object, and the portrait of the target object represents the ironic text score information corresponding to the target object; The detection unit is used to detect the current text information to be detected by using the current text recognition model to determine the detection result of the current text information to be detected.
17. A data detection device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps in the method according to any one of claims 1 to 15 are implemented.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.
Citation Information
Patent Citations
Image recognition system and image recognition method
CN107690659A
Irony text collaborative recognition method, device and equipment and computer readable medium
CN111859979A