Emotion category prediction method and device

Through the emotional category prediction method and device, the data to be tested is analyzed using the emotional prediction model, the emotional category value is output, and the fraud risk is reminded based on the preset threshold, which solves the problem of insufficient accuracy of fraud prediction in the prior art and improves the ability to identify fraud behavior.

CN120180231APending Publication Date: 2025-06-20VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510343493.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prediction accuracy of fraudulent behavior in the prior art is insufficient, making it difficult to deal with increasingly complex and advanced AI fraud threats.

Method used

A emotion category prediction method and device are provided, by obtaining the data to be tested and inputting it into the emotion prediction model, outputting the emotion category values ​​of at least two emotions dimensions corresponding to the data to be tested. When the emotional category value corresponding to any emotion dimension is greater than or equal to the preset threshold, prompt information is output to remind you of possible fraud risks.

Benefits of technology

The accuracy of prediction of fraud behavior is improved. Through the use of emotional prediction models, the emotional tendencies related to fraud can be more accurately identified, thereby reducing the probability of fraud incidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180231A_ABST
    Figure CN120180231A_ABST
Patent Text Reader

Abstract

The invention discloses an emotion category prediction method and device, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring to-be-tested data; inputting to-be-tested data into the emotion prediction model, and outputting emotion category values of at least two emotion dimensions corresponding to the to-be-tested data; the emotion prediction model is obtained by training according to multiple groups of training data, and each group of training data comprises sample data and label emotion category values corresponding to the sample data; and when the emotion category value corresponding to any emotion dimension is greater than or equal to a preset threshold value, outputting prompt information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to a method and device for predicting emotion categories. Background Art

[0002] With the continuous and rapid progress of artificial intelligence (AI) technology, the quality and fidelity of synthetic audio and video have been continuously improved. Today's AI technology can already generate synthetic content that is almost indistinguishable from real human voices and real-scene videos.

[0003] The vast majority of fraud prevention and defense solutions on the market mainly focus on technical countermeasures. Taking the detection of the characteristics of forged audio and video as an example, such technical means attempt to identify synthetic audio and video content by analyzing many technical indicators such as the audio spectrum of audio and video, the details of video frame images, and the frame rate changes, and it has become difficult to cope with increasingly complex and advanced fraud threats.

[0004] Therefore, the current prediction accuracy of fraud behavior is insufficient. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a method and device for predicting emotion categories, which can solve the problem of insufficient prediction accuracy of fraud behavior.

[0006] In a first aspect, the embodiments of this application provide a method for predicting emotion categories, and the method includes:

[0007] Obtain data to be measured;

[0008] Input the data to be measured into an emotion prediction model, and output emotion category values of at least two emotion dimensions corresponding to the data to be measured; the emotion prediction model is trained based on multiple sets of training data, and each set of the training data includes: sample data and the labeled emotion category values corresponding to the sample data;

[0009] In the case where the emotion category value corresponding to any one of the emotion dimensions is greater than or equal to a preset threshold, output a prompt message.

[0010] In a second aspect, the embodiments of this application provide an emotion category prediction device, and the device includes:

[0011] An obtaining module, configured to obtain data to be measured;

[0012] An input module, configured to input the data to be measured into an emotion prediction model, and output emotion category values of at least two emotion dimensions corresponding to the data to be measured; the emotion prediction model is trained based on multiple sets of training data, and each set of the training data includes: sample data and the labeled emotion category values corresponding to the sample data;

[0013] An output module, configured to output a prompt message when the emotional category value corresponding to any one of the emotional dimensions is greater than or equal to a preset threshold.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0016] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0017] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0018] In the embodiments of the present application, by obtaining the data to be measured, inputting the data to be measured into the emotion prediction model, and outputting the emotion category values of at least two emotion dimensions corresponding to the data to be measured. Since the emotion prediction model is trained according to multiple sets of training data, each set of training data includes: sample data and the labeled emotion category value corresponding to the sample data. The emotion prediction model has learned the mapping relationship between the sample data and the labeled emotion category value during the training phase. When new data to be measured is input, the emotion prediction model calculates and outputs the emotion category values on at least two emotion dimensions corresponding to the data to be measured according to the internal algorithm of the model and the learned feature patterns. The emotion category value reflects the tendency degree of the data to be measured on different emotion dimensions. When the emotion category value corresponding to any one of the emotion dimensions is greater than or equal to a preset threshold, it indicates that the data to be measured shows a strong specific emotion tendency on this emotion dimension. This strong specific emotion tendency is very likely to be related to fraud behavior. Therefore, a prompt message is output to remind of the possible fraud risk, which can help users make more cautious decisions, thereby reducing the occurrence probability of fraud incidents. Description of the Drawings

[0019] Figure 1 is a flowchart of an emotion category prediction method provided by an embodiment of the present application;

[0020] Figure 2It is a schematic diagram of an emotional category prediction process provided by an embodiment of the present application;

[0021] Figure 3 It is a schematic diagram of a prompt message provided by an embodiment of the present application;

[0022] Figure 4 It is a structural diagram of an emotional category prediction device provided by an embodiment of the present application;

[0023] Figure 5 It is one of the schematic diagrams of the hardware structure of an electronic device provided by an embodiment of the present application;

[0024] Figure 6 It is the second schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0025] Next, the technical solutions of the embodiments of the present application will be clearly described in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0026] The terms "first", "second", etc. in the specification of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0027] The emotional category prediction method provided by the embodiments of the present application can be applied to at least the following application scenarios, which will be described below.

[0028] Currently, the vast majority of fraud prevention and defense solutions on the market mainly focus on technical countermeasures. Taking the detection of the characteristics of forged audio and video as an example, such technical means attempt to identify the synthesized audio and video content by analyzing many technical indicators such as the audio spectrum of the audio and video, the details of the video frame images, and the frame rate changes.

[0029] For example, the spectrum of normal human speech has certain natural fluctuations and characteristic distributions, while the synthetic speech may have abnormal energy distributions or unnatural frequency variations in certain frequency bands; in terms of videos, there are subtle differences between videos shot in reality and videos synthesized by AI in aspects such as light and shadow effects, object movement trajectories, and texture details of the pictures. Detection technologies are based on these differences to determine the authenticity of audio and video.

[0030] With the continuous and rapid progress of AI technology, the quality and fidelity of synthetic audio and video are constantly improving, and the difficulty of identifying the differences between synthetic and real content is also increasing. Today's AI technology can already generate synthetic content that is almost indistinguishable from real human voices and real-scene videos.

[0031] For example, some advanced speech synthesis technologies can accurately simulate the voices, intonations, speaking speeds, and emotional expressions of different people, making the synthetic speech sound vivid; video synthesis technologies have also made great breakthroughs in aspects such as image generation, scene rendering, and human motion simulation, and can produce highly realistic fake videos.

[0032] In this case, relying solely on traditional audio and video detection technologies is no longer sufficient to effectively defend against AI fraud. Because even if the detection technologies can be continuously upgraded, as long as there is a certain false positive rate, fraudsters may take advantage of these loopholes and find opportunities to successfully implement fraud through large-scale fraud attempts. Therefore, simply relying on audio and video detection means at the technical level has become difficult to cope with the increasingly complex and advanced AI fraud threats.

[0033] In response to the problems that occur in related technologies, the embodiments of the present application provide an emotional category prediction method and device, which can solve the problem of insufficient prediction accuracy of fraud behaviors in related technologies.

[0034] The following will combine the accompanying drawings and specifically illustrate the emotional category prediction method provided by the embodiments of the present application through specific embodiments and their application scenarios.

[0035] Figure 1 It is a flowchart of an emotional category prediction method provided by an embodiment of the present application.

[0036] As Figure 1 shown, the emotional category prediction method may include step 110 - step 130. This method is applied to an emotional category prediction device, and is specifically as follows:

[0037] Step 110, obtain the data to be measured;

[0038] Data to be measured: The data that needs to be analyzed and judged for fraud risks. This data may come from chat content or push content of instant messaging software, social platforms, etc.

[0039] The data to be tested can be read from a database or obtained from a communication software interface, and the data to be tested for fraud risk analysis is collected, providing input for subsequent model processing.

[0040] Step 120: Input the data to be tested into the sentiment prediction model, and output the sentiment category values of at least two sentiment dimensions corresponding to the data to be tested. The sentiment prediction model is trained based on multiple sets of training data, and each set of training data includes: sample data and the corresponding labeled sentiment category values.

[0041] Sentiment prediction model: A model constructed based on machine learning or deep learning algorithms. By training on sample data, the sentiment prediction model can learn the characteristic patterns of fraud-related chat data in different sentiment dimensions, so as to predict new input data to be tested and determine whether it may be involved in fraud.

[0042] Sentiment dimension: A classification method for the sentiment in chat data. For example, it can include different sentiment directions such as anger, fear, greed, excitement, etc., depicting the sentiment conveyed by chat data from multiple angles.

[0043] Sentiment category value: A numerical value output by the sentiment prediction model for the input chat data in each sentiment dimension, indicating the degree or probability that the chat data belongs to a certain sentiment category in that sentiment dimension.

[0044] Labeled sentiment category value: In the training data, for each sample data, the sentiment category values corresponding to each sentiment dimension are manually marked as the reference standard for model training.

[0045] Sample data: The sample data can be sourced from the chat content or push content of instant messaging software, social platforms, etc. The data types of the sample data and the data to be tested are the same.

[0046] Input the data to be tested into the sentiment prediction model. The sentiment prediction model has learned the mapping relationship between the sample data and the labeled sentiment category values during the training phase. When new data to be tested is input, the sentiment prediction model calculates and outputs the sentiment category values of the chat data in multiple sentiment dimensions according to its internal algorithm and the learned characteristic patterns. The sentiment category values reflect the tendency degree of the chat data in different sentiment dimensions.

[0047] Step 130: When the sentiment category value corresponding to any sentiment dimension is greater than or equal to the preset threshold, output a prompt message.

[0048] Preset Threshold: A numerically predefined standard used to determine whether to trigger the output of a prompt message based on the sentiment category value. When the sentiment category value on a certain sentiment dimension reaches or exceeds the preset threshold, it is considered that there is a risk.

[0049] Judge the sentiment category value output by the model. On each sentiment dimension, compare the output sentiment category value with the preset threshold. If the sentiment category value on any one sentiment dimension is greater than or equal to the preset threshold, it indicates that the data to be tested shows a strong specific sentiment tendency on this sentiment dimension, and this tendency may be related to fraud behavior. Therefore, output a prompt message to remind relevant personnel to pay attention to the possible fraud risk.

[0050] Exemplarily, the sentiment dimensions include two sentiment dimensions: "fear" and "greed". The preset thresholds are 0.6 for the "fear" dimension and 0.7 for the "greed" dimension. A set of chat data between users is obtained from the chat record database of a certain social platform as the data to be tested, and the content is "The opportunity is rare. You can make a fortune if you invest! Hahaha!" Input the above chat data into the sentiment prediction model that has been trained with a large number of sample data including fraud and non-fraud chat records. After calculation, the model outputs the sentiment category value of this chat data on the "fear" dimension as 0.7, and the sentiment category value on the "greed" dimension as 0.8. Compare the sentiment category values output by the model with the preset thresholds. The sentiment category value of the "fear" dimension is 0.7, which is greater than the preset threshold of 0.6, and the sentiment category value of the "greed" dimension is 0.8, which is greater than the preset threshold of 0.7. Therefore, output a prompt message such as "This chat may have a fraud risk. Please treat it with caution."

[0051] Among them, the sentiment dimension includes at least one of the following: trust, curiosity, fear, greed, slackness, social pressure, and information desire.

[0052] The following will be explained in turn:

[0053] Trust Dimension: In an actual chat scenario, if the data to be tested comes from a message sent by a seemingly authoritative person or an acquaintance, when the fraud prediction sentiment prediction model processes this data, it will judge the user's trust level in this information based on the features it has learned. For example, in the chat data, it is mentioned that "I am the financial director of the company. Due to the temporary change of the company's account, you need to transfer the funds to this new account immediately." The sentiment prediction model will analyze features such as the language style and identity identification in it. If it conforms to the characteristics of an authoritative person, it may output a relatively high sentiment category value on the trust dimension. Because people tend to trust such seemingly authoritative information, and fraudsters often take advantage of this to carry out fraud. The sentiment prediction model can capture this trust tendency and quantify it through learning a large number of similar data.

[0054] In the chat data of the deceived, the emotion of trust often manifests as an unreserved acceptance of the information provided by the scammer. For example, when the scammer communicates with the deceived in the guise of a "senior financial expert", the deceived may say, "I fully trust the investment plan you mentioned and will operate according to what you said." Such an expression directly reflects the high degree of trust of the deceived in the scammer. Based on the professional image created by the scammer, they abandon the doubt about the authenticity of the information.

[0055] Curiosity dimension: When the data to be measured contains tempting content, such as "Click on this link to receive a huge red envelope", "Exclusive revelation, click to view the unknown truth", etc., the emotion prediction model will identify these attractive words and expressions. Based on the learning of curiosity-triggering situations in the training data, it judges the possibility that the user may click or view due to curiosity. If the emotion prediction model detects similar tempting elements, it will output a relatively high emotion category value in the curiosity dimension. Since scammers often take advantage of people's curiosity to induce them to click on links, which may lead to information leakage or property loss, the emotion prediction model predicts risks by learning this pattern.

[0056] Curiosity may drive the deceived to communicate more deeply with the scammer. Suppose the scammer claims to have an "internal exclusive money-making opportunity", and the deceived may ask, "What exactly is this opportunity? Is the return really that high?" This curiosity about novel and high-yield opportunities makes the deceived gradually fall into the scam trap, constantly asking for details in the chat, thus giving the scammer more opportunities for induction.

[0057] Fear dimension: If statements creating fear or a sense of urgency appear in the chat data, such as "Your account has been hacked. If you don't click on the link immediately to modify your password, all your funds will be transferred away", "Limited-time offer, only 1 hour left, no chance if you miss it", etc., the emotion prediction model will analyze the emotional tendency and language features therein. Based on the mastery of fear-emotion-triggering characteristics during the training process, it judges the performance of this chat data in the fear dimension. If it conforms to the pattern of creating fear or a sense of urgency, it will output a relatively high emotion category value. Since scammers often use this means to make victims make wrong decisions in a hurry without thinking, the emotion prediction model evaluates the scam risk by identifying these features.

[0058] Scammers often take advantage of the fear psychology of the deceived, such as sending messages like "Your bank card is suspected of money laundering crime. If you don't handle it immediately, you will face legal sanctions." The deceived may reply at this time, "What should I do? I really haven't done anything illegal!" From this reply, it can be seen that the deceived is in a panic due to fear. Under the dominance of this emotion, it is easy to follow the instructions of the scammer and ignore the unreasonable points therein.

[0059] Greed Dimension: When the data to be measured contains expressions that take advantage of "free gift worth xxx" or "time-limited special offer, original price xxx, now only xxx", the sentiment prediction model will, based on the learning of greed-triggering features in the training data, determine the degree to which users may be lured due to greed. If similar greed-inducing elements are detected, a higher sentiment category value will be output on the greed dimension. Since scammers often set such traps to attract victims, the sentiment prediction model captures these features to predict the possibility of fraud using greed in chat data.

[0060] When scammers dangle the bait of "high returns", the greed of the deceived is easily aroused. For example, if a scammer promises a 50% return on investment in a project within a month, the deceived may excitedly respond: "If I invest 100,000, can I really make 50,000 in a month? Quickly tell me how to operate." This greed for the rapid appreciation of wealth makes the deceived lower their vigilance against risks and blindly believe the false promises of scammers.

[0061] Laziness Dimension: If the chat data involves content related to some daily tasks, such as "Please complete the review of this document and submit it as soon as possible without missing any details", and considering the actual situation, such as the user having a heavy workload recently, the sentiment prediction model will consider that people may neglect safety details due to being busy with tasks in this situation. The sentiment prediction model will analyze the possible lazy psychology of users when dealing with such tasks based on the learning of laziness in the training. If it is judged that the possibility of the user being in a lazy state is relatively high, a higher sentiment category value will be output on the laziness dimension. Since scammers may take advantage of people's lazy psychology and send some seemingly normal but actually risky information, the sentiment prediction model assesses the risk through the analysis of this situation.

[0062] Some of the deceived have a lazy psychology and are unwilling to spend the energy to verify information. In the chat, when scammers provide some seemingly complex documents or instructions, the deceived may say: "It's too troublesome. I trust you. Just tell me directly what to do." This laziness makes them skip the crucial verification link and easily accept the arrangements of scammers.

[0063] Social pressure dimension: When expressions such as "I am a senior executive of the company, and now there is an urgent task that requires your help to handle, don't refuse" and "This is a requirement from the leader, and you must do it" that exert pressure in the name of "senior executives of the company" or other authoritative identities appear in the chat data, the emotion prediction model will identify these language features and identity identifiers. Based on the learning of the situations that trigger social pressure in the training data, it judges the degree to which the user may be unable to refuse due to social pressure. If it conforms to the pattern of social pressure, a higher emotion category value will be output on the social pressure dimension. Since scammers often use this kind of authoritative identity to exert pressure on victims, making them make decisions that are disadvantageous to themselves reluctantly, the emotion prediction model analyzes these features to predict the fraud risk.

[0064] Scammers sometimes use social relationships to exert pressure. For example, on the grounds that "all your friends have participated in this project, and you will miss the opportunity if you don't participate". The deceived person may respond: "Everyone has done it, so I'll give it a try too." Under social pressure, in order not to appear out of place or miss the so-called "opportunity", the deceived person chooses to follow the scammer's lead and reveals the compromise made due to social pressure in the chat.

[0065] Information eagerness dimension: If the chat data contains expressions that take advantage of people's eagerness for unknown information, such as "Here is an industry secret you don't know, click to view" and "Exclusive news, which will be of great help to your career development after understanding", the emotion prediction model will analyze the emotional tendency and language characteristics therein. Based on the learning of the features that trigger the emotion of information eagerness in the training, it judges the degree to which the user may be attracted due to information eagerness. If elements similar to those taking advantage of information eagerness are detected, a higher emotion category value will be output on the information eagerness dimension. Since scammers often take advantage of people's curiosity and eagerness for unknown information to induce them to click on links or provide personal information, the emotion prediction model evaluates the fraud risk of the chat data by identifying these features.

[0066] The deceived person often has a strong eagerness for information that may change their situation. If the scammer uses the guise of "providing high-income part-time job information", the deceived person may eagerly ask: "What exactly does this part-time job do? What are the requirements? How is the income settled?" This eagerness for information makes the deceived person actively seek more details in the chat, ignoring the consideration of the authenticity of the information and being easily exploited by scammers.

[0067] By outputting emotion category values on multiple emotion dimensions, the emotion prediction model can analyze the data to be tested more comprehensively and meticulously, capture possible fraud risk factors from different angles, and provide more accurate fraud warnings for users.

[0068] By performing sentiment dimension analysis on chat data, it is possible to timely detect potential fraud risks before fraud occurs or further develops, provide early warnings for users or relevant institutions, and avoid or reduce potential losses. The output prompt information can help users, customer service personnel, security agencies, etc. handle relevant information more cautiously when facing chat scenarios, make more reasonable decisions, such as further verifying the authenticity of information, blocking possible fraud transactions, etc. It helps to enhance the security of the chat environment, reduce the occurrence probability of fraud incidents, and protect the property safety and personal information safety of users.

[0069] The training process of the sentiment prediction model is described below:

[0070] In a possible embodiment, before step 120, the following steps may be included:

[0071] Step 210, input the sample data into the initial prediction model, and output the sample sentiment category values on at least two sentiment dimensions corresponding to the sample data;

[0072] Step 220, calculate the loss value according to the sample sentiment category value and the label sentiment category value;

[0073] Step 230, adjust the model parameters of the initial prediction model according to the loss value until the preset training conditions are met, and obtain the sentiment prediction model.

[0074] Initial prediction model: A model architecture that has not been fully trained, with a basic prediction ability framework, but has not learned accurate enough features and patterns to accurately predict chat data on various sentiment dimensions. The structure of the initial prediction model can be a neural network, such as a Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), etc., or a machine learning or deep learning model such as a support vector machine.

[0075] Sample sentiment category value: After the initial prediction model processes the input sample data, the numerical values representing the sentiment tendency degree of the sample data on multiple sentiment dimensions are output. These numerical values are the prediction results obtained by the model based on the current parameters and algorithms.

[0076] Loss value: A numerical value used to measure the difference degree between the sample sentiment category value and the label sentiment category value. Common loss functions include mean squared error, cross-entropy loss, etc. The smaller the loss value, the closer the prediction result of the model is to the real situation.

[0077] Preset training conditions: Conditions that are preset for terminating the model training process. Common preset training conditions include reaching the maximum number of training epochs, the loss value being lower than a specific threshold, the accuracy on the validation set no longer improving, etc.

[0078] Input the sample data into the initial prediction model. The model will analyze and process the input sample data according to its internal parameters and calculation logic, and thus output the sample sentiment category values on multiple sentiment dimensions. This process is the initial prediction of the model for the sample data based on the current initial model parameters.

[0079] According to the sample sentiment category values output by the model and the pre-annotated label sentiment category values, use a specific loss function to calculate the loss value. The role of the loss function is to quantify the gap between the model's prediction result and the real situation. Through the loss value, we can understand the current performance of the model. The larger the loss value, the less accurate the model's prediction.

[0080] Based on the calculated loss value, use an optimization algorithm to adjust the parameters of the initial prediction model. The goal of the optimization algorithm is to gradually reduce the loss value by continuously adjusting the model parameters, that is, to make the model's prediction results closer and closer to the real labels. This process will be repeated continuously until the preset training conditions are met. At this time, the model parameters have been adjusted to a relatively optimal state, and the trained sentiment prediction model is obtained.

[0081] Exemplarily, use the two sentiment dimensions of "surprised" and "eager" to analyze the chat data. 1000 pieces of sample data are collected, and each piece of data is manually annotated with the label sentiment category values on the two sentiment dimensions of "surprised" and "eager". For example, there is a piece of sample data "Oh my god! This opportunity is fleeting. Act quickly!", and its label sentiment category value for the "surprised" dimension is annotated as 0.8, and the label sentiment category value for the "eager" dimension is annotated as 0.9.

[0082] Select a simple neural network as the initial prediction model. Input the above sample data into the initial prediction model batch by batch. For the data "Oh my god! This opportunity is fleeting. Act quickly!", the model outputs the sample sentiment category value of 0.6 on the "surprised" dimension and the sample sentiment category value of 0.7 on the "eager" dimension.

[0083] Use the mean squared error (MSE) as the loss function to calculate the loss value. For this piece of data, the loss for the "surprised" dimension is (0.8 - 0.6) 2 = 0.04, and the loss for the "eager" dimension is (0.9 - 0.7) 2 = 0.04. The total loss value is the average of the losses of the two dimensions, that is, (0.04 + 0.04) / 2 = 0.04.

[0084] Use the Adam optimization algorithm to adjust the parameters of the model according to the calculated loss value. Set the preset training condition that the maximum number of training rounds is 100 rounds. Each round of training will traverse all the sample data and continuously adjust the model parameters. When training reaches the 100th round, the preset training condition is met, the training ends, and a trained sentiment prediction model is obtained.

[0085] By continuously calculating the loss value and adjusting the model parameters, the sentiment prediction model can gradually learn the fraud-related feature patterns in the sample data, thereby improving the accuracy and reliability of the sentiment prediction model in predicting the sentiment category values of chat data. The trained sentiment prediction model can better adapt to different types of chat data, and can make more accurate judgments on various expressions and sentiment tendencies that may involve fraud, improving the generalization ability of the model. The trained sentiment prediction model can accurately predict new data to be tested in actual scenarios, provide effective technical support for preventing fraud, help users and relevant institutions timely discover potential fraud risks, and ensure property safety and information security.

[0086] In a possible embodiment, step 210 includes:

[0087] Input the sample data into the initial prediction model to extract the sample sentiment features of the sample data;

[0088] Classify the sample sentiment features and output the sample sentiment category values on at least two sentiment dimensions corresponding to the sample data.

[0089] Sample sentiment features: refer to the feature information extracted from the sample data that can reflect its sentiment tendency. The sample sentiment features can be at the lexical level, such as specific sentiment words, keywords, etc., at the semantic level, such as the sentiment tendency and theme expressed by the sentence, or at the syntactic level, such as the sentence structure, etc. It is an abstract representation of the sample data, used to help the model understand the sentiment information in the data.

[0090] Sample data: refers to the chat record data used to train the sentiment prediction model, and has been manually labeled with the label sentiment category values on multiple sentiment dimensions as a reference basis for training the model.

[0091] Initial prediction model: is a model architecture that has not been fully trained, with a basic functional framework for feature extraction and classification of input data. However, before training, its parameter settings have not been optimized, and the accuracy of extracting and classifying the sentiment features of the sample data needs to be improved.

[0092] After the sample data is input into the initial prediction model, the model processes the data according to its internal algorithms and structure. For example, if the initial prediction model is based on a neural network, the initial prediction model performs operations such as word embedding and feature extraction on the input text data through the calculations of multiple layers of neurons, converting the text data into a series of feature vectors that can represent its sentiment tendency, and the feature vectors constitute the sample sentiment features. Different model structures and algorithms will adopt different ways to extract sample sentiment features, and the purpose is to capture the key information related to sentiment in the chat data.

[0093] After obtaining the sample sentiment features, the model further analyzes and processes these features, maps the features to different sentiment dimensions through a classifier, and calculates the probabilities or degrees belonging to different sentiment categories on each sentiment dimension, thereby outputting the sample sentiment category values on multiple sentiment dimensions. This process is a preliminary classification prediction based on the parameter settings and algorithm logic before the model training.

[0094] By extracting the sample sentiment features, the original chat data can be transformed into more representative and analyzable feature vectors, enabling the model to more effectively process and understand the sentiment information in the data, providing a basis for subsequent classification and prediction. Classifying the sample sentiment features and outputting the sample sentiment category values on multiple sentiment dimensions provides preliminary prediction results for the model training. By comparing with the labeled sentiment category values, the loss value can be calculated, and then the model can be optimized and trained to gradually improve the prediction accuracy of the model for the sentiment tendency of chat data.

[0095] Among them, inputting the sample data into the initial prediction model and extracting the sample sentiment features of the sample data include at least one of the following:

[0096] In the case where the sample data includes sample text data, identifying the sample sentiment keywords in the sample text data; and extracting the sample text sentiment features in the sample sentiment keywords;

[0097] In the case where the sample data includes sample voice data, extracting the sample voice sentiment features of the sample voice data; the sample voice sentiment features include: timbre features, pitch features, and speech rate features;

[0098] In the case where the sample data includes sample video data, identifying the key frames in the sample video data; and extracting the sample video sentiment features in the key frames, the sample video sentiment features include: expression features and body features.

[0099] Sample sentiment keywords: Words in sample text data that can directly or indirectly express a specific sentiment tendency. For example, in "I'm very happy today", "happy" is a sample sentiment keyword expressing a positive sentiment; in "I'm very angry about this matter", "angry" is a keyword expressing a negative sentiment.

[0100] Sample text sentiment features: Specific feature information extracted from sample sentiment keywords that can reflect the sentiment represented by the keyword. For example, the part of speech of the word, sentiment polarity, such as positive, negative, neutral, semantic intensity, etc. can all be used as sample text sentiment features.

[0101] Sample voice sentiment features: Include timbre features, such as the sound quality of the voice, the uniqueness of the timbre, etc. Different people have different timbres, and the timbre may also change under different emotional states;

[0102] Pitch features, the rise and fall of pitch. For example, the pitch may rise when angry and may fall when sad;

[0103] Speech rate features, the speed of speaking. The speech rate may increase when excited and may be slower when calm. These features can reflect the emotions contained in the sample voice data.

[0104] Sample video data in sample data: Refers to the video information included in the chat process, which may be short video clips sent by users, etc., or the video data of a video call between both parties.

[0105] Key frame: In sample video data, a specific frame that contains important information or can represent the main content and emotion of the video. For example, in a video expressing anger, the frame with the most exaggerated facial expression and the most obvious body movements of the person may be the key frame.

[0106] Sample video sentiment features: Sentiment features extracted from key frames, including facial expression features and body features.

[0107] Facial expression features, such as the facial expressions of people. A smile indicates happiness, and frowning indicates dissatisfaction or anger, etc.; Body features, such as the amplitude and posture of body movements. A large wave of the hand may indicate excitement, and crossing the arms may indicate defense or dissatisfaction.

[0108] For sample text data, natural language processing technology can be used to perform word segmentation, part-of-speech tagging, and other processing on the sample text data, and then the matching emotional keywords can be found in the segmented text according to the pre-established emotional keyword dictionary. For example, a rule-based method or a machine learning method can be used to determine whether each word belongs to an emotional keyword. After identifying the sample emotional keywords, these keywords are further analyzed. The word's part of speech, emotional polarity, and other information can be obtained by querying the dictionary, or the machine learning model can be used to analyze the semantics of the keywords and extract more complex emotional features, such as semantic strength. These features can help the model understand the emotions expressed in the text more accurately.

[0109] For the sample speech data, the speech signal can be preprocessed, such as noise reduction, frame division, etc. Then, the timbre features, pitch features and speech speed features are extracted respectively. The timbre features can be obtained by analyzing the spectral characteristics of the speech signal, the pitch features can be determined by detecting the fundamental frequency changes of the speech signal, and the speech speed features can be obtained by calculating the number of words or syllables in the speech per unit time. The above features can be combined to reflect the emotional information contained in the speech.

[0110] For sample video data, video processing technology can be used to analyze the sample video data frame by frame. Key frames can be determined based on factors such as the degree of change in video content and motion information. For example, if a frame has a large change in picture content compared to adjacent frames, or contains important actions or expressions, then this frame may be identified as a key frame. For the identified key frames, computer vision technology, such as face detection and posture estimation, is used to extract expression features and body features. By analyzing the facial expressions and body movements of the characters, the emotional tendencies expressed in the video can be judged. For example, the expression of a person's mouth corners raised and eyes narrowed may indicate a happy emotion; the body movements of a person clenching his fists and leaning forward may indicate anger or excitement.

[0111] For example, there is a text message in the sample data: "This activity is so bad that I can't stand it anymore!" First, the text is segmented to obtain words such as "this time", "of", "activity", "too", "bad", "simply", "let", "me", "can't stand it anymore". Then, "bad" and "can't stand it anymore" are identified as sample emotional keywords through the emotional keyword dictionary. Next, the emotional features of the sample text are extracted. "Bad" is an adjective with a negative emotional polarity; "can't stand it anymore" has a high semantic intensity and a negative emotional polarity.

[0112] There is a voice sample data sent by a user, and the content is the user complaining angrily. After preprocessing the voice, timbre features are extracted and it is found that the voice is relatively sharp; pitch features are extracted and it is detected that the pitch is high and has large fluctuations; speech rate features are extracted and it is calculated that the speech rate is fast. Combining these features, it is judged that the voice expresses an angry emotion.

[0113] In a sample video data, the user is expressing an excited emotion. Through video processing technology, the frame with the richest facial expressions and the most obvious actions of the person is identified as the key frame. Then the key frame is analyzed, facial expression features are extracted and it is found that the person's eyes are wide open and the mouth is open; limb features are extracted and it is seen that the person's hands are waving. These features indicate that the user in the video is in an excited emotional state.

[0114] By processing the sample text, voice and video data, emotional information can be extracted from multiple modalities, enabling the model to more comprehensively understand the emotions expressed in the sample data and improving the accuracy and reliability of emotional analysis. Adopting corresponding feature extraction methods for different modalities of data can accurately capture the key features related to emotions in each modality, providing richer and more accurate information for subsequent emotional classification and fraud prediction. Extracting various types of emotional features enables the model to adapt to different forms of chat data, enhancing the generalization ability and adaptability of the model, being able to better handle the complex and diverse data forms in actual chat scenarios, and improving the overall performance of the emotional prediction model.

[0115] Among them, the sample emotional features include at least two of the sample text emotional features, the sample voice emotional features and the sample video emotional features. Classifying the sample emotional features outputs the sample emotional category values on at least two emotional dimensions corresponding to the training data, including:

[0116] Performing weighted fusion on at least two of the sample text emotional features, the sample voice emotional features and the sample video emotional features to obtain the sample fused emotional features;

[0117] Classifying the sample fused emotional features outputs the sample emotional category values on at least two emotional dimensions corresponding to the training data.

[0118] Sample fused emotional features: The comprehensive emotional features obtained after performing weighted fusion processing on at least two of the sample text emotional features, the sample voice emotional features and the sample video emotional features. The sample fused emotional features integrate the emotional information from different modality data.

[0119] The importance of emotional features in different modalities may vary when expressing emotions. Therefore, different weights need to be assigned to the emotional features of sample texts, sample speech, and sample videos respectively. The determination of weights can be set based on experience or automatically learned through machine learning algorithms during the training process.

[0120] For example, in certain scenarios, the pitch feature in speech may be more crucial for expressing anger. Then, the weight of the pitch feature in the sample speech emotional features can be set relatively high. By means of weighted summation and other methods, these emotional features in different modalities are fused together to form a comprehensive sample fused emotional feature, which contains emotional information from multiple modalities and can more comprehensively reflect the emotional state of the sample data.

[0121] The obtained sample fused emotional feature is input into the classifier. The classifier analyzes and processes the sample fused emotional feature according to its internal algorithm and the learned patterns, maps it to multiple emotional dimensions, and calculates the probability or degree of belonging to different emotional categories on each emotional dimension, thereby outputting the sample emotional category values on multiple emotional dimensions.

[0122] Exemplarily, there is a piece of chat data of a user, including a text message "This news is so exciting!", a voice data of the user excitedly telling, and a video data of the user dancing and having an excited expression on the face.

[0123] Through text analysis, the emotional keyword "excited" is extracted, and its emotional polarity is determined to be positive with a relatively high semantic intensity. The speech is processed to extract the timbre feature as a relatively bright voice, the pitch feature as a high and rising pitch, and the speech rate feature as a fast speech rate. The key frames of the video are identified, and the expression feature is extracted as the corners of the mouth turned up and the eyes widened, and the body feature is extracted as the hands waving.

[0124] Suppose it is determined through training that the weight of the sample text emotional feature is 0.3, the weight of the sample speech emotional feature is 0.4, and the weight of the sample video emotional feature is 0.3. The emotional features of each modality are weighted and fused to obtain the sample fused emotional feature. For example, the semantic intensity score of "excited" in the text, the pitch feature score in the speech, the expression feature score in the video, etc. are weighted and summed according to the weights.

[0125] The sample fused emotional feature is input into the trained classifier, and the classifier analyzes it. On the "excited" emotional dimension, the output sample emotional category value is 0.8, indicating that the chat data has a high probability of belonging to the excited emotional category on the "excited" emotional dimension; on the "calm" emotional dimension, the output sample emotional category value is 0.2, indicating a low probability of belonging to the calm emotional category.

[0126] It can be understood that when the sample emotion features include at least two of the sample text emotion features, sample speech emotion features, and sample video emotion features, the at least two sample emotion features can be weighted and fused to obtain the sample fused emotion features. When the sample emotion features include any one of the sample text emotion features, sample speech emotion features, and sample video emotion features, any one of the sample emotion features is classified to output the sample emotion category values on multiple emotion dimensions.

[0127] By weighted-fusing the emotion features of sample text, speech, and video, the emotion information contained in different modality data is fully utilized, avoiding the problem of incomplete information that may exist in single-modality data, enabling the model to more comprehensively and accurately understand the emotion state of chat data. Classifying based on comprehensive multi-modal emotion features can capture richer emotion clues compared to only using single-modal features, thereby improving the accuracy of the model's prediction of the sample emotion category values on multiple emotion dimensions, and helping to more accurately determine whether there is an emotion tendency related to fraud in the chat data.

[0128] Data of different modalities may be affected by noise or interference in some cases. Multi-modal feature fusion can reduce this impact to a certain extent. When the data of a certain modality is inaccurate or incomplete, the data of other modalities can provide supplementary information, enhancing the robustness and stability of the emotion prediction model, enabling the emotion prediction model to also perform well in complex chat scenarios.

[0129] The following explains the specific application process of the emotion prediction model:

[0130] In a possible embodiment, step 120 may specifically include the following steps:

[0131] Step 310, input the data to be measured into the emotion prediction model to extract the emotion features to be measured of the data to be measured;

[0132] Step 320, classify the emotion features to be measured and output the emotion category values of at least two emotion dimensions corresponding to the data to be measured.

[0133] Emotion features to be measured: Feature information extracted from the data to be measured that can reflect the emotion tendency of the data to be measured, which may include keywords in the text, pitch of the speech, expressions in the video, etc.

[0134] The sentiment prediction model processes the data to be measured according to its internal algorithms and structures. For different forms of chat data, corresponding feature extraction methods are adopted. For example, for text data, operations such as word segmentation, part-of-speech tagging, and keyword extraction may be performed to identify sentiment-related words and semantic information; for voice data, features such as pitch, timbre, and speech rate are extracted; for video data, key frames are identified and features such as expressions and body movements are extracted. These extracted features are combined to form the sentiment features to be measured, which can reflect the sentiment information contained in the data to be measured.

[0135] The classifier in the sentiment prediction model analyzes and judges the sentiment features to be measured according to the learned patterns and rules. These features are mapped to each sentiment dimension, and the sentiment category values on each sentiment dimension are calculated. For example, by comparing the sentiment features to be measured with the feature patterns of different sentiment dimensions in the training data, the tendency degree of the chat data on each sentiment dimension such as "trust" and "curiosity" is determined and output in numerical form.

[0136] Exemplarily, a piece of data to be measured is a text message: "You have won the grand prize of our company. You need to pay a handling fee first to receive the bonus." The text is segmented into words such as "you", "have won", "our", "company", "the", "grand prize", "need", "pay first", "a sum of", "handling fee", "in order to", "receive", "bonus". Sentiment keywords "grand prize" and "handling fee" are identified, and these keywords may be related to emotions such as "greed" and "fear".

[0137] Through semantic analysis, it is found that the data to be measured creates a situation where people feel there is an opportunity to obtain huge benefits but also need to pay a certain price. Combining this information, the sentiment features to be measured are extracted, such as positive semantic features related to "grand prize" and cautious semantic features related to "handling fee". The extracted sentiment features to be measured are input into the classifier of the sentiment prediction model. After calculation, the classifier outputs the sentiment category values on each sentiment dimension:

[0138] On the "greed" dimension, the sentiment category value is 0.8, indicating that the chat data has, to a certain extent, stimulated the recipient's greed for the huge bonus.

[0139] On the "fear" dimension, the sentiment category value is 0.3 because the mention of the need to pay a handling fee may make the recipient have certain concerns and fears.

[0140] On the "trust" dimension, the sentiment category value is 0.2, indicating that the authenticity and reliability of the text are relatively low and it is difficult for the recipient to generate trust.

[0141] In the "curiosity" dimension, the emotional category value is 0.6 because the news of "winning a big prize" is likely to arouse the curiosity of the recipient.

[0142] In the "laziness" dimension, the emotional category value is 0.1, and there are no obvious factors in this text that can arouse lazy emotions.

[0143] In the "social pressure" dimension, the emotional category value is 0.1, and no social pressure is reflected.

[0144] In the "information eagerness" dimension, the emotional category value is 0.4. The recipient may have a certain eagerness for information about further understanding the authenticity of the big prize and the handling fees.

[0145] In step 320, a fully connected layer, a support vector machine, or a Softmax classifier can be used to classify the emotional features to be measured, and the emotional category values on multiple emotional dimensions are output.

[0146] Among them, when using the fully connected layer in the neural network for emotional classification, the weighted emotional features are input into the fully connected layer for calculation, and finally the probability distribution of various emotions is obtained. The support vector machine (Support Vector Machine, SVM) is a commonly used binary classification algorithm, which can also be extended to a multi-classification problem. It separates different emotional categories by constructing a hyperplane for emotional classification. The Softmax classifier is usually used in multi-class emotional classification to map the weighted and fused feature vectors into the probability values of different emotional categories, and finally determine the emotional category.

[0147] By extracting the emotional features to be measured and classifying them on multiple emotional dimensions, the emotional tendency contained in the data to be measured can be analyzed more carefully and accurately. Different fraud means often trigger different emotional reactions. Analyzing from at least two dimensions can capture these emotional features more comprehensively, which helps to identify potential fraud risks. The emotional category values on at least two emotional dimensions output provide a quantitative basis for judging whether there is fraud. For example, if the emotional category values on dimensions such as "greed" or "fear" are relatively high, it may mean that there is a greater suspicion of fraud in this chat data. It can help users, customer service staff, or security agencies make more scientific decisions and take corresponding preventive measures.

[0148] The method based on multi-emotional dimension analysis can enhance the ability to identify fraud behaviors, discover potential fraud risks in a timely manner, thereby improving the security of the chat environment and protecting the property safety and personal information safety of users.

[0149] In a possible embodiment, step 310 specifically includes at least one of the following:

[0150] When the data to be measured includes text data, extract the text sentiment features of the text data;

[0151] When the data to be measured includes voice data, extract the voice sentiment features of the voice data. The voice sentiment features include: timbre features, pitch features, and speech rate features;

[0152] When the data to be measured includes video data, extract the video sentiment features of the video data. The video sentiment features include: expression features and body features.

[0153] Text sentiment features: Feature information that can reflect the sentiment conveyed by the text, extracted from the text part of the data to be measured. For example, specific sentiment words, the sentiment polarity of words, the tone and semantic tendency of sentences, etc. These features can help understand the sentiment state in the text.

[0154] Voice sentiment features: Existing in the voice data of the data to be measured, used to reflect the features of the sentiment expressed by the voice. Include timbre features, pitch features, and speech rate features.

[0155] Video sentiment features: Features related to sentiment extracted from the video data of the data to be measured. Mainly include expression features and body features, which can intuitively show the sentiment state of the people in the video.

[0156] First, preprocess the text data, such as operations like word segmentation and stop word removal. Then, through natural language processing techniques, use sentiment dictionaries or machine learning models to identify sentiment words in the text. For example, use rule-based methods to judge the sentiment tendency of the text according to the sentiment polarity of words in the sentiment dictionary; or use deep learning models, such as RNN, LSTM, etc., to model the text and learn the semantic features and sentiment expressions of the text.

[0157] Further analyze the grammatical structure of the text, the tone of the sentence, etc., and comprehensively extract features that can accurately reflect the sentiment of the text, such as the frequency of sentiment words, the intensity of sentiment, etc.

[0158] Preprocess the voice data, including operations such as noise reduction and frame division to improve the quality of the voice signal. When extracting timbre features, obtain the unique sound quality information of the voice by analyzing methods such as the spectral characteristics of the voice signal. For pitch features, use the fundamental frequency detection algorithm to determine the pitch changes of the voice signal. Calculate the speech rate feature by counting the number of words or syllables in the voice per unit time. Combine these extracted timbre, pitch, and speech rate features to reflect the sentiment expressed by the voice data.

[0159] Process the video data frame by frame. Through computer vision technologies such as face detection and pose estimation, identify the key frames in the video, that is, those frames containing important emotional information. Extract the facial expression features from the key frames, use facial expression recognition algorithms to detect the facial expressions of the people, and analyze the types and intensities of the expressions. When extracting the body features, determine the body movements and postures of the people through pose estimation algorithms, and analyze the amplitudes, directions, etc. of the body movements, so as to judge the emotional states of the people in the video.

[0160] By separately extracting the emotional features of text data, speech data, and video data, it is possible to comprehensively capture the emotions expressed by the data to be measured from multiple dimensions. The data of different modalities can complement each other, providing richer emotional information and making the emotional analysis of the data to be measured more accurate and comprehensive. The rich emotional features provide more detailed input information for the emotion prediction model. Since fraudulent behaviors often trigger specific emotional reactions, by accurately extracting these emotional features, the emotion prediction model can better judge whether there is a fraud risk in the chat data, thus improving the accuracy of fraud prediction. In the actual chat scenario, the data forms are diverse, including various forms such as text, speech, and video. This multi-modal emotional feature extraction method can adapt to different data forms. No matter what form the data exists in, it can effectively extract the emotional features, enhancing the applicability and robustness of the system in complex chat scenarios.

[0161] In a possible embodiment, the emotional features to be measured include at least two of text emotional features, speech emotional features, and video emotional features. In step 320, it may specifically include the following steps:

[0162] Perform weighted fusion on at least two of the text emotional features, speech emotional features, and video emotional features to obtain the fused emotional features;

[0163] Classify the fused emotional features and output the emotional category values of at least two emotional dimensions corresponding to the data to be measured.

[0164] Fused emotional features: The comprehensive emotional features obtained by fusing text emotional features, speech emotional features, and video emotional features according to certain weights. The fused emotional features integrate the emotional information of multi-modal data.

[0165] The importance of emotional features in different modalities may vary when expressing emotions. Therefore, different weights are assigned to text emotional features, speech emotional features, and video emotional features respectively. The determination of weights can be based on experience or automatically learned through machine learning algorithms during model training. For example, in some cases, the pitch feature of speech may be more crucial for expressing fear emotions, so the weight of speech emotional features can be relatively high. These emotional features in different modalities are fused together through methods such as weighted summation to form a comprehensive fused emotional feature, which can more comprehensively reflect the emotional state of the data to be measured.

[0166] The fused emotional feature is input into a classifier. Based on its internal classification algorithm and the learned patterns, the classifier analyzes and processes the fused emotional feature. It maps the fused emotional feature to multiple preset emotional dimensions and calculates the probability or degree of belonging to different emotional categories on each emotional dimension, and finally outputs the emotional category values on multiple emotional dimensions.

[0167] As Figure 2 shown, the data to be measured includes: text data 111, speech data 112, and video data 113. The text emotional feature 121 of the text data is extracted, the speech emotional feature 122 in the speech data is extracted, and the video emotional feature 123 of the video data is extracted. The text emotional feature 121, speech emotional feature 122, and video emotional feature 123 are fused to obtain the fused emotional feature 131, and the fused emotional feature 131 is classified to obtain the emotional classification value 141.

[0168] Exemplarily, a piece of data to be measured containing text, speech, and video is received. The content is a person claiming to have discovered an investment opportunity to get rich quickly. Words such as "get rich quickly" and "excellent opportunity" are used in the text. After analysis, features such as positive semantic orientation and high emotional intensity are extracted. The pitch of the speech is high and the speaking speed is fast, and the tone color appears excited. Based on this, excited pitch features, fast speaking speed features, and excited tone color features are extracted. In the video, the person's eyes are bright and the hands are waving, and excited expression features and excited body features are extracted.

[0169] Suppose it is determined through training that the weight of text emotional features is 0.3, the weight of speech emotional features is 0.4, and the weight of video emotional features is 0.3. The emotional features of each modality are weighted and fused according to the weights. For example, the positive semantic orientation score in the text, the excited pitch feature score in the speech, the excited expression feature score in the video, etc. are weighted and summed to obtain the fused emotional feature.

[0170] The fused emotional feature is input into the trained classifier. After calculation, the classifier outputs the emotional category values on each emotional dimension:

[0171] In the "greed" dimension, the sentiment category value is 0.8, which shows that the chat data has stimulated people's greed for wealth to a great extent.

[0172] On the “curiosity” dimension, the sentiment category value is 0.7, indicating that the message is likely to arouse the curiosity of others.

[0173] On the “Trust” dimension, the sentiment category value is 0.2, because the credibility of such get-rich-quick claims is low and it is difficult for the recipient to trust them.

[0174] In the “fear” dimension, the sentiment category value is 0.1, and the message does not obviously arouse fear.

[0175] In the "slackness" dimension, the emotion category value is 0.1, which does not reflect the emotions related to slackness.

[0176] In the “social pressure” dimension, the emotional category value is 0.1, and there is no suggestion of social pressure.

[0177] On the dimension of “information desire”, the sentiment category value is 0.6, and the receiver may have a certain desire for specific information about this investment opportunity.

[0178] By weighted fusion of the sentiment features of text, voice, and video, the sentiment information of multimodal data is fully integrated, avoiding the problem of incomplete information that may exist in single-modal data, making the analysis of sentiment more comprehensive and accurate. The comprehensive fusion sentiment features provide the classifier with richer and more comprehensive information, which helps the classifier to more accurately judge the sentiment tendency of the test data in various sentiment dimensions, thereby improving the prediction accuracy of the sentiment category value. Since fraudulent behavior often triggers specific emotional reactions, accurate sentiment classification can more effectively identify these sentiment features, thereby improving the ability of the sentiment prediction model to judge fraud risks and providing a more reliable basis for preventing fraud.

[0179] In a possible embodiment, after step 130, the following steps may also be included:

[0180] Receive feedback information from the user regarding the prompt information input;

[0181] According to the feedback information and the data to be tested, the model parameters of the emotion prediction model are updated.

[0182] Among them, the prompt information is used to indicate that the chat data may contain fraud risks.

[0183] like Figure 3As shown, the prompt message is: In this call, your emotional scores for [Curiosity] and [Desire for Information] exceed 80%. You need to handle unknown information with caution to avoid irrational decision-making and fraud risks. The prompt message includes a first control 151 "Confirm Risk" and a second control 152 "No Risk".

[0184] Feedback information: The response information input by the user after receiving the prompt message output by the system. The feedback information can be the user's judgment on whether the chat data is really involved in fraud, such as confirming fraud, believing it is not fraud, etc., or can include other relevant descriptions or opinions of the user on the chat data.

[0185] As Figure 3 shown, when a control input to the first control 151 is received, the feedback information is determined to be "Confirm Risk", and when a control input to the second control 152 is received, the feedback information is determined to be "No Risk".

[0186] Model parameters: Various parameters used to determine the behavior and output of the emotion prediction model. For example, in a neural network model, the model parameters can be the weight values, bias values, etc. connecting each neuron; in a machine learning model, they may be some hyperparameters, etc. These parameters are continuously adjusted during the model training process to optimize the performance of the model.

[0187] After outputting the prompt message to remind the user that the data to be tested may have a fraud risk, in order to obtain more accurate information about whether the chat data is really involved in fraud, it will wait for the user to input feedback information. The user provides corresponding feedback based on their own judgment and understanding, and this feedback information will be an important basis for subsequent model updates.

[0188] Combine the user's feedback information with the original data to be tested, and analyze the relationship between the two. According to the analysis result, use a specific algorithm, such as the backpropagation algorithm, etc. to calculate the direction and amplitude of the adjustment of the model parameters. Then update the model parameters of the emotion prediction model so that the model can better adapt to the new situation and improve the accuracy of the model in predicting fraud for similar chat data in the future. By continuously receiving user feedback and updating the model parameters, the model can gradually learn more accurate fraud patterns and characteristics, thereby continuously optimizing its performance.

[0189] Exemplarily, a prompt message is output, indicating that a certain chat data "You have a high - amount tax refund to claim. Just click on the link to fill in your personal information to handle it" may have a fraud risk. After further verification and judgment by the user, the feedback information is input, clearly indicating that this chat data is fraud information. The system integrates the user's feedback information with the original data to be tested "You have a high - amount tax refund to claim. Just click on the link to fill in your personal information to handle it".

[0190] Using machine learning algorithms, analyze the characteristics of the chat data, such as the relationship between features that are easy to attract people, like containing words such as "high - amount tax refund", and requirements to click on links to fill in personal information, and fraud judgment. According to the analysis results, calculate the adjustment amount of the model parameters. The model parameters can be the weight values and bias values in the neural network model.

[0191] Update the model parameters of the sentiment prediction model so that when the model encounters similar chat data containing words like "high - amount tax refund" and requires filling in personal information in the future, it can more accurately judge it as fraud information and improve the fraud prediction ability of the model.

[0192] By receiving the feedback information from users and updating the model parameters in combination with the data to be tested, the model can continuously learn new fraud patterns and characteristics, thereby gradually improving the prediction accuracy of whether the chat data involves fraud and reducing the situations of misjudgment and missed judgment. The feedback information from users can reflect various complex situations in the actual application scenario and the real judgment of users. Updating the model parameters based on these feedbacks enables the model to better adapt to different users' judgment criteria and various complex chat scenarios, enhancing the adaptability and generalization ability of the model. This mechanism for updating model parameters based on user feedback is a continuous process. As users continuously provide feedback, the model can be continuously optimized and improved, making its performance continuously enhanced to better meet the actual needs of preventing fraud and providing more reliable fraud prediction services for users.

[0193] In the embodiment of the present application, by obtaining the data to be tested and inputting the data to be tested into the sentiment prediction model, the sentiment category values of at least two sentiment dimensions corresponding to the data to be tested are output. Since the sentiment prediction model is trained based on multiple sets of training data, each set of training data includes: sample data and the corresponding labeled sentiment category values. The sentiment prediction model has learned the mapping relationship between the sample data and the labeled sentiment category values during the training stage. When new data to be tested is input, the sentiment prediction model calculates and outputs the sentiment category values on at least two sentiment dimensions corresponding to the data to be tested according to the internal algorithm of the model and the learned feature patterns. The sentiment category values reflect the tendency degree of the data to be tested on different sentiment dimensions. When the sentiment category value corresponding to any sentiment dimension is greater than or equal to the preset threshold, it indicates that the data to be tested shows a strong specific sentiment tendency on this sentiment dimension. This strong specific sentiment tendency is very likely related to fraud behavior. Therefore, a prompt message is output to remind of the possible fraud risk, which can help users make more cautious decisions and thus reduce the occurrence probability of fraud events.

[0194] The emotion category prediction method provided by the embodiments of this application may be executed by an emotion category prediction device. In the embodiments of this application, taking the emotion category prediction device executing the emotion category prediction method as an example, the emotion category prediction device provided by the embodiments of this application will be described.

[0195] Figure 4 It is a block diagram of an emotion category prediction device provided by the embodiments of this application. The device 400 includes:

[0196] An acquisition module 410, configured to acquire data to be measured;

[0197] An input module 420, configured to input the data to be measured into an emotion prediction model and output emotion category values of at least two emotion dimensions corresponding to the data to be measured; the emotion prediction model is obtained by training based on multiple groups of training data, and each group of the training data includes: sample data and the labeled emotion category value corresponding to the sample data;

[0198] An output module 430, configured to output a prompt message when the emotion category value corresponding to any one of the emotion dimensions is greater than or equal to a preset threshold.

[0199] In a possible embodiment, the input module 420 includes:

[0200] An extraction module, configured to input the data to be measured into the emotion prediction model and extract the emotion feature to be measured of the data to be measured;

[0201] A classification module, configured to classify the emotion feature to be measured and output emotion category values of at least two emotion dimensions corresponding to the data to be measured.

[0202] In a possible embodiment, the extraction module is specifically configured to:

[0203] When the data to be measured includes text data, extract the text emotion feature of the text data;

[0204] When the data to be measured includes voice data, extract the voice emotion feature of the voice data, and the voice emotion feature includes: timbre feature, pitch feature and speech rate feature;

[0205] When the data to be measured includes video data, extract the video emotion feature of the video data, and the video emotion feature includes: expression feature and body feature.

[0206] In a possible embodiment, the emotion feature to be measured includes at least two of text emotion feature, voice emotion feature and video emotion feature. The classification module is specifically configured to:

[0207] Perform weighted fusion on at least two of the text emotion features, the speech emotion features, and the video emotion features to obtain fused emotion features;

[0208] Classify the fused emotion features and output emotion category values of at least two emotion dimensions corresponding to the data to be measured.

[0209] In a possible embodiment, the apparatus 400 may further include:

[0210] A receiving module, configured to receive feedback information input by a user for the prompt information;

[0211] An updating module, configured to update model parameters of the emotion prediction model according to the feedback information and the data to be measured.

[0212] In a possible embodiment, the input module 420 is further configured to input the sample data into an initial prediction model and output sample emotion category values on at least two emotion dimensions corresponding to the sample data;

[0213] The apparatus 400 may further include:

[0214] A calculation module, configured to calculate a loss value according to the sample emotion category values and the labeled emotion category values;

[0215] An adjustment module, configured to adjust model parameters of the initial prediction model according to the loss value until a preset training condition is met, and obtain the emotion prediction model.

[0216] In an embodiment of the present application, by obtaining data to be measured, inputting the data to be measured into an emotion prediction model, and outputting emotion category values of at least two emotion dimensions corresponding to the data to be measured. Since the emotion prediction model is trained according to multiple sets of training data, each set of training data includes: sample data and labeled emotion category values corresponding to the sample data. The emotion prediction model has learned the mapping relationship between the sample data and the labeled emotion category values during the training phase. When new data to be measured is input, the emotion prediction model calculates and outputs emotion category values on at least two emotion dimensions corresponding to the data to be measured according to the algorithms inside the model and the learned feature patterns. The emotion category values reflect the tendency degrees of the data to be measured on different emotion dimensions. When the emotion category value corresponding to any emotion dimension is greater than or equal to a preset threshold, it indicates that the data to be measured shows a strong specific emotion tendency on this emotion dimension. This strong specific emotion tendency is very likely to be related to fraud behavior. Therefore, a prompt information is output to remind of possible fraud risks, which can help users make more cautious decisions and thus reduce the occurrence probability of fraud events.

[0217] The emotion category prediction device in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than terminals. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0218] The emotion category prediction device in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0219] The emotion category prediction device provided in the embodiments of the present application can implement each process implemented in the above method embodiments. To avoid repetition, it will not be described in detail here.

[0220] Optionally, as Figure 5 shown, the embodiments of the present application further provide an electronic device 510, including a processor 511, a memory 512, a program or instruction stored on the memory 512 and executable on the processor 511. When the program or instruction is executed by the processor 511, it implements each step of any of the above emotion category prediction method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described in detail here.

[0221] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0222] Figure 6 It is a schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0223] The electronic device 600 includes, but is not limited to, components such as a radio frequency unit 601, a network module 602, an audio output unit 603, an input unit 604, a sensor 605, a display unit 606, a user input unit 607, an interface unit 608, a memory 609, and a processor 610, etc.

[0224] Those skilled in the art can understand that the electronic device 600 may further include a power supply (such as a battery) for powering each component. The power supply can be logically connected to the processor 610 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 6 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0225] Among them, the processor 610 is used to obtain the data to be measured;

[0226] The processor 610 is further used to input the data to be measured into the emotion prediction model and output the emotion category values of at least two emotion dimensions corresponding to the data to be measured; the emotion prediction model is trained according to multiple sets of training data, and each set of the training data includes: sample data and the label emotion category value corresponding to the sample data;

[0227] The processor 610 is further used to output a prompt message when the emotion category value corresponding to any one of the emotion dimensions is greater than or equal to a preset threshold.

[0228] Optionally, the processor 610 is further used as an extraction module to input the data to be measured into the emotion prediction model and extract the emotion features to be measured of the data to be measured;

[0229] The processor 610 is further used to classify the emotion features to be measured and output the emotion category values of at least two emotion dimensions corresponding to the data to be measured.

[0230] Optionally, when the data to be measured includes text data, the processor 610 is further used to extract the text emotion features of the text data;

[0231] When the data to be measured includes voice data, the processor 610 is further used to extract the voice emotion features of the voice data, and the voice emotion features include: timbre features, pitch features, and speech rate features;

[0232] When the data to be measured includes video data, the processor 610 is further used to extract the video emotion features of the video data, and the video emotion features include: expression features and body features.

[0233] Optionally, the emotional features to be measured include at least two of text emotional features, speech emotional features, and video emotional features. The processor 610 is further configured to perform weighted fusion on at least two of the text emotional features, the speech emotional features, and the video emotional features to obtain a fused emotional feature;

[0234] The processor 610 is further configured to classify the fused emotional feature and output emotional category values of at least two emotional dimensions corresponding to the data to be measured.

[0235] Optionally, the user input unit 607 is configured to receive feedback information input by the user for the prompt information;

[0236] The processor 610 is further configured to update the model parameters of the emotion prediction model according to the feedback information and the data to be measured.

[0237] Optionally, the processor 610 is further configured to input the sample data into the initial prediction model and output sample emotional category values on at least two emotional dimensions corresponding to the sample data;

[0238] The processor 610 is further configured to calculate a loss value according to the sample emotional category value and the labeled emotional category value;

[0239] The processor 610 is further configured to adjust the model parameters of the initial prediction model according to the loss value until a preset training condition is met, and obtain the emotion prediction model.

[0240] In the embodiments of the present application, by obtaining the data to be measured, inputting the data to be measured into the emotion prediction model, and outputting emotional category values of at least two emotional dimensions corresponding to the data to be measured. Since the emotion prediction model is trained according to multiple sets of training data, each set of training data includes: sample data and the labeled emotional category value corresponding to the sample data. The emotion prediction model has learned the mapping relationship between the sample data and the labeled emotional category value during the training phase. When new data to be measured is input, the emotion prediction model calculates and outputs emotional category values on at least two emotional dimensions corresponding to the data to be measured according to the algorithm inside the model and the learned feature patterns. The emotional category value reflects the tendency degree of the data to be measured on different emotional dimensions. When the emotional category value corresponding to any emotional dimension is greater than or equal to a preset threshold, it indicates that the data to be measured shows a strong specific emotional tendency on this emotional dimension. This strong specific emotional tendency is likely to be related to fraud behavior. Therefore, a prompt information is output to remind of the possible fraud risk, which can help users make more cautious decisions, thereby reducing the occurrence probability of fraud incidents.

[0241] It should be understood that in the embodiments of the present application, the input unit 604 may include a Graphics Processing Unit (GPU) 6041 and a microphone 6042. The GPU 6041 processes the image data of static pictures or video images obtained by an image capturing device (such as a camera) in a video image capturing mode or an image capturing mode. The display unit 606 may include a display panel 6061, and the display panel 6061 may be configured in the form of, for example, a liquid crystal display, an organic light emitting diode, etc. The user input unit 607 includes at least one of a touch panel 6071 and other input devices 6072. The touch panel 6071 is also referred to as a touch screen. The touch panel 6071 may include two parts: a touch detection device and a touch controller. The other input devices 6072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an action bar, which will not be elaborated here. The memory 609 may be used to store software programs and various data, including but not limited to application programs and operating systems. The processor 610 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 610.

[0242] The memory 609 can be used to store software programs and various data. The memory 609 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 609 can include volatile memory or non-volatile memory, or the memory 609 can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 609 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.

[0243] The processor 610 may include one or more processing units; optionally, the processor 610 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 610 either.

[0244] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above-mentioned embodiment of the emotion category prediction method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0245] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.

[0246] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned embodiment of the emotion category prediction method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0247] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0248] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above-mentioned embodiment of the emotion category prediction method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0249] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0250] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0251] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A sentiment category prediction method, characterized in that: The method comprises: Get the data to be tested; The test data is input into a sentiment prediction model, and sentiment category values ​​of at least two sentiment dimensions corresponding to the test data are output; the sentiment prediction model is obtained by training based on multiple sets of training data, each set of training data includes: sample data and label sentiment category values ​​corresponding to the sample data; When the emotion category value corresponding to any one of the emotion dimensions is greater than or equal to a preset threshold, a prompt message is output.

2. The method according to claim 1, characterized in that The step of inputting the test data into a sentiment prediction model and outputting sentiment category values ​​of at least two sentiment dimensions corresponding to the test data comprises: Inputting the data to be tested into the emotion prediction model to extract the emotion features to be tested of the data to be tested; The emotion features to be tested are classified, and emotion category values ​​of at least two emotion dimensions corresponding to the data to be tested are output.

3. The method according to claim 2, characterized in that The step of inputting the data to be tested into the emotion prediction model and extracting the emotion features of the data to be tested comprises: In the case where the data to be tested includes text data, extracting text sentiment features of the text data; In the case where the data to be tested includes speech data, extracting speech emotion features of the speech data, wherein the speech emotion features include: timbre features, pitch features, and speech speed features; In the case that the data to be tested includes video data, video emotion features of the video data are extracted, and the video emotion features include: facial expression features and body features.

4. The method according to claim 2, characterized in that: The emotion features to be tested include: at least two of text emotion features, speech emotion features and video emotion features, and the classifying of the emotion features to be tested and outputting emotion category values ​​of at least two emotion dimensions corresponding to the data to be tested includes: Performing weighted fusion on at least two of the text emotion feature, the voice emotion feature and the video emotion feature to obtain a fused emotion feature; The fused emotional features are classified, and emotional category values ​​of at least two emotional dimensions corresponding to the data to be tested are output.

5. The method according to claim 1, characterized in that In the case where the emotion category value on any of the emotion dimensions is greater than or equal to a preset threshold, after outputting the prompt information, the method further includes: receiving feedback information input by the user in response to the prompt information; The model parameters of the emotion prediction model are updated according to the feedback information and the data to be tested.

6. The method according to claim 1, characterized in that Before inputting the test data into the emotion prediction model and outputting the emotion category values ​​of at least two emotion dimensions corresponding to the test data, the method further includes: Inputting the sample data into an initial prediction model, and outputting sample emotion category values ​​on at least two emotion dimensions corresponding to the sample data; Calculate the loss value according to the sample sentiment category value and the label sentiment category value; The model parameters of the initial prediction model are adjusted according to the loss value until preset training conditions are met to obtain the emotion prediction model.

7. An emotion prediction device, characterized in that: The device comprises: An acquisition module is used to obtain the data to be tested; An input module is used to input the test data into the emotion prediction model, and output the emotion category values ​​of at least two emotion dimensions corresponding to the test data; the emotion prediction model is trained based on multiple sets of training data, each set of training data includes: sample data and label emotion category values ​​corresponding to the sample data; The output module is used to output prompt information when the emotion category value corresponding to any of the emotion dimensions is greater than or equal to a preset threshold.

8. The device according to claim 7, characterized in that The input module comprises: An extraction module, used for inputting the data to be tested into the emotion prediction model to extract the emotion features to be tested of the data to be tested; The classification module is used to classify the emotional features to be tested and output emotional category values ​​of at least two emotional dimensions corresponding to the data to be tested.

9. The device according to claim 8, characterized in that The extraction module is specifically used for: In the case where the data to be tested includes text data, extracting text sentiment features of the text data; In the case where the data to be tested includes speech data, extracting speech emotion features of the speech data, wherein the speech emotion features include: timbre features, pitch features, and speech speed features; In the case that the data to be tested includes video data, video emotion features of the video data are extracted, and the video emotion features include: facial expression features and body features.

10. The device according to claim 8, characterized in that The emotion features to be tested include: at least two of text emotion features, voice emotion features and video emotion features. The classification module is specifically used for: Performing weighted fusion on at least two of the text emotion feature, the voice emotion feature and the video emotion feature to obtain a fused emotion feature; The fused emotional features are classified, and emotional category values ​​of at least two emotional dimensions corresponding to the data to be tested are output.