Apparatus for evaluating sentiment cognition capability of large language model
By constructing an evaluation device for the emotional cognitive ability of a large language model, and utilizing an evaluation data generation module and multiple recognition and assessment modules, the subjectivity and simplification problems of existing evaluation methods are solved, and a comprehensive, objective, and quantitative evaluation of the emotional cognitive ability of a large language model is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-17
- Publication Date
- 2026-03-20
AI Technical Summary
Existing methods for assessing affective cognition using large language models are highly subjective, overly simplistic, lack depth and precision, and lack unified and objective assessment standards in real-world emotional communication.
A device for evaluating the emotional cognitive ability of a large language model is provided, including an evaluation data generation module, a key event, mixed event, implicit emotion and intent recognition and evaluation module, which evaluates the PASS rate and WIN rate of response statements through a two-dimensional binary classification model, generates evaluation scores for key events, mixed events, implicit emotion and intent, and obtains a comprehensive evaluation score for emotional cognitive ability through a comprehensive evaluation calculation module.
It enables a comprehensive and objective quantitative assessment of the emotional cognitive ability of large language models, providing a unified standard that can quantify the model's ability to recognize key events, mixed events, implicit emotions, and intentions.
Smart Images

Figure CN118673117B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of affective cognition, and specifically relates to a device for evaluating the affective cognition capability of a large language model. BACKGROUND
[0002] In recent years, evaluating the affective cognition of large language models has attracted increasing attention. Early efforts have mainly focused on atomic tasks of sentiment recognition, such as aspect-based sentiment analysis (ABSA) by Tang et al. in 2015, particularly in the context of single text passages, such as target-related sentiment classification. With the improvement of dialogue models, the work by Li et al. in 2023 extended the evaluation range to sentiment analysis in a dialogue environment. In addition, some evaluations by Zhao et al. in 2023 even explored empathetic dialogue.
[0003] Currently, although large language models have made significant achievements in handling emotional dialogue and exhibiting emotional resonance, the overall evaluation of the affective cognition of these models is still limited, and existing evaluations often have subjective or simplification problems. For example, the evaluation by Elyoseph et al. in 2023 relies on subjective evaluation, while the evaluation by Schaaff et al. in 2023 uses low-difficulty tasks, failing to fully measure the affective cognition of the model.
[0004] In summary, current technologies mostly remain at a basic level, lacking depth and precision. More importantly, in handling real-world emotional communication, there is a lack of unified and objective standards or benchmarks to quantify and evaluate the performance of large language models in empathetic ability. SUMMARY
[0005] The present application is made to solve the above problems, and aims to provide a device for evaluating the affective cognition capability of a large language model.
[0006] The application provides a large language model emotion recognition ability evaluation device, which has the following characteristics: an evaluation data generation module stores a plurality of test statements for generating a reply statement corresponding to each test statement by a large language model to be evaluated, and the reply statement includes a key event reply statement, a mixed event reply statement, an implied emotion reply statement, and an intent reply statement; a key event recognition evaluation module includes a first evaluation model for evaluating each key event reply statement to obtain a key event evaluation score; a mixed event recognition evaluation module includes a second evaluation model for evaluating each mixed event reply statement to obtain a mixed event evaluation score; an implied emotion recognition evaluation module includes a third evaluation model for evaluating each implied emotion reply statement to obtain an implied emotion evaluation score; an intent recognition evaluation module includes a fourth evaluation model for evaluating each intent reply statement to obtain an intent evaluation score; and a comprehensive evaluation calculation module is used to obtain a comprehensive evaluation score according to the key event evaluation score, the mixed event evaluation score, the implied emotion evaluation score, and the intent evaluation score, wherein the emotion recognition ability result includes the key event evaluation score, the mixed event evaluation score, the implied emotion evaluation score, the intent evaluation score, and the comprehensive evaluation score.
[0007] In the large language model emotion recognition ability evaluation device provided by the application, the test statement can include a key event recognition statement, a mixed event recognition statement, an implied emotion recognition statement, and an intent recognition statement, the key event recognition statement includes one important event content and one daily event content, the mixed event recognition statement includes two important event contents, the implied emotion recognition statement includes one implied emotion, and the intent recognition statement includes one real purpose.
[0008] In the large language model emotion recognition ability evaluation device provided by the application, the reply statement generated by the large language model to be evaluated according to the key event recognition statement is a key event reply statement, the reply statement generated by the large language model to be evaluated according to the mixed event recognition statement is a mixed event reply statement, the reply statement generated by the large language model to be evaluated according to the implied emotion recognition statement is an implied emotion reply statement, and the reply statement generated by the large language model to be evaluated according to the intent recognition statement is an intent reply statement.
[0009] In the device for evaluating the sentiment cognitive ability of the large language model provided by the application, the first evaluation model can be a two-dimensional binary classification model, which is used to generate corresponding first PASS rates and first WIN rates according to the key event reply statements, the key event evaluation score includes the sum of the first PASS rates and the sum of the first WIN rates, the first evaluation model is used to determine whether the key event reply statement contains reply content to important event content, if yes, the first PASS rate is 1, if not, the first PASS rate is 0, the first evaluation model is used to determine whether the key event reply statement is an empathetic reply to important event content, if yes, the first WIN rate is 1, if not, the first WIN rate is 0.
[0010] In the device for evaluating the sentiment cognitive ability of the large language model provided by the application, the second evaluation model can be a two-dimensional binary classification model, which is used to generate corresponding second PASS rates and second WIN rates according to the mixed event reply statements, the mixed event evaluation score includes the sum of the second PASS rates and the sum of the second WIN rates, the second evaluation model is used to determine whether the mixed event reply statement contains reply content to two important event contents, if yes, the second PASS rate is 1, if not, the second PASS rate is 0, the second evaluation model is used to determine whether the mixed event reply statement is an empathetic reply to two important event contents, if yes, the second WIN rate is 1, if not, the second WIN rate is 0.
[0011] In the device for evaluating the sentiment cognitive ability of the large language model provided by the application, the third evaluation model can be a two-dimensional binary classification model, which is used to generate corresponding third PASS rates and third WIN rates according to the implicit sentiment reply statements, the implicit sentiment evaluation score includes the sum of the third PASS rates and the sum of the third WIN rates, the third evaluation model is used to determine whether the implicit sentiment reply statement contains reply content to implicit sentiment, if yes, the third PASS rate is 1, if not, the third PASS rate is 0, the third evaluation model is used to determine whether the implicit sentiment reply statement is an empathetic reply to implicit sentiment, if yes, the third WIN rate is 1, if not, the third WIN rate is 0.
[0012] In the device for evaluating the sentiment cognitive ability of a large language model provided by the application, the fourth evaluation model can be a two-dimensional binary classification model, which is used to generate a corresponding fourth PASS rate and a fourth WIN rate according to the intention reply statement, the intention evaluation score includes the sum of the fourth PASS rate and the sum of the fourth WIN rate, the fourth evaluation model is used to determine whether the intention reply statement contains the reply content for the true purpose, if yes, the fourth PASS rate is 1, if not, the fourth PASS rate is 0, the fourth evaluation model is used to determine whether the intention reply statement contains the suggestion or help for the true purpose, if yes, the fourth WIN rate is 1, if not, the fourth WIN rate is 0.
[0013] In the device for evaluating the sentiment cognitive ability of a large language model provided by the application, the content of the test statement can be constructed according to a plurality of different scenarios, and the scenarios include achievements, family and friends, health conditions, economic conditions and accidents.
[0014] In the device for evaluating the sentiment cognitive ability of a large language model provided by the application, the comprehensive evaluation calculation module can calculate the average value as the comprehensive evaluation score according to the key event evaluation score, the mixed event evaluation score, the implied sentiment evaluation score and the intention evaluation score.
[0015] Effects of the application
[0016] According to the device for evaluating the sentiment cognitive ability of a large language model, the reply statement of the large language model to be evaluated in four aspects of key event recognition, mixed event recognition, implied sentiment recognition and intention recognition is obtained through the evaluation data generation module; then, the PASS rate and the WIN rate of the corresponding reply statement are scored through the key event recognition evaluation module, the mixed event recognition evaluation module, the implied sentiment recognition evaluation module and the intention recognition evaluation module, so as to obtain the quantitative score of the large language model to be evaluated in the four aspects; finally, the comprehensive quantitative score of the sentiment cognitive ability is obtained by calculating the four quantitative scores through the comprehensive evaluation calculation module. Therefore, the device for evaluating the sentiment cognitive ability of a large language model can obtain a comprehensive and objective quantitative result of the sentiment cognitive ability of the large language model. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is a block diagram of the sentiment cognitive ability evaluation device in the embodiment of the application;
[0018] Figure 2 is a schematic diagram of sentiment cognitive evaluation in the embodiment of the application;
[0019] Figure 3is a flowchart of the evaluation of the sentiment cognitive ability of a large language model in an embodiment of the present application. DETAILED DESCRIPTION
[0020] In order to make the technical means, creative features, purposes and effects of the present application easy to understand, the following embodiments will be described in detail in combination with the drawings to explain the evaluation device of the sentiment cognitive ability of a large language model.
[0021] In this embodiment, an evaluation device of the sentiment cognitive ability of a large language model, hereinafter referred to as a sentiment cognitive ability evaluation device, is provided for obtaining the sentiment cognitive ability result of a large language model to be evaluated. The sentiment cognitive ability result includes a key event evaluation score, a mixed event evaluation score, an implied sentiment evaluation score, an intention evaluation score, and a comprehensive evaluation score.
[0022] Figure 1 is a block diagram of the sentiment cognitive ability evaluation device in an embodiment of the present application.
[0023] As shown in Figure 1 , the sentiment cognitive ability evaluation device 100 includes a data input module 10, an evaluation data generation module 20, a key event identification evaluation module 30, a mixed event identification evaluation module 40, an implied sentiment identification evaluation module 50, an intention identification evaluation module 60, a comprehensive evaluation calculation module 70, and a control module 80 for controlling the operation of the above-mentioned modules.
[0024] The data input module 10 is used to input a large language model to be evaluated.
[0025] The evaluation data generation module 20 stores a plurality of test statements for the large language model to be evaluated to generate a reply statement corresponding to each test statement.
[0026] The test statements include key event identification statements, mixed event identification statements, implied sentiment identification statements, and intention identification statements. The content of the test statements is constructed according to different situations such as achievements, family and friends, health status, economic status, and accidents. The reply statements include key event reply statements, mixed event reply statements, implied sentiment reply statements, and intention reply statements.
[0027] The key event identification statement includes one important event content and one daily event content. The mixed event identification statement includes two important event contents. The content of the implied sentiment identification statement contains one implied sentiment. The content of the intention identification statement contains one real purpose.
[0028] Figure 2 is a flowchart of the evaluation of the sentiment cognitive ability of a large language model in an embodiment of the present application.
[0029] As shown in Figure 2As shown, affective cognitive assessment includes critical event identification, mixed event identification, implicit sentiment identification, and intent identification. In critical event identification, the large language model to be evaluated needs to generate critical event response statements related to the critical events based on the critical event identification statement. In mixed event identification, the large language model to be evaluated needs to generate mixed event response statements related to the mixed events based on the mixed event identification statement. In implicit sentiment identification, the large language model to be evaluated needs to generate implicit sentiment response statements related to the implicit sentiment based on the implicit sentiment identification statement. In intent identification, the large language model to be evaluated needs to generate intent response statements related to the intent based on the intent identification statement.
[0030] The critical event identification and evaluation module 30 includes a first evaluation model, which is used to evaluate each critical event response statement separately to obtain a critical event evaluation score.
[0031] The first evaluation model is a two-dimensional binary classification model, used to generate corresponding first PASS rate and first WIN rate based on the key event response statements. The key event evaluation score includes the sum of the first PASS rates and the sum of the first WIN rates. In this embodiment, the performance quantification score of the large language model to be evaluated in key event recognition, i.e., the key event evaluation score, is obtained based on 100 first PASS rates and 100 first WIN rates.
[0032] The first evaluation model is used to determine whether the response statement for a critical event contains a response to the content of the important event. If so, the first pass rate is 1; otherwise, the first pass rate is 0.
[0033] The first evaluation model is also used to determine whether the response statement to a critical event is an empathetic response to the content of the important event. If so, the first win rate is 1; if not, the first win rate is 0.
[0034] In this embodiment, the critical event identification and evaluation module 30 is used to evaluate the critical event identification capability of the large language model to be evaluated. Critical event identification focuses on identifying and understanding the important events expressed by the user and their emotional impact. Therefore, for critical event identification statements, the large language model to be evaluated should identify the important event content and respond accordingly.
[0035] For example, the key event recognition statement is "I went to the hospital to visit my sick mother today, and then went to the supermarket", the important event content is "went to the hospital to visit the sick mother", and the daily event is "went to the supermarket". When the key event reply statement is "Is your mother all right?", it contains important event content and is a sympathetic reply, so the first PASS rate is 1 and the first WIN rate is 1. When the key event reply statement is "It's really troublesome that your mother is sick, I hope she won't keep you too busy", it contains important event content but is not a sympathetic reply, so the first PASS rate is 1 and the first WIN rate is 0.
[0036] The mixed event recognition evaluation module 40 comprises a second evaluation model for evaluating each mixed event reply statement respectively to obtain a mixed event evaluation score.
[0037] The second evaluation model is a two-dimensional binary classification model for generating corresponding second PASS rates and second WIN rates according to the mixed event reply statements. The mixed event evaluation score includes the sum of the second PASS rates and the sum of the second WIN rates. In this embodiment, according to 100 second PASS rates and 100 second WIN rates, the performance quantification score of the large language model to be evaluated in mixed event recognition, i.e., the mixed event evaluation score, is obtained.
[0038] The second evaluation model is used to determine whether the mixed event reply statement contains reply content to two important event contents. If yes, the second PASS rate is 1, and if no, the second PASS rate is 0.
[0039] The second evaluation model is also used to determine whether the mixed event reply statement is a sympathetic reply to the two important event contents. If yes, the second WIN rate is 1, and if no, the second WIN rate is 0.
[0040] In this embodiment, the mixed event recognition evaluation module 40 is used to evaluate the mixed event recognition capability of the large language model to be evaluated. Mixed event recognition focuses on responding to two events with similar importance in the user statement, i.e., the mixed event recognition statement, which is different from key event recognition, which deals with one important event. Therefore, for the mixed event recognition statement, the large language model to be evaluated should reply to all important event contents with rich empathetic ability.
[0041] For example, the mixed event recognition statement is "I was promoted, but also need to move to a new city", and the important event content therein is the two events of "being promoted" and "moving to a new city". When the mixed event reply statement is "Congratulations on your promotion! Is moving to a new city a challenge for you?", it contains a co-empathetic reply to the two important events, and thus the second PASS rate is 1 and the second WIN rate is 1. When the mixed event reply statement is "Great, you were promoted! That's a huge achievement!", it contains a co-empathetic reply to one important event, and thus the second PASS rate is 0 and the second WIN rate is 0.
[0042] The implied sentiment recognition evaluation module 50 comprises a third evaluation model for evaluating each implied sentiment reply statement respectively to obtain an implied sentiment evaluation score.
[0043] The third evaluation model is a two-dimensional binary classification model for generating a corresponding third PASS rate and a third WIN rate according to the implied sentiment reply statement. The implied sentiment evaluation score comprises a sum of the third PASS rates and a sum of the third WIN rates. In this embodiment, according to 100 third PASS rates and 100 third WIN rates, the performance quantification score of the large language model to be evaluated in implied sentiment recognition, i.e., the implied sentiment evaluation score, is obtained.
[0044] The third evaluation model is used to determine whether the implied sentiment reply statement contains reply content to the implied sentiment. If yes, the third PASS rate is 1, and if no, the third PASS rate is 0.
[0045] The third evaluation model is also used to determine whether the implied sentiment reply statement is a co-empathetic reply to the implied sentiment. If yes, the third WIN rate is 1, and if no, the third WIN rate is 0.
[0046] In this embodiment, the implied sentiment recognition evaluation module 50 is used to evaluate the implied sentiment recognition capability of the large language model to be evaluated. Implied sentiment recognition refers to recognizing the potential deep emotion in the user statement, i.e., the implied sentiment recognition statement. In some scenarios, although the statement only contains one event, the emotion is implied by language rather than directly expressed. Therefore, for a mixed event recognition statement, the large language model to be evaluated should recognize the implied sentiment behind the implied sentiment recognition statement and provide targeted emotional support.
[0047] For example, the implicit emotion recognition statement is "I feel like I have gained weight", and the implicit emotion is "sad". When the implicit emotion reply statement is "Do you feel you have not gained weight? Your belly seems to have shrunk?", the large language model to be evaluated correctly identifies "sad" and gives a targeted comfort, i.e., an empathetic reply, so the third PASS rate is 1 and the third WIN rate is 1. When the implicit emotion reply statement is "Exercise more", the large language model to be evaluated does not correctly identify "sad" and is a normal reply, i.e., without empathy, so the third PASS rate is 0 and the third WIN rate is 0.
[0048] The intent recognition evaluation module 60 comprises a fourth evaluation model for evaluating each intent reply statement to obtain an intent evaluation score.
[0049] The fourth evaluation model is a two-dimensional binary classification model for generating a corresponding fourth PASS rate and a fourth WIN rate according to the intent reply statement. The intent evaluation score includes the sum of the fourth PASS rate and the sum of the fourth WIN rate. In this embodiment, according to 100 fourth PASS rates and 100 fourth WIN rates, the performance quantification score of the large language model to be evaluated in intent recognition, i.e., the intent evaluation score, is obtained.
[0050] The fourth evaluation model is used to determine whether the intent reply statement contains reply content to the true purpose. If yes, the fourth PASS rate is 1, and if no, the fourth PASS rate is 0.
[0051] The fourth evaluation model is also used to determine whether the intent reply statement contains suggestions or help for the true purpose. If yes, the fourth WIN rate is 1, and if no, the fourth WIN rate is 0.
[0052] In this embodiment, the intent recognition evaluation module 60 is used to evaluate the intent recognition capability of the large language model to be evaluated. Intent recognition aims to understand the underlying intent or needs behind the user statement, i.e., the intent recognition statement, and provide specific help or solutions. Therefore, for the intent recognition statement, the large language model to be evaluated should give a response related to the intent or needs.
[0053] For example, the intent recognition statement is "I am busy with work all day", and the true purpose is "not eating". When the intent reply statement is "Have you eaten? I can help you order some takeout.", the true intent is recognized and specific help is given, so the fourth PASS rate is 1 and the fourth WIN rate is 1. When the intent reply statement is "Remember to eat", it is only a simple emotional support without specific help, so the fourth PASS rate is 1 and the fourth WIN rate is 0.
[0054] The first evaluation model, the second evaluation model, the third evaluation model and the fourth evaluation model are all large language models, and the large language model in the embodiment is GPT-4. According to the different scoring standards, four training data sets corresponding to key event recognition, mixed event recognition, implicit sentiment recognition and intent recognition are constructed, and each training data set contains a plurality of reply sentences, which are divided into positive samples containing correct key reply information and negative samples not containing correct key information. The corresponding large language model is trained through each training data set, and the cross-entropy function is used as the loss function, so as to complete the training of the large language model.
[0055] In the embodiment, the score generated by the trained GPT-4 is subjected to Pearson correlation coefficient calculation with the score of the reply sentence given by the evaluation personnel, and the Pearson correlation coefficient is 0.991. It can be seen that the trained GPT-4 can realize accurate evaluation of the reply sentence, so as to realize reasonable quantification of the sentiment cognitive ability of the large language model to be evaluated.
[0056] The comprehensive evaluation calculation module 70 is configured to obtain a comprehensive evaluation score according to the key event evaluation score, the mixed event evaluation score, the implicit sentiment evaluation score and the intent evaluation score.
[0057] The comprehensive evaluation calculation module 70 calculates the average value as the comprehensive evaluation score according to the key event evaluation score, the mixed event evaluation score, the implicit sentiment evaluation score and the intent evaluation score.
[0058] The control module 80 stores a control program for controlling the operation of each module.
[0059] The process of evaluating the sentiment cognitive ability of the large language model by using the sentiment cognitive ability evaluation device 100 will be described below with reference to the accompanying drawings.
[0060] Figure 3 is a flowchart of the process of evaluating the sentiment cognitive ability of the large language model in the embodiment of the present application.
[0061] As shown in Figure 3 , the evaluation of the sentiment cognitive ability of the large language model includes the following steps:
[0062] Step S1, input the large language model to be evaluated by using the data input module 10.
[0063] Step S2, make the large language model to be evaluated generate reply sentences corresponding to each test statement by using the evaluation data generation module 20.
[0064] Step S3, evaluate each key event reply sentence by using the key event recognition evaluation module 30, and obtain a key event evaluation score.
[0065] Step S4, each mixed event reply sentence is respectively evaluated by using the mixed event recognition evaluation module 40 to obtain a mixed event evaluation score.
[0066] Step S5, each implicit sentiment reply sentence is respectively evaluated by using the implicit sentiment recognition evaluation module 50 to obtain an implicit sentiment evaluation score.
[0067] Step S6, each intent reply sentence is respectively evaluated by using the intent recognition evaluation module 60 to obtain an intent evaluation score.
[0068] Step S7, the comprehensive evaluation calculation module 70 obtains a comprehensive evaluation score according to the key event evaluation score, the mixed event evaluation score, the implicit sentiment evaluation score and the intent evaluation score.
[0069] Effects of the embodiments
[0070] According to the big language model sentiment cognition ability evaluation device, first, the reply sentences of the big language model to be evaluated in key event recognition, mixed event recognition, implicit sentiment recognition and intent recognition are obtained by the evaluation data generation module; then, the PASS rate and WIN rate scores of the corresponding reply sentences are scored by the key event recognition evaluation module, the mixed event recognition evaluation module, the implicit sentiment recognition evaluation module and the intent recognition evaluation module, so as to obtain the quantitative scores of the big language model to be evaluated in the four aspects; finally, the comprehensive quantitative score of the sentiment cognition ability is obtained by calculating the four quantitative scores by the comprehensive evaluation calculation module. In summary, the method can obtain the quantitative results of the comprehensive and objective big language model sentiment cognition ability.
[0071] Those skilled in the art should understand that the present application is not limited to the above embodiments, and the above embodiments and descriptions in the specification are only to illustrate the principles of the present application, and various changes and improvements can be made without departing from the spirit and scope of the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A device for evaluating the affective cognitive ability of a large language model, used to obtain the affective cognitive ability results of the large language model to be evaluated, characterized in that, include: The evaluation data generation module stores multiple test statements, which are used by the large language model to be evaluated to generate response statements corresponding to each test statement. The response statements include key event response statements, mixed event response statements, implied sentiment response statements, and intention response statements. The critical event identification and evaluation module includes a first evaluation model, which is used to evaluate each of the critical event response statements to obtain a critical event evaluation score. The hybrid event identification and evaluation module includes a second evaluation model, which is used to evaluate each of the hybrid event response statements to obtain a hybrid event evaluation score. The implicit sentiment recognition and evaluation module includes a third evaluation model, which is used to evaluate each of the implicit sentiment response statements to obtain an implicit sentiment evaluation score. The intent recognition and evaluation module includes a fourth evaluation model, which is used to evaluate each of the intent response statements to obtain an intent evaluation score. The comprehensive evaluation calculation module is used to obtain a comprehensive evaluation score based on the key event evaluation score, the mixed event evaluation score, the implicit emotion evaluation score, and the intention evaluation score. The emotional cognitive ability results include the critical event assessment score, the mixed event assessment score, the implicit emotion assessment score, the intention assessment score, and the comprehensive assessment score. The test statements include key event identification statements, mixed event identification statements, implicit sentiment identification statements, and intent identification statements. The key event identification statement includes one important event and one routine event. The hybrid event identification statement includes the contents of both of the important events. The implicit emotion recognition statement contains an implicit emotion. The intent recognition statement contains a genuine purpose. The response statement generated by the large language model to be evaluated based on the key event identification statement is the key event response statement. The response statement generated by the large language model to be evaluated based on the hybrid event recognition statement is the hybrid event response statement. The response statement generated by the large language model to be evaluated based on the implicit sentiment recognition statement is the implicit sentiment response statement. The response statement generated by the large language model to be evaluated based on the intent recognition statement is the intent response statement. The first evaluation model is a two-dimensional binary classification model, used to generate the corresponding first PASS rate and first WIN rate based on the key event response statements. The critical event assessment score includes the sum of the first PASS rate and the first WIN rate. The first evaluation model is used to determine whether the key event response statement contains a response to the important event content. If yes, the first PASS rate is 1; otherwise, the first PASS rate is 0. The first evaluation model is used to determine whether the key event response statement is an empathetic response to the content of the important event. If yes, the first WIN rate is 1; otherwise, the first WIN rate is 0. The second evaluation model is a two-dimensional binary classification model, used to generate the corresponding second PASS rate and second WIN rate based on the mixed event response statements. The mixed event evaluation score includes the sum of the second PASS rate and the sum of the second WIN rate. The second evaluation model is used to determine whether the mixed event response statement contains response content to the two important events. If yes, the second PASS rate is 1; otherwise, the second PASS rate is 0. The second evaluation model is used to determine whether the mixed event response statement is an empathetic response to the content of the two important events. If yes, the second WIN rate is 1; if no, the second WIN rate is 0.
2. The device for evaluating the emotional cognitive ability of a large language model according to claim 1, characterized in that: in, The third evaluation model is a two-dimensional binary classification model, used to generate the corresponding third PASS rate and third WIN rate based on the implicit sentiment response statements. The implicit sentiment assessment score includes the sum of the third pass rate and the sum of the third win rate. The third evaluation model is used to determine whether the implicit emotion response statement contains content responding to the implicit emotion. If yes, the third PASS rate is 1; otherwise, the third PASS rate is 0. The third evaluation model is used to determine whether the implicit emotion response statement is an empathetic response to the implicit emotion. If yes, the third WIN rate is 1; if no, the third WIN rate is 0.
3. The evaluation device for large language model emotional cognition ability according to claim 1, characterized in that: in, The fourth evaluation model is a two-dimensional binary classification model, used to generate the corresponding fourth PASS rate and fourth WIN rate based on the intended response statement. The intent assessment score includes the sum of the fourth PASS rate and the sum of the fourth WIN rate. The fourth evaluation model is used to determine whether the intended response statement contains content that addresses the true purpose. If yes, the fourth pass rate is 1; otherwise, the fourth pass rate is 0. The fourth evaluation model is used to determine whether the intention response statement contains suggestions or help for the true purpose. If yes, the fourth WIN rate is 1; otherwise, the fourth WIN rate is 0.
4. The evaluation device for large language model emotional cognition ability according to claim 1, characterized in that: in, The test statements are constructed based on multiple different scenarios. The scenarios include achievements, family and friends, health status, financial situation, and accidents.
5. The evaluation device for large language model emotional cognitive ability according to claim 1, characterized in that: in, The comprehensive evaluation calculation module calculates the average value of the key event evaluation score, the mixed event evaluation score, the implicit emotion evaluation score, and the intention evaluation score as the comprehensive evaluation score.