Method, device and equipment for evaluating dialogue memory ability of large model and medium

By adding markers to the dialogue sequence text and calculating the similarity score of the large-model dialogue content, the problem of difficulty in quantifying the large-model dialogue memory ability in the prior art is solved, and the accurate evaluation of the large-model memory ability is achieved.

CN120045656APending Publication Date: 2025-05-27CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411982453.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively quantify the dialogue memory capabilities of large models over different time intervals.

Method used

By obtaining the conversation sequence text, adding marks to the problems in it, and inputting the marked text into the big model for multiple rounds of dialogue processing. The similarity scores between the first and second answers corresponding to the target conversation round were calculated to evaluate the memory ability of the big model.

Benefits of technology

Accurate quantitative evaluation of the memory ability of the large model is achieved, which can effectively measure the memory ability of the large model in multiple rounds of conversations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045656A_ABST
    Figure CN120045656A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a large model dialogue memory ability evaluation method and device, equipment and a medium, and the method comprises the steps: obtaining a dialogue sequence text in which multiple groups of questions and answers are sequentially arranged; adding a mark for at least one question contained in the dialogue sequence text, and inputting the dialogue sequence text added with the mark into the to-be-tested large model, so that the to-be-tested large model is subjected to multiple rounds of dialogue processing, and each round of dialogue processing realizes that first answer content corresponding to the question is output for a group of questions and answers; determining a target dialogue round from dialogue rounds subjected to dialogue processing of the to-be-tested large model, and inputting a question corresponding to the target dialogue round into the to-be-tested large model to obtain second answer content output by the to-be-tested large model; and calculating a similarity score between the first answer content and the second answer content corresponding to the target dialogue round so as to evaluate the memory ability of the to-be-tested large model. Therefore, the memory ability of the large model can be accurately evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology. Specifically, it relates to a method, device, electronic device, and computer-readable storage medium for evaluating the dialogue memory ability of large models. Background Art

[0002] Currently, in the field of artificial intelligence technology, the ability of large models to remember and utilize previous information during conversations is an important part of multi-turn dialogue capabilities. For example, in the middle and late stages of a conversation, whether the model can accurately recall and utilize key information mentioned earlier. Although existing technologies consider the performance of large models in multi-turn conversations to a certain extent, there is a lack of effective quantitative evaluation methods for the memory ability of large models at different time intervals. Summary of the Invention

[0003] To solve the above technical problems, embodiments of this application provide a method, device, electronic device, and computer-readable storage medium for evaluating the dialogue memory ability of large models, which can accurately measure the memory ability of large models in multi-turn conversations.

[0004] According to one aspect of the embodiments of this application, a method for evaluating the dialogue memory ability of a large model is provided. The method includes: obtaining a dialogue sequence text, where multiple groups of questions and answers are arranged in sequence in the dialogue sequence text; adding marks to at least one question included in the dialogue sequence text, and inputting the dialogue sequence text with the added marks into the large model to be tested, so that the large model to be tested performs multi-turn dialogue processing, where each round of dialogue processing outputs a first answer content corresponding to the question for a group of questions and answers; when the large model to be tested outputs the corresponding first answer content for the dialogue sequence text with marks added, determining a target dialogue round from the dialogue rounds that the large model to be tested has performed dialogue processing on, and inputting the question corresponding to the target dialogue round into the large model to be tested to obtain a second answer content output by the large model to be tested; calculating a similarity score between the first answer content corresponding to the target dialogue round and the second answer content, and evaluating the memory ability of the large model to be tested based on the similarity score.

[0005] In another exemplary embodiment, a dialogue topic library containing dialogue data sets corresponding to different topics and a grammar template library corresponding to each topic are pre-constructed, and the grammar template is used to specify the format of questions and answers in the dialogue; the obtaining of the dialogue sequence text includes: obtaining a dialogue data set corresponding to a selected topic from the dialogue topic library; generating the dialogue sequence text according to the selected grammar template corresponding to the selected topic and the dialogue data set corresponding to the selected topic.

[0006] In another exemplary embodiment, adding a mark to at least one question included in the dialogue sequence text includes: selecting a first number of question-answer groups from the dialogue sequence text, and selecting a second number of question-answer groups from the first number of question-answer groups, where the first number is greater than the second number; adding a mark to the questions included in the second number of question-answer groups.

[0007] In another exemplary embodiment, determining the target dialogue turn from the dialogue turns in which the large model to be tested has performed dialogue processing includes: determining the current turn of the large model to be tested for the dialogue sequence text with the mark added; determining the target dialogue turn according to a set interval turn and the current turn.

[0008] In another exemplary embodiment, the method further includes: repeatedly executing replacing the mark position inserted in the dialogue sequence text, and inputting the dialogue sequence text with the replaced mark position into the large model to be tested to obtain a similarity score determined based on the replaced mark until the number of repetitions reaches a first preset number threshold; evaluating the memory ability of the large model to be tested according to the similarity score obtained each time.

[0009] In another exemplary embodiment, the method further includes: repeatedly executing to obtain a new dialogue sequence text and obtaining a corresponding similarity score based on the new dialogue sequence text until the number of repetitions reaches a second preset number threshold; evaluating the memory ability of the large model to be tested according to the similarity score obtained each time.

[0010] In another exemplary embodiment, adding a mark to at least two question texts included in the dialogue sequence text; calculating the similarity score between the first answer content and the second answer content corresponding to the target dialogue turn to evaluate the memory ability of the large model to be tested according to the similarity score includes: determining the similarity score obtained based on each mark; evaluating the memory ability of the large model to be tested according to the similarity score obtained based on each mark.

[0011] According to one aspect of the embodiments of the present application, there is provided an evaluation device for the dialogue memory ability of a large model, including: a dialogue sequence text acquisition module configured to acquire a dialogue sequence text, in which multiple groups of questions and answers are arranged in sequence; a first answer content acquisition module configured to add a mark to at least one question included in the dialogue sequence text, and input the dialogue sequence text with the mark added into the large model to be tested, so that the large model to be tested performs multi-round dialogue processing, wherein in each round of dialogue processing, a first answer content corresponding to the question is output for a group of questions and answers; a second answer content acquisition module configured to, when the large model to be tested outputs the corresponding first answer content for the dialogue sequence text with the mark added, determine a target dialogue round from the dialogue rounds in which the large model to be tested has performed dialogue processing, and input the question corresponding to the target dialogue round into the large model to be tested to obtain a second answer content output by the large model to be tested; an evaluation module configured to calculate a similarity score between the first answer content corresponding to the target dialogue round and the second answer content, so as to evaluate the memory ability of the large model to be tested according to the similarity score.

[0012] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the electronic device to implement the evaluation method for the dialogue memory ability of the large model as described above.

[0013] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer-readable instructions are stored, when the computer-readable instructions are executed by a processor of a computer, enabling the computer to execute the evaluation method for the dialogue memory ability of the large model as described above.

[0014] In the technical solution provided by the embodiments of the present application, since when the large language model outputs the first answer content of the target dialogue round, it uses the answer to the question as the answer reference, and when outputting the second answer content, only the question is input without inputting the question answer, this makes the second answer content be output by the large language model based on its memory of the question answer. Therefore, the similarity score between the first answer content corresponding to the target dialogue round and the second answer content accurately represents the memory ability of the large language model, enabling the memory ability of the large language model to be quantitatively evaluated.

[0015] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:

[0017] Figure 1 is a flowchart of a method for evaluating the dialogue memory ability of a large model shown in an exemplary embodiment of the present application;

[0018] Figure 2 is a flowchart of a method for obtaining a dialogue sequence text shown in an exemplary embodiment of the present application;

[0019] Figure 3 is a schematic diagram of selecting a topic from a dialogue topic library to generate a corresponding grammar template shown in an exemplary embodiment of the present application;

[0020] Figure 4 is a flowchart of adding marks to at least one question included in a dialogue sequence text shown in an exemplary embodiment of the present application;

[0021] Figure 5 is a schematic diagram of a dialogue sequence text with marks added shown in an exemplary embodiment of the present application;

[0022] Figure 6 is a flowchart of determining a target dialogue turn from the dialogue turns that have undergone dialogue processing by a large model to be tested shown in an exemplary embodiment of the present application;

[0023] Figure 7 is a flowchart of calculating a similarity score shown in an exemplary embodiment of the present application;

[0024] Figure 8 is a flowchart of a method for evaluating the dialogue memory ability of a large model shown in another exemplary embodiment of the present application;

[0025] Figure 9 is a flowchart of a method for evaluating the dialogue memory ability of a large model shown in another exemplary embodiment of the present application;

[0026] Figure 10 is a block diagram of an apparatus for evaluating the dialogue memory ability of a large model shown in an exemplary embodiment of the present application;

[0027] Figure 11 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed Embodiments

[0028] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numerals in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments identical to the present application. On the contrary, they are merely examples of apparatuses and methods that are the same as some aspects of the present application as detailed in the appended claims.

[0029] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in the form of application programs, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0030] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0031] It should be noted that "a plurality of" mentioned in the present application means two or more. " / " describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0032] Currently, in the field of artificial intelligence technology, the ability of large models to remember and utilize previous information during the conversation process is an important part of the multi-round conversation ability. For example, when the conversation reaches the middle and late stages, whether the model can accurately recall and utilize the key information mentioned earlier. Although the prior art to some extent considers the performance of large models in multi-round conversations, there is a lack of effective quantitative evaluation means for the memory ability of large models at different time intervals.

[0033] Based on this, the embodiments of the present application propose an evaluation method, apparatus, electronic device, and computer-readable storage medium for the conversation memory ability of large models.

[0034] Please refer to Figure 1 , Figure 1 which is a flowchart of the evaluation method for the conversation memory ability of large models shown in an exemplary embodiment of the present application.

[0035] Next, the evaluation method for the conversation memory ability of large models proposed in the embodiments of the present application will be introduced in detail with a computer as the specific execution entity.

[0036] AsFigure 1 As shown, in an exemplary embodiment, the method for evaluating the dialogue memory ability of a large model at least includes S110 to S140, which are introduced in detail as follows:

[0037] S110, obtain the dialogue sequence text, in which multiple groups of questions and answers are arranged in sequence.

[0038] S120, add marks to at least one question included in the dialogue sequence text, and input the dialogue sequence text with marks added into the large model to be tested, so that the large model to be tested performs multi-round dialogue processing, where each round of dialogue processing realizes outputting the first answer content corresponding to the question for a group of questions and answers.

[0039] S130, when the large model to be tested outputs the corresponding first answer content for the dialogue sequence text with marks added, determine the target dialogue round from the dialogue rounds that the large model to be tested has performed dialogue processing, and input the question corresponding to the target dialogue round into the large model to be tested to obtain the second answer content output by the large model to be tested.

[0040] S140, calculate the similarity score between the first answer content and the second answer content corresponding to the target dialogue round, so as to evaluate the memory ability of the large model to be tested according to the similarity score.

[0041] The above S110 - S140 are elaborated in detail below.

[0042] In S110, one dialogue sequence text can be selected from multiple dialogue sequence texts. In the dialogue sequence text, multiple groups of questions and answers are arranged in sequence.

[0043] The dialogue sequence text can be the following Example 1:

[0044] "Question 1: What are the famous tourist attractions in Area A?

[0045] Answer 1: There are many famous attractions in Area A, such as Attraction 1, Attraction 2, and Attraction 3. Attraction 1 is the landmark building in Area A, Attraction 2 houses many precious artworks, and Attraction 3 is a classic of architecture.

[0046] Question 2: Are there any recommended local delicacies nearby?

[0047] Answer 2: There are many local delicacies in Area A, such as Delicacy 1 and Delicacy 2. Delicacy 1 has a sweet taste, and Delicacy 2 is delicious."

[0048] Similarly, the dialogue sequence text can be the following Example 2:

[0049] "Question 1: What are some simple and effective ways to improve sleep quality?

[0050] Answer 1: First, maintain a regular schedule. Try to go to bed and get up at the same time every day to let your body form a biological clock. Avoid using electronic devices before going to bed because the blue light emitted by the screen will inhibit the secretion of melatonin and affect falling asleep. You can also create a quiet, comfortable and properly temperatureed sleeping environment, such as choosing a suitable mattress and pillow, and drawing the curtains.

[0051] Question 2: If you don't feel well the next day due to poor sleep, how can you quickly refresh yourself?

[0052] Answer 2: You can wash your face with cold water first to stimulate the facial nerves and make yourself more awake. You can also drink a cup of coffee or tea. The caffeine in them has the effect of refreshing the mind. You can also do some simple stretching exercises appropriately, move your body, promote blood circulation and drive away drowsiness.

[0053] Question 3: What harm will long-term poor sleep do to the body?

[0054] Answer 3: It may lead to memory loss, affecting the storage and retrieval of information in the brain, and reducing learning and work efficiency. In addition, it also has an adverse effect on the cardiovascular system. For example, it is easy to have unstable blood pressure, arrhythmia, etc. It may even affect emotions, making people become anxious, irritable or depressed.

[0055] In Example 1, Question 1 and Answer 1 can be combined into a group of question and answer. Similarly, Question 2 and Answer 2 can also be combined into a group of question and answer.

[0056] In the dialogue sequence text, there are multiple groups of questions and answers arranged in sequence. In Example 1, the dialogue sequence text includes two groups of questions and answers arranged in sequence. In Example 2, the dialogue sequence text includes three groups of questions and answers arranged in sequence. It should be noted that the questions and answers included in the dialogue sequence text can exceed three groups.

[0057] Please refer to Figure 2 , Figure 2 which is a flowchart for obtaining the dialogue sequence text shown in an exemplary embodiment of the present application. In another exemplary embodiment of the present application, there is a pre-constructed dialogue topic library containing dialogue data sets corresponding to different topics, and a grammar template library corresponding to each topic. The grammar template is used to specify the format of questions and answers in the dialogue. The steps S110 for obtaining the dialogue sequence text include the following S210 - S220:

[0058] S210, obtain the dialogue data set corresponding to the selected topic from the dialogue topic library.

[0059] S220, generate the dialogue sequence text according to the selected grammar template corresponding to the selected topic and the dialogue data set corresponding to the selected topic.

[0060] The above S210 - S220 will be elaborated in detail below.

[0061] Before S210, please refer to Figure 3 , Figure 3 which is a schematic diagram showing the selection of a topic from a dialogue topic library and generating a corresponding grammar template in an exemplary embodiment of the present application. As Figure 3 shown, a dialogue topic library containing dialogue datasets corresponding to different topics is pre - constructed. The dialogue topic library aims to cover various possible topic areas to provide basic data support for generating diverse dialogue sequence texts subsequently.

[0062] Different topics can be reasonably classified and labeled for subsequent quick and accurate retrieval. Exemplarily, they can be divided into topics such as travel planning, health and fitness, educational tutoring, technology consultation, and common sense of life.

[0063] After constructing the dialogue topic library containing dialogue datasets corresponding to different topics, a grammar template library corresponding to each topic is constructed. The grammar template library is used to unify and standardize the expression form of the dialogue, making the generated dialogue sequence text more reasonable, standardized, and in line with people's expression habits in terms of language structure.

[0064] In some embodiments of the present application, the grammar template library contains multiple grammar templates. Exemplarily, for the travel planning topic, corresponding grammar template 1, grammar template 2, and grammar template 3, etc. can be constructed; for the common sense of life topic, corresponding grammar template 1, grammar template 2, and grammar template 3, etc. can also be constructed.

[0065] In some embodiments of the present application, for the travel planning topic, the grammar template can be the following Example 3:

[0066] "What are the famous tourist attractions?

[0067] [Name of the scenic spot]

[0068] Is the transportation there convenient?

[0069] [Description of the transportation situation]."

[0070] Thus, using the test data of dialogue sequence texts based on multiple topics and multiple grammar templates can ensure the diversity and logic of the dialogue, and can more comprehensively evaluate the memory ability of the large model to be tested.

[0071] In S210, corresponding dialogue data sets can be divided according to different themes in the dialogue theme library. Exemplarily, the dialogue theme library includes a dialogue data set corresponding to travel planning; the dialogue theme library includes a dialogue data set corresponding to health and fitness; the dialogue theme library includes a dialogue data set corresponding to educational tutoring; the dialogue theme library includes a dialogue data set corresponding to technology consultation; the dialogue theme library includes a dialogue data set corresponding to common sense of life. Thus, after the selected theme is determined, the dialogue data set corresponding to the selected theme can be obtained from the dialogue theme library.

[0072] In some embodiments of the present application, an effective retrieval mechanism is set up to achieve fast and accurate acquisition of the dialogue data set corresponding to the selected theme. Exemplarily, methods such as keyword matching and theme label screening are adopted.

[0073] In addition, if the scale of the dialogue theme library is large, an indexing system is established to improve the efficiency of data search, reduce the search time and resource consumption, and ensure the smoothness of the entire process.

[0074] In S220, the obtained dialogue data set is combined with the selected grammar template corresponding to the selected theme, and the actual questions and answer contents in the data set are used to arrange and generate a dialogue sequence text according to the format specified by the grammar template.

[0075] In some embodiments of the present application, the question-and-answer contents in the obtained dialogue data set are appropriately adjusted and optimized to better adapt to the selected grammar template, improving the readability and fluency of the text.

[0076] In S120, in some embodiments of the present application, the mark can be a specific symbol (such as "[MARK]") or a specific keyword (such as "memory mark"). The mark can also be other specific symbols.

[0077] In some embodiments of the present application, marks can be randomly added to the questions included in the dialogue sequence text. Marks can be added to the questions included in the dialogue sequence text according to a set probability. The set probability can be 10%-20%.

[0078] In some embodiments of the present application, if marks are added to at least two questions included in the dialogue sequence text, the marks at different positions are distinguished.

[0079] Please refer to Figure 4 , Figure 4 which is a flowchart showing the addition of marks to at least one question included in the dialogue sequence text shown in an exemplary embodiment of the present application. In some other embodiments of the present application, the steps of adding marks to at least one question included in the dialogue sequence text include the following S310-S320:

[0080] S310, select a first number of question-answer groups from the dialogue sequence text, and select a second number of question-answer groups from the first number of question-answer groups, where the first number is greater than the second number.

[0081] S320, add marks to the questions included in the second number of question-answer groups.

[0082] In S310 - S320, in some embodiments of the present application, select 10 question-answer groups from the dialogue sequence text, and select 2 question-answer groups from the 10 question-answer groups, and add marks to the questions included in the 2 question-answer groups.

[0083] In other embodiments of the present application, select 5 question-answer groups from the dialogue sequence text, and select 1 question-answer group from the 5 question-answer groups, and add marks to the questions included in the 1 question-answer group.

[0084] It should be noted that the questions with marks added cannot be the first questions in the dialogue sequence text.

[0085] Please refer to Figure 5 , Figure 5 is a schematic diagram of the dialogue sequence text with marks added shown in an exemplary embodiment of the present application. As Figure 5 shown, in the above Example 1, the dialogue sequence text with marks added can be as follows in Example 4:

[0086] "Question 1: What are the famous tourist attractions in Area A?

[0087] Answer 1: There are many famous attractions in Area A, such as Attraction 1, Attraction 2, and Attraction 3. Attraction 1 is the landmark building in Area A, Attraction 2 houses many precious artworks, and Attraction 3 is a classic architectural work.

[0088] [MARK]Question 2: Are there any recommended local delicacies nearby?

[0089] Answer 2: There are many local delicacies in Area A, such as Delicacy 1 and Delicacy 2. Delicacy 1 has a sweet taste, and Delicacy 2 is delicious."

[0090] In the above Example 2, the dialogue sequence text with marks added can be as follows in Example 5:

[0091] "Question 1: What are some simple and effective ways to improve sleep quality?

[0092] Answer 1: First, maintain a regular schedule. Try to go to bed and wake up at the same time every day to let your body form a biological clock. Avoid using electronic devices before bedtime because the blue light emitted by the screen inhibits the secretion of melatonin and affects sleep. You can also create a quiet, comfortable and temperature-appropriate sleeping environment, such as choosing a suitable mattress and pillow and drawing the curtains.

[0093] Question 2: If you don't feel well the next day due to poor sleep, how can you quickly refresh yourself?

[0094] Answer 2: You can wash your face with cold water first to stimulate the facial nerves and make yourself more awake. You can also drink a cup of coffee or tea, as the caffeine in them has the effect of refreshing the mind. You can also do some simple stretching exercises appropriately to move your body, promote blood circulation and drive away drowsiness.

[0095] Question 3: What harms will long-term poor sleep cause to the body?

[0096] [MARK]Answer 3: It may lead to memory loss, affecting the storage and retrieval of information in the brain, and reducing learning and work efficiency. In addition, it also has an adverse effect on the cardiovascular system, such as being prone to unstable blood pressure, arrhythmia, etc. It may even affect emotions, making people become anxious, irritable or depressed.

[0097] In some embodiments of the present application, after adding marks to at least one question included in the dialogue sequence text, the marked dialogue sequence text is input into the large model to be tested, so that the large model to be tested starts a multi-round dialogue processing flow.

[0098] Thus, the large model to be tested will output the first answer content corresponding to Question 1 according to the first group of questions and answers. The large model to be tested will output the first answer content corresponding to Question 2 according to the second group of questions and answers. The large model to be tested will output the first answer content corresponding to Question 3 according to the third group of questions and answers.

[0099] In S130, when the large model to be tested inputs the marked dialogue sequence text into the large model to be tested, the large model to be tested will process the questions arranged in sequence in the dialogue sequence text and output the corresponding first answer content.

[0100] The target dialogue turn should be before the turn corresponding to the marked question.

[0101] Referring to the above Example 4, if the mark is added to Question 2, that is, when the large model to be tested outputs the corresponding first answer content for Question 2, the dialogue turns that have undergone dialogue processing are the first dialogue turn, and the target dialogue turn is also the first dialogue turn.

[0102] Referring to the above Example 5, if the mark is added to Question 3, that is, the first response content corresponding to Question 3 is output by the large model to be tested, then the conversation turns that have undergone conversation processing are the first conversation turn and the second conversation turn, and the first conversation turn and / or the second conversation turn can be selected as the target conversation turn.

[0103] Please refer to Figure 6 , Figure 6 FIG. is a flowchart for determining a target conversation turn from the conversation turns that have undergone conversation processing by the large model to be tested, shown in an exemplary embodiment of the present application. In some other embodiments of the present application, the steps for determining a target conversation turn from the conversation turns that have undergone conversation processing by the large model to be tested include the following S410-S420:

[0104] S410, determine the current turn of the large model to be tested for the conversation sequence text with a mark added.

[0105] S420, determine the target conversation turn according to the set interval turns and the current turn.

[0106] The above S410-S420 will be elaborated in detail below.

[0107] In S410, referring to the above Example 4, the current turn is the second conversation turn. Similarly, referring to the above Example 5, the current turn is the third conversation turn.

[0108] In S420, in some embodiments of the present application, the set interval turns is 3 times. The set interval turns is 5 times. The set interval turns can be other values.

[0109] In some embodiments of the present application, the target conversation turn = the current turn - the set interval turns.

[0110] After determining the target conversation turn, input the question corresponding to the target conversation turn into the large model to be tested again, and the second response content output by the large model to be tested can be obtained.

[0111] In some embodiments of the present application, referring to the above Example 4, if the target conversation turn is the first conversation turn, then input the question "What are the famous tourist attractions in Place A?" in the first conversation turn into the large model to be tested again to obtain the second response content output by the large model to be tested.

[0112] In some other embodiments of the present application, referring to the above Example 5, if the target conversation turn is the first conversation turn, then input the question "What are some simple and effective ways to improve sleep quality?" in the first conversation turn into the large model to be tested again to obtain the second response content output by the large model to be tested.

[0113] In S140, in some embodiments of the present application, referring to the above Example 4, if the target dialogue turn is the first dialogue turn, then for the question "What are the famous tourist attractions in Place A?", calculate the similarity score between the first answer content and the second answer content output by the large model to be tested.

[0114] In other embodiments of the present application, referring to the above Example 5, if the target dialogue turn is the first dialogue turn, then for the question "What are some simple and effective ways to improve sleep quality?", calculate the similarity score between the first answer content and the second answer content output by the large model to be tested.

[0115] It should be noted that the similarity score is proportional to the similarity. The more similar the first answer content and the second answer content corresponding to the target dialogue turn are, the higher the similarity score.

[0116] The memory ability of the large model to be tested is also proportional to the similarity score. The higher the similarity score between the first answer content and the second answer content corresponding to the target dialogue turn, the stronger the memory ability of the large model to be tested.

[0117] Please refer to Figure 7 , Figure 7 is a flowchart showing the calculation of the similarity score shown in an exemplary embodiment of the present application. In some embodiments of the present application, marks are added to at least two question texts included in the dialogue sequence text; the step S140 of calculating the similarity score between the first answer content and the second answer content corresponding to the target dialogue turn to evaluate the memory ability of the large model to be tested according to the similarity score includes the following S510 - S520:

[0118] S510, determine the similarity scores obtained based on each mark.

[0119] S520, evaluate the memory ability of the large model to be tested according to the similarity scores obtained based on each mark.

[0120] The above S510 - S520 will be elaborated in detail below.

[0121] In S510 - S520, in some embodiments of the present application, if two question texts included in the dialogue sequence text are added with marks, which are mark 1 and mark 2 respectively, then determine the similarity score obtained based on mark 1 and the similarity score obtained based on mark 2 respectively. After determining the similarity score obtained based on mark 1 and the similarity score obtained based on mark 2, the two similarity scores can be averaged, and the memory ability of the large model to be tested is evaluated according to the average value.

[0122] In some embodiments of the present application, if the three question texts included in the dialogue sequence text are added with marks, namely mark 3, mark 4, and mark 5 respectively, then the similarity scores obtained based on mark 3, the similarity score obtained based on mark 4, and the similarity score obtained based on mark 5 are determined respectively. After determining the similarity score obtained based on mark 3, the similarity score obtained based on mark 4, and the similarity score obtained based on mark 5, the average value of the three similarity scores can be taken, and the memory ability of the large model to be tested can be evaluated according to the average value.

[0123] As can be seen from the above S110 - S140, since when the large language model outputs the first answer content of the target dialogue turn, it uses the answer to the question as the answer reference, and when outputting the second answer content, only the question is input, and the question answer is not input. This makes the second answer content output by the large language model based on its memory of the question answer. Therefore, the similarity score between the first answer content and the second answer content corresponding to the target dialogue turn can more accurately characterize the memory ability of the large language model, enabling the memory ability of the large language model to be quantitatively evaluated.

[0124] Please refer to Figure 8 , Figure 8 which is a flowchart of an evaluation method for the dialogue memory ability of a large model shown in another exemplary embodiment of the present application. In some embodiments of the present application, the evaluation method for the dialogue memory ability of the large model further includes the following S610 - S620:

[0125] S610, repeatedly execute replacing the marked position of the inserted mark in the dialogue sequence text, and input the dialogue sequence text with the replaced marked position into the large model to be tested to obtain the similarity score determined based on the replaced mark until the number of repetitions reaches the first preset number threshold.

[0126] S620, evaluate the memory ability of the large model to be tested according to the similarity score obtained each time.

[0127] The above S610 - S620 will be elaborated in detail below.

[0128] In S610, the first preset number threshold can be 50 times. The first preset number threshold can be 30 times. The first preset number threshold can be other values.

[0129] In some embodiments of the present application, if the first preset number threshold is 50 times and the dialogue sequence text includes 110 rounds of dialogue, then the marked positions of the inserted mark can be replaced differently in this dialogue sequence. Exemplarily, the marked position of the first inserted mark is the question text in the 3rd round of dialogue, the marked position of the second inserted mark is the question text in the 5th round of dialogue, and the marked position of the nth inserted mark is the question text in the (2n + 1)th round of dialogue.

[0130] In some embodiments of the present application, if the first preset number threshold is 30 times and the dialogue sequence text includes 200 rounds of conversations, the marker positions for inserting markers can be changed at different positions in this dialogue sequence. Exemplarily, the marker position for the first insertion of the marker is the question text in the 4th round of conversation, the marker position for the second insertion of the marker is the question text in the 8th round of conversation, and the marker position for the nth insertion of the marker is the question text in the 4nth round of conversation.

[0131] In S620, in some embodiments of the present application, the average similarity score of the large model to be tested within all different round intervals is calculated, and the memory ability of the large model to be tested is evaluated according to the average similarity score.

[0132] Thus, by inserting markers into the dialogue sequence, multiple target time points for testing memory lengths can be flexibly set based on one dialogue sequence, enabling different memory length tests on the large model to be tested under the same test data, improving the effectiveness of the test. At the same time, it is possible to test the depth of the large model's memory ability as the number of dialogue rounds increases, especially to deeply evaluate the memory ability in the later stage of the dialogue rounds.

[0133] Please refer to Figure 9 , Figure 9 which is a flowchart of an evaluation method for the dialogue memory ability of a large model shown in another exemplary embodiment of the present application. In some embodiments of the present application, the evaluation method for the dialogue memory ability of the large model further includes the following S710 - S720:

[0134] S710, repeatedly execute the replacement to obtain a new dialogue sequence text, and obtain the corresponding similarity score based on the new dialogue sequence text until the number of repetitions reaches the second preset number threshold.

[0135] S720, evaluate the memory ability of the large model to be tested according to the similarity score obtained each time.

[0136] The above S710 - S720 will be elaborated in detail below.

[0137] In S710, the second preset number threshold can be 50 times. The second preset number threshold can be 30 times. The second preset number threshold can be greater than the first preset number threshold. The second preset number threshold can be equal to a preset number threshold. The second preset number threshold can be less than the first preset number threshold. The second preset number threshold can be other values.

[0138] In some embodiments of the present application, replace the dialogue sequence text 1 with the dialogue sequence text 2, and keep the marker positions of the original markers unchanged. Keep replacing until the number of repetitions of the replacement reaches the second preset number threshold, then terminate the replacement of the dialogue sequence text.

[0139] In other embodiments of the present application, the dialogue sequence text 1 is replaced with the dialogue sequence text 2, and the mark position of the original mark is also replaced. The replacement is continued until the number of repeated executions reaches a second preset number threshold, and then the replacement of the dialogue sequence text is terminated.

[0140] In S720, in some embodiments of the present application, the average similarity score of the large model to be tested in all different round intervals is calculated, and the memory ability of the large model to be tested is evaluated according to the average similarity score.

[0141] It should be noted that the above-mentioned evaluation method of the large model dialogue memory ability includes but is not limited to the application in the field of large model research and development and optimization and the field of intelligent customer service system improvement.

[0142] In the field of large model research and development, during the development process of large models, this method can be used to test different versions of the large model to be tested, and the model structure, training algorithm or parameter settings can be adjusted in a targeted manner according to the test results to optimize the model's memory and generation capabilities in multiple rounds of conversations and improve model performance.

[0143] In the area of ​​improving intelligent customer service systems, for intelligent customer service applications based on large models, this method is used to evaluate the performance of the customer service model when handling multiple rounds of customer inquiries, and to discover problems that may arise in the model during long conversations, such as forgotten information, inaccurate answers, etc., and then improve the customer service system so that it can better understand user intentions, maintain the logical coherence of the conversation, and perform better in scenarios such as continuous information query and completion of multi-step tasks.

[0144] In summary, the above-mentioned large-model dialogue memory ability evaluation method includes but is not limited to the following beneficial effects:

[0145] First, accurately evaluating the memory capabilities of large models for multiple rounds of conversations will help identify the model’s weak links in memory and generation, thereby guiding R&D personnel to make targeted improvements and improve the overall quality of the model, making it more reliable and stable in practical applications.

[0146] Second, in applications such as intelligent customer service and voice assistants, the model optimized based on this method can provide users with more consistent, accurate and personalized services, reduce incorrect or irrelevant answers caused by model memory problems, enhance users' trust in and willingness to use the product, and improve user retention and activity.

[0147] Third, in multi-round conversations, the recall task of the large model to be tested is triggered by inserting marked turns. Based on the target conversation turn, a strategy is set to enable the large model to recall the conversation information of the target conversation turn, so as to evaluate the memory ability of the large model to be tested in multi-round conversations.

[0148] Fourth, by inserting markers into the dialogue sequence text, it is possible to flexibly set multiple target dialogue turns for testing memory lengths based on a single dialogue sequence text. Different target dialogue turns mean different target dialogue time points, enabling the testing of the large model to be tested under the same test data with different memory lengths, thereby improving the effectiveness of the test. At the same time, it is possible to test the depth of the memory ability of the large model to be tested as the number of dialogue turns increases, especially to deeply evaluate the memory ability in the later stages of the dialogue turns.

[0149] Figure 10 FIG. 4 is a block diagram of an apparatus for evaluating the dialogue memory ability of a large model shown in an exemplary embodiment of the present application. The apparatus includes a dialogue sequence text acquisition module 810, a first response content acquisition module 820, a second response content acquisition module 830, and an evaluation module 840.

[0150] The dialogue sequence text acquisition module 810 is configured to acquire a dialogue sequence text, in which multiple sets of questions and answers are arranged in sequence.

[0151] The first response content acquisition module 820 is configured to add markers to at least one question included in the dialogue sequence text, and input the dialogue sequence text with markers added into the large model to be tested, so that the large model to be tested performs multi-turn dialogue processing. Among them, each turn of dialogue processing realizes outputting the first response content corresponding to the question for a set of questions and answers.

[0152] The second response content acquisition module 830 is configured to, when the large model to be tested outputs the corresponding first response content for the dialogue sequence text with markers added, determine the target dialogue turn from the dialogue turns that the large model to be tested has performed dialogue processing on, and input the question corresponding to the target dialogue turn into the large model to be tested to obtain the second response content output by the large model to be tested.

[0153] The evaluation module 840 is configured to calculate the similarity score between the first response content and the second response content corresponding to the target dialogue turn, so as to evaluate the memory ability of the large model to be tested according to the similarity score.

[0154] In another exemplary embodiment, the dialogue sequence text acquisition module 810 further includes a dialogue dataset acquisition unit and a dialogue sequence text generation unit. The dialogue dataset acquisition unit is configured to acquire a dialogue dataset corresponding to a selected topic from a dialogue topic library. The dialogue sequence text generation unit is configured to generate a dialogue sequence text according to a selected grammar template corresponding to the selected topic and the dialogue dataset corresponding to the selected topic.

[0155] In another exemplary embodiment, the evaluation device for the large model dialogue memory ability further includes a question-answer group selection unit and a marking addition unit. The question-answer group selection unit is configured to select a first number of question-answer groups from the dialogue sequence text, and select a second number of question-answer groups from the first number of question-answer groups, where the first number is greater than the second number. The marking addition unit is configured to add marks to the questions included in the second number of question-answer groups.

[0156] In another exemplary embodiment, the evaluation module 840 further includes a current round determination unit and a target dialogue round determination unit. The current round determination unit is configured to determine the current round of the large model to be tested for the dialogue sequence text with marks added. The target dialogue round determination unit is configured to determine the target dialogue round according to the set interval rounds and the current round.

[0157] In another exemplary embodiment, the evaluation device for the large model dialogue memory ability further includes a mark replacement module and a first evaluation module. The mark replacement module is configured to repeatedly execute replacing the mark positions of the inserted marks in the dialogue sequence text, and input the dialogue sequence text with the replaced mark positions into the large model to be tested to obtain the similarity scores determined based on the replaced marks until the number of repetitions reaches the first preset number threshold. The first evaluation module is configured to evaluate the memory ability of the large model to be tested according to the similarity scores obtained each time.

[0158] In another exemplary embodiment, the evaluation device for the large model dialogue memory ability further includes a dialogue sequence text replacement module and a second evaluation module. The dialogue sequence text replacement module is configured to repeatedly execute replacing to obtain a new dialogue sequence text, and obtain the corresponding similarity scores based on the new dialogue sequence text until the number of repetitions reaches the second preset number threshold. The second evaluation module is configured to evaluate the memory ability of the large model to be tested according to the similarity scores obtained each time.

[0159] In another exemplary embodiment, the evaluation device for the large model dialogue memory ability further includes a similarity score determination unit and an evaluation unit. The similarity score determination unit is configured to determine the similarity scores obtained based on each mark. The evaluation unit is configured to evaluate the memory ability of the large model to be tested according to the similarity scores obtained based on each mark.

[0160] It should be noted that the evaluation device for the large model dialogue memory ability provided in the above embodiments and the evaluation method for the large model dialogue memory ability provided in the above embodiments belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiments, and will not be elaborated here. In practical applications, the evaluation device for the large model dialogue memory ability provided in the above embodiments can, as needed, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here either.

[0161] An embodiment of the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the evaluation method for the large model dialogue memory ability provided in each of the above embodiments.

[0162] Figure 11 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. It should be noted that Figure 11 The computer system 1200 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0163] As Figure 11 shown, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage section 1208 into the random access memory (RAM) 1203, such as executing the method described in the above embodiments. In the RAM 1203, various programs and data required for system operation are also stored. The CPU 1201, ROM 1202, and RAM 1203 are connected to each other through a bus 1204. The input / output (I / O) interface 1205 is also connected to the bus 1204.

[0164] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including such as a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc. and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as required. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 1210 as required so that a computer program read from it is installed into the storage section 1208 as required.

[0165] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including a computer program for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by a central processing unit (CPU) 1201, various functions defined in the system of the present application are executed.

[0166] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0167] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0168] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0169] Another aspect of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for evaluating the large model dialogue memory ability as described above. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device.

[0170] Another aspect of this application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for evaluating the large model dialogue memory ability provided in the above various embodiments.

[0171] The above content is only a preferred exemplary embodiment of this application and is not used to limit the implementation of this application. Those of ordinary skill in the art can easily make corresponding adaptations or modifications according to the main ideas and spirits of this application. Therefore, the protection scope of this application should be subject to the protection scope required by the claims.

Claims

1. A method for evaluating the dialogue memory ability of a large model, characterized in that: The method comprises: Obtaining a dialogue sequence text, wherein a plurality of groups of questions and answers are sequentially arranged in the dialogue sequence text; Adding a mark to at least one question included in the dialogue sequence text, and inputting the dialogue sequence text with the mark into the large model to be tested, so that the large model to be tested performs multiple rounds of dialogue processing, wherein each round of dialogue processing realizes outputting a first answer content corresponding to a set of questions and answers; When the large model to be tested outputs the corresponding first answer content for the marked dialogue sequence text, a target dialogue turn is determined from the dialogue turns that have been processed by the large model to be tested, and the question corresponding to the target dialogue turn is input into the large model to be tested, so as to obtain the second answer content output by the large model to be tested; Calculate the similarity score between the first answer content and the second answer content corresponding to the target dialogue turn, so as to evaluate the memory ability of the large model to be tested according to the similarity score.

2. The method according to claim 1, characterized in that A conversation topic library containing conversation data sets corresponding to different topics and a grammar template library corresponding to each topic are pre-built. The grammar template is used to specify the format of questions and answers in the conversation. The obtaining of the dialogue sequence text comprises: Acquire a conversation data set corresponding to the selected topic from the conversation topic library; The dialogue sequence text is generated according to the selected grammar template corresponding to the selected topic and the dialogue data set corresponding to the selected topic.

3. The method according to claim 1, characterized in that The adding a mark to at least one question included in the dialogue sequence text includes: Selecting a first number of question answer groups from the dialogue sequence text, and selecting a second number of question answer groups from the first number of question answer groups, the first number being greater than the second number; Marks are added to the questions included in the second number of question answer groups.

4. The method according to claim 1, characterized in that: The step of determining a target dialogue turn from the dialogue turns that have been dialogue-processed by the large model to be tested includes: Determine the current round performed by the large model to be tested on the dialogue sequence text with the mark added; A target dialogue turn is determined according to the set interval turn and the current turn.

5. The method according to claim 1, characterized in that The method further comprises: Repeating the process of changing the mark position of the insertion mark in the dialogue sequence text, and inputting the dialogue sequence text with the changed mark position into the large model to be tested, so as to obtain a similarity score determined based on the changed mark, until the number of repetitions reaches a first preset number threshold; The memory capacity of the large model to be tested is evaluated according to the similarity score obtained each time.

6. The method according to claim 1, characterized in that The method further comprises: Repeating the replacement to obtain a new dialogue sequence text, and obtaining a corresponding similarity score based on the new dialogue sequence text, until the number of repetitions reaches a second preset number threshold; The memory capacity of the large model to be tested is evaluated according to the similarity score obtained each time.

7. The method according to claim 1, characterized in that Adding marks to at least two question texts included in the dialogue sequence text; calculating a similarity score between the first answer content and the second answer content corresponding to the target dialogue turn, so as to evaluate the memory ability of the large model to be tested according to the similarity score, includes: Determining a similarity score based on each tag; The memory ability of the large model to be tested is evaluated based on the similarity score obtained based on each marker.

8. A device for evaluating the dialogue memory ability of a large model, characterized in that: include: A dialogue sequence text acquisition module is configured to acquire a dialogue sequence text, wherein a plurality of groups of questions and answers are sequentially arranged in the dialogue sequence text; A first answer content acquisition module is configured to add a mark to at least one question included in the dialogue sequence text, and input the dialogue sequence text with the mark added to the large model to be tested, so that the large model to be tested performs multiple rounds of dialogue processing, wherein each round of dialogue processing realizes outputting a first answer content corresponding to a set of questions and answers; The second answer content acquisition module is configured to, when the large model to be tested outputs the corresponding first answer content for the marked dialogue sequence text, determine the target dialogue turn from the dialogue turns that have been dialogue-processed by the large model to be tested, and input the question corresponding to the target dialogue turn into the large model to be tested, so as to obtain the second answer content output by the large model to be tested; An evaluation module is configured to calculate a similarity score between the first answer content and the second answer content corresponding to the target dialogue turn, so as to evaluate the memory ability of the large model to be tested according to the similarity score.

9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the method for evaluating the large model dialogue memory ability as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the method for evaluating the large model dialogue memory ability according to any one of claims 1 to 7.