Dialogue data generation method, electronic device, storage medium and program product
By generating problem statements containing value dimensions and generating responses based on real-time state values, the problem of lack of value state alignment in existing conversation data sets is solved, and high-accurate conversation data sets and agent responses are achieved.
Patent Information
- Application Number
- CN202510592554.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing dialogue dataset lacks alignment of value states, which causes the agent to be unable to respond according to the agent's real-time state when answering questions, and the existing retrieval schemes cannot perceive the priority of timestamps, resulting in recall deviations.
Construct a conversation dataset by generating problem statements containing value dimensions and generating aligned value and state values based on real-time state values. At the same time, by adding timestamps to the problem statement and retrieving the status using timestamps, the accuracy of dialogue data is improved.
The dialogue data set is generated and aligned with the value set and the value state values, which improves the accuracy of the dialogue data and the response quality of the agent, and avoids recall deviations.
Smart Images

Figure CN120144723A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for generating dialogue data, an electronic device, a storage medium, and a program product. Background Art
[0002] Existing open-source dialogue datasets mainly include open-domain dialogue datasets (chatting type) and vertical-domain dialogue datasets. The methods for generating datasets include crawling dialogue data from public websites, manually annotating dialogue content, and generating dialogue data through methods such as data conversion and expansion.
[0003] Existing dialogue datasets only perform simple question-and-answer according to settings. The answers corresponding to the same question are the same, and they do not answer according to the corresponding state of the scenario where the agent is located.
[0004] In the prior art, most similar document retrievals are directly based on BM2.5 or vector recall. This retrieval scheme cannot perceive the priority of timestamps, resulting in a recall bias of "semantically relevant but time-effectively incorrect". For example, there are two pieces of log information stored in the satiety value dimension: at 10:01, the hunger state is very hungry; at 10:30, the hunger state is not hungry. When the query statement is: Are you very hungry? Directly recalling through similarity will result in recalling old data instead of the latest data, thus affecting the accuracy of the agent. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the present invention provides a method for generating dialogue data, an electronic device, a storage medium, and a program product, which at least partially solve the problems of the lack of value and value status values in the existing dialogue set.
[0006] In a first aspect, an embodiment of the present disclosure provides a method for generating dialogue data, including: Generating a question statement including a value dimension; Retrieving a corresponding real-time status value based on the value dimension of the question statement; Generating a response that aligns the value and the status value based on the real-time status value; Constructing dialogue data based on the question and the response of the question statement.
[0007] Optionally, the generating of the question including a value dimension includes: Adding a timestamp to the question statement; Performing vector retrieval on the constructed statement and the value dimension database based on the question statement to obtain similar statements, and thus obtaining the value dimension corresponding to the question statement based on the similar statements; Extract a time window based on timestamps, retrieve in the constructed value and status database based on the value dimension and the time window, obtain the real-time status corresponding to the time window, and thus generate a response that aligns the value and status values corresponding to the time window.
[0008] Optionally, the conversation data includes a prompt part and a response part; the prompt part includes a question field and a personal information field. The question field is used to store the question statement, and the personal information field is used to store the value and status information related to the current question statement retrieved from the value and status database. The response part is a reply to the question field that is generated by calling a language model based on the status value in the personal information field and is aligned with the personal information field.
[0009] Optionally, generating the question statement including the value dimension includes: Calling a language model based on all combination data of enumerated value dimensions and status values to generate a data pair including a question statement and a corresponding response.
[0010] Optionally, generating the response that aligns the value and status values based on the real-time status value includes: Modifying and adapting the data pair of the question statement and the corresponding response based on the real-time status value to obtain the response that aligns the value and status values.
[0011] Optionally, calling a language model based on all combination data of enumerated value dimensions and status values to generate a data pair including a question statement and a corresponding response includes: Constructing a prompt for calling a language model in a way based on a small number of sample examples; Entering a value description in the prompt description field, and the language model generates a data pair including a question statement and a corresponding response based on the value description and all combination data of the enumerated value dimensions and status values in the prompt.
[0012] Optionally, the prompt includes question-value response data, question-status response data, and question-status-value response data.
[0013] In a second aspect, an embodiment of the present disclosure further provides an electronic device, which includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the conversation data generation method according to any one of the first aspects.
[0014] In a third aspect, embodiments of the present disclosure further provide a computer-readable storage medium storing computer instructions for causing a computer to execute the dialogue data generation method according to any one of the first aspect.
[0015] In a fourth aspect, embodiments of the present disclosure further provide a computer program product including a computer program / instructions which, when executed by a processor, implement the dialogue data generation method according to any one of the first aspect.
[0016] The dialogue data generation method, electronic device, storage medium and program product provided by the present invention. In the dialogue data generation method, value dimensions are set in a question, and the real-time state of an agent is adapted to the set value dimensions to obtain a value aligned with the real-time state, so as to achieve the purpose of generating a dialogue data set aligned with the set value and value state values.
[0017] By setting a timestamp and retrieving the state according to the timestamp, the states at different times can be distinguished, thereby achieving the purpose of improving the accuracy of dialogue data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features and advantages of the present disclosure will become more obvious. In the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0019] Figure 1 It is a flowchart of a dialogue data generation method provided by an embodiment of the present disclosure; Figure 2 It is a flowchart of another dialogue data generation method provided by an embodiment of the present disclosure; Figure 3 It is a user interaction flowchart provided by an embodiment of the present disclosure; Figure 4 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The embodiments of the present disclosure will be described in detail below in conjunction with the accompanying drawings.
[0021] It should be clear that the following illustrates the implementation modes of the present disclosure through specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation modes, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present disclosure.
[0022] It should be noted that the following describes various aspects of embodiments within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. Additionally, this apparatus and / or method can be implemented using other structures and / or functionality in addition to one or more of the aspects described herein.
[0023] It also should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure schematically. The diagrams only show the components related to the present disclosure and are not drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in its actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0024] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the aspects can be practiced without these specific details.
[0025] In some application scenarios of humanoid embodied intelligent robots, engineers will configure humanoid physiological values (such as thirst, fullness, tiredness, sleepiness, etc.), intrinsic values (such as tidiness, safety, curiosity, etc.) for the robots, as well as the real-time status values corresponding to these values. As the agent moves in the virtual scenario, such as tidying up toys, eating bread, etc., the status values corresponding to these value dimensions of the agent will change in real time.
[0026] When the conversation scenario is a dialogue between a human user and a robot regarding the robot's own value, the robot needs to be able to reply with content aligned with the value and status settings.
[0027] For example, the physiological value set by the engineer for the robot's "being sleepy" is: when feeling sleepy, unable to tolerate sleepiness and especially in need of sleep, and the real-time status of the robot in the dimension of "being sleepy" is: feeling a bit sleepy. At this time, when the user's question is "Do you feel sleepy?", the fine-tuned robot needs to reply that I feel a bit sleepy, rather than "As an artificial intelligence, I have no perception or emotion, so I don't feel sleepy."
[0028] In order to enable the replies generated by the robot to be aligned with the values and statuses given to the robot, a batch of value-aligned dialogue data needs to be generated, then this batch of data is mixed into the existing dialogue dataset, and then the model is post-trained (fine-tuning techniques such as SFT, RL, etc.) so that the model can generate value-aligned replies.
[0029] The solution for generating the dialogue dataset in this implementation includes two parts: 1) Retrieving and recalling relevant values and their status information for the user query; 2) Constructing a reply aligned with the recalled value and status descriptions.
[0030] As Figure 1 shown, this embodiment discloses a dialogue data generation method, including: Generating question statements containing value dimensions; Optionally, the generating of question statements containing value dimensions includes: Adding timestamps to the question statements; Performing vector retrieval in the constructed statement and value dimension database based on the question statements to obtain similar statements, and thus obtaining the value dimensions corresponding to the question statements based on the similar statements; Extracting a time window based on the timestamp, retrieving in the constructed value and status database based on the value dimension and the time window to obtain the real-time status corresponding to the time window, and thus generating a response aligned with the value and status values corresponding to the time window.
[0031] Retrieving the corresponding real-time status value based on the value dimension of the question statement; Generating a response aligned with the value and status values based on the real-time status value; Constructing dialogue data based on the question and response of the question statement.
[0032] Optionally, the conversation data includes a prompt part and a response part; the prompt part includes a question field and a personal information field. The question field is used to store the question statement, and the personal information field is used to store the value and status information related to the current question statement retrieved from the value and status database. The response part is a reply to the question field aligned with the personal information field, generated by calling a language model based on the status value in the personal information field.
[0033] Specifically, for example, there are two log messages stored in the satiety value dimension: at 10:01, the hunger status is very hungry; at 10:30, the satiety status is not hungry.
[0034] When the question is "Are you very hungry?", the set value dimensions are that when hungry, answer "hungry", when full, answer "not hungry", and when a little hungry but not wanting to eat yet, etc.
[0035] When the language model generates question statements according to the value dimension, they are "Are you hungry now?", "Are you thirsty now?", "Are you tired now?", etc.
[0036] After generating the question statement, first retrieve it in the database of statements and value dimensions. For example, retrieve the answer related to the value of "Are you hungry now?". When hungry, answer "hungry", when full, answer "not hungry", and when a little hungry but not wanting to eat yet.
[0037] Then retrieve the real-time status during the conversation according to the timestamp of the question statement. For example, if the real-time status retrieved at 10:30 is the satiety status, then the response obtained is "I'm not hungry yet". If the real-time status retrieved at 10:01 is retrieved, the response is "I'm very hungry and need to eat now". The real-time status of the conversation can be obtained according to the sensor's perception of the surrounding environment or updated according to the time setting. For example, when the intelligent agent moves from the office scene to the restaurant, its real-time status can change from not hungry to more hungry. It is obtained that 7:00 to 7:30 is breakfast time, 12:00 to 12:30 is lunch time, 6:00 to 6:30 is dinner time. When the time is within the above time periods, it is in the hungry state, and when the time is outside the above time periods, it is in the satiety state. Among them, it can also be set that 4:00 to 4:30 is afternoon tea time. The conversation dataset constructed according to this technical solution can give appropriate answers according to the real-time status during the conversation and the set values, rather than only answering "As a robot, I don't feel sleepy" when asking "Are you hungry?".
[0038] First, call the language model through preset value dimensions (such as thirst quenching, satiety, tiredness, sleepiness, tidiness, safety, curiosity, etc.) to generate question queries about these value dimensions. Then, retrieve the preset real-time status values related to the query from the memory through relevance. Finally, call the language model according to the preset real-time status values to generate the corresponding response.
[0039] The format of the dialogue dataset for training constructed based on the template specifically includes a prompt part and a response part. The prompt part consists of two fields: the
Question
Personal Information
Question
Personal Information
Personal Information
[0040] In this embodiment, when retrieving, the time-aware relevance retrieval method is used, aiming to surpass traditional retrieval methods such as BM25 or document recall relying on static vectors. This embodiment achieves more accurate retrieval results by comprehensively utilizing time information and value dimensions. First, in the query analysis stage, the concept of a time window is introduced. For queries that do not explicitly contain time information, the system defaults to using the current time as the time window. This step ensures the timeliness of the retrieval. Next, the language model is used to generate similar queries to construct a vector library containing similar queries. Each similar query is associated with a specific value dimension, enabling subsequent retrieval to better capture the potential value intention of the query. In the retrieval stage, the system first recalls the top 10 most similar queries through the vector retrieval mechanism. Subsequently, based on the value dimensions of these similar queries, the system determines the target value dimension of the user's query. This step ensures that the retrieval results are not only relevant to the user's query content but also consistent with their potential value requirements. Finally, the system obtains the relevant value and status information from the corresponding time range according to the identified value dimension and the specified time window. This method effectively improves the accuracy and timeliness of the recall results, providing a more valuable retrieval experience for users.
[0041] In the timestamp, the termination time of the question is the time of the timestamp, and the start time can be set according to different scenarios. For example, if the question is "Were you hungry 10 minutes ago?", the termination time of the question is 10 minutes before the current time, and the start time of the question can be 11 minutes before the current time. If the current time is 11 o'clock and the question is "Were you hungry 10 minutes ago?", the real-time status of the query can be the real-time status from 10:49 to 10:50, or it may be the real-time status from 10:45 to 10:50, etc. This time window can be set according to different application scenarios.
[0042] In this embodiment, for the above two-stage dialogue generation scheme, when retrieving the value and status values related to the query, on the one hand, in order to ensure recall, the retrieval system may recall information in multiple value dimensions. For example, when asking if one is hungry, the retrieval system may recall the value status information in the fullness and fatigue dimensions. On the other hand, our retrieval system will also recall other information within the scenario (non-value dimension information). Therefore, when generating the response, the language model may sometimes generate descriptions that cannot align with the relevant values and status values in the face of complex inputs. For example, the generated response may only express "I'm a little sleepy" without accurately reflecting the state of "especially in need of sleep".
[0043] In another implementation scenario, generating question statements containing value dimensions includes: Invoking the language model based on all combination data of enumerated value dimensions and status values to generate data pairs including question statements and corresponding responses.
[0044] Invoking the language model based on all combination data of enumerated value dimensions and status values to generate data pairs including question statements and corresponding responses, including: Constructing prompt words for invoking the language model in the way of a small number of sample examples; Inputting value descriptions in the prompt word description field, and the language model generates data pairs including question statements and corresponding responses based on the value descriptions and all combination data of the enumerated value dimensions and status values in the prompt words.
[0045] The prompt words include question value response data, question status response data, and question status value response data.
[0046] Then, when generating a response that aligns values and status values based on real-time status values, modify and adapt the data pairs of question statements and corresponding responses based on the real-time status values to obtain a response that aligns values and status values.
[0047] Such as Figure 2As shown, in this scenario, all possible combinations of value dimensions and status values are enumerated first, and then queries and responses that align with this description are generated simultaneously based only on the enumerated value descriptions, ensuring that each pair of dialogue data accurately reflects the preset value status. Then, relevant value status information is retrieved from the repository for the queries. However, the [Personal Information] in the prompt for constructing the dialogue at the end is not the originally retrieved value status information, but rather, after further modification, the value description information used in the first step to generate the query and response pairs is used to replace the value status information of the retrieved target dimension. In this way, the generated dialogue dataset not only contains accurately aligned value descriptions but also ensures that the responses highly fit the actual situation, effectively improving the understanding and response capabilities of the dialogue system.
[0048] By first enumerating the value dimensions and status value combinations and then generating query and response pairs based on them, it can ensure that each pair of dialogue data is accurately aligned with the preset value status. This method effectively avoids problems such as ambiguous value dimensions and unclear status information that may occur in traditional generation methods, making the dialogue data more accurate and clear in value transmission and semantic expression. The replaced
Personal Information
[0049] Key elements of few-shot learning: A small number of labeled samples: The core of few-shot learning lies in using a small number of labeled samples to train the model. These samples are usually referred to as the "support set". The size of the support set can be determined according to the specific task and the availability of data, generally ranging from a few to dozens of samples. For example, in a text classification task, 5 labeled samples may be provided for each category.
[0050] Prompt design: The prompt is an important tool for guiding the model to complete tasks. In the field of language processing, the prompt can be a natural language instruction, template, or example. Designing an effective prompt can significantly improve the performance of the model. For example, for a sentiment analysis task, the prompt can be "Please judge whether the following comment is positive or negative." Model selection and fine-tuning: Selecting an appropriate pre-trained model is the basis of few-shot learning. The pre-trained model has learned a large amount of language knowledge and patterns and can quickly adapt to new tasks. In few-shot learning, the pre-trained model is usually fine-tuned to better adapt to a specific task. When fine-tuning, some model parameters can be frozen and only some layers are trained to prevent overfitting.
[0051] Application scenarios of few-shot learning: Text classification: Few-shot learning can be used to classify text, such as news classification, sentiment analysis, etc. When the labeled data is limited, through few-shot learning, the model can use a small number of labeled samples to learn the classification boundary and classify new text.
[0052] Question Answering System: In a question answering system, few-shot learning can help the model quickly adapt to question answering tasks in new domains. By providing a small number of labeled question-answer pairs, the model can learn how to extract answers from the context and generate accurate responses.
[0053] Text Generation: Few-shot learning can be used for text generation tasks, such as story generation, summary generation, etc. By providing a small number of examples, the model can learn to generate text that conforms to a specific style and theme.
[0054] """Please generate questions and answers based on the input description. The answers need to be semantically aligned with the input description.
[0055] The prompt format is as follows: Example 1: Description: [Value] When feeling hungry, there is a strong need for satiety and cannot tolerate hunger.
[0056] Your generation: Question: Are you resistant to hunger? Answer: I'm not resistant to hunger.
[0057] Question: Do you have a need for satiety? Answer: I have a strong need for satiety.
[0058] Example 1 is the question-value response data.
[0059] Example 2: Description: [State] Feeling a bit hungry.
[0060] Your generation: Question: Are you hungry? Answer: I'm a bit hungry.
[0061] Question: Are you a bit hungry? Answer: Yes, I'm feeling a bit hungry.
[0062] Example 2 is the question-state response data.
[0063] Example 3: Description: [Value + State] When feeling hungry, there is a strong need for satiety and cannot tolerate hunger. Not hungry Your generation: Question: Do you want to eat? Answer: I'm not hungry now, so I don't want to eat.
[0064] Question: Are you hungry in your stomach? Answer: I'm not hungry in my stomach! Example 3 is the question-state-value response data.
[0065] Description: {desc} Your generation: """ Examples 1, 2, and 3 above are all inputs to the language model.
[0066] In the obtained dialogue set, users will ask various questions of the robot. One type of question is whether the robot can be aware of its own value and corresponding status information. Therefore, the application of the generated dialogue dataset with value alignment is to first be mixed into other previously constructed dialogue datasets, and then jointly perform SFT (supervised fine-tuning) on the base model. Through this mixed training, the model is enabled to have the ability to reply to questions about value status while not losing other dialogue capabilities.
[0067] The Base model is a basic pre-trained model that has not been fine-tuned for specific tasks. It masters general language knowledge through self-supervised learning and requires secondary fine-tuning to adapt to specific tasks.
[0068] Language model tasks: This is the most common pre-training objective, such as predicting the next word or the word probability distribution in text. Taking the GPT series as an example, its pre-training is carried out by predicting the next word in the text sequence. By inputting a large amount of text, the model learns text patterns and grammar rules and can generate smooth and natural text.
[0069] Masked language model tasks: Such as BERT, during pre-training, some words are randomly masked, and the model is required to predict the masked words. When inputting text, some words are masked, and the model needs to predict the masked words based on the context. Learning to understand the meaning of words in a sentence and the associations between words is important for understanding text and generating answers.
[0070] Pre-training data: The pre-training data for the Base model needs to be large and diverse. For example, BERT is pre-trained using an extremely large amount of text (books, web pages, etc.). The advantage of diverse data is that it enables the model to see texts of different topics, styles, and domains, and learn general language knowledge and patterns. The quality of pre-training data is crucial for the performance of the model. Only when the accuracy and relevance of data annotation are high can the pre-training effect be good.
[0071] Model architecture: Transformer architecture: It is a commonly used architecture for modern Base models, consisting of multiple layers of Transformer encoders or decoders. The encoder has an attention mechanism and a feed-forward neural network, which can capture the dependencies between words in the sequence; in addition to the attention mechanism and the feed-forward network, the decoder also has an input embedding layer and an output layer, which can generate sequence outputs. The advantage of the Transformer architecture is parallel computing and capturing long-distance dependencies, which is suitable for processing natural language tasks.
[0072] Parameter scale: The parameter scale of the Base model affects the model performance and the computational resource requirements. Small models (with millions of parameters) are suitable for resource-constrained scenarios, while large models (with hundreds of millions or even hundreds of billions of parameters) have strong representation capabilities and can capture complex text patterns and semantic information, but have high requirements for computational and storage resources.
[0073] Application Scenarios: Text Generation: The Base model can generate text, providing a preliminary text framework or creative inspiration, such as news reports, story beginnings, and product descriptions.
[0074] Text Understanding: It can handle text classification, sentiment analysis, and question - answering tasks, analyze text content, extract key information, judge sentiment tendencies, and provide answers.
[0075] As a Basis for Fine - Tuning: During fine - tuning for specific tasks, the parameters of the Base model are optimized and adjusted to improve task performance. Fine - tuning enables the model to adapt to specific tasks, enhancing accuracy and relevance.
[0076] The Base model is the foundation for the development and application of language models. The pre - training objectives, data, model architecture, and application scenarios are interrelated. The pre - training objectives and data determine the model's capabilities, the architecture provides a computational framework, and the application scenarios demonstrate its functions. The Base model provides general language knowledge and patterns for various natural language processing tasks and plays a greater role in combination with fine - tuning and other techniques.
[0077] Supervised Fine - Tuning (SFT) is a technique for further training based on a pre - trained model.
[0078] Supervised fine - tuning is to further train on a labeled dataset for a specific task on the basis that the pre - trained model has completed large - scale unsupervised pre - training. Through learning on large - scale general data, the pre - trained model already has good general language representation capabilities and general knowledge, equivalent to completing a kind of "pre - learning".
[0079] When introducing labeled data for a specific task, the knowledge and capabilities of the pre - trained model can be further focused on that specific task, making it more suitable for the task requirements. By adjusting the model's parameters, the model can make more accurate predictions and decisions under the data distribution and patterns of the specific task, thereby improving the model's performance on that task.
[0080] Prepare a Labeled Dataset: Collect and organize a large amount of labeled data related to the target task. This data should be representative and diverse, fully covering various situations and variations of the task, so that the model can learn comprehensive task knowledge and patterns during the fine - tuning process.
[0081] Select a Pre - trained Model: According to the characteristics and requirements of the task, select a suitable base model from among many publicly available pre - trained models. For example, for natural language processing tasks, models such as BERT, GPT, etc. can be selected.
[0082] Model Adjustment and Adaptation: Adapt the output of the pre-trained model to the output requirements of the target task. Some adjustments to the model structure may be needed, such as adding specific output layers, modifying the loss function, etc., to enable it to output results that meet the task requirements, such as classification labels, sequence annotation results, or generated text, etc.
[0083] Fine-tuning Training: Further train the pre-trained model using the prepared labeled dataset. Adjust the model's parameters through optimization algorithms to minimize the loss function on the task data, thereby learning the patterns and regularities of the specific task. During the training process, appropriate hyperparameters can be set, such as the learning rate, batch size, number of training epochs, etc., to control the training process and improve the convergence speed and performance of the model.
[0084] Evaluation and Testing: After completing the fine-tuning training, use independent validation sets and test sets to evaluate and test the fine-tuned model, and calculate relevant metrics such as accuracy, recall, F1-score, mean squared error, etc., to objectively measure the performance of the model on the target task. According to the evaluation results, the model structure, training process, or dataset, etc., can be further adjusted to optimize the model performance.
[0085] When the user asks a question, the system will retrieve relevant value and status values according to the query, then construct a dialogue template consistent with the offline training, input it into the fine-tuned language model, and then the language model generates a response and replies to the user. The overall interaction logic is as Figure 3 shown.
[0086] The language model in this embodiment is not limited and can be a self-developed language model or a large language model such as GPT (Generative Pre-trained Transformer), etc.
[0087] The electronic device disclosed in this embodiment includes a memory and a processor. The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0088] The processor can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the electronic device executes all or part of the steps of the dialogue data generation method of the various embodiments of the present disclosure described above.
[0089] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain a good user experience effect, well-known structures such as communication buses and interfaces may also be included in this embodiment, and these well-known structures should also be included in the protection scope of the present disclosure.
[0090] As Figure 4 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. It shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present disclosure. Figure 4 The illustrated electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0091] As Figure 4 As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the electronic device are also stored. The processing device, ROM, and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0092] Generally, the following devices can be connected to the I / O interface: an input device including, for example, a sensor or a visual information acquisition device; an output device including, for example, a display screen; a storage device including, for example, a magnetic tape, a hard disk, etc.; and a communication device. The communication device can allow the electronic device to communicate with other devices (such as edge computing devices) wirelessly or wiredly to exchange data. Although Figure 4 the illustrated electronic device shows various devices, it should be understood that it is not required to implement or have all the illustrated devices. More or fewer devices can be alternatively implemented or had.
[0093] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, all or part of the steps of the dialogue data generation method according to the embodiments of the present disclosure are performed.
[0094] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated here.
[0095] The computer-readable storage medium disclosed in this embodiment stores non-temporary computer-readable instructions. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the dialogue data generation method according to the foregoing embodiments of the present disclosure are performed.
[0096] The above-mentioned computer-readable storage medium includes, but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).
[0097] For a detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details are not repeated here.
[0098] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for illustrative and facilitating understanding purposes, and not for limitation. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.
[0099] In this disclosure, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the phrase "and / or", and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with it.
[0100] In addition, as used herein, "or" in the listing of items starting with "at least one" indicates a disjunctive listing, so that for example, the listing of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Further, the term "exemplary" does not mean that the examples described are preferred or better than other examples.
[0101] It should also be noted that in the systems and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure.
[0102] Various changes, substitutions, and alterations to the technologies described herein can be made without departing from the teachings defined by the appended claims. In addition, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and acts described above. Current or later-developed processes, machines, manufactures, compositions of events, means, methods, or acts that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or acts within their scope.
[0103] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
[0104] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit embodiments of the present disclosure to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some variations, modifications, alterations, additions, and subcombinations thereof.
Claims
1. A method for generating conversation data, characterized in that: include: Generate question statements that include value dimensions; Retrieve the corresponding real-time status value based on the value dimension of the question statement; Generate a response that aligns the value and the state value based on the real-time state value; Conversation data is constructed based on questions and responses of question sentences.
2. The method for generating conversation data according to claim 1, characterized in that: The generation of value-related issues includes: Add timestamps to question statements; Based on the question sentence, vector retrieval is performed in the constructed sentence and value dimension database to obtain similar sentences, thereby obtaining the value dimension corresponding to the question sentence based on the similar sentences; The time window is extracted based on the timestamp, and the constructed value and status database is searched based on the value dimension and the time window to obtain the real-time status corresponding to the time window, thereby generating a response with aligned value and status values corresponding to the time window.
3. The method for generating conversation data according to claim 1, characterized in that: The dialogue data includes a prompt part and a response part; the prompt part includes a question field and a personal information field, the question field is used to store the question statement, and the personal information field is used to store the value and status information related to the current question statement retrieved from the value and status database; The response part is a reply to the question field that is generated by calling a language model based on the status value in the personal information field and is aligned with the personal information field.
4. The method for generating conversation data according to claim 1, characterized in that: The generating of the question statement containing the value dimension includes: Based on all the combined data of the enumerated value dimensions and the status values, the language model is called to generate data pairs including question sentences and corresponding responses.
5. The method for generating conversation data according to claim 4, characterized in that: The step of generating a response of aligning the value and the status value based on the real-time status value comprises: The data pairs of question statements and corresponding responses are modified and adapted based on the real-time status value to obtain a response with aligned value and status value.
6. The method for generating conversation data according to claim 4, characterized in that: The method of calling the language model based on all the combined data of the enumerated value dimension and the state value to generate a data pair including a question statement and a corresponding response includes: Constructing prompt words for calling language models based on a small number of sample examples; Enter the value description in the prompt word description field, and the language model generates a data pair including a question statement and a corresponding response based on the value description and all combinations of value dimensions and status values enumerated in the prompt word.
7. The method for generating conversation data according to claim 6, characterized in that: The prompt words include question value response data, question status response data and question status value response data.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the conversation data generating method described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the conversation data generating method described in any one of claims 1-7.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the method for generating conversation data described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and device for generating reply information based on robot emotional state
CN109033179A
Multi-modal fusion natural interaction method and system of intelligent robot and medium
CN114995657A
Interaction method and device based on reinforcement learning
CN116521850A
Dynamic sampling dialogue generation model training method and device, equipment and medium
CN116719920A
Intelligent dialogue method and device, equipment and storage medium
CN117009469A