Dialogue data generation method and device, electronic equipment and readable storage medium
By obtaining the text data and user types of the anchor during live broadcast, subject division and language style analysis are carried out, dialogue data that conforms to the anchor's language style is generated, which solves the problem of lack of personalized characteristics of dialogue data in the existing technology, and realizes personalized expression of dialogue data and efficient interaction of virtual clone model.
Patent Information
- Application Number
- CN202510306055.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The existing dialogue data generation methods lack personalized characteristics, and the generated dialogue data does not have the unique style and personalized expression of the anchor.
By obtaining the text data of the anchor during live broadcast and multiple user types, the topic division and language style analysis are carried out to generate dialogue data that matches the anchor's language style.
The generated dialogue data has the unique style and personalized expression of the anchor, which improves the personalized interaction ability of the virtual clone model.
Smart Images

Figure CN120218244A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device and readable storage medium for generating dialogue data. Background Art
[0002] In the fields of artificial intelligence and machine learning in recent years, the development and application of virtual avatar models have become a hot topic. Virtual avatar models show great potential in multiple industries such as entertainment, education, and business. Especially in the live streaming industry, virtual avatar models can achieve personalized interaction experiences. However, the key to training these large-scale models lies in effective dialogue data. The dialogue data not only needs to reach a certain scale in quantity but also must meet the learning needs of the virtual avatar model in terms of quality to ensure that the virtual avatar model can accurately capture and simulate the real language features and communication habits of the live streamer.
[0003] Currently, common methods for generating dialogue data include extracting dialogue corpus from movies and TV shows through Automatic Speech Recognition (ASR) technology, and extracting dialogue text from paper materials through Optical Character Recognition (OCR) technology. However, the above methods have significant limitations, that is, they lack personalized features. Since movies, TV shows, and paper materials are usually created for the general public, the generated dialogue data does not have the unique style and personalized expression of the live streamer. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, device, electronic device and readable storage medium for generating dialogue data to solve the problem that the dialogue data does not have the unique style and personalized expression of the live streamer.
[0005] In a first aspect, the present invention provides a method for generating dialogue data, the method comprising: Obtaining text data during the live broadcast of a live streamer and multiple user types; the text data is obtained according to the voice data during the live broadcast of the live streamer; Performing theme partitioning on the text data to obtain multiple dialogue themes and sub-text data corresponding to each dialogue theme; Determining the language style of the live streamer according to the sub-text data corresponding to different dialogue themes; Generating dialogue data that conforms to the language style according to multiple user types and all dialogue themes.
[0006] In an alternative embodiment, the performing theme partitioning on the text data to obtain multiple dialogue themes includes: Perform topic division on the text data to obtain multiple original conversation topics; Perform topic expansion on the multiple original conversation topics to obtain a preset number of new conversation topics for supplementation; the conversation topics include the original conversation topics and the new conversation topics.
[0007] In an alternative embodiment, the determining the host's language style according to the sub-text data corresponding to different conversation topics includes: Perform text proofreading processing on the sub-text data corresponding to different conversation topics to obtain the proofread sub-text data corresponding to different conversation topics; Determine the host's language style according to the proofread sub-text data corresponding to different conversation topics.
[0008] In an alternative embodiment, the determining the host's language style according to the proofread sub-text data corresponding to different conversation topics includes: Analyze the proofread sub-text data corresponding to different conversation topics to obtain the host's conversation style for different conversation topics; Determine the host's language style according to the host's conversation style for different conversation topics.
[0009] In an alternative embodiment, the method further includes: Generate multiple user attributes for each of the user types; The generating conversation data that conforms to the language style according to the multiple user types and all the conversation topics in the conversation topic set includes: Generate conversation data that conforms to the language style according to the multiple user attributes corresponding to each user type and all the conversation topics in the conversation topic set.
[0010] In an alternative embodiment, the method further includes: Review the conversation data and delete the data that does not conform to the language style in the conversation data.
[0011] In a second aspect, the present invention provides a conversation data generation device, and the device includes: A text data acquisition module, configured to acquire the text data during the host's live broadcast and multiple user types; the text data is obtained according to the voice data during the host's live broadcast; A conversation topic acquisition module, configured to perform topic division on the text data to obtain multiple conversation topics and the sub-text data corresponding to each conversation topic; A language style determination module, configured to determine the host's language style according to the sub-text data corresponding to different conversation topics; A dialogue data generation module, configured to generate dialogue data conforming to the language style according to multiple types of the users and all the dialogue topics.
[0012] In an alternative embodiment, the dialogue topic acquisition module is further configured to perform topic division on the text data to obtain a plurality of original dialogue topics; perform topic expansion on the plurality of original dialogue topics to obtain a preset number of supplementary new dialogue topics; the dialogue topics include the original dialogue topics and the new dialogue topics.
[0013] In a third aspect, the present invention provides an electronic device, including a processor and a memory, where the memory stores a computer program capable of being executed by the processor, and the computer program executable by the processor is used to implement the dialogue data generation method according to any one of the foregoing embodiments.
[0014] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the dialogue data generation method according to any one of the foregoing embodiments.
[0015] The dialogue data generation method, device, electronic device, and readable storage medium provided by the embodiments of the present invention, the method includes: obtaining text data during the live broadcast of the anchor and multiple types of users, the text data is obtained according to the voice data during the live broadcast of the anchor, performing topic division on the text data to obtain a plurality of dialogue topics and sub-text data corresponding to each dialogue topic, determining the language style of the anchor according to the sub-text data corresponding to different dialogue topics, and generating dialogue data conforming to the language style according to multiple types of users and all dialogue topics. By performing style analysis on the text data converted from the voice data during the live broadcast of the anchor and simulating the language style of the anchor to generate dialogue data, the generated dialogue data has the unique style and personalized expression of the anchor. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0017] Figure 1 Shows a schematic flowchart of a dialogue data generation method provided by an embodiment of the present invention; Figure 2 Shows another schematic flowchart of a dialogue data generation method provided by an embodiment of the present invention; Figure 3It shows another schematic flowchart of the dialogue data generation method provided by an embodiment of the present invention; Figure 4 It shows still another schematic flowchart of the dialogue data generation method provided by an embodiment of the present invention; Figure 5 It shows a block diagram of a dialogue data generation device provided by an embodiment of the present invention; Figure 6 It shows a schematic block diagram of an electronic device provided by an embodiment of the present invention.
[0018] Icons: 100 - electronic device; 110 - memory; 120 - processor; 130 - communication module; 200 - dialogue data generation device; 210 - text data acquisition module; 220 - dialogue topic acquisition module; 230 - language style determination module; 240 - dialogue data generation module. Detailed implementation manners
[0019] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0020] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but is merely representative of selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0021] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non - exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or also elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0022] Please refer to Figure 1 , Figure 1A flow chart of a method for generating conversation data provided by an embodiment of the present invention is shown. The method for generating conversation data can be applied to an electronic device. The specific flow of this embodiment is described below by taking conversation data generation as an example. Figure 1 The process shown is described in detail, and the method for generating conversation data may specifically include the following steps: Step 110: Acquire text data of the anchor during live broadcast and multiple user types.
[0023] Among them, the text data is obtained based on the voice data of the anchor during the live broadcast.
[0024] In some implementations, the electronic device may obtain the voice data of the anchor during the live broadcast, recognize the voice data through voice recognition technology, and obtain text data. The voice recognition technology may include but is not limited to ASR technology.
[0025] It should be noted that the prompt refers to the text or question input into the generative language model, which guides the generative language model to generate relevant output or response.
[0026] In some implementations, developers can design user type generation prompt words to drive the first generative language model to generate multiple user types. The electronic device can input the user type generation prompt words into the first generative language model to obtain multiple user types output by the first generative language model.
[0027] Among them, the first generative language model may include a generative pre-trained language model (Generative Pre-trained Transformer, GPT), a recurrent neural network (Recurrent Neural Network, RNN) and a PaLM (Pathways Language Model) language model, etc., which are not limited here.
[0028] User types are classified according to different user characteristics, interest preferences and other factors, so that different answers can be generated for different user types in the anchor's conversation style. User types can be classified according to professional roles, etc. Specifically, according to professional roles, they can be divided into students, freelancers, entrepreneurs, professional makeup enthusiasts, retirees and professionals.
[0029] For example, the user type generation prompt designed by the developer is "design n possible user types". The electronic device can input the user type generation prompt into the first generative language model, and the language model can generate n user types such as students, retirees, and professionals.
[0030] As an implementation, when it is necessary to generate dialogue data, the user can input the user type generation prompt word into the first generative language model through the electronic device, and obtain multiple user types output by the first generative language model.
[0031] As another implementation, after the electronic device obtains the text data during the host's live broadcast, it can input the user type generation prompt word into the first generative language model, and obtain multiple user types output by the first generative language model.
[0032] Step 120: Perform topic division on the text data to obtain multiple dialogue topics and the sub-text data corresponding to each dialogue topic.
[0033] In some implementations, developers can design a topic classification prompt to drive the language model to perform topic division on the input text data, and output multiple dialogue topics and the sub-text data corresponding to each dialogue topic.
[0034] For example, the topic classification prompt can be "Please perform topic classification on the input text data", "Please divide the following content into different dialogue topics, and summarize each dialogue topic with a phrase", etc., which are not limited here.
[0035] As an implementation, when the electronic device obtains the text data during the host's live broadcast, the user can input the text data and the topic classification prompt into the first generative language model through the electronic device, and obtain multiple dialogue topics output by the first generative language model and the sub-text data corresponding to each dialogue topic.
[0036] As another implementation, when the electronic device obtains the text data during the host's live broadcast, the electronic device can input the text data and the topic classification prompt into the first generative language model, and obtain multiple dialogue topics output by the first generative language model and the sub-text data corresponding to each dialogue topic.
[0037] For example, the electronic device inputs the following content: "Please divide the following content (text data) into different topics, and summarize each topic with a phrase (topic classification prompt): 'Welcome everyone to today's beauty live stream. First, let's talk about how to choose the right foundation for yourself. When choosing a foundation, pay attention to the matching of skin tone and skin type. Next, I'll recommend a sunscreen that's especially suitable for summer. This product is not only light but also non-greasy, and is especially suitable for oily skin. Finally, let's share some tips on evening skincare and how to use essence correctly.'" to the first generative language model. The conversation topics output by the first generative language model include makeup techniques, product recommendations, and skincare advice. The sub-text data corresponding to makeup techniques is 'First, let's talk about how to choose the right foundation for yourself. When choosing a foundation, pay attention to the matching of skin tone and skin type.'; the sub-text data corresponding to product recommendations is 'I'll recommend a sunscreen that's especially suitable for summer. This product is not only light but also non-greasy, and is especially suitable for oily skin.'; the sub-text data corresponding to skincare advice is 'Let's share some tips on evening skincare and how to use essence correctly.'
[0038] In some embodiments, the electronic device can divide the text data into topics through a topic classification model to obtain multiple conversation topics and the sub-text data corresponding to each conversation topic. Among them, the topic classification model can include a topic modeling model (Latent Dirichlet Allocation, LDA), a probabilistic topic classification model (Probabilistic Latent Semantic Analysis, PLSA), and the lda2vec model, etc., which are not limited here.
[0039] Step 130: Determine the host's language style according to the sub-text data corresponding to different conversation topics.
[0040] Among them, the host's language style refers to the language expression characteristics demonstrated by the host during the live stream.
[0041] In some embodiments, developers can design language style analysis prompts to drive the language model to output the host's language style according to the sub-text data corresponding to different input conversation topics.
[0042] For example, the language style analysis prompt can be 'Please analyze the language style of the following host when discussing different topics with the audience', etc., which are not limited here.
[0043] As an implementation, the user can input the sub-text data corresponding to different conversation topics and the language style analysis prompt to the first generative language model through the electronic device to obtain the host's language style output by the first generative language model.
[0044] As another implementation, after the electronic device obtains the sub-text data corresponding to different conversation topics, it obtains the preset language style analysis prompt words, and inputs the sub-text data corresponding to different conversation topics and the language style analysis prompt words into the first generative language model, and the first generative language model outputs the language style of the host.
[0045] For example, the electronic device inputs "Analyze the following content (sub-text data corresponding to different conversation topics) to determine the language style of the host when discussing different topics with the audience: The sub-text data corresponding to the game topic is: "Hey, everyone! I've been playing Game A recently. It's really great! The open-world design of the game has made me addicted, especially those puzzles. They're just so addictive. Has anyone else been playing?"; The sub-text data corresponding to daily life is: "Today has been going well. I went for a run this morning and the weather was great. Then I had a cup of coffee and felt really energetic. How about you? Has anything interesting happened?" into the first generative language model, and the first generative language model outputs that the language style of the host is short sentences, friendly tone and strong interactivity. Step 140: Generate conversation data that conforms to the language style according to multiple user types and all conversation topics.
[0046] In some implementations, developers can design a conversation data generation prompt word template to drive the language model to generate conversation data that conforms to the language style for multiple user types and multiple conversation topics included in the prompt words.
[0047] For example, the conversation data generation prompt word can be "The conversation topics are 'travel' and'makeup', the user types are 'young backpackers' and'makeup beginners', and the host's language style is relaxed and humorous. Simulate the conversation between the host and young backpackers and makeup beginners for the above conversation topics." etc., which is not limited here.
[0048] The electronic device can write the language style, multiple user types and all conversation topics into the conversation data generation prompt word template, generate a conversation data generation prompt word, and input the conversation data generation prompt word into the first generative language model to obtain the conversation data that conforms to the language style output by the first generative language model.
[0049] For example, the prompt word template generated by the conversation data is: the conversation topic is Y, the user type is X, the host's language style is Z, and the conversation between the host and the user type on the above conversation topic is simulated. The conversation topics obtained in step 120 include travel topics, beauty topics and food topics. The user types obtained in step 110 are college students, office workers and young backpackers. The host's language style is light-hearted and humorous. The electronic device writes the above content into the conversation data generation prompt word template, and the generated conversation data generates prompt words as "The conversation topics are travel topics, beauty topics and food topics. The user types are college students, office workers and backpackers. The host's language style is light-hearted and humorous. The conversation between the host and the user type on the above conversation topics is simulated."
[0050] Taking the conversation topic as travel, the user type as backpacker, and the anchor style as relaxed and humorous as an example, the first generative language model will generate the following content: Backpacker: Hey, I just arrived in Thailand, any recommendations for places? Host: Hello! Welcome to the Land of Smiles! You have to go to the night market in Bangkok. There are not only delicious food, but also all kinds of strange gadgets. It is a shopping paradise! Backpacker: That sounds great! I'd also like to go somewhere less crowded and not get overwhelmed by the crowds.
[0051] Host: Then you must go to Chiang Mai in the north! Hiking, cycling, and making friends with elephants is definitely an adventure. And the cafes in Chiang Mai are also great places to take photos and check in.
[0052] In some embodiments, the first generative language model pre-sets corresponding regulations on vocabulary selection, sentence structure, politeness, etc. for different user types and conversation topics, and configures a framework of some basic questions and their answers around each conversation topic. The first generative language model can generate conversation data that conforms to the language style for multiple user types and different conversation topics based on the above regulations and framework.
[0053] The conversation data generated by the conversation data generation method provided by the embodiment of the present invention is used to train the virtual avatar model. It can be understood that the conversation data is generated based on the user type, conversation topic and the host's language style. Training the virtual avatar model with the conversation data helps the virtual avatar model to respond in a personalized manner according to the preferences and needs of different users, so that users can feel more real and natural interaction, thereby improving the user experience.
[0054] The dialogue data generation method provided by the embodiments of the present invention obtains the text data during the host's live broadcast and multiple user types. The text data is obtained based on the voice data during the host's live broadcast. The text data is subject to theme division to obtain multiple dialogue themes and the sub-text data corresponding to each dialogue theme. According to the sub-text data corresponding to different dialogue themes, the host's language style is determined. According to multiple user types and all dialogue themes, dialogue data conforming to the language style is generated. By performing style analysis on the text data converted from the voice data during the host's live broadcast and simulating the host's language style to generate dialogue data, the generated dialogue data has the unique style and personalized expression of the host.
[0055] In order to make the dialogue data more comprehensive and cover a variety of dialogue themes, the original dialogue themes included in the text data during the host's live broadcast are expanded. Step 120 specifically further includes the following steps: performing theme division on the text data to obtain multiple original dialogue themes; expanding the multiple original dialogue themes to obtain a preset number of new dialogue themes.
[0056] Among them, the dialogue themes include the original dialogue themes and the new dialogue themes.
[0057] In some embodiments, developers can design dialogue theme expansion prompt words to drive the language model to expand the existing original dialogue themes and generate a preset number of new dialogue themes. The dialogue theme expansion prompt words include the preset number of supplements.
[0058] After obtaining the text data, the electronic device can input the text data and the theme classification prompt words into the first generative language model to obtain multiple original dialogue themes output by the first generative language model and the sub-text data corresponding to each original dialogue theme. The electronic device then inputs the dialogue theme expansion prompt words and the multiple original dialogue themes into the first generative language model to obtain a preset number of new dialogue themes output by the first generative language model.
[0059] For example, the electronic device inputs the following content "You are an assistant who helps expand chat themes. The existing original dialogue themes include: healthy diet, work-life balance, and the impact of technology on society. Please analyze their content based on the above original dialogue themes and expand 4 related dialogue themes." into the first generative language model, and the first generative language model outputs 4 dialogue themes such as psychological factors in healthy diet, the impact of remote work on work-life balance, the application of artificial intelligence in education, and the relationship between nutrition and exercise.
[0060] The electronic device can generate dialogue data conforming to the language style according to all dialogue themes (original dialogue themes and new dialogue themes) and multiple user types, so that the generated dialogue data is richer.
[0061] To ensure the accuracy and personalization of the dialogue data, as Figure 2 shown, step 130 specifically further includes the following steps: Step 131: Perform text proofreading on the sub-text data corresponding to different dialogue topics to obtain the proofread sub-text data corresponding to different dialogue topics.
[0062] In some embodiments, developers can design proofreading prompt words to drive the language model to perform text proofreading on the sub-text data corresponding to different dialogue topics, and obtain the proofread sub-text data corresponding to different dialogue topics.
[0063] The electronic device can input the proofreading prompt words and the sub-text data corresponding to different dialogue topics into the first generative language model. The first generative language model performs text proofreading on the sub-text data corresponding to different dialogue topics to obtain the proofread sub-text data corresponding to different dialogue topics output by the first generative language model.
[0064] Specifically, the first generative language model can check for typos in the sub-text data corresponding to different dialogue topics and correct the typos in the sub-text data corresponding to different dialogue topics; the first generative language model can check the grammar of the sub-text data corresponding to different dialogue topics and adjust the grammar in the sub-text data corresponding to different dialogue topics, etc., which are not limited here.
[0065] It can be understood that in the above step "expand multiple original dialogue topics to obtain a preset number of new dialogue topics", the new dialogue topics are not the dialogue topics divided from the text data. Therefore, there is no corresponding sub-text data for the new dialogue topics. The first generative language model performs text proofreading on the sub-text data corresponding to different original dialogue topics and will not perform text proofreading on the sub-text data corresponding to the new dialogue topics. The new dialogue topics are used as a reference when generating dialogue data later.
[0066] Step 132: Determine the host's language style according to the proofread sub-text data corresponding to different dialogue topics.
[0067] In some embodiments, the electronic device can analyze the proofread sub-text data corresponding to different dialogue topics to obtain the dialogue styles of the host for different dialogue topics, and determine the host's language style according to the dialogue styles of the host for different dialogue topics.
[0068] For example, during a technology conversation topic, the conversation style is more formal and logical; during a health conversation topic, the host's conversation style is compassionate and encouraging; during an entertainment conversation topic, the host's conversation style is lively, humorous, and relaxed, and likes to use pop culture terms. Through the above analysis, it can be concluded that the overall language style of this host is to use friendly language for most conversation topics and more formal language for a few conversation topics.
[0069] It can be understood that by analyzing the styles of different conversation topics, one can understand the host's language characteristics in different situations more meticulously. This enables the generated conversation data to better reflect the host's personalized style, making the conversation data more natural and practical.
[0070] To make the conversation data more targeted so that the trained virtual avatar model can output different responses for different user attributes of different types, this conversation data generation method further includes: generating multiple user attributes for each user type.
[0071] Among them, user attributes include user age, user gender, etc., which are not limited here.
[0072] In some embodiments, developers can design user attribute generation prompt words to drive the first generative language model to generate multiple user attributes for each user type. The electronic device can input the user attribute generation prompt words into the first generative language model and obtain multiple user attributes corresponding to each user type output by the first generative language model.
[0073] For example, the electronic device inputs "Design 5 different user types and divide 3 different age groups for each type" into the first generative language model, and the output of the first generative language model is as follows: The user type is a fitness enthusiast: Age group 1 is 18 - 25 years old; Age group 2 is 26 - 35 years old; Age group 3 is 36 - 50 years old.
[0074] Step 140 specifically further includes the following steps: Generate conversation data that conforms to the language style based on the multiple user attributes corresponding to each user type and all the conversation topics in the conversation topic set.
[0075] The specific process of generating conversation data that conforms to the language style is the same as the process in step 140 above.
[0076] It can be understood that by generating multiple user attributes for each user type, the electronic device can generate more targeted conversation data based on the multiple user attributes corresponding to each user type. This enables the trained virtual avatar model to handle more complex and diverse user needs and improve the personalization level of user interaction.
[0077] To improve the accuracy of the dialogue data, after the dialogue data is generated, the dialogue data is audited. As Figure 3 shown, after step 140, the following steps are specifically included: Step 150: Audit the dialogue data and delete the data in the dialogue data that does not conform to the language style.
[0078] It can be understood that the electronic device audits the dialogue data and deletes the data that does not conform to the language style to ensure the quality of the finally generated dialogue data, avoid the occurrence of responses that do not conform to the host's language style or are unnatural, thereby improving the user experience.
[0079] In some embodiments, the electronic device inputs the dialogue data into a second generative language model. The second generative language model audits the dialogue data and deletes the data in the dialogue data that does not conform to the language style, and obtains the dialogue data output after being audited by the second generative language model.
[0080] Among them, the first generative language model and the second generative language model are different generative language models. The second generative language model may include a Generative Pre-trained Transformer (GPT), a Recurrent Neural Network (RNN), and a PaLM (Pathways Language Model), etc., which are not limited herein.
[0081] It can be understood that each generative language model has a bias towards the dialogue data generated by itself. Using different generative language models to audit the dialogue data can reduce the bias towards the dialogue data, thereby improving the fairness and accuracy of the audit. And different generative language models have different performances when generating dialogue data, and can provide different audit rules, so as to more objectively evaluate the quality of the dialogue data, and further improve the accuracy of the dialogue data.
[0082] The dialogue data generation method provided by the embodiments of the present invention divides the text data of the host during the live broadcast into dialogue topics and proofreads the text data according to the prompt words through a generative language model, reducing the workload of manual preprocessing; analyzes and summarizes the host's language style according to the prompt words through the generative language model, making the subsequent generated dialogue data conform to the host's language style and making the dialogue data more natural; expands the dialogue topics and generates user types according to the prompt words through the generative language model to improve the richness and diversity of the subsequent generated dialogue data; audits the dialogue data through the generative language model to ensure the quality and style matching of the generated dialogue data, reducing the need for manual review and modification. Therefore, the above dialogue data generation method can improve the efficiency of generating dialogue data and ensure the quality of the dialogue data.
[0083] Taking the data processing using the prompt words (prompt) designed by developers and the generative language model as an example, as Figure 4 shown, the dialogue data generation method provided by the embodiments of the present invention solves the problem that the generated dialogue data does not have the unique style and personalized expression of the host through the following process, and the specific process is as follows: Step 210: Obtain the voice data of the host during the live broadcast, and identify the voice data through voice recognition technology to obtain text data.
[0084] Step 220: Input the user attribute generation prompt words into the first generative language model to obtain multiple user types output by the first generative language model and multiple user attributes corresponding to each user type.
[0085] Step 230: Input the text data and the theme classification prompt words into the first generative language model to obtain multiple original dialogue topics output by the first generative language model and the sub-text data corresponding to each original dialogue topic.
[0086] Step 240A: Input the dialogue topic expansion prompt words and multiple original dialogue topics into the first generative language model to obtain a preset number of new dialogue topics output by the first generative language model.
[0087] Step 240B: Input the proofreading prompt words and the sub-text data corresponding to different original dialogue topics into the first generative language model, and the first generative language model performs text proofreading processing on the sub-text data corresponding to different original dialogue topics to obtain the proofread sub-text data corresponding to different original dialogue topics output by the first generative language model.
[0088] Step 250B: Input the proofread sub-text data corresponding to different original conversation topics and the language style analysis prompt words into the first generative language model to obtain the language style of the host output by the first generative language model.
[0089] Step 260: Generate conversation data that conforms to the language style according to multiple user attributes corresponding to each user type and all conversation topics in the conversation topic set.
[0090] Step 270: Input the conversation data into the second generative language model. The second generative language model audits the conversation data and deletes the data that does not conform to the language style in the conversation data to obtain the conversation data output after being audited by the second generative language model.
[0091] The conversation data generation method provided by the embodiments of the present invention performs style analysis on the text data converted from the voice data during the host's live broadcast, and simulates the host's language style to generate conversation data, so that the generated conversation data has the unique style and personalized expression of the host.
[0092] To execute the corresponding steps in the above embodiments and all possible ways, an implementation manner of a conversation data generation device is given below. Further, please refer to Figure 5 , the figure is a functional module diagram of a conversation data generation device provided by the embodiments of the present invention. It should be noted that the basic principle and the technical effects generated by the conversation data generation device provided in this embodiment are the same as those in the above embodiments. For the sake of brief description, for the parts not mentioned in this embodiment, reference can be made to the corresponding content in the above embodiments. The conversation data generation device 200 includes: a text data acquisition module 210, a conversation topic acquisition module 220, a language style determination module 230, and a conversation data generation module 240, where: The text data acquisition module 210 acquires the text data during the host's live broadcast and multiple user types; the text data is obtained according to the voice data during the host's live broadcast.
[0093] The conversation topic acquisition module 220 is used to divide the topic of the text data to obtain multiple conversation topics and the sub-text data corresponding to each conversation topic.
[0094] The language style determination module 230 is used to determine the language style of the host according to the sub-text data corresponding to different conversation topics.
[0095] The conversation data generation module 240 is used to generate conversation data that conforms to the language style according to multiple user types and all conversation topics.
[0096] Optionally, the dialogue topic acquisition module 220 is further specifically configured to perform topic division on the text data to obtain multiple original dialogue topics; perform topic expansion on the multiple original dialogue topics to obtain a preset number of new dialogue topics; the dialogue topics include the original dialogue topics and the new dialogue topics.
[0097] Optionally, the language style determination module 230 is further specifically configured to perform text proofreading processing on the sub-text data corresponding to different dialogue topics to obtain the proofread sub-text data corresponding to different dialogue topics; determine the host's language style according to the proofread sub-text data corresponding to different dialogue topics.
[0098] Optionally, the language style determination module 230 is further specifically configured to analyze the proofread sub-text data corresponding to different dialogue topics to obtain the host's dialogue style for different dialogue topics; determine the host's language style according to the host's dialogue style for different dialogue topics.
[0099] Optionally, the text data acquisition module 210 is further specifically configured to generate multiple user attributes for each user type.
[0100] The dialogue data generation module 240 is further specifically configured to generate dialogue data that conforms to the language style according to the multiple user attributes corresponding to each user type and all the dialogue topics in the dialogue topic set.
[0101] Optionally, the dialogue data generation device 200 further includes: a dialogue data review module, where: The dialogue data review module is used to review the dialogue data and delete the data in the dialogue data that does not conform to the language style.
[0102] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0103] In several embodiments provided by the present invention, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0104] In addition, in each embodiment of the present invention, the various functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0105] Please refer to Figure 6, which is a block diagram of the electronic device 100 provided by an embodiment of the present invention. The electronic device 100 includes a memory 110, a processor 120, and a communication module 130. Each element of the memory 110, the processor 120, and the communication module 130 is electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines.
[0106] Among them, the memory 110 is used to store programs or data. The memory 110 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0107] The processor 120 is used to read / write the data or programs stored in the memory and execute corresponding functions. For example, when the computer program stored in the memory 110 is executed by the processor 120, the dialogue data generation methods disclosed in the above embodiments can be implemented.
[0108] The communication module 130 is used to establish a communication connection between the electronic device 100 and the server through the network and is used to transmit and receive data through the network.
[0109] It should be understood that Figure 6 The structure shown is only a schematic diagram of the structure of the electronic device, and the electronic device may also include more or fewer components than those shown Figure 6 in the figure. Figure 6 Each component shown in the figure can be implemented by hardware, software, or a combination thereof.
[0110] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the dialogue data generation method described in the above method embodiment can be implemented.
[0111] A computer-readable storage medium may be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has a storage space for program code for performing any of the method steps in the above-described methods. Such program code may be read out from, or written into, one or more computer program products. The program code may be compressed, for example, in a suitable form.
[0112] In several embodiments provided by the present invention, it should be understood that the disclosed apparatus and methods may also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of apparatuses, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.
[0113] If a function is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods according to the various embodiments of the present invention. The aforementioned storage medium includes: various media such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, which can store program code.
[0114] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for generating conversation data, characterized in that: The method comprises: Acquire text data of a host during live broadcast and multiple user types; the text data is obtained according to the voice data of the host during live broadcast; Dividing the text data into topics to obtain a plurality of conversation topics and sub-text data corresponding to each of the conversation topics; Determining the host's language style according to the subtext data corresponding to different conversation topics; According to the multiple user types and all the conversation topics, conversation data conforming to the language style is generated.
2. The method according to claim 1, characterized in that The subject division of the text data to obtain a plurality of conversation subjects includes: Dividing the text data into topics to obtain a plurality of original conversation topics; The multiple original conversation topics are expanded to obtain a preset number of new conversation topics; the conversation topics include the original conversation topics and the new conversation topics.
3. The method according to claim 1, characterized in that The determining the host's language style according to the subtext data corresponding to different dialogue topics includes: Performing text proofreading on the sub-text data corresponding to different conversation topics to obtain proofread sub-text data corresponding to the different conversation topics; The host's language style is determined according to the proofread sub-text data corresponding to different dialogue topics.
4. The method according to claim 3, characterized in that The determining the host's language style according to the proofread sub-text data corresponding to different dialogue topics includes: Analyze the proofread sub-text data corresponding to different dialogue topics to obtain the host's dialogue style for different dialogue topics; The host's language style is determined according to the host's conversation style for different conversation topics.
5. The method according to claim 1, characterized in that: The method further comprises: For each of the user types, generating a plurality of user attributes; Generating the conversation data conforming to the language style according to the multiple user types and all the conversation topics in the conversation topic set includes: Conversation data conforming to the language style is generated according to the multiple user attributes corresponding to each user type and all the conversation topics in the conversation topic set.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: The conversation data is reviewed, and data in the conversation data that does not conform to the language style is deleted.
7. A conversation data generating device, characterized in that: The device comprises: A text data acquisition module, used to acquire text data of the anchor during live broadcast and multiple user types; the text data is acquired based on the voice data of the anchor during live broadcast; A dialogue topic acquisition module, used to divide the text data into topics, obtain multiple dialogue topics and sub-text data corresponding to each dialogue topic; A language style determination module, used to determine the host's language style according to the subtext data corresponding to different conversation topics; The dialogue data generation module is used to generate dialogue data that conforms to the language style according to the multiple user types and all the dialogue topics.
8. The device according to claim 7, characterized in that The dialogue topic acquisition module is also used to divide the text data into topics to obtain multiple original dialogue topics; to expand the topics of the multiple original dialogue topics to obtain a preset number of new dialogue topics; the dialogue topics include the original dialogue topics and the new dialogue topics.
9. An electronic device, characterized in that: The device comprises a processor and a memory, wherein the memory stores a computer program executable by the processor, and the computer program executable by the processor is used to implement the method for generating conversation data according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating dialog data according to any one of claims 1 to 6 is implemented.