Dialogue voice generation method and device, equipment and medium

Through large language model and speech synthesis technology, text summary is converted into a dialogue text structure with multi-character interactive features, and each character is matched with acoustic feature parameters to generate dialogue voice with deep content and expressiveness, solving the problem of poor generation results in the existing technology and improving the quality of podcast programs.

CN120375802APending Publication Date: 2025-07-25PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510696766.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing dialogue voice generation technology has problems such as insufficient content depth, poor voice expression, lack of character interaction and insufficient personalization when generating podcast programs, which cannot meet the audience's needs for in-depth discussions and professional insights.

Method used

A large language model is used to convert text summary into a dialogue text structure with multi-role interaction features, and each role is assigned a unique label feature, matching consistent acoustic feature parameters from the preset voice library, and generating dialogue voice through the speech synthesis model.

Benefits of technology

It realizes the generation of dialogue voices that combine content depth and expressiveness, meets the audience's needs for in-depth discussions and professional insights, and enhances the diversity and appeal of the program.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375802A_ABST
    Figure CN120375802A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, the scheme can be applied to the fields of finance and medical treatment, and the invention provides a dialogue voice generation method, device and equipment and a medium, and the method comprises the steps: converting an input text abstract into a dialogue type text structure with a multi-role interaction feature through a large language model; distributing a unique label feature for each proxy role in the dialogue text structure; automatically matching acoustic characteristic parameters conforming to the agent roles from a preset voice library according to the label characteristics; and converting the dialogue text structure into dialogue voice through a voice synthesis model according to the acoustic characteristic parameters of each agent role, and outputting the dialogue voice. According to the embodiment of the invention, the input text abstract can be converted into the dialogue type text structure with the multi-role interaction characteristic, the requirements of audiences for deep discussion and professional insight are met, and the dialogue type text structure can be converted into dialogue voice with content depth and expressive force according to the acoustic characteristic parameter of each agent role.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a method, apparatus, device, and medium for generating dialogue voice. Background Art

[0002] Currently, there are various deficiencies in the automatic audio generation technology when generating podcast-like programs:

[0003] 1. Lack of content depth: Many automatically generated audio contents lack in-depthness and often can only provide superficial information, unable to meet the needs of listeners for in-depth discussions and professional insights;

[0004] 2. Poor voice expressiveness: Although existing speech synthesis technologies can generate understandable speech, they are still insufficient in aspects such as emotional expression and intonation changes, resulting in a poor listener experience;

[0005] 3. Lack of role interaction: Traditional audio generation methods are difficult to simulate the natural interaction between a host and guests, lacking a sense of reality and unable to create the atmosphere of a real conversation;

[0006] 4. Lack of personalization: Existing technologies are difficult to generate personalized content according to different topics or listener needs, restricting the diversity and attractiveness of programs.

[0007] Therefore, the existing methods for generating dialogue voice have the problem of poor generation effects. Summary of the Invention

[0008] Embodiments of the present invention provide a method, apparatus, device, and medium for generating dialogue voice, aiming to solve the problem of poor generation effects existing in the existing methods for generating dialogue voice.

[0009] In a first aspect, embodiments of the present invention provide a method for generating dialogue voice, the method comprising:

[0010] Using a large language model to convert the input text summary into a dialogue text structure with multi-role interaction characteristics;

[0011] Assigning a unique label feature to each agent role in the dialogue text structure;

[0012] Automatically matching acoustic feature parameters that match each agent role from a preset voice library according to the label feature;

[0013] Converting the dialogue text structure into dialogue voice according to the acoustic feature parameters of each agent role through a speech synthesis model and outputting it.

[0014] In a second aspect, embodiments of the present invention further provide a device for generating dialogue voice, the device comprising:

[0015] A first conversion unit, configured to use a large language model to convert an input text summary into a conversational text structure with multi-role interaction characteristics;

[0016] An allocation unit, configured to assign a unique label feature to each agent role in the conversational text structure;

[0017] A matching unit, configured to automatically match acoustic feature parameters corresponding to each agent role from a preset voice library according to the label features;

[0018] A second conversion unit, configured to convert the conversational text structure into conversational speech according to the acoustic feature parameters of each agent role through a speech synthesis model and output the conversational speech.

[0019] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0020] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium. The storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the method described in the first aspect above can be implemented.

[0021] The present invention provides a method, device, equipment and medium for generating conversational speech. The method includes: using a large language model to convert an input text summary into a conversational text structure with multi-role interaction characteristics; assigning a unique label feature to each agent role in the conversational text structure; automatically matching acoustic feature parameters corresponding to each agent role from a preset voice library according to the label features; converting the conversational text structure into conversational speech according to the acoustic feature parameters of each agent role through a speech synthesis model and outputting the conversational speech. Embodiments of the present invention can convert an input text summary into a conversational text structure with multi-role interaction characteristics, meet the needs of listeners for in-depth discussions and professional insights, and can also convert the conversational text structure into conversational speech with both content depth and expressiveness according to the acoustic feature parameters of each agent role. Description of the Drawings

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1Schematic flowchart of the dialogue voice generation method provided by an embodiment of the present invention;

[0024] Figure 2 Schematic block diagram of the dialogue voice generation device provided by an embodiment of the present invention;

[0025] Figure 3 Schematic block diagram of the electronic device provided by an embodiment of the present invention;

[0026] Figure 4 Schematic diagram of the application environment of the dialogue voice generation method provided by an embodiment of the present invention. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0029] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0030] It should be further understood that the term " / and / " used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Embodiments of the present invention provide a dialogue voice generation method, device, device and medium. For the dialogue voice generation method, please refer to Figure 4 , Figure 4 is a schematic diagram of the application environment of the dialogue voice generation method provided by an embodiment of the present invention. The dialogue voice generation method is applied to, for example, Figure 4In the application environment, the client communicates with the server through the network. The server converts the text summary from the client into a conversational text structure with multi-role interaction characteristics; assigns a unique label feature to each agent role in the conversational text structure; automatically matches acoustic feature parameters that match each agent role from a preset speech library according to the label feature; and converts the conversational text structure into conversational speech according to the acoustic feature parameters of each agent role through a speech synthesis model and outputs it. Embodiments of the present invention can convert the input text summary into a conversational text structure with multi-role interaction characteristics, meet the needs of the audience for in-depth discussion and professional insights, and can also convert the conversational text structure into conversational speech with both content depth and expressiveness according to the acoustic feature parameters of each agent role. Among them, the client can be, but is not limited to, various intelligent devices such as personal computers, laptops, smartphones, and tablets. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail through specific embodiments below.

[0031] Figure 1 It is a schematic flowchart of the conversational speech generation method provided by the embodiment of the present invention. As Figure 1 shown, the method includes the following steps S110-S140.

[0032] S110. Use a large language model to convert the input text summary into a conversational text structure with multi-role interaction characteristics.

[0033] In this embodiment, the conversational text structure includes questions from the host agent, answers from at least one guest agent, and transition segments written by the screenwriter agent, etc.

[0034] Specifically, the user inputs the text summary content on the operation interface of the client and customizes the configuration information of the agent roles. The agent roles include a screenwriter agent, a host agent (interviewer), and at least one guest agent (interviewee); among them, the screenwriter agent is responsible for writing the program script to ensure the coherence and logic of the content; the host agent is responsible for guiding the program process, asking questions, and controlling the discussion rhythm; the guest agent is responsible for providing professional insights and in-depth discussions according to the questions raised by the host.

[0035] Embodiments of the present invention allow users to customize the configuration information of agent roles. For example, by setting exclusive fields and opinion tendencies for different guest agents, the generated conversational text structure can present multi-perspective professional discussions. This configurable mechanism not only enhances the depth and diversity of the conversation content, but also can flexibly adjust the discussion direction according to user needs to achieve personalized content generation.

[0036] Example 1: Taking the medical scenario as an example, when the input text summary is "Diet management suggestions for diabetic patients", the server can automatically generate professional conversation content (i.e., conversational text structure) among endocrinologists, dietitians, and patient representatives according to the configuration information of the proxy roles. Specifically, by configuring corresponding professional knowledge bases for each proxy role (for example, doctors focus on medical principles and dietitians focus on dietary collocation), the generated conversation content can cover both professional content such as blood glucose control mechanisms and provide personalized diet plans, and finally output conversation content with a multi-disciplinary perspective.

[0037] Example 2: Taking the financial scenario as an example, when the input text summary is "Analysis of personal pension investment strategies", the server can automatically generate a multi-party professional conversation including financial advisors, fund managers, and retirement planners according to the configuration information of the proxy roles. Specifically, by configuring domain expertise for each proxy role (financial advisors focus on risk preference assessment, fund managers focus on asset allocation, and retirement planners focus on long-term cash flow planning), the generated conversation can cover both investment portfolio suggestions for different risk levels and provide personalized solutions for different age groups, and finally output conversation content with multi-dimensional professional insights.

[0038] In one embodiment, step S110 includes: analyzing the core content, semantic logic, and key information of the text summary through the large language model to obtain corresponding analysis results; using the large language model, generating conversation content among the proxy roles according to the analysis results and the role configuration information of each proxy role, where the proxy roles include a screenwriter proxy, a host proxy, and at least one guest proxy; organizing and arranging the conversation content among the proxy roles to form a conversational text structure with multi-role interaction characteristics.

[0039] In this embodiment, the large language model is used to analyze the core content, semantic logic, and key information of the text summary to obtain corresponding analysis results; and the large language model is used to generate conversation content among the proxy roles according to the analysis results and the role configuration information of each proxy role.

[0040] For ease of explanation, taking the medical scenario as an example: analyzing the input text summary (such as "Diet management suggestions for diabetic patients"), accurately identifying the core content of "diabetes management", constructing a medical logic chain of "etiology - diagnosis - treatment - prognosis" to achieve semantic logic construction, and at the same time marking professional data such as HbA1c standard values and key information such as highlighting precautions for insulin use; then, using the large language model, according to the analysis results and the role configuration information of each proxy role, while ensuring the accuracy of medical expressions, achieving a natural transition from pathological mechanisms to life suggestions, and outputting medical conversation content with both professional depth and practical value.

[0041] In one embodiment, generating the conversation content among the agent roles according to the analysis result and the role configuration information of each agent role includes: generating an overall program script according to the scriptwriter configuration information of the scriptwriter agent and the analysis result; formulating a corresponding discussion outline according to the overall program script and the host agent's host configuration information; generating the conversation content between the host agent and the guest agents according to the overall program script, the discussion outline and the guest configuration information of each guest agent.

[0042] In this embodiment, generating an overall program script according to the scriptwriter configuration information of the scriptwriter agent and the analysis result, and the scriptwriter configuration information; formulating a corresponding discussion outline according to the overall program script and the host agent's host configuration information, where the host configuration information includes a question strategy matrix and a rhythm control parameter; generating the conversation content between the host agent and the guest agents according to the overall program script, the discussion outline and the guest configuration information of each guest agent, where the guest configuration information includes a domain knowledge graph and a view tendency weight.

[0043] For the convenience of description, taking "Personal Pension Investment Strategy" in the financial scenario as an example: configure a financial product knowledge graph (including the risk-return characteristics of investment tools such as deposits, funds, and insurance) for guest agent 1 (in the role of a financial advisor), and set the view tendency weights of conservative (60%), balanced (30%), and aggressive (10%); configure a capital market knowledge graph (including stock-bond allocation models, risk hedging strategies, etc.) for guest agent 2 (in the role of a fund manager), and set a neutral and optimistic (70%) view tendency.

[0044] In one embodiment, after generating the conversation content between the host agent and the guest agents according to the overall program script, the discussion outline and the guest configuration information of each guest agent, it further includes: performing a coherence check on the generated conversation content. If a logical break or a role expression conflict is detected, according to the preset content correction rules, by adjusting the discussion outline or inserting supplementary content, regenerating the conversation content that meets the coherence requirements.

[0045] In this embodiment, when it is detected that the semantic association degree of the conversation context is lower than the preset threshold (i.e., a logical break), extracting relevant arguments from the overall program script to generate supplementary / transition content, and regenerating the conversation content that meets the coherence requirements based on the supplementary / transition content. When the fluctuation of the view tendency of the front and back speeches of the same guest agent exceeds the threshold (i.e., a role expression conflict), adjusting the guiding direction of the host agent's questions and correcting the speech content of the guest agent to achieve logical repair.

[0046] S120. Assign a unique label feature to each agent role in the dialogical text structure.

[0047] In this embodiment, a unique label feature is assigned to each agent role according to the conversational text structure, and the label feature includes at least three attributes: gender (male / female / neutral), age (youth / middle age / old age), and language style (rigorous / lively / authoritative).

[0048] In one embodiment, step S120 includes: assigning a unique label feature to each agent role according to the conversational text structure and the role configuration information of each agent role; wherein, the label feature includes at least three categories of attributes: gender, age, and language style.

[0049] In this embodiment, a unique label feature is assigned to each agent role according to the conversational text structure and the role configuration information of each agent role; wherein, the label feature includes at least three categories of attributes: gender, age, and language style. Taking the discussion on pension investment in the financial scenario as an example: Guest Agent 1 (fund manager role) can be marked as [male, middle age, authoritative]; 2) Guest Agent 2 (financial advisor) can be marked as [female, youth, lively].

[0050] S130. Automatically match acoustic feature parameters that match each agent role from a preset voice library according to the label feature.

[0051] In this embodiment, acoustic feature parameters that match each agent role are automatically matched from a preset voice library according to the label feature, and the acoustic feature parameters include at least: basic tone color parameters, prosody feature parameters, and pronunciation style parameters.

[0052] In one embodiment, step S130 includes: establishing a multi-dimensional mapping relationship between the role label feature and the acoustic feature parameter system; automatically matching acoustic feature parameters that match each agent role from a preset voice library according to the multi-dimensional mapping relationship, and the acoustic feature parameters include at least: basic tone color parameters, prosody feature parameters, and pronunciation style parameters.

[0053] In this embodiment, a multi-dimensional mapping relationship between the role label feature and the acoustic feature parameter system is established; acoustic feature parameters that match each agent role are automatically matched from a preset voice library according to the multi-dimensional mapping relationship, and the acoustic feature parameters include at least: basic tone color parameters, prosody feature parameters, and pronunciation style parameters. Specifically, the basic tone color parameters are determined based on the gender attribute, the prosody feature parameters are optimized according to the age feature, and the pronunciation style parameters are adjusted in combination with the language style.

[0054] In one embodiment, after automatically matching the acoustic feature parameters that match each agent role from a preset speech database according to the multi-dimensional mapping relationship, the method further includes: calculating a matching deviation degree between the role label feature and the corresponding acoustic feature parameter; if the matching deviation degree is greater than a preset threshold, iteratively optimizing the acoustic feature parameter until the matching deviation degree between the optimized acoustic feature parameter and the corresponding role label feature is not greater than the preset threshold.

[0055] In this embodiment, the matching deviation degree is a multi-dimensional matching deviation degree; specifically, calculate the matching deviation degree between the acoustic feature parameter and the agent role label feature, and trigger the calibration process only when the matching deviation degree exceeds the preset threshold; when the timbre dimension deviation degree > threshold T1, automatically adjust the formant parameter; when the prosody dimension deviation degree > threshold T2, dynamically adjust the fundamental frequency curve; when the style dimension deviation degree > threshold T3, iteratively optimize the pronunciation feature.

[0056] S140. Convert the conversational text structure into conversational speech according to the acoustic feature parameters of each agent role through a speech synthesis model and output the speech.

[0057] In this embodiment, convert the conversational text structure into conversational speech with both content depth and expressiveness according to the acoustic feature parameters of each agent role through a speech synthesis model and output the speech.

[0058] In summary, the embodiment of the present invention can convert the input text summary into a conversational text structure with multi-role interaction characteristics, meet the needs of the audience for in-depth discussions and professional insights, and can also convert the conversational text structure into conversational speech with both content depth and expressiveness according to the acoustic feature parameters of each agent role.

[0059] Figure 2 It is a schematic block diagram of a conversational speech generation device provided by an embodiment of the present invention. As Figure 2 shown, corresponding to the above conversational speech generation method, the present invention further provides a conversational speech generation device, and the device is configured in an application environment such as Figure 4 wherein the user terminal communicates with the server through a network. Specifically, please refer to Figure 2 , the conversational speech generation device 700 includes:

[0060] A first conversion unit 701, configured to convert an input text summary into a conversational text structure with multi-role interaction characteristics by using a large language model;

[0061] An allocation unit 702, configured to allocate a unique label feature to each agent role in the conversational text structure;

[0062] A matching unit 703, configured to automatically match acoustic feature parameters corresponding to each agent role from a preset speech library according to the label features;

[0063] A second conversion unit 704, configured to convert the dialog text structure into dialog voice according to the acoustic feature parameters of each agent role through a speech synthesis model and output the dialog voice.

[0064] In some embodiments, when the first conversion unit 701 executes the step of converting the input text summary into a dialog text structure with multi-role interaction features by using a large language model, it is specifically configured to:

[0065] Analyze the core content, semantic logic, and key information of the text summary through the large language model to obtain corresponding analysis results; use the large language model to generate the dialog content among the agent roles according to the analysis results and the role configuration information of each agent role; wherein, the agent roles include a screenwriter agent, a host agent, and at least one guest agent; organize and arrange the dialog content among the agent roles to form a dialog text structure with multi-role interaction features.

[0066] In some embodiments, when the first conversion unit 701 executes the step of generating the dialog content among the agent roles according to the analysis results and the role configuration information of each agent role, it is specifically configured to:

[0067] Generate an overall program script according to the screenwriter configuration information of the screenwriter agent and the analysis results; formulate a corresponding discussion outline according to the overall program script and the host configuration information of the host agent; generate the dialog content between the host agent and the guest agents according to the overall program script, the discussion outline, and the guest configuration information of each guest agent.

[0068] In some embodiments, after the first conversion unit 701 executes the step of generating the dialog content between the host agent and the guest agents according to the overall program script, the discussion outline, and the guest configuration information of each guest agent, it is further configured to:

[0069] Perform coherence verification on the generated dialog content. If a logical break or role expression conflict is detected, according to the preset content correction rules, regenerate the dialog content that meets the coherence requirements by adjusting the discussion outline or inserting supplementary content.

[0070] In some embodiments, when the allocation unit 702 executes the step of allocating a unique label feature to each agent role in the dialog text structure, it is specifically configured to:

[0071] Assign a unique label feature to each agent role according to the dialog text structure and the role configuration information of each agent role; wherein, the label feature includes at least three types of attributes: gender, age, and language style.

[0072] In some embodiments, when the matching unit 703 executes the step of automatically matching acoustic feature parameters that match each agent role from a preset speech library according to the label feature, it specifically is used for:

[0073] Establish a multi-dimensional mapping relationship between the role label feature and the acoustic feature parameter system; automatically match acoustic feature parameters that match each agent role from the preset speech library according to the multi-dimensional mapping relationship, and the acoustic feature parameters at least include: basic tone color parameters, prosody feature parameters, and pronunciation style parameters.

[0074] In some embodiments, after the matching unit 703 executes the step of automatically matching acoustic feature parameters that match each agent role from the preset speech library according to the multi-dimensional mapping relationship, it is further used for:

[0075] Calculate the matching deviation degree between the role label feature and the corresponding acoustic feature parameter; if the matching deviation degree is greater than a preset threshold, perform iterative optimization on the acoustic feature parameter until the matching deviation degree between the optimized acoustic feature parameter and the corresponding role label feature is not greater than the preset threshold.

[0076] It should be noted that those skilled in the art can clearly understand that the specific implementation processes of the above dialog voice generation device and each unit can refer to the corresponding descriptions in the foregoing method embodiments. For the sake of convenience and conciseness of description, they will not be elaborated here.

[0077] The above dialog voice generation device can be implemented in the form of a computer program, and this computer program can run on an electronic device as shown in Figure 3 shown.

[0078] Please refer to Figure 3 , Figure 3 is a schematic block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 800 can be a terminal or a server. Among them, the terminal can be an electronic device with communication functions. The server can be an independent server or a server cluster composed of multiple servers.

[0079] Refer to Figure 3 , the electronic device 800 includes a processor 802, a memory, and a network interface 805 connected through a system bus 801. Among them, the memory can include a non-volatile storage medium 803 and an internal memory 804.

[0080] The non-volatile storage medium 803 can store an operating system 8031 and a computer program 8032. The computer program 8032 includes program instructions that, when executed, enable the processor 802 to execute a method for generating dialogue voice.

[0081] The processor 802 is used to provide computing and control capabilities to support the operation of the entire electronic device 800.

[0082] The internal memory 804 provides an environment for the operation of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, it enables the processor 802 to execute a method for generating dialogue voice.

[0083] The network interface 805 is used for network communication with other devices. Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the electronic device 800 to which the solution of the present invention is applied. The specific electronic device 800 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0084] Among them, the processor 802 is used to run the computer program 8032 stored in the memory to implement the following steps:

[0085] Using a large language model to convert the input text summary into a dialogue text structure with multi-role interaction characteristics; assigning a unique label feature to each agent role in the dialogue text structure; automatically matching acoustic feature parameters that match each agent role from a preset voice library according to the label feature; converting the dialogue text structure into dialogue voice according to the acoustic feature parameters of each agent role through a voice synthesis model and outputting it.

[0086] In some embodiments, when the processor 802 implements the step of using a large language model to convert the input text summary into a dialogue text structure with multi-role interaction characteristics, the specific implementation steps are as follows:

[0087] Analyzing the core content, semantic logic, and key information of the text summary through the large language model to obtain a corresponding analysis result; using the large language model to generate dialogue content between each agent role based on the analysis result and the role configuration information of each agent role; among them, the agent roles include a screenwriter agent, a host agent, and at least one guest agent; organizing and arranging the dialogue content between each agent role to form a dialogue text structure with multi-role interaction characteristics.

[0088] In some embodiments, when the processor 802 implements the step of generating the conversation content among the agent roles according to the analysis result and the role configuration information of each agent role, the specific implementation is as follows:

[0089] Generate an overall program script according to the screenwriter configuration information of the screenwriter agent and the analysis result; formulate a corresponding discussion outline according to the overall program script and the host configuration information of the host agent; generate the conversation content between the host agent and the guest agents according to the overall program script, the discussion outline and the guest configuration information of each guest agent.

[0090] In some embodiments, after the processor 802 implements the step of generating the conversation content between the host agent and the guest agents according to the overall program script, the discussion outline and the guest configuration information of each guest agent, the following steps are further implemented:

[0091] Perform coherence verification on the generated conversation content. If a logical break or a conflict in role expression is detected, according to the preset content correction rules, regenerate the conversation content that meets the coherence requirements by adjusting the discussion outline or inserting supplementary content.

[0092] In some embodiments, when the processor 802 implements the step of assigning a unique label feature to each agent role in the dialog text structure, the specific implementation is as follows:

[0093] Assign a unique label feature to each agent role according to the dialog text structure and the role configuration information of each agent role; wherein, the label feature includes at least three types of attributes: gender, age and language style.

[0094] In some embodiments, when the processor 802 implements the step of automatically matching the acoustic feature parameters that match each agent role from the preset voice library according to the label feature, the specific implementation is as follows:

[0095] Establish a multi-dimensional mapping relationship between the role label feature and the acoustic feature parameter system; automatically match the acoustic feature parameters that match each agent role from the preset voice library according to the multi-dimensional mapping relationship, and the acoustic feature parameters include at least: basic timbre parameters, prosody feature parameters and pronunciation style parameters.

[0096] In some embodiments, after the processor 802 implements the step of automatically matching the acoustic feature parameters that match each agent role from the preset voice library according to the multi-dimensional mapping relationship, the following steps are further implemented:

[0097] Calculate the matching deviation degree between the character label feature and the corresponding acoustic feature parameter; if the matching deviation degree is greater than a preset threshold, iteratively optimize the acoustic feature parameter until the matching deviation degree between the optimized acoustic feature parameter and the corresponding character label feature is not greater than the preset threshold.

[0098] It should be understood that in the embodiment of the present invention, the processor 802 may be a central processing unit (CPU), and the processor 802 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0099] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.

[0100] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, where the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the following steps:

[0101] Use a large language model to convert the input text summary into a conversational text structure with multi-role interaction features; assign a unique label feature to each agent role in the conversational text structure; automatically match acoustic feature parameters that match each agent role from a preset voice library according to the label features; convert the conversational text structure into conversational speech according to the acoustic feature parameters of each agent role through a speech synthesis model and output it.

[0102] In an embodiment, when the processor executes the program instructions to implement the step of using a large language model to convert the input text summary into a conversational text structure with multi-role interaction features, the following steps are specifically implemented:

[0103] Analyze the core content, semantic logic, and key information of the text summary through the large language model to obtain corresponding analysis results; use the large language model to generate the conversation content among the agent roles according to the analysis results and the role configuration information of each agent role; where the agent roles include a screenwriter agent, a host agent, and at least one guest agent; organize and arrange the conversation content among the agent roles to form a conversational text structure with multi-role interaction characteristics.

[0104] In one embodiment, when the processor executes the program instructions to implement the step of generating the conversation content among the agent roles according to the analysis results and the role configuration information of each agent role, the specific implementation steps are as follows:

[0105] Generate an overall program script according to the screenwriter configuration information of the screenwriter agent and the analysis results; formulate a corresponding discussion outline according to the overall program script and the host configuration information of the host agent; generate the conversation content between the host agent and the guest agents according to the overall program script, the discussion outline, and the guest configuration information of each guest agent.

[0106] In one embodiment, after the processor executes the program instructions to implement the step of generating the conversation content between the host agent and the guest agents according to the overall program script, the discussion outline, and the guest configuration information of each guest agent, the following steps are further implemented:

[0107] Perform coherence verification on the generated conversation content. If a logical break or role expression conflict is detected, according to the preset content correction rules, regenerate the conversation content that meets the coherence requirements by adjusting the discussion outline or inserting supplementary content.

[0108] In one embodiment, when the processor executes the program instructions to implement the step of assigning a unique label feature to each agent role in the conversational text structure, the specific implementation steps are as follows:

[0109] Assign a unique label feature to each agent role according to the conversational text structure and the role configuration information of each agent role; where the label feature includes at least three types of attributes: gender, age, and language style.

[0110] In one embodiment, when the processor executes the program instructions to implement the step of automatically matching the acoustic feature parameters that match each agent role from the preset voice library according to the label feature, the specific implementation steps are as follows:

[0111] Establish a multi-dimensional mapping relationship between the role label features and the acoustic feature parameter system; automatically match the acoustic feature parameters that match each agent role from a preset speech library according to the multi-dimensional mapping relationship, and the acoustic feature parameters at least include: basic timbre parameters, prosody feature parameters, and pronunciation style parameters.

[0112] In one embodiment, after the processor executes the program instructions to implement the step of automatically matching the acoustic feature parameters that match each agent role from the preset speech library according to the multi-dimensional mapping relationship, the following steps are further implemented:

[0113] Calculate the matching deviation degree between the role label features and the corresponding acoustic feature parameters; if the matching deviation degree is greater than a preset threshold, iterate and optimize the acoustic feature parameters until the matching deviation degree between the optimized acoustic feature parameters and the corresponding role label features is not greater than the preset threshold.

[0114] The storage medium can be various computer-readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes.

[0115] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0116] In several embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0117] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0118] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0119] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for generating dialogue voice, characterized in that, The method includes: Using a large language model to convert the input text summary into a conversational text structure with multi-role interaction features; Assigning a unique label feature to each agent role in the conversational text structure; Automatically matching acoustic feature parameters that match each agent role from a preset voice library according to the label features; Converting the conversational text structure into conversational speech according to the acoustic feature parameters of each agent role through a speech synthesis model and outputting it.

2. The method for generating dialogue voice according to claim 1, wherein The using a large language model to convert the input text summary into a conversational text structure with multi-role interaction features includes: Analyzing the core content, semantic logic, and key information of the text summary through the large language model to obtain corresponding analysis results; Using the large language model, based on the analysis results and the role configuration information of each agent role, generating the conversation content between each agent role; where the agent roles include a screenwriter agent, a host agent, and at least one guest agent; Organizing and arranging the conversation content between each agent role to form a conversational text structure with multi-role interaction features.

3. The method for generating dialogue voice according to claim 2, wherein The generating the conversation content between each agent role based on the analysis results and the role configuration information of each agent role includes: Generating an overall program script based on the screenwriter configuration information of the screenwriter agent and the analysis results; Formulating a corresponding discussion outline based on the overall program script and the host configuration information of the host agent; Generating the conversation content between the host agent and the guest agents based on the overall program script, the discussion outline, and the guest configuration information of each guest agent.

4. The method for generating dialogue voice according to claim 3, characterized in that, After the generating the conversation content between the host agent and the guest agents based on the overall program script, the discussion outline, and the guest configuration information of each guest agent, it further includes: Performing a coherence check on the generated conversation content. If a logical break or role expression conflict is detected, according to the preset content correction rules, by adjusting the discussion outline or inserting supplementary content, regenerating the conversation content that meets the coherence requirements.

5. The method for generating dialogue voice according to claim 1, wherein The assigning a unique label feature to each agent role in the conversational text structure includes: Assigning a unique label feature to each agent role based on the conversational text structure and the role configuration information of each agent role; where the label features at least include three categories of attributes: gender, age, and language style.

6. The method for generating dialogue voice according to claim 1, characterized in that The automatically matching acoustic feature parameters that match each agent role from a preset voice library according to the label features includes: Establishing a multi-dimensional mapping relationship between the role label features and the acoustic feature parameter system; Automatically matching acoustic feature parameters that match each agent role from a preset voice library according to the multi-dimensional mapping relationship. The acoustic feature parameters at least include: basic tone color parameters, prosody feature parameters, and pronunciation style parameters.

7. The method for generating dialogue voice according to claim 6, wherein After the automatically matching acoustic feature parameters that match each agent role from a preset voice library according to the multi-dimensional mapping relationship, it further includes: Calculating the matching deviation degree between the role label features and the corresponding acoustic feature parameters; If the matching deviation degree is greater than a preset threshold, the acoustic feature parameters are iteratively optimized until the matching deviation degree between the optimized acoustic feature parameters and the corresponding role label features is not greater than the preset threshold.

8. A dialogue voice generation device, characterized in that, The device includes: A first conversion unit, configured to convert an input text summary into a conversational text structure with multi-role interaction features by using a large language model; An allocation unit, configured to allocate unique label features to each agent role in the conversational text structure; A matching unit, configured to automatically match acoustic feature parameters that match each agent role from a preset speech library according to the label features; A second conversion unit, configured to convert the conversational text structure into conversational speech according to the acoustic feature parameters of each agent role through a speech synthesis model and output the conversational speech.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the conversational speech generation method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by the processor, the processor is caused to execute the conversational speech generation method according to any one of claims 1-7.