Target content generation method and device based on large model, electronic equipment and medium

By using a large model-based approach to automatically generate podcast content, the problems of modality monotony and high creation threshold in podcast content generation are solved, and the fluency and quality of the content are improved.

CN120952004APending Publication Date: 2025-11-14BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511100458.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, the input modality in podcast content generation is singular, making it impossible to process mixed content such as text, links, and files simultaneously. The generated content is relatively bland, with a high barrier to entry and a large workload.

Method used

By using a large model-based approach, target data is acquired, scripts are generated, and target roles and acoustic parameters of text segments are determined, thus automatically generating podcast content, including role selection and acoustic parameter determination.

Benefits of technology

It simplifies the podcast content creation process, resulting in smoother and more coherent content, and improves creation efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952004A_ABST
    Figure CN120952004A_ABST
Patent Text Reader

Abstract

The invention provides a target content generation method and device based on a large model, electronic equipment, a computer readable storage medium and a computer program product, and relates to the field of artificial intelligence, in particular to the technical field of large models, data processing and content generation. According to the implementation scheme, target data used for generating target content are obtained, wherein the target content comprises audio; generating a script corresponding to the content based on the target data, wherein the script comprises at least one text segment; for each text segment in the at least one text segment, determining a target role for broadcasting the text segment; for each text segment, based on the text segment and the target role corresponding to the text segment, determining acoustic parameters corresponding to the target role when the text segment is broadcasted; and generating target content based on the script, the target role corresponding to each text segment and the acoustic parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, and in particular to the fields of large models, data processing, and content generation technology. Specifically, it relates to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating target content based on a large model. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] Human-computer interaction (HCI) is a way for humans to interact with machines using natural language. With the continuous development of artificial intelligence technology, machines have become capable of understanding human-generated information, comprehending its inherent meaning, and providing corresponding feedback. In these operations, the accuracy of semantic understanding, the speed of feedback, and the provision of appropriate opinions or suggestions all influence the smoothness of HCI interaction. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating target content based on a large model.

[0005] According to one aspect of this disclosure, a method for generating target content based on a large model is provided, comprising: acquiring target data for generating the target content, wherein the target content includes audio; generating a script corresponding to the content based on the target data, wherein the script includes at least one text segment; for each of the at least one text segment, determining a target character for broadcasting the text segment; for each text segment, determining acoustic parameters corresponding to the target character when broadcasting the text segment based on the text segment and the target character corresponding to the text segment; and generating the target content based on the script, the target character corresponding to each text segment, and the acoustic parameters.

[0006] According to another aspect of this disclosure, a data processing apparatus based on a large model is provided, comprising: a content outline generation module configured to acquire target data for generating the target content, wherein the target content includes audio; a script generation module configured to generate a script corresponding to the content based on the target data, wherein the script includes at least one text segment; a role determination module configured to determine a target role for broadcasting each of the at least one text segment; a feature determination module configured to determine, for each text segment, acoustic parameters corresponding to the target role when broadcasting the text segment based on the text segment and the target role corresponding to the text segment; and a content generation module configured to generate the target content based on the script, the target role corresponding to each text segment, and the acoustic parameters.

[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor to enable the at least one processor to perform the methods described in this disclosure.

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods described in this disclosure.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described in this disclosure.

[0010] According to one or more embodiments of this disclosure, by automatically obtaining a script for generating target content based on the input target data, and further determining the role of each text segment in the script and the corresponding acoustic parameters when the role broadcasts the text segment, the creation process of generating target content is greatly simplified, and the generated target content is made smoother and more coordinated, thereby improving the creation efficiency and quality of target content.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0013] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown; Figure 2 A flowchart of a data processing method based on a large model according to an embodiment of the present disclosure is shown; Figure 3 A schematic diagram of a target content display page according to an embodiment of the present disclosure is shown; Figure 4 A schematic diagram of a dialog page for generating target content according to an embodiment of the present disclosure is shown; Figure 5 A structural block diagram of a large-model-based data processing apparatus according to embodiments of the present disclosure is shown; and Figure 6 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0014] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0015] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0016] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0017] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0018] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0019] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of target content generation methods based on large models.

[0020] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.

[0021] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0022] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to input target content generation requests, obtain generated target content, etc. The client devices can provide interfaces that allow users to interact with them. The client devices can also output information to the user through these interfaces. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0023] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0024] Network 110 can be any type of network well known to those skilled in the art, and can support data communication using any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0025] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0026] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0027] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0028] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0029] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store things such as podcast script files, conversation history, etc. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0030] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0031] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0032] For example, a podcast is a form of digital media that primarily publishes and distributes content in the form of audio or video. It allows creators to share information, stories, opinions, or educational content in the form of series, and allows listeners to listen or watch at their own pace. Podcast content is incredibly diverse, covering news, education, technology, culture, entertainment, and many other areas to meet the interests of different listeners.

[0033] In related technologies, current content such as podcasts has a single input modality and cannot simultaneously process mixed content such as text, links, and files; moreover, the generated content is relatively bland and monotonous, or the creation threshold is high, requiring users to participate in the operation of multiple features such as content generation and audio parameters, resulting in a high workload.

[0034] Therefore, embodiments of this disclosure provide a method for generating target content based on a large model. Figure 2 A flowchart of a target content generation method based on a large model according to an embodiment of the present disclosure is shown, such as... Figure 2 As shown, method 200 includes: acquiring target data for generating the target content, wherein the target content includes audio (step 210); generating a script corresponding to the content based on the target data, wherein the script includes at least one text segment (step 220); for each of the at least one text segment, determining a target character for broadcasting the text segment (step 230); for each text segment, determining acoustic parameters corresponding to the target character when broadcasting the text segment based on the text segment and the target character corresponding to the text segment (step 240); and generating the target content based on the script, the target character corresponding to each text segment, and the acoustic parameters (step 250).

[0035] Therefore, according to the embodiments of this disclosure, by automatically obtaining a script for generating target content based on the input target data, and further determining the role of each text segment in the script and the corresponding acoustic parameters when the role broadcasts the text segment, the creation process of generating target content is greatly simplified, and the generated target content is made smoother and more coordinated, thereby improving the creation efficiency and quality of target content.

[0036] According to embodiments of this disclosure, the target data includes at least one of the following: text data, web page links, document data, and voice data.

[0037] In some embodiments, target data can refer to information extracted from various data sources input by the user, such as text, URL links, files, and audio. A script can refer to a collection of structured text paragraphs that transform raw information, each with a clear theme and meaning. A target role can refer to the role used to deliver the corresponding paragraph. Due to different roles, such as a host, expert, or AI assistant, the content can be given a unique style and emotional tone based on their personality traits.

[0038] Therefore, the choice of characters can be determined based on the specific content of the text segment to ensure that the voice of the character most suitable for broadcasting the text segment is used to convey the content creation intention.

[0039] In some embodiments, acoustic parameters may include, but are not limited to, parameters such as speech rate, pitch, and tone, which collectively determine the quality and listening experience of the final output audio. Then, the script, the target character corresponding to each text segment, and the acoustic parameters can be integrated to generate the target content, such as a podcast.

[0040] According to embodiments of this disclosure, determining the target role for broadcasting the text segment includes: performing semantic analysis on the text segment to obtain semantic features; and determining the target role for broadcasting the text segment from a preset role pool based on the semantic features, wherein the preset role pool includes multiple roles.

[0041] In some embodiments, semantic analysis can be obtained through natural language processing (NLP), which involves understanding the text content to determine features such as the number of technical terms, the proportion of interrogative sentences, and the density of technical words. A pre-defined role pool can contain multiple roles with different attributes, such as host, expert, and AI assistant, each with its unique voice characteristics and expression style. For example, if a text segment contains more than 2 technical terms, it can be assigned to an expert; if a text segment has a high proportion of interrogative sentences, it can be assigned to a host, etc. Furthermore, role labels can be output based on pre-defined feature matching rules.

[0042] Therefore, according to embodiments of this disclosure, the semantic features include at least one of the following: topic, keywords, sentiment data, and interrogative sentences.

[0043] In some embodiments, the topic may refer to the core subject or issue discussed in the text segment. Keywords may be important words in the text segment that have special significance or appear frequently, such as technical terms, etc., which can be used to help understand the specific content and direction of the text segment. Sentiment data may reflect the emotional tendency expressed in the text segment, such as positive, negative, or neutral emotional attributes, and / or emotional intensity such as medium, low, or high, without limitation. In addition, identifying interrogative sentences in the text segment can help understand the author's or speaker's intention, guide the audience to think about questions, or meet their needs for answers.

[0044] Therefore, the method of determining target roles based on semantic analysis can not only improve the professionalism and attractiveness of the generated target content, but also enhance the user experience, making the generated audio content more in line with the needs and expectations of the audience, thereby improving the overall information delivery effect.

[0045] According to embodiments of this disclosure, determining a target role for broadcasting the text segment in a preset role pool based on the semantic features includes: determining at least one candidate role for broadcasting the text segment in the preset role pool based on the semantic features; for each candidate role, determining the target role corresponding to the first text segment closest to the text segment in the broadcast order, wherein the first text segment is a text segment among the at least one text segment; in response to determining that the target role corresponding to the text segment closest to the text segment is the same as the candidate role, determining the number of first text segments that the candidate role has continuously broadcast before the text segment, starting from the first text segment; and determining the target role for broadcasting the text segment from the at least one candidate role based on the number of first text segments and the semantic features.

[0046] In some embodiments, firstly, based on the semantic analysis results of the text segment, including factors such as topic, keywords, sentiment data, and interrogative sentences, at least one candidate character can be selected from a pre-defined character pool to ensure that the selected character can accurately convey the information and emotional tone of the text segment. Subsequently, to ensure the coherence and diversity of the audio content, each candidate character can be further evaluated.

[0047] Specifically, the target role corresponding to the first text segment that precedes and is closest to the current text segment can be identified. If the target role corresponding to the first text segment is found to be the same as a candidate role, the number of text segments that candidate role has continuously played before the first text segment can be calculated. This avoids audience fatigue or loss of interest caused by a single role continuously playing too many segments, or it can avoid confusion in the generated target content caused by frequent role changes. Based on the above calculations, the number of previously continuously played text segments and the original semantic features can be obtained, allowing for a final selection among multiple candidate roles. For example, if a candidate role has already continuously played multiple text segments, it may be inclined to choose another role to play the current text segment, even if the two are very similar in semantic features. Conversely, if a candidate role appears less frequently or has not been continuously played in previous text segments, it can be selected as the target role for the current text segment to ensure the coherence of the generated target content.

[0048] Therefore, by rationally allocating roles, the generated audio content is ensured to be both informative and avoids the use of a single or frequently changing broadcasting role, thereby improving the overall user experience.

[0049] According to embodiments of this disclosure, determining the acoustic parameters corresponding to a text segment based on the text segment and the target character corresponding to the text segment includes: determining the base weight corresponding to the target character; determining the emotional intensity corresponding to the text segment, and determining the emotional coefficient corresponding to the text segment based on the emotional intensity; determining the number of second text segments that the target character has continuously broadcast in at least one text segment, starting from the text segment; determining the context correction factor of the target character corresponding to the text segment based on the number of second text segments; determining the correction weight value corresponding to the text segment based on the base weight, the emotional coefficient, and the context correction factor; and determining the acoustic parameters corresponding to the target character when broadcasting the text segment based on the correction weight value.

[0050] In some embodiments, each target role can have a base weight, such as between 0.1 and 1, to represent the role's default voice characteristics or broadcasting style. For example, the base weight for an expert role could be set to 0.9, because their voice characteristics lean towards professionalism and authority; while the base weight for a presenter could be 0.7, because their voice characteristics lean towards friendliness and interactivity. The base weight can be set based on the role's basic attributes and is used to initially define the role's voice characteristics.

[0051] In some embodiments, emotional intensity can refer to data determined based on a text segment to characterize emotional fluctuations such as medium, low, and high. For example, it can be represented numerically to reflect the emotional tone contained in the text segment, and an emotional coefficient is determined accordingly. The emotional coefficient can be an adjustment parameter, for example, between 0 and 1, to reflect the normalized emotional intensity that the text segment needs to convey.

[0052] In some embodiments, the sigmoid function can be used to process emotional intensity to obtain an emotional coefficient. Since emotional intensity is often not linearly related—meaning the rate of emotional expression enhancement may not be constant as emotional intensity increases—the sigmoid function can be used to non-linearly map emotional intensity, non-linearly amplifying strong emotions. This allows the emotional coefficient to reasonably adjust the rate of change of vocal characteristics at different emotional intensity levels. Furthermore, by converting emotional intensity to a value within a range such as (0,1), it can be ensured that the emotional coefficient is not too large or too small, helping to maintain the consistency and stability of the audio output.

[0053] In some embodiments, starting from the current text segment, a context correction factor can be determined based on the number of second text segments that the target character has consecutively read. The context correction factor can be a dynamically adjusted item to reflect adaptive changes in the character's reading within the current context. For example, if a character has already read multiple segments consecutively, the context correction factor might be slightly reduced to introduce variation, making the voice more vivid and layered. Conversely, if this is the character's first or fewth reading, the correction factor might remain unchanged or increase slightly to highlight the character's unique style. Finally, a correction weight value can be determined by combining the base weight, sentiment coefficient, and context correction factor. The correction weight value can be used to determine the acoustic parameters corresponding to the target character reading the text segment.

[0054] In some embodiments, the roles in the role library can have their own preset initial acoustic parameters. For example, an expert role might have default settings for parameters such as tone and speaking speed, representing higher authority and professionalism. In practical applications, these preset acoustic parameters can be fine-tuned based on a correction weight value generated according to the specific content of the text segment, so that the acoustic parameters better match the content characteristics. For example, when the correction weight value is greater than the base weight value corresponding to the role, the speaking speed of the role when reading the text segment can be appropriately increased, and its tone can be made more authoritative.

[0055] According to embodiments of this disclosure, determining the corrected weight value corresponding to the text segment based on the base weight, the sentiment coefficient, and the context correction factor includes: determining the corrected weight value based on the following formula: Corrected weight value = base weight × (1 + sentiment coefficient) × context correction factor.

[0056] In some embodiments, by multiplying the base weight by (1 + sentiment coefficient), adjustments can be made appropriately based on the specific emotional tone of the text segment without altering the original style. This method preserves the character's fundamental traits while allowing for flexible variations based on content, thereby enhancing the richness and accuracy of the expression. Simultaneously, (1 + sentiment coefficient) ensures that the sentiment coefficient is always positive and greater than 0, preventing a reduction in weight due to sentiment.

[0057] According to embodiments of this disclosure, determining the corrected weight value corresponding to the text segment based on the base weight, the sentiment coefficient, and the context correction factor includes: determining the priority coefficient corresponding to the text segment, wherein the priority coefficient is determined by performing an importance analysis on the text segment; and determining the corrected weight value corresponding to the text segment based on the base weight, the sentiment coefficient, the context correction factor, and the priority coefficient.

[0058] In some embodiments, priority coefficients can be determined by performing an importance analysis on the text segments. Importance analysis may involve identifying the core role of a text segment within the entire script and its degree of importance to the overall message. Specifically, a priority coefficient, such as 0, 0.3, can be assigned based on the relative importance of each text segment by assessing its importance. For example, in an article about environmental protection, the section discussing the impact of global warming on ecosystems might be considered the most important part and therefore assigned a higher priority coefficient.

[0059] According to embodiments of this disclosure, determining the corrected weight value corresponding to the text segment based on the base weight, the sentiment coefficient, the context correction factor, and the priority coefficient includes: determining the corrected weight value based on the following formula: Corrected weight value = base weight × (1 + sentiment coefficient) × context correction factor + priority coefficient.

[0060] Therefore, by adding priority weights to the formula, it is ensured that the text segments carrying key information are appropriately emphasized, thereby improving the clarity of the generated content and its match with user needs.

[0061] According to embodiments of this disclosure, determining the acoustic parameters corresponding to the target character when broadcasting the text segment based on the modified weight value includes: determining the acoustic parameters corresponding to the target character when broadcasting the text segment using a large language model based on the target character corresponding to each text segment and the modified weight value corresponding to each text segment; and generating the target content based on the script, the target character corresponding to each text segment, and the acoustic parameters includes: generating the target content using a large language model based on the script, the target character corresponding to each text segment, and the acoustic parameters, wherein the acoustic parameters are used to guide the generation process of the target content.

[0062] In some embodiments, acoustic parameters may include, but are not limited to, parameters such as speech rate, pitch, and tone, which together determine the quality and listening experience of the final output audio.

[0063] In some examples, the steps of determining acoustic parameters and generating target operations described above can be implemented based on a large language model with a single input operation; alternatively, the output acoustic parameters can be obtained first based on the large language model, and then the data including the acoustic parameters can be input into the same or different large language models to generate the target content.

[0064] Therefore, by inputting scripts, target characters, and acoustic parameters into a large language model to generate target content, the automated process of generating target content based on the large language model can not only accurately convey the original information but also enhance the auditory experience through appropriate acoustic adjustments, further improving the standardization and consistency of content production.

[0065] According to embodiments of this disclosure, generating the target content based on the script, the target role corresponding to each text segment, and acoustic parameters includes: when the target roles corresponding to two consecutive text segments are different, generating the target content by adding a preset interval audio of a preset time period between the two consecutive text segments.

[0066] In some embodiments, when the target role changes between two adjacent text segments, a direct switch may result in abrupt content and negatively impact the user experience. Therefore, a preset interval audio can be inserted between text segments spoken by two different roles. The duration and content of this interval audio can be flexibly adjusted according to the actual situation. It can typically include short, silent or soft background music, and sometimes it may include transitional sound effects, such as ambient sound effects or gentle prompts, to achieve a natural transition.

[0067] In the above-described embodiment of generating target content through a large language model, a preset instruction can be defined to guide the large language model to add preset interval audio of a preset time period between two text segments when generating target content.

[0068] Therefore, adding intermittent audio significantly improved the professionalism and fluency of the audio content, making transitions between characters more natural and avoiding auditory shock caused by sudden character changes. At the same time, it enhanced the overall rhythm and layering of the content, further improving the quality of the target content and enhancing the listener experience.

[0069] According to embodiments of this disclosure, generating the target content based on the script, the target role corresponding to each text segment, and acoustic parameters includes: determining the theme information corresponding to the content; and determining the background information corresponding to the target content based on the theme information, wherein the background information includes at least one of the following: background music and animation.

[0070] In some embodiments, a theme may include the overall theme of the content. Additionally or alternatively, it may include a theme specific to each text segment within the content. For example, selecting appropriate background music based on a theme can create a sound environment that matches the atmosphere of the content.

[0071] In some embodiments, animation can refer to visual elements, such as charts, images, or motion graphics, that play in sync with text content. For example, animations matching a given theme can be selected from a pre-defined animation library. By adding animations that match the theme of the content when generating the target content, it is helpful to better convey the concepts or key information of the target content to the viewing user.

[0072] Therefore, appropriate background music and / or animation design enhance the expressiveness and appeal of the content, making information delivery more vivid and interesting; at the same time, it can also help users better understand some abstract or complex concepts.

[0073] According to embodiments of this disclosure, the script is subjected to key information identification to determine at least one key piece of information in the script and a timestamp corresponding to each of the at least one key piece of information; each of the at least one key piece of information is summarized to obtain at least one highlight piece of information; and the at least one highlight piece of information is associated with and stored in relation to the generated target content, wherein the at least one highlight piece of information is associated with the corresponding timestamp.

[0074] In some embodiments, the identified key information can be summarized using a large language model to obtain highlight information, such as a summary, designed to highlight its most important aspects or most attractive details. By associating the extracted highlight information with the generated target content and ensuring that each highlight information is associated with its corresponding timestamp, it is guaranteed that when a user accesses this content, they can not only hear the complete audio or see the complete video, but also quickly browse through each highlight information and quickly locate the specific position in the target content based on the corresponding timestamp.

[0075] Therefore, the combination of highlight information and timestamps significantly improves the usability and accessibility of content, enabling users to quickly locate and understand key information and enhancing the effectiveness of information delivery.

[0076] Figure 3 A schematic diagram of a target content display page according to an embodiment of the present disclosure is shown, wherein the target content is a podcast as an example. Figure 3 As shown, by selecting the "Highlights Summary" tab 310, the summarized highlights and their corresponding timestamps 320 can be displayed below. By clicking the play button, users can quickly locate the specific position within the target content based on the corresponding timestamp, thus enabling the playback of the specific key information corresponding to that highlight.

[0077] In some embodiments, such as Figure 3 As shown, the generated target content can also be associated with the links or data used to generate the target content (such as the data used to generate the script) for storage, so that when the generated target content is displayed on the content display page, its corresponding source data information can also be displayed synchronously.

[0078] In some examples, the script corresponding to the content is generated based on the target data. This can be done using a large language model, which generates the script based on the target data and other web page data found that match the target data. By synchronizing the source data with the target content as described above, the authority and objectivity of the target content can be improved, thus enhancing the user experience.

[0079] According to an embodiment of this disclosure, the method further includes: in response to obtaining target data for generating the target content on a dialog page, displaying reply information for identifying the target content in the form of a card on the dialog page; and in response to receiving a viewing operation for the reply information on the dialog page after the target content is generated, jumping to a details page for displaying the target content, so as to display the target content.

[0080] In some embodiments, target data for generating target content can be entered in the dialog page to generate target content based on the target data, such as through a large language model.

[0081] Furthermore, in some embodiments, after the target content is generated, when a request to view the reply information is received on the dialogue page, the user can be redirected to a details page to display the complete target content. For example, this details page could be a detailed display page of the target content (e.g., a...). Figure 3 The target content display page shown can also be any other page that can be used to jump to a detailed display page of the target content, such as a channel page. In other words, based on this detail page, more detailed target content information can be provided than that of the chat page, including but not limited to the complete target content, the script corresponding to the target content, highlight information and / or other relevant background information (such as target data) used to generate the target content, etc., without any restrictions.

[0082] Therefore, through the above embodiments, the synchronization of target content generation information across multiple pages or platforms can be achieved. Furthermore, it facilitates the generation of target content for users without requiring them to navigate to a specific page. The target page generated based on the dialogue page operation can be automatically synchronized to the details page used to display the target content. Thus, whenever a request to view the reply information is received on the dialogue page, the user is automatically redirected to the details page to display the complete target content, facilitating user operation and avoiding interruptions in user interaction.

[0083] According to an embodiment of this disclosure, the method further includes: during the generation of the target content, obtaining generation progress information of the target content, so that the response information displayed in the form of a card on the dialogue page includes the generation progress information.

[0084] In some embodiments, during the generation of the target content, the dialog page can display real-time generation progress information to the user. Specifically, the dialog page can display response information identifying the target content in the form of cards. The response information may include the generation progress information. In some examples, the response information may further include other basic information about the target content, such as topic, summary, etc.

[0085] For example, only after the target content has been generated, indicating that the progress information has been completed, can the user's request to view the content be accepted on the chat page, allowing the user to jump to the details page to display the complete target content. This instant feedback mechanism reduces user anxiety while waiting.

[0086] Figure 4A schematic diagram of a dialog page for generating target content according to an embodiment of the present disclosure is shown. Figure 4 As shown, the response information used to identify the generated target content is displayed in the form of a card on the chat page. This card includes generation progress information 410, as well as information such as the title and cover image. Users can click "Podcast Details" 420 to jump to the detailed display page of the target content, or click "Go to Podcast Channel Page" 430 to jump to the channel page that allows access to the detailed display page of the target content, thus achieving multi-page information synchronization.

[0087] According to embodiments of this disclosure, such as Figure 5 As shown, a data processing device 500 based on a large model is also provided, including: a content outline generation module 510, configured to acquire target data for generating the target content, wherein the target content includes audio; a script generation module 520, configured to generate a script corresponding to the content based on the target data, wherein the script includes at least one text segment; a role determination module 530, configured to determine a target role for broadcasting each text segment in the at least one text segment; a feature determination module 540, configured to determine the acoustic parameters corresponding to the target role when broadcasting the text segment, based on the text segment and the target role corresponding to the text segment; and a content generation module 550, configured to generate the target content based on the script, the target role corresponding to each text segment, and the acoustic parameters.

[0088] Here, the operation of each of the above units 510 to 550 of the data processing device 500 based on the large model is similar to the operation of steps 210 to 250 described above, and will not be repeated here.

[0089] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0090] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0091] refer to Figure 6The present invention describes a structural block diagram of an electronic device 600 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0092] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0093] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, output unit 607, storage unit 608, and communication unit 609. Input unit 606 can be any type of device capable of inputting information to electronic device 600. Input unit 606 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and can include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 607 can be any type of device capable of presenting information, and can include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 608 can include, but is not limited to, disk and optical disk. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and can include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0094] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform method 200 by any other suitable means (e.g., by means of firmware).

[0095] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0096] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0097] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0098] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0099] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0100] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0101] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0102] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A method for generating target content based on a large model, comprising: Obtain target data for generating the target content, wherein the target content includes audio; A script corresponding to the content is generated based on the target data, wherein the script includes at least one text segment; For each of the at least one text segment, determine the target role for broadcasting that text segment; For each text segment, based on the text segment and the target role corresponding to the text segment, determine the acoustic parameters corresponding to the target role when broadcasting the text segment; as well as The target content is generated based on the script, the target role corresponding to each text segment, and the acoustic parameters.

2. The method of claim 1, wherein, The target roles for broadcasting this text segment include: Perform semantic analysis on this text segment to obtain semantic features; and Based on the semantic features, a target role for broadcasting the text segment is determined from a preset role pool, wherein the preset role pool includes multiple roles.

3. The method as described in claim 2, wherein, The semantic features include at least one of the following: topic, keywords, sentiment data, and interrogative sentences.

4. The method as described in claim 2 or 3, wherein, Based on the semantic features, the target roles for broadcasting the text segment are determined from the preset role pool, including: Based on the semantic features, at least one candidate role for broadcasting the text segment is determined from the preset role pool; For each candidate role, determine the target role corresponding to the first text segment that is closest to the text segment before the text segment in the broadcast order, wherein the first text segment is a text segment among the at least one text segment; In response to determining that the target role corresponding to the text segment closest to the text segment is the same as the candidate role, starting from the first text segment, determine the number of first text segments that the candidate role has continuously broadcast before that text segment; and Based on the number of the first text segments and the semantic features, a target role for broadcasting the text segment is determined from the at least one candidate role.

5. The method of claim 1, wherein, Based on the text segment and the target role corresponding to the text segment, the acoustic parameters corresponding to the text segment are determined as follows: Determine the base weights corresponding to the target role; Determine the emotional intensity corresponding to the text segment, and determine the emotional coefficient corresponding to the text segment based on the emotional intensity; Starting from this text segment, determine the number of second text segments that the target character has continuously broadcast in at least one text segment; Based on the number of the second text segments, determine the context correction factor of the target role corresponding to the text segment; Based on the base weights, the sentiment coefficients, and the context correction factor, the corrected weight value corresponding to the text segment is determined; and The acoustic parameters corresponding to the target character's recitation of the text segment are determined based on the corrected weight value.

6. The method as described in claim 1 or 5, wherein, Determining the acoustic parameters corresponding to the target character when broadcasting the text segment based on the modified weight value includes: determining the acoustic parameters corresponding to the target character when broadcasting the text segment through a large language model based on the target character corresponding to each text segment and the modified weight value corresponding to each text segment; Generating the target content based on the script, the target role corresponding to each text segment, and the acoustic parameters includes: generating the target content through a large language model based on the script, the target role corresponding to each text segment, and the acoustic parameters, wherein the acoustic parameters are used to guide the generation process of the target content.

7. The method of claim 5, wherein, Based on the base weight, the sentiment coefficient, and the context correction factor, the correction weight value corresponding to this text segment is determined as follows: The corrected weight value is determined based on the following formula: Corrected weight value = base weight × (1 + sentiment coefficient) × context correction factor.

8. The method of claim 5 or 6, wherein, Based on the base weight, the sentiment coefficient, and the context correction factor, the correction weight value corresponding to this text segment is determined as follows: Determine the priority coefficient corresponding to the text segment, wherein the priority coefficient is determined by performing an importance analysis on the text segment; and Based on the base weight, the sentiment coefficient, the context correction factor, and the priority coefficient, the correction weight value corresponding to the text segment is determined.

9. The method of claim 8, wherein, Based on the base weight, the sentiment coefficient, the context correction factor, and the priority coefficient, the correction weight value corresponding to the text segment is determined as follows: The corrected weight value is determined based on the following formula: Corrected weight value = base weight × (1 + sentiment coefficient) × context correction factor + priority coefficient.

10. The method of claim 1, wherein, Based on the script, the target role corresponding to each text segment, and the acoustic parameters, generating the target content includes: When the target characters corresponding to the two preceding and following text segments are different, the target content is generated by adding audio at a preset interval for a preset time period between the two preceding and following text segments.

11. The method of claim 1, wherein, Based on the script, the target role corresponding to each text segment, and the acoustic parameters, generating the target content includes: Determine the topic corresponding to the content; and Based on the theme, the background information corresponding to the target content is determined, wherein the background information includes at least one of the following: background music and animation.

12. The method of claim 1, further comprising: The script is subjected to key information identification to determine at least one key piece of information in the script and the timestamp corresponding to each of the at least one key piece of information; Summarize each of the at least one key piece of information to obtain at least one highlight piece of information; as well as The at least one highlight information is associated with the generated target content and stored, wherein the at least one highlight information is associated with the corresponding timestamp.

13. The method of claim 1, further comprising: In response to obtaining target data for generating the target content on the chat page, reply information for identifying the target content is displayed in the chat page in the form of a card; as well as Upon receiving a view operation for the reply information on the dialog page after the target content has been generated, the system redirects to a details page to display the target content.

14. The method of claim 13, further comprising: During the generation of the target content, the generation progress information of the target content is obtained so that the response information displayed in the form of a card on the dialog page includes the generation progress information.

15. The method of claim 1, wherein, The target data includes at least one of the following: text data, web page links, document data, and voice data.

16. A target content generation device based on a large model, comprising: The data acquisition module is configured to acquire target data for generating the target content, wherein the target content includes audio. The script generation module is configured to generate a script corresponding to the content based on the target data, wherein the script includes at least one text segment; The role determination module is configured to determine the target role for broadcasting each text segment in the at least one text segment; The feature determination module is configured to, for each text segment, determine the acoustic parameters corresponding to the target role when the target role broadcasts the text segment, based on the text segment and the target role corresponding to the text segment; as well as The content generation module is configured to generate the target content based on the script, the target role corresponding to each text segment, and the acoustic parameters.

17. An electronic device comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-15.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-15.

19. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 1-15.