Play obtaining method, device, equipment, medium and program product

By receiving podcast playback requests and using a target model to generate podcast content that matches the user's style, the system solves the problems of real-time performance and personalization in in-vehicle podcast systems, enabling the generation and playback of real-time personalized podcast content and improving the user experience.

CN121728147APending Publication Date: 2026-03-24STARRY SKY PLAN (SHANGHAI) AUTOMOBILE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing in-vehicle podcast systems cannot meet users' needs for real-time updates and personalization. Traditional podcast content has a long update cycle and cannot provide the latest information in a timely manner. Furthermore, the broadcast mode cannot be customized according to user preferences, especially for humor podcasts, which are difficult to match the humor preferences of different users.

Method used

By receiving podcast playback requests, the system uses a target model to generate podcast content that matches the user's style. The target model expands based on the reasoning ability of thought chains, generates multiple jokes, and filters high-quality jokes through vector similarity and reward mechanisms. The model is then optimized by combining feedback information to achieve real-time generation and playback of personalized podcast content.

Benefits of technology

It enables real-time and personalized customization of in-car podcast content, improving user experience and satisfaction, especially for humor podcasts, which can better cater to the humor preferences of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121728147A_ABST
    Figure CN121728147A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a broadcast obtaining method and device, equipment, a medium and a program product, and relates to the technical field of vehicles. The method comprises the steps that firstly, a broadcaster playing request containing theme information of a broadcaster to be played is received, then a target broadcaster generated based on a target model matched with the style of the requested broadcaster is obtained, the theme information serves as input of the target model, and it is ensured that the generated target broadcaster content is matched with a preset style. According to the method, the problems that the real-time performance of existing vehicle-mounted broadcasting is not high, and customized services cannot be provided when all users receive the same content are solved, so that the comprehensive requirements of the users on the timeliness and individuation of the content are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle technology, and more particularly to a podcast acquisition method, apparatus, device, medium, and program product. Background Technology

[0002] With the rapid development of intelligent vehicle technology, in-vehicle entertainment systems have gradually become an important component in enhancing the user's driving experience. In-vehicle podcasts, as a form of audio content that requires no visual interaction, have become a core way for users to access news, entertainment, and knowledge due to their safety and convenience in driving scenarios.

[0003] However, current in-car podcast applications face multiple challenges: First, there is a high demand for real-time updates. Users expect to hear the latest breaking news and events while driving, but traditional podcast content is mostly pre-recorded, fixed programs with long update cycles, making it difficult to meet dynamic needs. Second, there is a strong demand for personalization. Existing podcasts use a "one-to-many" broadcast model, where all users receive the same content, failing to provide customized services. This is especially true for humor podcasts, where the existing model struggles to cater to diverse user preferences, resulting in low user satisfaction.

[0004] Therefore, there is an urgent need for a car podcast solution that can generate content in real time, personalize it, and combine fun and professionalism to meet users' comprehensive needs for timely and personalized content. Summary of the Invention

[0005] This application provides a podcast acquisition method, apparatus, device, medium, and program product to meet users' comprehensive needs for the timeliness and personalization of podcast content.

[0006] In a first aspect, embodiments of this application provide a podcast acquisition method, the method comprising:

[0007] Receive a podcast playback request, the podcast playback request including the topic information of the podcast to be played and information indicating the style of the podcast;

[0008] Based on the podcast playback request, a target podcast is obtained. The target podcast is generated based on a target model, which is a model that matches the style of the podcast. The input of the target model includes the topic information, and the content of the target podcast is adapted to the style of the podcast.

[0009] In one possible implementation, the podcast style includes a joke style, and the target model is specifically used for:

[0010] Based on the reasoning ability of the thought chain, the topic information is diverged to obtain divergent information, which includes at least contradictory information;

[0011] Based on the contradiction information and the preset joke generation logic, multiple jokes are generated;

[0012] Based on the scores of the multiple jokes, the target joke is determined;

[0013] The target podcast is obtained based on the target joke.

[0014] In one possible implementation, the step of diverging the topic information based on the ability of thought chain reasoning to obtain divergent information includes:

[0015] Based on the reasoning ability of the thought chain, the topic information is diverged at multiple levels to obtain the divergent information;

[0016] The multi-level divergence includes a first-level divergence, a second-level divergence, and a third-level divergence;

[0017] The first layer of divergence is used to obtain core concept information related to the topic information;

[0018] The second layer of divergence is used to derive specific scenario information from each of the core concept information;

[0019] The third layer of divergence is used to identify the contradiction information in each of the specific scene information.

[0020] In one possible implementation, the preset joke generation logic includes logic for establishing expectations, breaking expectations, and interpreting them.

[0021] In one possible implementation, the target model includes at least one of the following mechanisms:

[0022] A vector similarity mechanism is used to deduplicate the multiple jokes based on vector similarity;

[0023] A reward mechanism is used to obtain individual scores for each of the multiple jokes based on manually labeled data;

[0024] A reinforcement learning mechanism is used to maximize the expected reward based on the scores of the multiple jokes.

[0025] In one possible implementation, the scores for the multiple jokes are calculated using a weighted algorithm of multiple parameters, wherein the multiple parameters include at least: humor, unexpectedness, relevance, comprehensibility, and offensiveness; and the weighting coefficient for offensiveness is negative.

[0026] In one possible implementation, obtaining the target podcast based on the target joke includes:

[0027] A dialogue structure outline is obtained based on the target joke;

[0028] Based on the preset character modeling, a leading role is assigned to each stage in the dialogue structure outline, and the dialogue flow of each stage is generated to obtain the dialogue of each stage; the character modeling includes the following multiple dimensions of various characters: basic attributes, personality traits, knowledge domain, language style and dialogue habits.

[0029] The target podcast is obtained based on the dialogue at each stage.

[0030] In one possible implementation, obtaining the target podcast based on the dialogue at each stage includes:

[0031] The dialogue at each stage is checked for consistency with the persona and corrected to obtain the target podcast.

[0032] In one possible implementation, the method further includes:

[0033] Play the target podcast and obtain user feedback on the target podcast;

[0034] Based on the feedback information, instructions are given to process the target model.

[0035] In one possible implementation, obtaining the target podcast based on the podcast playback request includes:

[0036] Send the podcast playback request to the cloud device;

[0037] Receive the target podcast from the cloud device.

[0038] Secondly, embodiments of this application provide a podcast acquisition device, the device comprising:

[0039] The receiving module is used to receive podcast playback requests, which include the topic information of the podcast to be played and information indicating the style of the podcast;

[0040] The acquisition module is used to acquire a target podcast based on the podcast playback request. The target podcast is generated based on a target model, which is a model that matches the podcast style. The input of the target model includes the topic information, and the content of the target podcast is adapted to the podcast style.

[0041] In one possible implementation, the device further includes: a divergence module;

[0042] The divergent module is used to diverge the topic information based on the reasoning ability of the thought chain to obtain divergent information, which includes at least contradiction information.

[0043] The device further includes: a generation module;

[0044] The generation module is used to generate multiple jokes based on the contradiction information and the preset joke generation logic;

[0045] The device further includes: a determining module;

[0046] The determining module is used to determine the target joke based on the scores of the plurality of jokes;

[0047] The acquisition module is specifically used to obtain the target podcast based on the target joke.

[0048] In one possible implementation, the divergence module is specifically used to diverge the topic information at multiple levels based on the reasoning ability of the thought chain to obtain the divergent information;

[0049] The multi-level divergence includes a first-level divergence, a second-level divergence, and a third-level divergence;

[0050] The first layer of divergence is used to obtain core concept information related to the topic information;

[0051] The second layer of divergence is used to derive specific scenario information from each of the core concept information;

[0052] The third layer of divergence is used to identify the contradiction information in each of the specific scene information.

[0053] In one possible implementation, the preset joke generation logic includes logic for establishing expectations, breaking expectations, and interpreting them.

[0054] In one possible implementation, the target model includes at least one of the following mechanisms:

[0055] A vector similarity mechanism is used to deduplicate the multiple jokes based on vector similarity;

[0056] A reward mechanism is used to obtain individual scores for each of the multiple jokes based on manually labeled data;

[0057] A reinforcement learning mechanism is used to maximize the expected reward based on the scores of the multiple jokes.

[0058] In one possible implementation, the scores for the multiple jokes are calculated using a weighted algorithm of multiple parameters, wherein the multiple parameters include at least: humor, unexpectedness, relevance, comprehensibility, and offensiveness; and the weighting coefficient for offensiveness is negative.

[0059] In one possible implementation, the determining module is further configured to obtain a dialogue structure outline based on the target joke;

[0060] The generation module is also used to assign a leading role to each stage in the dialogue structure outline based on the preset character modeling, and to generate the dialogue flow of each stage to obtain the dialogue of each stage; the character modeling includes the following dimensions of various characters: basic attributes, personality traits, knowledge domain, language style and dialogue habits.

[0061] The acquisition module is specifically used to obtain the target podcast based on the dialogue at each stage.

[0062] In one possible implementation, the acquisition module is specifically used to perform persona consistency detection and correction on the dialogue at each stage to obtain the target podcast.

[0063] In one possible implementation, the device further includes: a playback module;

[0064] The playback module is used to play the target podcast;

[0065] The acquisition module is also used to acquire user feedback information regarding the target podcast;

[0066] The device further includes: an indicator module;

[0067] The instruction module is used to instruct the processing of the target model based on the feedback information.

[0068] In one possible implementation, the device further includes: a transmitting module;

[0069] The sending module is used to send the podcast playback request to the cloud device;

[0070] The receiving module is specifically used to receive the target podcast from the cloud device.

[0071] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.

[0072] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.

[0073] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.

[0074] The podcast acquisition method, apparatus, device, medium, and program product provided in this application embodiment receive a podcast playback request. The podcast playback request includes topic information of the podcast to be played and information indicating the podcast style. Based on the podcast playback request, a target podcast is acquired. The target podcast is generated based on a target model, which is a model that matches the podcast style. The input of the target model includes topic information, and the content of the target podcast is adapted to the podcast style. This method, by introducing a style-matching generation model, solves the problems of insufficient real-time podcast content and the inability to meet user customization needs due to unified push in existing solutions. It enables podcast content to be adjusted in real time according to the style preferences of different users, thereby improving user experience and satisfaction. Attached Figure Description

[0075] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0076] Figure 1 Flowchart of the podcast acquisition method provided in the embodiments of this application Figure 1 ;

[0077] Figure 2 The flowchart of the podcast acquisition method provided in this embodiment Figure 2 ;

[0078] Figure 3 The flowchart of the podcast acquisition method provided in this embodiment Figure 3 ;

[0079] Figure 4 This is a schematic diagram of the structure of the podcast acquisition device provided in the embodiments of this application;

[0080] Figure 5 A schematic diagram of the structure of the electronic device provided in this application.

[0081] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0082] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0083] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.

[0084] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0085] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0086] In the embodiments of this application, the use of terms such as "first" and "second" is to distinguish between identical or similar items that have essentially the same function and effect. For example, "first electronic device" and "second electronic device" are merely used to distinguish different electronic devices and do not limit their order of execution. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.

[0087] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following associated objects have an "or" relationship.

[0088] The following is an explanation of some terms used in the embodiments of this application:

[0089] A Large Language Model (LLM) is a deep learning model trained on large amounts of text data, enabling it to generate natural language text or understand the meaning of language text. A LLM consists of multiple layers, including an input layer, hidden layers, and an output layer. The input layer receives and encodes the input text, converting it into a numerical representation. The hidden layers process the input information through multiple layers of neural networks, including self-attention mechanisms, to capture complex semantics and contextual relationships in the text. The output layer converts the processed information into understandable language output, such as generating sentences or answering questions.

[0090] Chain-of-Thought (CoT) is a hinting engineering technique for Large Language Models (LLMs) and a core concept describing the complex reasoning process of the model. Its core is to guide the model to imitate the step-by-step thinking logic of humans, breaking down complex tasks into a series of ordered and interpretable intermediate reasoning steps, and finally deriving a conclusion.

[0091] Thought chain reasoning ability is the capacity of Large Language Models (LLMs) to simulate human thinking by generating multi-step reasoning processes when solving complex tasks. This approach enables the model to break down problems, deduce step by step, and ultimately arrive at the answer, performing particularly well in areas such as mathematical, logical, and common-sense reasoning.

[0092] Sentence-BERT (SBERT for short) is a model based on BERT (Bidirectional Encoder Representations from Transformers) designed for sentence embedding computation. Compared to the traditional BERT model, Sentence-BERT performs better in handling sentence-level semantic relevance tasks. This model can learn semantic representations at the sentence and paragraph levels by fine-tuning the original BERT model, thereby calculating the similarity between sentences.

[0093] DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a density-based clustering algorithm designed to discover clusters of arbitrary shapes by identifying high-density regions of data points. Compared to deduplication methods that directly compare pairs of data points to calculate similarity, this algorithm offers significant advantages when processing large volumes of joke text. It can efficiently group similar texts, greatly reduce computational costs, and improve the comprehensiveness and accuracy of deduplication.

[0094] With the widespread adoption of intelligent in-vehicle systems, in-vehicle podcasts have become an important form of entertainment and information access for users during commutes and long-distance driving. In the in-vehicle environment, users are typically on the move, making the demand for real-time, personalized, and engaging content particularly prominent.

[0095] Existing in-vehicle podcasts mainly adopt a pre-recorded and uniformly distributed generation model. The content is mostly pre-produced fixed programs, which are then distributed to in-vehicle users through a "one-to-many" broadcast-style push mechanism.

[0096] However, this generation method has significant drawbacks: on the one hand, because the content is pre-recorded and the update cycle is long, it cannot meet the needs of users to listen to the latest hot news and breaking events while driving, and it is difficult to adapt to the dynamically changing information environment. On the other hand, the "one-to-many" broadcasting model means that all users can only passively receive the same content and cannot obtain customized services according to their own preferences. Especially for humor podcasts, different users have different understandings and preferences for humor, and the existing model is difficult to match these differences, resulting in low user satisfaction.

[0097] Therefore, there is an urgent need for an in-vehicle podcast solution that can generate content in real time, personalize it, and combine fun and professionalism to meet users' comprehensive needs for timely, diverse, and immersive content.

[0098] To address the aforementioned issues, this application provides a podcast acquisition method. First, it receives a podcast playback request containing the topic information of the podcast to be played. Then, it acquires a target podcast generated based on a target model that matches the style of the requested podcast. This target model takes the topic information as input, ensuring that the generated target podcast content is adapted to a preset style. This method solves the problems of low real-time performance of existing in-vehicle podcasts and the inability to provide customized services when all users receive the same content, thus meeting users' comprehensive needs for timely, diverse, and immersive content.

[0099] The execution subject of this application embodiment can be, for example, a vehicle-mounted terminal.

[0100] The technical solutions of this application will be described in detail below with reference to specific embodiments. The specific embodiments described below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0101] Figure 1 Flowchart of the podcast acquisition method provided in the embodiments of this application Figure 1 ,like Figure 1 As shown, the method includes:

[0102] S101. Receive a podcast playback request. The podcast playback request includes the topic information of the podcast to be played and information indicating the style of the podcast.

[0103] Understandably, the topic information refers to the content area covered by the podcast to be played, such as technology, entertainment, health, and current social events.

[0104] Podcast style refers to users' preferences for the presentation format of podcast content, such as humor, seriousness, interviews, storytelling, etc.

[0105] By receiving podcast requests, it's possible to clarify not only the content areas users want to listen to, but also their preferences for the podcast's presentation format. For example, when a user requests "tell me a joke about eating at a restaurant," the topic information is "eating at a restaurant," and the podcast style information is "joke," meaning the user wants the podcast content to revolve around a restaurant dining scenario and be presented in a humorous and entertaining way.

[0106] There are several ways to receive podcast playback requests that include topic and style information: First, voice interaction, where users can issue commands directly through the in-vehicle voice assistant, such as "tell me a joke about eating at a restaurant," and the in-vehicle terminal will recognize the topic and style information in the command and complete the reception; second, touch operation, where users can initiate requests by selecting preset topic categories and style tags on the podcast application interface of the in-vehicle central control screen, and the requests will be directly received by the in-vehicle terminal; and third, linking to mobile terminals, where users can set topic and style preferences on the podcast application software on their mobile phones and synchronize the requests to the in-vehicle terminal through the vehicle-machine interconnection function.

[0107] S102. Based on the podcast playback request, obtain the target podcast. The target podcast is generated based on the target model, which is a model that matches the style of the podcast. The input of the target model includes topic information and the content of the target podcast is adapted to the podcast style.

[0108] Understandably, different podcast styles correspond to different target models. For example, humorous podcasts correspond to models with content generation capabilities, while news podcasts correspond to models with information retrieval and integration capabilities.

[0109] By matching the target model with the podcast style and inputting the topic information into the model, podcast content that fits the specified topic and conforms to the preset style can be generated to meet users' personalized listening needs. At the same time, it effectively solves the problem that traditional podcast content is mostly pre-recorded fixed programs that cannot be customized and generated in real time, and greatly improves the personalization and scene adaptability of podcast content.

[0110] In this step, firstly, the target model is determined based on the podcast style information contained in the podcast playback request: if the podcast style information is "joke", the model corresponding to the joke is triggered, and then the topic information contained in the podcast playback request is input into the model, which generates joke-type podcast content that matches the topic; if the podcast style information is "news", the model corresponding to the news is triggered, which can directly capture news content related to the topic and organize it into a podcast.

[0111] The podcast acquisition method provided in this application first receives a podcast playback request, which includes the topic information of the podcast to be played and information indicating the podcast style. Finally, based on the podcast playback request, a target podcast is acquired. The target podcast is generated based on a target model, which is a model that matches the podcast style. The input of the target model includes topic information, and the content of the target podcast is adapted to the podcast style. This method, by introducing a style-matching generation model, solves the problems of insufficient real-time podcast content and the inability to meet user customization needs due to unified push notifications in existing solutions. It enables podcast content to be adjusted in real time according to different users' style preferences, thereby improving user experience and satisfaction.

[0112] Figure 2 The flowchart of the podcast acquisition method provided in this embodiment Figure 2 .like Figure 2 As shown. This embodiment is... Figure 1 Based on the embodiments, the podcast acquisition method is described in detail below. The podcast acquisition method provided in this embodiment includes:

[0113] S201. Receive a podcast playback request. The podcast playback request includes the topic information of the podcast to be played and information indicating the style of the podcast.

[0114] The explanation of step S201 is the same as that in the above embodiments, and will not be repeated here.

[0115] S202, Send a podcast playback request to the cloud device.

[0116] Understandably, in-vehicle terminals are limited by hardware computing power, storage resources, and model deployment costs, making it impossible to directly support the operation and computation of complex target models; while cloud devices have sufficient resources to support model operation and content processing, and can stably run target models corresponding to various styles.

[0117] Therefore, by sending podcast playback requests to cloud devices, various models deployed in the cloud can be used to complete the output of podcast content. For example, content can be generated for jokes and retrieved and organized for news, thereby achieving personalized output of in-vehicle podcasts and reducing the hardware load on in-vehicle terminals.

[0118] S203. Receive the target podcast from the cloud device. The target podcast is generated based on the target model, which is a model that matches the podcast style. The input of the target model includes topic information and the content of the target podcast is adapted to the podcast style.

[0119] Understandably, by receiving target podcasts from cloud devices, in-vehicle terminals can acquire podcast content resources that are both relevant to the user's specified topics and conform to the preset style, thereby meeting the user's personalized listening needs in driving scenarios.

[0120] S204. Play the target podcast and obtain user feedback on the target podcast.

[0121] Understandably, feedback refers to the evaluations users give after listening to a target podcast, focusing on aspects such as the content's relevance to the theme, style, and entertainment value. Feedback includes two categories: explicit and implicit. Explicit feedback includes directly perceptible information such as subjective evaluations expressed verbally and intuitive feelings conveyed through facial expressions. Implicit feedback includes behavioral data that indirectly reflects satisfaction, such as listening duration, whether the podcast was replayed, and whether the user switched content midway through.

[0122] By playing target podcasts and obtaining user feedback, we can not only present personalized podcast content to users to meet their listening needs in driving scenarios, but also obtain user evaluation information on podcast content, providing data support for subsequent optimization of model matching accuracy and improvement of content personalization.

[0123] The way to play the target podcast is that after the in-vehicle terminal receives the target podcast content transmitted from the cloud, it processes the data through the built-in audio decoding program and then plays the content out through the speaker.

[0124] Methods for obtaining user feedback on a target podcast include, for example, collecting real-time interactive feedback from users, such as capturing user facial expressions through image recognition or extracting user language evaluations through speech recognition; or automatically collecting user listening behavior data, such as listening duration, number of plays, pause or skip actions.

[0125] S205. Based on the feedback information, instruct the processing of the target model.

[0126] Understandably, user feedback can directly reflect their satisfaction with the target podcast content. Therefore, by using feedback to guide the processing of the target model, data support can be provided for the iterative optimization of the model, thereby driving the target model to complete updates and optimizations, making it more in line with users' personalized listening preferences in the subsequent generation of podcast content.

[0127] In this step, firstly, the vehicle terminal sends feedback information to the cloud device based on the user's feedback information, instructing the cloud device to iteratively update the target model; then, the cloud device receives the feedback information and updates and optimizes the target model based on the feedback information.

[0128] The podcast acquisition method provided in this application receives a playback request containing the topic information and style indication of the podcast to be played, sends it to a cloud device, and the cloud device generates a target podcast with content adapted to the user's specified style in real time, and returns it to the local playback. Simultaneously, user feedback is collected to instruct the processing of the target model. This method solves the problems of high real-time requirements for in-vehicle podcasts and the inability to provide customized services for all users receiving the same content. Especially for humor podcasts, it effectively overcomes the problem that existing models struggle to cater to different user preferences, thus providing personalized in-vehicle podcast content and significantly improving user satisfaction with in-vehicle podcasts.

[0129] Figure 3 The flowchart of the podcast acquisition method provided in this embodiment Figure 3 .like Figure 3 As shown. This embodiment, based on the above embodiments, provides a detailed explanation of the process for generating target podcasts from target models when the podcast style includes joke styles. The podcast acquisition method provided in this embodiment includes:

[0130] S301. Based on the reasoning ability of the thinking chain, diverge the topic information to obtain divergent information, which includes at least contradictory information.

[0131] Understandably, the ability to reason through thought processes refers to the ability of a target model to mimic human thought processes when processing topical information, to deduce and solve problems step by step according to preset logical steps, and to output results.

[0132] Directly inputting topic information into the target model to generate jokes can easily result in empty content, forced jokes, or deviation from the topic; while thought chain reasoning can guide the model to break down the topic step by step and logically, avoiding the blindness of creative divergence.

[0133] Therefore, by leveraging the reasoning ability of the thought chain, the target model can hierarchically diverge the topic information to obtain divergent information containing contradictions, which can provide a foundation for the generation of subsequent jokes.

[0134] In one possible implementation, this application provides a possible implementation method for diverging topic information based on the ability of thought chain reasoning to obtain divergent information, including: diverging topic information at multiple levels based on the ability of thought chain reasoning to obtain divergent information; wherein, the multi-level divergence includes a first-level divergence, a second-level divergence, and a third-level divergence; the first-level divergence is used to obtain core concept information related to the topic information; the second-level divergence is used to diverge specific scene information from each core concept information; and the third-level divergence is used to identify contradictions in each specific scene information.

[0135] Understandably, contradictory information may contain counterintuitive content. Counterintuitive content refers to phenomena or viewpoints that contradict common sense, everyday experience, or conventional logic. For example, in the scenario of "eating at a restaurant," the statement that "the extra-spicy dish ordered by the customer has no spiciness at all" is counterintuitive.

[0136] The target model utilizes the reasoning ability of the thought chain to perform a three-level progressive divergence on the input topic information. The specific operations are as follows: First, the first level of divergence is performed to extract core concept information that is strongly related to the topic information; then, the second level of divergence is performed to diverge the corresponding specific daily life scenario information for each core concept obtained in the first level; finally, the third level of divergence is performed to identify contradictions or counterintuitive information with humorous potential from each specific scenario in the second level.

[0137] For example, assuming the input topic is "eating at a restaurant," the contextual reasoning layer of the target model performs three layers of divergence as follows: First, it performs the first layer of divergence on "eating at a restaurant" to obtain concepts such as "ordering food," "serving food," "paying the bill," and "food taste." Then, based on the conceptual information obtained from the first layer of divergence, it performs the second layer of divergence, diverging specific scenarios such as "struggling to choose dishes for half an hour" and "being recommended expensive dishes by the waiter" for "ordering food," and specific scenarios such as "the extra spicy dish ordered is not spicy at all" and "the signature dish is not as good as the home-style dish" for "food taste." Finally, it performs the third layer of divergence for each specific scenario, identifying the counterintuitive contradiction "the extra spicy dish the waiter said was only mildly spicy, which is inconsistent with expectations" from the scenario "the extra spicy dish ordered was not spicy at all."

[0138] S302. Generate multiple jokes based on the contradiction information and the preset joke generation logic.

[0139] Understandably, contradiction information is the basis for generating jokes. It can provide the target model with contrasting inspirations that have humorous potential to generate jokes, and avoid the target model generating jokes that are empty, far-fetched, or off-topic. The preset joke generation logic is the creative execution logic of the target model to generate jokes.

[0140] Therefore, by combining contradiction information with pre-defined joke generation logic, the target model can generate multiple candidate jokes with complete structure and clear humor.

[0141] In one possible implementation, the pre-defined joke generation logic includes logic for establishing expectations, breaking expectations, and interpreting them.

[0142] Understandably, building expectations means describing a common scenario that is close to life or a view that is generally accepted by the public, so that the audience can make a psychological prediction that conforms to conventional logic.

[0143] Breaking expectations refers to introducing an unexpected, illogical, or counterintuitive twist based on established expectations, thereby disrupting the audience's preconceived notions and creating a sense of cognitive dissonance.

[0144] The logic of explanation refers to providing a seemingly reasonable and self-consistent explanation for the cognitive dissonance caused by the unexpected step. This explanation should not only resolve the previous contradiction, but also fit the theme and contradiction of the joke, so that the audience can feel the humor at the same time as suddenly realizing the truth, and finally complete the logical loop of the joke.

[0145] In the process of generating jokes, the target model first generates common scenarios or widely accepted viewpoints based on the contradiction information. Then, it introduces unexpected twists and turns based on the scenario or viewpoint, using the contradiction information to create a humorous effect. Next, it provides a clever and seemingly reasonable explanation for the twist, completing the joke loop. Finally, it splices these three text fragments into a complete joke, thereby generating multiple candidate jokes in batches.

[0146] S303. Based on the scores of multiple jokes, determine the target joke.

[0147] It is understandable that the quality of the multiple jokes generated by the target model according to the preset joke generation logic and contradiction information varies, and there may be problems such as forced jokes and deviation from the topic; while the score is a quantitative evaluation of the joke from multiple dimensions such as humor, relevance, and comprehensibility.

[0148] Therefore, when selecting high-quality jokes, the target model uses the scores of multiple jokes as a basis to objectively and reasonably eliminate low-quality jokes, ensuring that the final target jokes meet the user's expected style and needs.

[0149] In one possible implementation, the target model includes at least one of the following mechanisms:

[0150] The first method is the vector similarity mechanism, which is used to deduplicate multiple jokes based on vector similarity.

[0151] Understandably, the vector similarity mechanism involves: first, using the Sentence-BERT model to convert multiple joke texts into high-dimensional vectors, and then measuring the semantic similarity of jokes by calculating the cosine similarity between vectors; subsequently, using the DBSCAN clustering algorithm to automatically group semantically similar jokes into the same category based on a similarity threshold; finally, filtering jokes within each category and retaining the most representative entries (e.g., retaining the joke with the highest score), thereby achieving joke deduplication.

[0152] The target model first uses the Sentence-BERT model to convert multiple joke texts into high-dimensional vectors, then calculates the cosine similarity between these vectors to quantify the degree of semantic similarity; finally, based on the similarity results, the DBSCAN clustering algorithm is used to automatically classify jokes with high semantic similarity into the same category, thereby completing the deduplication of jokes.

[0153] The second type is a reward mechanism, which is used to obtain scores for multiple jokes based on manually labeled data.

[0154] Understandably, the process begins with determining training samples based on manually labeled data. Then, a large language model is used as the foundation, with the manually labeled data serving as the training samples. A regression head is added to the end of the large language model to convert text features into specific numerical scores. Finally, the large language model is trained using the manually labeled data, with a ranking loss function used to optimize model parameters until the model can directly output the overall quality score of jokes, achieving automated quality assessment of large-scale jokes without the need for manual annotation.

[0155] The ranking loss function is a loss function used to optimize the ranking ability of a model. Its core goal is to enable the model to learn to distinguish the order of the quality of samples. During the training of the model, the model parameters are adjusted by calculating the degree of matching of the logic that "high-quality jokes score higher than low-quality jokes". Specifically, the larger the difference between the scores of high-quality jokes and low-quality jokes predicted by the model, the lower the loss value, and vice versa. This helps the model gradually master human rating preferences for jokes.

[0156] The functional expression for the ranking loss function is:

[0157]

[0158] in, This represents the loss value calculated in a single training iteration. The lower the loss value, the more the model's ranking judgment aligns with human preferences. This represents the model's predicted score for the output of high-quality joke A; This represents the model's predicted score for the poor joke B. It is an activation function that is used to... The difference is mapped between 0 and 1, thereby quantifying the probability that "a high-quality joke scores higher than a low-quality joke"; It is a logarithmic function used to transform probability values ​​into loss values, enabling backpropagation optimization of model parameters.

[0159] In one possible implementation, when calculating scores for multiple jokes, the scores are obtained by a weighted algorithm of multiple parameters, wherein the multiple parameters include at least: humor, unexpectedness, relevance, comprehensibility, and offensiveness; the weighting coefficient for offensiveness is negative.

[0160] Understandably, humor is used to assess whether a joke is funny, with a score range of 1-10. The higher the humor score, the funnier the joke; conversely, the lower the humor score, the less funny the joke.

[0161] Unexpectedness is used to assess whether a joke is surprising, with a score range of 1-10. The higher the unexpectedness score, the more novel and surprising the plot twist of the joke; conversely, the lower the unexpectedness score, the more mundane and predictable the plot of the joke.

[0162] Relevance is used to assess whether a joke is relevant to the topic, with a score range of 1-10. The higher the relevance score, the better the joke fits the topic; conversely, the lower the relevance score, the weaker the connection between the joke and the topic, or even the more off-topic it may be.

[0163] Comprehensibility is used to assess how easy a joke is to understand, with a score range of 1-10. The higher the comprehensibility score, the more accessible the language and the clearer the logic of the joke, making it easier for the audience to understand; conversely, the lower the comprehensibility score, the more obscure the joke, and the more difficult it is for the audience to understand.

[0164] Offensiveness is used to assess whether a joke contains inappropriate content, with a score range of 1-10. The higher the offensiveness score, the greater the likelihood that the joke contains vulgar, discriminatory, or other inappropriate content; conversely, the lower the offensiveness score, the healthier and more suitable the joke is for a wider audience.

[0165] Different parameters have different weighting coefficients. For example, the weighting coefficient for humor is 0.5, the weighting coefficient for unexpectedness is 0.2, the weighting coefficient for relevance is 0.15, the weighting coefficient for comprehensibility is 0.1, and the weighting coefficient for offensiveness is -0.05.

[0166] For example, if the weighting coefficient for humor is 0.5, the weighting coefficient for unexpectedness is 0.2, the weighting coefficient for relevance is 0.15, the weighting coefficient for comprehensibility is 0.1, and the weighting coefficient for offensiveness is -0.05, and assuming that a joke has a humor score of 8, an unexpectedness score of 7, a relevance score of 8, a comprehensibility score of 8, and an offensiveness score of 2, then based on the above information, the joke can be determined to have a score of 7.3.

[0167] The third type is a reinforcement learning mechanism, which is used to maximize the expected reward based on the scores of multiple jokes.

[0168] Understandably, the implementation steps of a reinforcement learning mechanism include:

[0169] The first step is to model the task, transforming the joke generation task into a reinforcement learning problem, including:

[0170] State: Input topic keywords, contradiction information, and joke setup and punchline text generated by the model.

[0171] Action: The decision-making behavior that generates the next word or sentence in the current state.

[0172] Reward: Input the generated complete joke into the reward mechanism to get the score of this joke, which is the reward value.

[0173] Policy: Construct a generation policy based on the parameters of a large language model.

[0174] The second step is to select the optimal algorithm: the Proximal Policy Optimization (PPO) algorithm is adopted as the optimization strategy. PPO ensures that the optimized strategy does not deviate too far from the original pre-trained model by maximizing the expected reward and incorporating KL divergence constraints.

[0175] The third step is to define the optimization objective: the optimization objective is the maximum expected reward, which can be represented as follows:

[0176]

[0177] in, As a reference model, To control the coefficient of KL penalty.

[0178] Step 4: Implement the PPO update rule: Calculate the loss function. And update the strategy, which can be done in the following ways:

[0179]

[0180] in, This represents the probability ratio between the current strategy and the old strategy. The dominant function; These are the trimming parameters.

[0181] Fifth step, iterative training: By repeatedly iterating the above steps and continuously updating the strategy, the quality of the generated jokes is gradually improved, and finally a model that can generate humorous jokes that are relevant to the topic is obtained.

[0182] S304. Obtain the target podcast based on the target joke.

[0183] Understandably, by converting target jokes into target podcasts, jokes that originally existed only in text form can be listened to by users through auditory channels, thereby satisfying users' entertainment needs in scenarios such as driving.

[0184] In one possible implementation, this application provides a possible way to obtain a target podcast based on a target joke, including:

[0185] The first step is to obtain a dialogue structure outline based on the target joke.

[0186] Understandably, the outline of a dialogue includes, but is not limited to: opening remarks, introduction of the topic, in-depth discussion, clash of viewpoints, differences in creative content, summary, and closing remarks.

[0187] Because traditional jokes are usually presented in a single narrative format, listeners need to piece together the logical chain themselves, making it easy to miss the punchlines due to distraction. Dialogue structure outlines, however, break down jokes into character interactions, guiding listeners with more intuitive dialogue clues. For example, transforming a one-person joke into a two-person dialogue format, using questions, rhetorical questions, or role-playing to enhance foreshadowing and plot twists, allows users, even while driving, to quickly grasp key information and easily enjoy the entertainment.

[0188] The second step involves modeling the pre-defined personas, assigning a leading role to each stage of the dialogue structure outline, and generating the dialogue flow for each stage to obtain the dialogue for each stage. The persona modeling includes the following dimensions of various personas: basic attributes, personality traits, knowledge domain, language style, and dialogue habits.

[0189] Understandably, basic attributes include, but are not limited to: name, gender, age, occupation, and educational background.

[0190] Personality traits: extroverted, agreeable, conscientious, neurotic, open-minded.

[0191] Knowledge domain: areas of expertise and hobbies.

[0192] Language style: vocabulary preferences, sentence structure characteristics, verbal tics, and usage habits of interjections.

[0193] Conversation habits: frequency of asking questions, tendency to interrupt, and ability to guide the conversation.

[0194] First, the pre-defined character modeling clearly identifies the characteristics of different roles in the dialogue. Therefore, by using the pre-defined character modeling, a leading role can be assigned to each stage of the dialogue structure outline, determining who will advance the current segment. Subsequently, generating the dialogue flow for each stage can determine the speaking order and specific content of each role, helping to ensure that the dialogue content is logically coherent, humorous, and suitable for the auditory context of audio podcasts, reducing the cognitive burden on listeners.

[0195] The third step is to identify the target podcast based on the conversations at each stage.

[0196] Understandably, since the conversations at each stage exist only in text form, in order to meet the listening needs of users in driving scenarios, these conversations need to be converted into podcast content that can be listened to directly, so that users can easily obtain information while driving.

[0197] In one possible implementation, this application provides a possible way to obtain the target podcast based on the dialogue at each stage, including: performing persona consistency detection and correction on the dialogue at each stage to obtain the target podcast.

[0198] Understandably, the consistency detection rules include: judging one by one whether the deviation of each feature extracted from the dialogue from the corresponding preset feature benchmark value is greater than a preset threshold; if the deviation is greater than the threshold, the corresponding penalty value is deducted from the initial score of the dialogue and the feature is marked as a deviation item; until all features have been detected, the final consistency score can be obtained.

[0199] The correction rules include: for each generated dialogue, if the consistency score is less than the preset score, then the marked deviation items in the dialogue are corrected accordingly; for example, an instruction-based correction method can be used, such as marking "The following dialogue deviates from the character's personality: [deviation item]. Please modify the dialogue to better conform to the style requirements of [character's personality]". Conversely, if the consistency score of the dialogue is greater than or equal to the preset score, then the dialogue is output, indicating that the dialogue does not require correction.

[0200] Before conducting consistency testing, an initial baseline score can be set for the dialogue to be evaluated. This score represents the ideal score when the dialogue fully conforms to the preset character traits. Subsequently, feature extraction is performed on the dialogue to be evaluated, including but not limited to: measuring lexical complexity through word frequency statistics, analyzing sentence structure through average sentence length and clause usage frequency, judging the emotional tendency of the text through the proportion of positive, negative, and neutral expressions, and defining the professionalism of the text through the density of professional terms, thereby obtaining a complete set of extracted features. Next, for each extracted feature of the dialogue, its deviation from the corresponding preset feature baseline value is calculated one by one. If the deviation is greater than a preset threshold, the corresponding penalty value is deducted from the initial baseline score and the feature is marked as a deviation item. After all features have been detected, the consistency score of the dialogue can be obtained. Then, if the consistency score is lower than the preset threshold, the marked deviation items in the dialogue are specifically corrected, and finally, a corrected dialogue that conforms to the character requirements is obtained.

[0201] The podcast acquisition method provided in this application firstly involves expanding the podcast topic information based on thought chain reasoning to obtain expanded content containing at least contradictory information; then, multiple jokes are created based on the contradictory information and a preset joke generation logic; subsequently, target jokes are selected based on their scores; and finally, the target jokes are transformed into target podcasts. This method, by using thought chain reasoning to uncover thematic contradictions and combining a scoring mechanism to select high-quality jokes, not only generates humorous podcast content that fits the theme and is highly entertaining, but also ensures the stability of content quality, effectively improving the uniqueness and adaptability of podcast content and meeting users' needs for personalized, high-quality in-car humor podcasts.

[0202] Optionally, based on the above embodiments, this application provides a detailed description of the training method for the target model, including:

[0203] The first step is sample construction and preprocessing of the constructed samples, including:

[0204] First, by calling multiple application programming interfaces or using targeted web scraping technology, multi-source data containing podcast topic information, podcast style, and content relationships are crawled in real time from multiple publicly available platforms.

[0205] Subsequently, leveraging the natural language understanding capabilities of the large language model, entity recognition (person names, place names, organizations), event extraction (time, location, people, causes, results), and opinion extraction (stance, sentiment) are performed on multi-source data, transforming unstructured text data into structured text data. Next, the Sentence-BERT model is used to vectorize the text. By calculating cosine similarity and combining it with the DBSCAN clustering algorithm, text content with a similarity higher than 0.9 is grouped into one category. For each text cluster, the large language model is then used to perform multi-document summarization, generating a structured summary containing core events, multiple viewpoints, and a timeline.

[0206] Based on the basic samples, it is necessary to expand the samples through the ability of reasoning through thought chains: extract multiple core concepts around the theme information of the podcast, match multiple daily life scenarios for each concept, explore the contradictions or counterintuitive points in the scenarios that can reflect the style of the podcast, and then generate extended samples that conform to the style of the podcast in batches according to the preset joke generation logic.

[0207] The second step involves manual annotation and the construction of a preference dataset, including:

[0208] The generated samples were manually labeled, and a preference dataset was constructed. Assuming the podcast style is jokes, the labeling dimensions included humor, unexpectedness, relevance, comprehensibility, and offensiveness, each scored on a scale of 1-10 (where lower offensiveness scores indicate safer content). After labeling, the average score for each dimension of each joke was calculated. Then, a weighted formula was designed to calculate the overall score, taking into account the needs of the in-car podcast scenario. Humor had the highest weight, while offensiveness was given a negative weight to ensure content safety. Subsequently, multiple jokes on the same topic were sorted according to their overall scores, constructing preference pairs of "high-scoring jokes - low-scoring jokes," forming the preference dataset.

[0209] The third step involves training a reward model based on the preference dataset, including:

[0210] The goal of the reward model is to learn human preferences for podcast content and output a corresponding comprehensive score based on the input joke text.

[0211] This model is based on a pre-trained large language model, with a regression head added to the end of the model to map the input joke text to a comprehensive score prediction. The training process uses the preference dataset as supervised data and employs ranking loss as the loss function, as shown in the formula: Where R(joke A) represents the model's predicted score for the output of high-quality joke A; R(joke B) represents the model's predicted score for the output of low-quality joke B. The training objective is to enable the model to learn the preference relationship that "joke A scores higher than joke B", and finally obtain a reward model that can accurately evaluate the quality of podcast content.

[0212] IV. Generation Strategy Optimization Based on Proximity Policy Optimization (PPO)

[0213] Based on the reward model, a reinforcement learning framework is used to optimize the generation strategy of the large language model, achieving the goal of outputting high-scoring podcast content from the input podcast topic. First, the joke generation task is modeled as a reinforcement learning problem: the definition and state are the topic keywords, contradictions, and the already generated expected and transitional text; the action is to generate the next word or sentence; the reward is calculated by the trained reward model, such as R(joke A) representing the model's predicted score for the output of a high-quality joke A; R(joke B) representing the model's predicted score for the output of a low-quality joke B; the policy is the generation strategy of the large language model. The optimization algorithm uses Proximal Policy Optimization (PPO), whose optimization objective is to maximize the expected reward, while incorporating KL divergence constraints to prevent the optimized model from deviating too far from the original pre-trained model. The formula is as follows: For the original pre-trained model, This is the KL penalty coefficient (typically ranging from 0.01 to 0.1). The PPO update rule is as follows: ),

[0214] This rule uses probability ratios and dominance function Implement and introduce clipping parameters. (For example, 0.2) Limit the magnitude of policy updates to obtain a target model that can stably generate high-quality, highly adaptable podcast content.

[0215] Figure 4 This is a schematic diagram of the structure of the podcast acquisition device provided in the embodiments of this application, such as... Figure 4 As shown, this application embodiment provides a podcast acquisition device, the device 400 including:

[0216] The receiving module 401 is used to receive a podcast playback request, which includes the topic information of the podcast to be played and information indicating the style of the podcast.

[0217] The acquisition module 402 is used to acquire the target podcast based on the podcast playback request. The target podcast is generated based on the target model, which is a model that matches the podcast style. The input of the target model includes topic information and the content of the target podcast that is adapted to the podcast style.

[0218] In one possible implementation, the device further includes: a divergence module 403;

[0219] The divergence module 403 is used to diverge the topic information based on the reasoning ability of the thought chain to obtain divergent information, which includes at least the information of contradictions.

[0220] The device also includes: a generation module 404;

[0221] The generation module 404 is used to generate multiple jokes based on the contradiction information and the preset joke generation logic;

[0222] The device also includes: a determination module 405;

[0223] Module 405 is used to determine the target joke based on the scores of multiple jokes;

[0224] Module 402 is used to obtain the target podcast based on the target joke.

[0225] In one possible implementation, the divergence module 403 is specifically used to diverge the topic information at multiple levels based on the reasoning ability of the thought chain to obtain divergent information.

[0226] Among them, multi-level divergence includes first-level divergence, second-level divergence, and third-level divergence;

[0227] The first layer of divergence is used to obtain core conceptual information related to the topic;

[0228] The second layer of divergence is used to derive specific scenario information from each core concept information;

[0229] The third layer of divergence is used to identify contradictory information in specific scenarios.

[0230] In one possible implementation, the pre-defined joke generation logic includes logic for establishing expectations, breaking expectations, and interpreting them.

[0231] In one possible implementation, the target model includes at least one of the following mechanisms:

[0232] Vector similarity mechanism, used to deduplicate multiple jokes based on vector similarity;

[0233] A reward mechanism is used to obtain individual scores for multiple jokes based on manually labeled data;

[0234] A reinforcement learning mechanism is used to maximize the expected reward based on the scores of multiple jokes.

[0235] In one possible implementation, when calculating scores for multiple jokes, a weighted algorithm is used to obtain scores based on multiple parameters, including at least: humor, unexpectedness, relevance, comprehensibility, and offensiveness; the weighting coefficient for offensiveness is negative.

[0236] In one possible implementation, the determining module 405 is further configured to obtain a dialogue structure outline based on the target joke;

[0237] The generation module 404 is also used for character modeling based on preset personas, assigning leading roles to each stage in the dialogue structure outline, and generating the dialogue flow for each stage to obtain the dialogue for each stage; the persona modeling includes the following dimensions of various personas: basic attributes, personality traits, knowledge domain, language style and dialogue habits.

[0238] The acquisition module 402 is specifically used to obtain the target podcast based on the dialogue at each stage.

[0239] In one possible implementation, the acquisition module is specifically used to perform persona consistency detection and correction on the dialogue at each stage to obtain the target podcast.

[0240] In one possible implementation, the device further includes: a playback module 406;

[0241] Playback module 406 is used to play the target podcast;

[0242] Module 402 is also used to obtain user feedback information regarding the target podcast;

[0243] The device also includes: an indicator module 407;

[0244] Instruction module 407 is used to instruct the processing of the target model based on feedback information.

[0245] In one possible implementation, the device further includes: a transmitting module 408;

[0246] Sending module 408 is used to send podcast playback requests to cloud devices;

[0247] The receiving module 401 is specifically used to receive target podcasts from cloud devices.

[0248] The podcast acquisition device provided in this application embodiment can be used to execute the podcast acquisition method in any of the above embodiments of this application. Its implementation principle and technical effect are similar, and will not be described again in this embodiment.

[0249] Figure 5 A schematic diagram of the structure of the electronic device provided in this application. Figure 5 As shown, the electronic device 500 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the device 50 also includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus 504.

[0250] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above-described method.

[0251] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0252] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0253] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0254] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0255] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0256] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0257] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0258] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0259] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0260] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0261] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0262] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0263] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0264] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A method for acquiring podcasts, characterized in that, The method includes: Receive a podcast playback request, the podcast playback request including the topic information of the podcast to be played and information indicating the style of the podcast; Based on the podcast playback request, a target podcast is obtained. The target podcast is generated based on a target model, which is a model that matches the style of the podcast. The input of the target model includes the topic information, and the content of the target podcast is adapted to the style of the podcast.

2. The method according to claim 1, characterized in that, The podcast style includes a joke style, and the target model is specifically used for: Based on the reasoning ability of the thought chain, the topic information is diverged to obtain divergent information, which includes at least contradictory information; Based on the contradiction information and the preset joke generation logic, multiple jokes are generated; Based on the scores of the multiple jokes, the target joke is determined; The target podcast is obtained based on the target joke.

3. The method according to claim 2, characterized in that, The process of diverging the topic information based on the ability of thought chain reasoning to obtain divergent information includes: Based on the reasoning ability of the thought chain, the topic information is diverged at multiple levels to obtain the divergent information; The multi-level divergence includes a first-level divergence, a second-level divergence, and a third-level divergence; The first layer of divergence is used to obtain core concept information related to the topic information; The second layer of divergence is used to derive specific scenario information from each of the core concept information; The third layer of divergence is used to identify the contradiction information in each of the specific scene information.

4. The method according to claim 3, characterized in that, The preset joke generation logic includes: logic for establishing expectations, breaking expectations, and explanation.

5. The method according to any one of claims 2-4, characterized in that, The target model includes at least one of the following mechanisms: A vector similarity mechanism is used to deduplicate the multiple jokes based on vector similarity; A reward mechanism is used to obtain individual scores for each of the multiple jokes based on manually labeled data; A reinforcement learning mechanism is used to maximize the expected reward based on the scores of the multiple jokes.

6. The method according to claim 5, characterized in that, When calculating the scores for the multiple jokes, a weighted algorithm is used to obtain the scores based on multiple parameters, which include at least: humor, unexpectedness, relevance, comprehensibility, and offensiveness; the weighting coefficient for offensiveness is negative.

7. The method according to any one of claims 2-4, characterized in that, The process of obtaining the target podcast based on the target joke includes: A dialogue structure outline is obtained based on the target joke; Based on the pre-defined character modeling, a leading role is assigned to each stage in the dialogue structure outline, and the dialogue flow for each stage is generated to obtain the dialogue for each stage; the character modeling includes the following dimensions of various characters: basic attributes, personality traits, knowledge domain, language style and dialogue habits. The target podcast is obtained based on the dialogue at each stage.

8. The method according to claim 7, characterized in that, The process of obtaining the target podcast based on the dialogue at each stage includes: The dialogue at each stage is checked for consistency with the persona and corrected to obtain the target podcast.

9. The method according to any one of claims 1-4, characterized in that, The method further includes: Play the target podcast and obtain user feedback on the target podcast; Based on the feedback information, instructions are given to process the target model.

10. The method according to any one of claims 1-4, characterized in that, The step of obtaining the target podcast based on the podcast playback request includes: Send the podcast playback request to the cloud device; Receive the target podcast from the cloud device.

11. A podcast acquisition device, characterized in that, The device includes: The receiving module is used to receive podcast playback requests, which include the topic information of the podcast to be played and information indicating the style of the podcast; The acquisition module is used to acquire a target podcast based on the podcast playback request. The target podcast is generated based on a target model, which is a model that matches the podcast style. The input of the target model includes the topic information, and the content of the target podcast is adapted to the podcast style.

12. An electronic device, characterized in that, include: Memory and processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-10.

14. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.