Human-AI Interactive Hybrid Audio Broadcast Production System and Method
Patent Information
- Application Number
- TR202611181
- Authority / Receiving Office
- TR · TR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-21
Smart Images

Figure 00000007_0000
Abstract
Description
1 TARIFF Human-AI Interactive Hybrid Audio Broadcast Production System and Method TECHNICAL AREA 5 The invention is based on radio broadcasting, audio content production, and human-machine interaction media. It is related to the production system. The invention specifically involves a content stream, guided by a real human server, over a certain 10-minute period. in a dialogue structure with a synthetic host who has a persona and emotional tone the structuring, and the arrangement of the dialogues according to an interactive scenario structure, The main server and synthetic server input data are processed taking into account natural speech gaps. human and artificial intelligence that enables the merging of content and the production of the same content structure in different languages. It is related to the intelligence-interactive hybrid audio broadcast production system and method. 15 PREVIOUS TECHNIQUE In traditional dual-server publishing models today, content production takes place on two different sides. operational 20 because it depends on the server's simultaneous or coordinated participation. Processes can become more costly and complex. In such models, the servers must be the same. preparing it in accordance with the broadcast schedule, ensuring it's scheduling right, recording or establishing physical or technical coordination during a live broadcast and the broadcast language Factors such as dominance can limit the content production process, especially when working in different languages. When content needs to be prepared in different accents or different broadcast formats, human 25 A server-based infrastructure adds a burden in terms of both time and resource usage. It is able to create. On the other hand, existing AI-powered voiceover applications a linear text, usually a pre-prepared script, for vocal performance It remains limited to the reading process; these applications are context-sensitive with a human server. 30 It is not structured. Therefore, current solutions use a human server versus an AI-based system. interactive, natural, and contextually relevant within the virtual server's streaming environment. It falls short in terms of enabling adaptable communication. As a result of research conducted in the literature, the application number “2021 / 019768” and “Artificial Intelligence 35 A Turkish patent application titled "Smart Vehicle Assistant" has been found. The subject of the application is a device that can recognize the user's face and emotional state, and use natural conversational language. 2 Able to understand given commands, answer questions, read news, and understand weather conditions. their ability to provide information, translate into different languages, and communicate various information orally. It relates to an AI-powered smart assistant system. However, in the aforementioned application, In the broadcasting field, working alongside a human presenter, content is delivered depending on the broadcast schedule. by evaluating the context and establishing an asynchronous but dynamic dialogue with a human presenter, the pair 5 an indication of an AI-based virtual server system that performs server-side broadcasting None have been encountered. Ultimately, the problems mentioned above, which cannot be solved with current technology, are the subject of this technical analysis. This has made it necessary to make an innovation in the field. 10 A BRIEF DESCRIPTION OF THE INVENTION The present invention aims to eliminate the aforementioned disadvantages and introduce new technologies to the relevant technical field. a hybrid audio broadcast production system with human and artificial intelligence interaction to bring advantages and It is related to the method. 15 The main purpose of the invention is to enable a main server to perform an artificial attack within a predetermined scenario. a podcast or radio show that engages in a dialogue with a synthetic host generated by intelligence The goal is to develop a system and method for producing program content. Another purpose of the invention is to regulate the timing of dialogues between the main server and the synthetic server. An interactive scenario structure that allows for configuration based on intonation and response flow. to create. The other purpose of the invention is to record the main server and synthetic server audio at different times. by enabling the merging of data while respecting natural speech gaps, a dual-server system The goal is to ensure the technical creation of the broadcast stream. Another aim of the invention is to enable the synthetic presenter to embody a specific persona, voice character, and emotion. by enabling voice acting with intonation, the AI-based 30-person system combines a human presenter with a human-centered speaker. The aim is to create a natural dialogue effect between the synthetic server and the server. Another purpose of the invention is to synchronize the same content structure in English, German, or other languages. by enabling the production of podcast or radio program content in multiple languages The aim is to enable the creation of versions. 35 3 All the purposes mentioned above and those that will emerge from the detailed explanation below. The current invention aims to achieve interactive human and artificial intelligence sound broadcast production. It is a system, and its feature is; A voice recording unit that enables the acquisition of real voice data from the human presenter. Synthetic character voice data with a specific persona and emotional tone, artificial 5 A sound production unit that enables the production of sounds with the aid of intelligence. Dialogue between human voice data and synthetic character voice data structured text including sequence, timing information, intonation information, and response flow interactive scenario unit that makes up the series, The actual audio data obtained with the said audio recording unit and the said audio production 10 Synthetic character voice data produced by the unit, by the interactive scenario unit in accordance with the determined dialogue sequence and observing natural pauses in speech a hybrid audio mixing unit that combines and To enable the same content structure to be produced synchronously in different languages. Multilingual phonetic language adaptation unit 15 that performs translation and speech synthesis processes It includes. The best way to utilize the advantages of the existing invention, together with its structure and additional elements. For it to be understood, it must be considered together with the figures explained below. BRIEF DESCRIPTION OF THE FIGURES Figure 1 shows the hybrid audio broadcast production system with human and artificial intelligence interaction that is the subject of the invention. It is a representative illustration. The drawings do not necessarily need to be scaled and are necessary for understanding the invention. Details that are not present may have been overlooked. Furthermore, at least to a large extent... Elements that are identical or at least have substantially identical functions are numbered the same. It is shown. REFERENCE NUMBERS 1. Sound recording unit 2. Sound production unit 3. Interactive scenario unit 35 4. Hybrid audio mixing unit 5. Multilingual phonetic language adaptation unit 4 DETAILED DESCRIPTION OF THE INVENTION This detailed explanation describes the invention, a hybrid audio broadcasting system that combines human and artificial intelligence interaction. The production system and method are solely for the purpose of better understanding the subject, with no limitations. 5 This is explained with examples that will not have an impact. The voice recording unit (1) in the system controls the content with real human voice and artificial intelligence. It is the unit from which the voice data of the main server that “passes” to the character is received. Voice production unit (2), For example, Erinome is characterized by her tone of voice, a specific persona, and an emotional intonation of 10 The interactive scenario unit (3) is the main unit that produces the output of the AI voice engine. Timing, intonation, and response in dialogues between the host and the synthetic character. It is a structured text sequence containing the keys. The hybrid audio mixing unit (4) is used at different times. The recorded master server and synthetic character voice data, natural speech gaps It is the technical layer in which the multilingual phonetic language adaptation unit (5) is combined by taking into consideration the same content 15 enabling the structure to be produced synchronously in different languages (English, German, etc.) It is a translation and voice synthesis unit. The system processes a server's predefined scenario. including, with a synthetic character generated by artificial intelligence (e.g., Google AI Studio) (within its Erinome voice) enables human and artificial intelligence to engage in dialogue. It relates to interactive podcast / radio program production systems. 20 In the invention: From the dialogues between the main host of the radio program / podcast and the synthetic character The text skeleton created by the interactive scenario unit (3), programme producer It is written by [author's name]. Where research is required in the written text, 25 is indicated in square brackets. By writing the necessary commands, the artificial intelligence is asked to locate the square brackets. The student is asked to complete the task according to the instructions inside the square brackets. Once the text is complete, the sections of the text that the synthetic character will voice, For example, voice production unit using Google AI Studio's text-to-speech function (2) It is voiced via and the relevant audio file(s) are recorded. 30 The main server's audio recordings are taken in a studio environment via the audio recording unit (1). Main server and synthetic character records are compatible with the interactive scenario unit (3). In this way, a radio is formed by combining the hybrid sound mixing unit (4) in succession. The program recording / podcast is created. After the process described above is completed, the same process is applied to multilingual phonetic language 35 through the adaptation unit (5) the radio programme / podcast in other languages It is repeatable for the production of its versions. The invention is incorporated into various radio programs. It is available for use. The invention works as follows: First, the content stream related to the podcast or radio program, Dialogue between a human presenter and a synthetic voice character generated by artificial intelligence 5 It is created through the interactive scenario unit (3) to include the sequence. In the script, there will be segments voiced by a human presenter and segments voiced by a synthetic voice character. The sections to be voiced are determined. The sections belonging to the synthetic voice character are handled by voice production. While the parts belonging to the human presenter are voiced with the help of artificial intelligence through the unit (2), The human voice data obtained is recorded via the voice recording unit (1). Synthetic voice 10 data, natural speech gaps, pauses and by hybrid voice mixing unit (4) The dialogues are combined in a sequence that respects the order of the dialogue. This allows a human presenter to interact with an AI-based synthetic server. A podcast or radio program recording that gives the impression of a conversation between voice characters. is created. If necessary, the same content structure, multilingual phonetic language adaptation unit (5) through which podcast or radio program versions are adapted into different languages. 15 It can be produced.
Claims
6 REQUESTS 1. It is a human-artificial intelligence interactive voice broadcast production system, and its feature is; 5 voice recording unit (1) that enables the acquisition of real voice data of the human presenter. synthetic character voice data with a specific persona and emotional tone voice production unit (2) that enables the production of sound with artificial intelligence support. Dialogue between human voice data and synthetic character voice data structural 10 including sequence, timing information, intonation information and response flow interactive scenario unit (3) which forms the text sequence The actual sound data obtained with the said sound recording unit (1) and the said sound Synthetic character voice data produced by production unit (2), interactive scenario unit in accordance with the sequence of dialogue determined by (3) and natural speech hybrid sound mixing unit (4) and 15 which combine by taking into account the gaps To enable the same content structure to be produced synchronously in different languages. Multilingual phonetic language adaptation performs translation and speech synthesis processes. unit (5) It includes.
2. It is a method of producing voice broadcasts that interacts with humans and artificial intelligence, and its characteristic feature is; Human host and synthetic character related to podcast or radio program the sequence of dialogue, timing information, tone information, and response flow between them The interactive scenario unit (3) is created by the structured text sequence containing, Receiving the actual voice data of the human presenter through the voice recording unit (1), 25 Synthetic character voice data with a specific persona and emotional tone Production with artificial intelligence support through the production unit (2), hybrid sound is created by combining real voice data with synthetic character voice data. determined by the interactive scenario unit (3) through the merging unit (4) in accordance with the dialogue sequence and observing natural pauses in speech 30 combining, Creating a podcast or radio program recording and replicating the same content structure synchronized in different languages through the phonetic language adaptation unit (5) production It includes the steps of the process. 35