Character-based licensed AI-powered audio conversion and broadcasting system for VOD / OTT content.

TR202612121A2Pending Publication Date: 2026-09-21TURKCELL TEKNOLOJI ARASTIRMA & GELISTIRME AS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TR202612121
Authority / Receiving Office
TR · TR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-21

Smart Images

  • Figure 00000016_0000
    Figure 00000016_0000
Patent Text Reader

Abstract

The invention relates to an AI-powered audio conversion and publishing system that reproduces character voices from films, series, animations, documentaries, educational content, or similar audiovisual VOD / OTT content on a character-by-character basis using licensed voice actor identities. This system is time-coded, emotionally and prosodiously compatible, lip-synced, quality-controlled, and watermarked / origin information marked. It is used in the fields of digital media processing, AI-based voice synthesis, voice conversion, speaker separation, licensed voice ID management, VOD / OTT content packaging, and digital copyright management.
Need to check novelty before this filing date? Find Prior Art

Description

- 1 - TARIFF Character-based licensed artificial intelligence in VOD / OTT content. AUDIO CONVERSION AND BROADCASTING SYSTEM WITH ASSISTED RESISTANCE TECHNICAL FIELD The invention involves digital media processing, AI-based voice synthesis, and audio. conversion, speaker parsing, licensed audio ID management, VOD / OTT content In the field of packaging and digital copyright management, for films, series, animations, documentaries, 10 characters in educational content or similar audiovisual VOD / OTT content Their voices are created using licensed voice actor voice identities and character-based voice acting. reproducing, time-code protected, emotion and prosody compatible, lip Synchronized, quality-controlled, and AI-powered with watermark / origin information marking. It relates to an audio conversion and broadcasting system. The system that is the subject of the invention is VOD / OTT 15. digital streaming platforms, online video services, AI-powered dubbing, and localization infrastructures, alternative voice channel management, and licensed voice model. It can be used in management platforms. PREVIOUS TECHNIQUE 20 On current VOD / OTT platforms, users are usually only able to select a language. Alternatively, users can choose their preferred voiceover from standard dubbing options. If a different voice actor's voice is requested for the character, then the classic approach is used. These methods necessitate studio recording; this process is time-consuming and costly. 25 and scaling is difficult. General purpose AI voice cloning. The solutions include license control, character-based time code management, and lip-syncing. synchronization, emotion / prosody preservation, quality control, synthetic voice verification, and OTT. It does not meet packaging requirements in an integrated manner. Document number US10930263B1, 30 for media content localization. It describes an automated voice dubbing system. This solution is neural network-based for voice dubbing. The synthesis includes the use of speaker embedding and prosody / emotion preservation. - 2 - It involves media processing similar to VOD / OTT. However, it is character-based speech. parsing, license-controlled voice ID vault, watermark / origin information generation, time Code-protected lip-sync adaptation and OTT alternate audio channel Packaging elements are missing. Document number WO2020242662A1, multilingual speech synthesis and 5 It describes cross-lingual voice cloning. Speaker embedding, prosody. It includes approaches to preservation and the synthesis of multilingual writing and speech. However VOD / OTT content analysis, character speech timeline database, licensing lack of control engine, segment-based conversion and broadcast packaging. It is located. 10 Document WO2019209569A1, end-to-end speaker parsing. It addresses the model. However, licensed voice identity management deals with emotion / prosody. It does not include protected mixing and OTT packaging elements. Current technical documentation offers fragmented teachings; license-controlled audio Identity vault, character-based speech timeline, emotion / prosody protected 15 Segment-based conversion, time code and lip-sync adaptation, quality. Controlled automated reproduction, with watermark / origin information integration. End-to-end integrated OTT / VOD compatible alternative audio channel packaging processes. It does not combine them in an architecture. Therefore, existing solutions cost studio space. reduction, scalable personalization, minimization of legal risk and publication 20 It cannot provide technical features such as quality assurance in a single system. A BRIEF DESCRIPTION OF THE INVENTION Digital media processing, AI-based speech synthesis, voice 25 conversion, speaker parsing, license management, and VOD / OTT packaging. licensed voice actors for character voices in VOD / OTT content in the field. Voice identifiers with character-based, time-code protected, emotion and prosody matching, Reproduced in lip-sync and with watermark / origin information marking. An AI-powered audio conversion and broadcasting system has been developed. 30 The developed system includes content input, media parsing, character parsing, and audio processing. preprocessing, license verification, audio conversion, emotion / prosody preservation, time - 3 - Code / lip sync adaptation, remixing, quality control, By working through the watermark / origin information generation and OTT / VOD packaging steps. It offers the user character-based licensed voice options and VOD / OTT streaming. It produces alternative audio channels that can be integrated into the blockchain. The developed system includes a license-controlled voice ID database and character-based speech processing. Timeline database, emotion / prosody preservation module, time code protected. Unique features include a lip-sync module and watermark / origin information integration. It differs from classic dubbing and general voice cloning systems with its components; Reduces the need for studio recording, provides a personalized user experience. It provides and technically protects copyrights. 10 DESCRIPTION OF THE FIGURES Figure 1. Character-Based Licensed AI in VOD / OTT Content. Supported Audio Conversion and Broadcasting System Architecture 15 The corresponding part numbers shown in the figures are given below. 1. VOD / OTT Content Source 2. Media Parsing Module 20 3. Character and Speaker Separation Module 4. Character Speech Timeline Database 5. Licensed Audio Data Collection Module 6. Audio Preprocessing and Feature Extraction Module 7. License-Controlled Voice ID Safe 25 8. Voice Cloning / Voice Conversion Model 9. Emotion and Prosody Preservation Module 10. Time Code and Lip Synchronization Coordination Module 11. Music / Effects / Ambient Sound Preservation and Remixing Module 12. Quality Control and Automated Remanufacturing Module 30 13. Synthetic Audio Watermark and Origin Information Module 14. OTT / VOD Packaging and Manifest Generation Module - 4 - 15. User Preferences and Playback Interface 16. Usage Analytics and Revenue Sharing Module DETAILED DESCRIPTION OF THE INVENTION Digital media processing, AI-based speech synthesis, voice conversion, speaker parsing, license management, and VOD / OTT packaging. licensed voice actors for character voices in VOD / OTT content in the field. Voice identifiers with character-based, time-code protected, emotion and prosody matching, EN 10 reproduces lip-synced and watermarked / origin information marked. a small processor, memory, data storage units, graphics processing unit (GPU), and network on one or more computer / server devices containing the interface The invention, developed for use in character-based VOD / OTT content, Licensed AI-powered audio conversion and broadcasting system; VOD / OTT The content input module, which receives data from the content source (1), processes media streams 15 media parsing module (2) which separates and makes it processable, character-based Character and speaker parsing module that detects speech segments (3), The character that creates and stores the dialogue timeline for each character. Speech timeline database (4), licensed collecting audio data. audio data collection module (5), cleans raw audio recordings and extracts features audio 20 Preprocessing and feature extraction module (6), audio IDs together with the license terms License-controlled voice ID vault (7) that manages the source character speech. voice cloning / voice conversion model that reproduces with target voice ID (8), The emotion and prosody preservation module (9), which preserves the emotion and prosody structure, time Time code and lip sync synchronization synchronization 25 module (10) preserves and reassembles the original music / effects / ambient sounds Music / effects / ambient sound preservation and remixing module (11), improves sound quality. Quality control and automated remanufacturing module that monitors and automatically improves. (12), synthetic voice watermark and origin information which adds verification trace to synthetic voice module (13), alternative audio channel creating OTT / VOD packaging and manifest 30 creation module (14), user preference and playback that manage user selection - 5 - Usage analytics and revenue sharing that records usage data with the interface (15) It includes module (16). The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. VOD / OTT content included in the supported audio conversion and broadcasting system source (1); audiovisual content such as film, series, animation, documentary or educational content 5 media files, video stream, original audio stream, subtitle file, time code, It represents the data source containing scene transition information and metadata. This resource includes accessible network storage or processor and memory resources. VOD / OTT content is provided to the system via file systems. from its source (1) visual-10 such as film, series, animation, documentary or educational content audio media files, video stream, original audio stream, subtitle file, time The code, scene transition information, and metadata information are entered via the content input module. It is being integrated into the system. The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. Media parsing included in the supported audio conversion and broadcasting system 15 module (2); video, audio, subtitles, dubbing, time code and VOD / OTT content This module separates metadata streams using a multiplexing solution process. Speech-based channels, music / effects channels, and time codes can be processed. It converts the data into a standard input for subsequent stages. The invention concerns character-based licensed artificial intelligence in VOD / OTT content. characters included in the supported audio conversion and broadcasting system and Speaker parsing module (3); which character the speeches in the content belong to This module determines the speaker parsing function. speech recognition, subtitle-timecode alignment, face / image matching, and lip tracking. Character-based speech segments using motion analysis methods 25 It is created. If metadata is available, it is used directly. The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. character speech included in the supported audio conversion and broadcasting system timeline database (4); speech start and end for each character time, speech text, phoneme sequence, emotion tag, speech rate, baseline frequency 30 contour, accent points, breath intervals, volume, stage acoustics, and environment It is a data structure that stores information. This database contains different voiceovers for the same content. - 6 - This eliminates the need for re-analysis once the artists are selected, and It reduces calculation costs. The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. licensed audio data included in the supported audio conversion and broadcasting system Collection module (5); explicit consent and license agreement from voice actors 5 high-quality audio recordings (WAV, FLAC or Broadcast) obtained within the scope WAV formats (with a sampling rate of 48 kHz or higher) included in the system It is a module. The recordings do not consist of plain reading sentences, but rather reflect different emotional states. speech speeds, question-answer intonations, shouting, whispering, changes in emphasis, It includes phonetic variations and characteristic speech patterns. 10 The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. audio preprocessing included in the supported audio conversion and broadcasting system. Feature extraction module (6); noise reduction, silence of raw audio recordings cleaning, removal of clicking / popping noises, sound level normalization, 15 that undergo channel balancing and the elimination of low-quality segments. and mel-spectrogram, MFCC, speaker embedding (speaker vector), basic This module extracts frequency, energy contour, phoneme alignment, and prosody features. module, Voice Activity Detection, forced alignment, automatic It uses techniques such as speech recognition and pitch tracking. The invention concerns character-based licensed artificial intelligence in VOD / OTT content. License-controlled audio included in the supported audio conversion and broadcasting system. identity vault (7); voice identity vectors, model adaptation parameters, License terms (voice actor's name, types of content it can be used in, permission) Character types provided, language, country / region, usage period, platform scope, subscription type, commercial use model, revenue sharing ratio, license start / end date, 25 cryptographic model identifier), usage permissions and watermark / origin information It is a secure data vault that stores data together. The audio model is only available from authorized manufacturers. It is run on their servers and after license verification. The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. Audio cloning / audio 30 included in the supported audio conversion and broadcasting system conversion model (8); source character speech target voiceover It is an artificial intelligence model that recreates the artist's voice identity. The model is VALL-E. - 7 - similar neural codec language model, YourTTS-like zero-shot or low-shot writing. speech synthesis, convolutional neural network-based acoustic feature extraction, LSTM / Transformer-based time series modeling, speaker encoder, prosody encoder, acoustic decoder and neural sound decoder components It may include speech-to-speech conversion and / or 5. Transcript-to-speech synthesis modes are supported; mode selection is available. based on source audio quality, background noise, subtitle accuracy, and quality score. It is done automatically. The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. Emotion and prosody included in the supported audio conversion and broadcasting system 10 Protection module (9); the emotional state of the source character on stage, speech speed, emphasis structure, intonation curve, breath intervals, and fundamental frequency by analyzing its changes, it resynthesizes it in a way that is compatible with the target voice identity. This module uses artificial intonation that exceeds the target artist's natural vocal range. It also includes voice identity boundary protection functionality that prevents unauthorized access. 15 The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. time code and lip syncing included in the supported audio conversion and broadcasting system synchronization module (10); the generated sound of the original speech segment ensuring the start and end times are consistent and that the mouth in the video content is aligned. movement, lip closure moment and plosive phoneme points analysis with 20 It is the module that performs the synchronization. The generated sound segment is longer than the original. or if short, phoneme durations, speech rate and silence intervals are readjusted. It is being organized. The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. Music / effects / media included in the supported audio conversion and broadcasting system 25 Sound preservation and remixing module (11); music, effects and in the original content a speech program that changes only the voice of a selected character without distorting ambient sounds. separating its component from other sound components and the target sound from the original stage reverb, reconfigured based on channel position, volume level, dynamic range, and EQ parameters. It is the mixing module. 30 The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. Quality control involved in the supported audio conversion and broadcasting system. - 8 - Automatic reproduction module (12); likeness to the target voice actor score, speech intelligibility, retention rate of source speech content, lip synchronization deviation, sound naturalness, sound intensity level, true peak level, interrupt, noise ratio, channel balance, segment transition continuity, and broadcast quality. segments that perform technical measurements such as the score and fall below the quality threshold are categorized as 5. This module automatically redirects to reproduction or manual inspection. Sound intensity level and true peak level measurements are based on ITU-R BS.1770. It is calculated using this approach. The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. Synthetic audio watermark 10 included in the supported audio conversion and broadcasting system. and origin information module (13); model ID for each synthetic voice segment produced, license ID, content ID, character ID, voice actor ID, production C2PA compliant, including time, regional usage rights and quality control results. This module adds origin information. This module is used for synthetic voice verification, copyright tracking, It is used for license verification and unauthorized use detection. 15 The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. OTT / VOD packaging included in the supported audio conversion and broadcasting system. and manifest generation module (14); generated sound HLS, MPEG-DASH or IMF packaging as an alternative voice channel in a distribution structure based on manifest updating, performing segment-based packaging, and selecting 20 players in the player. This module enables on-demand audio streaming with pre-generated audio channels. Segment-based manufacturing considers popularity, network latency, device capacity, and They can make a choice based on the licensing cost. The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. User preferences and 25 included in the supported audio conversion and broadcasting system Playback interface (15); the user’s content, character and voice actor allows it to make a choice and after license verification, provides an alternative audio channel. It is the interface that provides the service. The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. Usage analytics and 30 included in the supported audio conversion and broadcasting system. Revenue sharing module (16); which voice actor in which content, which For the character, in which region, for how long, and by how many users? - 9 - This module records when a listening session is recorded. This module also handles license management, copyright calculation, and It generates data for revenue sharing. The invention focuses on character-based licensed artificial intelligence in VOD / OTT content. In the supported audio conversion and broadcasting system, VOD / OTT content is integrated into the system. video, audio, subtitles and metadata are obtained with the media parsing module (2) 5 The streams are being parsed, and the speech is parsed by the character and speaker parsing module (3). Segments are identified by time codes and character speech time The schedule database (4) is created. Licensed audio data is used for audio preprocessing and Prepared with feature extraction module (6), license controlled voice ID vault (7) is kept together with the license terms. User or platform 10 When character-artist matching is selected by the license check engine voice cloning / voice conversion model (8) comes into play, voice to voice Production in speech-to-speech or text-to-speech mode. being implemented, the source scene with the emotion and prosody preservation module (9) The emotional and prosodic structure is preserved, time code and lip synchronization are consistent. 15 Full compatibility with video image is ensured with module (10), music / effects / ambient original sound integrity with sound preservation and remixing module (11) kept, quality control and automatic re-production module (12) sound intensity level, true peak level, cutoff, lip sync deviation and Broadcast quality thresholds such as similarity score are monitored, and 20 that fall below the threshold are identified. Segments can be generated using different model parameters, vocoder outputs, or prosody transfer. Automatically reproduced with settings, synthetic audio watermark and origin information C2PA compliant verification trace is added with module (13) and OTT / VOD packaging and manifest generation module (14) compatible with HLS, MPEG-DASH or IMF. It is broadcast as an alternative audio channel. This end-to-end architecture, processor, 25 licenses on memory, graphics processing unit (GPU), and data storage units validation, character-based segmentation, emotion / prosody transfer, time code technical layers such as protection, quality assurance and origin information integration by combining classic studio dubbing and general-purpose voice cloning approaches It is differentiating itself; scalable, customizable, copyright compliant, and publication 30 It provides a synthetic voiceover infrastructure that can be integrated into the chain. The system, Eliminates the need for studio recording from scratch for each new sound alternative. - 10 - removing, efficiently using multiple character-artist combinations for the same content. to produce in this way and to use synthetic voices within the scope of authorized production. It technically guarantees it. 10 20 30

Claims

- 11 - SYSTEMS 1. Digital media processing, AI-based voice synthesis, voice conversion, In the areas of speaker parsing, license management, and VOD / OTT packaging The character voices in VOD / OTT content are provided by licensed voice actors. Character-based identities, time-code protected, emotion and prosody compatible, the best way to reproduce lip-synced and watermarked / origin information marked a small processor, memory, data storage units, graphics processing unit (GPU), and network on one or more computer / server devices containing the interface The invention, developed for use in VOD / OTT content, involves 10 characters. It is a licensed, AI-powered audio conversion and broadcasting system based on [licensing information]. Feature; content input that receives data from VOD / OTT content source (1) The module performs media parsing, separating and processing media streams. module (2) detects character-based speech segments and character and The speaker parsing module (3) calculates the speech timeline for each character as 15 character speech timeline database (4) which creates and stores Licensed audio data collection module (5) collecting licensed audio data, raw audio Audio preprocessing and feature extraction module that cleans recordings and extracts features. (6), license-controlled voice identities governed by license terms. Identity vault (7), re-records the source character speech with the target voice ID 20 voice cloning / voice conversion model (8) that produces emotion and prosody structure Protecting emotion and prosody protection module (9), time code and lip Time code and lip sync compatibility module (10) that ensures synchronization. preserving and recombining original music / effects / ambient sounds Music / effects / ambient sound preservation and remixing module (11), sound 25 Quality control and automated rework that monitor and automatically improve quality. Production module (12), synthetic voice watermark that adds verification trace to synthetic voice and origin information module (13), OTT / VOD creating alternative audio channel Packaging and manifest creation module (14), managing user selection User preference and playback interface (15) and usage data recording 30 It is characterized by including a usage analytics and revenue sharing module (16). - 12 - 2. According to Claim 1, it is an AI-powered audio conversion and broadcasting system, Its feature is audiovisual content such as films, series, animations, documentaries or educational materials. media files, video stream, original audio stream, subtitle file, time code, VOD / OTT content source containing scene transition information and metadata information (1) It is characterized by containing 5 3. According to Claim 1, it is an AI-powered audio conversion and broadcasting system, Feature; movies, series, animation, documentaries or from VOD / OTT content source (1) educational content, audiovisual media files, video streaming, original audio stream, subtitle file, time code, scene transition information and metadata information It is characterized by including a content input module that integrates content into the system. 10 4. According to Claim 1, it is an AI-powered audio conversion and broadcasting system, This feature retrieves video, audio, subtitle, and time code data for VOD / OTT content. Media separation module (2) which performs the field and media separation process It is characterized by its inclusion.

5. According to Claim 1, it is an AI-powered voice conversion and broadcasting system, 15 Features include speaker parsing, automatic speech recognition, and subtitle timing. The code includes at least alignment, face / image matching, and lip movement analysis. Using one to create character-based speech segments, characters and It is characterized by containing a speaker parsing module (3).

6. According to Claim 1, it is an AI-powered voice conversion and broadcasting system, 20 Features include: start time, end time, and speech time for each speech segment. text, phoneme sequence, emotion tag, speech rate, baseline frequency contour, emphasis character speech that stores information about punctuation, breathing intervals, and stage acoustics. It is characterized by containing a timeline database (4).

7. According to Claim 1, it is an AI-powered voice conversion and broadcasting system, 25 Its feature is that it phonetically reproduces voice recordings from licensed voice actors. diversity, emotional diversity and tonal diversity are included in the system. It is characterized by containing a licensed audio data collection module (5).

8. It is an AI-powered audio conversion and broadcasting system according to Claim 1, Features include noise reduction, silence cleaning, and volume level 30 for raw audio recordings. speech preprocessing that performs normalization and speech segmentation operations and is characterized by containing a feature extraction module (6). - 13 - 9. According to Claim 1, it is an AI-powered audio conversion and broadcasting system, Features; mel-spectrogram, MFCC, speaker vector, fundamental frequency, energy Audio preprocessing and features that extract contour, phoneme alignment, and prosody vector. It is characterized by containing a subtraction module (6).

10. AI-powered audio conversion and broadcasting system 5 according to Claim 1. Its features include: voice ID vector usage license, geographic region limit, and content. type permission, character type permission, usage period, and cryptographic pattern ID. together and the audio production model before license verification is complete by including a license-controlled voice ID vault (7) which prevents its operation It is characterized by: 10 11. AI-powered audio conversion and broadcasting system according to Claim 1. Its feature is speech-to-speech conversion and / or one or all of the transcript-to-speech synthesis methods It supports both, including source audio quality, background noise, and subtitles. Sound mode 15 automatically selects the mode based on accuracy and quality score parameters. It is characterized by its inclusion of a cloning / voice conversion model (8).

12. AI-powered audio conversion and broadcasting system according to Claim 1. Its feature is that it uses a VALL-E-like neural codec language model, similar to YourTTS. Text-to-speech synthesis, convolutional neural network-based acoustic model, LSTM / Transformer-based time series modeling, speaker encoder, 20 prosody encoder, acoustic decoder, and neural audio decoder voice cloning / voice conversion model (8) which includes at least one of its components It is characterized by its inclusion.

13. AI-powered audio conversion and broadcasting system according to Claim 1. and its characteristic feature is; the source character's emotion, intonation, emphasis, speech rate and 25 breath interval information is reconstructed to match the target voice ID. Synthesizing and preserving emotion and prosody, which includes voice identity boundary protection functions. It is characterized by containing module (9).

14. AI-powered audio conversion and broadcasting system according to Claim 1. Its characteristic feature is that the sound produced is 30 times the original time code and lip sync. adapting, redefining phoneme durations, speech rate, and silence intervals. edited by and includes mouth movement and lip closing moment in the video content. - 14 - Time code and lip sync matching that analyzes explosive phoneme points. It is characterized by containing module (10).

15. AI-powered audio conversion and broadcasting system according to Claim 1. Its feature is that it preserves the original music, effects, and ambient sounds, and captures the target sound. Stage reverb, channel position, volume level, dynamic range, and EQ 5 Preserving music / effects / ambient sound that remixes according to its parameters. It is characterized by containing a remixing module (11).

16. AI-powered audio conversion and broadcasting system according to Claim 1. Its features include: target similarity score, speech intelligibility, and lip sync. deviation, sound naturalness, sound intensity level, true peak level, cutoff, and 10 measuring noise ratio parameters and sound intensity based on ITU-R BS.1770. Quality control and automation that performs level and actual peak level calculations. It is characterized by containing a reproduction module (12).

17. AI-powered audio conversion and broadcasting system according to Claim 1. Its feature is that it offers different models for segments that fall below the quality threshold. automatic with parameters, vocoder outputs or prosody transfer settings. quality control and automatic re-production module (12) It is characterized by its inclusion. AI-powered audio conversion and broadcasting system according to Claim 18. Its features include; the generated sound has a model ID, license ID, content ID, and character 20. identity, voice actor identity, production time, and quality control result synthetic audio watermark and origin information that includes C2PA compliant origin information. It is characterized by containing module (13).

19. AI-powered audio conversion and broadcasting system according to Claim 1. Its feature is that the generated sound is an alternative audio 25 based on HLS, MPEG-DASH or IMF. packaging as a channel, updating the manifest and pre-generated and on demand. OTT / VOD packaging and instant segment-based production options It is characterized by containing a manifest generation module (14).

20. AI-powered audio conversion and broadcasting system according to Claim 1. Its feature is that it offers different licensed voiceovers for different characters within the same content. 30 the selection of artists and the same character speaking timeline multiple times - 15 - User preference that allows reuse with multiple voice ID models. It is characterized by containing a playback interface (15).

21. AI-powered audio conversion and broadcasting system according to Claim 1. Its characteristic feature is; which voice actor plays which character in which content. For this, the record shows in which region, for how long, and by how many users it was listened to. 5 by including usage analytics and revenue sharing module (16) It is characteristic. 15 25