System for script extraction and subtitle transmission synchronized with live performing art content

The system uses AR smart subtitle glasses to project real-time subtitles synchronized with live performances, addressing language barriers and enhancing comprehension by providing accurate, synchronized subtitles in the viewer's preferred language.

WO2026116526A1PCT designated stage Publication Date: 2026-06-04XPERTINC CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
XPERTINC CO LTD
Filing Date
2024-11-27
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing methods for providing subtitles in live performing arts events disperse audience attention, cause fatigue, and fail to effectively address language barriers, especially for hearing-impaired viewers and those unfamiliar with the language.

Method used

A system that uses AR smart subtitle glasses to project real-time, synchronized subtitles in the viewer's preferred language by removing background music, processing vocals through STT AI, and mapping dialogue to a pre-existing scenario, ensuring accurate synchronization and display.

Benefits of technology

Eliminates language barriers, enhances comprehension, and provides a comfortable viewing experience by allowing simultaneous performance and subtitle viewing, improving accessibility and audience satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019043_04062026_PF_FP_ABST
    Figure KR2024019043_04062026_PF_FP_ABST
Patent Text Reader

Abstract

According to an embodiment, a device for script extraction and subtitle transmission synchronized with live performing art content is provided. When a viewer views verbal performing arts such as plays or musicals, the device projects commentary subtitles synchronized in real time with performance content in a preferred language in front of eyes by using a display device including AR smart subtitle glasses, in order to eliminate a language barrier due to nationality or an auditory barrier caused by hearing impairment or hearing loss and enable viewing of the performing arts, so that the subtitles can be viewed in synchronization with the performance.
Need to check novelty before this filing date? Find Prior Art

Description

Script extraction and subtitle transmission system synchronized with live performing arts content

[0001] The technical concept of the present disclosure relates to a script extraction and subtitle transmission system synchronized with live performing arts content, and more specifically, to a script extraction and subtitle transmission system and device synchronized with live performing arts content that projects explanatory subtitles synchronized in real-time with the performance content in a preferred language onto an AR display device when viewing verbal performing arts such as plays or musicals.

[0002] Unless otherwise indicated in this specification, the contents described in this section are not prior art for the claims of this application, and are not to be recognized as prior art simply because they are included in this section.

[0003] Performing arts encompass all forms of art performed on stage; essentially, it is a form of art where the performer and the audience share the same time and space, with the substance of the work being created on the spot. To overcome the limitations of language barriers in the performing arts world, the success of nonverbal performances—which eliminate language barriers—has brought a new wave to the market since the late 20th century, and such performances can be found anywhere in the United States. However, since the performance methods used to eliminate linguistic borders until now have been nonverbal, there are limitations to entry for performance planners. Therefore, there is a need for a service that allows viewers to freely watch subtitles in their preferred language through glasses while watching original works with language.

[0004] Meanwhile, subtitles are an important element in enhancing the understanding of performing arts. They serve as a primary means to improve comprehension by overcoming language barriers, particularly for the hearing impaired or when watching original performances in foreign languages. Furthermore, they can also aid audience members who understand the language in comprehending the content of the performance.

[0005] As illustrated in Fig. 1, the method of providing subtitle services for language barriers in domestic and world-leading performing arts venues such as Broadway, West End, and Tokyo Shiki displays subtitles on screens fixed to both sides of the stage, above and below, or on the seats, which causes the audience's gaze to be dispersed and results in differences in understanding of the performance depending on the audience's seat location.

[0006] Furthermore, existing service delivery methods require a long eye movement between the live performance and the screen subtitles, leading not only to greater fatigue but also to a situation where viewers fail to fully perceive both performance and subtitle information due to visual dispersion. While the global verbal performing arts market is making various attempts to resolve the issue of language barriers, a clear solution has not yet been found.

[0007] A script extraction and subtitle transmission device synchronized with live performing arts content according to an embodiment enables viewing of performing arts with language (verbal), such as plays or musicals, by projecting explanatory subtitles synchronized with the performance content in real-time in front of the eyes in a preferred language using a display device including AR smart subtitle glasses, thereby allowing simultaneous viewing with the performance. This eliminates language barriers based on nationality or sound barriers caused by hearing loss.

[0008] In the embodiment, during live performances in the performing arts field such as plays, musicals, and operas, background music is removed from the actors' dialogue and music sounds to collect only the vocals. These vocals are then passed through a STT (Speech-to-Text) AI engine, and the converted dialogue information is compared with the original transcribed content to recognize the synchronization point. Subtitles are then projected onto AR glasses and a display device by mapping this point to a pre-existing scenario.

[0009] Additionally, in the embodiment, each actor captures voice using a microphone. The captured analog signals are transmitted to individual channels of an audio mixer, and the venue audio mixer mixes the voice channel signals, excluding unnecessary background sounds or music channels, and routes them to an output bus.

[0010] In addition, in the embodiment, the signal from the audio mixer is transmitted to an audio interface and the digital signal is transmitted to a desktop application via USB. Furthermore, the desktop application divides the audio into chunks and transmits the chunk data to a server.

[0011] In addition, in the embodiment, the server accumulates audio chunks for a certain period of time, inputs the accumulated audio data into an STT model to convert it into text, and performs a search on five script sentences after the previous matched index of the script. At this time, in the embodiment, the search range is limited to optimize the computation speed.

[0012] In addition, in the embodiment, distance-based similarity is measured, and if a script exceeding a threshold is found, the corresponding line is determined to be a matched line, and upon successful matching, the script is sent to a desktop application.

[0013] In addition, in the embodiment, the desktop application transmits real-time script data to smart glasses using the WebSocket protocol through an internal network (local network), and displays the script data received from the smart glasses on the screen in real time.

[0014] However, the problem to be solved according to one embodiment is not limited only to that mentioned above.

[0015] A script extraction and subtitle transmission system synchronized with live performing arts content according to an embodiment may include: a microphone that collects sounds including actors' dialogue and music during a live performance, including plays, musicals, and operas in the field of performing arts; a server that removes background music from the sound collected by the microphone, filters out only vocals, inputs them into a Speech-to-Text (STT) model, recognizes the synchronization point by comparing the converted dialogue information with the original transcribed content, maps it to a pre-stored scenario, and transmits subtitles to an output device according to the mapping result; an audio mixer that receives the sound, adjusts the volume, timbre, and pan of each signal, mixes them into a single output signal, and transmits it; and an output device that projects the subtitles received from the server onto a display to output the subtitles and outputs the output signal.

[0016] Additionally, the server includes a memory for storing at least one command for script extraction and subtitle transmission synchronized with live performing arts content; and a processor for performing an operation according to said command, wherein the processor captures voice from sound transmitted from a microphone and transmits said captured voice to an individual channel of an audio mixer, and said audio mixer can exclude unnecessary background sound or music channels and mix voice channel signals and route them to an output bus.

[0017] Additionally, the processor can receive chunk data in which audio is divided into chunks from a desktop application, accumulate the received audio chunks for a certain period of time, and input the accumulated audio data into an STT model to convert it into text.

[0018] In addition, the processor searches for a certain number of script sentences after the previously matched index of the script, and if it detects a script that exceeds a threshold by measuring distance-based similarity, it can determine the detected dialogue as a matched dialogue.

[0019] In addition, the processor can limit the search range for script sentences to fewer than a certain number.

[0020] Additionally, upon successful matching, the processor can send the matched script to the desktop application and initialize the existing audio chunk.

[0021] In addition, the desktop application transmits real-time script data to an output device including smart glasses using the WebSocket protocol over a local network, and the smart glasses can display the received script data on the screen in real time.

[0022] Additionally, to perform state machine-based speech chunk analysis, the processor attempts state transitions based on thresholds while accepting each chunk, and after checking transcription similarity based on the current state, can decide to transmit dialogue or initialize.

[0023] The script extraction and subtitle output device synchronized with live performing arts content according to the embodiment enables the removal of language barriers by providing subtitles in a language preferred by the audience in performances such as plays, musicals, and operas conducted in a foreign language.

[0024] Furthermore, the embodiment creates an environment where diverse audiences, regardless of nationality, can understand the same performance, and improves accessibility to the performance by enabling the hearing impaired or hard of hearing to understand the dialogue through real-time subtitles. In addition, by providing visual information complementarily to voice-based information delivery, the audience can check the dialogue synchronized in real-time right before their eyes, allowing them to understand the content of the performance immediately without missing anything.

[0025] In addition, the embodiment synchronizes the script and the performance so that the viewer can watch without missing the context of the scene, and provides accurate subtitles in real time through STT AI and scenario mapping.

[0026] In addition, the embodiment removes unnecessary background music and utilizes only the voice channel to increase the accuracy of the dialogue, and improves the reliability of subtitle provision by matching the optimal dialogue through distance-based similarity measurement.

[0027] In addition, through the embodiments, usability for viewers is maximized by utilizing display devices such as AR smart caption glasses, and fast and stable data transmission is enabled through real-time data transmission via WebSockets.

[0028] Furthermore, the embodiment enables real-time processing of various information occurring during the performance, allowing for the rapid reflection of actors' improvised lines or changes in the show in subtitles, and minimizes technical delays during the performance to maintain the flow of the viewing experience. Additionally, it provides a new form of experience that allows the audience to view subtitles synchronized with the performance via smart devices.

[0029] Furthermore, it enables the viewing of original performing arts content anywhere in the world without language barriers through AR subtitle glasses, and by providing a service that breaks down language barriers, it facilitates entry into new markets of the global performing arts industry.

[0030] In addition, through an embodiment, a service is provided that transmits subtitles synchronized with the content during live performances of performing arts by utilizing AI acoustic feature extraction and speech-to-subtitle conversion technology, and allows audience members to set their preferred language to enhance their understanding and immersion in the performance despite language barriers in the original work.

[0031] In addition, through the AR subtitle glasses worn in the embodiment, audience members can watch the performance while simultaneously observing the actors' acting and movements, enabling a more comfortable and enjoyable viewing experience. This not only increases audience satisfaction but also contributes to enhancing the popularity and reputation of the performance.

[0032] In addition, through the embodiments, it is possible to provide performances of performing arts content from various countries around the world in localized linguistic expressions.

[0033] Furthermore, since performing arts attract a large audience of diverse nationalities and languages, the introduction of subtitle glasses technology can provide a more convenient and comfortable viewing environment for multicultural audiences. By enabling understanding through subtitles translated into their native languages, it offers a more friendly and inclusive performance environment.

[0034] In addition, through the subtitle glasses provided in the embodiment, the audience can watch the performance while simultaneously observing the actors' acting and movements, enabling a more comfortable and enjoyable viewing experience. This not only increases audience satisfaction but also contributes to enhancing the popularity and reputation of the performance.

[0035] The effects obtainable from the exemplary embodiments of the present disclosure are not limited to those mentioned above, and other unmentioned effects can be clearly derived and understood by those skilled in the art to which the exemplary embodiments of the present disclosure belong from the description below. That is, unintended effects resulting from the implementation of the exemplary embodiments of the present disclosure can also be derived by those skilled in the art from the exemplary embodiments of the present disclosure.

[0036] Figure 1 is a drawing showing screen subtitles currently in use at a performing arts venue.

[0037] FIG. 2 is a drawing showing an architecture for a script extraction and subtitle transmission service synchronized with live performing arts content according to an embodiment.

[0038] FIG. 3 is a drawing showing a script extraction and subtitle transmission system synchronized with live performing arts content according to an embodiment.

[0039] FIG. 4 is a drawing showing a server configuration according to an embodiment.

[0040] FIG. 5 is a drawing for explaining the functions of a desktop application according to an embodiment.

[0041] FIG. 6 is a diagram illustrating the process of displaying script data received from smart glasses in real time on a display screen according to an embodiment.

[0042] FIG. 7 is a diagram showing a voice chunk state search algorithm according to an embodiment.

[0043] A script extraction and subtitle transmission device synchronized with live performing arts content according to an embodiment enables viewing of performing arts with language (verbal), such as plays or musicals, by projecting explanatory subtitles synchronized with the performance content in real-time in front of the eyes in a preferred language using a display device including AR smart subtitle glasses, thereby allowing simultaneous viewing with the performance. This eliminates language barriers based on nationality or sound barriers caused by hearing loss.

[0044] Hereinafter, various embodiments of the present disclosure are described in conjunction with the accompanying drawings. As various embodiments of the present disclosure may be subject to various modifications and may have various embodiments, specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the various embodiments of the present disclosure to specific forms, and it should be understood that they include all modifications and / or equivalents and substitutions that fall within the spirit and scope of the various embodiments of the present disclosure. In relation to the description of the drawings, similar reference numerals have been used for similar components.

[0045] In various embodiments of the present disclosure, terms such as “comprising” or “having” are intended to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0046] In various embodiments of the present disclosure, expressions such as “or” include any and all combinations of the words listed together. For example, “A or B” may include A, may include B, or may include both A and B.

[0047] Expressions such as "first," "second," "first," or "second" used in various embodiments of the present disclosure may modify various components of the various embodiments, but do not limit such components. For example, such expressions do not limit the order and / or importance of such components and may be used to distinguish one component from another.

[0048] When it is mentioned that a component is "connected" or "joined" to another component, it should be understood that the component may be directly connected or joined to the other component, but that a new component may also exist between the component and the other component.

[0049] In the embodiments of the present disclosure, terms such as "module," "unit," "part," etc. are used to refer to a component that performs at least one function or operation, and such component may be implemented in hardware or software, or in a combination of hardware and software. Additionally, a plurality of "modules," "units," "parts," etc. may be integrated into at least one module or chip and implemented as at least one processor, except where each needs to be implemented in specific individual hardware.

[0050] Terms such as those defined in commonly used dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in the various embodiments of the present disclosure.

[0051] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0052] FIG. 2 is a diagram showing an architecture for a script extraction and subtitle transmission service synchronized with live performing arts content according to an embodiment.

[0053] Referring to FIG. 2, the script extraction and subtitle transmission system synchronized with live performing arts content according to the embodiment removes background music from the actors' dialogue and music sounds during live performances in the performing arts field, such as plays, musicals, and operas, to collect only the vocals and input them into a Speech-to-Text (STT) model. Subsequently, by comparing the dialogue information converted into text with the original transcribed content, the system recognizes the synchronization point and maps it to a pre-existing scenario to project subtitles onto AR glasses and display devices. In the embodiment, each actor captures their voice using a microphone, and the captured analog signal is transmitted to an individual channel of an audio mixer. Additionally, the venue audio mixer excludes unnecessary background sounds or music channels, mixes the voice channel signals, and routes them to an output bus. Furthermore, in the embodiment, the signal from the audio mixer is transmitted to an audio interface. Additionally, in the embodiment, a digital signal is transmitted to a desktop application via USB.

[0054] FIG. 3 is a diagram illustrating a script extraction and subtitle transmission system synchronized with live performing arts content according to an embodiment. Referring to FIG. 3, the script extraction and subtitle transmission system synchronized with live performing arts content according to an embodiment may be configured to include a microphone (10), a server (100), an audio mixer (20), and an output device (30). The microphone (10) collects sounds including actors' dialogue and music during live performances, including plays, musicals, and operas in the field of performing arts. When the server (100) removes background music from the sound collected by the microphone (10) and filters out only the vocals to input them into a STT (Speech-to-Text) model, it compares the converted dialogue information with the original transcribed content to recognize the synchronization point, maps it to a pre-stored scenario, and transmits subtitles to the output device (30) according to the mapping result. The audio mixer (20) receives the sound, adjusts the volume, tone, and pan of each signal, and mixes them into a single output signal for transmission. The output device (30) projects the subtitles received from the server (100) onto the display to output the subtitles and outputs an output signal.

[0055] FIG. 4 is a block diagram of a server according to an embodiment. In the embodiment, a server is a computing system that provides services to other computers or devices in a computer network or stores and manages data. The server (100) accepts requests from other computers or devices called clients and provides responses or data to those requests. The configuration of the server (100) shown in FIG. 2 is merely a simplified example.

[0056] The communication module (110) can be configured regardless of the mode of communication, such as wired or wireless, and can be configured with various communication networks, such as a Personal Area Network (PAN) or a Wide Area Network (WAN). Additionally, the communication module (110) can operate based on the known World Wide Web (WWW) and may utilize wireless transmission technologies used for short-range communication, such as Infrared Data Association (IrDA) or Bluetooth. For example, the communication module (110) may be responsible for transmitting and receiving data necessary to perform a technique according to one embodiment of the present disclosure.

[0057] Memory (120) may refer to any type of storage medium. For example, memory (120) may include at least one type of storage medium among flash memory type, hard disk type, multimedia card micro type, card type memory (e.g., SD or XD memory, etc.), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), magnetic memory, magnetic disk, and optical disk. Such memory (120) may also constitute the database shown in FIG. 1.

[0058] The memory (120) can store at least one instruction that can be executed by the processor (130). Additionally, the memory (120) can store any form of information generated or determined by the processor (130) and any form of information received by the server (200). Additionally, the memory (120) stores various types of modules, instruction sets, or models.

[0059] The processor (130) can perform technical features according to embodiments of the present disclosure to be described below by executing at least one instruction stored in memory (120). In one embodiment, the processor (130) may be composed of at least one core and may include a processor for data analysis and / or processing, such as a central processing unit (CPU) of a computer device, a general purpose graphics processing unit (GPGPU), or a tensor processing unit (TPU).

[0060] This processor (130) can train a neural network or model designed in a machine learning or deep learning manner. To this end, the processor (130) can perform calculations for training the neural network, such as processing input data for training, extracting features from input data, calculating errors, and updating the weights of the neural network using backpropagation. Additionally, the processor (130) can perform inference for a specific purpose using a model implemented in an artificial neural network manner.

[0061] In an embodiment, the processor (130) captures voice from sound transmitted from a microphone and transmits the captured voice to individual channels of an audio mixer. To do this, the processor (130) receives analog sound signals transmitted from microphones worn or used by each actor during a performance. The signal may include various audio elements such as the actor's dialogue, background noise, and ambient music. Subsequently, the processor (130) analyzes the received sound signal to separate and capture the signal containing the voice. For example, the processor (130) may utilize specific frequency band filtering, noise removal algorithms, or Voice Activity Detection (VAD) technology. The aforementioned process is intended to ensure clarity of the dialogue signal. Additionally, the processor (130) converts the captured voice signal into a digital or analog form and transmits it to individual channels of an audio mixer. The signal from each microphone is assigned to a different channel, thereby allowing the audio mixer to independently control and mix the voice of each actor. In the embodiment, the processor (130) applies sampling rate adjustment, signal compression and restoration, or other audio signal processing techniques so that the quality of the voice is maintained during the signal transmission process. This is to ensure the stability of real-time processing during the performance.

[0062] Additionally, in the embodiment, the audio mixer excludes unnecessary background noise or music channels, mixes the voice channel signals, and routes them to the output bus. In the embodiment, the audio mixer receives signals generated from each microphone, instrument, or background noise during the performance and assigns them to individual channels. Each channel consists of an independently adjustable signal path, and among these, the actor's voice signal is identified as the voice channel. Subsequently, to exclude specific channels (background noise, music, etc.) from the input signal, the audio mixer first focuses on the frequency band of the voice channel to remove background noise or unnecessary noise. Additionally, it analyzes the signals within the channel using a Voice Activity Detection (VAD) algorithm to extract only the portions containing the voice signal. Afterward, it performs signal intensity adjustment to minimize unintentional noise and improve the clarity of the voice signal.

[0063] Subsequently, the audio mixer selectively mixes the identified voice channel signals to generate a single audio output signal. During this process, the audio mixer adjusts the volume balance between channels to ensure that each actor's voice is properly transmitted. Additionally, it optimizes the stereo image of the audio signal (e.g., left / right distribution) through pan adjustment and maintains the clarity of the voice signal by utilizing an equalizer. Furthermore, to route the mixed voice signal to the output bus, a path is established in the mixer's routing table so that the voice signal is transmitted to the output bus. In the embodiment, the routed signal is converted into a digital signal as needed and delivered in a form suitable for further processing or transmission. The converted signal is then transmitted to an audio interface or other output device and provided to a playback system or subtitle generation system within the performance venue.

[0064] In the embodiment, signals are processed in real-time during the performance to ensure that important voice information, such as actors' dialogue, is transmitted accurately without delay. To achieve this, the audio mixer utilizes a high-performance Digital Signal Processing (DSP) engine to minimize latency.

[0065] Additionally, the processor (130) receives chunk data in which audio is divided into chunk units from a desktop application, accumulates the received audio chunks for a certain period of time, and inputs the accumulated audio data into an STT model to convert it into text. The processor (130) receives audio chunk data transmitted from the desktop application via a network (local or cloud-based). FIG. 5 is a diagram illustrating the functions of a desktop application according to an embodiment. In the embodiment, an audio chunk is digital audio data divided into units of a certain time, and each chunk consists of a fixed size and a sampling rate. In the embodiment, to prevent loss of chunk data, the processor (130) applies an error verification technique (e.g., CRC, Checksum) upon receiving data to verify the integrity of the data. Additionally, the received chunk data is temporarily stored in a buffer within the processor (130). In the embodiment, the buffer is used to accumulate a certain amount of audio data by accumulating chunk data. In the embodiment, when the accumulated audio data reaches the minimum input length required by the STT model or when a specified time (e.g., 1 to 2 seconds) has elapsed, it is processed to the next step. Afterward, the processor (130) applies an algorithm to align data boundaries or maintain continuity between audio samples in order to prevent duplicate or missing issues that may occur at the boundaries of data divided into chunks.

[0066] The accumulated audio data is converted into the sampling rate, bit depth, and data format (e.g., PCM, WAV) required by the STT model. In the embodiment, the data may be converted into a form optimized for input to the STT model through decompression or normalization. Subsequently, the processor (130) transmits the prepared audio data to the STT model in real time, and the model analyzes the input audio data and converts it into text. The text data returned from the STT model is received by the processor (130). The conversion result is recorded along with the audio's timestamp to support synchronization processing. In the embodiment, after text conversion, the processor compares the STT result with a predefined script to calculate a match rate (similarity-based matching) and verifies the final dialogue. If the quality of the converted text is below a standard, additional processing (e.g., acoustic analysis or reprocessing) may be performed. The final text data is transmitted to a desktop application and, if necessary, stored in a local database to be used for subtitle generation, translation, or archiving purposes.

[0067] Additionally, the processor (130) performs a search for a certain number (e.g., 5) of script sentences after the previously matched index of the script, and if it detects a script exceeding a threshold by measuring distance-based similarity, it determines the detected dialogue as a matched dialogue. The process of determining the matched dialogue and displaying script data is explained in more detail below with reference to FIG. 6. FIG. 6 is a diagram illustrating the process of displaying script data received from smart glasses according to an embodiment on a screen in real time. In the embodiment, the processor (130) records the index of the previously matched script and sets a reference point to start a new search. The reference point set in the embodiment indicates the location of the recently identified dialogue within the script data. Additionally, the processor (130) minimizes the amount of computation by performing a search targeting only a certain number (e.g., 5 sentences) after the previously matched index. The search range can be dynamically adjusted considering the speed of the performance and the script structure. Subsequently, the processor (130) receives text data generated through a STT (Speech-to-Text) model and preprocesses it into an analyzable form. In the embodiment, unnecessary spaces and special characters are removed during the preprocessing process, case is standardized, normalization is performed to handle pronunciation and grammatical variations, and script sentence data is retrieved. Additionally, the processor (130) sequentially retrieves script sentences within the search range to generate a comparison target list. Subsequently, the processor (130) calculates distance-based similarity between the STT converted text and the retrieved script sentences. Methods such as edit distance calculation, coin similarity calculation, Jaccard similarity calculation, and Levenstein distance-based similarity calculation may be used for similarity calculation, but are not limited thereto.Edit Distance (Levenshtein Distance) measures the difference between two texts by calculating the number of insertion, deletion, and replacement operations, while Cosine Similarity measures similarity based on the angles between vectors after vectorizing the texts. Jaccard Similarity calculates the ratio of the intersection to the union of the word sets of two texts. Levenshtein Distance is a type of Edit Distance used to calculate the similarity between two strings. This distance represents the minimum number of edit operations (insertion, deletion, or replacement) required to transform one string into another.

[0068] In the embodiment, the processor (130) pre-sets a threshold for determining similarity, and if the similarity is, for example, 0.8 (80%) or higher, the two texts are considered to be matched. In the embodiment, the threshold can be adjusted according to the performance environment. Subsequently, the processor (130) selects the sentence with the highest similarity value among the searched script sentences, and if the selected sentence exceeds the threshold, it detects it as a matched dialogue. If there are no sentences exceeding the threshold, the processor (130) expands the search range to compare additional sentences. Additionally, the STT converted text can be reprocessed (noise removal, additional filtering, etc.) and a request for manual verification can be made to the user.

[0069] Subsequently, the processor (130) stores the index of the detected matching dialogue to use as a reference point for subsequent search, and the matched dialogue is stored along with a timestamp and used as a synchronized subtitle or record. In the embodiment, the final matching dialogue is transmitted to a desktop application and is displayed in real time on an AR subtitle device or display device.

[0070] Additionally, the processor (130) optimizes the computation speed by limiting the search range for script sentences to fewer than a certain number. The processor (130) records the index of the previously matched script and sets it as a reference point for a new search. The reference point reflects the synchronization status of the script and the progress of the performance. Subsequently, the processor (130) sets only a certain number (e.g., 5 or fewer) of script sentences after the reference point as search targets. This range may be set in advance considering the structure of the script and the average interval between lines, or may be dynamically adjusted during the performance. In the embodiment, the processor (130) loads only the script sentences corresponding to the search range after the reference point into memory for processing. This prevents memory overload that may occur when searching the entire script. Additionally, script sentences outside the search range are excluded from the computation target, thereby reducing unnecessary computations. To satisfy the real-time requirements of the performance, the processor (130) updates the search range in real time. In the embodiment, by limiting the search range, the processor minimizes the amount of data to be processed and shortens the time required for search and comparison operations. Additionally, since the processor (130) calculates distance-based similarity (e.g., Levenstein distance) only within the search range, there is no need to repeatedly calculate similarity for large datasets. The reduction in computational load improves the response speed of the real-time subtitle generation system.

[0071] Additionally, the processor (130) can dynamically adjust the search range according to the situation during the performance. For example, if the interval between lines is short, the search range can be reduced, and if the interval is long, it can be expanded. Through this, the system optimizes the computation speed while maintaining the accuracy of line matching. Furthermore, in the embodiment, if the processor (130) does not find a line with a similarity exceeding a threshold within the search range, it can expand the search range to call additional script sentences and proceed with the comparison. In the embodiment, when a matched line is confirmed within the searched range, the line is recorded by the processor and transmitted to a desktop application. Additionally, the index of the matched line is set as a new reference point, and the next search is prepared.

[0072] In an embodiment, the processor (130) transmits the matched script to a desktop application upon successful matching and initializes the existing audio chunk. To this end, the processor (130) checks the distance-based similarity calculation result between the Speech-to-Text (STT) converted text and the script, and determines that the match is successful if it exceeds a threshold. Additionally, the script sentence with the highest similarity is selected as the matched script, and the information is transmitted as data to the desktop application. In an embodiment, the matched script is transmitted to the desktop application via an internal network (local network) or other communication protocols (e.g., WebSockets). The transmitted data may include the matched script text, timestamp information of the dialogue, and metadata (e.g., actor name, scene number, etc.).

[0073] In the embodiment, the processor (130) transmits data in a lightweight format to minimize delay during the transmission process. For example, it is configured in JSON format to reduce data size and improve processing speed. Additionally, after a successful match, the processor (130) deletes or initializes existing audio chunk data accumulated in the internal buffer. To this end, the processor (130) empties the buffer memory to prepare for receiving the next audio data and optimizes memory usage by preventing unnecessary data retention.

[0074] In the embodiment, initialization is performed immediately after the script is transmitted, thereby securing space for new audio chunks received thereafter. After initialization, the processor (130) completes preparations to receive the next audio chunk data. In the embodiment, the newly received data is accumulated again in the initialization state, and a new matching operation is started based on this. In the embodiment, the index of the matched script is recorded as a reference point to set the starting point for subsequent search and matching operations. Additionally, if the matching fails or the similarity is below a threshold, the audio chunk is not initialized and is preserved in the buffer. The processor (130) retryes the search by accumulating additional data or proceeds with the next matching operation by expanding the search range.

[0075] Additionally, the desktop application transmits real-time script data to an output device, including smart glasses, using the WebSocket protocol over a local network, and the smart glasses display the received script data on the screen in real time. The desktop application acts as a WebSocket server, and the smart glasses or other output device is configured as a WebSocket client. In the embodiment, the client requests a connection to the server based on the local network (IP address and port number). In the embodiment, the WebSocket connection is established over the TCP protocol, and the client creates a continuous bidirectional communication channel through a handshake process with the server. Once successfully connected, both parties are ready to exchange data in real time. Subsequently, the desktop application converts the matched script data into JSON, XML, or other lightweight data formats. Additionally, a transmission queue is created to manage the real-time script data, ensuring data order and providing a loss prevention mechanism. In the embodiment, the prepared script data is transmitted to the output device (smart glasses) via the WebSocket protocol. The data is delivered quickly and reliably through bidirectional communication. In the embodiments, data is transmitted in UTF-8 encoded text or binary format. Text-based message formats (JSON or plain text) are primarily preferred. Additionally, the desktop application synchronizes the transmission cycle of the script data with the pace of the performance. For example, the corresponding data is delivered to an output device along with timestamps in accordance with the dialogue timing. In the embodiments, the smart glasses or other output device acts as a client and receives the transmitted data via a WebSocket connection. Subsequently, the received data is parsed and displayed on the screen. The script content is visually provided in real time through the smart glasses' Head-Up Display (HUD) or other display interface.Since WebSockets support bidirectional communication, smart glasses can periodically notify a desktop application of the connection status. For example, they check the connection status using Ping / Pong messages and automatically attempt to reconnect in the event of a network disconnection or error. Additionally, in the embodiment, if data is lost or the connection is interrupted, the desktop application retransmits unsent data or activates a recovery mechanism to recover missing data.

[0076] At the end of the performance, the desktop application properly terminates the WebSocket connection with the smart glasses. To this end, termination messages are exchanged between the client and the server.

[0077] After the connection is terminated, desktop applications and output devices release used memory and network resources to maintain system efficiency.

[0078] Additionally, the processor (130) attempts a state transition according to a threshold while accepting each chunk to perform state machine-based voice chunk analysis, and determines whether to transmit the dialogue or initialize after checking transcription similarity according to the current state. FIG. 7 is a diagram illustrating a voice chunk state search algorithm according to an embodiment. In the embodiment, the state machine includes the following main states that divide the voice chunk analysis process into steps.

[0079] Idle state: Initial state, waiting to start analysis

[0080] Accumulating status: Collects and accumulates voice chunk data

[0081] Processing Status: Analysis of accumulated data and enterprise

[0082] Matched status: Transcription result matches script

[0083] Error Status: Below threshold or matching failure

[0084] In the embodiment, the processor (130) starts in an Idle state upon system initialization and transitions to an Accumulating state upon receiving voice chunk data. Subsequently, the processor (130) continuously receives voice chunk data transmitted from the STT system or desktop application. Subsequently, when voice chunk data begins to be received, the processor (130) transitions from the Accumulating state to the Idle state, and transitions from the Accumulating state to the Processing state when chunks accumulate for a certain period of time or when the data size reaches a threshold. Additionally, the transition from the Processing state to the Matched state occurs when the transcription result matches the script and the similarity exceeds a threshold. Furthermore, if matching fails, the processor transitions from the Processing state to the Error state.

[0085] In the embodiment, in the Processing state, the processor (130) inputs the accumulated voice chunk data into a Speech-to-Text (STT) model to convert it into text. Additionally, it measures the similarity between the transcribed text and sentences within the current search range of the script. A distance-based algorithm (e.g., Levenstein distance) is used for similarity measurement, and if the similarity exceeds a threshold (e.g., 80%), it is considered a match. In the embodiment, if the transcription similarity exceeds the threshold, it transitions to the Matched state, and the matched script text and timestamp data are transmitted to a desktop application. Subsequently, the voice chunk data up to that point is initialized to prepare for receiving new data, and if the match fails, it transitions to the Error state. When transitioning to the Error state, the existing chunk data is preserved, additional data is accumulated for re-analysis, the search range is expanded, or the quality of the accumulated data is corrected to attempt the match again. Afterward, once the dialogue is successfully transmitted, the processor returns to the Idle state and waits for new voice chunk data. In the embodiment, if the chunk data is lost or incomplete, the processor (130) maintains the current state and attempts to re-receive. Subsequently, if a network error occurs during the transmission process, the processor temporarily stores the data locally and resumes transmission once the connection is restored.

[0086] The script extraction and subtitle output device synchronized with live performing arts content according to the embodiment enables the removal of language barriers by providing subtitles in a language preferred by the audience in performances such as plays, musicals, and operas conducted in a foreign language.

[0087] Furthermore, the embodiment creates an environment where diverse audiences, regardless of nationality, can understand the same performance, and improves accessibility to the performance by enabling the hearing impaired or hard of hearing to understand the dialogue through real-time subtitles. In addition, by providing visual information complementarily to voice-based information delivery, the audience can check the dialogue synchronized in real-time right before their eyes, allowing them to understand the content of the performance immediately without missing anything.

[0088] In addition, the embodiment synchronizes the script and the performance so that the viewer can watch without missing the context of the scene, and provides accurate subtitles in real time through STT AI and scenario mapping.

[0089] In addition, the embodiment removes unnecessary background music and utilizes only the voice channel to increase the accuracy of the dialogue, and improves the reliability of subtitle provision by matching the optimal dialogue through distance-based similarity measurement.

[0090] In addition, through the embodiments, usability for viewers is maximized by utilizing display devices such as AR smart caption glasses, and fast and stable data transmission is enabled through real-time data transmission via WebSockets.

[0091] Furthermore, the embodiment enables real-time processing of various information occurring during the performance, allowing for the rapid reflection of actors' improvised lines or changes in the show in subtitles, and minimizes technical delays during the performance to maintain the flow of the viewing experience. Additionally, it provides a new form of experience that allows the audience to view subtitles synchronized with the performance via smart devices.

[0092] Furthermore, it enables the viewing of original performing arts content anywhere in the world without language barriers through AR subtitle glasses, and by providing a service that breaks down language barriers, it facilitates entry into new markets of the global performing arts industry.

[0093] In addition, through an embodiment, a service is provided that transmits subtitles synchronized with the content during live performances of performing arts by utilizing AI acoustic feature extraction and speech-to-subtitle conversion technology, and allows audience members to set their preferred language to enhance their understanding and immersion in the performance despite language barriers in the original work.

[0094] In addition, through the AR subtitle glasses worn in the embodiment, audience members can watch the performance while simultaneously observing the actors' acting and movements, enabling a more comfortable and enjoyable viewing experience. This not only increases audience satisfaction but also contributes to enhancing the popularity and reputation of the performance.

[0095] In addition, through the embodiments, it is possible to provide performances of performing arts content from various countries around the world in localized linguistic expressions.

[0096] Furthermore, since performing arts attract a large audience of diverse nationalities and languages, the introduction of subtitle glasses technology can provide a more convenient and comfortable viewing environment for multicultural audiences. By enabling understanding through subtitles translated into their native languages, it offers a more friendly and inclusive performance environment.

[0097] In addition, through the subtitle glasses provided in the embodiment, the audience can watch the performance while simultaneously observing the actors' acting and movements, enabling a more comfortable and enjoyable viewing experience. This not only increases audience satisfaction but also contributes to enhancing the popularity and reputation of the performance.

[0098] Meanwhile, the methods according to the various embodiments of the present invention described above can be implemented in the form of an application or software program that can be installed on an existing electronic device.

[0099] In addition, the whole or part of the method may be composed of multiple software function modules and implemented on an operating system (OS). Alternatively, each step may be composed of a single software function module, or each step may be combined to form a single software function module and implemented on an operating system. Therefore, even if all of the embodiments of the present disclosure are not implemented as a single software function module, if multiple software function modules implement each step of the present disclosure and multiple software function modules are implemented on a single operating system, it can be understood that the method of the present disclosure has been implemented.

[0100] In addition, the methods according to the various embodiments of the present invention described above can be implemented solely through software upgrades or hardware upgrades of existing electronic devices. Furthermore, the various embodiments of the present invention described above can also be performed through an embedded server equipped in an electronic device or an external server of the electronic device.

[0101] Meanwhile, according to one embodiment of the present invention, the various embodiments described above may be implemented as software comprising instructions stored on a computer-readable recording medium using software, hardware, or a combination thereof. In some cases, the embodiments described herein may be implemented as the processor itself. According to the software implementation, embodiments such as the procedures and functions described herein may be implemented as separate software modules. Each of the software modules may perform one or more functions and operations described herein.

[0102] Meanwhile, a computer or a similar device may include a device according to the disclosed embodiments, which is capable of calling instructions stored from a storage medium and operating according to the called instructions. When said instructions are executed by a processor, the processor may perform a function corresponding to said instructions directly or by using other components under the control of said processor. The instructions may include code generated or executed by a compiler or an interpreter.

[0103] A computer-readable recording medium may be provided in the form of a non-transitory computer-readable recording medium. Here, "non-transitory" simply means that the storage medium does not contain a signal and is tangible, without distinguishing whether data is stored semi-permanently or temporarily on the storage medium. In this context, a non-transitory computer-readable medium refers to a medium that stores data semi-permanently and is readable by a device, rather than a medium that stores data for a short moment, such as registers, caches, or memory. Specific examples of non-transitory computer-readable media may include CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, and ROMs.

[0104] As described above, exemplary embodiments have been disclosed in the drawings and specification. Although specific terms have been used to describe the embodiments in this specification, they are used only for the purpose of explaining the technical concept of this disclosure and are not intended to limit the meaning or the scope of this disclosure as defined in the claims. Therefore, those skilled in the art will understand that various modifications and equivalent alternative embodiments are possible therefrom. Accordingly, the true technical scope of protection of this disclosure should be determined by the technical concept of the appended claims.

[0105] The script extraction and subtitle output device synchronized with live performing arts content according to the embodiment enables the removal of language barriers by providing subtitles in a language preferred by the audience in performances such as plays, musicals, and operas conducted in a foreign language.

[0106] Furthermore, the embodiment creates an environment where diverse audiences, regardless of nationality, can understand the same performance, and improves accessibility to the performance by enabling the hearing impaired or hard of hearing to understand the dialogue through real-time subtitles. In addition, by providing visual information complementarily to voice-based information delivery, the audience can check the dialogue synchronized in real-time right before their eyes, allowing them to understand the content of the performance immediately without missing anything.

[0107] 100: Server

Claims

1. In a script extraction and subtitle transmission system synchronized with live performing arts content, A microphone for collecting sound including actors' dialogue and music during live performances in the performing arts, including plays, musicals, and operas; A server that removes background music from the sound collected from the above microphone, filters out only vocals, inputs them into a STT (Speech-to-Text) model, compares the converted dialogue information with the original transcript, recognizes the synchronization point, maps it to a pre-stored scenario, and transmits subtitles to an output device according to the mapping result; An audio mixer that receives the above sound input, adjusts the volume, timbre, and pan of each signal, mixes them into a single output signal, and transmits it; A script extraction and subtitle transmission system synchronized with live performing arts content, comprising: an output device that projects a subtitle received from the server onto a display to output the subtitle and outputs the output signal.

2. In paragraph 1, the above server Memory for storing at least one command for script extraction and subtitle output synchronized with live performing arts content; and It includes a processor that performs an operation according to the above instruction, The above processor is, Capturing voice from sound transmitted from a microphone, and transmitting the captured voice to individual channels of an audio mixer, The above audio mixer A script extraction and subtitle transmission system synchronized with live performing arts content that excludes unnecessary background music or music channels, mixes voice channel signals, and routes them to an output bus.

3. In paragraph 2, the processor A script extraction and subtitle transmission system synchronized with live performing arts content, which receives chunk data in which audio is divided into chunk units from a desktop application, accumulates the received audio chunks for a certain period of time, and inputs the accumulated audio data into an STT model to convert it into text.

4. In paragraph 3, the processor A script extraction and subtitle transmission system synchronized with live performing arts content, which searches for a certain number of script sentences after the previously matched index of the script, measures distance-based similarity, and determines the detected dialogue as a matched dialogue when a script exceeding a threshold is detected.

5. In paragraph 4, the processor A script extraction and subtitling system synchronized with live performing arts content that limits the search range for script sentences to fewer than a certain number.

6. In paragraph 4, the processor A script extraction and subtitle transmission system synchronized with live performing arts content that transmits the matched script to a desktop application and initializes existing audio chunks upon successful matching.

7. In paragraph 6, the above desktop application Transmit real-time script data to an output device including smart glasses using the WebSocket protocol over a local network, and The smart glasses mentioned above A script extraction and subtitle transmission system synchronized with live performing arts content that displays received script data on the screen in real time.

8. In paragraph 3, the processor A script extraction and subtitle transmission system synchronized with live performing arts content, which attempts state transitions based on thresholds while accepting each chunk to perform state machine-based speech chunk analysis, checks transcription similarity based on the current state, and then decides to transmit dialogue or initialize.