In-game automatic subtitles and closed captions

JP2025504748A5Pending Publication Date: 2025-10-03ADVANCED MICRO DEVICES INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024535349
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-23
Filing Date
2022-12-01
Publication Date
2025-10-03

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An approach is provided for a gaming overlay application to provide automatic in-game subtitling and / or closed captioning for a video gaming application. The overlay application accesses an audio stream and a video stream generated by a running gaming application. The overlay application processes the audio stream with a text-to-text engine to generate at least one subtitle. The overlay application determines a display location to associate with the at least one subtitle. The overlay application generates a subtitle overlay including the at least one subtitle located at the associated display location. The overlay application causes a portion of the video stream to be displayed with the subtitle overlay.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Thus, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Further, it should not be assumed that any of the approaches described in this section are well understood, conventional, or conventional by virtue of their inclusion in this section.

[0002] Subtitles or closed captions for interactive content can provide important accessibility features for users with hearing impairments or difficult listening environments. Users who are hearing impaired, deaf, or suffer from tinnitus or other hearing conditions may not be able to fully understand audio cues and spoken dialogue. Noisy environments can exacerbate the problem, such as when a user is using public transportation, crossing a crowded space, or is in close proximity to buildings, traffic, musical performances, or other sources of background noise. Conversely, in environments where silence must be maintained, such as an office or library, or late at night when anti-noise ordinances may be in place, audio must be played at a low volume or muted, making it difficult to hear the audio clearly. Headphones can assist in hearing the audio, but the headphones may be misplaced, forgotten, or may not be compatible with hearing aids or other devices. Even if the spoken dialogue is clearly audible to the user, it may be spoken in a foreign language or dialect or accent that is not easily understood by the user. In these cases, subtitles or closed captions can help users understand the audio better.

[0003] Providing subtitles and closed captions for interactive content such as video and computer games can provide greater accessibility and more efficient gameplay interaction for a wider range of users. However, because video and computer games are programmed in different environments using different game engines and development methodologies, there is no universal standard for presenting subtitles and closed captions in games. Thus, games do not always natively support subtitles. Even if subtitles or closed captions are natively supported in games, only a limited number of languages ​​may be supported, or subtitles may be displayed only in limited portions of the game content, such as only in certain cut scenes. Thus, there is a need for an approach to provide subtitles or closed captions for computer and video games in a more flexible manner.

[0004] Embodiments are illustrated by way of example, and not by way of limitation, in the accompanying figures, in which like references refer to similar elements and in which: [Brief description of the drawings]

[0005] [Figure 1] FIG. 1 is a block diagram illustrating a system for implementing automatic in-game subtitling as described herein. [Figure 2A] FIG. 1 illustrates an example of a graphical user interface (GUI) for a video game application. [Figure 2B] FIG. 1 illustrates an example graphical user interface (GUI) for a video game application having in-game automatic subtitling. [Figure 2C] FIG. 1 illustrates an example graphical user interface (GUI) of a video game application having automatic in-game subtitles positioned in proximity to an audio source. [Diagram 3] FIG. 1 is a flow diagram illustrating an approach for implementing automatic in-game subtitling. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0006] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the embodiments. However, it will be apparent to one skilled in the art that the embodiments may be practiced without these specific details.

[0007] I. Overview II. Architecture III. Graphical User Interface of an Exemplary Gaming Application IV. In-game automatic subtitle generation process

[0008] (I. Overview) An approach is provided for a gaming overlay application to provide automatic in-game subtitling and / or closed captioning for a video gaming application. The overlay application accesses an audio stream and a video stream generated by a running gaming application. In an embodiment, the video stream includes frames of image data rendered during execution of the gaming application. The overlay application processes the audio stream with a text conversion engine, which in an embodiment includes a speech-to-text engine for generating at least one subtitle. The overlay application determines a display location to associate with the at least one subtitle. The overlay application generates a subtitle overlay including the at least one subtitle located at the associated display location. The overlay application causes at least a portion of the video stream to be displayed with the subtitle overlay.

[0009] The techniques described herein enable game overlay applications to analyze real-time audio streams from video games to generate subtitles that are displayed, even if the video game does not natively support subtitles. By using various cues, such as multi-channel surround sound information and machine learning-based voice profile matching, dialogue and audio cues are associated with specific characters, multi-player users, or other elements shown in the game, and subtitles are positioned on the screen at the user's preferred location or in close proximity to the associated sound source. In this way, users quickly identify speakers and their associated dialogue, even when the audio is hard to hear or muted. This allows users to react more quickly and efficiently by understanding and reacting to audio cues, even in hearing-impaired or hard-to-listen environments. Furthermore, since the techniques are applicable to any video game that generates audio, the described techniques can be used with video games that do not natively support subtitles. In an embodiment, subtitles are shown in various situations, including cut scenes, in matching lobbies, or during game play.

[0010] (II. Architecture) FIG. 1 is a block diagram illustrating a system 100 for implementing automatic in-game subtitling as described herein. Subtitling, as used in this disclosure, includes the transcription or translation of dialogue or speech of a video, video game, etc., and description of sound effects, musical cues, or other related audio information from the video / video game. Thus, reference to subtitling also includes closed captioning, or subtitling with additional context such as speaker identification and non-speech elements such as description of sound effects and audio cues. In an embodiment, the system 100 includes a computing device 110, a network 160, an input / output (I / O) device 170, and a display 180. In an embodiment, the computing device 110 includes a processor 120, a graphics processing unit (GPU) 122, a data bus 124, and a memory 130. In an embodiment, the GPU 122 includes memory for storing one or more frame buffers 123. In an embodiment, memory 130 stores game application 140 and game overlay application 150. In an embodiment, game application 140 outputs audio stream 142 and video stream 144. Game overlay application 150 includes text conversion engine 152, subtitle synthesizer 154, voice profile database 156, and user preferences 158. I / O device 170 includes microphone 172 and speaker 174. Display 180 includes an interface for receiving game graphics 182 from computing device 110. In an embodiment, game graphics 182 includes subtitle overlay 190. The components of system 100 are merely exemplary and any configuration of system 100 can be used according to the requirements of game application 140.

[0011] The gaming application 140 is executed on the computing device 110 by one or more of the processor 120, the GPU 122, or other computing resources not specifically shown. The processor 120 may be any type of general purpose single or multi-core processor, or a dedicated processor such as an application-specific integrated circuit (ASIC) or a field programmable gate array (FPGA). In an embodiment, there are two or more processors 120. The GPU 122 is any type of dedicated hardware for graphics processing that is addressable using various graphics application programming interfaces (APIs), such as DirectX, Vulkan, OpenGL, and OpenCL. In an embodiment, the GPU 122 includes a frame buffer 123, where finalized video frames are stored before output to the display 180. The data bus 124 is any high-speed interconnect for communication between components of the computing device 110, such as a Peripheral Component Interconnect (PCI) Express bus, Infinity Fabric, or Infinity Architecture, etc. Memory 130 may be any type of memory, such as random access memory (RAM) or other storage device.

[0012] As shown in FIG. 1, the gaming application 140 generates an audio stream 142 and a video stream 144 corresponding to real-time audio and video content. In some embodiments, the audio stream 142 and the video stream 144 are combined into a single audiovisual stream. The audio stream 142 corresponds to internally generated in-game audio, and in embodiments includes multiple channels for surround sound and / or 3D positional audio information. In embodiments, the gaming application 140 supports multiplayer gaming over a network 160. In embodiments, a voice chat stream from game participants is embedded in the audio stream 142 as a separate channel that is combined with existing in-game audio or mixed by the operating system. For example, a microphone 172 is used to record voice chat from participants. Although the gaming overlay application 150 is shown as receiving the audio stream 142 from the gaming application 140, in embodiments, the audio stream 142 is received from an audio mixer output provided by the operating system of the computing device 110.

[0013] In an embodiment, video stream 144 corresponds to in-game footage generated by GPU 122 and exposed for access via a video capture service provided by GPU 122. For example, completed frame buffer 123 is buffered in memory 130 for access by a video streaming application. For simplicity, game overlay application 150 is shown as accessing video stream 144 from game application 140.

[0014] In an embodiment, game overlay application 150 corresponds to any program that includes functionality for displaying an overlay over in-game video content, including programs provided by the manufacturer of GPU 122, such as Radeon Software Crimson ReLive Edition or GeForce Experience, game clients such as Steam with Steam Overlay, voice chat tools such as Discord, or operating system features such as Windows® Xbox Game Bar. In an embodiment, the game overlay application allows a user to enable options such as displaying an in-game overlay for configuring video capture, video streaming, audio mixing, voice chat, game profile settings, friends list, and other options.

[0015] In an embodiment, the game overlay application 150 includes functionality for video capture and audio capture and streaming. In an embodiment, this functionality is utilized to capture the audio stream 142 and the video stream 144 from the game application 140. In an embodiment, the game overlay application 150 is further extended to support automatic in-game subtitling by implementing or accessing a text conversion engine 152 and a subtitle synthesizer 154. In an embodiment, the text conversion engine 152 accesses the audio stream 142 and generates text corresponding to detected speech or sound effects. For example, the text conversion engine 152 includes a speech-to-text engine and a video game sound effects detection engine. Examples of speech-to-text engines include DeepSpeech, Wav2Letter++, OpenSeq2Seq, Vosk, and ESPnet. By using alternative models trained on video game sound effects and other non-dialogue speech cues, the speech-to-text engine is adaptable for use as a video game sound effects detection engine.

[0016] In an embodiment, to provide real-time or near real-time processing, the audio stream 142 is loaded into a buffer of limited size for processing by the text conversion engine 152. For example, the buffer is limited in maximum size or length, such as 5 seconds or less, and the buffer is split on the fly according to pauses or breaks detected in the audio stream 142. In this manner, the dialogue is processed in buffers containing short dialogue phrases and processed for display as quickly as possible.

[0017] In an embodiment, once the subtitle text is obtained from the text conversion engine 152, the subtitle synthesizer 154 determines a display location associated with the subtitle. For example, in an embodiment, the user preferences 158 define a preferred area of ​​the screen for displaying the subtitle, such as near the bottom of the screen. In an embodiment, the video stream 144 is scanned for user interface elements of the game application 140, such as health indicators or other in-game indicators that are preferably kept unobscured, and these areas are marked as excluded regions or no-go zones where the subtitle should not be displayed. For example, computer vision models are used to detect common video game user interface elements, such as health indicators, mini-maps, compasses, quest arrows, ammo and resource counters, ranking or score information, timers or clocks, and other heads-up display (HUD) elements. In an embodiment, the subtitle synthesizer 154 positions the subtitle in proximity to an in-game object associated with the in-game speaker, as described below in conjunction with FIG. 2C. In embodiments, to determine the identity of an in-game speaker, voices detected in audio stream 142 are matched against machine learning classifications stored in voice profile database 156. In embodiments, spatial audio cues from audio stream 142 are utilized to triangulate the location of in-game objects associated with the in-game speaker.

[0018] Although the text conversion engine 152 and the voice profile database 156 are shown as integral to the gaming overlay application 150, in an embodiment, the components of the gaming overlay application 150 are implemented by remote services (e.g., cloud servers) accessed over a network 160. This allows various tasks, such as text conversion, foreign language translation, and / or machine learning matching tasks, to be offloaded to external cloud services.

[0019] After the subtitle synthesizer 154 determines the display location of the subtitles generated from the text conversion engine 152, the subtitle overlay 190 is generated accordingly. Display characteristics of the subtitles, such as font color and size, are set according to one or more of user preferences 158, readability considerations, or speaker intent detected from the audio stream 142, as described further herein. To combine the subtitle overlay 190 with the corresponding portion of the video stream 144, the subtitle overlay 190 is blended with data from one or more frame buffers 123 that is finalized prior to output to the display 180, for example, as one or more processing steps in a rendering pipeline in the GPU 122 or by a desktop compositor of an operating system running on the computing device 110. In this manner, subtitle support is provided via the game overlay application 150, even if the game application 140 does not natively support subtitles.

[0020] III. Illustrative Game Application Graphical User Interface 2A, an exemplary display 280A is shown that corresponds to display 180 of FIG. As shown in display 280A, game graphics 282 are shown that correspond to game graphics 182. Display 280A represents a display of game application 150 when subtitle overlay 190 has not been generated or is disabled, or when game overlay application 140 is not running. In these cases, no subtitles appear and only in-game elements are shown, including character 284A positioned on the left side of display 280A, character 284B positioned on the right side of display 280A, and user interface elements 286 that display game play status, including the user's health and ammunition.

[0021] 2B, an exemplary display 280B is shown, which corresponds to display 180 of FIG. As shown in display 280B, subtitle overlay 290B is overlaid on top of game graphics 282 and includes the subtitles "(explosion from right)" and "This is bad. Let's take the left hallway instead." Note that subtitle overlay 290B is positioned near the bottom of display 280B, which in an embodiment is set according to user preferences 158. Note further that subtitle overlay 290B avoids placing subtitles over user interface elements 286, thereby maintaining visibility of important in-game information.

[0022] Referring to FIG. 2C, an exemplary display 280C is shown, which corresponds to display 180 of FIG. 1. As shown in display 280C, subtitle overlays 290C and 290D are overlaid on top of game graphics 282. Subtitle overlay 290C includes a subtitle that reads, "This is bad. Let's take the left hallway instead." Additionally, subtitle overlay 290C is positioned proximate to an in-game object (e.g., character 284A) associated with an in-game speaker and appears within a speech bubble. Subtitle overlay 290D includes a closed caption "(explosion sound)" and is positioned proximate to the right side of display 280C. In this example, subtitle overlay 290D points off-screen because it was determined that the explosion itself occurs to the user's right, not visible within game graphics 282.

[0023] In an embodiment, the location of the audio source in the game world is estimated according to position cues in the audio stream 142. For example, the panning position of the stereo sound is used to determine whether the audio source is located to the left, right, or center of the user's current viewpoint in the game world, as represented by the video stream 144. When multi-channel or positional 3D audio is available, the location of the audio source is estimated with greater accuracy, such as in front of, behind, above, or below the user's current viewpoint. In an embodiment, referring to FIG. 1, the multi-channel or positional 3D audio in the audio stream 142 indicates that the current in-game speaker is heard primarily from the left channel of the speaker 174. Thus, the in-game object associated with the in-game speaker is likely to be the character 284A on the left side, rather than the character 284B on the right side. Similarly, the audio stream 142 indicates that the explosion sound is heard primarily from the right channel of the speaker 174. However, since no explosion graphic is detected in the video stream 144, the explosion itself is determined to be off-screen and further to the right. These positional audio cues are factors used to determine the positioning of subtitle overlays 290C and 290D within the display so that they are in proximity to the in-game objects associated with their sound sources or in-game speakers. For example, sounds heard primarily from the center or rear surround channels indicate sound sources that are positioned in front-center or behind the user in the game world rendered by game application 140, while sounds heard primarily from the height channels indicate sound sources that are positioned above the user.

[0024] (IV. In-game automatic subtitle generation process) To illustrate an example process for implementing automatic in-game subtitling in a game overlay application, flow diagram 300 of Figure 3 will be described with respect to Figures 1 and 2B-C. As discussed above, displays 280B and 280C reflect an example of display 180 after game overlay application 150 has generated subtitle overlay 190 for display with game graphics 182.

[0025] Flow diagram 300 illustrates an approach for implementing in-game automatic subtitling in a gaming overlay application. In an embodiment, blocks 302, 304, 306, 308, 310 are performed by one or more processors. In an embodiment, blocks 302, 304, 306, 308, 310 are performed by a single processor of a computing device, similar to FIG. 1. In an embodiment, one or more of the blocks of flow diagram 300 are performed by one or more cloud servers or other computing devices distributed across a wireless or wired network.

[0026] At block 302, audio stream 142 and video stream 144 generated as a result of executing game application 140 are accessed. In an embodiment, a game overlay application executing on a processor receives the audio and video streams. In an embodiment, the processor executes game overlay application 150 concurrently with the game application. In some embodiments, game application 140 executes on a remote server. For example, when using a cloud-based game streaming service, audio stream 142 and video stream 144 are received from the remote server over network 160.

[0027] At block 304, the audio stream 142 is processed through a text conversion engine 152 to generate at least one subtitle. As described above, in an embodiment, the text conversion engine 152 is part of the game overlay application 150, while in other embodiments, the text conversion engine 152 is accessed using a cloud-based service via the network 160. Alternatively, both a cloud-based and an internal text conversion engine 152 are provided, with the internal version being utilized when the network 160 is unavailable or disconnected. In an embodiment, the text conversion engine 152 supports translation of text into the user's preferred native language and local dialect, as defined in user preferences 158. Because the translation function requires significant processing resources, in an embodiment, offloading the text conversion engine 152 to a cloud-based service helps minimize processing overhead that is detrimental to the performance of the game application 140.

[0028] At block 306, a display location is determined to associate with at least one subtitle from block 304. In an embodiment, the subtitle compositor 154 determines the display location using one or more factors. One factor includes a user-defined preference for subtitle location, such as near the bottom of the screen. This user preference is retrieved from user preferences 158. Another factor includes avoiding excluded regions detected in the video stream 144. For example, as described above, the video stream 144 is scanned for user interface elements generated by the gaming application 140, and portions of the display that include these user interface elements are marked as excluded regions that should not include subtitles.

[0029] Yet another factor involves positioning the subtitles in proximity to the source of the sound or in-game speaker. For example, computer vision processing is performed to identify in-game characters, multi-player users, and other objects within the video stream 144 that are potential sources of sound associated with the subtitles or closed captions. Once the characters and objects are identified, at least one subtitle from block 304 is matched to the most likely source of the sound and positioned in proximity to that source within the video stream 144.

[0030] Matching the at least one subtitle to the most likely sound source is based on various considerations. As described above, in an embodiment, the matching is based on triangulation using spatial audio cues from the audio stream 142. Thus, in-game objects (e.g., characters) located in the in-game world that match the spatial audio cues are more strongly correlated with the sound source.

[0031] Another consideration includes matching the voice characteristics with classifications in the voice profile database 156 and ascertaining whether the matched classification matches the visual characteristics of the potential audio source. For example, the voice profile database 156 includes classifications such as age range, gender, and dialect. Using machine learning techniques, the characteristics analyzed from the audio stream 142 and matched with the voice profile database 156 are used to classify the in-game speaker as likely or unlikely to be a child, adult, elderly, male, female, or a speaker with a regional dialect. The computer vision processing described above is used to ascertain whether the potential audio source or in-game character matches the matched classification. For example, if the audio stream 142 is classified as likely to be "female" in the voice profile database 156 and computer vision processing of the video stream 144 identifies the potential in-game character as likely to be a female character, matching the potential in-game character to at least one subtitle will be more strongly correlated.

[0032] Further considerations include matching the audio stream 142 to a particular user. For example, as described above, in an embodiment, the game application 140 is a multiplayer game in which participants use voice chat to communicate with other participants. In this case, the audio stream 142 includes multiple voice chat streams associated with particular users, and therefore the user speaking at any given time is easily determined according to the original voice chat stream. If the audio stream 142 is only available as a single mixed stream, the other considerations described above can still be used to determine the in-game speaker. Furthermore, since the game overlay application 150 includes identifying information, such as a username or handle, for each participant, the subtitles will include such identifying information, if available.

[0033] In block 308, a subtitle overlay 190 is generated that includes at least one subtitle from block 304 positioned at the associated display location by block 306. As described above, the subtitle compositor 154 generates the subtitle overlay 190 with various visual characteristics of the subtitle. In an embodiment, these visual characteristics include font attributes (e.g., italics, bold, outline), font color, font size, and callout type. Callout type includes, for example, callout, floating text, or other text presentation methods. The visual characteristics are set according to user preferences 158, such as the font size and color that the user prefers. The visual characteristics are set according to readability considerations, for example, by ensuring that the subtitle has high contrast according to the colors in the associated region of the video stream 144. For example, if the subtitle is positioned in an area that has predominantly vivid or bright colors, the subtitle uses a darker color or a darker outline to enhance visibility and readability. Visual characteristics are also set according to the in-game speaker, for example by mapping a specific font color to each in-game character.

[0034] In an embodiment, the visual characteristics are set according to the speaker's intent detected from the audio stream 142. For example, the audio stream 142 is analyzed for loudness, speaking tempo, syllable strength, voice pitch, and other factors to determine if the in-game speaker is calm, in which case the display characteristics use default values. On the other hand, if the analysis of the audio stream 142 determines that the in-game speaker is excited or conveying an urgent message, the display characteristics highlight this by using a bold font, a larger font size, or a speech bubble that is highlighted using a pointy line or other visual indicator. Thus, the speaker's intent is better understood in a visual manner.

[0035] In block 310, a portion of the video stream 144 is displayed with a subtitle overlay 190. In an embodiment, as described above, this is accomplished by modifying the rendering pipeline in the GPU 122 or by using the operating system's desktop compositor, among other methods. Thus, the display 180 outputs the game graphics 182 with the subtitle overlay 190. As shown in FIG. 2B, the subtitle overlay 290B is positioned according to a user preference for subtitle placement. Alternatively, as shown in FIG. 2C, the subtitle overlays 290C and 290D are positioned according to their proximity to the sound source. In this way, subtitle support is provided via the game overlay application 150, even if the game application 140 does not natively support subtitles.

Claims

1. accessing multi-channel or positional audio and video streams generated by one or more running gaming applications; generating at least one subtitle based on the multi-channel or positional audio stream in a text conversion engine circuit; analyzing the multi-channel or positional audio stream to identify a speaker associated with the at least one subtitle; locating an object associated with the speaker by triangulation from the multi-channel or positional audio stream; determining a display location proximate to an object associated with the speaker; generating a subtitle overlay including the at least one subtitle positioned at the display location; and displaying a portion of the video stream with the subtitle overlay. method.

2. The method includes loading the multi-channel or positional audio stream into a size-limited buffer for real-time or near real-time processing.

10. The method of claim 1.

3. determining the display position includes analyzing the video stream for an exclusion region that includes a user interface element of the one or more running applications; 10. The method of claim 1.

4. The speaker is an in-game speaker.

10. The method of claim 1.

5. analyzing the multi-channel or positional audio stream to identify the speaker includes matching at least one characteristic of the speaker with an associated classification in a voice profile database.

10. The method of claim 1.

6. The at least one characteristic of the speaker includes age, gender, and dialect. The method of claim 5.

7. determining the display position includes processing the video stream with computer vision to identify the speaker; 10. The method of claim 1.

8. generating the subtitle overlay includes configuring one or more visual characteristics of the at least one subtitle; 10. The method of claim 1.

9. the one or more visual characteristics include at least one of a font attribute, a font color, a font size, and a callout type; 9. The method of claim 8.

10. the one or more running applications is a multiplayer game, and the multi-channel or positional audio stream includes voice chat from participants in the multiplayer game.

10. The method of claim 1.

11. determining the display position includes accessing stored user preferences for subtitle positioning; 10. The method of claim 1.

12. 1. A system comprising: comprising one or more processors configured to perform the method of any of claims 1 to 11, system.

13. One or more computer-readable storage media containing instructions executable by one or more processors, The instructions cause the one or more processors to perform a method according to any one of claims 1 to 11. A computer-readable storage medium.