Audio control system and method
Patent Information
- Application Number
- JP2023006652
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-01-27
- Filing Date
- 2023-01-19
- Publication Date
- 2025-12-04
AI Technical Summary
Audio tracks in audio-visual entertainment, such as video games, often interfere with each other, causing acoustic confusion and detracting from the user experience, particularly when dialogue is obscured by music or other audio elements.
A system with an event detection unit, separation unit, and audio output unit dynamically adjusts audio elements by detecting significant events and performing source separation, ensuring key audio characteristics are preserved while minimizing interference.
The system enhances user comprehension and immersion by allowing important dialogue and audio elements to be heard clearly, reducing confusion and maintaining a seamless audio-visual experience.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to systems and methods for processing audio within audio-visual entertainment, and more particularly to systems and methods for processing audio within a video game environment. Computer programs, systems, and devices for implementing this method will also be described.
Background Art
[0002] Audio-visual entertainment such as movies and video games combines a large number of audio, visual contents, and / or sensory contents to provide a multimedia experience to users. Each sensory medium uses a vast array of assets to create and provide the desired user experience. For example, a video game environment involves various audio tracks such as background music, character dialogues, a wide range of sound effects, etc. Video games may be played by multiple players connected via a network, or conversations may be carried out through audio voice chat.
Summary of the Invention
Problems to be Solved by the Invention
[0003] Each audio track is used for the purpose of enhancing immersion or creating a suitable combination to form the overall user experience. On the other hand, a part or all of the audio materials may interfere with each other. For example, when a video game environment involves music with vocal tracks, the conversations within the game (or related to the game) may become unclear due to the words in the vocal tracks. Such conversations and simultaneous (or nearly simultaneous) audio may cause acoustic confusion to the user. This may lead to the user missing one or more music or conversations and feeling bad.
[0004] When dialogue is present in a game environment or other audio-visual media, the volume of music and other audio elements can be reduced during the mixing process. This technique is known as ducking. Ducking is a relatively forceful solution, and it can cause the volume of music tracks to constantly fluctuate. This can impair immersion and the overall user experience, potentially making the user feel uncomfortable.
[0005] Therefore, it is desirable to provide users with an audio-visual experience that maintains a complete audio-visual experience while allowing them to hear and understand important conversations and audio elements. [Means for solving the problem]
[0006] In a first aspect, the disclosure provides a system comprising an event detection unit, a separation unit, and an audio output unit. The event detection unit is configured to detect significant events related to a video game environment and to selectively output indications of the detected events. The separation unit is configured to perform source separation of playback music in response to indications from the event detection unit. The audio output unit is configured to output playback audio derived from the results of source separation by the separation unit.
[0007] By detecting critical events and performing source isolation on specific elements of the audio experience, pre-generated audio can be dynamically adapted to other audio elements in (or related to) the video game environment. This allows for the seamless removal of specific elements of the music that interfere with other audio elements in the game, while preserving the key characteristics of the music being played. Furthermore, the system can be configured to dynamically detect and adjust specific audio elements in real time, providing an optimal solution for any pre-generated music. Consequently, there is no need to adjust the music scene by scene.
[0008] A significant event within a video game environment may be (or related to) any number of other events. For example, dialogue such as voiceovers, narration, voice tracks in other secondary music, and dialogue from game-related voice chat may be considered significant events. Accordingly, an event detection unit may be configured to detect dialogue related to the video game environment as a significant event. As disclosed herein, music and other audio tracks can obscure dialogue being played, potentially causing confusion (and other undesirable effects) for the user. By detecting the presence of dialogue in (or related to) the game environment and altering the characteristics of other audio elements (e.g., music being played) through source isolation, the audio experience can be dynamically adjusted to improve dialogue comprehension while providing a desired user experience.
[0009] If an important event involves dialogue, dialogue related to the video game environment may include dialogue that has an audio source within that video game environment. For example, an event may be related to dialogue spoken by in-game characters or voiceover narration. An event may also be related to pre-recorded dialogue by the user or other users. In some examples, dialogue related to a video game may include audio from voice chat related to (or associated with) the video game environment. Dialogue from voice chat should be understood as user-generated audio (e.g., user speech, not computer-generated audio). Generally, such dialogue has an audio source in a microphone connected to a computer or gaming system. Such a microphone is configured to receive the user's voice and transmit that voice over a network or other means.
[0010] If different languages are used for character dialogue and song vocals, the user is less likely to be confused. The system may further include a language detection module configured to detect the language of dialogue detected by a detection unit. The language detection module may also be configured to detect the language of the music being played or the language of a separate vocal track. In such cases, the separation unit may be configured to perform sound source separation depending on the language detected by the language detection module. For example, suppose the language detection module detects character dialogue in a first language. If the music being played is also sung in the first language, then sound source separation may be performed only on the music being played (and furthermore, the volume of the vocal track may be lowered, for example). On the other hand, if the music being played is sung in a second language different from the first language, sound source separation may not be performed. Alternatively, sound source separation may be performed, but a different effect may be applied to the vocal track than if the same language were used for the dialogue and music. Such language information may be included as part of the indication from the detection unit.
[0011] Selectively, instead of detecting the language, the language of the lyrics may be indicated, for example, using metadata. On the other hand, languages that may potentially interfere with in-game conversations may also be indicated by metadata or user selection (if conversations in multiple languages are possible). Meanwhile, the language spoken by the user may, as a first approximation, be inferred from pre-selected language settings through the system user interface. For words entered by other users, similarly inferred languages may be sent as metadata at least once per line of dialogue. Selectively, the language detection module may use such metadata or contextual indicators as an alternative to or complement to language detection.
[0012] The indicators output from the event detection unit may include audio characteristics related to the detected important event. For example, the indicators may include signals that convey data about audio characteristics related to the detected important event. For instance, if the detected important event is a game-related conversation, the audio characteristics may include, for example, one or more pitches, durations, and volumes, and may also include information about the character who spoke the conversation, information about the event, and the state of the in-game environment. The isolation unit may be configured to perform source separation according to the audio characteristics. That is, the method of performing source separation may be influenced by certain characteristics of the detected event (e.g., the vocal characteristics of the detected conversation). For example, if a low-pitched conversation is detected in a low-frequency audio signal, source separation for music may be performed by separating and filtering (e.g., removing or lowering the volume) the low-frequency voice track, while retaining and leaving the high-frequency voice track unchanged. Alternatively, user confusion with low-frequency conversations may be reduced by separating and modifying the low-frequency layer of music.
[0013] The signal may characterize the duration of the detected event. This allows source separation to be performed only while the event is ongoing, or during overlapping "dubtail" intervals with lead-up and lead-out periods (for example, if the event lasts 5 seconds, the music can be modified for 7 seconds, including 1 second before the event starts and 1 second after it ends). Thus, the signal may be sent only once per event (for example, at the start of the event or before the event starts). Alternatively, the signal may be a continuous signal indicating that events are occurring (or not occurring) consecutively. For example, the signal may be present while an in-game character is speaking and stop as soon as the character stops speaking.
[0014] The separation unit may be configured to perform source separation to separate one or more vocal tracks from the music being played. The separation unit may also be configured to lower the volume of one or more vocal tracks in response to indications from the event detection unit. The separation unit may also be configured to modify the audio characteristics of the separated one or more vocal tracks in response to indications from the event detection unit. In another example, the separation unit may be configured to perform source separation to separate one or more other tracks (tracks other than vocal tracks) from the music being played.
[0015] The audio output unit may be configured to generate multiple audio channels based on the source separation results of the separation unit and to output multi-channel audio for playback. Source separation may generate multiple playable tracks from multiple original playback source music. In a simple example, music may be separated into instrument tracks and vocal tracks. The audio output unit may send the instrument tracks to one channel and the vocal tracks to another. This may be done before or after modifications (e.g., volume reduction, pitch adjustment) are made to each track. Thus, the audio output unit may be configured to output multi-channel audio for playback. This multi-channel audio includes vocal tracks separated by channel and dialogue from the game in other channels. Such multi-channel output can give the user a choice of how each track is output. For example, the audio output unit may be configured to send dialogue from the game to one channel and vocals from the music to another channel. The vocals and dialogue may be played from different speakers. This reduces the chance of confusion.
[0016] The audio output unit may be configured to perform quality checks on the corrections made by the separation unit. For example, the audio output unit may be configured to detect artifacts in the results of source separation and adjust the audio characteristics of the output audio for playback.
[0017] The system may further include an audio input unit configured to identify or generate music being played. The system may also include a microphone. The microphone may be part of the audio input unit or a separate peripheral device. The microphone may be positioned to receive user-inputted dialogue. An event detection unit may be configured to detect when dialogue is input via the microphone as an important event. This would be useful, for example, if audio is output through a speaker (placed near the microphone, and the audio is acquired via the microphone) and the user is recording their dialogue via the microphone for voice chat. If the audio being played includes voice, the recorded user dialogue might be confused with the voice in the playing audio. When it is detected that the microphone is being used, the volume of the voice in the audio may be lowered or modified. This reduces the risk of confusion with the recorded dialogue. Alternatively, the event that generates new audio may be delayed or modified. This prevents the playing sound from interfering with the currently input voice.
[0018] Certain types of events cause more severe auditory confusion for the user compared to others. In this regard, the system may further include a user recognition unit. This user recognition unit is configured to detect one or more events that cause confusion for the user in the video game environment. The user recognition unit may be configured to monitor the user's behavior in response to trigger events in the game environment and to detect the events that cause confusion. For example, if it is clear that the user missed a command from a voiceover narration, the user recognition unit will detect this. In some examples, the user recognition unit employs a machine learning model to learn the user's behavior and trains the model to predict the types of events that cause confusion. Such events that cause confusion are then designated as important events and are subject to detection by an event detection unit. Recognition of the type of event that causes confusion can be performed concurrently with (or prior to) the detection of specific audio elements associated with a particular event and source isolation (some events may be pre-configured before the user recognition unit has a chance to learn or predict them). In another example, the user recognition unit may be configured to generate user input requests in (or related to) the game environment. This allows users to flag specific events as difficult to hear. For example, the dialogue of a particular character (e.g., a character with a specific voice pitch) may be particularly difficult to hear when it is overlaid with music. By recognizing specific important events that are more likely to cause confusion, the system can adjust the music (or other elements) to mitigate these problems. Thus, the user recognition unit may be configured to manipulate the detection unit so that the detection unit detects events that cause confusion for the user as important events. This would be particularly useful for users with partial hearing loss or significant hearing loss in one ear, for example, because these users may have difficulty hearing in certain frequency ranges due to audio interference, or when the sound is located to the left or right.Therefore, the system may learn one or more frequency domains and / or one or more directional (or angular) domains to be confusing.
[0019] In a second aspect, the disclosure provides a method, which includes the steps of detecting important events related to a video game environment, performing source separation on playback music in response to the detection of important events, and outputting playback audio derived from the results of source separation.
[0020] It goes without saying that one or more features relating to the first embodiment described above can also be applied to the second embodiment. For example, the method of the second embodiment may include a step of performing any function relating to the system of the first embodiment. In this case, the same or similar technical advantages can be obtained.
[0021] In a third aspect, the Disclosure provides a computer program that includes instructions for a computer in an audio-visual entertainment system. The computer controls the audio-visual entertainment system by means of these instructions so that the audio-visual entertainment system performs the method of the second aspect. [Brief explanation of the drawing]
[0022] The present invention will be described with reference to the attached drawings. [Figure 1] This is a schematic diagram of an audio-visual entertainment system in which the method according to the present invention is implemented. [Figure 2] This is a schematic diagram of an example of a system in an assembled configuration. [Figure 3] This is a schematic diagram illustrating an example of a method for modifying audio characteristics related to audio-visual entertainment. [Modes for carrying out the invention]
[0023] Certain aspects of the present disclosure relate to a system for adjusting audio characteristics within (or related to) audio-visual entertainment. FIG. 1 shows a typical environment within such an audio-visual entertainment system.
[0024] As an example, multimedia environment 1 includes various sound sources. Each of these sound sources can generate sound that is reproduced via an audio output and conveyed to a user. Such an environment can be, for example, a scene within a video such as a movie or a television program, or a video game environment.
[0025] In this example, the scene is within a video game and includes a character speaking (uttering dialogue 4 from the mouth), an animal barking (producing bark 5), leaves producing ambient sound 6, and a weather phenomenon (e.g., thunder) producing thunderclap 7. Additionally, background music 2 is being played, and a voice chat 3 related to the game environment is active.
[0026] Background music 2 is related to the game environment. Typically, background music 2 includes one or more instrumental elements (e.g., strings, percussion, brass) and one or more vocal elements. Also, generally, background music 2 is loaded as a pre-mixed single track. Background music 2 may be pre-recorded, stored, and / or loaded as part of the game environment. Alternatively or additionally, background music 2 may be stored in another location (e.g., within a user's personal data storage separate from the game storage or within a database accessible via a network). Background music 2 may also be generated procedurally, for example, using a neural network.
[0027] Character dialogue 4 includes lines of dialogue. Generally, dialogue is derived from pre-recorded audio (e.g., voiceover artists) and can be played back along with the character model's animation. Dialogue can also be generated step-by-step using machine learning models, such as neural networks. In this case, the dialogue may be generated according to input factors (e.g., events occurring in the game world or user actions).
[0028] Dialogue is usually in a single language, but it can sometimes include multiple languages. For example, the first part of a line of dialogue might be in the first language, the second part in the second language, and then it might return to the first language or move to a third language that is different from the first and second languages.
[0029] In the example scene in Figure 1, the character is visible within the frame, but in another example, conversation 4 may be narration, and the character (or voiceover actor) may not be visible within the scene. Game environment 1 may include multiple characters engaging in various different conversations 4.
[0030] The environment may also include other sound sources. For example, Scene 1 includes animals that emit barks and various other animal-related sound effects 5. Similar to character dialogue 4, such animal sound effects may also be accompanied by animations of the animal's visual model. Generally, animal sound effects 5 are not dialogue or lines, but may be treated as such. In some examples, animal sound effects 5 may contain conversational elements and / or may be full voiceover (or procedurally generated) human language.
[0031] Other elements of the scene (such as weather or leaves) can also generate audible sound effects. For example, rain, wind, and the rustling of leaves all contribute to these ambient sound effects.6 While these ambient sound effects6 generally do not include spoken language, they may, like animal sound effects, include conversational elements that the user may perceive as dialogue. Weather sound effects (such as rain) can be constant and last for a set period of time, or they can be one-time events or responses to user actions. For example, thunder can be a temporal event, a response to user actions, or it can be accompanied by a one-time sound effect such as thunderclap7. Examples of one-time sound effects7 include gunshots fired by the user and footsteps when the user moves in the game. These one-time sound effects may include dialogue or may sound like dialogue.
[0032] As described above, video game environment 1, as an example, relates to audio from voice chat 3. The video chat function may be built into the video game environment or it may be a standalone process that runs concurrently with the game. Voice chat provides audio derived from conversations entered by the user. Typically, such dialogue is collected by an input device (e.g., a microphone or other related peripheral connected to the game system). Generally, voice chat audio 3 includes the user's dialogue or conversation.
[0033] Figure 2 is a schematic block diagram of system 10 as an example. System 10 is configured to modify certain audio characteristics related to an audio-visual entertainment system (for example, as described with reference to Figure 1 above). The features in Figure 2 are embodied in a video game system such as a console or computer. In this example, the system is part of a video game configuration mechanism relating to a gaming system 20, an audio system 30, and a connected microphone 40.
[0034] As an example, gaming system 20 is a video game console configured to generate graphics, audio, and control functions for loading a game environment (e.g., as shown in Figure 1). Audio system 30 can take the processing results of system 10 and output the final audio through multiple speakers 31, 32. In Figure 2, system 10 is shown separately from gaming system 20, but they may be integrated as part of a single system, or integrated with audio system 30 and microphone 40. System 10 may also be part of a cloud network connected to multiple gaming systems 20 via a network connection, and may be configured to provide remote audio processing to multiple such gaming systems 20.
[0035] As an example, system 10 is configured to modify audio related to an audio-visual entertainment system and includes an event detection unit 11, an isolation unit 12, and an audio output unit 13. In this example, system 10 is configured to modify audio related to a video game environment.
[0036] The event detection unit 11 is configured to detect important events related to the video game environment. The purpose of the event detection unit 11 is to identify when audio corrections are needed to mitigate user audience confusion. The event detection unit 11 in this example is also configured to selectively output indicators of detected events. Such indicators may be utilized by other units in the system 10 to perform relevant processing only when necessary and to an appropriate (or desirable) degree.
[0037] As an example, the separation unit 12 within system 10 is configured to perform source separation in the music being played. Typically, source separation refers to the process of separating one or more source signals from one or more mixed signal sets. This process is usually performed without access to information about the source signals or knowledge of the mixing details. This process can be used to reconstruct the original signal set (which consists of individual tracks such as one or more instrument tracks or vocal tracks). Generally, music 2 played within an audio-visual scene 1 is labeled as such. Therefore, the separation unit 12 can simply take the music being played as input and perform appropriate source separation. The system may also include an audio input unit configured to identify or generate the music being played. For example, if the music is part of ambient sound 6 (e.g., a radio playing in the background) or plays simultaneously with dialogue (e.g., a song sung a cappella or with instrumental accompaniment by a character), the audio input unit can be configured to identify the music currently being played. In this case, the separation unit can perform source separation on the music identified by the audio input unit. The audio input unit may also be configured to generate new music. In this example, source separation is performed on the music being played, but there are also examples where the separation unit 12 is configured to perform source separation on other acoustic objects in the environment (e.g., conversation 4 or ambient sound 6).
[0038] The audio output unit 13 is configured to output the playback sound derived from the sound source separation result by the separation unit 12. In other words, the audio output unit 13 is configured to acquire the audio output from the separation unit 12 and configure the audio for playback.
[0039] When the separation unit 12 performs sound source separation and generates one or more separated audio tracks, these separated audio tracks can be modified. The modified tracks can then be combined with other tracks to generate the playback audio to be output. The separation unit 12 itself can perform any modifications on the separated audio tracks. In contrast, the audio output unit 13 is typically configured to perform specific modifications on the separated audio tracks. For example, the audio output unit 13 can be configured to identify a separated vocal track and reduce its volume (or mute it). The audio output unit 13 can also be configured to split the sound into multiple channels and play them on different speakers (e.g., right speaker 31 and left speaker 32).
[0040] During use, the gaming system 20 generates a game environment 1 with multiple sound sources (e.g., background music 2 during playback). When the detection unit 11 detects an important event (e.g., audible dialogue 4 spoken by a game character), the separation unit 12 acquires the background music 2 and performs sound source separation in real time. The separation unit 12 separates one or more vocal tracks from the background music 2. The audio output unit 13 lowers the volume of the vocal tracks so that dialogue 4 can be heard. In this example, the detection unit 11 indicates to the separation unit 12 that dialogue 4 will continue. For this continuation, the volume of the vocal tracks in the music is lowered. The audio output unit 13 acquires the results of sound source separation and volume reduction and generates an output mix. This output mix is the output for playback and replaces the original background music 2 (at least while the dialogue is continuing). The audio is output via the sound system 30. If the audio output unit 13 generates a multi-channel output mix, each channel can be output via separate speakers 31, 32.
[0041] Figure 3 is a schematic flowchart illustrating the steps of an example method relating to this disclosure.
[0042] In step S110, a significant event related to the video game environment is detected. In some examples, the significant event to be detected is a conversation or other audible conversation within the video game environment. In some examples, the significant event is a conversation or other audible conversation within a video chat related to the video game environment. This step can be performed using a detection unit of the type described with reference to Figure 2. In other words, step S110 can be performed by configuring or controlling the detection unit 11 to detect significant events related to the video game environment. In some examples, this step may include a substep of generating a signal representing the detected significant event. Such a signal may include information indicating the characteristics, duration, and other details of the detected event. In some examples, the signal continues to be output while the significant event is occurring. For example, if a character conversation is considered a significant event, a signal indicating this conversation is generated and output the moment a character conversation is detected within the game environment. This signal then continues to be output until the time when the conversation is determined to have ended (e.g., by the detection unit). The signal stops at this end time. In another example, the first "start" signal pulse may be output at the beginning of the event, and the second "stop" signal pulse may be output at the end of the event.
[0043] In step S120, source separation is performed on the playback music. Source separation is performed in response to the detection of a significant event in step S110. In particular, the playback music may be separated into at least one vocal track and at least one instrument track. In some examples, this step may include a substep of receiving a signal indicating the detected event and performing source separation in response to the received signal. Source separation may continue to be performed as long as the detected event persists. For example, if a significant event such as character dialogue is detected in step S110 and the dialogue lasts for 20 seconds, selectively, source separation may be performed for the 20 seconds of the dialogue. Alternatively, the time for which source separation is performed may be dovetailed (or offset) around the duration of the event. This step can be performed using the type of separation unit described with reference to Figure 2. In other words, step S120 can be performed by configuring or controlling the separation unit 12 to perform source separation on the playback music in response to the indication of a detected significant event.
[0044] In step S130, the audio derived from the source separation result is output for playback. This step may further include a substep of receiving the source separation result from step S120 and modifying the audio characteristics of one or more tracks included in the received result. For example, if the source separation result includes a vocal track and an instrument track, the volume of the vocal track may be reduced in this step. This reduction in volume is intended to sustain the event detected in step S110. This step can be performed using an audio output unit of the type described with reference to Figure 2. In other words, step S130 can be performed by configuring or controlling the audio output unit 13 to output the playback audio derived from the source separation result.
[0045] Referring again to Figure 2, in the summary embodiment of this specification, the system includes an event detection unit, a separation unit, and an audio output unit. The event detection unit is configured to detect important events related to the video game environment and to selectively output signs of the detected events. The separation unit is configured to perform source separation of the playback music in response to signs from the event detection unit. The audio output unit is configured to output the playback audio derived from the results of source separation by the separation unit. -In one example of a summary embodiment, the event detection unit is configured to detect conversations related to the video game environment as important events. -In this example, selectively, conversations related to the video game environment include conversations that have sound sources within the video game environment. -In this example, selectively, conversations related to the video game environment include voice chat related to the video game environment. -In one example of a summary embodiment, the indicators output from the event detection unit include audio characteristics related to the detected significant event, and the separation unit is configured to perform source separation according to the audio characteristics. -In one example of a summary embodiment, the separation unit is configured to perform source separation in order to separate one or more vocal tracks from the music being played and to change one or more audio characteristics of the one or more vocal tracks in response to an indication from the event detection unit. -In this example, selectively, the isolation unit is configured to reduce the volume of one or more isolated vocal tracks in response to indications from the event detection unit. -In one example of a summary embodiment, the audio output unit is configured to generate multiple audio channels based on the result of sound source separation by the separation unit, and to output the generated multiple audio channels for playback. -In this example, the audio output unit is selectively configured to output multi-channel audio for playback, and the multi-channel audio is configured to include separate vocal tracks for each channel and dialogue from the game in the other channels. -In one example of a summary embodiment, the audio output unit is configured to detect artifacts included in the result of source separation and adjust the characteristics of the audio output for playback. -In one example of a summary embodiment, the system further includes an audio input unit configured to identify or generate music for playback. -In one example of a summary embodiment, the system further includes a user recognition unit, which detects one or more events that cause confusion to a user in a video game environment and is configured to operate a detection unit so that the detection unit detects the events that cause confusion to the user as important events. -In one example of a summary embodiment, the system further includes a microphone, and a detection unit is configured to detect when dialogue is input via the microphone as an important event. - An example of a summary embodiment is as described above.
[0046] Referring to Figure 3, in one example of a summary embodiment of this specification, the method includes the steps of: detecting important events related to a video game environment; performing source separation on playback music in response to the detection of important events; and outputting playback audio derived from the results of source separation.
[0047] It will be apparent to those skilled in the art that various modifications of the above method corresponding to the operations of the various embodiments described in the present specification and claims also fall within the scope of the present invention.
[0048] It will be understood that the above methods can be executed using ordinary hardware to which suitable software instructions can be applied, or (in addition to or instead of) dedicated hardware.
[0049] Implementation using existing parts of a typical equivalent device is possible in the form of a computer program product with a processor capable of executing instructions recorded on a non-temporary computer-readable medium (e.g., floppy disk®, optical disk, hard disk, solid disk, PROM, RAM, flash memory, or a combination thereof), or it can be done using hardware (e.g., ASIC (application specific integrated circuit), FPGA (field programmable gate array), or other configurable circuits suitable for a typical device). Such computer programs may be transmitted via data signals over a network (e.g., Ethernet®, wireless network, internet, or a preferred combination thereof).
[0050] The above discussion merely discloses and describes examples of embodiments of the present invention. Those skilled in the art will understand that the present invention can be realized in other specific forms without departing from the spirit and essential features of the invention. Accordingly, the disclosure of the present invention is for illustrative purposes only and is not intended to limit the scope of the invention or the claims. This disclosure, including any identifiable modifications of the above teachings, partially defines the scope of the terms of the claims. The subject matter of the invention is not dedicated to the public.
Claims
1. an event detection unit; A separation unit; an audio output unit; Including, the event detection unit is configured to detect significant events related to the video game environment and selectively output indicia of the detected events; the separation unit is configured to perform sound source separation on the played music in response to an indication from the event detection unit; The audio output unit is configured to output a reproduced audio derived from a result of the sound source separation by the separation unit.
2. The system of claim 1 , wherein the event detection unit is configured to detect conversations related to the video game environment as significant events.
3. 3. The system of claim 2, wherein the conversations related to the video game environment include conversations having audio sources within the video game environment.
4. 4. The system of claim 2 or 3, wherein the conversation related to the video game environment includes a voice chat related to the video game environment.
5. the indicia output from the event detection unit includes audio features associated with the detected significant event; The system of claim 1 , wherein the separation unit is configured to perform source separation according to the audio characteristics.
6. 2. The system of claim 1, wherein the separation unit is configured to perform sound source separation to separate one or more vocal tracks from a playback music stream and to alter one or more audio characteristics of the one or more vocal tracks in response to an indication from the event detection unit.
7. The system of claim 6 , wherein the separation unit is configured to reduce the volume of one or more separated vocal tracks in response to an indication from the event detection unit.
8. 2. The system of claim 1, wherein the audio output unit is configured to generate multiple channels of audio based on a result of the sound source separation by the separation unit, and output the generated multiple channels of audio for playback.
9. the audio output unit outputs multi-channel audio for playback; 10. The system of claim 8, wherein the multi-channel audio is configured to include a separate vocal track per channel and dialogue from the game in other channels.
10. The system of claim 1 , wherein the audio output unit is configured to detect artifacts in the results of the sound source separation and adjust characteristics of the audio output for playback.
11. 10. The system of claim 1, further comprising an audio input unit configured to specify or generate music to be played.
12. further comprising a user recognition unit; 2. The system of claim 1, wherein the user recognition unit is configured to detect one or more events that cause confusion to a user of the video game environment and to operate the detection unit such that the detection unit detects the events that cause confusion to the user as significant events.
13. further comprising a microphone; The system of claim 1 , wherein the detection unit is configured to detect input of dialogue via the microphone as a significant event.
14. detecting a significant event associated with the video game environment; performing sound source separation on the played music in response to detecting the significant event; outputting a playback audio derived from the result of the sound source separation; A method comprising:
15. 1. A computer program comprising instructions for a computer in an audio-visual entertainment system, comprising:
15. A computer program product for controlling the audiovisual entertainment system in accordance with the instructions, such that the audiovisual entertainment system performs the method of claim 14.