Audio analysis and accessibility across applications and platforms

User audio profiles with real-time audio modification improve audio accessibility by prioritizing and modifying sound characteristics, addressing the challenge of overwhelming audio environments for users with hearing impairments.

JP7808525B2Active Publication Date: 2026-01-29SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022128872
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-08-31
Filing Date
2022-08-12
Publication Date
2026-01-29
Estimated Expiration
2042-08-12

AI Technical Summary

Technical Problem

Users, especially those with hearing impairments or disabilities, struggle to navigate complex audio environments in digital content and social interactions across multiple platforms, leading to a diminished experience due to overwhelming sound combinations.

Method used

A user audio profile is created with custom audio parameter prioritizations, allowing real-time modification of audio streams based on user preferences, enhancing audio accessibility by prioritizing and modifying sound characteristics in real-time.

Benefits of technology

Enhances user experience by allowing individuals to focus on relevant audio streams, improving comprehension and enjoyment of digital content and social interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007808525000001
    Figure 0007808525000001
  • Figure 0007808525000002
    Figure 0007808525000002
  • Figure 0007808525000003
    Figure 0007808525000003
Patent Text Reader

Abstract

To provide improved systems and methods for audio analysis and audio-based accessibility.SOLUTION: A method performed by a processor of an electronic entertainment system includes the steps of: storing a user audio profile; monitoring one or more audio streams associated with a user device; analyzing the audio streams based on the user audio profile; and modifying the audio streams. The user audio profile includes custom prioritization of one or more audio parameters. At least one sound characteristic of the audio stream is modified in real time based on the prioritization of the audio parameters detected and ranked in the analysis of the audio stream before at least one audio stream is provided to the user device.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to audio analysis and, more particularly, to audio analysis and audio-based accessibility across applications and platforms. [Background technology]

[0002] Currently available digital content may include audiovisual and other types of data. Accordingly, interacting with various titles of such digital content may involve using a device to present such audiovisual and other data in a digital environment. For example, playing an interactive game title may involve presenting a variety of different audiovisual effects within an associated virtual environment. Audiovisual effects may include a soundtrack, score, background noises associated with the virtual environment (e.g., the in-game environment), sounds associated with virtual characters and objects, etc. During a gameplay session of an interactive game title, multiple different types of audio may be presented simultaneously. For example, an in-game scene may be associated with a musical score, ambient sounds associated with a particular location within the virtual environment, and noises spoken or otherwise made by one or more virtual characters and objects. Different users may find such sound combinations confusing, distracting, annoying, or otherwise undesirable.

[0003] Additionally, many users may enjoy playing such digital content titles in social settings that enable simultaneous interaction with other users (e.g., friends, teammates, opponents, spectators). Such social interaction (which may include voice or video chat) may occur across one or more different platforms (e.g., game platform servers, lobby servers, chat servers, other service providers). Thus, users may be presented with additional audio streams as they simultaneously attempt to decipher and understand the content-related audio. Furthermore, such sound combinations may overwhelm a user's ability to hear what is happening in the real-world environment. Some users (especially those with hearing loss or other conditions and disabilities affecting hearing or cognition) may find such situations difficult to navigate, thereby negatively impacting the user's enjoyment and experience with the interactive game title.

[0004] Therefore, there is a need in the art for improved systems and methods for audio analysis and audio-based accessibility across applications and platforms. Summary of the Invention

[0005] Embodiments of the present invention include systems and methods for audio analysis and audio-based accessibility across applications and platforms. A user audio profile may be stored in a user's memory. The user audio profile may include a custom prioritization of one or more audio parameters associated with one or more audio modifications. Audio streams associated with the user's user device may be monitored based on the user audio profile during a current session. Audio parameters may be detected as present in the monitored audio streams, and the detected audio parameters may be prioritized based on the custom prioritization of the user audio profile. At least one sound characteristic of at least one audio stream may be modified in real time based on the prioritization of the detected audio parameters by applying the audio modifications of the user audio profile to the at least one audio stream before the at least one audio stream is provided to the user device. [Brief explanation of the drawings]

[0006] [Figure 1] 1 illustrates a network environment in which a system for audio analysis and audio-based accessibility may be implemented. [Figure 2] 1 illustrates an exemplary Uniform Data System (UDS) that can be used to provide data to systems for audio analysis and audio-based accessibility. [Figure 3] 1 is a flowchart illustrating an exemplary method for audio analysis and audio-based accessibility. [Figure 4] FIG. 1 illustrates an exemplary implementation of audio analysis and audio-based accessibility features. [Figure 5] FIG. 1 is a block diagram of an exemplary electronic entertainment system that may be used with embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0007] Embodiments of the present invention include systems and methods for audio analysis and audio-based accessibility across applications and platforms. A user audio profile may be stored in a user's memory. The user audio profile may include a custom prioritization of one or more audio parameters associated with one or more audio modifications. Audio streams associated with the user's user device may be monitored based on the user audio profile during a current session. Audio parameters may be detected as present in the monitored audio streams, and the detected audio parameters may be prioritized based on the custom prioritization of the user audio profile. At least one sound characteristic of the audio stream may be modified in real time based on the prioritization of the detected audio parameters by applying the audio modifications of the user audio profile to the at least one audio stream before the at least one audio stream is provided to the user device.

[0008] FIG. 1 illustrates a network environment 100 in which a system for audio analysis and audio-based accessibility may be implemented. The network environment 100 may include one or more content source servers 110 that provide digital content for distribution (e.g., games, other applications, and services), one or more content provider server application program interfaces (APIs) 120, a content delivery network server 130, an audio analysis server 140, and one or more user devices 150. The devices in the network environment 100 communicate with each other using one or more communication networks, which may include dedicated local networks (e.g., intranets) and / or may be part of a larger wide-area network. A communication network may be a local area network (LAN) that may be communicatively coupled to a wide area network (WAN) such as the Internet. The Internet is a broad network of interconnected computers and servers that enables the transmission and exchange of Internet Protocol (IP) data between connected users via a network service provider. Examples of network service providers include the public switched telephone network, a cable service provider, a digital subscriber line (DSL) service provider, or a satellite service provider. One or more communication networks enable communication between the various components of network environment 100 .

[0009] The servers described herein may include any type of server known in the art, including standard hardware computing components such as network and media interfaces, non-transitory computer-readable media (memory), and a processor for executing instructions or accessing information that may be stored in the memory. The functionality of multiple servers may be integrated into a single server. Any of the foregoing servers (or integrated servers) may exhibit certain client-side, cache, or proxy server characteristics. These characteristics may depend on the particular network placement of the server or the particular configuration of the server.

[0010] Content source server 110 may maintain and provide a variety of digital content and services that are available for distribution over a communications network. Content source server 110 may be associated with any content provider that makes its content available for access over a communications network. Thus, content source server 110 may host a variety of different content titles that may be further associated with object data (e.g., activity information, zone information, character information, player information, other game media information, etc.) regarding digital or virtual objects that are displayed in a digital or virtual environment during an interactive session.

[0011] Such content may include not only digital video and games, but also other types of digital applications and services. Such applications and services may include any of a variety of different digital content and functionality that may be provided to user device 150, including providing and supporting chat and other communication channels. Chat and communication services may include voice-based, text-based, and video-based messaging. Thus, user device 150 may participate in one or more communication sessions and simultaneously in a gameplay session, and the gameplay and communication sessions may be hosted on one or more content source servers 110.

[0012] Content from content source servers 110 may be provided via content provider server API 120, which enables various types of content source servers 110 to communicate with other servers (e.g., user devices 150) in network environment 100. Content provider server API 120 may be specific to the particular operating language, system, platform, protocol, etc. of the content source server 110 providing the content, as well as the user devices 150 and other devices in network environment 100. In network environment 100 that includes multiple different types of content source servers 110, there may likewise be a corresponding number of content provider server APIs 120, which enable various formats, conversions, and other cross-device and cross-platform communication processes for providing content and other services to different user devices 150, each of which may process such content using different operating systems, protocols, etc. Accordingly, applications and services may be made available in different formats to be compatible with a variety of different user devices 150. In a network environment 100 that includes multiple different types of content source servers 110, content delivery network servers 130, conversion filter servers 140, user devices 150, and databases 160, there may likewise be a corresponding number of APIs managed by the content provider server API 120.

[0013] Content provider server API 120 may further facilitate each of user devices 150's access to content hosted by or services provided by content source server 110, either directly or via content delivery network server 130. Additional information, such as metadata about the accessed content or service, may also be provided to user device 150 by content provider server API 120. As described below, the additional information (e.g., object data, metadata) may be available to provide details about the content or service provided to user device 150. In some embodiments, services provided by content source server 110 to user device 150 via content provider server API 120 may include supporting services associated with other content or services, such as chat services, ratings, and profiles associated with particular games, teams, communities, etc. In such cases, content source servers 110 may also communicate with each other via content provider server API 120.

[0014] The content delivery network servers 130 may include servers that provide resources, files, etc. related to content from the content source servers 110, including various content and service configurations, to the user devices 150. The content delivery network servers 130 may also be invoked by user devices 150 requesting access to particular content or services. The content delivery network servers 130 may include universe management servers, game servers, streaming media servers, servers hosting downloadable content, and other content delivery servers known in the art.

[0015] The audio analysis server 140 may include any data server known in the art that is capable of communicating with the different content source servers 110, the content provider server API 120, the content delivery network server 130, the user devices 150, and the database 160. Such an audio analysis server 140 may be implemented in one or more cloud servers that execute instructions associated with interactive content (e.g., games, activities, videos, podcasts, user-generated content (“UGC”), publisher content, etc.). The audio analysis server 140 may further execute instructions, for example, to monitor one or more audio streams based on a user audio profile. Specifically, the audio analysis server 140 may monitor one or more audio parameters specified by the user audio profile. When an audio parameter is detected, the audio analysis server 140 may prioritize such audio parameters according to a custom prioritization specified by the user audio profile and modify at least one of the audio streams based on the prioritization.

[0016] The user device 150 may include several different types of computing devices. The user device 150 may be a server that provides internal services (e.g., to other servers) in the network environment 100. In such a case, the user device 150 may correspond to one of the content servers 110 described herein. Alternatively, the user device 150 may be a computing device, including any number of different game consoles, mobile devices, laptops, and desktops. Such user devices 150 may also be configured to access data from other storage media, such as, but not limited to, a memory card or disk drive, which may be appropriate for downloaded services. Such user devices 150 may include standard hardware computing components, such as, but not limited to, network and media interfaces, non-transitory computer-readable storage (memory), and a processor for executing instructions that may be stored in memory. These user devices 150 may also run a variety of different operating systems (e.g., iOS, Android®), applications, or computing languages ​​(e.g., C++, JavaScript®). An exemplary client device 150 is described in detail herein with respect to FIG. 5. Each of the user devices 150 may be associated with a participant (e.g., a player) or other type of user (e.g., a spectator) associated with a collection of digital content streams.

[0017] Although depicted separately, database 160 may be stored on the same server, a different server, or any of the servers and devices shown in network environment 100 of any of user devices 150. Such database 160 may store or link various sources and services used for audio analysis and modification. Additionally, database 160 may store user audio profiles, as well as audio analysis models that may be specific to a particular user, user group or team, user category, game title, game genre, sound type, etc. The models may be used to identify specific speech sounds (e.g., the user's voice, the voices of specific other users), which may serve as audio parameters that may be prioritized and subjected to audio modification. One or more user audio profiles may also be stored in database 160 for each user. In addition to gameplay data related to the user (e.g., the user's progress in an activity and / or media content title, user ID, user game character, etc.), the user audio profile may include a set of audio parameters specified or adopted by the user in a custom prioritization scheme.

[0018] For example, certain users with hearing loss may find it difficult to distinguish between multiple different types of sounds. In a movie or game context, the inability to understand speech (e.g., in a movie, in a game, or from other players in the current game session) may represent an accessibility barrier that prevents a user from fully understanding session events. Such a user may specify that dialogue be prioritized over other types of sounds (e.g., background noise, environmental (real-world or virtual) noise, musical score, or background music). Such prioritization may further be associated with audio modification, where higher priority streams (and associated audio parameters) are boosted relative to other simultaneous streams (and associated audio parameters) with lower priority.

[0019] In addition to sound type, other audio parameters may include specific user voices, the number of user voices speaking simultaneously, their relative locations (e.g., in-game or real-world), total and individual stream volume, frequency response (e.g., pitch), speed, and other detectable or measurable audio parameters known in the art. For example, a user may wish to prioritize communications from teammates or opponents with whom they may be playing during a session over communications from non-playing friends or other spectators. A user may be a spectator and may wish to prioritize communications from specific players they follow over other players they do not follow. Also, different auditory preferences (e.g., frequency response range, speed range, volume range) may be specified by the user and stored in a user audio profile for reference when applying them to incoming audio streams. Preferences may apply to multiple streams, individual streams, and / or specific sound / speech combinations within a stream.

[0020] The location of the audio may also be associated with prioritization. For example, real-world audio originating from the user's real-world environment may be prioritized. Related audio modifications may include increasing the level of noise transparency associated with other audio (e.g., in-game audio, online chat audio) and / or boosting the real-world audio. Thus, the user may be able to clearly hear the real-world audio even when the user may be engrossed in gameplay or online chat. The real-world audio may be captured by a microphone associated with the user device and provided to the audio analysis server 140 for analysis and processing.

[0021] A microphone may also be used to capture data regarding speech by the user. The user's speech may also be analyzed by the audio analysis server 140 according to the user audio profile. In some cases, a user may wish to apply modifications to their voice to one or more audiences. Different modifications may be applied to different streams. For example, the user's voice may be modified in one way in relation to one stream (e.g., modified to sound like an in-game character associated with the game title being played in the current session) while being modified in a different way in relation to another stream (e.g., modified low frequency response, increased or decreased volume, slowed down speed, modified to sound like a favorite character, celebrity, etc.).

[0022] Different prioritization schemes may also be specified depending on the particular content title, content type (e.g., movies vs. games), peer group (e.g., friends, teammates, opponents, spectators), channel or platform (e.g., in-game chat, external chat service such as Discord), and other characteristics of the session. In a given session, a user may receive multiple different audio streams via one or more user devices. The audio analysis and modifications described herein may be applied to the entire combination of audio streams or to one or more audio streams specified by the user. In some implementations, a user may specify all incoming audio streams, each of which may be associated with a different source, channel, or platform. The audio analysis server 140 may act as an intermediate device that modifies one or more audio streams before delivering the streams to the user device(s). Alternatively, the audio analysis server 140 may operate in conjunction with one or more local applications to provide audio-related insights and modifications to the user device(s) and local application(s) that apply them.

[0023] Different combinations of audio parameters (and associated audio modifications) may be stored for each user in a user audio profile. Thus, collectively, the user audio profile provides a custom audio modification setting for a particular user. Customization may further be based on current session conditions (e.g., current gameplay status, current game title, current in-game conditions). Because the audio stream is personalized by the audio analysis server 140 for a particular user based on their respective user audio profiles, sessions involving multiple different users may be modified to result in as many different versions of the audio stream(s) as there are users. In some embodiments, the audio analysis server 140 may generate customized bots programmed to apply a user's custom user audio profile to online communication sessions in which the user participates via their respective user devices.

[0024] In an exemplary embodiment, a user audio profile may be stored in a user's memory. The user audio profile may include a custom prioritization of one or more audio parameters associated with one or more audio modifications. Audio streams associated with the user's user device may be monitored based on the user audio profile during a current session. Audio parameters may be detected as present in the monitored audio streams, and the detected audio parameters may be prioritized based on the custom prioritization of the user audio profile. At least one sound characteristic of the audio streams may be modified in real time based on the prioritization of the detected audio parameters by applying the audio modifications of the user audio profile to the at least one audio stream before the at least one audio stream is provided to the user device.

[0025] FIG. 2 illustrates an exemplary Uniform Data System (UDS) 200 that can be used to provide data to a system for audio analysis and audio-based accessibility. Based on the data provided by the UDS 200, the conversion filter server 140 can recognize current session conditions, such as the in-game objects, entities, activities, and events the user is engaged in, which can then support the analysis and adjustment of conversion and filtering by the conversion filter server 140 along with the current gameplay and in-game activity. Each user interaction can be associated with metadata, such as the type of in-game interaction, its location within the in-game environment, its point in time within the in-game timeline, and other players, objects, and entities involved. Thus, metadata can be tracked for any of a variety of user interactions that may occur during a game session, including associated activities, entities, settings, results, actions, effects, locations, character statistics, and the like. Such data can be further aggregated, applied to a data model, and subjected to analysis. Such a UDS data model can be used to assign contextual information to pieces of information in a uniform manner across multiple games.

[0026] For example, various content titles may depict one or more objects with which users can interact (e.g., engage in in-game activities) and / or UGC (e.g., screenshots, videos, play-by-play commentary, mashups, etc.) created by peers, the publisher of the media content title, and / or third-party publishers. Such UGC may include metadata for searching such UGC. Such UGC may also include information about the media and / or peers. Such peer information may be derived from data collected during a peer's interaction with an object in an interactive content title (e.g., a video game, an interactive book, etc.) and may be "bound" to and stored with the UGC. Such binding enhances UGC because the UGC may deep link (e.g., launch directly) to the object, provide information about the object and / or peer in the UGC, and / or enable a user to interact with the UGC.

[0027] 2, an exemplary console 228 (e.g., user device 130) and exemplary servers 218 (e.g., streaming server 220, activity feed server 224, user-generated content (UGC) server 232, and object server 226) are shown. In one example, console 228 may be implemented on either platform server 120, cloud server, or server 218. In one illustrative example, content recorder 202 may be implemented on either platform server 120, cloud server, or server 218. Such content recorder 202 receives content (e.g., media) from interactive content title 230 and records it on a content ring buffer 208. Such a ring buffer 208 may store multiple content segments (e.g., v1, v2, and v3), a start time for each segment (e.g., V1_START_TS, V2_START_TS, V3_START_TS), and an end time for each segment (e.g., V1_END_TS, V2_END_TS, V3_END_TS). Such segments may be stored by the console 228 as media files 212 (e.g., MP4, WebM, etc.). Such media files 212 may be uploaded to a streaming server 220 for storage and subsequent streaming or use, although the media files 212 may be stored on any server, cloud server, any console 228, or any user device 130. Such start and end times for each segment may be stored by the console 228 as a content timestamp file 214. Such a content timestamp file 214 may also include a stream ID that matches the stream ID of the media file 212, thereby associating the content timestamp file 214 with the media file 212.Such content timestamp files 214 may be uploaded and stored on the activity feed server 224 and / or the UGC server 232, but the content timestamp files 214 may be stored on any server, cloud server, any console 228, or any user device 130.

[0028] At the same time that the content recorder 202 receives and records content from the interactive content title 230, the object library 204 receives data from the interactive content title 230, and the object recorder 206 tracks the data to determine the start and end times of the activity. The object library 204 and the object recorder 206 may be implemented on either the platform server 120, the cloud server, or the server 218. When the object recorder 206 detects the start of an activity, it receives object data (e.g., if the object is an activity, data on the activity, user interaction with the activity, activity ID, activity start time, activity end time, activity result, activity type, etc.) from the object library 204 and records this activity data (e.g., ActivityID1, START_TS; ActivityID2, START_TS; ActivityID3, START_TS) in the object ring buffer 210. Such activity data recorded on the object ring buffer 210 may be stored in the object file 216. Such object files 216 may also include activity start time, activity end time, activity ID, activity result, activity type (e.g., competition, quest, task, etc.), and user or peer data related to the activity. For example, object files 216 may store data regarding items used during the activity. Such object files 216 may be stored on an object server 226, although object files 216 may be stored on any server, cloud server, any console 228, or any user device 130.

[0029] Such object data (e.g., object file 216) can be associated with content data (e.g., media file 212 and / or content timestamp file 214). In one example, UGC server 232 stores content timestamp file 214 with object file 216 and associates content timestamp file 214 with object file 216 based on a match between the stream ID of content timestamp file 214 and the corresponding activity ID of object file 216. In another example, object server 226 can store object file 216 and receive queries for object file 216 from UGC server 232. Such queries can be performed by searching for an activity ID of object file 216 that matches the stream ID of content timestamp file 214 transmitted with the query. In yet another example, queries of stored content timestamp file 214 can be performed by matching the start and end times of content timestamp file 214 with the start and end times of the corresponding object file 216 transmitted with the query. Such an object file 216 may also be associated with a matching content timestamp file 214 by the UGC server 232, although this association may be performed by any server, cloud server, any console 228, or any user device 130. In another example, the object file 216 and the content timestamp file 214 may be associated by the console 228 during the creation of each of the files 216, 214.

[0030] In an exemplary embodiment, the media files 212 and activity files 216 may provide the audio analysis server 140 with information about current session conditions, which may also be used as another criterion for prioritizing and applying audio modifications to different audio streams. Thus, the audio analysis server 140 may use such media files 212 and activity files 216 to identify specific conditions of the current session, including currently speaking or noise-producing players, characters, and objects in specific locations and events. Based on such files 212 and 216, for example, the audio analysis server 140 may identify the significance level of an in-game event (e.g., a major battle, proximity to a record-breaking event) and use the significance level to apply audio modifications that prioritize in-game audio over other audio streams and allow the user to focus on the current session. Such session conditions may determine how audio parameters of different audio streams may be prioritized, thereby determining whether and which audio modifications are applied to which audio parameters / streams.

[0031] 3 is a flowchart illustrating an exemplary method 300 for audio analysis and audio-based accessibility. The method 300 of FIG. 3 may be embodied as executable instructions on a non-transitory computer-readable storage medium, including, but not limited to, non-volatile memory such as a CD, DVD, or hard drive. The instructions on the storage medium may be executed by a processor (or multiple processors) to cause various hardware components of a computing device that hosts or otherwise accesses the storage medium to perform the method. The steps (and their order) identified in FIG. 3 are exemplary and may include various alternatives, equivalents, or derivations thereof, including, but not limited to, the order of execution thereof.

[0032] At step 310, a user audio profile may be stored in memory (e.g., database 160). The user audio profile may include a custom prioritization of different audio parameters and associated audio modifications. The different audio parameters may include any detectable characteristics of sounds that may be provided to the user's user device. In some embodiments, a new user may develop the user's user audio profile by specifying the user's personal preferences and priorities with respect to different sounds and streams. In some implementations, audio analysis server 140 may query the new user to identify how to generate a custom prioritization of the user audio profile. As noted above, different user audio profiles may be stored for each user.

[0033] At step 320, one or more audio streams associated with user device 150 may be monitored by audio analysis server 140. A user using user device 150 to stream and play digital content in a current session may be presented with an audio stream associated with the digital content. If the digital content is a multiplayer game, a separate audio stream associated with a voice chat feature may also include sounds and speech associated with other users. In some cases, other chat services (e.g., Discord servers) may be used to provide additional audio streams. Audio analysis server 140 may monitor all such audio streams (or any particular subset) based on the user's user audio profile.

[0034] In step 330, the audio streams are analyzed by audio analysis server 140 against the user audio profile. In particular, audio parameters specified by the user audio profile are identified within each of the audio streams and prioritized according to a custom prioritization. For example, certain audio parameters may be preferred or dispreferred by a user, and each user's priority may reflect how the audio parameters are modified. Also, different session conditions may affect the degree to which an audio parameter is preferred or dispreferred relative to other audio parameters associated with user device 150.

[0035] At step 340, at least one audio stream associated with user device 150 may be modified based on the user audio profile. Based on the analysis and prioritization performed in step 330, audio analysis server 140 may identify relevant audio modifications and apply the modifications to the audio stream and one or more of the associated audio parameters. Preferred or high-priority audio parameters may be boosted or amplified, while less preferred or low-priority audio parameters may be attenuated, have their transparency adjusted, or canceled, according to the relevant audio modifications specified by the user audio profile.

[0036] In step 350, the modified audio stream(s) may be provided to user device 150 for real-time or near-real-time playback and presentation. Thus, a user of user device 150 may be presented with audio stream(s) that may emphasize or de-emphasize different audio parameters according to custom prioritizations specified by the user audio profile. The user may continue to specify refinements and other changes to the user audio profile over time, allowing audio modifications to be applied to future audio streams to better reflect the user's preferences and priorities.

[0037] 4 illustrates an exemplary implementation of audio analysis and audio-based accessibility features. As shown, different audio streams 410A-C may be associated with a user device's current session and provided to audio analysis server 140 for analysis and processing. Interactive audio stream 410A may be associated with an interactive content title and may include audio associated with the playback of that interactive content title. Social audio stream 410B may be associated with a social chat service or platform and may include audio associated with other users participating in the social chat session. Real-world audio 410C may be sounds occurring in the user's real-world environment that are captured by a microphone on user device 150 and provided to audio analysis server 140 for analysis and processing.

[0038] The audio analysis server 140 may retrieve the user audio profile from one of the databases 160, which may include a specific database that stores the user audio profile 420. Using the user audio profile associated with the user of the user device 150, the audio analysis server 140 may analyze various audio parameters that occur simultaneously across the different audio streams 410A-C. For example, an in-game sound from the interactive audio stream 410A may occur simultaneously with a voice chat message from the social audio stream 410B, which may also occur simultaneously with a real-world sound from the real-world audio stream 410C. The different audio parameters may be identified and prioritized by the audio analysis server according to a custom prioritization specified by the user audio profile. The audio analysis server 140 may further identify what audio modifications are specified by the user audio profile in relation to the prioritized audio parameters. The identified audio modifications may then be applied by the audio analysis server 140 to generate one or more modified audio streams 430. The modified audio stream(s) 430 may then be provided to the user device 150.

[0039] Figure 5 is a block diagram of an exemplary electronic entertainment system that may be used with embodiments of the present invention. The entertainment system 500 of Figure 5 includes a main memory 505, a central processing unit (CPU) 510, a vector unit 515, a graphics processing unit 520, an input / output (I / O) processor 525, an I / O processor memory 530, a controller interface 535, a memory card 540, a universal serial bus (USB) interface 545, and an IEEE interface 550. The entertainment system 500 may further include an operating system read-only memory (OS ROM) 555, an audio processing unit 560, an optical disc control unit 570, and a hard disk drive 565, which are connected to the I / O processor 525 via a bus 575.

[0040] Entertainment system 500 may be an electronic gaming console. Alternatively, entertainment system 500 may be implemented as a general-purpose computer, a set-top box, a handheld gaming device, a tablet computing device, a mobile computing device, or a mobile phone. Entertainment systems may include more or fewer operating components depending on the particular form factor, purpose, or design.

[0041] The CPU 510, vector unit 515, graphics processing unit 520, and I / O processor 525 of FIG. 5 communicate via a system bus 585. Additionally, the CPU 510 of FIG. 5 communicates with main memory 505 via a dedicated bus 580, while the vector unit 515 and graphics processing unit 520 may communicate via a dedicated bus 590. The CPU 510 of FIG. 5 executes programs stored in the OS ROM 555 and the main memory 505. The main memory 505 of FIG. 5 may include pre-stored programs and programs transferred via the I / O processor 525 from a CD-ROM, DVD-ROM, or other optical disk (not shown) using the optical disk control unit 570. The I / O processor 525 of FIG. 5 may also enable the introduction of content transferred over wireless or other communication networks (e.g., 4G, LTE, 3G, etc.). The I / O processor 525 of FIG. 5 primarily controls the exchange of data between various devices of the entertainment system 500, including the CPU 510, the vector unit 515, the graphics processing unit 520, and the controller interface 535.

[0042] The graphics processing unit 520 of Figure 5 executes graphics instructions received from the CPU 510 and the vector unit 515 to generate images for display on a display device (not shown). For example, the vector unit 515 of Figure 5 may convert an object from three-dimensional coordinates to two-dimensional coordinates and send the two-dimensional coordinates to the graphics processing unit 520. Additionally, the audio processing unit 560 executes instructions to generate audio signals, which are output to an audio device such as a speaker (not shown). Other devices may be connected to the entertainment system 500 via the USB interface 545 and the IEEE 1394 interface 550, such as a wireless transceiver; these interfaces may also be embedded in the system 500 or as part of some other component, such as a processor.

[0043] 5 provides instructions to CPU 510 via controller interface 535. For example, the user may instruct CPU 510 to store certain game information on memory card 540 or other non-transitory computer-readable storage medium, or may instruct a character in the game to perform some particular action.

[0044] The present invention may be implemented in an application that may be operable by a variety of end-user devices. For example, the end-user device may be a personal computer, a home entertainment system (e.g., Sony PlayStation2® or Sony PlayStation3® or Sony PlayStation4®), a portable gaming device (e.g., Sony PSP® or Sony Vita®), or a home entertainment system from a different, but subordinate, manufacturer. It is fully contemplated that the methods of the present invention described herein will be operable on a variety of devices. The present invention may also be implemented in a cross-title neutral manner, such that embodiments of the inventive system may be utilized across a variety of titles from a variety of publishers.

[0045] The present invention can be implemented in applications that can operate using a variety of devices. A non-transitory computer-readable storage medium refers to any medium or media that participates in providing instructions to a central processing unit (CPU) for execution. Such media can take many forms, including, but not limited to, non-volatile and volatile media, such as optical or magnetic disks and dynamic memory, respectively. Common forms of non-transitory computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROM disks, digital video disks (DVDs), any other optical media, RAM, PROM, EPROM, FLASHEPROM, and any other memory chip or cartridge.

[0046] Various forms of transmission media may be involved in carrying one or more sequences of one or more instructions to the CPU for execution. A bus carries the data to system RAM, and the CPU reads and executes the instructions from the system RAM. The instructions received by the system RAM may optionally be stored on a fixed disk either before or after execution by the CPU. Various forms of storage may be implemented similarly, as may network interfaces and network topologies necessary to implement the storage.

[0047] The foregoing detailed description of the present technology has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the technology to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The described embodiments have been chosen to best explain the principles of the technology, its practical application, and to enable those skilled in the art to utilize the technology in various embodiments and with various modifications suitable for the particular uses contemplated. It is intended that the scope of the present technology be defined by the claims.

Claims

1. 1. A computer-implemented method for audio analysis and audio-based accessibility, comprising: The computer storing a user audio profile in a user memory, the user audio profile including a custom prioritization of one or more audio parameters associated with one or more audio modifications; monitoring one or more audio streams associated with the user device of the user during a current session, the monitoring of the audio streams being based on the user audio profile; monitoring data relating to sounds of a real-world environment associated with the user captured by a microphone associated with the user device; Detecting one or more of the audio parameters as being present in the monitored audio stream, wherein the detected audio parameters are further based on the sounds of the real-world environment; prioritizing the detected audio parameters based on the custom prioritization of the user audio profile; modifying at least one sound characteristic of the audio stream in real time based on the prioritization of the detected audio parameters, the sound characteristic of the at least one audio stream including at least one of volume, speed, and frequency response relative to one or more others of the audio streams, and modifying the sound characteristic includes applying the audio modifications of the user audio profile to the at least one audio stream before the at least one audio stream is provided to the user device; A method comprising:

2. 2. The method of claim 1 , wherein monitoring the audio streams includes analyzing each of the audio streams to distinguish speech from other sounds, and modifying the sound characteristics includes boosting the sound characteristics of the at least one audio stream that includes the speech relative to one or more others of the audio streams.

3. The method of claim 1 , wherein the detected audio parameters include speech by a plurality of different users, and the custom prioritization includes a prioritization order for one or more user groups.

4. The method of claim 1 , wherein the sounds of the real-world environment include speech by the user, and detecting the audio parameter as present is based on identifying the speech by the user.

5. 5. The method of claim 4, wherein the user audio profile further includes data regarding the user's voice, and further comprising providing the at least one modified audio stream to one or more other user devices associated with the current session.

6. The method of claim 1 , wherein the audio streams are associated with multiple different applications or platforms operating simultaneously during the current session.

7. The method of claim 1 , wherein the audio modification comprises at least one of adjusting one of the audio parameters, attenuation, adjusting transparency, and noise cancellation.

8. 1. A system for audio analysis and audio-based accessibility, comprising: a memory for storing a user audio profile in the memory, the user audio profile including a custom prioritization of one or more audio parameters associated with one or more audio modifications; a communications interface for communicating over a communications network, the communications interface receiving one or more audio streams associated with a user device of the user during a current session; a processor that executes instructions stored in a memory, the processor executing the instructions to monitoring the audio stream based on the user audio profile; monitoring data relating to sounds of a real-world environment associated with the user captured by a microphone associated with the user device; Detecting one or more of the audio parameters as being present in the monitored audio stream, wherein the detected audio parameters are further based on the sounds of the real-world environment; prioritizing the detected audio parameters based on the custom prioritization of the user audio profile; modifying at least one sound characteristic of the audio stream in real time based on the prioritization of the detected audio parameters, the sound characteristic of the at least one audio stream including at least one of volume, speed, and frequency response relative to one or more others of the audio streams, and modifying the sound characteristic includes applying the audio modifications of the user audio profile to the at least one audio stream before the at least one audio stream is provided to the user device; Including, the system.

9. 9. The system of claim 8, wherein the processor monitors the audio streams by analyzing each of the audio streams to distinguish speech from other sounds, and the processor modifies the sound characteristics by boosting the sound characteristics of the at least one audio stream that includes the speech relative to one or more others of the audio streams.

10. The system of claim 8 , wherein the detected audio parameters include speech by a plurality of different users, and the custom prioritization includes a prioritization order for one or more user groups.

11. The system of claim 8 , wherein the sounds of the real-world environment include speech by the user, and the processor further detects the presence of the audio parameter based on identifying the speech by the user.

12. 12. The system of claim 11, wherein the user audio profile further includes data regarding the user's voice, and the communication interface further provides the at least one modified audio stream to one or more other user devices associated with the current session.

13. The system of claim 8 , wherein the audio streams are associated with multiple different applications or platforms operating simultaneously during the current session.

14. The system of claim 8 , wherein the audio modification includes at least one of adjusting one of the audio parameters, attenuation, adjusting transparency, and noise cancellation.

15. 1. A non-transitory computer-readable storage medium embodying a program executable by a processor to perform a method for audio analysis and audio-based accessibility, the method comprising: storing a user audio profile in a user memory, the user audio profile including a custom prioritization of one or more audio parameters associated with one or more audio modifications; monitoring one or more audio streams associated with the user device of the user during a current session, the monitoring of the audio streams being based on the user audio profile; monitoring data relating to sounds of a real-world environment associated with the user captured by a microphone associated with the user device; Detecting one or more of the audio parameters as being present in the monitored audio stream, wherein the detected audio parameters are further based on the sounds of the real-world environment; prioritizing the detected audio parameters based on the custom prioritization of the user audio profile; modifying at least one sound characteristic of the audio stream in real time based on the prioritization of the detected audio parameters, the sound characteristic of the at least one audio stream including at least one of volume, speed, and frequency response relative to one or more others of the audio streams, and modifying the sound characteristic includes applying the audio modifications of the user audio profile to the at least one audio stream before the at least one audio stream is provided to the user device; 1. A non-transitory computer-readable storage medium comprising:

Citation Information

Patent Citations

  • Sound volume controller

    JP2002116045A

  • Visual instruction of current voice speaker

    JP2005100420A