Sound Editing
The method addresses unnatural sound reflections in AR/VR by fragmenting and attenuating sound clips based on a historical corpus, enhancing the audio experience with natural attenuation and user-adjustable modifications.
Patent Information
- Application Number
- JP2025514168
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-07
- Filing Date
- 2023-08-02
- Publication Date
- 2025-09-11
AI Technical Summary
In augmented or virtual reality environments, audio clips often require acoustic attenuation due to the acoustical nature of three-dimensional space, where objects reflect and absorb sound, leading to unnatural sound experiences.
A method and system that identifies sound clips for a user's location, fragments them if necessary, applies acoustic attenuation, and stitches them back together, using a historical corpus to mimic natural sound attenuation, while providing a visual representation for modification.
Enhances the audio experience in AR/VR by ensuring natural sound attenuation and allowing users to adjust sound clips, improving the realism and coherence of the audio track.
Smart Images

Figure 2025530173000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to audio editing, and more particularly to audio editing in augmented or virtual reality environments.
[0002] Augmented reality (AR) is the integration of digital information, such as audio, with the physical environment in which a user is located. Virtual reality (VR) is a simulation of a three-dimensional digital environment in which a user can experience and interact with various digital objects within the three-dimensional digital environment. Both AR and VR require the use of electronic devices to present digital information and the three-dimensional digital environment to the user. Audio in an AR or VR situation typically includes multiple audio clips stitched together to create an aggregated and complete audio track for the user's position relative to the physical or digital environment. Due to the acoustical nature of three-dimensional space, there may be objects or materials in the three-dimensional space that reflect and absorb audio from various locations relative to the user's position in the AR or VR environment. Summary of the Invention
[0003] Embodiments of the present invention disclose a method, computer program product, and computer system for acoustic attenuation of sound clips, which can identify an audio clip for a user's location in an environment. The method, computer program product, and computer system can fragment the audio clip into multiple sound clips. In response to determining that at least one sound clip from the audio clip requires acoustic attenuation, the method, computer program product, and computer system can perform the acoustic attenuation on the at least one sound clip, where an attenuation ratio for the at least one sound clip is changed. In response to determining to stitch the multiple sound clips, the method, computer program product, and computer system can stitch the multiple sound clips to form the audio clip, where the multiple sound clips include the at least one sound clip with the acoustic attenuation. The method, computer program product, and computer system can display a visual representation of the audio clip comprising the multiple sound clips. [Brief explanation of the drawings]
[0004] [Figure 1] FIG. 1 is a functional block diagram illustrating a computing environment according to one embodiment of the present invention.
[0005] [Figure 2] 1 is a flowchart of an audio editing program for stitching multiple sound clips to create an aggregated sound clip, according to one embodiment of the present invention.
[0006] [Figure 3]1 is a flowchart of an audio editing program for modifying aggregated sound clips and developing a historical corpus with a plurality of aggregated sound clips, according to one embodiment of the present invention.
[0007] [Figure 4A] FIG. 2 shows an illustrative example of multiple sound clips before stitching, according to one embodiment of the present invention.
[0008] [Figure 4B] FIG. 2 shows an illustrative example of multiple sound clips after stitching to create an aggregated sound clip according to one embodiment of the present invention.
[0009] [Figure 4C] FIG. 10 shows an illustrative example of modifying aggregated sound clips based on directional sources per sound clip, according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] Although detailed embodiments of the claimed structures and methods are disclosed herein, it should be understood that the disclosed embodiments are merely exemplary of the claimed structures and methods, which may be embodied in various forms. The present invention may, however, be embodied in many different forms and should not be construed as limited to the exemplary embodiments set forth herein. In the description, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments. The singular forms "a," "an," and "the" should be understood to include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a "component surface" includes reference to one or more of such surfaces unless the context clearly dictates otherwise.
[0011] An embodiment of the present invention provides an audio editing program that identifies sound clips within an audio clip in an augmented or virtual reality environment, provides abrupt termination with limited or no attenuation, and creates appropriate attenuation effects for each sound clip within the audio clip. The audio editing program allows a user to opt in as an administrative user for the augmented or virtual reality environment to modify audio clips with appropriate attenuation for the sound clip. The audio editing program can create a sonic profile for each sound clip within the audio clip by analyzing the sound profile of the audio clip, identifying whether any sound clips in the audio clip require attenuation, and, if necessary, splitting the audio clip into individual sound clips to perform the attenuation. Because multiple sound clips may overlay each other, the audio editing program splits each sound clip so that appropriate attenuation can be performed. The audio editing program can identify the source of each sound clip from the audio clip and identify where appropriate stitching is required after performing the attenuation.
[0012] The sound editing program can be incorporated into an augmented or virtual reality electronic device (e.g., a wearable headset with speakers), where an administrator user can modify each sound clip of an audio clip. The sound editing program can identify unattenuated or under-attenuated sound clips from the audio by comparing each sound clip to a historical corpus, and can perform attenuation based on the historical corpus to attenuate each sound clip to mimic natural sound attenuation in a physical environment. The sound editing program can modify each sound clip's attenuation coefficient and each sound clip's attenuation type (e.g., full attenuation, no attenuation, variable attenuation) to mimic natural sound attenuation in a physical environment. The sound editing program can display a visual representation of the sound waves for each sound clip of the audio clip in an augmented or virtual reality electronic device associated with the administrator user, along with the position of each sound clip relative to the user's location, where the user can modify the position of each sound clip via an input in the augmented or virtual reality electronic device. The sound editing program can provide selection of modifications for each sound wave of a sound clip in a user collaborative environment, where multiple users can edit an audio clip as a designated administrator user.
[0013] Various aspects of the present disclosure are described through text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program product (CPP) embodiments. For any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, again depending on the technology involved, two operations shown in successive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in an at least partially overlapping manner.
[0014] A computer program product embodiment ("CPP embodiment" or "CPP") is a term used in this disclosure to describe any set of one or more storage media (also referred to as "media") collectively included in a set of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. The computer-readable storage medium may be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on a major surface of a disk), or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in this disclosure, is not to be construed as storage in the form of a transitory signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through fiber optic cables, electrical signals communicated through wires, and / or other transmission media.As will be appreciated by those skilled in the art, data is typically moved at some infrequent time during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but the above does not make the storage device temporary, as the data is not temporary while it is stored.
[0015] Figure 1 is a functional block diagram illustrating a computing environment, generally designated 100, in accordance with one embodiment of the present invention. Figure 1 is intended only as an illustration of one implementation and does not imply any limitation with regard to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be made by one skilled in the art without departing from the scope of the present invention as recited by the claims.
[0016] Computing environment 100 comprises an example of an environment for the execution of at least a portion of computer code involved in performing the method of the present invention, such as audio editing program 200. In addition to block 200, computing environment 100 includes, for example, computer 101, wide area network (WAN) 102, end user device (EUD) 103, remote server 104, public cloud 105, and private cloud 106. In this embodiment, computer 101 includes a processor set 110 (including processing circuitry 120 and cache 121), communication fabric 111, volatile memory 112, persistent storage 113 (including operating system 122 and block 200 as identified above), a peripheral device set 114 (including a user interface (UI) device set 123, storage 124, and an Internet of Things (IoT) sensor set 125), and network module 115. Remote server 104 includes a remote database 130. The public cloud 105 includes a gateway 140, a cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.
[0017] Computer 101 may take the form of a desktop computer, a laptop computer, a tablet computer, a smartphone, a smartwatch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device now known or later developed that is capable of executing programs, accessing a network, or querying a database, such as remote database 130. As is well understood in the art of computer technology, and depending on the technology, execution of a computer-implemented method may be distributed among multiple computers and / or among multiple locations. While in this presentation of computing environment 100, to keep the presentation as concise as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not depicted in the cloud of FIG. 1 , it may be located in a cloud. However, computer 101 is not required to reside within a cloud except to any extent that may be expressly indicated.
[0018] Processor set 110 includes one or more computer processors of any type now known or later developed. Processing circuitry 120 may be distributed across multiple packages, e.g., multiple coordinated integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on relative proximity to the processing circuitry. Alternatively, some or all caches for a processor set may be located “off-chip.” In some computing environments, processor set 110 may be designed to operate with qubits and perform quantum computing.
[0019] Computer-readable program instructions are typically loaded onto computer 101 to cause processor set 110 of computer 101 to perform a series of operational steps, thereby realizing a computer-implemented method, whereby the instructions so executed instantiate the method specified in the flowcharts and / or descriptions of the computer-implemented method contained herein (collectively referred to as the "methods of the present invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by processor set 110 to control and direct the execution of the methods of the present invention. In computing environment 100, at least some of the instructions for performing the methods of the present invention may be stored in block 200 within persistent storage 113.
[0020] Communications fabric 111 is the signal-conducting pathway that allows various components of computer 101 to communicate with one another. Typically, this fabric is made up of switches and conductive pathways, such as those that make up buses, bridges, physical input / output ports, and the like. Other types of signal communication pathways may be used, such as fiber optic and / or wireless communication pathways.
[0021] Volatile memory 112 may be any type of volatile memory now known or later developed. Examples include dynamic type random access memory (RAM) or static type RAM. Typically, volatile memory 112 is characterized by random access, although this is not required unless expressly indicated. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.
[0022] Persistent storage 113 is any form of non-volatile storage for a computer, now known or later developed. The non-volatility of this storage means that stored data remains whether or not power is supplied to computer 101 and / or to persistent storage 113 directly. Persistent storage 113 may be read-only memory (ROM), but typically at least a portion of persistent storage allows data to be written, data to be deleted, and data to be rewritten. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. Operating system 122 may take several forms, such as various known proprietary operating systems or open-source Portable Operating System Interface-type operating systems that employ a kernel. The code contained in block 200 typically includes at least a portion of the computer code involved in performing the methods of the present invention.
[0023] The peripheral device set 114 includes a set of peripheral devices of the computer 101. Data communication connections between the peripheral devices and other components of the computer 101 may be implemented in various ways, such as Bluetooth® connections, Near-Field Communication (NFC) connections, connections made by cable (such as a universal serial bus (USB)-type cable), insertion-type connections (e.g., a secure digital (SD) card), connections made over a local area communication network, and even connections made over a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. The storage 124 may be external storage, such as an external hard drive, or insertable storage, such as an SD card. The storage 124 may be persistent and / or volatile. In some embodiments, storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 is required to have a large amount of storage (e.g., computer 101 locally stores and manages large databases), this storage may be provided by a peripheral storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. IoT sensor set 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.
[0024] Network module 115 is a collection of computer software, hardware, and firmware that enables computer 101 to communicate with other computers over WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi® signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing software-defined networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for implementing the methods of the present invention can be downloaded to computer 101 from an external computer or external storage device, typically through a network adapter card or network interface included in network module 115.
[0025] WAN 102 is any wide area network (e.g., the Internet) capable of communicating computer data over non-local distances by any technology for communicating computer data, now known or later developed. In some embodiments, WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi® network. WANs and / or LANs typically include copper transmission cables, optical fiber transmissions, wireless transmissions, and computer hardware such as routers, firewalls, switches, gateway computers, and edge servers.
[0026] End-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and may take any of the forms discussed above in connection with computer 101. EUD 103 typically receives useful and useful data from the operation of computer 101. For example, in the hypothetical case where computer 101 is designed to provide recommendations to the end user, the recommendations would typically be communicated from network module 115 of computer 101 over WAN 102 to EUD 103. In this manner, EUD 103 can display or otherwise present the recommendations to the end user. In some embodiments, EUD 103 may be a client device such as a thin client, a heavy client, a mainframe computer, a desktop computer, etc.
[0027] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 may be controlled and used by the same entity that operates computer 101. Remote server 104 represents a machine that collects and stores useful and useful data for use by other computers, such as computer 101. For example, in the hypothetical case where computer 101 is designed and programmed to provide recommendations based on historical data, this historical data may be provided to computer 101 from remote database 130 of remote server 104.
[0028] Public cloud 105 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computer functionality, particularly data storage (cloud storage) and computing power, without direct active management by users. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct active management of public cloud 105's computing resources is performed by computer hardware and / or software in cloud orchestration module 141. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments (VCEs) running on various computers that comprise host physical machine set 142, the universe of physical computers in and / or available to public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 143 and / or containers from container set 144. It is understood that these VCEs may be stored as images and can be transferred among and between various physical machine hosts, either as images or after instantiation of the VCEs. Cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of VCE, and manages active instantiations of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that enables public cloud 105 to communicate over WAN 102.
[0029] Some further description of a virtualized computing environment (VCE) is now provided. A VCE can be stored as an "image." A new, active instance of a VCE can be instantiated from the image. Two well-known types of VCE are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows the existence of multiple isolated user space instances called containers. These isolated user space instances typically behave as actual computers from the perspective of programs running in them. A computer program running on a normal operating system can utilize all of the computer's resources, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and of the devices assigned to the container; this feature is known as containerization.
[0030] A private cloud 106 is similar to a public cloud 105, except that the computing resources are available only for use by a single enterprise. While the private cloud 106 is shown as communicating with the WAN 102, in other embodiments, the private cloud may be completely disconnected from the Internet and accessible only through a local / private network. A hybrid cloud is a composite of multiple clouds of different types (e.g., private, community, or public cloud types), often each implemented by a different vendor. While each of the multiple clouds remains a separate, discrete entity, the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.
[0031] FIG. 2 shows a flow chart of an audio editing program for stitching multiple sound clips to create an aggregated sound clip according to one embodiment of the present invention.
[0032] The sound editing program 200 receives user approval to opt in to sound editing (202). In one embodiment, the sound editing program 200 allows a user to opt in to sound editing of sound clips for a virtual reality (VR) environment utilizing a VR electronic device associated with the user, where the user utilizes a VR electronic device (e.g., a headset) to access the VR environment. In another embodiment, the sound editing program 200 allows a user to opt in to sound editing of sound clips for an augmented reality environment utilizing an AR electronic device associated with the user, where the user utilizes an AR electronic device (e.g., smart glasses) to access the AR environment. Once the user opts in to sound editing, the sound editing program 200 assigns the user an administrator role to edit various sound clips in the VR or AR environment. The sound editing program 200 manages a historical corpus of edits for various sound clips in the VR or AR environment, as discussed in more detail with respect to (312) in FIG. 3 .
[0033] The sound editing program 200 identifies audio clips in the augmented or virtual reality environment (204). When a user enters a physical or virtual environment using an AR or VR device, the sound editing program 200 determines a location for the user in the physical or virtual reality environment. In one embodiment, the sound editing program 200 identifies existing audio clips for locations in the physical or virtual reality environment, where the sound editing program 200 can utilize integrated audio editing software to extract track information. In another embodiment, the sound editing program 200 can identify audio clips in the augmented reality environment that include sound clips from the physical environment and existing digital sound clips for locations in the physical environment. The portions of the audio clips that include sound clips from the physical environment identified by the sound editing program 200 can be stored in a history corpus and subsequently made available for other users (i.e., non-administrator users) who enter that location in the augmented reality environment. For example, if an audio clip includes a sound clip of rustling leaves from the physical environment and an existing digital sound clip of blowing wind for that location, the sound editing program 200 can update the historical corpus by storing the sound clip of rustling leaves from the physical environment as the audio clip for that location that already includes the existing digital sound clip of blowing wind.
[0034] The sound editing program 200 analyzes a sound profile for the audio clip (206). As mentioned above, the sound editing program 200 utilizes integrated audio editing software to extract track information for the AR or VR environment, and the sound editing program 200 analyzes the sound profile for the audio clip based on the extracted track information to identify each sound clip in the audio clip. The audio clip includes one or more sound clips for a location in the AR or VR environment, where each sound clip is associated with an audio source from which the sound clip originates. Each sound clip generates a sound wave with an attenuation that determines the duration and intensity of the sound clip at the location where the user is positioned in the AR or VR environment. The sound editing program 200 can analyze the sound waves for each sound clip of the audio clip via a Fourier transform using equation (A) provided below:
number
[0035] The sound editing program 200 determines whether multiple sound clips are present in the audio clip (decision 208). If the sound editing program 200 determines that multiple sound clips are present in the audio clip (the "yes" branch of decision 208), the sound editing program 200 fragments the multiple sound clips for the audio clip (210). If the sound editing program 200 determines that multiple sound clips are not present in the audio clip (the "no" branch of decision 208), the sound editing program 200 determines whether attenuation of the single sound clip representing the audio clip is required (decision 212).
[0036] The sound editing program 200 fragments the multiple sound clips for an audio clip (210). Because an audio clip may include at least two overlapping sound clips, the sound editing program 200 fragments the multiple sound clips for an audio clip by pulling each sound clip from the multiple sound clips that form the audio clip for the AR or VR environment. By fragmenting the multiple sound clips, the sound editing program 200 can identify an attenuation profile for each sound clip that forms the audio clip, where the attenuation profile is then viewable on an AR or VR electronic device associated with the user. Fragmenting the multiple sound clips allows a user of the sound editing program 200 to view the attenuation profile of each individual sound clip, because the attenuation profile of one sound clip may overwhelm or obscure the attenuation profile of another sound clip, thus affecting the user's experience of that location in the AR or VR environment.
[0037] The sound editing program 200 determines whether attenuation of the sound clip representing the audio clip is required (decision 212). If the sound editing program 200 determines that attenuation of the sound clip representing the audio clip is required (the "yes" branch of decision 212), the sound editing program 200 performs sound attenuation on the audio clip (214). If the sound editing program 200 determines that attenuation of the sound clip representing the audio is not required (the "no" branch of decision 212), the sound editing program 200 determines whether stitching of the audio clip is required (decision 216). The sound editing program 200 can utilize a historical corpus of various sounds in the AR or VR environment that can be continuously updated by the user. The sound editing program 200 utilizes the historical corpus to identify which sound clips from the audio clip require full attenuation, variable attenuation, or no attenuation. Different sound clips of the audio clip can contain different types of sounds, where full attenuation may not be captured or even required. For example, the sound editing program 200 may have previously identified a location for an audio clip as a cafe, where the audio clip may include portions of the sound clip that require full attenuation and other portions of the sound clip that do not require full attenuation. A sound clip of unclear speech from various customers involved in a conversation at the cafe represents continuous audio at that location, where the sound editing program 200, utilizing a historical corpus, may determine that full attenuation is not required for this sound clip. Instead, the sound editing program 200 may determine that varying degrees of attenuation are required because the unclear speech is coming from multiple sources (i.e., customers) at that location. A sound clip of a coffee cup being placed on a plate represents a sound clip that would require full attenuation because the coffee cup clashing with the plate is not a continuous sound in an AR or VR environment.The sound editing program 200 determines that full attenuation is desired for the sound clip of a coffee cup being placed on a plate.
[0038] The sound editing program 200 performs sound attenuation for the audio clip (214). For each sound clip in the audio clip, the sound editing program 200 utilizes the historical corpus to determine whether full attenuation, variable attenuation, or no attenuation is required for the sound waves. In the case of a sound clip for which the sound editing program 200 determines that the historical corpus does not contain an entry for the indicated level of attenuation (i.e., the attenuation ratio), the sound clip remains unchanged. The sound editing program 200 allows the user to modify the sound clip and then update the historical corpus for subsequent attenuation of similar sound clips. Updating the historical corpus is discussed in more detail with respect to (312) in FIG. 3. From the example described above in which the audio clip is about a cafe in an AR or VR environment, the sound editing program 200 has identified multiple sound clips for the audio clip that include garbled dialogue, coffee cups clinking against plates, background music, and footsteps. The sound editing program 200 utilizes the historical corpus to perform variable attenuation for garbled speech, full attenuation for coffee cups clinking against plates, and full attenuation for footsteps. However, the sound editing program 200 determines that background music sound clips do not exist in the historical corpus, and therefore the sound editing program 200 determines not to perform sound attenuation for the sound clips.
[0039] The sound editing program 200 determines whether stitching of an audio clip is required (decision 216). If the sound editing program 200 determines that stitching of an audio clip is required (the "yes" branch of decision 216), the sound editing program 200 stitches multiple sound clips for the audio clip (218). If the sound editing program 200 determines that stitching of an audio clip is not required (the "no" branch of decision 216), the sound editing program 200 displays a visual representation of the audio clip (220). If the audio clip has previously been fragmented due to multiple sound clips, the sound editing program 200 determines that stitching of an audio clip is required. Furthermore, if the sound editing program 200 previously performed sound attenuation of at least one sound clip from the multiple sound clips of the audio, the sound editing program 200 takes into account how the stitching of the multiple sound clips with at least one sound clip with sound attenuation will be performed.
[0040] The sound editing program 200 stitches multiple sound clips for the audio clip (218). The sound editing program 200 stitches multiple sound clips for the audio clip by analyzing various characteristics of each sound clip and the user's location in the AR or VR environment. The characteristics of each sound clip can indicate whether the sound clip is random or correlated to another item (e.g., sound clip, action) in the AR or VR environment. From the example above where the audio clip is about a cafe in the AR or VR environment, the sound editing program 200 stitches multiple sound clips for the audio clip including garbled dialogue, a coffee cup clinking against a plate, background music, and footsteps. As described above, the sound editing program 200 implemented variable attenuation for the garbled dialogue, full attenuation for the coffee cup clinking against a plate, and full attenuation for the footsteps. The sound editing program 200 stitches multiple sound clips for an audio clip by overlaying the indistinct dialogue on background music, while inserting random instances of coffee cups clinking against plates and footsteps into the overlay of the indistinct dialogue and background music. The sound editing program 200 determines that the indistinct dialogue should be overlaid with the background music based on the characteristics of the background music as it relates to the indistinct dialogue, because the background music should not overpower (i.e., be louder than) the conversations of customers in the cafe.
[0041] The sound editing program 200 displays a visual representation of the audio clip (220). The sound editing program 200 can display various visual representations of the audio clip, including at least one sound clip, on an AR or VR electronic device associated with the user. Through use of the AR or VR electronic device, the sound editing program 200 can receive various modifications to the audio clip, as discussed in more detail with respect to (306) in FIG. 3. In one embodiment, the sound editing program 200 displays a visual representation of the audio clip, including multiple sound clips stitched together to form the audio clip from (218). The sound editing program 200 can utilize distinct colors for each sound clip from the multiple sound clips of the audio clip for its location, along with distinct representations that differentiate the original portion of the sound clip and the modified attenuation portion of the sound clip, as discussed in more detail with respect to FIGS. 4A and 4B. The sound editing program 200 can include the option to individually view each sound clip with or without the attenuation performed in (214).
[0042] FIG. 3 shows a flow chart of an audio editing program for modifying aggregated sound clips and developing a historical corpus with a plurality of aggregated sound clips, according to one embodiment of the present invention.
[0043] The sound editing program 200 identifies, for the user, a location for each sound clip in the audio clip with sound attenuation (302). As previously described, when the user enters a physical or virtual environment using an AR or VR device, the sound editing program 200 determines a location for the user in the physical or virtual reality environment. In one embodiment, the audio clip with attenuation includes a sound clip from the physical environment and an existing digital sound clip for a location in the physical environment. The sound clip from the physical environment can be captured by a microphone on an AR electronic device associated with the user, where the sound editing program 200 can determine the location of the sound clip from the physical environment relative to the user's known location. For the existing digital sound clip from a location in the physical environment, the sound editing program 200 can utilize a history corpus to identify the location of the existing digital sound clip from the physical environment relative to the user's known location. In another embodiment, the sound editing program 200 identifies, for the user, a location for each sound clip in the audio clip with sound attenuation in the VR environment. From the example described above, where the audio clip is about a cafe in a VR environment, the sound editing program 200 identifies multiple sound clips for the audio clip, including indistinct conversation, a coffee cup clinking against a plate, background music, and footsteps. The sound editing program 200 identifies multiple locations for the clear conversation, where the multiple locations are 45°, 105°, and 285° relative to a user's front orientation of 0°. The sound editing program 200 identifies the location of the coffee clinking against the plate as 285° relative to a user's front orientation of 0°. The sound editing program 200 identifies the location of the background music as surrounding the user (i.e., 360°) and identifies the location of the footsteps as transitioning between 220° and 160° relative to a user's front orientation of 0°.
[0044] The sound editing program 200 displays the position of each sound in the audio clip with sound attenuation (304). The sound editing program 200 displays the position of each sound in the audio clip with sound attenuation on an AR or VR electronic device associated with the user. The sound editing program 200 displays the position of each sound clip (i.e., the original location) of the audio clip relative to the user's position; an example is discussed in more detail with respect to FIG. 4C . Based on the volume and degree of attenuation of each sound clip, the sound editing program 200 displays the position of each sound clip at an appropriate distance from the user's position. For example, the sound editing program 200 displays a louder sound clip at a location closer to the user than a lower volume sound clip at a location closer to the user.
[0045] The sound editing program 200 receives modifications to the audio clip with sound attenuation (306). In one embodiment, the sound editing program 200 receives modifications to the audio clip with sound attenuation via a VR electronic device associated with a user, where the modifications include a change to the position of the sound clip from the audio clip. For example, the sound editing program 200 previously identified and displayed the position of the sound clip for the coffee cup hitting the plate as 285 degrees relative to the user's front orientation of 0 degrees. The sound editing program 200 receives the modifications to the position via the VR electronic device, where the user drags the position of the sound clip for the coffee cup hitting the plate to 20 degrees relative to the user's front orientation of 0 degrees. In another embodiment, the sound editing program 200 receives modifications to the audio clip with sound attenuation via an AR electronic device associated with a user, where the modifications include a change to the attenuation of the sound clip from the audio clip. For example, the sound editing program 200 may have previously performed sound attenuation for a sound clip from an audio clip using a historical corpus of specific noises in the sound clip (e.g., unclear conversations in a cafe). However, the unclear conversations in the cafe typically vary in volume because different people speak from different locations in the cafe. The sound editing program 200 may receive modifications via the VR electronic device, including variable attenuation for the sound clip, where the volume increases and decreases at random intervals. In other embodiments, the sound editing program 200 may receive modifications including changes to the attenuation ratio for the sound clip of the audio clip, the attenuation type for the sound clip of the audio clip (e.g., full attenuation, variable attenuation, or no attenuation), the position for the sound clip of the audio clip relative to the user's location in the AR or VR environment, removing a sound clip from an audio clip, adding a sound clip to an audio clip, and stitching multiple sound clips together in an audio clip.
[0046] The sound editing program 200 displays the modified audio clip with sound attenuation (308). Similar to displaying a visual representation and location of each sound clip from the audio clip, the sound editing program 200 displays the modified audio clip with sound attenuation on an AR or VR electronic device associated with the user. An administrator user of the sound editing program 200 can view the modifications made to the audio clip with sound attenuation and can replay the modified audio clip with sound attenuation at the location. Replaying can include the sound editing program 200 continuously looping the modified audio clip with sound attenuation for the location as long as an administrator user of the sound editing program 200 is present at the location.
[0047] The sound editing program 200 stores the modified audio clips with sound attenuation for the environment (310). The sound editing program 200 stores each sound clip from the modified audio clips with sound attenuation for each location in the AR or VR environment. By storing the modified audio clips with sound attenuation for each location in the AR or VR environment, subsequently entering users experience the modified audio clips with sound attenuation for that location in the AR or VR environment. The user entering the location in the AR or VR environment can be either an administrator user with administrator privileges who can later modify the modified audio clips with sound attenuation, or a non-administrator user who is experiencing the modified audio clips with sound attenuation for that location in the AR or VR environment. The sound editing program 200 updates the history corpus with the modified audio clips with sound attenuation (312). The sound editing program 200 updates the history corpus with each sound clip from the modified audio clips with sound attenuation for that location. The sound editing program 200 utilizes the modifications received from the administrator user via the AR or VR electronic device for subsequent analysis and attenuation of the audio clip for other locations, thus reducing the amount of time the administrator user spends modifying the audio clip with attenuation.
[0048] 4A shows an illustrative example of multiple sound clips before stitching, according to one embodiment of the present invention. In this example, the sound editing program 200 identifies audio clips 400 in a VR environment for an administrator user using a VR electronic device at a particular location, determines the audio clip 400 includes multiple sound clips, and fragments the audio clip 400 into a first sound clip 402 and a second sound clip 404. Using the historical corpus, the sound editing program 200 determines when attenuation is required for both the first sound clip 402 and the second sound clip 404 based on both sound clips matching entries in the historical corpus. The first sound clip 402 represents a first person speaking indistinctly in the background, and the second sound clip 404 represents a second person speaking indistinctly in the background. The sound editing program 200 decides to perform sound attenuation on the audio clip 400 to reduce the moments of silence represented by areas 406 and 408, which may sound unnatural to other users due to the abrupt volume fluctuations between the first sound clip 402 and the second sound clip 404.
[0049] 4B shows an illustrative example of multiple sound clips after stitching to create an aggregated sound clip, according to one embodiment of the present invention. Continuing the example from FIG. 4A, the sound editing program 200 performs sound attenuation on the first sound clip 402 and the second sound clip 404 of the audio clip 400. Based on the historical corpus, the sound editing program 200 increases the oscillation of the attenuation of the first sound clip 402 and the second sound clip 404 by decreasing the attenuation ratio until moments of silence no longer exist between the first sound clip 402 and the second sound clip 404. The sound editing program 200 adds a first attenuation portion 410 to the first sound clip 402 and a second attenuation portion 412 to the second sound clip 404. In this example, the sound clip 400 is repeatedly played in a loop, so that the second attenuation portion 412 loops back to the first sound clip 402, completing a loop cycle for the audio clip 400 for the VR environment.
[0050] 4C shows an illustrative example of modifying aggregated sound clips based on a directional source for each sound clip, according to one embodiment of the present invention. Continuing the example from FIGS. 4B and 4C, the sound editing program 200 displays a visual 414 at a VR electronic device associated with an administrator user, where the visual 414 provides the positions of a first sound clip 402 with a first attenuation portion 410 and a second sound clip 404 with a second attenuation portion 412 relative to the administrator user's location. In this example, the sound editing program 200 identifies the position for the first sound clip 402 with the first attenuation portion 410 as 310° relative to the user's front orientation of 0°, and identifies the position for the second sound clip 404 with the second attenuation portion 412 as 135° relative to the user's front orientation of 0°. However, the sound editing program 200 receives modifications to the first sound clip 404 with the second attenuation portion 412 by the user dragging the visual representation via the VR device from a 135° position to a 45° position relative to the user's front orientation of 0°. Additionally, the sound editing program 200 can increase or decrease the volume of the first sound clip 402 with the first attenuation portion 410 and the second sound clip 404 with the second attenuation portion 412 by dragging each visual representation toward or away from the user's view 416, respectively.
[0051] The description of various embodiments of the present invention has been presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been selected to best explain the principles, practical applications, or technical improvements of the embodiments over techniques found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. identifying an audio clip for a location of the user in the environment; fragmenting the audio clip into a plurality of sound clips; performing the sound attenuation on at least one sound clip from the audio clips in response to determining that at least one sound clip requires sound attenuation, wherein an attenuation ratio for the at least one sound clip is changed; In response to determining to stitch the plurality of sound clips, stitching the plurality of sound clips to form the audio clip, wherein the plurality of sound clips includes the at least one sound clip with the acoustic attenuation; and displaying a visual representation of the audio clip comprising the plurality of sound clips. A computer-implemented method comprising:
2. The computer-implemented method of claim 1 , wherein the environment is an augmented reality environment or a virtual reality environment.
3. The computer-implemented method of claim 2 , wherein the visual representation of the audio clip comprising the plurality of sound clips is displayed in an augmented reality electronic device associated with the user or a virtual reality electronic device associated with the user.
4. 2. The computer-implemented method of claim 1, further comprising determining, based on a historical corpus, that the at least one sound clip from the audio clips requires the sound attenuation, wherein the sound attenuation includes a complete attenuation of sound waves for the at least one sound clip.
5. 2. The computer-implemented method of claim 1, further comprising determining, based on a historical corpus, that the at least one sound clip from the audio clips requires the sound attenuation, wherein the sound attenuation comprises a varying attenuation of sound waves for the at least one sound clip.
6. displaying, at the augmented reality electronic device, the location of each sound clip from the plurality of sound clips of the audio clip in response to identifying the location of each sound clip from the plurality of sound clips of the audio clip relative to the location of the user; receiving, via the augmented reality electronic device, modifications to the audio clip; storing a modified audio clip for the location of the user in the augmented reality environment; and updating a historical corpus with the modified audio clip for the location of the user in the augmented reality environment; The computer-implemented method of claim 3 further comprising:
7. displaying, at the virtual reality electronic device, the location of each sound clip from the plurality of sound clips of the audio clip in response to identifying the location of each sound clip from the plurality of sound clips of the audio clip relative to the location of the user; receiving, via the virtual reality electronic device, modifications to the audio clip; storing a modified audio clip for the location of the user in the virtual reality environment; and updating a historical corpus with the modified audio clip for the location of the user in the virtual reality environment; The computer-implemented method of claim 3 further comprising:
8. 1. A computer program product comprising: One or more computer readable storage media and program instructions stored on the one or more computer readable storage media capable of carrying out a method, the method comprising: identifying an audio clip for a location of the user in the environment; fragmenting the audio clip into a plurality of sound clips; performing the sound attenuation on at least one sound clip from the audio clips in response to determining that at least one sound clip requires sound attenuation, wherein an attenuation ratio for the at least one sound clip is changed; In response to determining to stitch the plurality of sound clips, stitching the plurality of sound clips to form the audio clip, wherein the plurality of sound clips includes the at least one sound clip with the acoustic attenuation; and displaying a visual representation of the audio clip comprising the plurality of sound clips.
1. A computer program product comprising:
9. The computer program product of claim 8 , wherein the environment is an augmented reality environment or a virtual reality environment.
10. 10. The computer program product of claim 9, wherein the visual representation of the audio clip comprising the plurality of sound clips is displayed in an augmented reality electronic device associated with the user or a virtual reality electronic device associated with the user.
11. 9. The computer program product of claim 8, further comprising determining, based on a historical corpus, that the at least one sound clip from the audio clips requires the sound attenuation, wherein the sound attenuation comprises a complete attenuation of sound waves for the at least one sound clip.
12. 9. The computer program product of claim 8, further comprising determining, based on a historical corpus, that the at least one sound clip from the audio clips requires the sound attenuation, wherein the sound attenuation comprises a varying attenuation of sound waves for the at least one sound clip.
13. displaying, at the augmented reality electronic device, the location of each sound clip from the plurality of sound clips of the audio clip in response to identifying the location of each sound clip from the plurality of sound clips of the audio clip relative to the location of the user; receiving, via the augmented reality electronic device, modifications to the audio clip; storing a modified audio clip for the location of the user in the augmented reality environment; and updating a historical corpus with the modified audio clip for the location of the user in the augmented reality environment; The computer program product of claim 10, further comprising:
14. displaying, at the virtual reality electronic device, the location of each sound clip from the plurality of sound clips of the audio clip in response to identifying the location of each sound clip from the plurality of sound clips of the audio clip relative to the location of the user; receiving, via the virtual reality electronic device, modifications to the audio clip; storing a modified audio clip for the location of the user in the virtual reality environment; and updating a historical corpus with the modified audio clip for the location of the user in the virtual reality environment; The computer program product of claim 10, further comprising:
15. 1. A computer system comprising: a computer-readable storage medium for executing a method for a computer-implemented system, the computer-implemented system comprising: one or more computer processors; one or more computer-readable storage media; and program instructions stored on the one or more of the one or more computer-readable storage media for execution by at least one of the one or more processors capable of performing a method, the method comprising: identifying an audio clip for a location of the user in the environment; fragmenting the audio clip into a plurality of sound clips; performing the sound attenuation on at least one sound clip from the audio clips in response to determining that at least one sound clip requires sound attenuation, wherein an attenuation ratio for the at least one sound clip is changed; In response to determining to stitch the plurality of sound clips, stitching the plurality of sound clips to form the audio clip, wherein the plurality of sound clips includes the at least one sound clip with the acoustic attenuation; and displaying a visual representation of the audio clip comprising the plurality of sound clips. A computer system comprising:
16. The computer system of claim 15 , wherein the environment is an augmented reality environment or a virtual reality environment.
17. 17. The computer system of claim 16, wherein the visual representation of the audio clip comprising the plurality of sound clips is displayed in an augmented reality electronic device associated with the user or a virtual reality electronic device associated with the user.
18. 16. The computer system of claim 15, further comprising determining, based on a historical corpus, that the at least one sound clip from the audio clips requires the sound attenuation, wherein the sound attenuation includes a complete attenuation of sound waves for the at least one sound clip.
19. 16. The computer system of claim 15, further comprising determining, based on a historical corpus, that the at least one sound clip from the audio clips requires the sound attenuation, wherein the sound attenuation comprises a varying attenuation of sound waves for the at least one sound clip.
20. displaying, at the augmented reality electronic device, the location of each sound clip from the plurality of sound clips of the audio clip in response to identifying the location of each sound clip from the plurality of sound clips of the audio clip relative to the location of the user; receiving, via the augmented reality electronic device, modifications to the audio clip; storing a modified audio clip for the location of the user in the augmented reality environment; and updating a historical corpus with the modified audio clip for the location of the user in the augmented reality environment; 20. The computer system of claim 17, further comprising: