Audio Emulation
The system addresses the lack of audio emulation in augmented and virtual reality by using digital twin simulation and contextual information to recreate location-specific sounds, enhancing user experience and noise optimization.
Patent Information
- Application Number
- JP2023524360
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-06
- Filing Date
- 2021-10-20
- Publication Date
- 2025-11-20
- Estimated Expiration
- 2041-10-20
AI Technical Summary
Current augmented and virtual reality systems lack a comprehensive approach to emulating audio and sound effects of physical locations, failing to accurately convey typical sounds associated with an area during certain conditions or times of day.
The system utilizes digital twin simulation and contextual information to emulate audio within a generated user interface, incorporating data from IoT devices and public databases to recreate how sound reacts to a location's layout, occupancy, and environmental conditions.
Enables a realistic representation of location-specific audio, allowing users to experience accurate sound simulations and providing recommendations for optimizing noise levels or sound coverage.
Smart Images

Figure 0007773837000001 
Figure 0007773837000002 
Figure 0007773837000003
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to audio emulation, and more particularly to emulating audio using one or more Internet of Things (IoT) devices. [Background technology]
[0002] Virtual reality (VR) typically refers to a simulated experience that may resemble the real world or may be completely different from it. Uses of virtual reality can include entertainment and educational purposes. Other different types of VR technology include augmented reality and mixed reality. A person using a virtual reality device can look around an artificial world, move around the artificial world, and interact with virtual features or items. This effect is typically created by a VR headset, which consists of a head-mounted display with a small screen in front of the eyes, but can also be created by a specially designed room with multiple large screens. Virtual reality typically includes auditory and video feedback, but can also enable other types of sensory and force feedback through haptic technology.
[0003] Augmented reality (AR) generally refers to the interactive experience of a real-world environment in which real-world objects are augmented with computer-generated sensory information across multiple sensory modalities, sometimes including vision, hearing, touch, somatosensation, and smell. AR can be defined as a system that achieves three fundamental characteristics: the coupling of real and virtual worlds, real-time interaction, and precise 3D positioning of virtual and real objects. The overlaid sensory information can be constructive (i.e., additive to the natural environment) or destructive (i.e., occluding the natural environment). The experience is seamlessly interwoven with the physical world to be perceived as an immersive aspect of the real environment. In this way, augmented reality alters one's ongoing perception of a real-world environment, while virtual reality completely replaces a user's real-world environment with a simulated one.
[0004] A digital twin is a digital replica of a living or non-living physical entity. In general, a digital twin refers to a digital replica of potential and actual physical assets (physical twins), processes, people, places, systems, and equipment that can be used for a variety of purposes. This digital representation provides both the elements and dynamics of how an Internet of Things device operates and persists throughout its life cycle.
[0005] Digital twins have two key characteristics: a connection between a physical model and its corresponding virtual model or counterpart, and this connection is established by generating real-time data using sensors. Digital twins generally integrate IoT, artificial intelligence, machine learning, and software analytics with spatial network graphs to create living digital simulation models that update and change as their physical counterparts change. Digital twins continuously learn and update themselves from multiple sources to represent their near-real-time situation, operating state, or location. This learning system learns from itself using sensor data that conveys various aspects of the learning system's operating conditions from human experts, such as engineers with deep knowledge of the relevant industry domain, other similar machines, other similar machines, and the larger system and environment of which the learning system may be a part. Digital twins also integrate historical data from past machine use to incorporate into their digital models.
[0006] Virtual surround is an audio system that attempts to create the perception of more sound sources than are actually present. Most recent examples of such systems are designed to simulate a true (physical) surround sound experience using one, two, or three loudspeakers. Such systems are popular among consumers who want to enjoy a surround sound experience without the large number of speakers traditionally required for this purpose.
[0007] 3D sound effects are a group of sound effects that manipulate the sound produced by stereo speakers, surround sound speakers, speaker arrays, or headphones. They often involve virtually placing sound sources somewhere in three-dimensional space, including behind, above, or below the listener. Summary of the Invention
[0008] According to one aspect of the present invention, a computer-implemented method is provided that includes dynamically generating audio for one or more images associated with a location based on contextual information that fulfills a request, embedding the generated audio into the one or more images, and displaying the one or more images along with the embedded audio on a user device.
[0009] Preferred embodiments of the present invention will now be described, by way of example only, with reference to the following drawings: [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram of a computing environment according to an embodiment of the present invention. [Figure 2] 1 is a flowchart illustrating operational steps for creating a multivariate experience, according to an embodiment of the present invention. [Figure 3] 4 is a flowchart illustrating operational steps for generating and simulating speech in accordance with an embodiment of the present invention. [Figure 4] 1 is a block diagram of an exemplary system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0011] Embodiments of the present invention recognize the deficiencies of current augmented and virtual reality systems. Specifically, embodiments of the present invention recognize that current augmented and virtual reality systems lack a comprehensive approach to emulating the audio and sound effects of a physical location. For example, traditional augmented and virtual reality systems typically do not focus on the manner in which audio is transmitted or otherwise represented and experienced by a user. As a result, users typically lack audio presented in a recreated location and are unable to experience the audio presented in the recreated location. For example, while augmented and virtual reality systems can recreate the layout of a building (e.g., rooms in a residence), traditional augmented and virtual reality systems lack a means to show how sound (e.g., voice) reverberates within the structure and are unable to accurately convey typical sounds associated with an area during certain conditions (e.g., during a rain shower), during certain events (e.g., noise levels heard from the street), or at different times of day. Therefore, embodiments of the present invention provide a solution to the deficiencies of augmented and virtual reality systems by emulating audio within a generated user interface designed to provide a user with a realistic representation of the location. For example, embodiments of the present invention may emulate audio by using digital twin simulation and collected contextual information, as discussed in more detail later herein. For example, some embodiments of the present invention may simulate how different noise levels affect a location and generate suggestions for improving (e.g., reducing) the noise level.
[0012] As used herein, context information refers to information about a location (e.g., an intended destination). As used herein, a location refers to a physical structure having one or more structural layouts, each containing one or more objects (e.g., furniture, decorations, etc.). Examples of locations include residential structures (e.g., homes, apartment buildings, condominiums, etc.) and commercial structures (e.g., buildings in retail districts). Context information can also include materials used to construct the structure and the structure's layout (e.g., the use of wood vs. carpeted flooring, sound-deadening materials, wall thickness, etc.).
[0013] The context information may also include audio data collected from one or more Internet of Things (IoT) devices associated with the location and one or more public or otherwise accessible databases. Examples of audio data may include one or more audio files (e.g., pre-recorded sounds, such as a stored audio file library that can be reproduced and played for a particular layout or location).
[0014] The context information may also include weather data (e.g., sunshine / rain / snow, humidity, cloud index, UV index, wind, dew point, pressure, visibility, etc.), brightness (e.g., sun position), time of day, GPS location, and quantity of users in a location. The context information may further include information about objects at or near a location (e.g., geotags for certain street signs, lights, billboards, benches, etc.). In this embodiment, the weather data may be correlated with one or more audio files to simulate the weather experienced at a particular location.
[0015] Context information can also include information about a location (e.g., location information) such as building hours of operation, road closures, expected traffic based on scheduled events such as concerts, real-time traffic, queue conditions at a location such as restaurant wait times, user preferences, etc.
[0016] Embodiments of the present invention may utilize contextual information via data from cloud sources with permission from the user. For example, embodiments of the present invention may provide a user with a opt-out mechanism that allows embodiments of the present invention to collect and use information provided by the user (e.g., user-generated audio, user-uploaded images, user-generated tags, user-copyrighted images, etc.). Some embodiments of the present invention may send a notification to the user each time information is collected or otherwise used.
[0017] Figure 1 is a functional block diagram illustrating a computing environment, generally designated computing environment 100, in accordance with one embodiment of the present invention. Figure 1 illustrates only one implementation and is not intended to imply limitations with respect to the environments in which different embodiments may be implemented. Those skilled in the art may implement many modifications to the illustrated environment without departing from the scope of the present invention, as defined by the claims.
[0018] Computing environment 100 includes client computing devices 102 and server computers 108, all interconnected via network 106. Client computing devices 102 and server computers 108 can be stand-alone computing devices, management servers, web servers, mobile computing devices, or any other electronic devices or computing systems capable of receiving, transmitting, and processing data. In other embodiments, client computing devices 102 and server computers 108 can represent server computing systems that utilize multiple computers as server systems, such as in a cloud computing environment. In another embodiment, client computing devices 102 and server computers 108 can be laptop computers, tablet computers, netbook computers, personal computers (PCs), desktop computers, personal digital assistants (PDAs), smart phones, or any programmable electronic devices capable of communicating with various components and other computing devices (not shown) within computing environment 100. In another embodiment, the client computing device 102 and the server computer 108 each represent a computing system utilizing clustered computers and components (e.g., database server computers, application server computers, etc.) that function as a single pool of seamless resources when accessed within the computing environment 100. In some embodiments, the client computing device 102 and the server computer 108 are a single device. The client computing device 102 and the server computer 108 may include internal and external hardware components capable of executing machine-readable program instructions, as shown and described in further detail with respect to FIG.
[0019] In this embodiment, client computing device 102 is a user device associated with a user and includes application 104. Application 104 communicates with server computer 108 to access sound emulator 110 (e.g., using TCP / IP) to access content, user information, and database information. Application 104 can further communicate with sound emulator 110 to send instructions to generate and subsequently display a computer-rendered view including an audio simulation of the user's viewpoint and location contextually related to the current perspective, as discussed in more detail with respect to FIGS. 2-3 .
[0020] Network 106 may be, for example, a telecommunications network, a local area network (LAN), a wide area network (WAN) such as the Internet, or a combination of the three, and may include wired, wireless, or fiber optic connections. Network 106 may include one or more wired or wireless networks, or both, capable of receiving and transmitting data, voice, or video signals, including multimedia signals, including voice, data, and video information, or a combination thereof. In general, network 106 may be any combination of connections and protocols that support communication between client computing device 102 and server computer 108, as well as other computing devices (not shown) in computing environment 100.
[0021] The server computer 108 is a digital device that hosts the sound emulator 110 and the database 112. In this embodiment, the sound emulator 110 resides on the server computer 108. In other embodiments, the sound emulator 110 may comprise an instance of a program (not shown) stored locally on the client computing device 102. For example, the sound emulator 110 may be integrated with an existing augmented reality or virtual reality system installed on the client device. In other embodiments, the sound emulator 110 may be a standalone program or system that generates one or more contextually relevant interfaces for a user to experience, which display layouts and accompanying sound simulations based on received user requests. In other embodiments, the sound emulator 110 may be stored on any number of computing devices.
[0022] In this embodiment, sound emulator 110 generates and subsequently displays a computer-rendered view that includes an audio simulation of the location contextually related to the user's viewpoint and current field of view. In this embodiment, sound emulator 110 may include a digital twin system (not shown) that is used to replicate the physical location.
[0023] For example, sound emulator 110 may receive information about a location that is a three-story home and the accompanying layout of each floor of the home. Specifically, the home includes a home theater and surround sound system on one of the floors of the home. Sound emulator 110 may then recreate the received layout using a digital twin. In this embodiment, sound emulator 110 may generate or otherwise recreate the sound output by a stereo system, using the digital twin system to collect varying sound levels within the physical layout of the home, including how the sound reacts to other materials in the home. Sound emulator 110 may then adjust the volume level to account for and simulate how the sound will sound as a result of other variables, such as the number of objects placed in the room, the room's occupancy rate, and the surface materials of the walls. In some embodiments, sound emulator 110 may recreate the sound as it would sound inside a home with one or more people present in the room, and may simulate background noise both inside and outside the home (e.g., traffic, traffic density, weather, etc.).
[0024] In this embodiment, the sound emulator 110 dynamically simulates the audio of a location using the location's digital twin and integrates the simulated audio into one or more virtual and augmented reality systems. For example, the sound emulator 110 can simulate the voices of one or more users within the location's changing layout based on information received from the users. Using the digital twin system, the sound emulator 110 can passively or actively collect and subsequently play audio (e.g., background noise, rain, lightning, sounds emanating from a light source, garbled voices, etc.). In some embodiments, the sound emulator 110 can also utilize the digital twin system to recreate certain lighting effects shown in the location (e.g., mimicking a lighting setup or simulating other lighting options).
[0025] In this embodiment, the received information generally refers to a received request to simulate sounds experienced at the intended location. For example, the received information may include a request to simulate variables such as weather-related noises (e.g., rain, wind, weather-related data, etc.), voices / sounds generated by nearby locations (e.g., residences), voices experienced both inside and outside the location, rush hour and off-peak traffic sounds, etc. The received information may also include location information (e.g., building hours of operation, road closures, predicted traffic based on scheduled events such as concerts, real-time traffic, location queue conditions such as restaurant wait times, user preferences, etc.), and changes to information about the intended location (e.g., location information from cloud sources including road closures, predicted and actual traffic, and changes to business hours).
[0026] In other embodiments, the received information may be actively collected by the sound emulator 110. For example, the sound emulator 110 may invoke applications (to which the sound emulator 110 is permitted to access) such as one or more cameras, smart devices, or audio equipment located within the location, and record or otherwise capture a series of images and audio throughout different time periods. Finally, the received information may also include user-generated content and publicly available content associated with the location. Specifically, the received information may include one or more image and audio files associated with the location from one or more multiple views and respective time periods. For example, the user-generated content associated with the location may include multiple views (e.g., different angles of the same location showing multiple entry points and multiple street views) at different time periods (e.g., daytime or nighttime).
[0027] Content may include one or more of text information, pictures, audio, visual, and graphic information. Content may also include one or more files and extensions (e.g., file extensions such as .doc, .docx, .odt, .pdf, .rtf, .txt, .wpd, etc.). Content may further include audio (e.g., .m4a, .flac, .mp3, .mp4, .wave, .wma, etc.) and video / images (e.g., .jpeg, .tiff, .bmp, .pdf, .gif, etc.).
[0028] In this embodiment, sound emulator 110 can then use the received information to emulate the sounds experienced in the location. For example, sound emulator 110 generates the sound emulation by determining contextually relevant information, prioritizing the relevant information, and generating images and respective sound emulations consistent with the contextual information, as discussed in more detail with respect to FIGS. 2 and 3. For example, a user may request a sound emulation for location A during the day. In this scenario, sound emulator 110 can dynamically generate and emulate the sounds the user may experience at location A during the day. Optionally, sound emulator 110 can generate and emulate sounds for that same user to experience noise at location A during the night.
[0029] In some embodiments, the sound emulator 110 can further generate one or more images or a series of images simulating an event taking place at a location. For example, the sound emulator 110 can generate images and associated sounds including an event, such as a celebration, game night, party, or dinner event, involving one or more users. In this embodiment, the sound emulator 110 can optionally generate recommendations for placing one or more user devices to optimize sound coverage. For example, the sound emulator 110 can recognize certain objects shown with the received image of the location and target those recognized objects for determining the optimal placement to maximize sound coverage. For example, location A can include one or more speakers in its lobby. The sound emulator 110 can recognize those speakers and target those speakers for calculating an optimal placement, which determines the optimal placement of the speakers based on the number of users present and the users' current locations.
[0030] In other embodiments, the sound emulator 110 may generate one or more graphic icons associated with the identified objects (e.g., speakers or any other identifiable audio source) associated with the location. The sound emulator 110 may then display the generated one or more graphic icons for display in a virtual or augmented reality interface or otherwise overlay the generated one or more graphic icons on an image representing the location. Continuing with the above example, the sound emulator 110 may generate an icon that highlights or otherwise flags one or more speakers (e.g., identified objects) associated with the location. In response to a user selecting the generated graphic icon, the sound emulator 110 may play audio associated with the identified objects. In other embodiments, the sound emulator 110 may display (in response to a selection of the generated graphic icon) an option for optimizing object placement (e.g., if the objects are speakers, optimizing the speaker positions to provide optimal coverage for the location given the number of users present and the current user positions).
[0031] Optionally, the sound emulator 110 may then refine the generated image. In this embodiment, the sound emulator 110 may refine the image using an iterative feedback loop. For example, the sound emulator 110 may include a mechanism for soliciting feedback from the user to indicate satisfaction (e.g., whether the generated image was accurately reproduced) or dissatisfaction (e.g., whether the generated image was not accurately reproduced). The sound emulator 110 may further solicit feedback based on the user's perceived accuracy of the generated image. For example, the sound emulator 110 may solicit feedback regarding the accuracy of the colors used, the filters used, the generated graphic icons, etc. In other embodiments, the sound emulator 110 may utilize one or more IoT devices to collect user reactions and satisfaction levels (when the user provides permission to do so).
[0032] In this embodiment, the sound emulator 110 can automatically generate or otherwise identify a threshold level of noise (e.g., speech) that a user will tolerate. For example, in this embodiment, the sound emulator 110 can identify each user and categorize them based on their preferences (e.g., age group, needs, eyesight, sound level preferences, color preferences, room acoustic preferences, etc.).
[0033] The sound emulator 110 can further index the acoustic properties (i.e., acceptable levels) of detected objects and materials (e.g., wood type, marble, epoxy, lighting, fixtures, such as blinds, curtains, AC units, wall paint, etc.) shown within a location to generate recommendations for increasing or decreasing sound attenuation. For example, the sound emulator 110 can generate a score (e.g., meeting or exceeding a threshold score for noise, or both) indicating that a particular material will produce more reverberation or reverberation than the user prefers, and can subsequently recommend a different material to replace or otherwise modify the material to reduce the noise level. In this embodiment, the sound emulator 110 utilizes a numerical scale in which a higher number indicates a higher noise score and a lower number indicates a lower noise score (e.g., quieter noise). In some examples, the sound emulator 110 can further recommend one or more vendors to facilitate modifications to the location that will reduce the noise level or otherwise bring the location within an acceptable threshold score for noise levels. If the location is under design (e.g., not yet constructed and equipped with objects), the sound emulator 110 can generate recommendations for materials to use to make the location into an area that meets the user's noise level requirements.
[0034] In some embodiments, the sound emulator 110 can take into account the user's health parameters by simulating sounds (e.g., voices) that may affect the user's health and alerting the user to possible sounds that the user should be aware of (before entering the location). The sound emulator 110 can simultaneously alert the user of alternative locations that would meet the respective user's tolerance threshold score and generate recommendations to change the location to one that is within the tolerance threshold score (e.g., threshold score for noise).
[0035] Database 112 stores the received information and may represent one or more databases to which sound emulator 110 is permitted to access, or a publicly available database. Generally, database 112 may be implemented using any non-volatile storage medium known in the art. For example, database 112 may be implemented by a tape library, an optical library, one or more independent hard disk drives, or multiple hard disk drives in a redundant array of independent disks (RAID). In this embodiment, database 112 is stored on server computer 108.
[0036] FIG. 2 is a flowchart 200 illustrating operational steps for navigating a user to an intended location according to an embodiment of the present invention.
[0037] In step 202, the sound emulator 110 receives information. In this embodiment, the sound emulator 110 receives a request from the client computing device 102. In other embodiments, the sound emulator 110 may receive information from one or more other components of the computing environment 100.
[0038] In this embodiment, the information may include a request (e.g., by a user) to emulate sound for a location. The request may specify other contextual information, user preferences, location layout, or in other embodiments, sound emulator 110 may access other authorized or otherwise publicly available databases of contextual information. Examples of user preferences may include a user preference for certain audio noise levels (e.g., a preference for muted outdoor sounds such as rain, thunder, traffic, dogs barking, neighborhood noise, etc.).
[0039] The request may also include the location layout, materials used to furnish the location, and objects within the location. For example, the request may include that the location is a neighborhood containing 600 homes in 10 available models ranging from 1,500 square feet to 4,500 square feet single family homes.
[0040] At step 204, the sound emulator 110 uses the received information to simulate sound and generate one or more images. In this embodiment, the sound emulator 110 can reference existing images associated with the intended location and can utilize one or more artificial intelligence algorithms and generative adversarial networks (GANs) to modify the existing images or generate entirely new images of the intended location based on the contextual information. The sound emulator 110 can then use a digital twin to represent and recreate each received location.
[0041] In this embodiment, the sound emulator 110 can emulate or otherwise reproduce the sound for one or more generated images by prioritizing the contextual information, generating audio consistent with the contextual information, and then embedding the generated audio for display in an interface that enables the user to experience the location, as discussed in more detail with respect to FIG. 3.
[0042] For example, a user may submit a request to generate a virtual reality representation of a physical location and emulate sounds within the generated representation that mimic sounds experienced within the physical location. The sound emulator 110 may receive information (e.g., contextual information) indicating that the user prefers daytime and nighttime views. The sound emulator 110 may then modify the daytime view of location A to show what location A looks like at night and generate sounds associated with objects detected within the location to simulate how certain sounds (i.e., noises) sound in the location.
[0043] In another example, the sound emulator 110 can take into account contextual information such as snow and modify the displayed image to show what the location and associated objects at the location look like with newly fallen or removed snow, as well as generate audio that accompanies what snow removal (a noise generated outdoors) sounds like from inside the location.
[0044] In step 206, the sound emulator 110 generates an interface including simulated sounds. In this embodiment, the sound emulator 110 generates an interface including simulated sounds associated with the generated images. In this embodiment, the sound emulator 110 can generate an interface for displaying an augmented or virtual reality display. The sound emulator 110 can then display one or more dynamically generated images and accompanying emulated sounds on the user device. In examples where the sound emulator 110 has modified an image to better show an object (e.g., an illuminated object, a sign, text, etc.), the sound emulator 110 can use the generated image in place of the original image. In examples where the identified object includes sounds, the sound emulator 110 can generate a graphic icon that, when selected, can play a sound associated with the object. In examples where the object is stationary (e.g., a car parked outdoors), the sound emulator 110 can generate an icon, which can then be overlaid on the object (e.g., a parked car) to generate the engine start-up noise of the parked car.
[0045] In step 208, the sound emulator 110 refines the generated interface and simulated sounds. In this embodiment, the sound emulator 110 can refine the image using an iterative feedback loop. For example, the sound emulator 110 can include a mechanism to solicit feedback from the user to indicate satisfaction (e.g., whether the generated image was an accurate reproduction) or dissatisfaction (e.g., whether the generated image was an inaccurate reproduction). Figure 3 is a flowchart 300 illustrating operational steps for generating a context image according to an embodiment of the present invention.
[0046] The sound emulator 110 can refine the generated interface and simulated voices in response to user requests to simulate a sound at a particular time. For example, the sound emulator 110 can receive a request to simulate the noise produced by rain experienced from within a location. The sound emulator 110 can then generate voices that are contextually related to the acoustics of the location to mimic what rain would sound like if the user were physically present in the location.
[0047] In step 210, the sound emulator 110 generates recommendations. In this embodiment, the sound emulator 110 can generate recommendations based on a user profile. In some embodiments, the sound emulator 110 can generate recommendations to reduce noise levels (e.g., to meet a noise threshold) by suggesting alternative materials to use (e.g., carpet vs. hardwood floors or marble) that will minimize noise levels, or to modify the appearance of the location by suggesting different objects (e.g., furniture) that will meet the noise level requirements and the user's personal style. In yet other embodiments, the sound emulator 110 can also suggest contractors to facilitate further action (e.g., to achieve reduced noise levels).
[0048] FIG. 3 is a flowchart 300 illustrating operational steps for generating and simulating speech in accordance with an embodiment of the present invention.
[0049] In step 302, the sound emulator 110 prioritizes the context information. In this embodiment, the sound emulator 110 prioritizes the context information according to user preferences. For example, the sound emulator 110 may access user preferences, including a ranking of the user's audio preferences and concerns for one or more audio systems. For example, the sound emulator 110 may access the user's preferences to identify the user's concerns about neighborhood noise, creaking floors, and the acoustics of a room containing a home entertainment system. The sound emulator 110 may then prioritize noise and sound emulations associated with the user's concerns as part of a sound package embedded in an augmented or virtual reality interface display.
[0050] In other embodiments, the sound emulator 110 may have access to a list of standard sound preferences, including noise tolerance threshold levels for each sound (e.g., noise levels for neighborhood noise, traffic, creaks, ventilation fan noise, AC noise, weather noise, etc.). In yet other embodiments, the sound emulator 110 may use one or more artificial intelligence and machine learning algorithms to prioritize contextual information.
[0051] In step 304, the sound emulator 110 generates images and sounds consistent with the contextual information. In this embodiment, the sound emulator 110 generates images consistent with the contextual information by matching the identified contextual factors to one or more images displaying the contextual factors. For example, the sound emulator 110 may receive a request to generate an augmented reality or virtual reality display of a particular room within a location. The sound emulator 110 may receive the contextual information and subsequently prioritize the received contextual information. For example, in an example where the sound emulator 110 receives contextual information detailing objects in a room capable of producing sound and a preference for listening to nearby sounds (outside the location) while in the room, the sound emulator 110 may recreate the layout of the room, including one or more devices capable of emitting sound, and may embed appropriate sounds coming from the one or more objects. Specifically, in examples where the sound emulator 110 can identify the type of stereo / speaker, the sound emulator 110 can match the identified stereo type to the sound emitted by the stereo / speaker. The sound emulator 110 can then modify the generated sound to match the acoustics of the room (e.g., the sound emulator 110 can then modify the sound emitted from the stereo to account for the furniture in the room, the type of material used for the floor, the wall materials, e.g., wallpaper, paint, insulation, window type, etc.). In examples where stored audio exists around the location (e.g., raindrops, wind, neighborhood noise), the sound emulator 110 can retrieve the stored audio and embed the stored audio into the augmented or virtual reality display. In examples where stored audio does not exist, the sound emulator 110 can retrieve stock or default noises, simulate those noises, and simulate the interaction of the generated noise with materials associated with the location.
[0052] In examples where the sound emulator 110 does not receive a room layout in advance, the sound emulator 110 can identify objects shown in a picture of the location using an object recognition algorithm and generate corresponding images in an augmented or virtual reality display. The sound emulator 110 can then generate one or more images by utilizing one or more artificial intelligence algorithms and generative adversarial networks (GANs). For example, if a nighttime image of the location is not available, the sound emulator 110 can apply one or more filters to mimic the nighttime environment of the location and subsequently modify the image to better reveal objects (e.g., illuminated objects, signs, text, etc.).
[0053] In some embodiments, the sound emulator 110 can modify the generated images based on the context information. For example, the sound emulator 110 can modify the generated images to display daytime and nighttime views of the location that are consistent with or otherwise correspond to the physical layout of the location. In some examples, the sound emulator 110 can modify the furnishings of the location based on the context information. For example, if the sound emulator 110 mimics the furnishings of a physical room decorated with midcentury modern furnishings, the sound emulator 110 can modify the display of the room to appear furnished with contemporary furnishings. In other embodiments, the sound emulator 110 can modify the furnishings to other styles, such as modern, Scandinavian, country, bohemian, etc.
[0054] The sound emulator 110 can further modify the generated images to simulate different events at the location. As used herein, an event refers to a series of planned or unplanned activities and gatherings. For example, the sound emulator 110 can modify the generated images of a location (e.g., a room) to simulate a birthday celebration, a dinner, a game night, a move night, etc. In some examples, the sound emulator 110 can generate images of one or more users to simulate user interactions at the location and appropriately emulate the sounds associated with the generated images (e.g., to emulate the sounds of multiple users in a room and to emulate sounds emanating from one or more devices).
[0055] In step 306, the sound emulator 110 displays the generated images and sounds on the interface. In this embodiment, the sound emulator 110 displays the generated images and embeds the generated sounds associated with the generated images on an augmented reality or virtual reality device. In other embodiments, the sound emulator 110 can display the generated images and sounds on a display screen and play the associated sounds using the user's speakers.
[0056] Figure 4 illustrates a block diagram of computing system components within computing environment 100 of Figure 1 in accordance with an embodiment of the present invention. It should be understood that Figure 4 illustrates only one implementation and is not intended to imply any limitation with respect to the environments in which different embodiments may be implemented. Many modifications to the depicted environment may be implemented.
[0057] The programs described herein are identified based on the application in which they are implemented in particular embodiments of the invention. However, it should be understood that any particular program name herein is used merely as a matter of convenience and therefore should not limit the invention to the particular application identified or implied by such name or for use with the particular application identified and implied by such name.
[0058] Computing system 400 includes a communications fabric 402 that provides communications between cache 416, memory 406, persistent storage 408, communications unit 412, and input / output (I / O) interface 414. Communications fabric 402 may be implemented by any architecture designed to communicate data and / or control information between processors (e.g., microprocessors, communications processors, and network processors), system memory, peripheral devices, and any other hardware components in the system. For example, communications fabric 402 may be implemented by one or more buses or crossbar switches.
[0059] Memory 406 and persistent storage 408 are computer-readable storage media. In this embodiment, memory 406 includes random access memory (RAM). In general, memory 406 may include any suitable volatile or non-volatile computer-readable storage medium. Cache 416 may include high-speed memory that enhances performance of computer processor 404 by retaining recently and nearly recently accessed data from memory 406.
[0060] The sound emulator 110 (not shown) can be stored in persistent storage 408 and memory 406 for execution by one or more of the respective computer processors 404 via cache 416. In one embodiment, persistent storage 408 includes a magnetic hard disk drive. Alternatively, or in addition to a magnetic hard disk drive, persistent storage 408 can include a solid-state hard drive, a semiconductor memory device, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, or any other computer-readable storage medium capable of storing program instructions or digital information.
[0061] The media used by persistent storage 408 may also be removable. For example, a removable hard drive may be used for persistent storage 408. Other examples include optical and magnetic disks, thumb drives, and smart cards inserted into a drive for transfer to another computer-readable storage medium that is also part of persistent storage 408.
[0062] In these examples, the communications unit 412 provides for communication with other data processing systems or devices. In these examples, the communications unit 412 includes one or more network interface cards. The communications unit 412 may provide for communication through the use of either or both physical and wireless communications links. The sound emulator 110 may be downloaded to the persistent storage 408 through the communications unit 412.
[0063] The I / O interface 414 enables input and output of data to and from other devices, which may be connected to the client computing device, the server computer, or both. For example, the I / O interface 414 may provide connection to external devices 420, such as a keyboard, a keypad, a touch screen, or some other suitable input device or combination thereof. The external devices 420 may also include portable computer-readable storage media, such as thumb drives, portable optical or magnetic disks, and memory cards. Software and data used to implement embodiments of the present invention, such as the sound emulator 110, may be stored on such portable computer-readable storage media and loaded into persistent storage 408 via the I / O interface 414. The I / O interface 414 also connects to a display 422.
[0064] Display 422 provides a mechanism for displaying data to a user and may be, for example, a computer monitor.
[0065] The present invention may be a system, a method, or a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to carry out aspects of the present invention.
[0066] The computer-readable storage medium may be any tangible device capable of retaining and storing instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, or semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed as being a transitory signal itself, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating in a waveguide or other transmission body (e.g., a light pulse traveling in a fiber optic cable), or an electrical signal transmitted over an electrical wire.
[0067] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium into a respective computing / processing device, or can be downloaded to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to a computer-readable storage medium within the respective computing / processing device for storage.
[0068] The computer-readable program instructions for carrying out the operations of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, or state-setting data, or may be source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and traditional procedural programming languages such as the "C" programming language or the like. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or remote server. In the last scenario above, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry to carry out aspects of the present invention.
[0069] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0070] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, generate means for performing the functions / acts specified in one or more blocks of the flowchart and / or block diagram illustrations. These computer-readable program instructions may also be stored on a computer-readable storage medium, instructing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, such that the computer-readable storage medium having the instructions stored thereon comprises an article of manufacture containing instructions for performing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram illustrations.
[0071] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing device, or other device to cause the computer, other programmable device, or other device to perform a series of operational steps to produce a computer-implemented process such that the instructions, executed on the computer, other programmable device, or other device, perform the functions / operations specified in one or more blocks of the flowchart and / or block diagram illustrations.
[0072] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions that implement the specified logical function(s). In some alternative implementations, the functions shown in the blocks may be performed in an order different from that shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations or implements a combination of dedicated hardware and computer instructions.
[0073] The description of various embodiments of the present invention has been provided for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art that do not depart from the scope of the present invention. The terms used herein have been selected to best explain the principles, practical applications, or technical improvements of the embodiments over commercially available technology, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. In response to receiving a request to emulate audio for a location recreated in virtual reality, dynamically generating audio for one or more images associated with the location based on contextual information that fulfills the request; embedding the generated audio into the one or more images; emulating sounds within a generated user interface using the embedded audio to provide a realistic representation of the location, simulating the interaction of generated noises related to different times and weather conditions, traffic conditions, health conditions with material accompanying the one or more images in response to a user request for the one or more images; displaying the one or more images together with the embedded audio on a user device; and in response to receiving a feedback request indicating a perceived accuracy of the one or more images accompanying the emulated audio, improving the accuracy of the one or more images and the corresponding emulated audio to recommend alternative materials, redirect to alternative locations, and alert the user to health parameters.
11. A computer-implemented method comprising:
2. Optionally, improving said one or more images The computer-implemented method of claim 1 , further comprising:
3. dynamically generating audio for one or more images associated with the location based on contextual information satisfying the request; prioritizing the location-related contextual information; generating one or more images consistent with the contextual information; and generating audio associated with the generated one or more images consistent with the contextual information; 3. The computer-implemented method of claim 1, comprising:
4. Modifying at least one object of the identified plurality of objects based on the context information. The computer-implemented method of claim 3 further comprising:
5. indexing the identified plurality of objects based on an acoustic attribute of each identified object of the identified plurality of objects; The computer-implemented method of claim 4 further comprising:
6. generating one or more graphical icons representing at least one object of said plurality of objects for overlaying on said generated one or more images; overlaying the at least one or more generated graphic icons over a generated image of the one or more generated images displayed on the user device; and playing a sound associated with a respective object of the plurality of objects in response to selecting at least one generated graphic icon of the one or more generated graphic icons. The computer-implemented method of claim 4 or 5, further comprising:
7. generating a score indicative of a noise level associated with an object shown in each of the one or more images; and recommending an action to modify an acoustic attribute of the object in response to the generated score meeting or exceeding a threshold score for the noise level. The computer-implemented method of any one of claims 1 to 6, further comprising:
8. A computer-executable program for causing a computer system to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium storing the computer-executable program according to claim 8.
10. A method for dynamically generating audio for one or more images associated with a virtual reality location in response to receiving a request to emulate audio for the virtual reality location based on contextual information that satisfies the request; means for embedding the generated audio into the one or more images; means for emulating sounds within a generated user interface to provide a realistic representation of the location using the embedded audio, simulating the interaction of generated noises related to different times and weather conditions, traffic conditions, health conditions with material accompanying the one or more images in response to a user request for the one or more images; means for displaying the one or more images together with the embedded audio on a user device; means for, in response to receiving a feedback request indicating a perceived accuracy of the one or more images accompanying the emulated audio, improving the accuracy of the one or more images and the corresponding emulated audio to recommend alternative materials, redirect to alternative locations, and alert the user to health parameters; 1. A computer system comprising:
11. Optionally, means for improving said one or more images.
11. The computer system of claim 10, further comprising:
12. The means for dynamically generating audio for one or more images associated with a location based on context information satisfying a request comprises: means for prioritizing the location-related contextual information; means for generating one or more images consistent with the context information; means for generating audio associated with the generated one or more images consistent with the context information; 12. The computer system of claim 10 or 11, comprising:
13. means for modifying at least one object of the identified plurality of objects based on the context information; 13. The computer system of claim 12, further comprising:
14. means for indexing the identified plurality of objects based on an acoustic attribute of each identified object of the identified plurality of objects; 14. The computer system of claim 13, further comprising:
15. Furthermore, means for generating one or more graphic icons representing at least one object of said plurality of objects to overlay on said generated one or more images; means for overlaying the at least one or more generated graphic icons over a generated image of the one or more generated images displayed on the user device; means for playing a sound associated with a respective object of the plurality of objects in response to selecting at least one generated graphic icon of the one or more generated graphic icons; 14. The computer system of claim 13, further comprising:
Citation Information
Patent Citations
Virtual reality apparaus mainly based on vision
JP1995073338A
Method for creating sound environment prediction data, program for creating sound environment prediction data, and sound environment prediction system
JP2004085665A
Felt sound volume filter and sound simulation system
JP2005266639A
Sound insulating structure selecting apparatus
JP2008051935A
Imaging controller, imaging control method, program for imaging control method, and imaging apparatus
JP2013106298A