Method for enhancing communication, computer program for enhancing communication, and augmented reality system (identifying voice command boundaries)

The augmented reality system calculates and visually indicates sound boundaries to adjust speech volume, ensuring intended recipients understand while minimizing disturbance, by dynamically switching communication modes.

JP7740836B2Active Publication Date: 2025-09-17INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021199842
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-11
Filing Date
2021-12-09
Publication Date
2025-09-17
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

In crowded or noisy environments, it is challenging to adjust the volume of speech so that only the intended recipient can hear and understand, while minimizing disturbance to others.

Method used

An augmented reality system calculates sound boundaries using environmental factors and ambient noise levels to predict the maximum distance at which a communication can be understood or heard, and visually indicates these boundaries to the user, automatically adjusting communication mode if necessary.

Benefits of technology

Ensures the intended recipient understands the communication without disturbing others by dynamically switching between unassisted and assisted communication modes based on real-time environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007740836000001
    Figure 0007740836000001
  • Figure 0007740836000002
    Figure 0007740836000002
  • Figure 0007740836000003
    Figure 0007740836000003
Patent Text Reader

Abstract

To provide a method for augmenting communication, a computer program product for augmenting communication, and an augmented reality system.SOLUTION: A method for augmenting communication may comprise the steps of calculating a sound boundary within which a communication can be heard, generating a visualization of the sound boundary on an augmented reality device, and presenting the visualization on the augmented reality device. The sound boundary may represent a predicted maximum distance at which the communication can be understood.SELECTED DRAWING: Figure 7A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE The present disclosure relates to augmented reality systems, and more particularly to identifying voice command boundaries with augmented reality systems. [Background technology]

[0002] The development of the EDVAC system in 1948 is often cited as the beginning of the computer age. Since that time, computer systems have evolved into extremely complex devices. Today's computer systems typically include a combination of sophisticated hardware and software components, application programs, operating systems, processors, buses, memory, input / output devices, etc. As advances in semiconductor processing and computer architectures push performance ever higher, ever more advanced computer software has evolved to take advantage of these higher performance capabilities, resulting in today's computer systems being far more powerful than they were just a few years ago.

[0003] One application of these capabilities is augmented reality ("AR"). AR generally refers to technology that uses computer-generated material (e.g., text or graphics overlaid on a visual presentation) to augment or otherwise enhance a real-world environment. The AR presentation may be direct, such as when a user views through a transparent screen with computer-generated material overlaid on the screen. The AR presentation may also be indirect, such as a presentation of a sporting event where computer-generated graphics highlight key actions, time, score information, etc., overlaid on the game action.

[0004] The AR presentation is not necessarily viewed by the user at the same time that the visual image is being captured, nor is the AR presentation necessarily viewed in real time. For example, the AR presentation may be in the form of a snapshot that depicts a single moment in time, but the snapshot may be re-viewed by the user over a relatively long period of time. Summary of the Invention [Problem to be solved by the invention]

[0005] According to an embodiment of the present disclosure, there is a method for enhancing communications. [Means for solving the problem]

[0006] An embodiment may include calculating a sound boundary within which the communication can be heard, generating a visualization of the sound boundary on an augmented reality device, and presenting the visualization on the augmented reality device. In some embodiments, the sound boundary may represent a predicted maximum distance over which the communication can be understood.

[0007] According to an embodiment of the present disclosure, there is provided a computer program product for enhancing communications, the computer program product including a computer-readable storage medium having program instructions embodied thereon. The program instructions may be executable by a processor to cause a processor to calculate a first sound boundary representing a predicted maximum distance from which a communication can be understood, calculate a second sound boundary representing a predicted maximum distance from which the communication can be heard, and predict an intended recipient from among a plurality of people at a location. The prediction may include determining a direction of the communication and analyzing content of the communication. The program instructions may further cause the processor to determine, based on the first sound boundary and the location of the intended recipient, that the intended recipient will be unable to understand the communication, and in response, overlay a graphical indication on a view of the locale from a user's perspective indicating that the intended recipient will be unable to understand the communication and automatically electronically transmit the communication to the intended recipient. The program instructions may further cause the processor to determine that an unintended recipient may be able to hear the communication based on the second sound boundary and the location of the unintended recipient, and in response, overlay a graphical indication over a view of the locale from the user's perspective indicating that the unintended recipient may be able to hear the communication. Calculating the first sound boundary and the second sound boundary may include measuring a volume of the communication, measuring an ambient noise level in the locale, and calculating a sound attenuation rate based on one or more environmental factors of the locale.

[0008] According to an embodiment of the present disclosure, there is provided an augmented reality system. One embodiment may include a wearable frame, a processor coupled to the wearable frame, and a display coupled to the wearable frame. The processor may calculate a sound boundary within which a communication can be heard. The display may overlay a visualization of the sound boundary over a user's field of view. The sound boundary may represent a predicted maximum distance at which the communication can be understood, and the processor may determine that the intended recipient will be unable to understand the communication based on the sound boundary and the intended recipient's location.

[0009] The above summary is not intended to describe each illustrated embodiment or every implementation of the present disclosure. [Brief explanation of the drawings]

[0010] The drawings included herein are incorporated into and constitute a part of this specification. They illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the disclosure. The drawings are only illustrative of particular embodiments and are not intended to limit the disclosure.

[0011] [Figure 1] 1 illustrates an embodiment of a data processing system (DPS) consistent with some embodiments.

[0012] [Figure 2] 1 illustrates a cloud computing environment consistent with some embodiments.

[0013] [Figure 3] 1 illustrates abstraction model layers consistent with some embodiments.

[0014] [Figure 4] FIG. 1 is a perspective view of a head-mounted display system ("AR system") having an augmented reality glasses display, consistent with some embodiments.

[0015] [Figure 5A] 1 is a diagram of an AR system in operation consistent with some embodiments of the present disclosure. [Figure 5B] 1 is a diagram of an AR system in operation consistent with some embodiments of the present disclosure.

[0016] [Figure 6] 1 is a diagram of an AR system in operation consistent with some embodiments of the present disclosure.

[0017] [Figure 7A] FIG. 1 is a process flow diagram consistent with some embodiments of the present disclosure. [Figure 7B] FIG. 1 is a process flow diagram consistent with some embodiments of the present disclosure. [Figure 7C] FIG. 1 is a process flow diagram consistent with some embodiments of the present disclosure.

[0018] While the invention is susceptible to various modifications and alternative forms, specifics of which have been shown by way of example in the drawings and may be described in detail, it is to be understood, however, that it is not intended to limit the invention to the particular embodiments described. On the contrary, it is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the invention. DETAILED DESCRIPTION OF THE INVENTION

[0019] Aspects of the present disclosure relate to augmented reality systems, and more particularly to identifying voice command boundaries with an augmented reality system. The present disclosure is not necessarily limited to such applications, and various aspects of the present disclosure may be understood through the description of various examples using this context.

[0020] Some embodiments of the present disclosure may include a head-mounted AR system that allows a user to visualize their actual physical environment. Digital augmentations may be projected directly onto the user's retina, such that computer-generated materials may be presented alongside and on top of those actual physical environments. Additionally or alternatively, in some embodiments, the digital augmentations may be presented on a screen, such as a heads-up display mounted in front of the user, a virtual reality headset, the user's mobile device, etc.

[0021] In some applications of the present disclosure, the primary user may be located in a crowded or noisy environment, or both. It may be important to appropriately adjust the volume of a communication, spoken command, or a combination thereof, to ensure that the intended recipient (e.g., person, device) can clearly hear the speech. At the same time, the primary user's voice may disturb others sharing the same space, be heard by people not intended for the communication, or both. For example, a person may speak to a nearby person in a library, but others in the same space may easily hear and be distracted by the conversation. As another example, the primary user may be in a noisy environment, such as on a train or in a manufacturing facility, where speaking at a normal volume may not be sufficient for others to hear and understand the primary user.

[0022] More generally, when speaking to others in a shared space, it is often difficult to properly measure the volume of one's voice so that only the intended recipient can hear and / or understand one's speech, especially in situations where speaking louder is physically or socially difficult.Accordingly, some embodiments of the present disclosure include methods and systems that can help a primary user of an AR device understand whether the intended recipient can hear one's voice.Some embodiments of the present disclosure can also help a primary user of an AR device understand whether other people sharing the same space are disturbed by one's conversation.

[0023] Some embodiments may calculate a sound decay rate and then use the sound decay rate to predict who can hear the primary user, who will be disturbed by the primary user, or both. In some embodiments, the calculated decay rate may be location-specific. In these embodiments, the system may detect or receive from external sensors one or more environmental parameters, such as humidity, temperature, wind direction, etc., or both. In some embodiments, these predictions may also be based on the ambient noise level of the surrounding area. The AR system in these embodiments may measure the ambient noise, or receive from external sensors, or both. This decay rate may be used to calculate the maximum distance at which the primary user can be heard and understood.

[0024] In some embodiments, the AR system presents the primary user with an indicator of who can and cannot hear them. In some embodiments, the indicator may include a green or red icon superimposed on the head of a nearby person indicating that the person can hear and / or understand the primary user. In some embodiments, the indicator may include a glowing circle superimposed on the ground or a glowing cylinder superimposed in space indicating how far the primary user is likely to be heard and / or understood. Some embodiments may use a microphone integrated into the AR system to detect the loudness (e.g., in decibels) of the primary user's voice.

[0025] Some embodiments may predict who the intended recipient of a particular utterance from the primary user is and who may hear it and be disturbing. Some embodiments may use the primary user's direction of focus as input. This direction may be determined using a camera system integrated into the AR system. Some embodiments may also analyze the content of the utterance (e.g., name, command, etc.). This analysis may utilize a historical knowledge corpus that may be customized for the primary user using analyzed content such as their past utterances, social media contacts, facial recognition, etc.

[0026] Some embodiments may calculate a listening profile for the intended recipient that is customized for a particular locale using environmental parameters, an ambient noise profile, and the current distance between the two users. In some embodiments, this profile may further include modifiers for any equipment the intended recipient may use, such as hearing aids or hearing protection devices.

[0027] If the intended recipient is unlikely to hear the utterance, some embodiments may automatically use the AR system's electronic messaging capabilities to transmit and / or retransmit the primary user's utterance to the expected recipient. This may include dynamically initiating a phone call, shortwave radio broadcast, or the like, between the primary user and the intended recipient. Additionally or alternatively, some embodiments may transcribe the primary user's utterance into text format and then transmit the text to the intended recipient (e.g., as an SMS message or email). If the primary user is speaking with a group of people, at least some of whom are located beyond hearing distance, some embodiments may initiate a group call or message to a subset of the group that is outside hearing range.

[0028] Some embodiments may continuously track the distance between the primary user and the intended recipient and continuously monitor the ambient noise and environmental parameters of the location to detect a change from an audible distance to an inaudible distance. In response, some embodiments may dynamically change the communication mode from an unassisted communication mode to an assisted communication mode (e.g., phone call, SMS, etc.) and back to an unassisted communication mode. [Data Processing System]

[0029] FIG. 1 illustrates one embodiment of data processing systems (DPSs) 100a, 100b (collectively referred to herein as DPS 100) consistent with some embodiments. FIG. 1 illustrates only representative major components of DPS 100; these individual components may be much more complex than depicted in FIG. 1. In some embodiments, DPS 100 may be implemented as a personal computer; a server computer; a portable computer, such as a laptop or notebook computer, a PDA (personal digital assistant), a tablet computer, or a smartphone; a processor integrated into a larger device, such as an automobile, an airplane, a teleconferencing system, or an appliance; a smart device; or any other suitable type of electronic device. Furthermore, components other than or in addition to those illustrated in FIG. 1 may be present, and the number, type, and configuration of such components may vary.

[0030] The data processing system 100 of FIG. 1 may include multiple central processing units 110a-110d (collectively referred to as processors 110 or CPUs 110), which may be connected by a system bus 122 to a main memory unit 112, a mass storage interface 114, a terminal / display interface 116, a network interface 118, and an input / output ("I / O") interface 120. The mass storage interface 114 of this embodiment may connect the system bus 122 to one or more mass storage devices, such as a direct access storage device 140 or a readable / writable optical disk drive 142. The network interface 118 may enable the DPS 100a to communicate with another DPS 100b over the network 106. The main memory 112 may also include an operating system 124, multiple application programs 126, and program data 128.

[0031] The embodiment of DPS 100 of FIG. 1 may be a general-purpose computing device. In these embodiments, processor 110 may be any device capable of executing program instructions stored in main memory 112, which may itself be constructed of one or more microprocessors, integrated circuits, or both. In some embodiments, DPS 100 may include multiple processors and / or processing cores, as is typical in larger, more capable computer systems, while in other embodiments, computing system 100 may include only a single processor system designed to mimic a multiprocessor system, a single processor, or both. Furthermore, processor 110 may be implemented using several heterogeneous data processing systems 100, in which main processor 110 resides on a single chip along with secondary processors. As another example, processor 110 may be a symmetric multiprocessor system including multiple processors 110 of the same type.

[0032] When DPS 100 is booted, the associated processor 110 may first execute program instructions that make up operating system 124. Operating system 124 may then manage the physical and logical resources of DPS 100. These resources may include main memory 112, mass storage interface 114, terminal / display interface 116, network interface 118, and system bus 122. As with processor 110, some DPS 100 embodiments may utilize multiple system interfaces 114, 116, 118, 120 and bus 122, each of which may include its own separate, fully programmed microprocessor.

[0033] Instructions for the operating system 124 and / or application programs 126 (collectively referred to as "program code," "computer-usable program code," or "computer-readable program code") may initially reside in a mass storage device that is in communication with the processor 110 through the system bus 122. The program code in different embodiments may be embodied on different physical or tangible computer-readable media, such as the memory 112 or the mass storage device. In the example of FIG. 1, the instructions may be stored in a functional form in persistent storage on the direct-access storage device 140. These instructions may then be loaded into the main memory 112 for execution by the processor 110. However, the program code may also be located in a functional form on a computer-readable medium 142, which in some embodiments is selectively removable. The program code may be loaded or transferred to the DPS 100 for execution by the processor 110.

[0034] 1, the system bus 122 may be any device that facilitates communication between the processor 110, the main memory 112, and the interfaces 114, 116, 118, 120. Additionally, although the system bus 122 in this embodiment is a relatively simple single bus structure that provides a direct communication path between the system bus 122, other bus structures are consistent with this disclosure, including, but not limited to, point-to-point links in hierarchical, star, or web configurations, multiple hierarchical buses, parallel and redundant paths, etc.

[0035] Main memory 112 and mass storage device 140 may cooperate to store operating system 124, application programs 126, and program data 128. In some embodiments, main memory 112 may be a random-access semiconductor memory device (“RAM”) capable of storing data and program instructions. While FIG. 1 conceptually illustrates main memory 112 as a single monolithic entity, main memory 112 in some embodiments may be a more complex configuration, such as a hierarchy of caches and other memory devices. For example, main memory 112 may exist in multiple levels of caches, and these caches may be further divided by function, such as one cache holding instructions while another cache holds non-instruction data used by processors 110. Main memory 112 may also be distributed and associated with different processors 110 or sets of processors 110, as known in any of a variety of so-called non-uniform memory access (NUMA) computer architectures. Additionally, some embodiments may utilize a virtual addressing mechanism that allows DPS 100 to behave as if it has access to a large, single storage entity rather than access to multiple smaller storage entities (e.g., main memory 112 and mass storage device 140).

[0036] 1 as being contained within main memory 112 of DPS 100a, in some embodiments some or all of them may be physically located on a different computer system (e.g., DPS 100b) and may be accessed remotely, for example, via network 106. Furthermore, operating system 124, application programs 126, and program data 128 are not necessarily all contained within the exact same physical DPS 100a at the same time, and may even reside in the physical or virtual memory of another DPS 100b.

[0037] The system interface units 114, 116, 118, 120 in some embodiments may support communication with a variety of storage and I / O devices. The mass storage interface unit 114 may support the attachment of one or more mass storage devices 140, which may include rotating magnetic disk drive storage devices, solid-state storage devices (SSDs) that use integrated circuit assemblies as memory to persistently store data, typically using flash memory, or a combination of the two. Additionally, the mass storage device 140 may also include an array of disk drives configured to appear to a host as a single large storage device (commonly referred to as a RAID array), or other devices and assemblies including hard disk drives, tape (e.g., mini-DV), recordable compact discs (e.g., CD-R and CD-RW), digital versatile discs (e.g., DVD, DVD-R, DVD+R, DVD+RW, DVD-RAM), holographic storage systems, Blue Laser Discs, recordable storage media such as IBM Millipede devices, or combinations thereof.

[0038] Terminal / display interface 116 may be used to connect one or more display units 180 directly to data processing system 100. These display units 180 may be non-intelligent (i.e., dumb) terminals, such as LED monitors, or may themselves be fully programmable workstations that allow IT administrators and users to communicate with DPS 100. However, it should be noted that while display interface 116 may be provided to support communication with one or more displays 180, computer system 100 does not necessarily require a display 180, as all necessary interaction with users and other processes may occur over network 106.

[0039] Network 106 may be any suitable network or combination of networks and may support any appropriate protocol suitable for communicating data, code, or both to and from multiple DPSs 100. Accordingly, network interface 118 may be any device that facilitates such communication, whether the network connection is made using current analog or digital techniques or both, or via some future network mechanism. Suitable networks 106 include, but are not limited to, networks implemented using one or more of the “Infiniband” or IEEE (Institute of Electrical and Electronics Engineers) 802.3x “Ethernet” specifications, cellular transmission networks, wireless networks implemented with one of the IEEE 802.11x, IEEE 802.16, General Packet Radio Service (“GPRS”), Family Radio Service (FRS), or Bluetooth® specifications, ultra-wideband (“UWB”) technology as described in FCC 02-48, and the like. Those skilled in the art will appreciate that many different network and transport protocols may be used to implement network 106. The Transmission Control Protocol / Internet Protocol ("TCP / IP") suite includes preferred network and transport protocols. [Cloud Computing]

[0040] 2 illustrates one embodiment of a cloud environment suitable for an edge-enabled, scalable, and dynamic transfer learning mechanism. While this disclosure includes detailed descriptions related to cloud computing, it should be understood that implementation of the teachings recited herein is not limited to a cloud computing environment. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0041] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0042] The characteristics are as follows: On-Demand Self-Service: Cloud consumers can unilaterally provision computing capacity, such as server time and network storage, automatically as needed, without requiring human interaction with the service provider. Broad Network Access: Capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (eg, cell phones, laptops, and PDAs). Resource Pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources being dynamically allocated and reallocated according to demand. While consumers generally have no control or knowledge of the exact location of the resources provided, there is an implication of location independence in that they may be able to specify location at a higher level of abstraction (e.g., country, state, or data center). Rapid Elasticity: Capacity can be rapidly and elastically provisioned, in some cases automatically, for rapid scale out, and rapidly released for rapid scale in. To the consumer, the capacity available for provisioning often appears unlimited, and can be purchased in any quantity at any time. Measured Services: Cloud systems automatically control and optimize resource usage by leveraging measurement capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of utilized services.

[0043] The service model is as follows: Software as a Service (SaaS): The consumer is offered the ability to use a provider's applications running on a cloud infrastructure. The applications are accessible from a variety of client devices through a thin-client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings. Platform as a Service (PaaS): The ability offered to consumers is to deploy applications they create or acquire, formed using programming languages ​​and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the application hosting environment configuration. Infrastructure as a Service (IaaS): The ability offered to consumers is to provision processing, storage, network, and other basic computing resources onto which they can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but does have control over the operating systems, storage, deployed applications, and, in some cases, limited control of selected network components (e.g., host firewalls).

[0044] The deployment model is as follows: Private Cloud: Cloud infrastructure is operated solely for the organization. It can be managed by the organization or a third party and can be on-site or off-site. Community Cloud: Cloud infrastructure is shared by multiple organizations to support a specific community with shared interests (e.g., roles, security requirements, policies, and compliance considerations). It may be managed by the organization or a third party and may reside on-premises or off-premises. Public Cloud: Cloud infrastructure is made available to the general public or large industry groups and is owned by an organization that sells cloud services. Hybrid Cloud: The cloud infrastructure remains a unique entity, but is a combination of two or more clouds (private, community, or public) bound together by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data and application portability.

[0045] A cloud computing environment is service oriented with a focus on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0046] Referring now to FIG. 2, an exemplary cloud computing environment 50 is illustrated. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which local computing devices used by cloud consumers may communicate, such as, for example, a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, or an automobile computer system 54N, or combinations thereof. The nodes 10 may communicate with each other. They may be physically or virtually grouped in one or more networks (not shown), such as the private, community, public, or hybrid clouds described above, or combinations thereof. This enables the cloud computing environment 50 to provide infrastructure, platform, or software, or combinations thereof, as a service without the need for cloud consumers to maintain resources on their local computing devices. It should be understood that the types of computing devices 54A-N illustrated in FIG. 2 are for illustrative purposes only, and that the computing nodes 10 and the cloud computing environment 50 may communicate with any type of computerized device over any type of network or network-addressable connection, or both (e.g., using a web browser).

[0047] Referring now to Figure 3, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 2) is shown. It should be understood that the components, layers, and functions shown in Figure 3 are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:

[0048] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, RISC (Minimum Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and network components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0049] The virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities may be provided: virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75.

[0050] In one example, management layer 80 may provide the functions described below. Resource provisioning 81 provides dynamic procurement of computing resources and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 82 provides cost tracking as resources are utilized within the cloud computing environment and charging or billing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides cloud computing resource allocation and management so that required service levels are met. Service level agreement (SLA) planning and achievement 85 provides advance arrangement and procurement of cloud computing resources in anticipation of future requirements according to SLAs.

[0051] The workload tier 90 provides examples of functions for which a cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this tier include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and recipient prediction modules 96. [AR System]

[0052] 4 is a perspective view of a head-mounted display system including an augmented reality glasses display (“AR system 400”) consistent with some embodiments. As shown, AR system 400 may include a processor 416, a positioning device 404, a camera 406, a display 408, and a set of lenses 410 incorporated into a wearable frame 402 that a primary user may wear like regular glasses. AR system 400 may also include a DPS 100a that performs some or all of the calculations for processor 416.

[0053] The processor 403 and the positioning device 404 may cooperate to determine the physical location and field of view (based on directional orientation) of the primary user. For example, the positioning device 404 may include both a geographic locator (e.g., a Global Positioning System (GPS) device that utilizes signals from GPS satellites to determine the user / wearer's current location) and an orientation device (e.g., an accelerometer, a three-axis gravity detector, etc.) that determines the direction the user / wearer is looking (the "field of view") while wearing the AR system 400.

[0054] Camera 406 may capture an image or video of the primary user's field of view in the current locale, i.e., camera 406 may capture an electronic image of what the primary user sees while wearing AR system 400. This image or video may be displayed on display 408 along with the augmentations.

[0055] The processor 403 may generate an augmentation to overlay information onto the primary user's field of view. In one embodiment, this information may be overlaid on the display 408 so that what the person sees through the lens 410 is augmented by the overlaid information.

[0056] Display 408 may be a small display device (e.g., a video screen such as a micro LED display) that the wearer views directly, or it may be a projection device that displays images on lenses 410. In one embodiment, display 408 may present images / video from camera 406 to the user / wearer of AR system 400. In other embodiments, display 408 may be semi-transparent so that the primary user can view their location through display 408 as a backdrop to the augmentation.

[0057] 5A-5B, 6, and 7A-7C, processor 403 generates visual information about the particular locale the wearer is observing (i.e., within the field of view of the wearing vehicle), which is then communicated by and overlays the visual information on display 408. [Operating environment]

[0058] 5A and 5B are diagrams illustrating environments 500a, 500b of an AR system 400 in operation, consistent with certain embodiments of the present disclosure, and described with reference to an illustrative conversation taking place between a primary user 510 of the AR system 400 and an intended recipient 530 in a locale 550 where social norms dictate relative quiet, such as a library. One or more other people 520 who share the same locale 550 are also shown, who may be disturbed if the primary user 510 speaks too loudly to the intended recipient 530.

[0059] 5A, the AR system 400 may first detect that the primary user 510 is speaking. In response, the AR system 400 may prompt the user to identify an intended recipient 530, or may predict one or more intended recipients 530 to whom the primary user 510 is speaking, or both. This may be based on the direction the primary user is facing, or the content of the utterance, or both.

[0060] The AR system 400 may calculate two distances: (i) a shorter distance representing the maximum distance at which the primary user 510 can be clearly understood, and (ii) a longer distance representing the maximum distance at which the primary user 510 can be heard. As part of these calculations, the AR system may input the measured volume of the primary user's 510's voice, the measured volume of ambient noise in the locale 550, the estimated distances between the primary user 510, the intended recipient 530, and everyone else 520 in the locale 550, and one or more environmental factors to calculate a sound attenuation rate customized for the locale 550 using an appropriate physical model. Additionally, some embodiments may allow the primary user 510 to manually increase each distance for additional privacy protection or manually decrease each distance to ensure the primary user's 510's voice can be heard (e.g., for safety-related utterances).

[0061] The AR system 400 may use the calculated shorter distance to expand the user's field of view with a first graphic icon 525 (e.g., a green light versus a red light above the intended recipient's head) indicating whether that particular intended recipient should have been able to hear and understand the primary user's 510's speech. Additionally, the AR system 400 may use the calculated longer distance to expand the user's field of view with a second graphic indicator 535 (e.g., a glowing circle) indicating how far the primary user's 510's voice is likely to be heard. The AR system 400 may use a third graphic indicator 545 (e.g., a glowing exclamation point) to indicate that the other person(s) 520 may be obstructed.

[0062] 5A , if AR system 400 determines that one of intended recipients 530 is outside the calculated shorter distance, AR system 400 may initiate its electronic messaging function to relay the primary user's voice to intended recipient 530. In this way, the intended recipient can hear what the primary user is saying without the primary user having to increase their speaking volume to a level that could disturb other persons 520 in locale 550.

[0063] 5B, the AR system 400 may determine that the intended recipient 530 has subsequently moved closer to the primary user 510. In response, the AR system may terminate the electronic message function, allowing the primary user to speak to the intended recipient unassisted.

[0064] 6 is another diagram 600 of the AR system 400 in operation, consistent with some embodiments of the present disclosure, described with reference to an illustrative conversation between a primary user 510 and an intended recipient 530 in a locale 650 with high levels of ambient noise, such as a construction site. In this example, the primary user 510 may need to communicate important information, safety commands, etc. The AR system 400 in this example may calculate and present to the primary user 510 a graphic indicator indicating who will be able to hear the speech. This graphic indicator may be of a different style or color than that described with reference to FIGS. 5A-5B.

[0065] If the AR system 400 determines that the intended recipient 530 is outside the calculated shorter distance, the system may automatically use an electronic message to convey the utterance. As in the embodiment of Figures 5A-5B, the AR system 400 of Figure 6 may dynamically switch from unassisted mode to assisted mode and back to unassisted mode as each person 510, 530 moves within the locale 650. [Process flow diagram]

[0066] 7A-7C each illustrate a portion of a process flow diagram 700 consistent with several embodiments of the present disclosure. At operation 705, the AR system 400 may prompt the primary user 510 to opt in to enable recipient prediction and enable collection of a historical corpus. The historical corpus may then contain information useful in identifying who the intended recipient 530 is and what common interactions with the intended recipient 530 are. In some embodiments, boundary calculation and display may be implemented in software integrated with the AR system 400, while any recipient prediction and historical corpus functions may be implemented on the DPS 100 running in the cloud computing environment 50.

[0067] In response to opt-in by the primary user 510, the AR system 400 may initialize a history corpus and begin data collection at operation 710. This may include the primary user's conversation history, the identities of people nearby when those conversations occurred, social media contacts, facial recognition information, etc. Additionally, if people 520, 530 near the primary user 510 are wearing "Internet of Things" enabled devices (e.g., smart devices), those devices may be identified and used at operation 715 to better identify the intended recipient 530.

[0068] Next, the AR system 400 may initialize hardware within the AR system 400 in operation 720. This may include determining whether the hardware includes microphone, video, camera, or remote messaging (e.g., phone) capabilities, or a combination thereof.

[0069] The primary user 510 may then begin speaking. In response, at operation 722, the AR system 400 may measure the volume of the speech using a microphone. The AR system 400 may then calculate a first, shorter distance (“maximum intelligible distance”) indicating how far the speech remains intelligible (operation 725) and a second, longer distance (“maximum audible distance”) indicating how far the speech remains audible (operation 727). In some embodiments, these two calculations may involve identifying one or more environmental parameters, such as temperature, humidity, wind direction, etc., and then using an appropriate physical model to calculate how quickly speech degrades over distance. In some embodiments, these two calculations may include measuring ambient noise in the locale 550, 650, estimating how much ambient noise will interfere with understanding the speech content (e.g., similar frequencies, similar directions), and calculating how far the spoken content is audible / intelligible above the ambient noise.

[0070] At operations 730-738, the AR system 400 may use the two calculated distances to display a visual indication to the primary user 510 of who can hear and / or understand the speech. This may include using the longer calculated distance at operation 730 to create a first visual indication (e.g., a semi-transparent bounding cylinder, a cylindrical disk added to the floor surface, a hemisphere, etc.) to indicate how far spoken content will remain audible. Some embodiments may use the shorter calculated distance to create a visual indication that may visually indicate how far spoken content is likely to remain intelligible. In some embodiments, this second visual indication may be graduated (e.g., color, intensity, or both) to visually indicate the intelligibility level decreasing with distance as the volume dissipates.

[0071] At operation 732, the AR system 400 may prompt the user to identify the recipient and / or indicate how loud the speech should be at the location of the recipient 530. This information may be used to further enhance the display of the primary user 510 with a visual indication above the intended recipient 530 (e.g., a red cross or blue check superimposed over their head) with a prediction as to whether the identified recipient 500 is likely to have understood what was said.

[0072] Additionally or alternatively, some embodiments may predict intended recipients 530 in operation 734. While the primary user 510 is speaking, these embodiments may analyze the content of the utterance (e.g., names used and other factors derived from the corpus) and identify the primary user's direction of focus to predict the intended recipients of the utterance.

[0073] The AR system 400 may then mark any other persons within the calculated greater distance as unintended recipients in operation 736 and may visually indicate those unintended recipients with appropriate enhancements (e.g., red exclamation points superimposed over their heads). Additionally, some embodiments may verify what activities other persons are performing and indicate who may be disturbed. For example, if a nearby person is performing a delicate activity and may be disturbed by shouting, some embodiments may visually indicate that situation with appropriate enhancements.

[0074] At operation 738, the AR system 400 may develop and / or receive hearing profiles of the intended recipient 530 and / or other persons 520 in the locale 550, 650, and may then use those profiles to adjust the calculated distance. For example, one of the other persons 520 in the area 550, 650 may have its own AR system 400 or another IoT-enabled device and may use sensors on one of those devices to build a hearing profile of the recipient, e.g., the other person 520 and / or the intended recipient 530. Alternatively, if the AR system 400 detects a person wearing a hearing aid or hearing protection device, the AR system may increase or decrease the distance accordingly.

[0075] If the AR system 400 determines at operation 740 that the intended recipient 530's current location (predicted, indicated, or both) is beyond the calculated shorter distance, the AR system 400 may determine whether the intended recipient 530 has its own AR system 400 or another compatible messaging system (e.g., a smartwatch). In response to such a determination, the AR system 400 may automatically initiate electronic communication with that device at operation 742. In some embodiments, this may include automatically initiating a call with the intended recipient 530 so that no explicit action by the primary user 510 is required. Additionally, if the AR system determines that it is getting too loud on either side of the call, or if the AR system 400 determines that the ambient noise level in the locale 550, 650 is getting too high, the AR system 400 may automatically send a text message to the intended recipient's device.

[0076] At operation 750, the AR system 400 may continuously track the movements of all persons 510, 520, 530 in the locales 550, 650 and accordingly determine when each enters or leaves one of the hearing distances. Based on this determination, some embodiments may automatically change at operation 752 from an unassisted communication mode to an assisted communication mode (e.g., phone call, SMS), or from an assisted communication mode to an unassisted communication mode.

[0077] At operation 760, if the AR system 400 determines that the primary user 510 or intended recipient 530 is performing certain actions that cannot be disturbed, the communication of those utterances may be postponed for a later time. This may include observing the user's actions with the AR system 400 and identifying those actions by passing them through a convolutional neural network (CNN) classification system, or the like. Additionally, actions that are deemed "non-disturbing" and that are within one of the audible boundaries may be indicated by a visual warning to the primary user 510.

[0078] In operation 770, some embodiments may update the historical corpus with a log of who was the recipient of various utterances and a contextual analysis of those utterances. [Computer program product]

[0079] The present invention may be a system, method, apparatus, or computer program product, or combination thereof, at any possible level of technical detail of integration. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to perform aspects of the present invention.

[0080] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media may be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the above. A non-exhaustive list of more specific examples of computer-readable storage media may also include portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding devices such as punch cards or raised structures in grooves on which instructions are recorded, and any suitable combination of the above. Computer-readable storage media, as used herein, should not be interpreted as transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted through wires.

[0081] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device over a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.

[0082] Computer-readable program instructions for carrying out operations of the present invention may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk®, C++, or the like, or conventional procedural programming languages ​​such as the “C” programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet Service Provider). In some embodiments, to carry out aspects of the present invention, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry.

[0083] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0084] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to cause a machine, whereby the instructions executing via the processor of the computer or other programmable data processing apparatus form means for implementing the function(s) / act(s) specified in the block(s) of the flowcharts and / or block diagrams, and these computer-readable program instructions may be stored on a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium having the instructions stored thereon has an article of manufacture including instructions that implement aspects of the function(s) / act(s) specified in the block(s) of the flowcharts and / or block diagrams.

[0085] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device and cause a series of operational steps to be executed on the computer, other programmable apparatus, or other device to generate a computer-implemented process, whereby the instructions executing on the computer, other programmable apparatus, or other device implement the function / act specified in the flowchart and / or block diagram block or blocks.

[0086] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may actually be executed as a single step, may be executed concurrently or substantially concurrently, in a partially or fully overlapping manner, or the blocks may be executed in the reverse order depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or acts or executes a combination of dedicated hardware and computer instructions. [General]

[0087] Any particular program terminology used in this description is merely for convenience, and the present invention should not be limited to use solely in any particular application specified and / or implied by such terminology. Thus, for example, routines executed to implement embodiments of the present invention, whether implemented as part of an operating system or as a particular application, component, program, module, object, or sequence of instructions, may be referred to as a "program," "application," "server," or other meaningful terminology. Indeed, other alternative hardware and / or software environments may be used without departing from the scope of the present invention.

[0088] The presently described embodiments are therefore to be considered in all respects as illustrative and not restrictive, and reference should be made to the appended claims to determine the scope of the invention.

Claims

1. A method for augmenting a user's speech, comprising: calculating a first sound boundary within which the speech can be heard, the first sound boundary representing a predicted maximum distance over which the speech can be understood; generating a visualization of the first sound boundary on an augmented reality device; presenting the visualization on the augmented reality device; A method comprising:

2. The method of claim 1, further comprising determining that the intended recipient is unable to understand the utterance based on the boundary of the first sound and the position of the intended recipient.

3. The method of claim 2 , further comprising automatically electronically transmitting the utterance to the intended recipient.

4. The method of claim 3 , wherein automatically electronically transmitting the utterance comprises relaying the utterance over a cellular phone network.

5. 5. The method of claim 3 or 4, wherein automatically transmitting the utterance electronically comprises generating a transcript of the utterance and transmitting the transcript electronically.

6. The method of claim 2 , further comprising predicting the intended recipient from among a plurality of people in a locale.

7. The method of claim 6 , wherein the predicting the intended recipient comprises determining a direction of a user's attention.

8. The method of claim 6 or 7, wherein the step of predicting the intended recipient further comprises analyzing the content of the utterance.

9. The method according to claim 8, further comprising: calculating a second sound boundary, the second sound boundary representing a predicted maximum distance at which the speech can be heard; 9. The method of claim 1, further comprising determining that an unintended recipient may be able to hear the utterance based on the boundary of the second sound and a location of the unintended recipient.

10. The method of claim 9 , further comprising generating a visualization on the augmented reality device of the unintended recipient being able to hear the utterance.

11. The method of claim 1 , wherein the step of calculating the first sound boundary comprises measuring a user's volume.

12. The method of claim 11 , wherein the step of calculating the first sound boundary further comprises measuring an ambient noise level in a locale.

13. The method of claim 12 , wherein the step of calculating the first sound boundary comprises calculating a sound intensity dissipation rate based on one or more environmental factors.

14. 14. The method of claim 1, wherein the step of generating the visualization of the boundary of the first sound comprises overlaying a graphical indication on a view of a location from a user's viewpoint.

15. The processor calculating a first sound boundary representing a predicted maximum distance at which a user's speech can be understood; calculating a second sound boundary representing the maximum expected distance at which the speech can be heard; 1. A method for predicting an intended recipient from among a plurality of persons at a location, said prediction comprising: determining the direction of the speech; and a predicting step, which includes analyzing the content of the utterance; determining, based on the boundary of the first sound and the location of the intended recipient, that the intended recipient is unable to understand the utterance, and in response thereto: overlaying a graphical indication over a view of the locale from the user's perspective that indicates that the intended recipient is unable to understand the utterance; and determining whether to automatically electronically transmit the utterance to the intended recipient; determining that an unintended recipient may be able to hear the utterance based on the boundary of the second sound and the location of the unintended recipient, and in response, overlaying a graphical indication over a view of the locale from a user's perspective indicating that the unintended recipient may be able to hear the utterance; A computer program for expanding a user's utterance, the computer program comprising: The step of calculating the first note boundary and the second note boundary comprises: measuring the volume of the speech; measuring an ambient noise level in the locale; and calculating a rate of sound attenuation based on one or more environmental factors of the locale. Computer program.

16. A wearable frame and a processor coupled to the wearable frame that calculates a first sound boundary representing a predicted maximum distance within which a user's speech can be understood; a display coupled to the wearable frame, the display overlaying a visualization of the first sound boundary over a user's field of view. Augmented reality system.

17. The augmented reality system of claim 16 , further comprising a positioning device for determining the user's physical location and the field of view.

18. The processor determines that the intended recipient is unable to understand the utterance based on a boundary of the first sound and a location of the intended recipient.

17. The augmented reality system of claim 16.

19. 20. The augmented reality system of claim 18, further comprising a wireless communication interface that automatically electronically transmits the utterance to the intended recipient in response to determining that the intended recipient is unable to understand the utterance.

20. the processor calculates a second sound boundary representing a predicted maximum distance at which the speech can be heard; the processor determines that an unintended recipient may be able to hear the speech based on the boundary of the second sound and the location of the unintended recipient; the display further overlaying a visualization that the unintended recipient can hear the utterance.

20. The augmented reality system of claim 18.

Citation Information

Patent Citations

  • Information processing device and program

    JP2016187063A

  • Augmented Reality Noise Visualization

    US20200202626A1

  • Information processing device, information processing method, and program

    WO2018135304A1

  • Information processing device, information processing method, and program

    WO2020213292A1