Identify voice command boundaries

Calculating the sound attenuation rate and predicting the voice propagation range through augmented reality system, the problem of difficulty in calibration of voice loudness in crowded or noisy environments is solved, and effective speech communication and reduced disturbances are achieved.

CN114627857BActive Publication Date: 2025-06-06INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111430888.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-11
Filing Date
2021-11-29
Publication Date
2025-06-06
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

In a crowded or noisy environment, it is difficult to properly calibrate the loudness of the voice, so that the expected recipient may not be able to hear or understand the main user's voice, and may also disturb others.

Method used

Through an augmented reality system, the sound attenuation rate is calculated and who can hear and understand the voice of the main user, presents indicators of who can and cannot hear the voice to the main user, and automatically adjusts the communication mode to ensure that the expected recipient can hear it.

Benefits of technology

Effective calibration of voice loudness in a shared space is achieved, ensuring that the intended recipient can hear and understand the main user's voice while reducing disturbance to others.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627857B_ABST
    Figure CN114627857B_ABST
Patent Text Reader

Abstract

A method for enhancing communication, a computer program product for enhancing communication, and an augmented reality system. A method for enhancing communication may include calculating an acoustic boundary within which communication can be heard, generating a visualization of the acoustic boundary on an augmented reality device, and presenting the visualization on the augmented reality device. The acoustic boundary may represent a predicted maximum distance at which the communication can be understood.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to augmented reality systems; and more particularly, to identifying voice command boundaries through augmented reality systems. Background Art

[0002] The development of the EDVAC system in 1948 is often cited as the beginning of the computer age. Since that time, computer systems have evolved into extremely complex devices. Today's computer systems typically include a combination of complex hardware and software components, applications, operating systems, processors, buses, memory, input / output devices, and more. As advances in semiconductor processing and computer architecture have driven performance ever higher, even more advanced computer software has been developed to take advantage of these capabilities, resulting in today's computer systems being much more powerful than they were just a few years ago.

[0003] One application of these new capabilities is augmented reality ("AR"). AR generally refers to technology that uses computer-generated material (e.g., text or graphics overlaid on a visual presentation) to add to or enhance the real-world environment. AR presentations can be direct, such as a user viewing through a transparent screen with computer-generated material superimposed on the screen. AR presentations can also be indirect, such as a presentation of a sporting event with computer-generated graphics superimposed on the game action highlighting key moves, time and score information, etc.

[0004] The AR presentation need not be viewed by the user while the visual image is being captured, nor need the AR presentation be viewed in real time. For example, the AR presentation may be in the form of a snapshot showing a single moment in time, but the snapshot may be viewed by the user for a relatively long period of time. Summary of the invention

[0005] According to an embodiment of the present disclosure, a method for enhancing communication. One embodiment may include calculating a sound boundary within which communication can be heard, generating a visualization of the sound boundary on an augmented reality device, and presenting the visualization on the augmented reality device. In some embodiments, the sound boundary may represent a predicted maximum distance at which the communication can be understood.

[0006] According to an embodiment of the present disclosure, a computer program product for enhancing communication includes a computer-readable storage medium having program instructions embodied therein. The program instructions can be executed by a processor to enable the processor to: calculate a first sound boundary representing a predicted maximum distance at which communication can be understood, calculate a second sound boundary representing a predicted maximum distance at which communication can be heard, and predict an intended recipient from among multiple people at a location. The prediction may include determining the direction of the communication and analyzing the content of the communication. The program instructions may also enable the processor to determine that the intended recipient cannot understand the communication based on the first sound boundary and the location of the intended recipient, and in response, superimpose a graphical indication indicating that the intended recipient cannot understand the communication on a view of the site from the user's perspective, and automatically transmit the communication electronically to the intended recipient. The program instructions may also enable the processor to determine that an unintended recipient may be able to hear the communication based on the second sound boundary and the location of the unintended recipient, and in response, superimpose a graphical indication indicating that the unintended recipient can hear the communication on a view of the site from the user's perspective. Calculating the first sound boundary and the second sound boundary may include measuring the volume level of the communication, measuring the ambient noise level in the site, and calculating the sound attenuation rate based on one or more environmental factors at the site.

[0007] According to an embodiment of the present disclosure, an augmented reality system. One embodiment may include a wearable frame, a processor coupled to the wearable frame, and a display coupled to the wearable frame. The processor may calculate a sound boundary within which a communication can be heard. The display may overlay a visualization of the sound boundary onto the user's field of view. The sound boundary may represent a predicted maximum distance at which the communication can be understood, and the processor may determine that the intended recipient cannot understand the communication based on the sound boundary and the position of the intended recipient.

[0008] The above summary is not intended to describe each illustrated embodiment or every implementation of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The accompanying drawings included in this application are incorporated into the specification and form a part of the specification. They show embodiments of the present disclosure and are used to explain the principles of the present disclosure together with the specification. The accompanying drawings only illustrate certain embodiments and do not limit the present disclosure.

[0010] Figure 1 An embodiment of a data processing system (DPS) consistent with some embodiments is shown.

[0011] Figure 2 A cloud computing environment consistent with some embodiments is depicted.

[0012] Figure 3 Abstract model layers consistent with some embodiments are depicted.

[0013] Figure 4 is a perspective view of a head-mounted display system (“AR system”) with an augmented reality glass display consistent with some embodiments.

[0014] Figure 5A and Figure 5B is a diagram of an AR system in operation consistent with some embodiments of the present disclosure.

[0015] Figure 6 is a diagram of an AR system in operation consistent with some embodiments of the present disclosure.

[0016] Figure 7A-7C is a process flow diagram consistent with some embodiments of the present disclosure.

[0017] Although the present invention may have various modifications and alternative forms, its details have been shown by way of example in the accompanying drawings and can be described in detail. However, it should be understood that its purpose is not to limit the present invention to the specific embodiments described. On the contrary, the present invention covers all modifications, equivalents and substitutions that fall within the spirit and scope of the present invention. DETAILED DESCRIPTION

[0018] Aspects of the present disclosure relate to augmented reality systems; more particularly, aspects relate to identifying voice command boundaries through augmented reality systems. Although the present disclosure is not necessarily limited to such applications, various aspects of the present disclosure may be understood through discussion of various examples using this context.

[0019] Some embodiments of the present disclosure may include a head-mounted AR system that allows a user to visualize their actual physical environment. Digital augmentations may be projected directly onto the user's retina so that computer-generated material may be presented with and on top of those actual physical environments. Additionally or alternatively, in some embodiments, digital augmentations may be presented on a screen, such as a heads-up display attached to the front of the user, a virtual reality headset, the user's mobile device, etc.

[0020] In some applications of the present disclosure, a primary user may be located in a crowded and / or noisy environment. Appropriately modulating the loudness of communications and / or verbal commands may be important to ensure that the intended recipient (e.g., person, device) can clearly hear the spoken words. At the same time, the primary user's voice may disturb others sharing the same space and / or be overheard by someone who is not intended to communicate with it. For example, someone may be in a library talking to someone nearby, but others in the same space may easily hear and be distracted by the conversation. As another example, the primary user may be in a noisy environment, such as on a train or in a manufacturing workplace, where speaking at a normal volume may not be enough for others to hear and understand the primary user.

[0021] More generally, when talking with other people in a shared space, it is often difficult to properly calibrate the loudness of a person's voice, enough so that only the intended recipient can hear and / or understand what the person is saying, especially where it is physically or socially difficult for the person to speak in a louder voice. Therefore, some embodiments of the present disclosure include methods and systems that can help a primary user of an AR device understand whether (one or more) intended recipients can hear their voice. Some embodiments of the present disclosure can also help the primary user (one or more) of an AR device understand whether others sharing the same space will be disturbed by their conversation.

[0022] Some embodiments may calculate a sound attenuation rate and then use the sound attenuation rate to predict who can hear the primary user's voice and / or who will be disturbed by the primary user's voice. In some embodiments, the calculated attenuation rate may be location specific. In these embodiments, the system may detect and / or receive one or more environmental parameters such as humidity, temperature, wind flow direction, etc. from external sensors. In some embodiments, the prediction may also be based on the ambient noise level of the surrounding area. The AR system in these embodiments may measure ambient noise, and / or receive ambient noise from external sensors. The attenuation rate may be used to calculate the maximum distance at which the primary user can be heard and understood.

[0023] In some embodiments, the AR system presents an indicator to the primary user of who can and cannot hear their voice. In some embodiments, the indicator may include a green or red icon superimposed on the head of a nearby person to indicate whether that person should be able to hear and / or understand the primary user. In some embodiments, the indicator may include a glowing circle superimposed on the ground or a glowing cylinder superimposed in space that indicates how far the primary user's voice is likely to be heard and / or understood. Some embodiments may use a microphone integrated into the AR system to detect the loudness (e.g., in decibels) of the primary user's voice.

[0024] Some embodiments may predict who is the intended recipient(s) of a particular utterance from the primary user, and who may hear it and be interrupted. Some embodiments may use the direction of the primary user's focus as input. A camera system integrated into the AR system may be used to determine the direction. Some embodiments may also analyze the content of the utterance (e.g., name, command, etc.), which may utilize a historical knowledge corpus that may be customized to the primary user using analyzed content of the primary user's past utterances, social media contacts, facial recognition, etc.

[0025] Some embodiments can use environmental parameters, ambient noise profile and the current distance between two users to calculate a hearing profile of the intended recipient customized for a specific venue. In some embodiments, the profile can also include modifications to any equipment that the intended recipient can use, such as hearing aids or hearing protection devices.

[0026] If the intended recipient is unlikely to hear the utterance, some embodiments may automatically use the electronic messaging capabilities of the AR system to send and / or resend the primary user's utterance(s) to the predicted recipient(s). This may include dynamically initiating a phone call, shortwave radio broadcast, and the like between the primary user and the intended recipient(s). Additionally or alternatively, some embodiments may transcribe the primary user's utterance(s) into text format and then send the text to the intended recipient (e.g., as an SMS message or email). If the primary user is attempting to speak to a group of people, at least some of whom are out of hearing distance, some embodiments may initiate a group phone call or message to a subset of the group that is out of hearing range.

[0027] Some embodiments may continuously track the distance between the primary user and the intended recipient, and continuously monitor local ambient noise and environmental parameters to detect a change from an audible distance to an inaudible distance. In response, some embodiments may dynamically change the communication mode from an unassisted communication mode to an assisted communication mode (e.g., phone, SMS, etc.) and back to an unassisted communication mode.

[0028] Data processing system

[0029] Figure 1 One embodiment of a data processing system (DPS) 100a, 100b (generally referred to herein as DPS 100) consistent with some embodiments is shown. Figure 1 Only representative major components of DPS 100 are shown, and those individual components may have different Figure 1 In some embodiments, DPS 100 may be implemented as a personal computer; a server computer; a portable computer, such as a laptop or notebook computer, a PDA (personal digital assistant), a tablet computer, or a smart phone; a processor embedded in a larger device, such as an automobile, an airplane, a teleconferencing system, an appliance; a smart device; or any other suitable type of electronic device. Figure 1 Components shown or components in addition to those shown, and the number, type, and configuration of these components may vary.

[0030] Figure 1The data processing system 100 in FIG. 1 may include a plurality of central processing units 110 a-110 d (collectively referred to as processors 110 or CPUs 110), which may be connected to a main memory unit 112, a mass storage interface 114, a terminal / display interface 116, a network interface 118, and an input / output (“I / O”) interface 120 via a system bus 122. In this embodiment, the mass storage interface 114 may connect the system bus 122 to one or more mass storage devices, such as a direct access storage device 140 or a readable / writable optical drive 142. The network interface 118 may allow the DPS 100 a to communicate with other DPSs 100 b via a network 106. The main memory 112 may also contain an operating system 124, a plurality of application programs 126, and program data 128.

[0031] Figure 1 The DPS 100 embodiments in the present invention may be a general purpose computing device. In these embodiments, the processor 110 may be any device capable of executing program instructions stored in the main memory 112, and may itself be comprised of one or more microprocessors and / or integrated circuits. In some embodiments, the DPS 100 may include multiple processors and / or processing cores, which is typical for larger, more powerful computer systems; however, in other embodiments, the computing system 100 may include only a single processor system and / or a single processor designed to emulate a multi-processor system. In addition, the processor(s) 110 may be implemented using multiple heterogeneous data processing systems 100, with the main processor 110 present on a single chip along with secondary processors. As another illustrative example, the processor(s) 110 may be a symmetric multiprocessor system comprising multiple processors 110 of the same type.

[0032] When the DPS 100 boots up, the associated processor(s) 110 may initially execute program instructions that make up an operating system 124. The operating system 124, in turn, may manage the physical and logical resources of the DPS 100. These resources may include a main memory 112, a mass storage interface 114, a terminal / display interface 116, a network interface 118, and a system bus 122. As with the processor(s) 110, some DPS 100 embodiments may utilize multiple system interfaces 114, 116, 118, 120 and buses 122, which in turn may each include their own separate, fully programmed microprocessor.

[0033] Instructions (generally, "program code," "computer usable program code," or "computer readable program code") for operating system 124 and / or application programs 126 may initially reside in a mass storage device that communicates with processor(s) 110 via system bus 122. Program code in different embodiments may be embodied on different physical or tangible computer readable media, such as memory 112 or a mass storage device. Figure 1 In the illustrative example of , the instructions may be stored in a functional form in permanent storage on the direct access storage device 140. These instructions may then be loaded into the main memory 112 for execution by the processor(s) 110. However, in some embodiments, the program code may also reside in a functional form on a selectively removable computer readable medium 142. It may be loaded or transferred to the DPS 100 for execution by the processor(s) 110.

[0034] Continue to refer Figure 1 , the system bus 122 may be any device that facilitates communication between the processor(s) 110; the main memory 112; and the interfaces 114, 116, 118, 120. In addition, while the system bus 122 in this embodiment is a relatively simple, single bus structure that provides a direct communication path between the system bus 122, other bus structures are consistent with the present disclosure, including but not limited to hierarchical point-to-point links, star or mesh configurations, multiple hierarchical buses, parallel and redundant paths, etc.

[0035] Main memory 112 and mass storage device 140 may work in cooperation to store operating system 124, application programs 126, and program data 128. In some embodiments, main memory 112 may be a random access semiconductor memory device ("RAM") capable of storing data and program instructions. Figure 1Main memory 112 is conceptually described as a single monolithic entity, but in some embodiments main memory 112 may be a more complex arrangement, such as a hierarchy of caches and other memory devices. For example, main memory 112 may exist in multiple levels of caches, and these caches may be further divided by function, such that one cache holds instructions while another cache holds non-instruction data used by (one or more) processors 110. Main memory 112 may be further distributed and associated with different processors 110 or sets of processors 110, as is known in any of a variety of so-called non-uniform memory access (NUMA) computer architectures. In addition, some embodiments may utilize a virtual addressing mechanism that allows DPS 100 to appear as if it accesses a large single storage entity rather than multiple smaller storage entities (such as main memory 112 and mass storage devices 140).

[0036] Although operating system 124, application programs 126, and program data 128 are Figure 1 100a, but in some embodiments, some or all of them may be physically located on a different computer system (e.g., DPS 100b) and may be remotely accessed, for example, via network 106. In addition, operating system 124, application programs 126, and program data 128 need not be fully contained in the same physical DPS 100a at the same time, and may even reside in physical or virtual memory of other DPSs 100b.

[0037] In some embodiments, the system interface units 114, 116, 118, 120 may support communications with various storage devices and I / O devices. The mass storage interface unit 114 may support the attachment of one or more mass storage devices 140, which may include rotating disk drive storage devices, solid-state storage devices (SSDs) that use integrated circuit components as memory to store data persistently, typically using flash memory, or a combination of the two. In addition, the mass storage devices 140 may also include other devices and assemblies, including disk drive arrays (commonly referred to as RAID arrays) configured to appear as a single mass storage device to the host, and / or archival storage media, such as hard disk drives, magnetic tapes (e.g., mini DV), writable compact discs (e.g., CD-R and CD-RW), digital versatile discs (e.g., DVD, DVD-R, DVD+R, DVD+RW, DVD-RAM), holographic storage systems, blue laser discs, IBM Millipede devices, and the like.

[0038] Terminal / display interface 116 may be used to connect one or more display units 180 directly to data processing system 100. These display units 180 may be non-intelligent (i.e., dumb) terminals, such as LED monitors, or they may themselves be fully programmable workstations that allow IT administrators and users to communicate with DPS 100. Note, however, that while display interface 116 may be provided to support communications with one or more displays 180, computer system 100 does not necessarily require display 180, as all desired interactions with users and other processes may occur via network 106.

[0039] The network 106 may be any suitable network or combination of networks and may support any suitable protocol suitable for communicating data and / or code to / from a plurality of DPSs 100. Thus, the network interface 118 may be any device that facilitates such communication, regardless of whether the network connection is made using current analog and / or digital technology or via some future networking mechanism. Suitable networks 106 include, but are not limited to, networks implemented using one or more of the "Infiniband" or IEEE (Institute of Electrical and Electronics Engineers) 802.3x "Ethernet" specifications; cellular transmission networks; wireless networks implementing one of the IEEE 802.11x, IEEE 802.16, General Packet Radio Service ("GPRS"), FRS (Home Wireless Service), or Bluetooth specifications; ultra-wideband ("UWB") technology, such as that described in FCC 02-48; or the like. Those skilled in the art will appreciate that many different network and transmission protocols may be used to implement the network 106. The Transmission Control Protocol / Internet Protocol ("TCP / IP") suite includes suitable network and transmission protocols.

[0040] cloud computing

[0041] Figure 2 An embodiment of a cloud environment suitable for edge-enabled scalable and dynamic transfer learning mechanisms is shown. It should be understood that although the present disclosure includes a detailed description of cloud computing, the implementation of the teachings set forth herein is not limited to a cloud computing environment. Instead, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0042] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be quickly provisioned and released with minimal management effort or interaction with the provider of the service. The cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0043] Features are as follows:

[0044] On-demand self-service: Cloud consumers can unilaterally and automatically provision computing capabilities, such as server time and network storage, as needed without requiring manual interaction with the service provider.

[0045] Wide Area Network Access: Capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0046] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. There is location independence in the sense that consumers typically do not control or know the exact location of the resources provided, but are able to specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0047] Rapid elasticity: In some cases, the ability to scale out quickly and in quickly can be provisioned quickly and elastically. To the consumer, the capacity available for provisioning often appears to be unlimited and can be purchased in any quantity at any time.

[0048] Metered Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the utilized services.

[0049] The service model is as follows:

[0050] Software as a Service (SaaS): The capability provided to the consumer is to use the provider's applications running on the cloud infrastructure. The applications are accessible from a variety of client devices through a thin client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure including the network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0051] Platform as a Service (PaaS): The capability provided to consumers is to deploy consumer-created or acquired applications onto cloud infrastructure, where the applications are created using programming languages ​​and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure including networks, servers, operating systems, or storage, but has control over the deployed applications and possible configuration of the application hosting environment.

[0052] Infrastructure as a Service (IaaS): The capabilities provided to consumers are processing, storage, networking, and other basic computing resources on which consumers can deploy and run arbitrary software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but have control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0053] The deployment model is as follows:

[0054] Private Cloud: The cloud infrastructure is operated only for the organization. It can be managed by the organization or a third party and can exist inside or outside the building.

[0055] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can exist on-premises or off-premises.

[0056] Public cloud: Cloud infrastructure is available to the general public or large industrial groups and is owned by an organization that sells cloud services.

[0057] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community or public) that remain a unique entity but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0058] The cloud computing environment is service-oriented, with a focus on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure consisting of a network of interconnected nodes.

[0059] Reference now Figure 2 , depicts an illustrative cloud computing environment 50. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 with which a local computing device used by a cloud consumer can communicate, such as a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, and / or an automobile computer system 54N. The nodes 10 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as a private cloud, community cloud, public cloud, or hybrid cloud, or a combination thereof, as described above. This allows the cloud computing environment 50 to provide infrastructure, platform, and / or software as a service for which the cloud consumer does not need to maintain resources on a local computing device. It should be understood that Figure 2The types of computing devices 54A-N shown in are intended to be illustrative only, and computing node 10 and cloud computing environment 50 may communicate with any type of computerized device over any type of network and / or network addressable connection (eg, using a web browser).

[0060] Reference now Figure 3 , showing the cloud computing environment 50 ( Figure 2 ) provides a set of functional abstraction layers. It should be understood in advance that Figure 3 The components, layers, and functions shown in are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0061] The hardware and software layer 60 includes hardware and software components. Examples of hardware components include: host 61; server 62 based on RISC (Reduced Instruction Set Computer) architecture; server 63; blade server 64; storage device 65; and network and network components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0062] Virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 71; virtual storage 72; virtual networks 73, including virtual private networks; virtual applications and operating systems 74; and virtual clients 75.

[0063] In one example, the management layer 80 may provide the functionality described below. Resource provisioning 81 provides dynamic procurement of computing resources and other resources for performing tasks within a cloud computing environment. Metering and pricing 82 provides cost tracking when resources are utilized in a cloud computing environment, as well as billing or invoicing for consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 provides cloud computing resource allocation and management so that the required service levels are met. Service level agreement (SLA) planning and fulfillment 85 provides pre-scheduling and procurement of cloud computing resources, where future demand is anticipated based on the SLA.

[0064] The workload layer 90 provides examples of functions that can take advantage of the cloud computing environment. Examples of workloads and functions that can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery 93; data analysis processing 94; transaction processing 95; and recipient prediction module 96.

[0065] AR System

[0066] Figure 44 is a perspective view of a head mounted display system with an augmented reality glass display ("AR system 400") consistent with some embodiments. As shown, AR system 400 may include a processor 416, a positioning device 404, a camera 406, a display 408, and a set of lenses 410 embedded in a wearable frame 402 that a primary user may wear like regular glasses. AR system 400 may also include a DPS 100a representing processor 416 performing some or all of the computations.

[0067] The processor 403 and the positioning device 404 can cooperate to determine the physical location and field of view (based on directional orientation) of the primary user. For example, the positioning device 404 may include a geolocator (e.g., a GPS device that uses signals from global positioning system (GPS) satellites to determine the current location of the user / wearer) and an orientation device (e.g., an accelerometer, a 3-axis gravity detector, etc.) that determines the direction ("field of view") that the user / wearer is looking at when wearing the AR system 400.

[0068] Camera 406 can capture an image or video of the primary user's field of view at the current location. That is, camera 406 can capture an electronic image of whatever the primary user is looking at while wearing AR system 400. The image or video can be displayed on display 408 along with the augmentations.

[0069] Processor 403 can generate enhancement to overlay information onto the primary user's field of view. In one embodiment, the information can be overlaid on display 408 so that anything that person sees through lens 410 is enhanced with the overlaid information.

[0070] Display 408 may be a small display device (e.g., a video screen such as a micro-LED display) that the wearer views directly, or it may be a projection device that displays images onto lens 410. In one embodiment, display 408 may present images / video from camera 406 to the user / wearer of AR system 400. In other embodiments, display 408 may be semi-transparent so that the primary user may see the location through display 408 as an augmented background.

[0071] If you will refer to Figure 5A-Figure 5B , Figure 6 and Figure 7A-7C As discussed in more detail, processor 403 generates visual information related to the particular location that the wearer is viewing (i.e., within his / her field of view). This visual information is then transmitted by display 408, which overlays the visual information on display 408.

[0072] Operating Environment

[0073] Figure 5Aand Figure 5B 5 is a diagram showing an environment 500A, 500B of the AR system 400 in operation, consistent with some embodiments of the present disclosure, and described with reference to an illustrative example of a conversation occurring between a primary user 510 and an intended recipient 530 of the AR system 400 in a place 550 where social norms dictate relative quiet, such as in a library. Also depicted are one or more other people 520 sharing the same place 550 who may be disturbed if the primary user 510 speaks too loudly to the intended recipient 530.

[0074] exist Figure 5A In the example of FIG. 4 , the AR system 400 may first detect that the primary user 510 is speaking. In response, the AR system 400 may prompt the user to identify the intended recipient(s) 530 and / or may predict the intended recipient(s) 530 that the primary user 510 is speaking to. This may be based on the direction the primary user is facing and / or the content of the speech.

[0075] The AR system 400 can calculate two distances: (i) a shorter distance, which represents the maximum distance at which the primary user 510 can be clearly understood; and (ii) a longer distance, which represents the maximum distance at which the primary user 510 can be heard. As part of these calculations, the AR system can input the measured volume of the primary user's 510 voice; the measured volume of the ambient noise in the venue 550; the estimated distances between the primary user 510, the intended recipient 530, and each other person 520 at the venue 550; and one or more environmental factors to calculate a sound attenuation rate customized for the venue 550 using an appropriate physical model. In addition, some embodiments can allow the primary user 510 to manually increase the corresponding distance for additional privacy protection, or manually decrease the corresponding distance to ensure that the primary user 510 will be heard (e.g., for safety-related speech).

[0076] The AR system 400 can use the calculated shorter distance to enhance the user's field of view with a first graphical icon 525 (e.g., a green light relative to a red light above the intended recipient's head), which indicates whether the particular intended recipient should be able to hear and understand the words of the primary user 510. In addition, the AR system 400 can use the calculated longer distance to enhance the user's field of view with a second graphical indicator 535 (e.g., a glowing circle), which indicates how far the primary user's 510 voice may be heard. The AR system 400 can use a third graphical indicator 545 (e.g., a glowing exclamation point) to indicate that one of the other persons 520 may be disturbed.

[0077] exist Figure 5A, if the AR system 400 determines that one of the intended recipients 530 is outside the calculated shorter distance, the AR system 400 can initiate its electronic messaging function to relay the primary user's voice to the intended recipient 530. In this way, the intended recipient can hear what the primary user is saying without the primary user having to increase the volume of his or her speech to a level that would disturb other people 520 in the venue 550.

[0078] exist Figure 5B , the AR system 400 may determine that the intended recipient 530 has subsequently moved closer to the primary user 510. In response, the AR system may terminate the electronic messaging functionality and allow the primary user to speak with the intended recipient unassisted.

[0079] Figure 6 600 is another diagram of an AR system 400 in operation, in accordance with some embodiments of the present disclosure, and described with reference to an illustrative example of a conversation between a primary user 510 and an intended recipient 530 at a location 650 with high ambient noise levels, such as a construction site. In this illustrative example, the primary user 510 may need to communicate important information, safety commands, etc. In this illustrative example, the AR system 400 may calculate and present to the primary user 510 a graphical indicator depicting who will be able to audibly hear the utterance. The graphical indicator may be similar to the one in the reference example. Figure 5A-Figure 5B Describes the different styles or colors.

[0080] If the AR system 400 determines that the intended recipient 530 is outside the calculated shorter distance, the system can automatically use electronic messaging to communicate the utterance. Figure 5A-Figure 5B In the embodiment, when the parties 510, 530 move around the place 650, Figure 6 The AR system 400 in can dynamically switch from an unassisted mode to an assisted mode, and back to an unassisted mode.

[0081] Process Flowchart

[0082] Figure 7A-7C Each part of the process flow diagram 700 is consistent with some embodiments of the present disclosure. At operation 705, the AR system 400 can prompt the primary user 510 to opt-in to allow recipient prediction and to allow the collection of a historical corpus. The historical corpus can then contain information that helps identify who the intended recipients 530 are and what the common interactions with those intended recipients 530 are. In some embodiments, the boundary calculation and display can be implemented in software integrated into the AR system 400, while the optional recipient prediction and historical corpus features can be implemented on the DPS 100 operating in the cloud computing environment 50.

[0083] In response to the primary user 510 opting in, the AR system 400 can initialize the history corpus and begin collecting data at operation 710. This can include a history of the primary user's conversations, the identities of people who were nearby when those conversations took place, social media contacts, facial recognition information, etc. In addition, if the people 520, 530 near the primary user 510 have "Internet of Things" enabled devices (e.g., smart devices) on them, these devices can be identified at operation 715 and used to better identify the intended recipient 530.

[0084] Next, at operation 720, the AR system 400 may initialize hardware in the AR system 400. This may include determining whether the hardware includes a microphone, a camera, and / or remote messaging (eg, telephony) capabilities.

[0085] The primary user 510 may then begin speaking. In response, the AR system 400 may measure the volume of the speech using a microphone at operation 722. The AR system 400 may then calculate (at operation 725) a first shorter distance indicating how far the speech will remain intelligible (“maximum intelligible distance”), and calculate (at operation 727) a second longer distance indicating how far the speech will remain audible (“maximum audible distance”). In some embodiments, these two calculations may include identifying one or more environmental parameters, such as temperature, humidity, wind direction, etc., and then using an appropriate physical model to calculate how quickly the speech will degrade relative to distance. In some embodiments, these two calculations may include measuring the ambient noise at the location 550, 650, estimating to what extent the ambient noise will interfere with understanding of the content of the speech (e.g., similar frequencies, similar directions), and calculating how far the spoken content will be audible / understandable relative to the ambient noise.

[0086] At operations 730-738, the AR system 400 can use the two calculated distances to display a visual indication to the primary user 510 of who can and cannot hear and / or understand the speech. This can include using the calculated longer distance to create a first visual indication (e.g., a semi-transparent bounding cylinder, a disk added to the floor surface, a hemisphere, etc.) at operation 730 to show how far away the spoken content will remain audible. Some embodiments can use the calculated shorter distance to create a visual indication that can visually indicate how far away the spoken content is likely to remain intelligible. In some embodiments, the second visual indication can be gradient (e.g., in color and / or in intensity) to visually indicate the level of intelligibility that decreases with distance as the volume dissipates.

[0087] At operation 732, the AR system 400 may prompt the user to identify the recipient and / or indicate how loud the speech should be at the location of the recipient 530. This information may be used to further enhance the display of the primary user 510 with a visual indication over the intended recipient 530 (e.g., a red cross or blue checkerboard superimposed over their head) with a prediction as to whether the identified recipient 500 is likely to have understood what was said.

[0088] Additionally or alternatively, some embodiments may predict the intended recipient 530 at operation 734. As the primary user 510 speaks, these embodiments may analyze the content of the utterance (e.g., the names used, and other factors derived from the corpus) and identify the direction of the primary user's focus to predict the intended recipient(s) of the utterance.

[0089] The AR system 400 may then mark any other persons within the calculated longer distance as unintended recipients at operation 736, and may visually indicate those unintended recipients with appropriate enhancements (e.g., a red exclamation mark superimposed above their heads). Additionally, some embodiments may verify which other persons are performing activities and will display who may be interrupted. For example, if someone nearby is performing an activity that requires close attention and may be interrupted if someone is yelling, some embodiments may visually indicate that with appropriate enhancements.

[0090] At operation 738, the AR system 400 may form and / or receive hearing profiles of the intended recipient 530 and / or other persons 520 in the venue 550, 650, and then use these profiles to adjust the calculated distance. For example, if one of the other persons 520 in the area 550, 650 has its own AR system 400 or another IoT-enabled device, then sensors on one of these devices may be used to build a hearing profile of the recipient, such as the other person 520 and / or the intended recipient 530. Alternatively, if the AR system 400 detects that someone is wearing a hearing aid or hearing protection device, the AR system may increase or decrease the distance accordingly.

[0091] At operation 740, if the AR system 400 determines that the current location (predicted and / or indicated) of the intended recipient 530 is beyond the calculated shorter distance, the AR system 400 may determine whether the intended recipient 530 has its own AR system 400 or another compatible messaging system (e.g., a smartwatch). In response to this determination, at operation 742, the AR system 400 may automatically initiate electronic communication with the device. In some embodiments, this may include automatically initiating a phone call with the intended recipient 530 so that no explicit action is required by the primary user 510. In addition, if the AR system determines that the phone call will be too loud on either end, or if the AR system 400 determines that the ambient noise level in the venue 550, 650 will be too high, the AR system 400 may automatically send a text message to the intended recipient's device.

[0092] At operation 750, the AR system 400 may continuously track the movement of each person 510, 520, 530 in the venue 550, 650, and accordingly determine when each enters or leaves one of the audible distances. Based on this determination, at operation 752, some embodiments may automatically change from an unassisted communication mode to an assisted communication mode (e.g., phone, SMS), or from an assisted communication mode to an unassisted communication mode.

[0093] At operation 760, if the AR system 400 determines that certain actions are being performed that the primary user 510 or the intended recipient 530 cannot be disturbed, then the communication of these utterances can be queued for a later time. This can include observing user actions by the AR system 400 and passing these actions through a convolutional neural network classification system, etc. to identify what these actions are. In addition, actions that are considered "do not disturb" and are within one of the audible boundaries can be indicated to the primary user 510 through a visual warning.

[0094] At operation 770 , some embodiments may update the historical corpus with a log of who the recipient(s) of various utterances were and contextual analysis of those utterances.

[0095] Computer program product

[0096] The present invention may be a system, method and / or computer program product at any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon, and the computer-readable program instructions are used to cause a processor to perform various aspects of the present invention.

[0097] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punch card or a raised structure in a groove on which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be interpreted as a temporary signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.

[0098] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.

[0099] The computer-readable program instructions for performing the operation of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data of an integrated circuit, or source code or object code written in any combination of one or more programming languages ​​(including object-oriented programming languages, such as Smalltalk, C++, etc.) and procedural programming languages ​​(such as "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider via the Internet). In some embodiments, in order to perform various aspects of the present invention, an electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute a computer-readable program instruction by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit.

[0100] Various aspects of the present invention are described herein with reference to the flow chart and / or block diagram of the method, device (system) and computer program product according to embodiments of the present invention. It will be understood that each frame of the flow chart and / or block diagram and the combination of frames in the flow chart and / or block diagram can be implemented by computer-readable program instructions.

[0101] These computer-readable program instructions can be provided to a processor of a computer or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can guide the computer, programmable data processing device and / or other equipment to work in a specific manner, so that the computer-readable storage medium having the instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0102] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0103] Flowcharts and block diagrams in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention.In this regard, each frame in the flow chart or block diagram can represent a module, segment or part of an instruction, which includes one or more executable instructions for realizing the specified logical function.In some alternative embodiments, the function noted in the frame may not occur in the order noted in the figure.For example, two frames shown continuously can actually be implemented as a step, and are performed simultaneously, substantially simultaneously, in a partially or entirely time-overlapping manner, or these frames can sometimes be performed in reverse order, depending on the function involved.It will also be noted that the combination of the frames in each frame of the block diagram and / or flow chart illustration and the block diagram and / or flow chart illustration can be implemented by a dedicated hardware-based system that performs a specified function or action or performs a combination of special-purpose hardware and computer instructions.

[0104] overall

[0105] Any specific program terminology used in this specification is for convenience only, and thus the present invention should not be limited to use only in any specific application identified and / or implied by such terminology. Thus, for example, the routines executed to implement embodiments of the present invention, whether implemented as part of an operating system or as a specific application, component, program, module, object, or sequence of instructions, may be referred to as "programs," "applications," "servers," or other meaningful terms. Indeed, other alternative hardware and / or software environments may be used without departing from the scope of the present invention.

[0106] It is therefore intended that the embodiments described herein be considered in all respects as illustrative and not restrictive, and that reference be made to the appended claims to determine the scope of the present invention.

Claims

1. A method for enhancing communication, include: Calculate the acoustic boundaries within which communications can be heard; generating a visualization of the sound boundary on an augmented reality device; as well as presenting the visualization on an augmented reality device, wherein the sound boundary comprises a first sound boundary representing a predicted maximum distance at which the communication can be heard; and the method further comprises determining whether the unintended recipient can hear the communication based on the first sound boundary and a position of the unintended recipient.

2. The method according to claim 1, in, The acoustic boundaries also include a second acoustic boundary representing a predicted maximum distance at which the communication can be understood; and the method further includes determining that the intended recipient cannot understand the communication based on the second acoustic boundary and the location of the intended recipient.

3. The method of claim 2, further comprising automatically electronically transmitting the communication to the intended recipient.

4. The method according to claim 3, in, Automatically electronically transmitting the communication includes relaying the communication over a cellular telephone network.

5. The method according to claim 3, in, Automatically electronically transmitting the communication includes generating a transcript of the communication and electronically transmitting the transcript.

6. The method of claim 2, further comprising predicting the intended recipient from a plurality of persons in a venue.

7. The method according to claim 6, in, Predicting the intended recipient includes determining the direction of the user's attention.

8. The method according to claim 7, in, Predicting the intended recipient also includes analyzing the content of the communication.

9. The method of claim 1, further comprising generating, on the augmented reality device, a visualization in which the unintended recipient is able to hear the communication.

10. The method according to claim 1, in, Calculating the sound boundary includes measuring a volume level of a user.

11. The method according to claim 10, in, Calculating the sound boundary also includes measuring the level of ambient noise in the venue.

12. The method according to claim 11, in, Calculating the sound boundary includes calculating a sound intensity dissipation rate based on one or more environmental factors.

13. The method according to claim 1, in, Generating the visualization of the sound boundary includes overlaying a graphical indication over a view of the location from a user's perspective.

14. A computer program product for enhancing communications, the computer program product comprising program instructions executable by a processor to cause the processor to: calculating a first acoustic boundary, the first acoustic boundary representing a predicted maximum distance at which communications can be understood; calculating a second sound boundary, the second sound boundary representing a predicted maximum distance at which the communication can be heard; Predicting an intended recipient from among a plurality of persons in a location, wherein The forecasts include: determining a direction of the communication; and analyzing the content of said communications; determining that the intended recipient is unable to understand the communication based on the first sound boundary and the location of the intended recipient, and in response: superimposing a graphical indication on the view of the venue from the user's perspective indicating that the intended recipient cannot understand the communication; and automatically electronically transmitting the communication to the intended recipient; and determining whether the communication is audible to the unintended recipient based on a second sound boundary and the location of the unintended recipient, and in response, superimposing a graphical indication on a view of the venue from a user's perspective indicating that the communication is audible to the unintended recipient; The calculating of the first sound boundary and the second sound boundary comprises: Measuring the volume level of communications; measuring the level of ambient noise in the location; and A sound attenuation rate is calculated based on one or more environmental factors at the location.

15. An augmented reality system, include: Wearable frame; a processor coupled to the wearable frame, wherein the processor calculates acoustic boundaries within which communications can be heard; as well as a display coupled to the wearable frame, wherein the display superimposes a visualization of the sound boundary onto a user's field of view, in, The acoustic boundaries include a first acoustic boundary representing a predicted maximum distance at which the communication can be heard; The processor determines whether the communication can be heard by the unintended recipient based on the first sound boundary and the location of the unintended recipient.

16. The system of claim 15, further comprising a positioning device for determining the user's physical location and the field of view.

17. The system according to claim 15, in: The acoustic boundaries also include a second acoustic boundary representing a predicted maximum distance at which the communication can be understood; and The processor determines that the intended recipient is unable to understand the communication based on the second sound boundary and the location of the intended recipient.

18. The system of claim 17, further comprising a wireless communication interface that automatically electronically transmits the communication to the intended recipient in response to determining that the intended recipient is unable to understand the communication.

19. The system according to claim 17, in: The display also overlays a visualization that the unintended recipient can hear the communication.

Citation Information

Patent Citations

  • Wearable electronic device for audio imaging and operating method thereof

    CN108594988A

  • Spatial audio for three-dimensional data sets

    CN110651248A