Speech Transcription Using Multiple Data Sources

By combining speech recognition and visual pattern recognition technology, the speech clips in multi-person interactive scenarios and the speaker are identified, the accuracy and naturalness of multi-person speech transcription in the prior art are solved, and efficient speech transcription and additional data generation are achieved.

CN114981886BActive Publication Date: 2025-06-17CTRL-LABS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080079550.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-20
Filing Date
2020-10-31
Publication Date
2025-06-17
Estimated Expiration
2040-10-31

AI Technical Summary

Technical Problem

Existing speech transcription systems are difficult to effectively transcribe multi-person voice, especially in multi-person interaction scenarios, and lack support for accurate identification and transcription of speaker identity.

Method used

Using a system combining speech recognition, speaker identification and visual pattern recognition technology, multiple voice fragments are identified through the capture and analysis of audio and image data, and the speaker associated with each voice fragment is identified based on the image data, complete transcription is generated and additional data is generated.

Benefits of technology

It realizes the accuracy and nature of speech transcription in multi-person interaction scenarios, and provides additional data, such as calendar invitations, topic information, task lists, etc., which improves the user's interactive experience with the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114981886B_ABST
    Figure CN114981886B_ABST
Patent Text Reader

Abstract

The present disclosure describes transcribing speech using audio, images, and other data. A system is described that includes an audio capture system configured to capture audio data associated with multiple speakers, an image capture system configured to capture images of one or more of the multiple speakers, and a speech processing engine. The speech processing engine can be configured to: identify multiple speech segments in the audio data, identify, for each of the multiple speech segments and based on the images, the speaker associated with the speech segment; transcribe each of the multiple speech segments to produce a transcription of the multiple speech segments, where for each of the multiple speech segments, the transcription includes an indication of the speaker associated with the speech segment; and analyze the transcription to produce additional data derived from the transcription.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to speech transcription systems and, more particularly, to transcribing the speech of multiple people. Background Art

[0002] Speech recognition is becoming increasingly popular and is being added to televisions (TVs), computers, tablets, smart phones, and speakers more and more. For example, many smart devices can perform services based on commands or questions spoken by a user. Such devices use speech recognition to identify the user's commands and questions based on the captured audio and then perform an operation or identify response information. Summary of the Invention

[0003] Generally, the present disclosure describes systems and methods for transcribing speech using audio, images, and other data. In some examples, a system can combine speech recognition, speaker identification, and visual pattern recognition techniques to produce a complete transcription of an interaction between two or more users. For example, such a system can capture audio data and image data, identify multiple speech segments in the audio data, identify the speaker associated with each speech segment based on the image data, and transcribe each speech segment of the multiple speech segments to produce a transcription that includes an indication of the speaker associated with each speech segment. In some examples, an artificial intelligence (AI) / machine learning (ML) model can be trained to recognize and transcribe speech from one or more identified speakers. In some examples, the system can identify speech and / or identify a speaker based on detecting one or more faces in the image data having moving lips. Such a system can also analyze the transcription to generate additional data from the transcription, including calendar invitations for a meeting or event described in the transcription, information related to a topic identified in the transcription, a task list including tasks identified in the transcription, a summary, notifications (e.g., to a person(s) not present in the interaction, to a user regarding a topic or person discussed in the interaction), statistical data (e.g., the number of words spoken by a speaker, the tone of a speaker, information regarding filler words used by a speaker, the percentage of time each speaker speaks, information regarding swear words used, information regarding the length of words used, the number of times "filler words" are used, the volume of a speaker, or the mood of a speaker, etc.). In some examples, speech transcription is performed while the speech, conversation, or interaction is occurring near real-time or seemingly near real-time. In some other examples, speech transcription is performed after the speech, conversation, or interaction has terminated.

[0004] In some examples, the techniques described herein are performed by a head-mounted display (HMD) or by a computing device having an image capture device (e.g., a camera) for capturing image data and an audio capture device (e.g., a microphone) for capturing audio data. In some examples, the HMD or the computing device may transcribe all of the speech segments in the speech segments captured for each user during an interaction between users. In some other examples, the HMD may transcribe only the speech segments for the user wearing the HMD, and the HMD, the computing device, and / or the transcription system may optionally combine the individual transcriptions received from other HMDs and / or computing devices.

[0005] According to a first aspect of the present invention, there is provided a system comprising: an audio capture system configured to capture audio data associated with a plurality of speakers; an image capture system configured to capture an image of one or more of the plurality of speakers; and a speech processing engine configured to: identify a plurality of speech segments in the audio data, identify a speaker associated with each speech segment of the plurality of speech segments and based on the image, transcribe each speech segment of the plurality of speech segments to produce a transcription of the plurality of speech segments, the transcription including an indication of the speaker associated with the speech segment for each speech segment of the plurality of speech segments, and analyze the transcription to produce additional data derived from the transcription.

[0006] To identify the plurality of speech segments, the speech processing engine may also be configured to identify the plurality of speech segments based on the image.

[0007] To identify a speaker for each speech segment of the plurality of speech segments, the speech processing engine may also be configured to detect one or more faces in the image.

[0008] The speech processing engine may also be configured to select one or more speech recognition models based on the identity of the speaker associated with each speech segment.

[0009] To identify a speaker for each speech segment of the plurality of speech segments, the speech processing engine may also be configured to detect one or more faces in the image having moving lips.

[0010] The speech processing engine may also be configured to access external data. To identify a speaker for each speech segment of the plurality of speech segments, the speech processing engine may also be configured to identify the speaker based on the external data.

[0011] The external data may include one or more of calendar information and location information.

[0012] The system may also include a head-mounted display (HMD) wearable by a user. One or more speech recognition models may include a voice recognition model for the user. The speech processing engine may also be configured to identify the user of the HMD as the speaker of the plurality of speech segments based on attributes of the plurality of speech segments. The HMD may be configured to output artificial reality content. The artificial reality content may include a virtual meeting application that includes a video stream and an audio stream.

[0013] The audio capture system may include a microphone array.

[0014] The additional data may include one or more of the following: a calendar invitation for the meeting or event described in the transcription, information related to the topic identified in the transcription, and / or a task list including the tasks identified in the transcription.

[0015] The additional data may include at least one of the following: statistical data about the transcription including the number of words spoken by the speaker, the tone of the speaker, information about the filler words used by the speaker, the percentage of time the speaker speaks, information about the swear words used, information about the length of the words used, a summary of the transcription, or the mood of the speaker.

[0016] The additional data may include an audio stream that includes a modified version of the speech segments associated with at least one of the plurality of speakers.

[0017] The method may further include: accessing external data; and for each of the plurality of speech segments, identifying the speaker based on the external data. The external data may include one or more of calendar information and location information.

[0018] The additional data may include one or more of the following: a calendar invitation for the meeting or event described in the transcription, information related to the topic identified in the transcription, and / or a task list including the tasks identified in the transcription.

[0019] The additional data may include at least one of the following: statistical data about the transcription including the number of words spoken by the speaker, the tone of the speaker, information about the filler words used by the speaker, the percentage of time the speaker speaks, information about the swear words used, information about the length of the words used, a summary of the transcription, or the mood of the speaker.

[0020] According to a second aspect of the present invention, there is provided a method comprising: capturing audio data associated with a plurality of speakers; capturing images of one or more of the plurality of speakers; identifying a plurality of speech segments in the audio data; identifying, for each speech segment of the plurality of speech segments and based on the images, the speaker associated with the speech segment; transcribing each speech segment of the plurality of speech segments to produce a transcription of the plurality of speech segments, the transcription including an indication of the speaker associated with the speech segment for each speech segment of the plurality of speech segments; and analyzing the transcription to produce additional data derived from the transcription.

[0021] According to a third aspect of the present invention, there is provided a computer-readable storage medium comprising instructions which, when executed, configure the processing circuitry of a computing system to: capture audio data associated with a plurality of speakers; capture images of one or more of the plurality of speakers; identify a plurality of speech segments in the audio data; identify, for each speech segment of the plurality of speech segments and based on the images, the speaker associated with the speech segment; transcribe each speech segment of the plurality of speech segments to produce a transcription of the plurality of speech segments, the transcription including an indication of the speaker associated with the speech segment for each speech segment of the plurality of speech segments; and analyze the transcription to produce additional data derived from the transcription.

[0022] These techniques have various technical advantages and practical applications. For example, techniques according to one or more aspects of the present disclosure can provide a speech transcription system that can generate additional data from the transcription. By automatically generating additional data, a system according to the techniques of the present disclosure can provide services to a user without the user having to say a specific word (e.g., a "wake" word) that signals to the system that a command or question has been spoken or will be spoken, and where there may be no specific command or instruction. This can facilitate interaction between the user and the system, making the interaction more consistent with the way the user might interact with another user, and thus making the interaction with the system more natural.

[0023] Details of one or more examples of the techniques of the present disclosure are set forth in the accompanying drawings and the following description. Other features, objects, and advantages of these techniques will become apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1A is a diagram depicting an example system that performs speech transcription according to the techniques of the present disclosure.

[0025] Figure 1B is a diagram depicting an example system that performs speech transcription according to the techniques of the present disclosure.

[0026] Figure 1CIs an illustration depicting an example system that performs speech transcription according to the techniques of the present disclosure.

[0027] Figure 2A Is an illustration depicting an example HMD according to the techniques of the present disclosure.

[0028] Figure 2B Is an illustration depicting an example HMD according to the techniques of the present disclosure.

[0029] Figure 3 Is an illustration depicting an example in which speech transcription is performed by an example instance of an HMD of an artificial reality system in which Figure 1A and Figure 1B of the present disclosure.

[0030] Figure 4 Is a block diagram showing an example implementation in which speech transcription is performed by an example instance of an HMD and a transcription system of an artificial reality system in which Figure 1A and Figure 1B of the present disclosure.

[0031] Figure 5 Is a block diagram showing an example implementation in which speech transcription is performed by an example instance of a computing device of a system in which Figure 1C of the present disclosure.

[0032] Figure 6 Is a flowchart illustrating example operations of a method for transcribing and analyzing speech according to aspects of the present disclosure.

[0033] Figure 7 Illustrates audio data and transcription according to the techniques of the present disclosure.

[0034] Figure 8 Is a flowchart illustrating example operations of a method for transcribing speech according to aspects of the present disclosure.

[0035] Figure 9 Is a flowchart illustrating example operations of a method for identifying a speaker of a speech segment according to aspects of the present disclosure.

[0036] Figure 10 Is a flowchart illustrating example operations of a method for identifying a potential speaker model according to aspects of the present disclosure.

[0037] Figure 11 Is a flowchart illustrating example operations of a method for transcribing speech for distributed devices according to aspects of the present disclosure.

[0038] Throughout the figures and the description, like reference numerals refer to like elements. Detailed Description

[0039] Figure 1A FIG. 10A is a diagram depicting a system 10A that performs speech transcription according to the techniques of the present disclosure. In Figure 1A this example, the system 10A is an artificial reality system that includes a head-mounted device (HMD) 112. As shown, the HMD 112 is typically worn by a user 110 and includes an electronic display and optical components for presenting artificial reality content 122 to the user 110. Additionally, for example, the HMD 112 includes one or more motion sensors (e.g., accelerometers) for tracking the motion of the HMD 112, one or more audio capture devices (e.g., microphones) for capturing audio data of the surrounding physical environment, and one or more image capture devices (e.g., cameras, infrared (IR) detectors, Doppler radar, line scanners) for capturing image data of the surrounding physical environment. The HMD 112 is illustrated as communicating with a transcription system 106 via a network 104, and the transcription system 106 can correspond to any form of computing resource. For example, the transcription system 106 can be a physical computing device or can be a component of a cloud computing system, server farm, and / or server cluster (or a portion thereof) that provides services to client devices and other devices or systems. Thus, the transcription system 106 can represent one or more physical computing devices, virtual computing devices, virtual machines, containers, and / or other virtualized computing devices. In some example implementations, the HMD 112 operates as a stand-alone mobile artificial reality system.

[0040] The network 104 can be the Internet or can include or represent any public or private communication network or other network. For example, the network 104 can be or can include cellular, Wi-Fi , ZigBee, Bluetooth, near field communication (NFC), satellite, enterprise, service provider, and / or other types of networks capable of transferring traffic data between computing systems, servers, and computing devices. One or more of the client devices, server devices, or other devices can use any suitable communication technology to transmit and receive data, commands, control signals, and / or other information across the network 104. The network 104 can include one or more network hubs, network switches, network routers, satellite antennas, or any other network equipment. Such devices or components can be operatively coupled to each other to provide an exchange of information between computers, devices, or other components (e.g., between one or more client devices or systems and one or more server devices or systems). Figure 1B Each device or system illustrated in FIG. 10A can be operatively coupled to the network 104 using one or more network links.

[0041] Generally, the artificial reality system 10A uses information captured from the real-world 3D physical environment to render the artificial reality content 122 for display to the user 110. InFigure 1A In the example of Figure 1A , user 110 views artificial reality content 122 constructed and rendered by an artificial reality application executed on HMD 112. The artificial reality content 122A can correspond to content rendered according to a virtual or video conferencing application, a social interaction application, a mobile instruction application, an alternative world application, a navigation application, an educational application, a gaming application, a training or simulation application, an augmented reality application, a virtual reality application, or other types of applications that implement artificial reality. In some examples, the artificial reality content 122 can include a mixture of real-world images and virtual objects, such as mixed reality and / or augmented reality.

[0042] During operation, the artificial reality application constructs the artificial reality content 122 for display to the user 110 by tracking and calculating the pose information of a reference frame, which is typically the viewing perspective of the HMD 112. Using the HMD 112 as the reference frame and based on the current field of view 130 determined by the current estimated pose of the HMD 112, the artificial reality application renders 3D artificial reality content that, in some examples, can be at least partially overlaid on the real-world 3D physical environment of the user 110. During this process, the artificial reality application uses the sensing data received from the HMD 112 (such as movement information and user commands) and, in some examples, data from any external sensors (such as an external camera) to capture 3D information within the real-world physical environment, such as the movement of the user 110. Based on the sensing data, the artificial reality application determines the current pose of the reference frame for the HMD 112 and renders the artificial reality content 122 according to the current pose of the HMD 112.

[0043] More specifically, as further described herein, the image capture device of the HMD 112 captures image data that represents objects in the real-world physical environment within the field of view 130 of the image capture device 138. These objects can include persons 101A and 102A. The field of view 130 typically corresponds to the viewing perspective of the HMD 112.

[0044] Figure 1A Depicts a scenario where the user 110 interacts with persons 101A and 102A. Both persons 101A and 102A are within the field of view 130 of the HMD 112, allowing the HMD 112 to capture the audio data and image data of persons 101A and 102A. The HMD 112A can display persons 101B and 102B in the artificial reality content 122 to the user 110, corresponding to persons 101A and 102A respectively. In some examples, persons 101B and / or 102B can be unchanged images of persons 101A and 102A respectively. In other examples, persons 101B and / or person 102B can be avatars (or any other virtual representation) corresponding to persons 101B and / or person 102B.

[0045] In Figure 1A the example shown, user 110 says "Hello Jack and Steve. How’s it going?" and person 101A responds "Where is Mary?" During this scenario, HMD 112 captures image data and audio data, and a speech processing engine (not shown) of HMD 112 can be configured to identify speech segments in the captured audio data and identify the speaker associated with each speech segment. For example, the speech processing engine can identify the speech segments "Hello Jack and Steve. How’s it going" and "Where is Mary?" in the audio data. In some examples, the speech processing engine can identify a single word (e.g., "Hello", "Jack", "and", "Steve", etc.) or any combination of one or more words as a speech segment. In some examples, based on a voice recognition model stored for user 110 (e.g., based on the attributes of the speech segment being similar to the stored voice recognition model) and / or sound intensity (e.g., volume), the speech processing engine can identify user 110 as the speaker of "Hello Jack and Steve. How’s it going".

[0046] In some examples, the speech processing engine can be configured to detect a face with moving lips in the image data to identify speech segments (e.g., the start and end of a speech segment) and / or identify the speaker. For example, the speech processing engine can detect the faces of persons 101A and 102A and detect that the mouth 103 of person 101A is moving while capturing audio associated with the speech segment "Where is Mary?". Based on this information, the speech processing engine can determine person 101A as the speaker of this speech segment. In another example, the speech processing engine can determine person 101A as the speaker because user 110 is paying attention to him while person 101A is speaking (e.g., while the lips of person 101A are moving and audio data is being captured). In some examples, the speech processing engine also obtains other information, such as, for example, location information (e.g., GPS coordinates) or calendar information, to identify the speaker or identify a potential speaker model. For example, the speech processing engine can use calendar meeting information to identify persons 101A and 102A.

[0047] The speech processing engine can transcribe each speech segment in the speech segments to produce a transcription that includes an indication of the speaker associated with each speech segment. The speech processing engine can also analyze the transcription to produce additional data derived from the transcription. For example, in Figure 1AIn the example shown, the speech processing engine can transcribe the speech segment "Where is Mary?", analyze the calendar information, and determine that Mary declined a meeting invitation. The speech processing engine can then generate alert 105 and display the alert to user 110 in the artificial reality content 122. In this way, the speech processing engine can assist user 110 in responding to person 101A.

[0048] The speech processing engine can generate other additional data, such as a calendar invitation for the meeting or event described in the transcription, information related to the topic identified in the transcription, or a task list including the tasks identified in the transcription. In some examples, the speech processing engine can generate a notification. For example, the processing engine can generate a notification indicating that person 101A is asking about Mary and transmit the notification to Mary. In some examples, the speech processing engine can generate statistical data about the transcription, including the number of words spoken by the speaker, the speaker's tone, the speaker's volume, information about the filler words used by the speaker, the percentage of time each speaker spoke, information about the swear words used, information about the length of the words used, a summary of the transcription, or the speaker's mood. The speech processing engine can also generate a modified version of the speech segment associated with at least one of the multiple speakers. For example, the speech processing engine can generate an audio or video file in which the voice of one or more speakers is replaced by another voice (e.g., the voice of a cartoon character or a celebrity), or replace one or more speech segments in the audio or video file.

[0049] In some examples, the speech processing engine can be included in the transcription system 106. For example, the HMD 112 can capture audio and image data and transmit the audio and image data to the transcription system 106 via the network 104. The transcription system 106 can identify the speech segments in the audio data, identify the speaker associated with each speech segment in the speech segments, transcribe each speech segment in the speech segments to produce a transcription including an indication of the speaker associated with each speech segment, and analyze the transcription to produce additional data derived from the transcription.

[0050] One or more of the techniques described herein can have various technical advantages and practical applications. For example, a speech transcription system according to one or more aspects of the present disclosure can generate additional data from a transcription. By automatically generating additional data, the systems according to the techniques of the present disclosure can provide services to users without the user having to say a "wake-up" word or even enter a command or instruction. This can facilitate the interaction between the user and the system, making the interaction more consistent with the way the user might interact with another user, and thus making the interaction with the system more natural.

[0051] Figure 1BFIG. is an illustration of an example system that performs speech transcription according to the techniques of the present disclosure. In this example, user 110 is wearing 112A, person 101A is wearing HMD 112B, and person 102A is wearing 112C. In some examples, user 110, 101A, and / or 103A may be in the same physical environment or in different physical environments. In Figure 1B the example shown, HMD 112A may display persons 101B and 102B in artificial reality content 123 to user 110. In this example, the artificial reality content 123 includes a virtual meeting application that includes video streams and audio streams from each of HMD 112B and HMD 112C. In some examples, persons 101B and / or 102B may be unchanged images of persons 101A and 102A, respectively. In some other examples, persons 101B and / or 102B may be avatars (or any other virtual representation) corresponding to persons 101B and / or 102B.

[0052] In Figure 1B the example shown, HMDs 112A, 112B, and 112C (collectively referred to as "HMD 112") communicate wirelessly with each other (e.g., directly or via network 104). Each of the HMDs 112 in HMD 112 may include a speech processing engine (not shown). In some examples, each of the HMDs 112 in HMD 112 may operate in substantially the same manner as Figure 1A the other HMDs 112. In some examples, HMD 112A may store a first speech recognition model corresponding to user 110, HMD 112B may store a second speech recognition model corresponding to user 101A, and HMD 112C may store a third speech recognition model corresponding to user 102A. In some examples, each of the HMDs 112 in HMD 112 may share and store copies of the first, second, and third speech recognition models.

[0053] In some examples, each HMD 112 in the HMDs 112 obtains audio data and / or image data. For example, each HMD 112 in the HMDs 112 can capture audio data and image data from its physical environment and / or obtain audio data and / or image data from other HMDs 112. In some examples, each HMD 112 can transcribe voice segments corresponding to the user wearing the HMD. For example, HMD 112A can transcribe only one or more voice segments corresponding to user 110, HMD 112B can transcribe only one or more voice segments corresponding to user 101A, and HMD 112C can transcribe only one or more voice segments corresponding to user 102A. For example, in such an example, HMD 112A will capture audio data and / or image data from its physical environment, identify voice segments in the audio data, identify the voice segments corresponding to user 110 (e.g., based on the voice recognition model stored for user 110), and transcribe each voice segment corresponding to user 110. Each HMD 112 in the HMDs 112 will transmit its respective transcription to the transcription system 106. The system 106 will combine the individual transcriptions to produce a complete transcription and analyze the complete transcription to produce additional data derived from the complete transcription. In this way, each HMD 112 does not need to store a voice recognition model for other users. Additionally, each HMD 112 that transcribes the voice of the corresponding user can improve transcription and / or speaker identification accuracy.

[0054] In some other examples, each HMD 112 in the HMDs 112 can capture audio and image data and transmit the audio and image data to the transcription system 106 via the network 104 (e.g., in audio and video streams). The transcription system 106 can identify voice segments in the audio data, identify the speaker associated with each voice segment in the voice segments, transcribe each voice segment in the voice segments to produce a transcription that includes an indication of the speaker associated with each voice segment, and analyze the transcription to produce additional data derived from the transcription.

[0055] Figure 1C is an illustration of an example system 10B that performs voice transcription in accordance with the techniques of the present disclosure. In this example, users 110, 101, and 102 are in the same physical environment and the computing device 120 captures audio and / or image data. In some other examples, one or more other users located in different physical environments can be part of the interaction with users 110, 101, and 102 facilitated by the computing device 120. Figure 1CThe computing device 120 therein is shown as a single computing device, which may correspond to a mobile phone, a tablet computer, a smart watch, a game console, a workstation, a desktop computer, a laptop computer, an assistive device, a dedicated desktop device, or other computing devices. In some other examples, the computing device 120 may be distributed across multiple computing devices.

[0056] In some examples, the computing device 120 may perform transcription operations similar to those described above with reference to Figure 1A and Figure 1B the HMD 112 therein. For example, a speech processing engine (not shown) of the computing device 120 may identify speech segments in audio data, identify the speaker associated with each speech segment in the speech segments, transcribe each speech segment in the speech segments to produce a transcription including an indication of the speaker associated with each speech segment, and analyze the transcription to produce additional data derived from the transcription. In another example, the computing device 120 captures audio and / or image data, transmits the audio and / or image data to a transcription system, and then a speech processing engine of the transcription system 106 identifies speech segments in the audio data, identifies the speaker associated with each speech segment in the speech segments, transcribes each speech segment in the speech segments to produce a transcription including an indication of the speaker associated with each speech segment, and analyzes the transcription to produce additional data derived from the transcription.

[0057] In examples where the computing device 120 facilitates interactions involving remote users and / or users in different physical environments, the computing device 120 may use any indication of audio information and image or video information from a device corresponding to a remote user (e.g., an audio and / or video stream) to identify speech segments in the (multiple) audio streams, identify the speaker (e.g., the remote user) associated with each speech segment in the (multiple) audio streams, transcribe each speech segment in the speech segments to produce a transcription including an indication of the speaker (including the remote speaker) associated with each speech segment, and analyze the transcription to produce additional data derived from the transcription.

[0058] Figure 2A is a diagram depicting an example HMD 112 configured to operate in accordance with one or more techniques of the present disclosure. Figure 2A The HMD 112 of Figure 1A may be an example of the HMD 112 of Figure 1B or the HMDs 112A, 112B, and 112C of Figure 1A , Figure 1B the system 10A. The HMD 112 may operate as a stand-alone mobile artificial reality system configured to implement the techniques described herein, or may be part of a system, such as

[0059] In this example, the HMD 112 includes a front rigid body and a strap for securing the HMD 112 to the user. Additionally, the HMD 112 includes an inward-facing electronic display 203 configured to present artificial reality content to the user. The electronic display 203 can be any suitable display technology, such as a liquid crystal display (LCD), quantum dot display, dot matrix display, light-emitting diode (LED) display, organic light-emitting diode (OLED) display, cathode ray tube (CRT) display, electronic ink, or monochrome, color, or any other type of display capable of generating a visual output. In some examples, the electronic display is a stereoscopic display for providing separate images to each eye of the user. In some examples, when tracking the position and orientation of the HMD 112 to render artificial reality content based on the current viewing perspective of the HMD 112 and the user, the known orientation and position of the display 203 relative to the front rigid body of the HMD 112 are used as a reference frame, also referred to as a local origin. The reference frame can also be used to track the position and orientation of the HMD 112. In some other examples, the HMD 112 can take the form of other wearable head-mounted displays, such as glasses or goggles.

[0060] As Figure 2AAs further shown, in this example, HMD 112 also includes one or more motion sensors 206, such as one or more accelerometers (also known as inertial measurement units or "IMUs") that output data indicating the current acceleration of HMD 112, a GPS sensor that outputs data indicating the position of HMD 112, a radar or sonar that outputs data indicating the distance between HMD 112 and various objects, or other sensors that provide an indication of the position or orientation of HMD 112 or other objects within the physical environment. Additionally, HMD 112 may include integrated image capture devices 208A and 208B (collectively referred to as "image capture system 208", which may include any number of image capture devices) (e.g., cameras, still cameras, IR scanners, UV scanners, laser scanners, Doppler radar scanners, depth scanners) and an audio capture system 209 (e.g., a microphone), which are respectively configured to capture raw image and audio data. In some aspects, image capture system 208 may capture image data from both the visible and invisible spectra of the electromagnetic spectrum (e.g., IR light). Image capture system 208 may include one or more image capture devices that capture image data from the visible spectrum and one or more separate image capture devices that capture image data from the invisible spectrum, or these may be combined in the same one or more image capture devices. More specifically, image capture system 208 captures image data representing objects in the physical environment within the field of view 130 of image capture system 208, which generally corresponds to the viewing perspective of HMD 112, and audio capture system 209 captures audio data in the vicinity of HMD 112 (e.g., within a 360-degree range of the audio capture device). In some examples, audio capture system 209 may include a microphone array that can capture information about the directionality of the audio source relative to HMD 112. HMD 112 includes an internal control unit 210, which may include an internal power supply and one or more printed circuit boards having one or more processors, memory, and hardware to provide an operating environment for performing programmable operations to process sensed data and present artificial reality content on display 203.

[0061] In one example, according to the techniques described herein, the control unit 210 is configured to identify speech segments in the audio data captured by the audio capture system 209, identify the speaker associated with each speech segment, transcribe each of the speech segments in the speech segments to produce a transcription of the plurality of speech segments, the transcription including an indication of the speaker associated with each speech segment, and analyze the transcription to produce additional data derived from the transcription. In some examples, the control unit 210 causes the audio data and / or image data to be transmitted over the network 104 to the transcription system 106 (e.g., in a manner that is near real-time or seemingly near real-time when the audio data and / or image data are captured, or after the interaction is completed).

[0062] Figure 2B is a diagram depicting an example HMD 112 according to the techniques of the present disclosure. As Figure 2B shown, the HMD 112 can take the form of glasses. Figure 2A The HMD 112 can be Figure 1A and Figure 1B Any example of the HMD 112 among the HMD 112s. The HMD 112 can be part of a system, such as Figure 1A - Figure 1B the system 10A, or can operate as a stand-alone mobile system configured to implement the techniques described herein.

[0063] In this example, the HMD 112 is glasses including a front frame that includes a bridge and temple arms (or "arms") that allow the HMD 112 to rest on the user's nose and extend over the user's ears to secure the HMD 112 to the user. Additionally, Figure 2B the HMD 112 includes inward-facing electronic displays 203A and 203B (collectively referred to as "electronic display 203") configured to present artificial reality content to the user. The electronic display 203 can be any suitable display technology, such as a liquid crystal display (LCD), a quantum dot display, a dot matrix display, a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a cathode ray tube (CRT) display, electronic ink, or a monochrome, color, or any other type of display capable of generating a visual output. In Figure 2B the example shown, the electronic display 203 forms a stereoscopic display for providing separate images to each eye of the user. In some examples, when tracking the position and orientation of the HMD 112 to render artificial reality content according to the current viewing perspective of the HMD 112 and the user, the known orientation and position of the display 203 relative to the front frame of the HMD 112 are used as a reference frame, also referred to as a local origin.

[0064] As Figure 2BAs further shown herein, in this example, the HMD 112 also includes one or more motion sensors 206, such as one or more accelerometers (also referred to as inertial measurement units or "IMUs") that output data indicative of the current acceleration of the HMD 112, a GPS sensor that outputs data indicative of the position of the HMD 112, a radar or sonar that outputs data indicative of the distance of the HMD 112 from various objects, or other sensors that provide an indication of the position or orientation of the HMD 112 or other objects within the physical environment. Additionally, the HMD 112 can include integrated image capture devices 208A and 208B (collectively referred to as "image capture system 208") (e.g., cameras, still cameras, IR scanners, UV scanners, laser scanners, Doppler radar scanners, depth scanners) and an audio capture system 209 (e.g., microphones), which are respectively configured to capture image and audio data. In some aspects, the image capture system 208 can capture image data from both the visible and non-visible spectra of the electromagnetic spectrum (e.g., IR light). The image capture system 208 can include one or more image capture devices that capture image data from the visible spectrum and one or more separate image capture devices that capture image data from the non-visible spectrum, or these can be combined in the same one or more image capture devices. More specifically, the image capture system 208 captures image data representing objects in the physical environment within the field of view 130 of the image capture system 208, which generally corresponds to the viewing perspective of the HMD 112, and the audio capture system 209 captures audio data in the vicinity of the HMD 112 (e.g., within a 360-degree range of the audio capture device). The HMD 112 includes an internal control unit 210, which can include an internal power supply and one or more printed circuit boards having one or more processors, memory, and hardware to provide an operating environment for performing programmable operations to process sensed data and present artificial reality content on the display 203. According to the techniques described herein, Figure 2B the control unit 210 is configured to operate similarly to Figure 2A the control unit 210.

[0065] Figure 3 is a block diagram depicting an example of speech transcription being performed by an example instance of the HMD 112 of an artificial reality system in which the techniques of the present disclosure are implemented by Figure 1A and Figure 1B In the example of Figure 3 the HMD 112 performs image and audio data capture, speaker identification, transcription, and analysis operations according to the techniques described herein.

[0066] In this example, the HMD 112 includes one or more processors 302 and a memory 304 which, in some examples, provide a computer platform for executing an operating system 305, which may be, for example, an embedded real-time multitasking operating system or other type of operating system. Further, the operating system 305 provides a multitasking operating environment for executing one or more software components 317. The processor 302 is coupled to one or more I / O interfaces 315 which provide I / O interfaces for communicating with other devices such as a display device, an image capture device, other HMDs, etc. In addition, one or more of the I / O interfaces 315 may include one or more wired or wireless network interface controllers (NICs) for communicating with a network such as network 104. Further, the (multiple) processors 302 are coupled to an electronic display 203, a motion sensor 206, an image capture system 208, and an audio capture system 209. In some examples, the processor 302 and the memory 304 may be separate, discrete components. In some other examples, the memory 304 may be on-chip memory collocated with the processor 302 within a single integrated circuit. The image capture system 208 and the audio capture system 209 are respectively configured to acquire image data and audio data.

[0067] Generally, the application engine 320 includes functionality for providing and presenting artificial reality applications such as transcription applications, voice assistant applications, virtual meeting applications, gaming applications, navigation applications, educational applications, training or simulation applications, etc. The application engine 320 may include, for example, one or more software packages, software libraries, hardware drivers, and / or application programming interfaces (APIs) for implementing artificial reality applications on the HMD 112. In response to control by the application engine 320, the rendering engine 322 generates 3D artificial reality content for display to a user by the application engine 340 of the HMD 112.

[0068] The application engine 340 and the rendering engine 322 construct artificial content for display to the user 110 based on the current pose information of the HMD 112 within a reference frame, which is typically the viewing perspective of the HMD 112 determined by the pose tracker 326. Based on the current viewing perspective, the rendering engine 322 constructs 3D artificial reality content, which in some cases can be at least partially overlaid on the user 110's real-world 3D environment. During this process, the pose tracker 326 operates on the sensed data and user commands received from the HMD 112 to capture 3D information within the real-world environment, such as the movement of the user 110, and / or feature tracking information regarding the user 110. In some examples, the application engine 340 and the rendering engine 322 can generate and render one or more user interfaces for a transcription application or a voice assistant application for display according to the techniques of the present disclosure. For example, the application engine 340 and the rendering engine 322 can generate and render a user interface for display to display transcription and / or additional data.

[0069] The software application 317 of the HMD 112 operates to provide an overall artificial reality application, including a transcription application. In this example, the software application 317 includes a rendering engine 322, an application engine 340, a pose tracker 326, a speech processing engine 341, image data 330, audio data 332, a speaker model 334, and a transcription 336. In some examples, the HMD 112 can store other data (e.g., in the memory 304), including location information, calendar event data for the user (e.g., invitees, confirmed persons, meeting topics), etc. In some examples, the image data 330, audio data 332, speaker model 334, and / or transcription 336 can represent a repository or cache.

[0070] The speech processing engine 341 performs functions related to transcribing speech in the audio data 332 and analyzing the transcription according to the techniques of the present disclosure. In some examples, the speech processing engine 341 includes a speech recognition engine 342, a speaker identifier 344, a speech transcriber 346, and a voice assistant application 348.

[0071] The speech recognition engine 342 performs functions related to recognizing one or more speech segments in the audio data 332. In some examples, the speech recognition engine 342 stores one or more speech segments in the audio data 332 (e.g., separate from the original analog data). The speech segments can include one or more spoken words. For example, a speech segment can be a single word, two or more words, or even a phrase or a complete sentence. In some examples, the speech recognition engine 342 uses any speech recognition technology to recognize one or more speech segments in the audio data 332. For example, the audio data 332 can include analog data and the speech recognition engine 342 can use an analog-to-digital converter (ADC) to convert the analog data to digital data, filter the noise in the digitized audio data, and apply one or more statistical models (e.g., hidden Markov models or neural networks) to the filtered digitized audio data to recognize one or more speech segments. In some examples, the speech recognition engine 342 can apply an artificial intelligence (AI) / machine learning (ML) model that is trained to recognize speech for one or more specific users (e.g., Figure 1A - Figure 1C user 110). In some examples, the AI / ML model can receive training feedback from the user to adjust the speech recognition determination. In some examples, the speech recognition engine 342 can recognize one or more speech segments in the audio data 332 based on the image data 330. For example, the speech recognition engine 342 can be configured to detect a face with moving lips in the image data to identify speech segments (e.g., the start and end of a speech segment).

[0072] The speaker identifier 344 performs functions related to identifying the speaker associated with each of one or more speech segments recognized by the speech recognition engine 342. For example, the speaker identifier 344 may be configured to detect a face with moving lips in the image data 330 to identify the speaker or potential speaker. In another example, the audio capture system 209 may include a microphone array that can capture information about the directionality of the audio source relative to the HMD 112, and the speaker identifier 344 may identify the speaker or potential speaker based on this directionality information and the image data 330 (e.g., the speaker identifier 344 may identify the person 101A in FIG. 1 based on the directionality information for the speech segment "Where is Mary?"). In yet another example, the speaker identifier 344 will identify the speaker based on who the user is looking at (e.g., based on the field of view of the HMD 112). In some examples, the speaker identifier 344 may determine a hash value or embedding value for each speech segment, obtain potential speaker models (e.g., from the speaker model 334), compare the hash value with the potential speaker models, and identify the speaker model that is closest to the hash value. The speaker identifier 344 may identify potential speaker models based on external data, the image data 330 (e.g., based on detected faces with moving lips), and / or user input. For example, the speaker identifier 344 may identify potential speaker models based on calendar information (e.g., information about confirmed or potential meeting invitees), one or more faces identified in the image data 330, location information (e.g., proximity information of people or devices associated with other people relative to the HMD 112), and / or based on potential speaker models selected via user input. In some examples, if the difference between the hash value for the speech segment and the closest speaker model is equal to or greater than a threshold difference, the speaker identifier 344 may create a new speaker model based on the hash value and associate the new speaker model with the speech segment. If the difference between the hash value of the speech segment and the closest speaker model is less than the threshold difference, the speaker identifier 344 may identify the speaker associated with the closest speaker model as the speaker of the speech segment. In some examples, the speaker model 334 may include hash values (or other voice attributes) for different speakers. In some examples, the speaker model 344 may include an AI / ML model that is trained to identify one or more speakers (e.g., Figure 1A - Figure 1Cvoices of people 110, 101, 102). In some examples, the AI / ML model can receive training feedback from the user to adjust speaker identification determination. The speaker model 334 can also include a speaker identifier (ID), name, or label that is automatically generated by the speaker identifier 344 (e.g., "Speaker 1", "Speaker 2", etc.) or manually entered by the user (e.g., the speaker) via the I / O interface 315 ("Jack", "Steve", "Boss", etc.). In some examples, each speaker model 344 can include one or more images of the speaker and / or a hash value for the speaker's face.

[0073] In some examples, the speaker identifier 344 can be configured to identify voice segments attributable to the user of the HMD 112. For example, the speaker identifier 344 can apply a speaker model specific to the user of the HMD 112 (e.g., user 110) to identify one or more voice segments associated with the user (e.g., identify the voice segments spoken by user 110 based on the attributes of the voice segments being similar to the user speaker model). In other words, the speaker identifier 344 can filter one or more speaker segments recognized by the speech recognition engine 342 for the (multiple) voice segments spoken by the user of the HMD 112.

[0074] The speech transcriber 346 performs functions related to transcribing the voice segments recognized by the speech recognition engine 342. For example, the speech transcriber 346 produces a text output of one or more voice segments recognized by the speech recognition engine 342, the text output having an indication of one or more speakers identified by the speaker identifier 344. In some examples, the speech transcriber 346 produces a text output of one or more voice segments recognized by the speech recognition engine 342 that are associated with the user of the HMD 112 (e.g., user 110). In other words, in some examples, the speech transcriber 346 only produces a text output for one or more voice segments spoken by the user of the HMD 112 as identified by the speaker identifier 344. Either way, the speech transcriber 346 then stores the text output in the transcription 336.

[0075] The voice assistant application 348 performs functions related to analyzing the transcription to produce additional data derived from the transcription. For example, the voice assistant application 348 can produce additional data such as a calendar invitation for a meeting or event described in the transcription (e.g., corresponding to the voice segment "Let's touch base again first thing Friday morning", information related to the topic identified in the transcription (e.g., a notification that a meeting invitee has declined a meeting invitation, such as Figure 1AAs shown, a notification for a person not present in the interaction), or a task list including the tasks identified in the transcription (e.g., a task item corresponding to the voice segment "Please send the sales report for last month after the meeting ends"). In some examples, the voice assistant application 348 can generate statistics about the transcription including the number of words spoken by the speaker, the tone of the speaker, information about the filler words used by the speaker (e.g., um, huh, oh, like, etc.), the percentage of time each speaker speaks, information about the swear words used, information about the length of the words used, a summary of the transcription, or the mood of the speaker. The voice assistant application 348 can also generate a modified version of the voice segment associated with at least one of the multiple speakers. For example, the voice assistant application 348 can generate an audio or video file in which the voice of one or more speakers is replaced by another voice (e.g., a cartoon voice or a celebrity voice), or replace the language of one or more voice segments in the audio or video file.

[0076] As described above, the speaker model 334 can include various AI / ML models. These AI / ML models can include artificial neural networks (ANNs), decision trees, support vector networks, Bayesian networks, genetic algorithms, linear regression, logistic regression, linear discriminant analysis, naive Bayes, k-nearest neighbors, learning vector quantization, support vector machines, random decision forests, or any other known AI / ML mathematical model. These AI / ML models can be trained to process audio data and identify voice segments and / or identify the speakers of the voice segments. For example, these AI / ML models can be trained to identify speech and / or specific voices in the audio data 332. In some examples, these AI / ML models can be trained to identify potential speakers in image data. For example, these AI / ML models can be trained to identify people (e.g., faces) and / or moving lips in the image data 330. In some examples, the speaker model 334 can be trained using a voice dataset for one or more users and / or an image set corresponding to one or more users. In one or more aspects, the information stored in each of the image data 330, audio data 332, speaker model 334, and / or transcription 336 can be stored in a repository, database, map, search tree, or any other data structure. In some examples, the image data 330, audio data 332, speaker model 334, and / or transcription 336 can be separate from the HMD 112 (e.g., can be (multiple) separate databases that communicate with the HMD 112 via Figure 1A the network 104).

[0077] The motion sensor 206 may include sensors such as one or more accelerometers (also known as an inertial measurement unit or “IMU”) that output data indicating the current acceleration of the HMD 112, radar or sonar 112 that outputs data indicating the distance of the HMD to various objects, or other sensors that provide an indication of the position or orientation of the HMD 112 or other objects within the physical environment.

[0078] Figure 4 is a diagram showing a technique according to the present disclosure wherein Figure 1A , Figure 1B A block diagram of an example implementation of an HMD and an example instance of a transcription system for an artificial reality system performing speech transcription. Figure 4 In the example of , the HMD 112 captures audio and / or image data and transmits the audio and / or image data to the transcription system 106. According to one or more techniques described herein, the speech recognition engine 441 of the transcription system 106 recognizes speech segments in the audio data, identifies a speaker associated with each speech in the speech segments, transcribes each of the speech segments to produce a transcription that includes an indication of the speaker associated with each speech segment, and analyzes the transcription to produce additional data derived from the transcription.

[0079] In this example, and with something like Figure 3 In some examples, HMD 112 includes one or more processors 302 and memory 304, which in some examples provide a computer platform for executing an operating system 305, which may be, for example, an embedded real-time multitasking operating system, or other types of operating systems. In turn, operating system 305 provides a multitasking operating environment for executing one or more software components 317. In addition, processor(s) 302 are coupled to electronic display 203, motion sensor 206, image capture system 208, and audio capture system 209. In some examples, HMD 112 also includes Figure 3 For example, HMD 112 may include a speech processing engine 341 (which includes a speech recognition engine 342, a speaker identifier 344, a speech transcriber 346, and a voice assistant application 348), image data 330, audio data 332, a speaker model 334, and transcription 336.

[0080] Generally, the transcription system 106 is a device that processes audio and / or image data received from the HMD 112 to produce a transcription that includes an indication of one or more speakers in a speech segment contained in the audio data, and additional data derived from the transcription. In some examples, the transcription system 106 is a single computing device, such as a server, workstation, desktop computer, laptop computer, or gaming system. In some other examples, at least a portion of the transcription system 106 (such as the processor 412 and / or the memory 414) may be distributed across a cloud computing system, a data center, or across a network, such as the Internet, another public or private communication network, such as broadband, cellular, Wi-Fi, and / or other types of communication networks used to transfer data between computing systems, servers, and computing devices.

[0081] In Figure 4 the example of, the transcription system 106 includes one or more processors 412 and a memory 414, which, in some examples, provide a computer platform for executing an operating system 416, which can be, for example, an embedded real-time multitasking operating system or other types of operating systems. Further, the operating system 416 provides a multitasking operating environment for executing one or more software components 417. The processor 412 is coupled to one or more I / O interfaces 415, which provide I / O interfaces for communicating with other devices such as a keyboard, mouse, game controller, display device, image capture device, HMD, etc. In addition, one or more of the I / O interfaces 415 may include one or more wired or wireless network interface controllers (NICs) for communicating with a network such as the network 104. Each of the processors 302, 412 may include any one or more of the following: a multi-core processor, a controller, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or equivalent discrete or integrated logic circuits. The memories 304, 414 may include any form of memory for storing data and executable software instructions, such as random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory.

[0082] The software application 417 of the transcription system 106 operates to provide a transcription application. In this example, the software application 417 includes a rendering engine 422, an application engine 440, a pose tracker 426, a speech processing engine 441, image data 430, audio data 432, a speaker model 434, and a transcription 436. Similar to Figure 3The voice processing engine 341, and the voice processing engine 441 includes a speech recognition engine 442, a speaker identifier 444, a speech transcriber 446, and a voice assistant application 448.

[0083] Generally, the application engine 440 includes providing and presenting the functionality of an artificial reality application, such as a transcription application, a voice assistant application, a virtual meeting application, a gaming application, a navigation application, an educational application, a training or simulation application, etc. The application engine 440 may include, for example, one or more software packages, software libraries, hardware drivers, and / or application programming interfaces (APIs) for implementing an artificial reality application on the computing system 120. In response to the control of the application engine 440, the rendering engine 422 generates 3D artificial reality content for display to the user by the application engine 340 of the HMD 112.

[0084] The application engine 440 and the rendering engine 422 perform functions related to constructing artificial content for display to the user 110 based on the current pose information of the HMD 112 within a reference frame, which is typically the viewing perspective of the HMD 112 determined by the pose determination tracker 426. Based on the current viewing perspective, the rendering engine 422 constructs 3D artificial reality content, which in some cases may be at least partially overlaid on the real-world 3D environment of the user 110. During this process, the pose tracker 426 operates on the sensed data received from the HMD 112 (such as image data 430 from sensors on the HMD 112), and in some examples on data from external sensors (such as an external camera), to capture 3D information within the real-world environment, such as the movement made by the user 110, and / or feature tracking information related to the user 110. Based on the sensed data, the computing system 120 constructs artificial reality content via one or more I / O interfaces 315, 415 for transmission to the HMD 112 for display to the user 110. In some examples, the application engine 440 and the rendering engine 422 may generate and render one or more user interfaces for a multimedia query application for display according to the techniques of the present disclosure. For example, the application engine 440 and the rendering engine 422 may generate and present a user interface for display for displaying transcription and / or additional data.

[0085] The speech recognition engine 442 performs functions related to recognizing one or more speech segments in the audio data 432 received from the HMD 112 (e.g., as described above with reference to Figure 3as described with respect to the speech recognition engine 342). In some examples, the speech recognition engine 442 stores one or more speech segments in the audio data 432 (e.g., separate from the original analog data). A speech segment can include one or more spoken words. For example, a speech segment can be a single word, two or more words, or even a phrase or a complete sentence.

[0086] The speaker identifier 444 performs functions related to identifying the speaker associated with each of the one or more speech segments recognized by the speech recognition engine 442. For example, the speaker identifier 444 can be configured to detect a face with moving lips in the image data 430 to identify the speaker or potential speaker. In another example, the audio capture system 209 of the HMD 112 can include a microphone array that can capture information about the directionality of the audio source relative to the HMD 112, and the speaker identifier 444 can identify the speaker or potential speaker based on this directionality information and the image data 430 (e.g., the speaker identifier 444 can identify the person 101A in FIG. 1 based on the directionality information for the speech segment "Where is Mary?"). In yet another example, the speaker identifier 444 will identify the speaker based on who the user is looking at (e.g., based on the field of view of the HMD 112).

[0087] In some examples, the speaker identifier 444 can determine a hash value or embedding value for each speech segment, obtain potential speaker models (e.g., from the speaker model 434), compare the hash value with the potential speaker models, and identify the speaker model that is closest to the hash value. The speaker identifier 444 can identify potential speaker models based on external data, the image data 430 received from the HMD 112 (e.g., based on a detected face with moving lips), and / or user input. For example, the speaker identifier 344 can identify potential speaker models based on calendar information (e.g., information about confirmed or potential meeting invitees), one or more faces identified in the image data 430 received from the HMD 112, location information (e.g., proximity information of people or devices associated with other people relative to the HMD 112), and / or based on potential speaker models selected via user input. In some examples, if the difference between the hash value of the speech segment and the closest speaker model is equal to or greater than a threshold difference, the speaker identifier 444 can create a new speaker model based on the hash value and associate the new speaker model with the speech segment. If the difference between the hash value of the speech segment and the closest speaker model is less than the threshold difference, the speaker identifier 444 will identify the speaker associated with the closest speaker model as the speaker of the speech segment. In some examples, the speaker model 434 can include hash values for different speakers.

[0088] In some examples, the speaker identifier 444 can be configured to identify speech segments attributable to a user of the HMD 112. For example, the speaker identifier 444 can apply a speaker model specific to a user of the HMD 112 (e.g., user 110) to identify one or more speech segments associated with the user (e.g., identify speech segments spoken by user 110 based on the attributes of the speech segments being similar to the user speaker model).

[0089] Similar to the speech transcriber 346 described above with respect to Figure 3 the speech transcriber 446 performs functions related to transcribing one or more speech segments recognized by the speech recognition engine 442. For example, the speech transcriber 446 generates a text output of one or more speech segments recognized by the speech recognition engine 442 and stores the text output in the transcription 436, the text output having an indication of one or more speakers identified by the speaker identifier 444. In some examples, the speech transcriber 346 generates only a text output of one or more speech segments spoken by a user of the HMD 112 as identified by the speaker identifier 444. In some examples, the speech processing engine 441 transmits the text output to the HMD 112.

[0090] The voice assistant application 448 performs functions related to analyzing the transcription to produce additional data derived from the transcription. For example, the voice assistant application 448 can produce additional data such as a calendar invitation for a meeting or event described in the transcription (e.g., corresponding to the speech segment "Let's touch base again first thing Friday morning"), information related to a topic identified in the transcription (e.g., a notification that a meeting invitee has declined a meeting invitation, as Figure 1A shown, a notification of a person who did not appear in the interaction), or a task list including tasks identified in the transcription (e.g., a task item corresponding to the speech segment "Please send out the sales report for last month after the meeting"). In some examples, the voice assistant application 448 can produce statistical data about the transcription including the number of words spoken by the speaker, the speaker's tone, information about filler words used by the speaker (e.g., um, uh, oh, like, etc.), the percentage of time each speaker spoke, information about swear words used, information about the length of the words used, a summary of the transcription, or the speaker's mood. The voice assistant application 448 can also produce a modified version of the speech segment associated with at least one of the multiple speakers. For example, the voice assistant application 348 can generate an audio or video file in which the voice of one or more speakers is replaced by another voice (e.g., the voice of a cartoon or a celebrity), or replace the language of one or more speech segments in the audio or video file. In some examples, the speech processing engine 441 transmits the additional data to the HMD 112.

[0091] Similar to the above Figure 3 Speaker model 334 described, speaker model 434 may include various AI / ML models. These AI / ML models may be trained to process audio data and recognize speech segments and / or identify speakers of speech segments. For example, these AI / ML models may be trained to recognize speech and / or specific voices in audio data 432. In some examples, these AI / ML models may be trained to identify potential speakers in image data. For example, these AI / ML models may be trained to recognize people (e.g., faces) and / or moving lips in image data 430. In some examples, speaker model 334 may be trained using a speech dataset of one or more users and / or an image set corresponding to one or more users. In some examples, the AI / ML model may receive training feedback from a user (e.g., via I / O interface 415) to adjust speaker identification determination. Speaker models 334 may also include speaker identifiers, names, or labels, which are automatically generated by speaker identifier 344 (e.g., "Speaker 1," "Speaker 2," etc.) or manually entered by a user via I / O interface 415 (e.g., "Jack," "Steve," "Boss," etc.). In some examples, speaker models 344 may each include one or more images of the speaker and / or a hash value for the speaker's face.

[0092] In some examples, the transcription system 106 receives data from two or more HMDs (e.g., Figure 1B In some examples, each HMD may transmit audio and / or image data of the same physical environment or from different physical environments (e.g., Figure 1B ). By capturing audio and / or image data about the same environment from two or more different sources, a greater amount of information can be captured. For example, image data can be captured from two or more different perspectives, or audio data can be captured from two different points of the environment, which can enable different sounds to be captured. In some examples, the transcription system 106 generates a single transcription based on the data received from all HMDs.

[0093] Figure 5 is a diagram showing that the technology according to the present disclosure is Figure 1C A block diagram of an example implementation of a computing device 120 of a system of FIG. 10 performing speech transcription. Figure 5 In the example of FIG. 1 , the computing device 120 executes the above reference Figure 3 The HMD 112 describes image and audio data capture, speaker identification, transcription, and analysis operations.

[0094] In this example, computing device 120 includes one or more processors 502 and a memory 504 which, in some examples, provide a computing platform for executing an operating system 505, such as an embedded real-time multitasking operating system or other type of operating system. Further, the operating system 505 provides a multitasking operating environment for executing one or more software components 517. The processor 502 is coupled to one or more I / O interfaces 515 which provide I / O interfaces for communicating with other devices such as a keyboard, a mouse, a game controller, a display device, an image capture device, other HMDs, etc. In addition, one or more of the I / O interfaces 515 may include one or more wired or wireless network interface controllers (NICs) for communicating with a network such as network 104. Further, the (one or more) processors 502 are coupled to an electronic display 503, an image capture system 508, and an audio capture system 509. The image capture system 208 and the audio capture system 209 are configured to acquire image data and audio data, respectively.

[0095] Figure 5 The computing device 120 in [the example] is shown as a single computing device which may correspond to a mobile phone, a tablet computer, a smartwatch, a game console, a workstation, a desktop computer, a laptop computer, or other computing device. In some other examples, the computing device 120 may be distributed across multiple computing devices such as a distributed computing network, a data center, or a cloud computing system.

[0096] The software application 517 of the computing system operates to provide a transcription application. Similar to the software applications 317 and 417 respectively similar to Figure 3 and Figure 4 the software application 517 includes a rendering engine 522, an application engine 540, a speech processing engine 541, image data 530, audio data 532, a speaker model 534, and a transcription 536. Similar to the speech processing engines 341 and 441 respectively similar to Figure 3 and Figure 4 the speech processing engine 541 includes a speech recognition engine 542, a speaker identifier 544, a speech transcriber 546, and a voice assistant application 548.

[0097] Similar to the way the HMD 112 processes audio and / or image data (e.g., as described above with respect to Figure 3As described above, the computing system 120 captures audio and / or image data and transmits the audio and / or image data to the transcription system 106. The speech recognition engine 441 of the transcription system 106 identifies speech segments in the audio data, identifies the speaker associated with each speech segment in the speech segments, transcribes each speech segment to produce a transcription including an indication of the speaker associated with each speech segment, and analyzes the transcription to produce additional data derived from the transcription.

[0098] In some examples, Figure 5 the computing device 120 simply captures the image data 530 and the audio data 532 and transmits the data to the transcription system 106. The transcription system 106 processes the audio and / or image data received from the computing device 120 in the same manner as it processes the audio and / or image data received from the HMD 112 to produce a transcription including an indication of one or more speakers in the speech segments contained in the audio data, and produces additional data from the additional data derived from the transcription (e.g., as described above with respect to Figure 4 ).

[0099] In some examples, the transcription system 106 receives audio and / or image data from Figure 4 the HMD 112 and Figure 5 the computing device 120. In some examples, the HMD 112 and the computing device 120 may transmit audio and / or image data of the same physical environment or from different physical environments. By capturing audio and / or image data about the same environment from two or more different sources, a greater amount of information can be captured. For example, image data can be captured from two or more different perspectives, or audio data can be captured from two different points in the environment, which can enable different sounds to be captured. In some examples, the transcription system 106 processes the data from the HMD 112 in the same or similar manner as it processes the data from the computing device 120, and vice versa, and produces a single transcription based on the data received from the HMD 112 and the computing device 120.

[0100] Figure 6 is a flowchart 600 illustrating example operations of a method for transcribing and analyzing speech according to aspects of the present disclosure. In some examples, Figure 6 one or more of the operations shown in

[0101] The audio capture system 209 and the image capture system 208 of the HMD 112 and / or the audio capture system 509 and the image capture system 508 of the computing device 120 capture audio and image data (602). In some examples, the audio and / or image data is captured automatically or manually. For example, the audio and / or image capture systems of the HMD 112 and / or the computing system 120 may be configured to always capture audio and / or image data when powered on. In some examples, the multimedia capture system 138 of the HMD 112 and / or the multimedia system 138 of the computing system 130 may be configured to capture multimedia data in response to a user input that initiates data capture and / or in response to initiating a transcription, a virtual meeting, or a voice assistant application. In some examples, the HMD 112 and / or the computing device 120 may transmit the audio and / or image data to the transcription system 106 (e.g., in real time, near real time, or after the interaction has terminated).

[0102] The speech processing engines 341, 441, or 541 use the image data to transcribe the audio data (604). For example, the speech processing engines 341, 441, or 541 may identify speech segments in the audio data, identify the speaker associated with each speech segment in the speech segments, and transcribe each speech segment to produce a transcription that includes an indication of the speaker associated with each speech segment.

[0103] The voice assistant applications 348, 448, or 548 then analyze the transcription to produce additional data derived from the transcription (606). For example, the voice assistant applications 348, 448, or 548 may produce additional data such as a calendar invitation for a meeting or event described in the transcription (e.g., corresponding to the speech segment "Let's touch base again first thing Friday morning"), information related to the topic identified in the transcription (e.g., a notice that a meeting invitee has declined a meeting invitation, as Figure 1A shown, a notice for a person who did not appear in the interaction), or a task list that includes the tasks identified in the transcription (e.g., a task item corresponding to the speech segment "Please send out the sales report for last month after the meeting").

[0104] In some examples, the additional data can include statistics about the transcription such as the number of words spoken by the speaker, the speaker's tone, information about filler words used by the speaker (e.g., um, huh, oh, like, etc.), the percentage of time each speaker speaks, information about swear words used, information about the length of the words used, a summary of the transcription, or the speaker's mood (e.g., for each segment or the entire transcription). The voice assistant applications 348, 448, or 548 can also produce a modified version of the voice segment associated with at least one of the multiple speakers. For example, the voice assistant applications 348, 448, or 548 can generate an audio or video file in which the voice of one or more speakers is replaced by another voice (e.g., a cartoon voice or a celebrity voice), or replace the language of one or more voice segments in the audio or video file. In some examples, the voice assistant applications 348, 448, or 548 analyze the transcription in real time (e.g., when the audio and image data are captured), near real time, after the interaction has terminated, or after the HMD 112 or the computing device 120 has stopped capturing images or image data.

[0105] Figure 7 Illustrated are audio data 702 and a transcription 706 in accordance with the techniques of the present disclosure. In Figure 7 the example shown, the audio data 702 corresponds to analog data captured by the audio capture system 209 of the HMD 112 or the audio capture system 509 of the computing device 120. The speech recognition engines 342, 442, or 552 identify speech segments 704A, 704B, 704C (collectively "speech segments 704") in the audio data 702 and generate corresponding transcribed speech segments 706A, 706B, and 706C (collectively "transcription 706"). Although each of the speech segments 704 includes an entire sentence, a speech segment can include one or more words. For example, a speech segment may not always include an entire sentence and can include a single word or phrase. In some examples, the speech recognition engines 342, 442, or 552 can combine one or more words to form a speech segment that includes a complete sentence, as Figure 7 shown in

[0106] In Figure 7 the example shown, the speaker identifier 344, 444, or 544 identifies "Speaker 1" as the speaker of speech segments 706A and 706B and "Speaker 2" as the speaker of speech segment 706C (e.g., based on the above reference to Figure 3 - Figure 5(described speaker models and / or image data). In some examples, the labels or identifiers "Speaker 1" and "Speaker 2" (inserted into the resulting transcription) can be automatically generated by speaker identifiers 344, 444, or 544. In some other examples, these identifiers or labels can be manually entered by the user and can include names (e.g., "Jack", "Steve", "Boss", etc.). Either way, these labels, identifiers, or names can provide an indication of the speaker who is the source of the speech segments in the transcription.

[0107] In some examples, voice assistant applications 348, 448, or 548 can analyze the transcription 706 to generate additional data. For example, voice assistant applications 348, 448, or 548 can generate notifications (e.g., a notification such as Figure 1A shown as "Mary declined the meeting invitation"). In some examples, the additional data can include statistical data about the transcription including the number of words spoken by the speaker, the speaker's tone, information about the filler words used by the speaker (e.g., um, uh, oh, like, etc.), the percentage of time each speaker spoke, information about swear words used, information about the length of the words used, a summary of the transcription, or the speaker's mood (e.g., for each segment or the entire transcription). In another example, voice assistant applications 348, 448, or 548 can generate audio or video data in which the voice of Speaker 1 and / or Speaker 2 is replaced by another voice (e.g., the voice of a cartoon or a celebrity), or replace the language of any of the speech segments 704 in the audio or video file with a different language.

[0108] Figure 8 is a flowchart 800 that illustrates an example operation of a method for transcribing speech according to aspects of the present disclosure. Flowchart 800 is an example of the functionality executed by speech processing engines 341, 441, or 541 at Figure 6 element 604 of flowchart 600 in

[0109] Initially, speech recognition engines 342, 442, or 542 identify one or more speech segments (802) in the audio data (e.g., audio data 332, 432, 532, or 702). For example, speech recognition engines 342, 442, or 542 can use an analog-to-digital converter (ADC) to convert analog audio data 702 to digital data, filter the noise in the digitized audio data, and apply one or more statistical models (e.g., hidden Markov models or neural networks) to the filtered digitized audio data to identify Figure 7 speech segments 706A. In some examples, speech recognition engines 342, 442, or 542 can apply models trained to recognize for one or more specific users (e.g., Figure 1A - Figure 1CThe AI / ML model for the voice of user 110 is applied to audio data 702. For example, the speech recognition engines 342, 442, or 542 may apply an AI / ML model that is trained to recognize only the voice of the user of the HMD 112 (user 110). In some examples, the AI / ML model may receive training feedback from the user to adjust the speech recognition determination. In some examples, the speech recognition engines 342, 442, or 542 may identify one or more speech segments in the audio data 332, 432, or 532 based on the image data 330, 430, or 530. For example, the speech recognition engines 342, 442, or 542 may be configured to detect a face with moving lips in the image data to identify speech segments (e.g., the start and end of a speech segment).

[0110] The speaker identifier 344, 444, or 544 identifies the speaker (804) associated with the recognized speech segment. For example, based on the sound intensity (e.g., volume) of the speech segment 704A (e.g., for speech from the user of the HMD 112A in Figure 1B the sound intensity would be greater), the speaker identifier 344, 444, or 544 may identify speaker 1 as the speaker of the segment 704A in Figure 7 . In another example, the speaker identifier 344, 444, or 544 may use the image data captured by the image capture system 208 of the HMD 112 and / or the image capture system 508 of the computing device 120 to identify speaker 2 as the speaker of the segment 704C in Figure 7 . For example, the speaker identifier 344, 444, or 544 may be configured to detect a face with moving lips in the image data 330, 430, or 530 to identify the speaker and may identify the speaker based on the detected face with moving lips and / or the focus of the image data (e.g., indicating that user 110 is looking at the speaker). In another example, the audio capture system 209 of the HMD 112 or the audio capture system 509 of the computing system 120 may each include a microphone array that may respectively capture information about the directionality of the audio source relative to the HMD 112 or the computing device 120, and the speaker identifier 344, 444, or 544 may identify the speaker or potential speaker based on this directionality information and the image data 330, 430, or 530.

[0111] The speaker identifier 344, 444, or 544 uses the speaker identifier to label the recognized speech segment (806). For example, the speaker identifier 344, 444, or 544 uses the identifier "speaker 1" in Figure 7 to label the speech segment 704A. As described above with respect to Figure 7As described, in some examples, the speaker identifier 344, 444, or 544 automatically generates the identifier "Speaker 1" to include in the transcription 706. In some other examples, the user, administrator, or other source enters identifiers, tags, or names for one or more segments. These tags, identifiers, or names can provide an indication of the speaker of the speech segment in the transcription.

[0112] The speech transcriber 346, 446, or 546 transcribes the speech segments (808) recognized by the speech recognition engine 342, 442, or 542. For example, the speech transcriber 346, 446, or 546 generates a text output 706A for the segment 704A in Figure 7 . The speech processing engine 341, 441, or 541 then determines whether the speech recognition engine 342, 442, or 542 has recognized one or more additional speech segments (810) in the audio data (e.g., audio data 332, 432, 532, or 702). If the speech recognition engine 342, 442, or 542 has recognized one or more additional speech segments (the "Yes" branch of 810), then the elements 804 to 810 are repeated. For example, the speech recognition engine 342, 442, or 542 recognizes the speech segment 704B (802), the speaker identifier 344, 444, or 544 then identifies Speaker 1 as the speaker of the speech segment 704B (804) and marks the speech segment 704B with an indication that Speaker 1 is the speaker, and then the speech transcriber 346, 446, or 546 transcribes the speech segment 704B. This process can continue until no additional speech segments are recognized (e.g., when the interaction terminates, when audio / image data is no longer being captured, or when the entire audio data has been processed) (the "No" branch of 810), and the transcription is complete (812) (e.g., the flowchart 600 can continue to Figure 6 606 in

[0113] In some examples, the flowchart 800 processes audio and / or image data (e.g., audio and / or video streams or files) from two or more sources (e.g., received from two or more HMDS 112 and / or computing devices 120). In that instance, the operations of the flowchart 800 can be repeated for each audio data stream or file. In some examples, the flowchart 800 will combine the transcriptions of each audio data stream or file and produce a single complete transcription that includes an indication of the speaker for each speech segment in the transcription. For example, the flowchart 800 can use timestamps from each audio data file or stream to combine the transcriptions.

[0114] Figure 9FIG. 900 is a flow chart illustrating an example operation of a method for identifying a speaker of a voice segment in accordance with aspects of the present disclosure. Flow chart 900 is an example of the functionality performed at element 804 of flow chart 800 in Figure 8 by speaker identifier 344, 444, or 544.

[0115] Speaker identifiers 344, 444, 544 may determine a voice segment hash value for the voice segment (902). For example, voice processing engines 341, 441, or 541 may store each recognized voice segment in a separate file (e.g., a temporary file). These files may contain analog audio data or a digitized version of the audio data (e.g., noise other than the voice has been filtered). The speaker identifier may apply a hash function to these individual files to determine the voice segment hash value for each voice segment. Speaker identifiers 344, 444, 544 may obtain potential speaker models from speaker models 334, 434, or 534 (904), and compare the voice segment hash value with the hash values of the potential speaker models (906). Speaker identifiers 344, 444, 544 identify the closest speaker model, the closest speaker having a hash value closest to the voice segment hash value (908).

[0116] If the difference between the voice segment hash value and the closest speaker model is equal to or greater than a threshold difference (the "no" branch of 910), then speaker identifier 344, 444, or 544 may create a new speaker model based on the voice segment hash value (916). For example, speaker identifier 344, 444, or 544 will determine a new speaker identifier (ID) for the voice segment hash value, and store the new speaker ID and the voice segment hash value as a new speaker model in speaker models 334, 434, or 534. Speaker identifier 344, 444, or 544 then returns the new speaker ID as the speaker of the voice segment (918) (e.g., flow chart 800 may continue to Figure 8 806 in

[0117] If the difference between the speech segment hash value of a speech segment and the hash value of the closest speaker model is less than a threshold difference (the "yes" branch of 910), then the speaker identifier 344, 444, or 544 updates the closest speaker model based on the speech segment hash value (912). For example, the hash value of the closest speaker model can include the average hash value of all speech segments associated with that speaker, and the speaker identifier 344, 444, or 544 can incorporate the speech segment hash value into that average. The speaker identifier 344, 444, or 544 then returns the speaker ID of the closest speaker model as the speaker of that speech segment (914) (e.g., the flowchart 800 can continue to Figure 8 806 associated with the closest speaker ID).

[0118] Figure 10 FIG. 1000 is a flowchart illustrating an example operation of a method for identifying potential speaker models in accordance with aspects of the present disclosure. Flowchart 1000 is an example of a function performed by the speaker identifier 344, 444, or 544 at element 904 of flowchart 900 in Figure 9 .

[0119] The speaker identifier 344, 444, or 544 can identify potential speaker models based on a number of inputs (1010). For example, the speaker identifier 344, 444, or 544 can obtain external data (1002) and process the external data to identify one or more potential speaker models (1010). In some examples, the external data can include location information of one or more users (e.g., GPS coordinates). For example, the speaker identifier 344, 444, or 544 can determine one or more users (or devices associated with one or more users) near the HMD 112 or the computing device 120 (e.g., within 50 feet) and use that information to obtain the speaker models associated with those users / devices (e.g., from speaker models 334, 434, or 534). In some examples, the external information can include calendar information, including invitee information for a meeting, location information for the meeting, and an indication of whether each invitee plans to attend the meeting. In some examples, the speaker identifier 344, 444, or 544 will identify the speaker models corresponding to all invitees in the calendar information. In some other examples, the speaker identifier 344, 444, or 544 will identify the speaker models corresponding to all invitees in the calendar information who plan to attend the meeting.

[0120] In some examples, the speaker identifier 344, 444, or 544 may obtain image data (1004) and process the image data to identify one or more potential speaker models (1010). For example, the speaker identifier 344, 444, or 544 may be configured to detect faces in the image data and identify the speaker model associated with the detected face (e.g., from speaker models 334, 434, or 534). In some other examples, the speaker identifier 344, 444, or 544 may be configured to detect a face with moving lips in the image data that corresponds to the speech segment recognized in the audio data, and identify the speaker model associated with the detected face with moving lips (e.g., speaker models 334, 434, or 534). In some examples, the speaker identifier 344, 444, or 544 may apply an AI / ML model trained to identify faces and / or faces with moving lips in images to the image data. In another example, the audio capture system 209 of the HMD 112 or the audio capture system 509 of the computing system 120 may each include a microphone array that may respectively capture information about the directionality of the audio source relative to the HMD 112 or the computing device 120, and the speaker identifier 344, 444, or 544 may identify the speaker or potential speaker based on the directionality information and the face detected in the image data. For example, based on the directionality information about the speech segment 704 and the correspondence of the directionality with the face of the person 101A in Figure 1C , the speaker identifier 344, 444, or 544 may identify the speaker 2 as Figure 7 the speaker of the speech segment 704C in. In yet another example, the speaker identifier 344, 444, or 544 will identify the speaker based on who the user is paying attention to (e.g., based on the field of view of the HMD 112).

[0121] In some examples, the speaker identifier 344, 444, or 544 may receive user input (1006) and process the user input to identify one or more potential speaker models (1010). For example, the speaker or speaker model may be identified (e.g., from speaker models 334, 434, or 534). In some other examples, the user may confirm the potential speaker model identified based on external data or image data.

[0122] Figure 11 FIG. 1100 is a flow chart illustrating example operations of a method for transcribing speech for distributed devices in accordance with aspects of the present disclosure. In some examples, Figure 11 one or more of the operations shown in may be performed by the HMD 112, the computing device 120, and / or the transcription system 106.

[0123] The audio capture system 209 and the image capture system 208 of the HMD 112 and / or the audio capture system 509 and the image capture system 508 of the computing device 120 capture audio and image data (1102). For example, two or more HMDs 112 and / or computing devices 120 may capture audio and / or image data (e.g., from the same or different physical environments).

[0124] The speech processing engines 341, 441, or 541 use a user speaker model (e.g., a speaker model specific to a user of a device) to transcribe the audio data using the image data for each device (1104). For example, in Figure 1B the speech processing engine of the HMD112A (e.g., using a speaker model specific to user 110) transcribes the speech segment corresponding to user 110, the speech processing engine of the HMD 112B (e.g., using a speaker model specific to user 101A) transcribes the speech segment corresponding to user 101A, and the speech processing engine of the HMD 112C (e.g., using a speaker model specific to user 102A) transcribes the speech segment corresponding to user 102A. In some examples, a user logs into the HMD 112 or the computing device 120 or otherwise identifies himself or herself as a user. In some other examples, the HMD 112 or the computing device 120 (e.g., using the voice and / or face recognition techniques described above) automatically identifies the user. For example, the speech processing engines 341, 441, or 541 transcribe each speech segment in the speech segments to produce a transcription including an indication of the speaker associated with each speech segment. In some examples, Figure 1C any of the HMDs 112A, 112B, and / or 112C of Figure 4 may capture audio and image data and transmit the audio and image data to the transcription system 106 for transcription (e.g., as described above with reference to Figure 1C . For example, the transcription system 106 may receive audio and image data from one or more of the HMDs 112A, 112B, and / or 112C of

[0125] The speech processing engines 341, 441, or 541 then combine all of the transcripts in the transcriptions corresponding to the speech segments in the audio data captured by two or more HMDs 112 and / or computing devices 120 to produce one complete transcription (1106) that includes an indication of the speaker / user associated with each transcribed speech segment. For example, each of the HMDs 112A, 112B, and 112C may transmit individual transcripts of the speech captured from users 110, 101A, and 102A, respectively, to the transcription system 106, which will combine the individual transcripts. In another example, the HMDs 112B and 112C may transmit individual transcripts of the speech captured from users 101A and 102A, respectively, to the HMD 112A, which will combine the individual transcripts. In some examples, the voice assistant applications 348, 448, or 548 then optionally analyze the individual transcriptions and / or the complete transcription to produce additional data derived from the transcription (e.g., as described above with reference to Figure 6 ).

[0126] The techniques described in this disclosure may be implemented, at least in part, in hardware, software, firmware, or any combination thereof. For example, aspects of the described techniques may be implemented within one or more processors that include one or more microprocessors, DSPs, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or any other equivalent integrated or discrete logic circuitry, as well as any combination of such components. The term "processor" or "processing circuitry" generally may refer to any of the foregoing logic circuitry, alone or in combination with other logic circuitry or any other equivalent circuitry. A control unit including hardware may also perform one or more of the techniques of this disclosure.

[0127] Such hardware, software, and firmware may be implemented within the same device or in separate devices to support the various operations and functions described in this disclosure. Additionally, any of the described units, modules, or components may be implemented together or separately as discrete but interoperable logic devices. Describing different features as modules or units is intended to highlight different functional aspects and does not necessarily imply that these modules or units must be implemented by separate hardware or software components. Rather, the functionality associated with one or more modules or units may be performed by separate hardware or software components, or may be integrated within common or separate hardware or software components.

[0128] The techniques described in this disclosure may also be embodied or encoded in a computer-readable medium, such as a computer-readable storage medium, that includes instructions. The instructions embedded or encoded in the computer-readable storage medium may cause a programmable processor or other processor to perform a method when the instructions are executed. The computer-readable storage medium may include random access memory (RAM), read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), flash memory, a hard disk, a CD-ROM, a floppy disk, magnetic tape, magnetic media, optical media, or other computer-readable media.

[0129] As described herein by various examples, the techniques of this disclosure may include or be implemented in conjunction with an artificial reality system. As described, artificial reality is a form of reality that has been adjusted in some manner before being presented to a user and may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include entirely generated content, or a combination of generated content and captured content (such as a photograph of the real world). Artificial reality content may include video, audio, haptic feedback, or some combination thereof, and any of these may be presented in a single channel or multiple channels (such as stereoscopic video that produces a three-dimensional effect for a viewer). Additionally, in some embodiments, artificial reality may be associated with, for example, applications, products, accessories, services, or some combination thereof that are used to create content in artificial reality and / or are used in artificial reality (such as to perform an activity therein). An artificial reality system that provides artificial reality content may be implemented on various platforms, including a head-mounted device (HMD) connected to a host computer system, a stand-alone HMD, a mobile device or computing system, or any other hardware platform capable of providing artificial reality content to one or more viewers.

[0130] In certain embodiments, one or more objects of a computing system (e.g., content or other types of objects) may be associated with one or more privacy settings. The one or more objects may be stored on or otherwise associated with any suitable computing system or application, such as, for example, a social networking system, a client system, a third-party system, a social networking application, a messaging application, a photo sharing application, or any other suitable computing system or application. Although the examples discussed herein are in the context of an online social network, these privacy settings may be applied to any other suitable computing system. Privacy settings for an object (or “access settings”) may be stored in any suitable manner, such as, for example, associated with the object, as an index on an authorization server, in another suitable manner, or any suitable combination thereof. Privacy settings for an object may specify how the object (or specific information associated with the object) may be accessed, stored, or otherwise used (e.g., viewed, shared, modified, copied, executed, surfaced, or identified) in an online social network. When the privacy settings of an object permit a particular user or other entity to access the object, the object may be described as “visible” to that user or other entity. By way of example and not limitation, a user of an online social network may specify privacy settings for a user profile page that identify a group of users who may access work experience information on the user profile page, thereby excluding other users from accessing that information.

[0131] In certain embodiments, privacy settings for an object can specify a "block list" of users or other entities that should not be permitted to access certain information associated with the object. In certain embodiments, the block list can include third-party entities. The block list can specify one or more users or entities to which the object is not visible. By way of example and not limitation, a user can specify a group of users who cannot access an album associated with that user, thereby excluding those users from accessing the album (while perhaps also allowing certain users not in the specified group of users to access the album). In certain embodiments, privacy settings can be associated with specific social graph elements. Privacy settings for a social graph element (such as a node or an edge) can specify how the social graph element, information associated with the social graph element, or an object associated with the social graph element can be accessed using an online social network. By way of example and not limitation, a particular concept node corresponding to a particular photo can have privacy settings that specify that the photo can only be accessed by the users tagged in the photo and friends of the users tagged in the photo. In certain embodiments, privacy settings can allow a user to opt in or opt out of having the social networking system store / record their content, information, or actions or share with other systems (e.g., third-party systems). Although this disclosure describes using specific privacy settings in a specific manner, this disclosure contemplates using any suitable privacy settings in any suitable manner.

[0132] In certain embodiments, privacy settings can be based on one or more nodes or edges of a social graph. Privacy settings can be specified for one or more edges or edge types of the social graph, or for one or more nodes or node types of the social graph. Privacy settings applied to a particular edge connecting two nodes can control whether the relationship between the two entities corresponding to the nodes is visible to other users of the online social network. Similarly, privacy settings applied to a particular node can control whether the user or concept corresponding to the node is visible to other users of the online social network. By way of example and not limitation, a first user can share an object with the social networking system. The object can be associated with a concept node of a user node that is connected to the first user by an edge. The first user can specify privacy settings for the particular edge connecting to the concept node of the object, or can specify privacy settings for all edges connecting to the concept node. As another example and not limitation, a first user can share a group of objects of a particular object type (e.g., a group of images). The first user can specify that the privacy settings for all objects associated with the first user of that particular object type have a particular privacy setting (e.g., specifying that all images posted by the first user are only visible to friends of the first user and / or users tagged in the images).

[0133] In certain embodiments, a social networking system may present a "privacy wizard" to a first user (e.g., within a web page, module, one or more dialog boxes, or any other suitable interface) to assist the first user in specifying one or more privacy settings. The privacy wizard may display instructions, appropriate privacy-related information, current privacy settings, one or more input fields for receiving one or more inputs for accepting changes or confirmations of the specified privacy settings from the first user, or any suitable combination thereof. In certain embodiments, the social networking system may provide "dashboard" functionality to the first user, which may display the first user's current privacy settings to the first user. The dashboard functionality may be displayed to the first user at any appropriate time (e.g., after an input from the first user invoking the dashboard functionality, after the occurrence of a particular event or trigger action). The dashboard functionality may allow the first user to modify one or more of the first user's current privacy settings in any suitable manner at any time (e.g., redirecting the first user to the privacy wizard).

[0134] The privacy settings associated with an object may specify any suitable granularity of access allowed or denied. By way of example and not limitation, access may be specified for particular users (e.g., only me, my roommate, my boss), users within a particular degree of separation (e.g., friends, friends of friends), user groups (e.g., gaming club, my family), user networks (e.g., employees of a particular employer, students or alumni of a particular university), all users ("public"), no users ("private"), users of third-party systems, particular applications (e.g., third-party applications, external websites), other suitable entities, or any suitable combination thereof. Although this disclosure describes particular granularities of access allowed or denied, this disclosure contemplates any suitable granularity of access allowed or denied.

[0135] In certain embodiments, one or more servers can be authorization / privacy servers for enforcing privacy settings. In response to a request from a user (or other entity) for a particular object stored in a data repository, the social networking system can send a request for the object to the data repository. The request can identify the user associated with the request, and the object can be sent to the user (or the user's client system) only if the authorization server determines that the user is authorized to access the object based on the privacy settings associated with the object. If the requesting user is not authorized to access the object, the authorization server can prevent the requested object from being retrieved from the data repository or can prevent the requested object from being sent to the user. In the context of a search query, an object can be provided as a search result only if the querying user is authorized to access the object, e.g., if the privacy settings of the object permit it to be presented to, discovered by, or otherwise made visible to the querying user. In certain embodiments, an object can represent content visible to a user via the user's newsfeed. By way of example and not limitation, one or more objects can be visible to a user's "trending topics" page. In certain embodiments, an object can correspond to a particular user. The object can be content associated with a particular user, or can be an account or information of a particular user stored on the social networking system or other computing system. By way of example and not limitation, a first user can view one or more second users of an online social network via the online social network's "People You May Know" feature or by viewing the first user's friends list. By way of example and not limitation, a first user can specify that they do not wish to see objects associated with a particular second user in their newsfeed or friends list. If the privacy settings of an object do not permit it to be presented to, discovered by, or made visible to a user, the object can be excluded from search results. Although this disclosure describes enforcing privacy settings in a particular manner, this disclosure contemplates enforcing privacy settings in any suitable manner.

[0136] In certain embodiments, different objects of the same type associated with a user can have different privacy settings. Different types of objects associated with a user can have different types of privacy settings. By way of example and not limitation, a first user can specify that status updates of the first user are public, but any images shared by the first user are visible only to friends of the first user on an online social network. As another example and not limitation, a user can specify different privacy settings for different types of entities (e.g., individual users, friends of friends, followers, user groups, or corporate entities). As another example and not limitation, a first user can specify a user group that can view a video posted by the first user while preventing the video from being visible to the first user's employer. In certain embodiments, different privacy settings can be provided for different user groups or user demographics. By way of example and not limitation, a first user can specify that other users who attended the same university as the first user can view the first user's photos, but other users who are family members of the first user cannot view those same photos.

[0137] In certain embodiments, a social networking system can provide one or more default privacy settings for each object of a particular object type. The privacy settings for an object that are set as the default can be changed by the user associated with the object. By way of example and not limitation, all images posted by a first user can have a default privacy setting of visible only to friends of the first user, and for a particular image, the first user can change the privacy setting for that image to visible to friends and friends of friends.

[0138] In certain embodiments, privacy settings can allow a first user to specify (e.g., by opting out, by not opting in) whether the social networking system can receive, collect, record, or store a particular object or information associated with the user for any purpose. In certain embodiments, privacy settings can allow a first user to specify whether a particular application or process can access, store, or use a particular object or information associated with the user. Privacy settings can allow a first user to opt in or opt out of having an object or information accessed, stored, or used by a particular application or process. The social networking system can access such information in order to provide a particular function or service to the first user, and the social networking system cannot access the information for any other purpose. Before accessing, storing, or using such an object or information, the social networking system can prompt the user to provide privacy settings that specify which applications or processes (if any) can access, store, or use the object or information before allowing any such action. By way of example and not limitation, a first user can transmit a message to a second user via an application associated with an online social network (e.g., a messaging application) and can specify a privacy setting that such a message should not be stored by the social networking system.

[0139] In certain embodiments, a user may specify whether a particular type of object or information associated with a first user can be accessed, stored, or used by a social networking system. By way of example and not limitation, the first user may specify that images sent by the first user via the social networking system cannot be stored by the social networking system. As another example and not limitation, the first user may specify that messages sent from the first user to a particular second user cannot be stored by the social networking system. As yet another example and not limitation, the first user may specify that all objects sent via a particular application can be saved by the social networking system.

[0140] In certain embodiments, privacy settings may allow a first user to specify whether a particular object or information associated with the first user can be accessed from a particular client system or third-party system. The privacy settings may allow the first user to opt in or opt out of having the object or information accessed from a particular device (e.g., the phone book on the user's smart phone), from a particular application (e.g., a messaging application), or from a particular system (e.g., an email server). The social networking system may provide default privacy settings for each device, system, or application, and / or the first user may be prompted to specify particular privacy settings for each context. By way of example and not limitation, the first user may utilize the location services feature of the social networking system to provide recommendations about restaurants or other places near the user. The first user's default privacy settings may specify that the social networking system can use location information provided from the first user's client device to provide location-based services, but that the social networking system cannot store the first user's location information or provide it to any third-party system. The first user may then update the privacy settings to allow a third-party image sharing application to use the location information in order to geotag photos.

Claims

1. A system, comprising: An audio capture system configured to capture audio data associated with a plurality of speakers; An image capture system configured to capture images of one or more of the plurality of speakers; And A speech processing engine configured to: Identify a plurality of speech segments in the audio data, For each of the plurality of speech segments, identify the speaker associated with the speech segment, Transcribe each of the plurality of speech segments to produce a transcription of the plurality of speech segments, wherein for each of the plurality of speech segments, the transcription includes an indication of the speaker associated with the speech segment, and Analyze the transcription to produce additional data derived from the transcription; Wherein the speech processing engine is further configured to access external data; And Wherein in order to identify the speaker associated with the speech segment, the speech processing engine is configured to identify the speaker based on the image and the external data, wherein the external data includes location information associated with the speaker.

2. The system according to claim 1, wherein, in order to identify the plurality of speech segments, the speech processing engine is further configured to identify the plurality of speech segments based on the image; wherein, in order to identify the speaker for each speech segment of the plurality of speech segments, the speech processing engine is further configured to detect one or more faces in the image.

3. The system according to claim 2, wherein the speech processing engine is further configured to select one or more speech recognition models based on the identity of the speaker associated with each speech segment.

4. The system according to claim 3, wherein, in order to identify the speaker for each speech segment of the plurality of speech segments, the speech processing engine is further configured to detect one or more faces in the image having moving lips.

5. The system according to claim 3 or 4, further comprising a head-mounted display (HMD) wearable by a user, and wherein the one or more speech recognition models include a speech recognition model for the user's voice; wherein the HMD is configured to output artificial reality content, and wherein the artificial reality content includes a virtual meeting application, and the virtual meeting application includes a video stream and an audio stream.

6. The system according to claim 3 or 4, further comprising a head-mounted display (HMD) wearable by a user, wherein the speech processing engine is further configured to identify the user of the HMD as the speaker of the plurality of speech segments based on the attributes of the plurality of speech segments.

7. The system according to any one of claims 1 to 4, wherein the audio capture system includes a microphone array; wherein the additional data includes an audio stream, and the audio stream includes a modified version of the speech segments associated with at least one of the plurality of speakers.

8. The system according to any one of claims 1 to 4, wherein the additional data includes one or more of the following: a calendar invitation for the meeting or event described in the transcription, information related to the topic identified in the transcription, or a task list including the tasks identified in the transcription.

9. The system according to any one of claims 1 to 4, wherein the additional data includes at least one of the following: statistical data about the transcription including the number of words spoken by the speaker, the tone of the speaker, information about the filler words used by the speaker, the percentage of time the speaker speaks, information about the swear words used, information about the length of the words used, a summary of the transcription, or the mood of the speaker.

10. A method, comprising: Capture audio data associated with a plurality of speakers; Capture images of one or more of the plurality of speakers; Identify a plurality of speech segments in the audio data; For each of the plurality of speech segments, identify the speaker associated with the speech segment; Transcribe each of the plurality of speech segments to produce a transcription of the plurality of speech segments, wherein for each of the plurality of speech segments, the transcription includes an indication of the speaker associated with the speech segment; Analyze the transcription to produce additional data derived from the transcription; And Access external data, and identifying the speaker associated with the speech segment includes identifying the speaker based on the image and the external data, the external data including location information associated with the speaker.

11. The method according to claim 10, wherein the additional data includes one or more of the following: a calendar invitation for the meeting or event described in the transcription, information related to the topic identified in the transcription, or a task list including the tasks identified in the transcription.

12. The method according to claim 10 or 11, wherein the additional data includes at least one of the following: statistical data about the transcription including the number of words spoken by the speaker, the tone of the speaker, information about the filler words used by the speaker, the percentage of time the speaker speaks, information about the swear words used, information about the length of the words used, a summary of the transcription, or the mood of the speaker.

13. A computer-readable storage medium including instructions that, when executed, configure the processing circuitry of a computing system to: Capture audio data associated with a plurality of speakers; Capture images of one or more of the plurality of speakers; Identify a plurality of speech segments in the audio data; For each of the plurality of speech segments, identify the speaker associated with the speech segment; Transcribe each of the plurality of speech segments to produce a transcription of the plurality of speech segments, wherein for each of the plurality of speech segments, the transcription includes an indication of the speaker associated with the speech segment; Analyze the transcription to produce additional data derived from the transcription; And Access external data, and in order to identify the speaker associated with the speech segment, the processing circuitry is configured to identify the speaker based on the image and the external data, the external data including location information associated with the speaker.

Citation Information

Patent Citations

  • Augmented Reality Conferencing System and Method

    US20180123813A1

  • Speech-inclusive device interfaces

    US8700392B1