Privacy aware meeting room transcription from audiovisual streams

By receiving audio-visual signals in the video conferencing system, segmenting audio data, and determining the speaker's identity based on image data, and applying corresponding conditions in combination with the participant's privacy request, the problems of speaker identification and privacy protection in video conferencing are solved, and transcript generation with high accuracy and privacy protection are achieved.

CN120091101APending Publication Date: 2025-06-03GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510148254.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2019-11-18
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify speakers and meet participants’ privacy needs while generating video conference transcripts, especially in a multi-participant environment.

Method used

By receiving an audio-visual signal including audio data and image data, the audio data is divided into a plurality of segments, and the speaker identity of each segment is determined based on the image data. Then, according to the participant's privacy request, the corresponding privacy conditions are applied, such as deleting or hiding the speaker's identity information.

Benefits of technology

More accurate speaker identification and transcript generation are achieved, while meeting participants’ privacy needs, improving the privacy protection of video conferencing and the accuracy of transcripts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091101A_ABST
    Figure CN120091101A_ABST
Patent Text Reader

Abstract

The invention relates to privacy aware meeting room transcription from an audiovisual stream. A method for privacy aware transcription includes receiving an audiovisual signal including audio data and image data of a voice environment and a privacy request from a participant in the voice environment, where the privacy request indicates a privacy condition of the participant. The method further includes segmenting the audio data into a plurality of segments. For each segment, the method includes determining an identity of a speaker of a corresponding segment of the audio data based on the image data, and determining whether the identity of the speaker of the corresponding segment includes a participant associated with a privacy condition. When the identity of the speaker of the corresponding segment includes the participant, the method includes applying the privacy condition to the corresponding segment. The method also includes processing the plurality of segments of the audio data to determine a transcript of the audio data.
Need to check novelty before this filing date? Find Prior Art

Description

Division Case Explanation

[0001] This application is a divisional application of Chinese Patent Application No. 201980102276.1 with an application date of November 18, 2019. Technical Field

[0002] This disclosure relates to privacy-aware conference room transcription from audiovisual streams. Background Art

[0003] Speaker diarization is the process of partitioning an input audio stream into homogeneous segments based on speaker identity. In an environment with multiple speakers, speaker diarization answers the question "who is speaking when" and has various applications, including multimedia information retrieval, speaker turn analysis, and audio processing, to name just a few. In particular, speaker diarization systems are capable of generating speaker boundaries that have the potential to significantly improve the accuracy of acoustic speech recognition. Summary of the Invention

[0004] One aspect of the present disclosure provides a method for generating a privacy-aware conference room transcript from a content stream. The method includes receiving, at data processing hardware, an audiovisual signal including audio data and image data. The audio data corresponds to voice utterances from multiple participants in a voice environment, and the image data represents the faces of the multiple participants in the voice environment. The method further includes receiving, at the data processing hardware, a privacy request from a participant among the multiple participants. The privacy request indicates privacy conditions associated with the participant in the voice environment. The method further includes segmenting, by the data processing hardware, the audio data into a plurality of segments. For each segment of the audio data, the method includes determining, by the data processing hardware, the identity of the speaker of the corresponding segment of the audio data from among the multiple participants based on the image data. For each segment of the audio data, the method further includes determining, by the data processing hardware, whether the identity of the speaker of the corresponding segment includes a participant associated with the privacy conditions indicated by the received privacy request. When the identity of the speaker of the corresponding segment includes a participant, the method includes applying the privacy conditions to the corresponding segment. The method further includes processing, by the data processing hardware, the plurality of segments of the audio data to determine a transcript of the audio data.

[0005] Embodiments of the present disclosure may include one or more of the following optional features. In some embodiments, applying the privacy conditions to the corresponding segment includes deleting the corresponding segment of the audio data after determining the transcript. Additionally or alternatively, applying the privacy conditions to the corresponding segment may include enhancing the corresponding segment of the image data to visually hide the identity of the speaker of the corresponding segment of the audio data.

[0006] In some examples, for each part of a transcript corresponding to a segment of audio data to which privacy conditions are applied, processing multiple segments of the audio data to determine the transcript of the audio data includes modifying the corresponding part of the transcript to not include the identity of the speaker. Optionally, for each segment of the audio data to which privacy conditions are applied, processing multiple segments of the audio data to determine the transcript of the audio data can include omitting the corresponding segment of the transcribed audio data. The privacy conditions can include content-specific conditions that indicate the type of content to be excluded from the transcript.

[0007] In some configurations, determining the identity of the speaker of a corresponding segment of audio data from among multiple participants includes determining multiple candidate identities of the speaker based on image data. Here, for each of the multiple candidate identities, a confidence score is generated that indicates the likelihood that the face corresponding to the candidate identity based on the image data includes the speaking face of the corresponding segment of the audio data. In this configuration, the method includes selecting the identity of the speaker of the corresponding segment of the audio data as the candidate identity among the multiple candidate identities associated with the highest confidence score.

[0008] In some embodiments, the data processing hardware resides on a device local to at least one of the multiple participants. The image data can include high-definition video processed by the data processing hardware. Processing multiple segments of the audio data to determine the transcript of the audio data can include processing the image data to determine the transcript.

[0009] Another aspect of the present disclosure provides a system for privacy-aware transcription. The system includes data processing hardware and memory hardware communicatively coupled to the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving an audiovisual signal including audio data and image data. The audio data corresponds to voice utterances from multiple participants in a voice environment, and the image data represents the faces of the multiple participants in the voice environment. The operations further include receiving a privacy request from a participant among the multiple participants, the privacy request indicating privacy conditions associated with the participant in the voice environment. The method further includes splitting the audio data into multiple segments. For each segment of the audio data, the operations include determining the identity of the speaker of the corresponding segment of the audio data from among the multiple participants based on the image data. For each segment of the audio data, the method also includes determining whether the identity of the speaker of the corresponding segment includes a participant associated with the privacy conditions indicated by the received privacy request. When the identity of the speaker of the corresponding segment includes a participant, the operations include applying the privacy conditions to the corresponding segment. The operations further include processing multiple segments of the audio data to determine the transcript of the audio data.

[0010] This aspect may include one or more of the following optional features. In some examples, applying a privacy condition to a corresponding segment includes deleting the corresponding segment of the audio data after determining the transcript. Optionally, applying a privacy condition to a corresponding segment may include enhancing the corresponding segment of the image data to visually hide the identity of the speaker of the corresponding segment of the audio data.

[0011] In some configurations, processing multiple segments of audio data to determine a transcript of the audio data includes, for each portion of the transcript corresponding to a segment of the audio data to which a privacy condition is applied, modifying the corresponding portion of the transcript to not include the identity of the speaker. Additionally or alternatively, processing multiple segments of audio data to determine a transcript of the audio data may include, for each segment of the audio data to which a privacy condition is applied, omitting the corresponding segment of the transcribed audio data. The privacy condition may include a content-specific condition that indicates the type of content to be excluded from the transcript.

[0012] In some embodiments, the operation of determining the identity of the speaker of a corresponding segment of audio data from a plurality of participants includes determining a plurality of candidate identities of the speaker based on image data. The embodiment includes, for each candidate identity of the plurality of candidate identities, generating a confidence score that indicates the likelihood that the face corresponding to the candidate identity based on the image data includes the speaking face of the corresponding segment of the audio data. The embodiment further includes selecting the identity of the speaker of the corresponding segment of the audio data as the candidate identity among the plurality of candidate identities associated with the highest confidence score.

[0013] In some examples, the data processing hardware resides on a device local to at least one of the plurality of participants. The image data may include high-definition video processed by the data processing hardware. Processing multiple segments of audio data to determine a transcript of the audio data may include processing the image data to determine the transcript.

[0014] Details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the following description. Other aspects, features, and advantages will be apparent from the specification, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1A is a schematic diagram of an exemplary assembly environment with a transcriber.

[0016] Figure 1B - 1E is an Figure 1A exemplary assembly environment with a privacy-aware transcriber.

[0017] Figure 2A and 2B are schematic diagrams of exemplary transcribers.

[0018] Figure 3 is a flowchart of an exemplary arrangement of operations of a method for transcribing content within an assembly environment Figure 1A .

[0019] Figure 4 is a schematic diagram of an exemplary computing device that can be used to implement the systems and methods described herein.

[0020] Figure 5 is a schematic diagram of an exemplary profile stored in a memory hardware accessible to a transcriber.

[0021] Like reference numerals in the various figures indicate like elements. DETAILED DESCRIPTION

[0022] The privacy of data used and generated by a video conferencing system is an important aspect of such a system. Conference participants can have their own personal views on the privacy of the audio and video data obtained during a conference. Thus, there is a technical problem of how to provide a video conferencing system that can accurately generate a transcript for a video conferencing session while also meeting such privacy requirements in a reliable and accurate manner. Embodiments of the present disclosure provide a technical solution by enabling participants in a conference to set their own privacy configurations (e.g., opt-in or opt-out of various functions of the video conferencing system), and then the video conferencing system accurately and effectively implements the expectations of the participants. Since the video conferencing system not only identifies the verbal contributions from participants based on the audio captured during the conference but also based on the video captured during the conference when generating a transcript - this ensures higher accuracy in the identification of contributors to the video conference, thereby enabling the improvement of the accuracy of the transcript while enabling the accurate and reliable implementation of the customized privacy requirements of the participants. In other words, a more accurate, reliable, and flexible video conferencing system is provided.

[0023] In addition, in some implementations, the process of generating a transcript of a video conference is performed locally for one or more participants of the video conference, e.g., by a device in the same room as those participants. In other words, in such embodiments, the process of generating the transcript is not performed remotely, such as via one or more remote / cloud servers. This helps to ensure that certain privacy expectations can be met while also ensuring that full / original resolution and full / original quality video data captured locally is available for use in identifying speakers during the video conference (as opposed to remote servers that may operate on lower resolution and / or lower quality video, which may reduce the accuracy of speaker identification).

[0024] In a gathering environment (commonly referred to as an environment), people come together to convey thoughts, ideas, schedules, or other concerns. The gathering environment serves as a shared space for its participants. This shared space can be a physical space, such as a meeting room or classroom, a virtual space (e.g., a virtual meeting room), or any combination thereof. The environment can be a centralized location (e.g., locally hosted) or a decentralized location (e.g., virtually hosted). For example, the environment is a single room where participants gather, such as a meeting room or classroom. In some implementations, the environment is more than one shared space linked together to form a gathering of participants. For example, a meeting has a host location (e.g., the location where the coordinator or speaker of the meeting may be) and one or more remote locations where participants attend (e.g., using a real-time communication application). In other words, a company hosts a meeting from an office in Chicago, but other offices of the company remotely attend the meeting (e.g., in San Francisco or New York). For example, there are many companies that have large meetings across several offices, where each office has a meeting space for participating in the meeting. This is especially true as it becomes more common for team members to be distributed across a company (i.e., in more than one location) or even work remotely. Additionally, as applications become more robust for real-time communication, environments can be hosted for remote offices, remote employees, remote partners (e.g., business partners), remote customers, etc. Thus, the environment has evolved to accommodate a wide variety of gathering and organizing efforts.

[0025] Typically, as a space for communication, the environment hosts multiple participants. Here, each participant can contribute audio content (e.g., audible utterances by speaking) and / or visual content (e.g., actions of the participant) while present in the environment. In the case where there are more than one participant in the environment, it is beneficial to track and / or record the participation of any or all of the participants. This is especially true when the environment accommodates extensive assembly organization work. For example, when the Chicago office hosts a meeting with remote attendance from both the New York office and the San Francisco office, someone in the Chicago office may have difficulty identifying the speaker in one of the remote locations. To illustrate, the Chicago office may include video feeds of the conference rooms of each office located far from the Chicago office. Even with the video feeds, the participants in the Chicago office may not be able to distinguish all the participants in the New York office. For example, the speaker in the New York office is located in a position far from the camera associated with the video feed, making it difficult for the participants in the Chicago office to identify who the speaker is in the New York office. This can also be difficult when the Chicago-based participant is not familiar with the other participants in the meeting (e.g., cannot identify the speaker by his / her voice). When the speaker cannot be identified, this can be problematic because the identity of the speaker may be a key component during the meeting. In other words, it may be important to identify the speaker (or the content source) to understand the key points / deliverables or generally understand who shared what content. For example, if Sally in the New York office took on an action item deliverable to Johnny in the Chicago office, but Johnny cannot identify that Sally took on the action item, Johnny may have difficulty following up on the action item later. In another scenario, because Johnny cannot identify that Sally took on the action item, Johnny may incorrectly identify Tracy (who is also in the New York office, for example) as having taken on the action item. This is also true at the basic level of a simple conversation between participants. If Sally talks about a certain topic, but Johnny thinks it is Tracy who is speaking, Johnny may cause confusion when he asks Tracy to join the conversation about the topic at a later moment in the meeting.

[0026] Another problem can arise when the speaker discusses names, acronyms, and / or jargon that are unfamiliar and / or not fully understood by another participant. In other words, Johnny might be discussing a problem with a conveyance used during transportation. Pete might chime in to help solve Johnny's problem by saying, "oh, you will want to speak with Teddy in logistics about that." If Johnny is not familiar with Teddy and / or the logistics team, Johnny might make a note to talk to Freddie instead of Teddy. This can also happen with acronyms or other jargon used in a given industry. For example, if the Chicago office is having a meeting with a Seattle company, where the Chicago office hosts the meeting and the Seattle company participates remotely, the participants in the Chicago office might use acronyms and / or jargon that are unfamiliar to the Seattle company. Unfortunately, without a record or transcription of what was presented by the Chicago office, the Seattle company might not be able to understand the meeting (e.g., resulting in a poor meeting outcome). Additionally or alternatively, a poor connection between locations or to the meeting hosting platform can also complicate matters for participants as they try to understand the content during the meeting.

[0027] To overcome these problems, there is a transcription device in the environment that generates (e.g., in real-time) a transcript of what is happening within the environment. When generating the transcript, the device can identify the speaker (i.e., the participant generating the audio content) and / or associate the content with the participants who are also present within the environment. Using the transcript of the content presented in the environment, the transcription device is able to remember key points and / or deliverables and provide a record of who initiated the content accessible to the participants for reference. For example, a participant can refer to the display of the transcript generated by the transcription device during the meeting (e.g., in real-time or substantially in real-time) or at some later time after the meeting. In other words, Johnny can refer to the display of the transcript generated by the transcription device to identify that Teddy (and not Freddie) is the person he needs to talk to in logistics and that he should follow up with Sally (and not Tracy) regarding that action item.

[0028] Unfortunately, although transcripts may solve some of the problems encountered in the environment, they raise questions about privacy. Here, privacy refers to the state of being free from observation on the transcripts generated by the transcription device. Although there may be many different types of privacy, some examples include content privacy or identity privacy. Here, content privacy is content-based, such that it is desired that certain sensitive content not be remembered in written or human-readable format (e.g., confidential content). For example, part of a meeting may include audio content about another employee who is not present in the meeting (e.g., a manager discussing a human resources issue that has arisen). In this example, the participants in the meeting would prefer not to transcribe or otherwise remember this part of the meeting about the other employee. This can also include not remembering the audio content that includes the content about the other employee. Here, since traditional transcription devices transcribe content indiscriminately, the meeting will not be able to utilize traditional transcription devices, at least during this part of the meeting.

[0029] Identity privacy refers to privacy that seeks to maintain the anonymity of the content source. For example, transcripts typically include markers within the transcript that identify the source of the transcribed content. For example, logging the speaker of the transcribed content can be referred to as speaker logging to answer "who said what" and "who spoke when". When the identity of the content source is sensitive or the source that generates the content (e.g., a participant) prefers to mask his / her identity for any reason (e.g., personal reasons), the source does not want the label to be associated with the transcribed content. Note that here, different from content privacy, the source does not mind the content being revealed in the transcript, but rather does not want the identifier (e.g., the label) to associate the content with the source. Since traditional transcription devices lack the ability to accommodate these privacy issues, participants may choose not to use transcription devices even if the above benefits are foregone. To maintain these benefits and / or protect the privacy of the participants, the environment can include a privacy-aware transcription device called a transcriptor. In another example, when a camera is capturing video of a speaker who wants to remain anonymous, the speaker can choose not to have the images (e.g., faces) they record be remembered. This can include distorting the video / image frames of the speaker's face and / or overlaying graphics that mask the identity of the speaker, such that other individuals in the meeting cannot visually identify the speaker. Additionally or alternatively, the audio of the speaker's voice may be distorted (e.g., by passing the audio through a vocoder) to mask the speaker's voice in a way that anonymizes the speaker.

[0030] In some embodiments, privacy issues are further enhanced by handling privacy on the device during transcription such that the transcript does not leave the scope of the assembly environment (e.g., a conference room or a classroom) that provides a shared space for its participants. In other words, by using a transcriber to generate a transcript on the device, speaker labels that identify speakers who want to remain anonymous can be removed on the device to alleviate any concerns that the identities of these speakers will be exposed / compromised if the transcript is processed on a remote system (e.g., a cloud environment). In other words, there are no unedited transcripts generated by the transcriber that could be shared or stored to endanger the privacy of the participants.

[0031] Another technical effect of performing audio-visual transcription (e.g., audio-visual automatic speech recognition (AVASR)) on the device is reduced bandwidth requirements because audio and image data (also referred to as video data) can be locally retained on the device without the need to transmit it to a remote cloud server. For example, if video data is to be transmitted to the cloud, it may first need to be compressed for transmission. Thus, another technical effect of performing video matching on the user device itself is that uncompressed (highest quality) video data can be used to perform video data matching. The use of uncompressed video data makes it easier to identify a match between the audio data and the speaker's face, such that speaker labels assigned to the transcribed portions of audio data spoken by speakers who do not want to be identified can be anonymized. At the same time, video data capturing the faces of individuals who do not want to be identified can be enhanced / distorted / blurred to mask these individuals such that they cannot be visually identified if the video recording is shared. Similarly, audio data representing the words spoken by these individuals may be distorted to anonymize the speech of these individuals who do not want to be identifiable. Reference Figure 1A - 1E, Environment 100 includes multiple participants 10, 10a - j. Here, Environment 100 is a hosting conference room, where six participants 10a - f are participating in a meeting (e.g., a video conference) in the hosting conference room. Environment 100 includes a display device 110 that receives a content feed 112 (also referred to as a multimedia feed, content stream, or feed) from a remote system 130 via a network 120. The content feed 112 can be an audio feed 218 (i.e., audio data 218 such as audio content, audio signal, or audio stream), a visual feed 217 (i.e., image data 217 such as video content, video signal, or video stream), or some combination of both (e.g., also referred to as an audiovisual feed, audiovisual signal, or audiovisual stream). The display device 110 includes a display 111 capable of displaying video content 217 and speakers for audible output of audio content 218, or communicates therewith. Some examples of the display device 110 include a computer, laptop, mobile computing device, television, monitor, smart device (e.g., smart speaker, smart display, smart appliance), wearable device, etc. In some examples, the display device 110 includes audiovisual feeds 112 from other conference rooms participating in the meeting. For example, Figure 1A - 1E depicts two feeds 112, 112a - b, where each feed 112 corresponds to a different remote conference room. Here, the first feed 112a includes three participants 10, 10g - i, while the second feed 112b includes a single participant 10, 10j (e.g., an employee working remotely from a home office). Continuing with the previous example, the first feed 112a can correspond to a feed 112 from the New York office, the second feed 112b corresponds to a feed 112 from the San Francisco office, and the hosting conference room 100 corresponds to the Chicago office.

[0032] The remote system 130 can be a distributed system (e.g., a cloud computing environment or storage abstraction) with scalable / elastic resources 132. The resources 132 include computing resources 134 (e.g., data processing hardware) and / or storage resources 136 (e.g., memory hardware). In some embodiments, the remote system 130 (e.g., on computing resources 132) hosts software that coordinates Environment 100. For example, the computing resources 132 of the remote system 130 execute software such as a real - time communication application or a professional conferencing platform.

[0033] Continuing to refer to Figure 1A - 1E, the environment 100 further includes a transcriber 200. The transcriber 200 is configured to generate a transcript 202 of the content occurring within the environment 100. The content can be from the location where the transcriber 200 is located (e.g., participant 10 in the meeting room 100 having the transcriber 200) and / or from a content feed 112 that transmits the content to the location of the transcriber 200. In some examples, the display device 110 transmits one or more content feeds 112 to the transcriber 200. For example, the display device 110 includes speakers that output the audio content 218 of the content feed 112 to the transcriber 200. In some embodiments, the transcriber 200 is configured to receive the same content feed 112 as the display device 110. In other words, the display device 110 can be used as an extension of the transcriber 200 by receiving both the audio and video feeds of the content feed 112. For example, the display device 110 can include hardware 210, such as data processing hardware 212 and memory hardware 214 that communicates with the data processing hardware 212, which causes the data processing hardware 212 to execute the transcriber 200. In this relationship, the transcriber 200 can receive the content feed 112 (e.g., audio and visual content / signals 218, 217) via a network connection, rather than only audibly capturing the audio content / signals 218 relayed through the peripheral devices (such as speakers) of the display device 110. In some examples, this connection between the transcriber 200 and the display device 110 enables the transcriber 200 to seamlessly display the transcript 202 on the display / screen 111 of the display device 110 within the local environment 100 (e.g., the hosting meeting room). In other configurations, the transcriber 200 is located in the same local environment 110 as the display device 110, but corresponds to a computing device separate from the display device 110. In these configurations, the transcriber 200 communicates with the display device 110 via a wired or wireless connection. For example, the transcriber 200 has one or more ports that allow for a wired / wireless connection, such that the display device 110 serves as a peripheral device of the transcriber 200. Additionally or alternatively, the application forming the environment 100 can be compatible with the transcriber 200. For example, the transcriber 200 is configured as an input / output (I / O) device within the application, such that the audio and / or visual signals coordinated by the application are diverted to the transcriber 200 (e.g., in addition to the display device 110).

[0034] In some examples, the transcriber 200 (and optionally the display device 110) is portable such that the transcriber 200 can be transferred between conference rooms. In some embodiments, the transcriber 200 is configured with processing capabilities (e.g., processing hardware / software) to process audio and video content 112 and generate a transcript 202 when the content 112 is presented in the environment 100. In other words, the transcriber 200 is configured to locally process the content 112 (e.g., audio and / or visual content 218, 217) at the transcriber 200 to generate the transcript 202 without any additional remote processing (e.g., at the remote system 130). Here, this type of processing is referred to as on-device processing. Different from remote processing that often uses low-fidelity compressed video for server-based applications due to bandwidth constraints, on-device processing can be free from bandwidth constraints, thus allowing the transcriber 200 to utilize more accurate high-definition video with high fidelity when processing visual content. Additionally, this on-device processing can allow real-time tracking of the speaker's identity without latency caused by the waiting time that would occur if the audio and / or visual signals 218, 217 were remotely processed (e.g., in a remote computing system 130 connected to the transcriber 200). To process the content at the transcriber 200, the transcriber 200 includes hardware 210, such as data processing hardware 212 and memory hardware 214 that communicates with the data processing hardware 212. Some examples of the data processing hardware 212 include a central processing unit (CPU), a graphics processing unit (GPU), or a tensor processing unit (TPU).

[0035] In some embodiments, the transcriber 200 is executed on the remote system 130 by receiving the content 112 (audio and video data 217, 218) from each of the first and second feeds 112a-b and receiving the feed 112 from the conference room environment 100. For example, the data processing hardware 134 of the remote system 130 can execute instructions stored on the memory hardware 136 of the remote system 130 for executing the transcriber 200. Here, the transcriber 200 can process the audio data 218 and the image data 217 to generate the transcript 202. For example, the transcriber 200 can generate the transcript 202 and transmit the transcript 202 to the display device 110 via the network 120 for display thereon. The transcriber 200 can similarly transmit the transcript 202 to the computing device / display device associated with the participants 10g-i corresponding to the first feed and / or the participant 10j corresponding to the second feed 10j.

[0036] In addition to processing hardware 210, the transcriber 200 also includes peripherals 216. For example, to process audio content, the transcriber 200 includes audio capture devices 216, 216a (e.g., microphones) that capture sounds (e.g., spoken utterances) about the transcriber 200 and convert the sounds into audio signals 218( Figure 2A and 2B )(or audio data 218). The transcriber 200 can then use the audio signal 218 to generate a transcript 202.

[0037] In some examples, the transcriber 200 also includes image capture devices 216, 216b as peripherals 216. Here, the image capture device 216b (e.g., one or more cameras) can capture image data 217( Figure 2A and Figure 2B ) as an additional input source (e.g., visual input) that is combined with the audio signal 218 to help identify which participant 10 in the multi - participant environment 100 is speaking (i.e., the speaker). In other words, by including both the audio capture device 216a and the image capture device 216b, the transcriber 200 can increase the accuracy of its speaker identification because the transcriber 200 can process the image data 217 captured by the image capture device 216b to identify visual features (e.g., facial features) that indicate which participant 10 among the multiple participants 10a - 10j is speaking (i.e., generating an utterance 12) in a particular instance. In some configurations, the image capture device 216b is configured to capture 360 degrees around the transcriber 200 to capture a panoramic view of the environment 100. For example, the image capture device 216b includes an array of cameras configured to capture a 360 - degree view.

[0038] Additionally or alternatively, using the image data 217 can improve the transcript 202 when a participant 10 has a speech impediment. For example, the transcriber 200 may have difficulty generating a transcript for a speaker with a speech impediment that causes the speaker to have problems with clearly enunciated speech. To overcome the inaccuracies of the transcript 202 caused by such enunciation problems, it can be made (e.g., Figure 2A and Figure 2BThe transcriber 200 (at the automatic speech recognition (ASR) module 230) becomes aware of an articulation problem during the generation of the transcript 202. By being aware of the problem, the transcriber 200 can adapt to the problem by utilizing image data 217 representing the face of the participant 10 while speaking to generate an improved or otherwise more accurate transcript 202, rather than the transcript 202 being based solely on the audio data 218 of the participant 10. Here, certain speech impairments may be apparent in the image data 217 from the image capture device 216b. For example, in the case of speech dysarthria, a neuromuscular disorder causing lip movements that affect articulation can be identified in the image 217. Additionally, techniques can be employed in which the image data 217 can be analyzed to correlate the lip movements of the participant 10 with a particular speech disorder with the speech intended by these participants 10, thereby improving automatic speech recognition in a way that would not be possible by using only the audio data 218. In some embodiments, by using the image 217 as an input to the transcriber 200, the transcriber 200 identifies potential articulation problems and takes the problems into account to improve the generation of the transcript 202 during ASR.

[0039] In some embodiments, such as Figure 1B - 1E , the transcriber 200 is privacy-aware such that the participant 10 can choose (e.g., in the transcript 202 or the visual feeds 112, 217) not to share any of his or her speech and / or image information. Here, one or more participants 10 transmit privacy requests 14 that indicate the privacy conditions of the participant 10 during participation in the video conferencing environment 100. In some examples, the privacy request 14 corresponds to a configuration setting of the transcriber 200. The privacy request 14 can occur before, during, or at the start of a meeting or a communication session with the transcriber 200. In some configurations, the transcriber 200 includes a profile (e.g., Figure 5 of the one or more privacy requests 14 of the participant 10 (e.g., Figure 5The profile 500 shown). Here, the profile 500 can be stored on the device (e.g., in the memory hardware 214) or outside the device (e.g., in the remote storage resource 136) and accessed by the transcriber 200. The profile 500 can be configured prior to a communication session and can include an image of the face of the corresponding participant 10 (e.g., image data 217), so that the participant 10 can be correlated with the corresponding portion of the received video content 217. That is, when the video content 217 of the participant 10 in the content feed 112 matches the face image associated with the personal profile 510, the personal profile 510 of the corresponding participant 10 can be accessed. Using the personal profile 510, the privacy settings of the participant can be applied during each communication session in which the participant 10 participates. In these examples, the transcriber 200 can identify the participant 10 (e.g., based on the image data 217 received at the transcriber 200) and apply the appropriate settings for the participant 10. For example, the profile 500 can include personal profiles 510, 510b for specific participants 10, 10b, which indicate that the specific participant 10b does not mind being seen (i.e., being included in the visual feed 217), but does not want to be heard (i.e., not included in the audio feed 218) nor have his / her words 12 transcribed (i.e., not included in the voice in the transcript 202), while another personal profile 510, 510c of another participant 10, 10c may not want to be seen (i.e., not included in the visual feed 217), but his / her words can be recorded and / or transcribed (i.e., included in the audio feed 218 and included in the transcript 202).

[0040] Reference Figure 1B , a third participant 10c has submitted a privacy request 14 with privacy conditions (i.e., a privacy request 14 for identity privacy), the privacy conditions indicating that the third participant 10c does not mind being seen or heard, but does not want the transcript 202 to include an identifier 204 for the third participant 10c (e.g., a label of the identity of the speaker) when the third participant 10c is speaking. In other words, the third participant 10c does not want to share or store his or her identity; thus, the third participant 10c selects that the transcript 202 does not include an identifier 204 that reveals his or her identity associated with the third participant 10c. Here, although Figure 1B FIG. shows a transcript 202 with an edited gray box in which there is an identifier 204 of the speaker 3, but the transcriber 200 can also completely remove the identifier 204 or obscure the identifier 204 in some other way that prevents the identity of the speaker associated with the privacy request 14 from being revealed by the transcriber 200. In other words, Figure 1B FIG. shows that the transcriber 200 modifies a portion of the transcript 202 to not include the identity of the speaker (e.g., by removing or obscuring the identifier 204).

[0041] Figure 1C Similar to Figure 1B , except that the third participant 10c who transmits the privacy request 14 requests not to be seen in any visual feeds 112, 217 of the environment 100 (e.g., another form of identity privacy). Here, the requesting participant 10c may not mind being heard, but preferably visually hides his or her visual identity (i.e., does not share or store his or her visual identity in the visual feeds 112, 217). In this case, the transcriber 200 is configured to obscure, distort, or otherwise mask the visual presence of the requesting participant 10c throughout the communication session among the participants 10, 10a - 10j. For example, in any instance, the transcriber 200 determines the location of the requester 10c based on the image data 217 received from one or more content feeds 112, and applies an abstraction 119 to any physical characteristics of the requester transmitted through the transcriber 200 (e.g., blurring). That is, when the image data 217 is displayed on the screen 111 of the display device 110 and on the screens in the remote environments associated with the participants 10g - 10j, the abstraction 119 at least covers the face of the requester 10c such that the requester 10c cannot be visually identified. In some examples, the personal profile 510 of the participant 10 identifies whether the participant 10 wants to be blurred or masked (i.e., distorted) or completely removed (e.g., as Figure 5 shown). Thus, the transcriber 200 is configured to enhance, modify, or remove portions of the video data 217 to hide the visual identity of the participant.

[0042] In contrast, Figure 1DIllustrated is an example of a privacy request 14 from a third participant 10c requesting that the transcriber 200 not track the visual representation of the third participant 10c or the voice information of the third participant 10c. As used herein, "voice information" refers to the audio data 218 corresponding to the utterances 12 spoken by the participant 10c and the transcript 202 identified from the audio data 218 corresponding to the utterances 12 spoken by the participant 10c. In this example, the participant 10c can be heard during the meeting, but the transcriber 200 does not remember the participant 10c either auditorily or visually (e.g., via the video feed 217 or in the transcript 202). The method can protect the privacy of the participant 10c by having no record of any voice information of the participant 10c in the transcript 202 or by having no identifier 204 in the transcript 202 that identifies the participant 10c. For example, the transcriber 200 can completely omit the text portion of the transcript 202 that transcribes the utterances 12 spoken by the participant 10c, or the transcriber 202 can leave these portions of the text, but not apply the identifier 204 that identifies the participant 10c. However, the transcriber 200 can apply some other arbitrary identifier that does not personally identify the participant 10c, but only demarcates these portions of the text in the transcript 202 from the other portions corresponding to the utterances 12 spoken by the other participants 10a, 10b, 10d - 10j. In other words, the participant 10 can request (e.g., via the privacy request 14) that the transcript 202 and any other records generated by the transcriber 200 do not have a record of the participant's participation in the communication session.

[0043] Compared with the identity privacy request 14, Figure 1EDepicts a content privacy request 14. In this example, a third participant 10c transmits a privacy request 14 that the transcriber 200 not include any content from the third participant 10c in the transcript 202. Here, the third participant 10c makes such a privacy request 14 because the third participant 10c will discuss sensitive content (e.g., confidential information) during the meeting. Due to the sensitive nature of the content, the third participant 10c takes the following precaution: the transcriber 200 does not remember the audio content 218 associated with the third participant 10c in the transcript 202. In some embodiments, the transcriber 200 is configured to receive a privacy request 14 that identifies (e.g., by keyword) the type of content that one or more participants 10 do not want to be included in the transcript 202 and determine when that type of content occurs during the communication session in order to exclude it from the transcript 202. In these embodiments, not all of the audio content 218 from a particular participant 10 is excluded from the transcript 202, only content-specific audio is excluded, such that the particular participant can still discuss other types of content and be included in the transcript 202. For example, the third participant 10c transmits a privacy request 14 that requests the transcriber 200 not to transcribe the audio content regarding Mike. In this case, when the third participant 10c discusses Mike, the transcriber 200 does not transcribe that audio content 218, but when the third participant talks about other topics (e.g., the weather), the transcriber 200 does transcribe that audio content 218. The participant 10c can similarly set a time boundary such that the transcriber 200 does not remember any audio content 218 for a period of time (e.g., the next 2 minutes).

[0044] Figure 2A and 2Bis an example of a transcriber 200. The transcriber 200 generally includes a logging module 220 and an ASR module 230 (e.g., an AVASR module). The logging module 220 is configured to receive audio data 218 corresponding to utterances 12 from participant 10 of a communication session (e.g., captured by audio capture device 216a) and image data 217 representing the face of participant 10 of the communication session, segment the audio data 218 into a plurality of segments 222, 222a-n (e.g., fixed-length segments or variable-length segments), and generate a logging result 224 that includes corresponding speaker labels 226 assigned to each segment 222 based on the audio data 218 and the image data 217 using a probability model (probabilistic generative model). In other words, the logging module 220 includes a series of speaker identification tasks with short utterances (e.g., segments 222) and determines whether two segments 222 of a given conversation are spoken by the same participant 10. At the same time, the logging module 220 may perform a face tracking routine to identify which participant 10 is speaking during which segment 222 to further optimize speaker identification. Then, the logging module 220 is configured to repeat the process for all segments 222 of the conversation. Here, the logging result 224 provides timestamped speaker labels 226, 226a-e for the received audio data 218, which not only identify who is speaking during a given segment 222 but also identify when a speaker change occurs between adjacent segments 222. Here, the speaker label 226 may be used as an identifier 204 within the transcript 202.

[0045] In some examples, the transcriber 200 receives a privacy request 14 at the logging module 220. Since the logging module 220 identifies the speaker tag 226 or identifier 204, the logging module 220 can advantageously resolve the privacy request 14 corresponding to the identity-based privacy request 14. In other words, when the participant 10 is the speaker, the logging module 220 receives the privacy request 14 when the privacy request 14 requests that the participant 10 not be identified by an identifier 204 such as the tag 226. When the logging module 220 receives the privacy request 14, the logging module 220 is configured to determine whether the participant 10 corresponding to the request 14 matches the tag 226 generated for a given segment 222. In some examples, an image of the face of the participant 10 can be used to associate the participant 10 with the tag 226 for that participant 10. When the tag 226 for the segment 222 matches the identity of the participant 10 corresponding to the request 14, the logging module 220 can prevent the transcriber 200 from applying the tag 226 or identifier 204 to the corresponding part of the resulting transcription 202 that transcribes the particular segment 222 into text. When the tag 226 for the segment 222 fails to match the identity of the participant 10 corresponding to the request 14, the logging module 220 can allow the transcriber to apply the tag 226 and identifier 204 to the part of the resulting transcript 202 that transcribes the particular segment into text. In some embodiments, when the logging module 220 receives the request 14, the ASR module 230 is configured to wait to transcribe the audio data 218 from the utterance 12. In other embodiments, the ASR module 230 transcribes in real time, and the resulting transcription 202 removes the tag 226 and identifier 204 in real time for any participant 10 who has provided a privacy request 14 to opt out of having their voice information transcribed. Optionally, the logging module 220 can further distort the audio data 218 associated with these participants 10 seeking privacy such that their spoken voice is altered in a way that cannot be used to identify the participants 10.

[0046] The ASR module 230 is configured to receive audio data 218 corresponding to utterance 12 and image data 217 representing the face of participant 10 while uttering utterance 12. Using the image data 217, the ASR module 230 transcribes the audio data 218 into a corresponding ASR result 232. Here, the ASR result 232 refers to the text transcription of the audio data 218 (e.g., transcript 202). In some examples, the ASR module 230 communicates with the logging module 220 to improve voice recognition based on utterance 12 using the logging result 224 associated with the audio data 218. For example, the ASR module 230 may apply different voice recognition models (e.g., language model, prosody model) for different speakers identified from the logging result 224. Additionally or alternatively, the ASR module 230 and / or the logging module 220 (or some other component of the transcriber 200) may use the timestamped speaker labels 226 predicted for each segment 222 obtained from the logging result 224 to index the transcription 232 of the audio data 218. In other words, the ASR module 230 uses the speaker labels 226 from the logging module 220 to generate identifiers 204 of the speakers within the transcript 202. As Figure 1A - 1E shown, the transcript 202 of the communication session within the environment 100 can be indexed by the speaker / participant 10 to associate portions of the transcript 202 with the corresponding speaker / participant 10 in order to identify what each speaker / participant 10 said.

[0047] In some configurations, the ASR module 230 receives a privacy request 14 for the transcriber 200. For example, whenever the privacy request 14 corresponds to a request 14 not to transcribe the voice of a particular participant 10, the ASR module 230 receives the privacy request 14 for the transcriber 200. In other words, whenever the request 14 is not a privacy request 14 based on a label / identifier, the ASR module 230 may receive the privacy request 14. In some examples, when the ASR module 230 receives the privacy request 14, the ASR module 230 first identifies the participant 10 corresponding to the privacy request 14 based on the speaker label 226 determined by the logging module 220. Then, when the ASR module 230 encounters the voice of that participant 10 to be transcribed, the ARS module 230 applies the privacy request 14. For example, when the privacy request 14 requests not to transcribe the voice of that particular participant 10, the ASR module 230 does not transcribe any voice of that participant and waits for the voice of a different participant 10 to appear.

[0048] Reference Figure 2B, in some embodiments, the transcriber 200 includes a detector 240 for performing a face tracking routine. In these embodiments, the transcriber 200 first processes the audio data 218 to generate one or more candidate identities for the speaker. For example, for each segment 222, the logging module 220 may include multiple tags 226, 226a1-3 as candidate identities of the speaker. In other words, the model may be a probability model that outputs multiple tags 226, 226a1-3 for each segment 222, where each tag 226 among the multiple tags 226, 226a1-3 is a potential candidate for identifying the speaker. Here, the detector 240 of the transcriber 200 uses the images 217, 217a-n captured by the image capture device 216b to determine which candidate identity has the best visual features indicating that he or she is the speaker of a particular segment 22. In some configurations, the detector 240 generates a score 242 for each candidate identity, where the score 242 indicates the confidence level that the candidate identity is the speaker based on the association between the audio signal (e.g., audio data 218) and the visual signal (e.g., the captured images 217a-n). Here, the highest score 242 may indicate the greatest likelihood that the candidate identity is the speaker. In Figure 2B , the logging module 220 generates three tags 226a1-3 at a particular segment 222. The detector 240 generates a score 242 (e.g., shown as three scores 2421-3) for each of these tags 226 based on the image 217 at the time the segment 222 appears in the audio data 218. Here, Figure 2B , the highest score 242 is indicated by the bold square around the third tag 226a3 associated with the third score 2423. When the transcriber 200 includes the detector 240, the best candidate identity can be transmitted to the ASR module 230 to form an identifier 204 of the transcript 202.

[0049] Additionally or alternatively, the process may be reversed such that the transcriber 200 first processes the image data 217 to generate one or more candidate identities of the speaker based on the image data 217. Then, for each candidate identity, the detector 240 generates a confidence score 242 that indicates the likelihood that the face corresponding to the candidate identity includes the speaking face for the corresponding segment 222 of the audio data 218. For example, the confidence score 242 for each candidate identity indicates the likelihood that the face corresponding to the candidate identity includes the speaking face during the image data 217 at the time instance corresponding to the segment 222 of the audio data 218. In other words, for each segment 222, the detector 240 may score 242 whether the image data 217 corresponding to the participant 10 has a facial expression similar to or matching the facial expression of the speaking face. Here, the detector 240 selects the identity of the speaker of the corresponding segment of the audio data 218 with the highest confidence score 242 as the candidate identity.

[0050] In some examples, the detector 240 is part of the ASR module 230. Here, the ASR module 230 performs a face tracking routine by implementing an encoder front-end with attention layers configured to receive multiple video tracks 217a-n of the image data 217, where each video track is associated with the face of a corresponding participant. In these examples, the attention layers at the ASR module 230 are configured to determine confidence scores indicating the likelihood that the face of the corresponding person associated with the video face track includes the speaking face of the audio track. Additional concepts and features related to an audiovisual ASR module including an encoder front-end with attention layers for multi-speaker ASR recognition can be found in U.S. Provisional Patent Application No. 62 / 923,096, filed on October 18, 2019, the entire contents of which are incorporated herein by reference.

[0051] In some configurations, the transcriber 200 (e.g., at the ASR module 230) is configured to support a multilingual environment 100. For example, when the transcriber 200 generates a transcript 202, the transcriber 200 is capable of generating the transcript 202 in different languages. This feature can enable the environment 100 to include a remote location having one or more participants 10 who speak a different language than the host location. Additionally, in some cases, the speakers in a meeting can be non-native speakers or speakers for whom the language of the meeting is not their first language. Here, the transcript 202 of the content from the speaker can assist other participants 10 in the meeting in understanding the presented content. Additionally or alternatively, the transcriber 200 can be used to provide feedback to the speaker regarding his or her pronunciation. Here, by combining video and / or audio data, the transcriber 200 can indicate incorrect pronunciation (e.g., allowing the speaker to learn and / or adapt with the help of the transcriber 200). In this way, the transcriber 200 can provide the speaker with a notification that provides feedback regarding his / her pronunciation.

[0052] Figure 3is an exemplary arrangement of operations of a method 300 of transcribing content (e.g., at the data processing hardware 212 of the transcriber 200). At operation 302, the method 300 includes receiving an audiovisual signal 217, 218 that includes audio data 218 and image data 217. The audio data 218 corresponds to voice utterances 12 from a plurality of participants 10, 10a-n in the voice environment 100, and the image data 217 represents the faces of the plurality of participants 10 in the voice environment 100. At operation 304, the method 300 includes receiving a privacy request 14 from a participant 10 among the plurality of participants 10a-n. The privacy request 14 indicates privacy conditions associated with the participant 10 in the voice environment 100. At operation 306, the method 300 segments the audio data 218 into a plurality of segments 222, 222a-n. At operation 308, the method 300 includes performing operations 308, 308a-c on each segment 222 of the audio data 218. At operation 308a, for each segment 222 of the audio data 218, the method 300 includes determining, based on the image data 217, the identity of the speaker of the corresponding segment 222 of the audio data 218 from among the plurality of participants 10a-n. At operation 308b, for each segment 222 of the audio data 218, the method 300 includes determining whether the identity of the speaker of the corresponding segment 222 includes the participant 10 associated with the privacy conditions indicated by the received privacy request 14. At operation 308c, for each segment 222 of the audio data 218, when the identity of the speaker of the corresponding segment 222 includes the participant 10, the method 300 includes applying the privacy conditions to the corresponding segment 222. At operation 310, the method 300 includes processing the plurality of segments 222a-n of the audio data 218 to determine a transcript 202 of the audio data 218.

[0053] In cases where certain embodiments discussed herein may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, and the user's activities and demographic information, relationships between users, etc.), the user is provided with one or more opportunities to control whether information is collected, whether personal information is stored, whether personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed herein collect, store, and / or use user personal information only when explicit authorization to do so is received from the relevant user.

[0054] For example, provide the user with control over whether a program or feature collects user information about that particular user or other users associated with the program or feature. Present one or more options to each user for whom personal information is to be collected to allow control over the information collection associated with that user, to provide permission or authorization regarding whether information is collected and regarding which portions of the information are to be collected. For example, one or more such control options can be provided to the user via a communication network. Additionally, certain data can be processed in one or more ways before it is stored or used such that personally identifiable information is removed. As an example, a user's identity can be processed such that personally identifiable information cannot be determined.

[0055] Figure 4 FIG. 400 is a schematic diagram of an exemplary computing device 400 that can be used to implement the systems and methods described in this document. Computing device 400 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown herein, their connections and relationships, and their functions are illustrative only and are not meant to limit the implementation of the inventions described and / or claimed in this document.

[0056] Computing device 400 includes a processor 410 (e.g., data processing hardware), a memory 420 (e.g., memory hardware), a storage device 430, a high-speed interface / controller 440 connected to the memory 420 and a high-speed expansion port 450, and a low-speed interface / controller 460 connected to a low-speed bus 470 and the storage device 430. Each of the components 410, 420, 430, 440, 450, and 460 is interconnected using various buses and can be mounted on a common motherboard or in other appropriate manners. Processor 410 is capable of processing instructions for execution within computing device 400, including instructions stored in memory 420 or on storage device 430, to display graphical information for a graphical user interface (GUI) on an external input / output device such as a display 480 coupled to the high-speed interface 440. In other embodiments, multiple processors and / or multiple buses can be appropriately used, as well as multiple memories and memory types. Moreover, multiple computing devices 400 can be connected, where each device provides a portion of the necessary operations (e.g., as a server group, blade server group, or multi-processor system).

[0057] The memory 420 stores information non - transiently within the computing device 400. The memory 420 can be a computer - readable medium, a volatile memory unit, or a non - volatile memory unit. The non - transient memory 420 can be a physical device for temporarily or permanently storing programs (e.g., sequences of instructions) or data (e.g., program state information) for use by the computing device 400. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., commonly used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and magnetic disks or tapes.

[0058] The storage device 430 is capable of providing large - capacity storage for the computing device 400. In some embodiments, the storage device 430 is a computer - readable medium. In various different embodiments, the storage device 430 can be a floppy disk device, a hard disk device, an optical disk device, or a magnetic tape device, a flash memory, or other similar solid - state storage devices, or an array of devices, including devices in a storage area network or other configurations. In additional embodiments, a computer program product is tangibly embodied as an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer or machine - readable medium, such as the memory 420, the storage device 430, or the memory on the processor 410.

[0059] The high - speed controller 440 manages bandwidth - intensive operations of the computing device 400, while the low - speed controller 460 manages lower - bandwidth - intensive operations. This division of responsibilities is merely illustrative. In some embodiments, the high - speed controller 440 is coupled to the memory 420, the display 480 (e.g., via a graphics processor or accelerator), and a high - speed expansion port 450 that can accept various expansion cards (not shown). In some embodiments, the low - speed controller 460 is coupled to the storage device 430 and a low - speed expansion port 470. The low - speed expansion port 470, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a network device, such as a switch or a router, for example, via a network adapter.

[0060] As shown, the computing device 400 can be implemented in many different forms. For example, the computing device 400 can be implemented as a standard server 400a or implemented multiple times in a group of such servers 400a, implemented as a laptop computer 400b, or implemented as part of a rack - server system 400c.

[0061] The various embodiments of the systems and techniques described herein can be implemented in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include an implementation in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be special purpose or general purpose, coupled to receive data and instructions from, and to send data and instructions to, a storage system, at least one input device, and at least one output device.

[0062] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0063] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special purpose logic circuitry, such as an FPGA (field programmable gate array) or ASIC (application specific integrated circuit). For example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing the instructions and one or more storage devices for storing the instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto - optical disks or optical disks, or be operatively coupled to the mass storage device to receive data therefrom or transfer data thereto, or both. However, a computer need not have such devices. Computer - readable media suitable for storing computer program instructions and data include all forms of non - volatile memory, media and memory devices, including semiconductor memory devices, such as EPROM, EEPROM and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto - optical disks; and CD ROM and DVD - ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0064] To provide for interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device and optionally a keyboard and a pointing device, the display device such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor or touch screen for displaying information to the user, the pointing device such as a mouse and a trackball by which the user can provide input to the computer. Other types of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback or tactile feedback; and input from the user can be received in any form, including voice, speech or tactile input. Additionally, a computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a web page to a web browser on a user client device in response to a request received from the web browser.

[0065] Numerous embodiments have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are within the scope of the appended claims.

Claims

1. A computer-implemented method performed by data processing hardware, the method causing the data processing hardware to perform operations that include: Receiving an audio signal comprising audio data, the audio data comprising a plurality of speech utterances in a speech environment; Receiving a privacy request indicating a privacy condition, the privacy condition including a content-specific condition indicating a type of content to be excluded from a transcript; Processing the audio data to identify one or more speech utterances corresponding to the type of content among the plurality of speech utterances; and Generating the transcript based on the audio data, the transcript excluding the one or more speech utterances corresponding to the type of content among the plurality of speech utterances.

2. The method according to claim 1, wherein, the data processing hardware resides on a device local to a user associated with the audio data.

3. The method according to claim 2, wherein, processing the audio data is performed locally on the device.

4. The method according to claim 1, wherein, the type of content includes content corresponding to a time duration.

5. The method according to claim 1, wherein, the type of content includes content associated with a specific person.

6. The method according to claim 1, wherein, the operations further include, for each speech utterance of the plurality of speech utterances of the audio data, associating the corresponding speech utterance with one of a first user or a second user.

7. The method according to claim 6, wherein, the privacy request applies only to each corresponding speech utterance of the plurality of speech utterances associated with the first user.

8. The method according to claim 6, wherein, the privacy request applies to: each corresponding speech utterance of the plurality of speech utterances associated with the first user; and each corresponding speech utterance of the plurality of speech utterances associated with the second user.

9. The method according to claim 1, wherein, the audio signal is received as part of an audiovisual signal, the audiovisual signal further including image data representing the faces of a plurality of users in the speech environment.

10. The method according to claim 9, wherein, the image data includes high-definition video processed by the data processing hardware.

11. A system, the system comprising: data processing hardware; and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that, when executed on the hardware processor, cause the hardware processor to perform operations including: Receiving an audio signal comprising audio data, the audio data comprising a plurality of speech utterances in a speech environment; Receiving a privacy request indicating a privacy condition, the privacy condition including a content-specific condition indicating a type of content to be excluded from a transcript; Processing the audio data to identify one or more speech utterances corresponding to the type of content among the plurality of speech utterances; and Generate the transcript based on the audio data, the transcript excluding the one or more voice utterances corresponding to the content type among the plurality of voice utterances.

12. The system according to claim 11, wherein, the data processing hardware resides on a device local to the user associated with the audio data.

13. The system according to claim 12, wherein, processing the audio data is performed locally on the device.

14. The system according to claim 11, wherein, the content type includes content corresponding to a time duration.

15. The system according to claim 11, wherein, the content type includes content associated with a specific person.

16. The system according to claim 11, wherein, the operation further includes, for each voice utterance among the plurality of voice utterances of the audio data, associating the corresponding voice utterance with one of a first user or a second user.

17. The system according to claim 16, wherein, the privacy request applies only to each corresponding voice utterance among the plurality of voice utterances associated with the first user.

18. The system according to claim 16, wherein, the privacy request applies to: each corresponding voice utterance among the plurality of voice utterances associated with the first user; and each corresponding voice utterance among the plurality of voice utterances associated with the second user.

19. The system according to claim 11, wherein, the audio signal is received as part of an audiovisual signal, the audiovisual signal further including image data representing the faces of a plurality of users in the voice environment.

20. The system according to claim 19, wherein, the image data includes high-definition video processed by the data processing hardware.