Transcription in a conference room with privacy considerations from an audio-visual stream

The privacy-aware transcription system addresses privacy concerns by processing audio and video data locally, enhancing speaker identification accuracy and reducing bandwidth, while ensuring secure and effective transcription.

JP7713504B2Active Publication Date: 2025-07-25GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023206380
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-07-25
Estimated Expiration
2039-11-18

AI Technical Summary

Technical Problem

Existing transcription systems fail to address privacy concerns of conference participants, particularly in identifying and anonymizing speakers, leading to potential disclosure of sensitive information and reduced accuracy in speaker identification.

Method used

A privacy-aware transcription system that processes audio and video data locally on-device, allowing participants to set privacy settings, identifies speakers using both audio and video data, and applies privacy conditions such as anonymization or exclusion of sensitive content, ensuring accurate and reliable transcripts while maintaining participant privacy.

Benefits of technology

Enhances speaker identification accuracy and reduces bandwidth requirements by using high-resolution video, while ensuring that sensitive information is not disclosed, thus providing a more secure and effective transcription solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713504000001
    Figure 0007713504000001
  • Figure 0007713504000002
    Figure 0007713504000002
  • Figure 0007713504000003
    Figure 0007713504000003
Patent Text Reader

Abstract

To provide a privacy-friendly transcription method and system.SOLUTION: A method receives an audiovisual signal including audio and image data related to a conversational environment and a privacy request from a participant in the conversational environment. The privacy request indicates the participant's privacy condition. The method also divides the audio data into a plurality of segments, and for each segment, determines the identity of a speaker of the corresponding segment of the audio data on the basis of the image data, determines whether speaker identification information of the corresponding segment includes a participant associated with the privacy condition, applies the privacy condition to the corresponding segment when the identity of the speaker of the corresponding segment includes the participant, and processes multiple segments of audio data to determine a transcript for the audio data.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to transcription in a conference room that takes into account privacy from an audio-visual stream.

Background Art

[0002] Speaker diarization is a process of dividing an input audio stream into segments of the same kind according to the identity of the speaker. In an environment with multiple speakers, speaker diarization answers the question of "who is speaking when", and has various applications such as retrieval of multimedia information, speaker turn-taking analysis, and audio processing. In particular, a speaker diarization system can generate speaker boundaries that have the potential to significantly improve the accuracy of acoustic speech recognition.

Summary of the Invention

[0003] One aspect of the present disclosure provides a method for generating a privacy - aware transcript in a conference room from a content stream. The method includes, in data - processing hardware, receiving an audiovisual signal including audio data and image data. The audio data corresponds to audio utterances from a plurality of participants in a conversation environment, and the image data represents the faces of the plurality of participants in the conversation environment. The method also includes, in data - processing hardware, receiving a privacy request from one of the plurality of participants. The privacy request indicates privacy conditions related to the participants in the conversation environment. The method further includes the data - processing hardware dividing the audio data into a plurality of segments. For each segment of the audio data, the method includes the data - processing hardware determining, from among the plurality of participants, identification information of the speaker of the corresponding segment of the audio data based on the image data. For each segment of the audio data, the method also includes the data - processing hardware determining whether the identification information of the speaker of the corresponding segment includes a participant related to the privacy conditions indicated by the received privacy request. If the identification information of the speaker of the corresponding segment includes a participant, the method includes applying the privacy conditions to the corresponding segment. The method further includes the data - processing hardware processing the plurality of segments of the audio data to determine a transcript for the audio data.

[0004] Embodiments of the present disclosure may include one or more of any of the following features. In some embodiments, applying the privacy conditions to the corresponding segment includes, after determining the transcript, deleting the corresponding segment of the audio data. Additionally or alternatively, applying the privacy conditions to the corresponding segment may include enhancing the corresponding segment of the image data to visually hide the identification information of the speaker of the corresponding segment of the audio data.

[0005] In some examples, for each part of a transcript corresponding to one of a plurality of segments of voice data to which privacy conditions are applied, processing the plurality of segments of voice data to determine a transcript for the voice data includes modifying the corresponding part of the transcript so as not to include speaker identification information. Optionally, for each segment of voice data to which privacy conditions are applied, processing the plurality of segments of voice data to determine a transcript for the voice data may include omitting to transcribe the corresponding segment of voice data. The privacy conditions include content-specific conditions, and the content-specific conditions indicate the type of content to be excluded from the transcript.

[0006] In some configurations, determining the speaker identification information of the corresponding segment of voice data from among a plurality of participants includes determining a plurality of candidate identification information of the speaker based on image data. Here, for each candidate identification information of the plurality of candidate identification information, a confidence score indicating the possibility that the face corresponding to the candidate identification information based on the image data includes the face speaking in the corresponding segment of voice data is generated. In this configuration, the method includes selecting, as the candidate identification information among the plurality of candidate identification information related to the highest confidence score, the speaker identification information of the corresponding segment of voice data.

[0007] In some embodiments, the data processing hardware is present on a device close to at least one of the plurality of participants. The image data may include high-resolution video processed by the data processing hardware. Processing the plurality of segments of voice data to determine a transcript for the voice data may include processing the image data to determine the transcript.

[0008] Another aspect of the present disclosure provides a system for privacy-conscious transcription. The system includes data processing hardware and memory hardware that communicates with the data processing hardware. The memory hardware stores instructions that cause the data processing hardware to execute operations when executed by the data processing hardware. The operations include receiving an audiovisual signal that includes audio data and image data. The audio data corresponds to audio utterances from a plurality of participants in a conversation environment, and the image data represents the faces of the plurality of participants in the conversation environment. The operations include receiving a privacy request from one of the plurality of participants, the privacy request indicating privacy conditions related to the participants in the conversation environment. The method further includes dividing the audio data into a plurality of segments. For each segment of the audio data, the operations include determining, from among the plurality of participants, identification information of the speaker of the corresponding segment of the audio data based on the image data. For each segment of the audio data, the operations also include determining whether the identification information of the speaker of the corresponding segment includes a participant related to the privacy conditions indicated by the received privacy request. If the identification information of the speaker of the corresponding segment includes a participant, the operations include applying the privacy conditions to the corresponding segment. The operations further include processing the plurality of segments of the audio data to determine a transcript for the audio data.

[0009] This aspect may include one or more of any of the following features. In some examples, applying the privacy conditions to the corresponding segment includes deleting the corresponding segment of the audio data after determining the transcript. Optionally, applying the privacy conditions to the corresponding segment may include enhancing the corresponding segment of the image data to visually obscure the identification information of the speaker of the corresponding segment of the audio data.

[0010] In some configurations, in order to determine a transcript for voice data, processing multiple segments of the voice data includes modifying corresponding portions of the transcript so that they do not include speaker identification information for each portion of the transcript corresponding to one of the multiple segments of voice data to which privacy conditions are applied. Additionally or alternatively, in order to determine a transcript for voice data, processing multiple segments of the voice data may include omitting the step of transcribing the corresponding segments of the voice data for each segment of voice data to which privacy conditions are applied. The privacy conditions include content-specific conditions, and the content-specific conditions indicate the types of content to be excluded from the transcript.

[0011] In some embodiments, the operation of determining the speaker identification information for the corresponding segment of voice data from among multiple participants includes determining multiple candidate identification information for the speaker based on image data. This embodiment includes generating a confidence score indicating the likelihood that the face corresponding to the candidate identification information based on the image data includes the face speaking in the corresponding segment of the voice data for each candidate identification information of the multiple candidate identification information. This embodiment also includes selecting, as the candidate identification information among the multiple candidate identification information associated with the highest confidence score, the speaker identification information for the corresponding segment of the voice data.

[0012] In some examples, the data processing hardware is present on a device that is close to at least one of the multiple participants. The image data may include high-resolution video processed by the data processing hardware. Processing multiple segments of the voice data to determine a transcript for the voice data may include processing the image data to determine the transcript.

[0013] Details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the following detailed description. Other aspects, features, and advantages will be apparent from the detailed description and drawings, as well as from the claims.

Brief Description of the Drawings

[0014]

Figure 1A

Figure 1B

Figure 1C

Figure 1D

Figure 1E

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Modes for Carrying Out the Invention

[0015] Like reference symbols in the various drawings indicate like components. The privacy of data used and generated by a video conferencing system is an important aspect of such a system. Conference participants may have personal views on the privacy of the audio and video data obtained during the conference. Therefore, there are technical problems regarding how to provide a video conferencing system that can accurately generate a transcript of a video conference while meeting such privacy requirements in a reliable and accurate manner. Embodiments of the present disclosure enable conference participants to set their own privacy settings (e.g., opt-in or opt-out of various functions of the video conferencing system), and the video conferencing system then accurately and effectively implements the participants' desires. When generating a transcript, the system identifies contributions from conversations by participants based not only on the audio captured during the conference but also on the video captured during the conference, which provides a technical solution that ensures higher accuracy in identifying contributors to the video conference, thereby improving the accuracy of the transcript while enabling accurate and reliable implementation of the privacy requirements of specific participant instructions. In other words, a more accurate, reliable, and flexible video conferencing system is provided.

[0016] Furthermore, in some embodiments, the process of generating a transcript of a video conference is performed locally, for example, by a device in the same room as one or more of the participants of the video conference. In other words, in such embodiments, the process of generating the transcript is not performed remotely via one or more remote servers / cloud servers, etc. This helps to meet specific privacy desires while enabling the use of fully / original resolution and fully / original quality video data captured locally when identifying speakers during the video conference (as opposed to remote servers that operate with low-resolution and / or low-quality video that may compromise the accuracy of speaker identification).

[0017] In an assembly environment (commonly also referred to as an environment), people gather, think, communicate about ideas, schedules, or other concerns. The assembly environment functions as a shared space for participants. This shared space can be a physical space such as a conference room or classroom, a virtual space (e.g., a virtual conference room), or any combination thereof. The environment can be a centralized location (e.g., hosted locally) or a decentralized location (e.g., hosted virtually). For example, the environment can be a single room where participants gather, such as a conference room or classroom. In some embodiments, the environment is a plurality of shared spaces linked together to form an assembly of participants. For example, for a meeting, there is a host location (e.g., the location where the meeting coordinator or presenter is located) and one or more remote locations (e.g., locations where real-time communication applications are used) participating in the meeting. In other words, a company hosts a meeting from an office in Chicago, but other offices of the company (e.g., in San Francisco or New York) participate in the meeting remotely. For example, there are many companies that hold large-scale meetings across multiple offices, and each office has a meeting space for participating in the meeting. This is especially true as it has become common for team members to be dispersed throughout the company (i.e., in multiple locations) and work remotely. Further, as applications become more robust for real-time communication, environments can be hosted for remote offices, remote employees, remote partners (such as business partners), remote customers, etc. Therefore, the environment has been developed to accommodate a wide range of assembly logistics.

[0018] Generally, as a communication space, the environment hosts multiple participants. Here, each participant can provide audio content (e.g., audible utterances by speaking) and / or video content (e.g., the actions of the participant) while present within the environment. When there are multiple participants in the environment, there are advantages to tracking and / or recording the participation of some or all of the participants. This is particularly true when the environment is equipped with a wide range of meeting equipment. For example, when a Chicago office hosts a meeting in which both a New York office and a San Francisco office participate remotely, it may be difficult for someone in the Chicago office to identify a speaker who is at one of the remote locations. For instance, the Chicago office may include video feeds that capture the conference rooms of each office that is remote from the Chicago office. Even using the video feeds, participants in the Chicago office may not be able to distinguish all of the participants in the New York office. For example, since the speaker in the New York office is located away from the camera associated with the video feed, it is difficult for participants in the Chicago office to identify who the speaker in the New York office is. This can also be difficult when the Chicago-based participant does not know the other participants in the meeting well (e.g., when the speaker cannot be identified by voice). When the speaker cannot be identified, problems can occur because the identification information of the speaker may be an important factor during the meeting. In other words, it may be important to identify the speaker (or the originator of the content) in order to understand the harvest / deliverable or generally who shared which content. For example, even if Sally in the New York office accepts an action item deliverable from Johnny in the Chicago office, if Johnny cannot identify that Sally has accepted the action item, Johnny may have trouble following up on the action item later. In another scenario, since Johnny cannot identify that Sally has accepted the action item, Johnny may incorrectly identify that Tracy (e.g., also in the New York office) has accepted the action item.The same can be said at the basic level of just having participants talk to each other. If Sally talked about a particular topic and Johnny thought it was something Tracey had said, Johnny might cause confusion when talking to Tracey about that topic in a later meeting.

[0019] Another problem can occur when a speaker talks about names, acronyms, and / or jargon that other participants are not familiar with and / or cannot fully understand. In other words, Johnny might discuss an issue that occurred with the carrier used at the time of shipment. Pete might interrupt Johnny's problem by saying, "Oh, you should talk to Teddy in Facilities about that." If Johnny doesn't know Teddy and / or the Facilities team well, Johnny might make a note to talk to Freddy instead of Teddy. This can also happen with acronyms or other jargon used in a particular industry. For example, if the Chicago office has a meeting with a company in Seattle, and the Chicago office hosts the meeting and the Seattle company participates remotely, the participants from the Chicago office might use acronyms or terms that the Seattle company is not familiar with. Without a record or transcription of what was presented by the Chicago office, the Seattle company may, unfortunately, have less understanding of the meeting (e.g., the quality of the meeting may decline). Additionally or alternatively, if the connection between locations or to the meeting hosting platform is insufficient, participants' problems can become complex when they are trying to understand the content during the meeting.

[0020] To address these issues, a transcription device exists (e.g., in real-time) within the environment that generates a transcript of the content that occurs within the environment. When generating the transcript, the device can identify the speaker (i.e., the participant generating the audio content) and / or associate the content with the participants present within the environment. Using the transcript of the content presented within the environment, the transcription device can record who created what content that can be harvested and / or the deliverables and is accessible for participants to reference. For example, a participant can reference the transcript during (e.g., in real-time or substantially real-time) or at some point after a meeting. In other words, Johnny can recognize, by referring to the display of the transcript generated by the transcription device, that the person who needs to speak at the facility is Teddy (not Freddy) and that he needs to follow up with Sally (not Tracy) regarding that action item.

[0021] Unfortunately, while transcripts may solve some of the problems encountered in the environment, they present privacy issues. Here, privacy means that the transcripts generated by the transcription device are not observed. There are various types of privacy, and some examples include content privacy or identifying information privacy. Here, content privacy is content-based where it is desired that certain sensitive content (e.g., confidential content) is not stored in a written or human-readable form. For example, a meeting may include audio content about another employee who is not present at the meeting (e.g., managers discussing a personnel issue that has arisen). In this example, the meeting participants would likely want this part of the meeting about the other employee not to be transcribed or otherwise stored. This could also include not storing audio content that includes information about the other employee. Here, since conventional transcription devices transcribe content indiscriminately, in a meeting, the conventional transcription device cannot be used, at least during that part of the meeting.

[0022] Identity information privacy refers to the privacy that attempts to maintain the anonymity of the content sender. For example, transcripts often contain labels within the transcript that identify the sender of the transcribed content. For example, labeling the speaker of the transcribed content can be referred to as speaker diarization to answer both "who said what" and "who said it when". When the identity information of the content sender is confidential, or when the sender (e.g., participant) who generates the content wishes to mask their identity information for some reason (e.g., personal reasons), the sender does not want a label to be associated with the transcribed content. Here, different from content privacy, it should be noted that the sender does not care that the content is made public in the transcript, but does not want an identifier (e.g., label) to associate the content with the sender. Since conventional transcription devices do not have functions to address these privacy concerns, participants may choose not to use the transcription device even if they forgo the aforementioned advantages. To maintain these advantages and / or protect the privacy of participants, the environment may include a privacy-conscious transcription device called a transcriptor. In an additional example, when a camera is capturing the video of a speaker who desires anonymity, the speaker can choose not to store the recorded image (e.g., face). This can include distorting the video frame / image frame of the speaker's face or overlaying graphics that hide the speaker's identity information so that other individuals participating in the meeting cannot visually identify the speaker. Additionally or alternatively, the voice of the speaker can be masked in a way that anonymizes the speaker (e.g., by passing the voice through a vocoder) to distort the voice of the speaker.

[0023] In some embodiments, by processing privacy on-device during transcription, concerns about privacy are further enhanced so that the transcript does not leave the scope of an assembly environment (e.g., a conference room or a classroom) that provides a shared space for participants. In other words, by using a transcriptor to generate a transcript on-device, speaker labels that identify speakers who desire anonymity are removed on-device, reducing the concern that the identification information of these speakers will be disclosed / leaked when the transcript is processed in a remote system (e.g., a cloud environment). Put another way, there is no unedited transcript generated by a transcriptor that could be shared or stored and put the privacy of participants at risk.

[0024] Another technical impact of performing audio-video transcription (e.g., audio-video automated speech recognition (AVASR)) on-device is that bandwidth requirements are reduced because audio data and image data (also referred to as video data) can be locally held on-device without the need to send it to a remote cloud server. For example, when sending video data to the cloud, it may first need to be compressed for transmission. Thus, another technical effect of performing video matching on the user device itself is that video matching can be performed using uncompressed (highest quality) video data. By using uncompressed video data, it becomes easier to recognize the match between the audio data and the speaker's face, and it is possible to anonymize the speaker label assigned to the transcribed portion of the audio data spoken by a speaker who does not wish to be identified. Similarly, video data that captures the face of an individual who does not wish to be identified can be enhanced / distorted / blurred to mask these individuals so that they cannot be visually identified even if the video recording is shared. Similarly, the audio data representing the utterances spoken by these individuals can be distorted to anonymize the voices of these individuals who do not wish to be made identifiable. Referring to FIGS. 1A-1E, environment 100 includes a plurality of participants 10, 10a-j. Here, environment 100 is a host conference room, and in the host conference room, six participants 10a-f are participating in a meeting (e.g., a video conference). Environment 100 includes a display device 110 that receives a content feed 112 (also referred to as a multimedia feed, content stream, or feed) from remote system 130 via network 120. Content feed 112 can be an audio feed 218 (i.e., audio data 218 such as audio content, audio signal, or audio stream), a video feed 217 (i.e., image data 217 such as video content, video signal, or video stream), or a combination of both (e.g., also referred to as an audio-video feed, audio-video signal, or audio-video stream).The display device 110 includes, or communicates with, a display 111 capable of displaying video content 217 and a speaker for audio output of audio content 218. Some examples of the display device 110 include a computer, a laptop, a mobile computing device, a television, a monitor, a smart device (e.g., a smart speaker, a smart display, a smart home appliance), a wearable device, etc. In some examples, the display device 110 includes an audiovisual feed 112 of other conference rooms participating in the meeting. For example, FIGS. 1A-1E show two feeds 112, 112a-b. Here, each feed 112 corresponds to a separate remote conference room. Here, the first feed 112a includes three participants 10, 10g-i, and the second feed 112b includes one participant 10, 10j (e.g., an employee working remotely from a home office). Continuing with the previous example, the first feed 112a corresponds to the feed 112 from the New York office, the second feed 112b corresponds to the feed 112 from the San Francisco office, and the host conference room 100 corresponds to the Chicago office.

[0025] The remote system 130 can be a distributed system (e.g., a cloud computing environment or storage abstraction) having scalable / elastic resources 132. The resources 132 include computing resources 134 (e.g., data processing hardware) and / or storage resources 136 (e.g., memory hardware). In some embodiments, the remote system 130 hosts software that adjusts the environment 100 (e.g., on the computing resources 132). For example, the computing resources 132 of the remote system 130 execute software such as a real-time communication application or a specialized conferencing platform.

[0026] Continuing to refer to FIGS. 1A-1E, environment 100 also includes a transcriptor 200. The transcriptor 200 is configured to generate a transcript 202 of the content occurring within environment 100. This content can be from where the transcriptor 200 is located (e.g., participant 10 within conference room 100 equipped with transcriptor 200), and / or from a content feed 112 that transmits the content to the location of the transcriptor 200. In some examples, display device 110 transmits one or more content feeds 112 to transcriptor 200. For example, display device 110 includes a speaker that outputs the audio content 218 of content feed 112 to transcriptor 200. In some embodiments, transcriptor 200 is configured to receive the same content feed 112 as display device 110. In other words, display device 110 can function as an extension of transcriptor 200 by receiving the audio feed and video feed of content feed 112. For example, display device 110 can include hardware 210 such as data processing hardware 212, and memory hardware 214 that communicates with data processing hardware 212 to cause data processing hardware 212 to execute transcriptor 200. In this relationship, transcriptor 200 not only audibly captures the audio content / audio signal 218 relayed through a peripheral device such as a speaker of display device 110, but can also receive content feed 112 (e.g., audio and video content / audio and video signals 218, 217) via a network connection. In some examples, due to this connectivity between transcriptor 200 and display device 110, transcriptor 200 can seamlessly display transcript 202 on the display / screen 111 of display device 110 locally within environment 100 (e.g., the host conference room).In other configurations, the transcriber 200 is located in the same local environment 110 as the display device 110, but corresponds to a computing device different from the display device 110. In these configurations, the transcriber 200 communicates with the display device 110 via a wired connection or a wireless connection. For example, the transcriber 200 has one or more ports that enable a wired / wireless connection such that the display device 110 functions as a peripheral device of the transcriber 200. Additionally or alternatively, the application forming the environment 100 may be compatible with the transcriber 200. For example, the transcriber 200 is configured as an input / output (I / O) device within the application such that an audio signal and / or a video signal adjusted by the application is sent to the transcriber 200 (e.g., in addition to the display device 110).

[0027] In some examples, the transcriptor 200 (and optionally the display device 110) is portable so that the transcriptor 200 can be moved between conference rooms. In some embodiments, the transcriptor 200 is configured with processing capabilities (e.g., processing hardware / software) to process audio and video content 112 to generate a transcript 202 when the content 112 is presented in the environment 100. In other words, the transcriptor 200 is configured to locally process the content 112 (e.g., audio and / or video content 218, 217) in the transcriptor 200 to generate the transcript 202 without additional remote processing (e.g., in the remote system 130). As used herein, this type of processing is referred to as on-device processing. Unlike remote processing, which often uses low-fidelity compressed video in server-based applications due to bandwidth constraints, on-device processing has no bandwidth constraints, so the transcriptor 200 can utilize high-fidelity and more accurate high-resolution video when processing video content. Further, this on-device processing can enable real-time tracking of speaker identification information without the latency due to waiting times that can occur when audio and / or video signals 218, 217 are remotely processed to some extent (e.g., in the remote computing system 130 connected to the transcriptor 200). To process content in the transcriptor 200, the transcriptor 200 includes hardware 210 such as data processing hardware 212 and memory hardware 214 that communicates with the data processing hardware 212. Some examples of the data processing hardware 212 include a central processing unit (CPU), a graphics processing unit (GPU), or a tensor processing unit (TPU).

[0028] In some embodiments, the transcriber 200 is executed on the remote system 130 by receiving content 112 (audio data and video data 217, 218) from each of the first and second feeds 112a - b, as well as the feed 112 from the conference room environment 100. For example, the data processing hardware 134 of the remote system 130 may execute instructions stored in the memory hardware 136 of the remote system 130 to execute the transcriber 200. Here, the transcriber 200 may process the audio data 218 and the image data 217 to generate a transcript 202. For example, the transcriber 200 may generate a transcript 202 and transmit it via the network 120 to the display device 110 for display on the display device 110. The transcriber 200 may similarly transmit the transcript 202 to the computing device / display device associated with the participants 10g - i corresponding to the first feed and / or the participant 10j corresponding to the second feed 10j.

[0029] In addition to the processing hardware 210, the transcriber 200 includes peripheral devices 216. For example, for processing audio content, the transcriber 200 includes audio capture devices 216, 216a (e.g., microphones) that capture the ambient sound (e.g., voice utterances) around the transcriber 200 and convert the sound into an audio signal 218 (FIGS. 2A and 2B) (or audio data 218). The audio signal 218 may then be used by the transcriber 200 to generate a transcript 202.

[0030] In some examples, the transcriber 200 includes, as peripheral devices 216, image capture devices 216, 216b. Here, the image capture device 216b (e.g., one or more cameras) can capture image data 217 (FIGS. 2A and 2B) as an additional input source (e.g., video input) that, in combination with the audio signal 218, helps identify which participant 10 is speaking (i.e., the speaker) within the multi-party participant environment 100. In other words, by including both the audio capture device 216a and the image capture device 216b, the transcriber 200 can process the image data 217 captured by the image capture device 216b to identify visual features (e.g., facial features) indicating which participant 10 of the plurality of participants 10a - 10j is speaking (i.e., generating the utterance 12) in a particular instance, so the transcriber 200 can improve the accuracy of speaker identification. In some configurations, the image capture device 216b is configured to capture 360 degrees around the transcriber 200 to capture a panoramic view of the environment 100. For example, the image capture device 216b includes an array of cameras configured to capture a 360-degree field of view.

[0031] Additionally or alternatively, by using the image data 217, if the participant 10 has a speech disability, the transcript 202 can be improved. For example, the transcriber 200 may have difficulty generating a transcript for a speaker with a speech disability who has problems clearly pronouncing utterances. To overcome the inaccuracy of the transcript 202 caused by such problems with clear pronunciation, the transcriber 200 (e.g., in the automatic speech recognition (ASR) module 230 of FIGS. 2A and 2B) can be made to recognize problems with clear pronunciation during the generation of the transcript 202. By recognizing the problem, the transcriber 200 can utilize the image data 217 representing the face of the participant 10 during the conversation to generate an improved or otherwise more accurate transcript 202 than if the transcript 202 were based only on the audio data 218 of the participant 10, thereby addressing the problem. Here, a particular speech disability may be prominent in the image data 217 from the image capture device 216b. For example, in the case of speech dysarthria, a neuromuscular disorder causing lip movements that affect clear pronunciation can be recognized in the image 217. Further, techniques can be used that analyze the image data 217 to correlate the lip movements of the participant 10 with a particular speech disability to the utterances intended by these participants 10, thereby improving automatic speech recognition in a way that is not possible using only the audio data 218. In some embodiments, by using the image 217 as an input to the transcriber 200, the transcriber 200 can identify potential problems with clear pronunciation and consider this problem to improve the generation of the transcription 202 during ASR.

[0032] In some embodiments, such as FIGS. 1B - 1E, the transcriber 200 is privacy - conscious so that participant 10 can opt out of sharing their audio and / or image information (e.g., in the transcript 202 or video feeds 112, 217). Here, one or more participants 10 communicate a privacy request 14 indicating the privacy conditions regarding participant 10 during participation in the video conferencing environment 100. In some examples, the privacy request 14 corresponds to the configuration settings of the transcriber 200. The privacy request 14 can occur before, during, or at the start of a meeting or communication session using the transcriber 200. In some configurations, the transcriber 200 includes a profile (e.g., profile 500 shown in FIG. 5) indicating one or more privacy requests 14 regarding participant 10 (e.g., individual profiles 510, 510a - n in FIG. 5). Here, the profile 500 is stored on the device (e.g., in the memory hardware 214) or outside the device (e.g., in the remote storage resource 136) and can be accessed by the transcriber 200. The profile 500 is configured before the communication session and can include an image of the face of individual participant 10 (e.g., image data 217) so that it can be correlated with individual portions of the video content 217 received by participant 10. That is, when the video content 217 of that participant 10 within the content feed 112 matches the face image associated with the individual profile 510, the individual profile 510 for individual participant 10 can be accessed. The individual profile 510 can be used to apply the participant's privacy settings during each communication session in which participant 10 participates. In these examples, the transcriber 200 can recognize participant 10 (e.g., based on the image data 217 received by the transcriber 200) and apply appropriate settings to participant 10.For example, profile 500 can include personal profiles 510, 510b for specific participants 10, 10b who don't care about being seen (i.e., don't care about being included in video feed 217), but don't want to be heard (i.e., want to be excluded from audio feed 218), and don't want their own utterances 12 to be transcribed (i.e., want conversations to be excluded from transcript 202). On the other hand, personal profiles 510, 510c for other participants 10, 10c may not want to be seen (i.e., want to be excluded from video feed 217), but it's okay for their own utterances to be recorded and / or transcribed (i.e., it's okay for them to be included in audio feed 218 and in transcript 202).

[0033] Referring to FIG. 1B, a third participant 10c submitted a privacy request 14 (i.e., a privacy request 14 regarding identifier privacy) with privacy conditions indicating that the third participant 10c doesn't care about being seen or heard, but doesn't want transcript 202 to include an identifier 204 (e.g., a label for the speaker's identification information) regarding the third participant 10c when the third participant 10c is speaking. In other words, the third participant 10c doesn't want their identification information to be shared or stored. Thus, the third participant 10c selects for transcript 202 not to include an identifier 204 associated with the third participant 10c that reveals their identification information. Here, FIG. 1B shows a transcript 202 with an edited gray box where the identifier 204 for speaker 3 would be present, but the transcriptor 200 may completely remove the identifier 204 or obscure it in another way that prevents the speaker's identification information associated with privacy request 14 from being revealed by the transcriptor 200. In other words, FIG. 1B shows that the transcriptor 200 modifies a part of transcript 202 so that it doesn't include the speaker's identification information (e.g., by removing or obscuring identifier 204).

[0034] Figure 1C is the same as Figure 1B, except that a third participant 10c who transmits privacy claim 14 requests not to be seen in any video feeds 112, 217 of the environment 100 (e.g., another form of identification information privacy). Here, the requesting participant 10c prefers to visually hide their video identification information (i.e., not share or store their video identification information within the video feeds 112, 217) while not being concerned about being heard. In this situation, the transcriber 200 is configured to blur, distort, or otherwise obscure the visual presence of the requesting participant 10c through the communication session among the participants 10, 10a - 10j. For example, the transcriber 200 determines the position of the requester 10c in any particular instance from the image data 217 received from one or more content feeds 112 and applies an abstraction (e.g., blur) 119 to any physical characteristics of the requester transmitted through the transcriber 200. That is, when the image data 217 is displayed on the screen 111 of the display device 110 and on the screen of the remote environment related to the participants 10g - 10j, the abstraction 119 at least overlaps the face of the requester 10c to make the requester 10c not visually identifiable. In some examples, the personal profile 510 regarding the participant 10 identifies whether the participant 10 desires to be blurred or obscured (i.e., distorted) or completely removed (e.g., as shown in Figure 5). Accordingly, the transcriber 200 is configured to augment, modify, or remove a portion of the video data 217 to hide the video identification information of the participant.

[0035] In contrast, FIG. 1D shows an example where a privacy request 14 from a third participant 10c requests that the transcriber 200 not track either the video representation of the third participant 10c or the conversation information of the third participant 10c. As used herein, "conversation information" refers to the audio data 218 corresponding to the utterance 12 spoken by the participant 10c, as well as the transcript 202 recognized from the audio data 218 corresponding to the utterance 12 spoken by the participant 10c. In this example, the participant 10c can be heard during the meeting, but the transcriber 200 does not remember the participant 10c either audibly or visually (e.g., by the video feed 217 or the transcript 202). This approach can protect the privacy of the participant 10c by not recording the conversation information of the participant 10c in the transcript 202 or recording the identifier 204 that identifies the participant 10c in the transcript 202. For example, the transcriber 200 can completely omit (omit) the text portion in the transcript 202 that transcribes (transcribe) the utterance 12 spoken by the participant 10c, or the transcript 202 can be made not to apply the identifier 204 that identifies the participant 10c while leaving these portions of the text intact. However, the transcriber 200 may apply any other identifier that simply distinguishes these portions of the text in the transcription 202 from other portions corresponding to the utterances 12 spoken by the other participants 10a, 10b, 10d - 10j without individually identifying the participant 10c. In other words, the participant 10 can request that the transcript 202 and any other records generated by the transcriber 200 not have a record of the participant's participation in the communication session (e.g., by the privacy request 14).

[0036] In contrast to the identification information privacy request 14, FIG. 1E shows the content privacy request 14. In this example, the third participant 10c transmits the privacy request 14 to the transcriber 200 so that the transcriber 200 does not include the content from the third participant 10c in the transcript 202. Here, the reason why the third participant 10c makes such a privacy request 14 is that the third participant 10c plans to discuss confidential content (e.g., confidential information) during the meeting. For the confidentiality of the content, the third participant 10c takes precautions to prevent the transcriber 200 from storing the voice content 218 associated with the third participant 10c in the transcript 202. In some embodiments, the transcriber 200 receives a privacy request 14 that identifies (e.g., by keywords) the types of content that one or more participants 10 do not want to be included in the transcript 202, and is configured to determine when content of that type occurs during the communication session in order to exclude content of that type from the transcript 202. In these embodiments, not all voice content 218 from a particular participant is excluded from the transcript 202 so that the particular participant can still discuss other types of content and may be included in the transcript 202; rather, only content-specific voice is excluded. For example, the third participant 10c transmits a privacy request 14 that requests that the transcriber 200 not transcribe the voice content about Mike. In this case, when the third participant 10c discusses Mike, the transcriber 200 does not transcribe this voice content 218, but when the third participant talks about other topics (e.g., the weather), the transcriber 200 transcribes this voice content 218. Similarly, the participant 10c may set a time boundary so that the transcriber 200 does not store the voice content 218 for a certain period of time (e.g., the next two minutes).

[0037] Figures 2A and 2B are examples of a transcriber 200. The transcriber 200 generally includes a diarization module 220 and an ASR module 230 (e.g., an AVASR module). The diarization module 220 receives voice data 218 corresponding to utterance 12 from participant 10 in a communication session (e.g., captured by voice capture device 216a), and image data 217 representing the face of participant 10 in the communication session, divides the voice data 218 into a plurality of segments 222, 222a-n (e.g., fixed-length segments or variable-length segments), and uses a probability model (e.g., a probabilistic generative model) based on the voice data 218 and the image data 217 to generate a diarization result 224 including corresponding speaker labels 226 assigned to each segment 222. In other words, the diarization module 220 includes a series of speaker recognition tasks with short utterances (e.g., segments 222) and determines whether two segments 222 of a given conversation are spoken by the same participant 10. At the same time, the diarization module 220 can execute a face tracking routine to identify which participant 10 is speaking between which segments 222 to further optimize speaker recognition. Next, the diarization module 220 is configured to repeat the process for all segments 222 of the conversation. Here, the diarization result 224 provides timestamped speaker labels 226, 226a-e for the received voice data 218 that not only identify the person speaking between a given segment 222 but also identify the time when a speaker change occurred between adjacent segments 222. Here, the speaker label 226 can function as an identifier 204 in the transcript 202.

[0038] In some examples, the transcriber 200 receives the privacy request 14 in the diarization module 220. Since the diarization module 220 identifies the speaker label 226 or identifier 204, the diarization module 220 can advantageously resolve the privacy request 14 corresponding to the privacy request 14 based on the identification information. In other words, the diarization module 220 receives the privacy request 14 when the privacy request 14 requires that the identifier 204 such as the label 226 does not identify the participant 10 when the participant 10 is the speaker. When the diarization module 220 receives the privacy request 14, the diarization module 220 is configured to determine whether the participant 10 corresponding to the request 14 matches the label 226 generated for a given segment 222. In some examples, an image of the face of the participant 10 can be used to associate the participant 10 with the label 226 regarding the participant 10. If the label 226 for the segment 222 matches the identification information of the participant 10 corresponding to the request 14, the diarization module 220 may prevent the transcriber 200 from applying the label 226 or identifier 204 to the corresponding part of the obtained transcription 202 which is the text transcription of the specific segment 222. If the label 226 for the segment 222 does not match the identification information of the participant 10 corresponding to the request 14, the diarization module 220 may allow the transcriber to apply the label 226 and identifier 204 to the part of the obtained transcription 202 which is the text transcription of the specific segment. In some embodiments, when the diarization module 220 receives the request 14, the ASR module 230 is configured to wait to transcribe the voice data 218 from the utterance 12.In other embodiments, the ASR module 230 performs real-time speech recognition, and the obtained transcription 202 removes the label 226 and the identifier 204 in real time for any participant 10 who provides a privacy request 14 to opt out of having their conversation information transcribed. Optionally, the diarization module 220 may further distort the voice data 218 associated with these privacy-seeking participants 10 such that the voices of the participants 10 are modified and cannot be used to identify the participants 10.

[0039] The ASR module 230 is configured to receive the voice data 218 corresponding to the utterance 12 and the image data 217 representing the face of the participant 10 while the utterance 12 is being made. Using the image data 217, the ASR module 230 transcribes the voice data 218 into the corresponding ASR result 232. Here, the ASR result 232 refers to the text transcription (e.g., the transcript 202) of the voice data 218. In some examples, the ASR module 230 communicates with the diarization module 220 and utilizes the diarization result 224 associated with the voice data 218 to improve the speech recognition based on the utterance 12. For example, the ASR module 230 can apply different speech recognition models (e.g., language model, prosody model) to different speakers identified from the diarization result 224. Additionally or alternatively, the ASR module 230 and / or the diarization module 220 (or some other components of the transcriptor 200) can index the transcription 232 of the voice data 218 using the timestamped speaker labels 226 predicted for each segment 222 obtained from the diarization result 224. In other words, the ASR module 230 uses the speaker label 226 from the diarization module 220 to generate an identifier 204 for the speaker in the transcript 202. As shown in FIGS. 1A - 1E, the transcript 202 for the communication session in the environment 100 can be indexed by the speaker / participant 10 so that multiple portions of the transcript 202 can be associated with the individual speaker / participant 10 to identify what each speaker / participant 10 said.

[0040] In some configurations, the ASR module 230 receives a privacy request 14 for the transcriber 200. For example, the ASR module 230 always receives the privacy request 14 for the transcriber 200 when the privacy request 14 corresponds to a request not to transcribe the conversation of a specific participant 10. In other words, the ASR module 230 always receives the privacy request 14 when the request 14 is not a label / identifier-based privacy request 14. In some examples, when the ASR module 230 receives the privacy request 14, the ASR module 230 first identifies the participant 10 corresponding to the privacy request 14 based on the speaker label 226 determined by the diarization module 220. Next, when the ASR module 230 encounters a conversation to be transcribed for that participant 10, the ARS module 230 applies the privacy request 14. For example, if the privacy request 14 requests not to transcribe the conversation of that specific participant 10, the ASR module 230 does not transcribe any of the participant's conversations and waits for a conversation to occur by another participant 10.

[0041] Referring to FIG. 2B, in some embodiments, the transcriber 200 includes a detector 240 for executing a face tracking routine. In these embodiments, the transcriber 200 first processes the audio data 218 to generate one or more candidate identification information for the speaker. For example, for each segment 222, the diarization module 220 may include a plurality of labels 226, 226a as candidate identification information for the speaker. 1-3 In other words, the model may be a probability model that outputs a plurality of labels 226, 226a for each segment 222. Here, the plurality of labels 226, 226a 1-3 1-3 ​Each label 226 is a potential candidate for identifying the speaker. Here, the detector 240 of the transcriber 200 uses the images 217, 217a-n captured by the image capture device 216b to determine which candidate identification information had the best visual features indicating that it was the speaker of a particular segment 22. In some configurations, the detector 240 generates a score 242 for each candidate identification information, and the score 242 indicates a confidence level that the candidate identification information is the speaker based on the correlation between the audio signal (e.g., audio data 218) and the video signal (e.g., the captured images 217a-n). Here, the highest score 242 may indicate that the candidate identification information is most likely to be the speaker. In FIG. 2B, the diarization module 220 generates three labels 226a 1-3 in a particular segment 222. The detector 240 generates a score 242 (e.g., three scores 242 1-3 shown as such) for each of these labels 226 based on the image 217 from the time within the audio data 218 where the segment 222 occurs. Here, FIG. 2B shows the highest score 242 by the bold square around the third label 226a3 associated with the third score 2423. If the transcriber 200 includes the detector 240, the best candidate identification information can be communicated to the ASR module 230 to form the identifier 204 of the transcript 202.

[0042] Additionally or alternatively, the process may be reversed such that the transcriptor 200 first processes the image data 217 to generate one or more candidate identification information for the speaker based on the image data 217. Next, for each candidate identification information, the detector 240 generates a confidence score 242 indicating the likelihood that the face corresponding to the corresponding candidate identification information contains the face of the person speaking during the corresponding segment 222 of the audio data 218. For example, the confidence score 242 for each candidate identification information indicates the likelihood that the face corresponding to the corresponding candidate identification information contains the face of the person speaking during the image data 217 corresponding to the instance in time of the segment 222 of the audio data 218. In other words, for each segment 222, the detector 240 can score 242 whether the image data 217 corresponding to the participant 10 has an expression similar to or matching the expression of the face of the person speaking. Here, the detector 240 selects the identification information of the speaker of the corresponding segment of the audio data 218 having the highest confidence score 242 as the candidate identification information.

[0043] In some examples, the detector 240 is part of the ASR module 230. Here, the ASR module 230 implements a face tracking routine by implementing an encoder front end having an attention layer configured to receive a plurality of video tracks 217a-n of the image data 217, whereby each video track is associated with the face of an individual participant. In these examples, the attention layer in the ASR module 230 is configured to determine a confidence score indicating the likelihood that the face of an individual person associated with the video face track contains the face of the person speaking on the audio track. Additional concepts and features related to an audiovisual ASR module including an encoder front end with an attention layer for multi-speaker ASR recognition are described in U.S. Provisional Patent Application No. 62 / 923,096, filed Oct. 18, 2019, which is incorporated herein by reference in its entirety.

[0044] In some configurations, the transcriber 200 (e.g., in the ASR module 230) is configured to support the multilingual environment 100. For example, when the transcriber 200 generates a transcript 202, the transcriber 200 can generate the transcript 202 in a plurality of distinct languages. This feature can enable the environment 100 to include a remote location having one or more participants 10 who speak a language different from the host location. Further, depending on the situation, a speaker in a meeting may be a non-native speaker or the language of the meeting may not be the speaker's first language. Here, the transcript 202 of the content from the speaker can assist other participants 10 in the meeting in understanding the content presented. Additionally or alternatively, the transcriber 200 may be used to provide feedback to the speaker regarding their pronunciation. Here, by combining video data and / or audio data, the transcriber 200 can indicate incorrect pronunciation (e.g., enabling the speaker to learn and / or adapt with the assistance of the transcriber 200). As such, the transcriber 200 can provide a notification to the speaker providing feedback regarding their pronunciation.

[0045] Figure 3 is an exemplary operational configuration regarding a method 300 for transcribing content (e.g., in the data processing hardware 212 of the transcriptor 200). In operation 302, method 300 includes receiving an audiovisual signal 217, 218 that includes audio data 218 and image data 217. The audio data 218 corresponds to the audio utterances 12 from a plurality of participants 10, 10a-n in the conversation environment 100, and the image data 217 represents the faces of the plurality of participants 10 in the conversation environment 100. In operation 304, method 300 includes receiving a privacy request 14 from one of the plurality of participants 10a-n, which is participant 10. The privacy request 14 indicates privacy conditions related to the participant 10 in the conversation environment 100. In operation 306, method 300 divides the audio data 218 into a plurality of segments 222, 222a-n. In operation 308, method 300 includes performing operations 308, 308a-c for each segment 222 of the audio data 218. In operation 308a, for each segment 222 of the audio data 218, method 300 includes determining, based on the image data 217, the identification information of the speaker of the corresponding segment 222 of the audio data 218 from among the plurality of participants 10a-n. In operation 308b, for each segment 222 of the audio data 218, method 300 includes determining whether the identification information of the speaker of the corresponding segment 222 includes a participant 10 related to the privacy conditions indicated by the received privacy request 14. In operation 308c, for each segment 222 of the audio data 218, if the identification information of the speaker of the corresponding segment 222 includes the participant 10, method 300 includes applying the privacy conditions to the corresponding segment 222. In operation 310, method 300 includes processing the plurality of segments 222a-n of the audio data 218 to determine a transcript 202 of the audio data 218.

[0046] In situations where the specific embodiments described herein may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, and the user's activities and demographic information, relationships between users, etc.), the user is provided with one or more opportunities to control whether to collect information, whether to store personal information, whether to use personal information, and how information about the user is collected, stored, and used. That is, the systems and methods described herein may collect, store, and / or use a user's personal information only if an explicit approval to do so has been received from the relevant user.

[0047] For example, the user is provided with control over whether a program or function collects user information about that particular user or other users related to the program or function. Each user from whom personal information is collected is presented with one or more options to enable control over the information collection related to that user and to provide permission or approval regarding whether information is collected and which portions of the information are collected. For example, one or more such control options can be provided to the user via a communication network. Further, specific data can be processed in one or more ways such that information that can identify an individual is removed before the data is stored or used. As an example, a user's identification information can be treated so as not to be able to identify an individual.

[0048] FIG. 4 is a schematic diagram of an exemplary computing device 400 that may be used to implement the systems and methods described herein. Computing device 400 is intended to represent various forms of digital computers, such as a laptop, desktop, workstation, personal digital assistant, server, blade server, mainframe, and other appropriate computers. The components shown herein, their connections and relationships, and their functions are exemplary only and are not intended to limit embodiments of the invention described herein and / or claimed in the claims.

[0049] Computing device 400 includes a processor 410 (e.g., data processing hardware), a memory 420 (e.g., memory hardware), a storage device 430, a high-speed interface / controller 440 connected to memory 420 and high-speed expansion port 450, and a low-speed interface / controller 460 connected to low-speed bus 470 and storage device 430. Each of the components 410, 420, 430, 440, 450, and 460 may be interconnected using various buses and may be mounted on a common motherboard or in other appropriate manners. Processor 410 processes instructions for execution within computing device 400, including instructions stored in memory 420 or storage device 430, and may display graphical information for a graphical user interface (GUI) on an external input / output device such as display 480 connected to high-speed interface 440. In other embodiments, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and multiple types of memory. Also, multiple computing devices 400 may be connected and each device may provide a portion of the necessary processing (e.g., as a server bank, a group of blade servers, or a multi-processor system).

[0050] Memory 420 stores information non - transiently within computing device 400. Memory 420 may be a computer - readable medium, a volatile memory unit(s), or a non - volatile memory unit(s). The non - transient memory 420 may be a physical device used to store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 400 on a temporary or permanent basis. Examples of non - volatile memory include, but are not limited to, flash memory and read - only memory (ROM) / programmable read - only memory (PROM) / erasable programmable read - only memory (EPROM) / electrically erasable programmable read - only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase - change memory (PCM), and disks or tapes.

[0051] Storage device 430 can provide a mass storage device for computing device 400. In some embodiments, storage device 430 is a computer - readable medium. In various different embodiments, storage device 430 can be a floppy disk (registered trademark) device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid - state memory device, or an array of devices including devices in a storage area network or other configurations. In additional embodiments, a computer program product is tangibly embodied in an information medium. The computer program product includes instructions that, when executed, perform one or more of the methods as described above. The information medium is a computer - readable medium or a machine - readable medium such as memory 420, storage device 430, or memory on processor 410.

[0052] The high-speed controller 440 manages processes that consume a large amount of the bandwidth of the computing device 400, and the low-speed controller 460 manages processes that consume a large amount of lower bandwidth. Such role distribution is merely exemplary. In some embodiments, the high-speed controller 440 is connected to a high-speed expansion port 450 that accepts a memory 420, a display 480 (e.g., via a graphics processor or accelerator), and various expansion cards (not shown). In some embodiments, the low-speed controller 460 is connected to a storage device 430 and a low-speed expansion port 470. The low-speed expansion port 470, including various communication ports (e.g., USB, Bluetooth®, Ethernet®, Wireless Ethernet®), can be connected to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a network device such as a switch or router via, for example, a network adapter.

[0053] As shown in the drawings, the computing device 400 can be implemented in several different forms. For example, it can be implemented as a standard server 400a, or multiple times within a group of such servers 400a, as a laptop computer 400b, or as part of a rack server system 400c.

[0054] The various embodiments of the systems and techniques described in this specification can be implemented in digital electronic circuitry and / or optical circuitry, integrated circuits, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include embodiments in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device, which may be of special or general purpose.

[0055] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly language / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any device and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including any computer program product, non-transitory computer-readable medium, and machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0056] The processes and logical flows described in this specification can be performed by one or more programmable processors executing one or more computer programs that process input data to generate output. The processes and logical flows can also be performed by special-purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer is operatively coupled to receive data from, transfer data to, or both, one or more mass storage devices (such as magnetic disks, magneto-optical disks, or optical disks) for storing data. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including semiconductor memory devices (such as EPROM, EEPROM, and flash memory devices), magnetic disks (such as internal hard disks or removable disks), magneto-optical disks, and CDROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.

[0057] To provide an interaction with a user, one or more aspects of the present disclosure can be implemented on a computer device having, for example, a display device for displaying information to the user such as a CRT (cathode ray tube), an LCD (liquid crystal display) monitor, or a touch screen, and optionally a keyboard and a pointing device (e.g., a mouse or a trackball) for the user to provide input to the computer. Other types of devices can be used to provide an interaction with the user, along with feedback provided to the user, which can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and the input from the user can be received in any form including acoustic, voice, or tactile input. Further, the computer can interact with the user by transmitting and receiving documents with the devices used by the user (e.g., by sending a web page to a web browser on the user's client device in response to a request received from a web browser).

[0058] Some embodiments have been described. Nevertheless, it will be understood that various changes can be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are also within the scope of the following claims.

Claims

1. A computer-executable method executed by data processing hardware, the method comprising causing the data processing hardware to: receive an audio signal including audio data, the audio data including a plurality of audio utterances in a conversation environment; receive a privacy request indicating privacy conditions, the privacy conditions including content-specific conditions indicating types of content to be excluded from a transcript; before generating a transcript of the audio data corresponding to the type of content, process the audio data to identify one or more first audio utterances not corresponding to the type of content among the plurality of audio utterances; process the audio data to identify one or more second audio utterances corresponding to the type of content among the plurality of audio utterances; generate the transcript based on the audio data, the transcript including the one or more first audio utterances not corresponding to the type of content among the plurality of audio utterances and excluding the one or more second audio utterances corresponding to the type of content among the plurality of audio utterances.

2. The method according to claim 1, wherein the data processing hardware is present on a device that is local to a user associated with the audio data.

3. The method according to claim 2, wherein processing the audio data is performed locally on the device.

4. The method according to claim 1, wherein the type of content includes content corresponding to a duration.

5. The method according to claim 1, wherein the type of content includes content associated with a specific person.

6. The method according to claim 1, wherein the operation further includes associating each individual audio utterance of the plurality of audio utterances of the audio data with one of a first user or a second user.

7. The method according to claim 6, wherein the privacy request is applied only to each individual audio utterance associated with the first user among the plurality of audio utterances.

8. The privacy request for each individual audio utterance associated with the first user among the plurality of audio utterances, The method according to claim 6, applied to each individual voice utterance associated with the second user among the plurality of voice utterances.

9. The method according to claim 1, wherein the voice signal is received as part of a voice video signal further including image data representing the faces of a plurality of users in the conversation environment.

10. The method according to claim 9, wherein the image data includes high-resolution video processed by the data processing hardware.

11. A system comprising: Data processing hardware; Memory hardware communicating with the data processing hardware, the memory hardware storing instructions which, when executed on the data processing hardware, cause the data processing hardware to: Receive a voice signal including voice data, the voice data including a plurality of voice utterances in a conversation environment; Receive a privacy request indicating privacy conditions, the privacy conditions including content-specific conditions indicating the types of content to be excluded from the transcript; Before generating a transcript of the voice data corresponding to the type of content, Process the voice data to identify one or more first voice utterances among the plurality of voice utterances that do not correspond to the type of content; Process the voice data to identify one or more second voice utterances among the plurality of voice utterances that correspond to the type of content; Generate the transcript based on the voice data, the transcript including the one or more first voice utterances among the plurality of voice utterances that do not correspond to the type of content and excluding the one or more second voice utterances among the plurality of voice utterances that correspond to the type of content.

12. The system according to claim 11, wherein the data processing hardware is present on a device that is local to the user associated with the voice data.

13. The system according to claim 12, wherein processing the voice data is performed locally on the device.

14. The system according to claim 11, wherein the type of content includes content corresponding to a duration.

15. The system according to claim 11, wherein the type of the content includes content associated with a specific person.

16. The system according to claim 11, wherein the operation further includes a step of associating each of the plurality of voice utterances of the voice data with one of a first user or a second user for each individual voice utterance.

17. The system according to claim 16, wherein the privacy requirement is applied only to each individual voice utterance associated with the first user among the plurality of voice utterances.

18. The privacy requirement is for each individual voice utterance associated with the first user among the plurality of voice utterances and for each individual voice utterance associated with the second user among the plurality of voice utterances, and is applied to the system according to claim 16.

19. The system according to claim 11, wherein the voice signal is received as part of a voice video signal further including image data representing the faces of a plurality of users in the conversation environment.

20. The system according to claim 19, wherein the image data includes high-resolution video processed by the data processing hardware.

Citation Information

Patent Citations

  • Minutes creating system, minutes data creating method, minutes data creating program

    JP2004080486A

  • Methods, programs and devices for representing meeting content

    JP2017229060A

  • Conference support system and conference support program

    JP2019061594A

  • Conditional disclosure of personally controlled content in a group context

    JP2019533247A