Voice reconstruction system for multimedia file
Through the combination of pseudo-audio-visual speech recognition, self-supervised language participle and symbol-to-speech model, visual and audio data are used to generate synthetic speech, which solves the problem of speech reconstruction under damaged audio, and improves the effectiveness of the speech recognition system and the speech quality of multimedia files.
Patent Information
- Application Number
- CN202480004883.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2024-02-21
- Publication Date
- 2025-08-19
AI Technical Summary
The prior art is difficult to effectively handle multiple forms of damaged audio, resulting in poor performance of speech recognition systems when handling different forms of damaged audio.
Using a pseudo-audiovisual speech recognition (P-AVSR) model combined with a self-supervised language (SSL) word segmenter and a symbol-to-speech (TTS) model, the visual data and audio data are used to generate pronunciation data through visual cues and convert them into encoded data to synthesize computer-generated synthesized speech, and play synchronously to replace corrupted audio.
It realizes efficient reconstruction of speech under a variety of damaged audio conditions, improves the accuracy and completeness of speech recognition, and enhances the voice comprehension effect of multimedia files.
Smart Images

Figure CN120513477A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to speech recognition, and more particularly, to reconstructing speech from a human recorded in a media file that also includes corrupted audio. Background Art
[0002] Various systems can be used to recognize speech, including spoken language, from multimedia files with corrupted audio. Typically, one system is constructed to handle a specific form of corrupted audio (e.g., speech restoration), while another system is constructed to handle another form of corrupted audio (e.g., speech with low intelligibility due to background noise). In this regard, a system may include a model trained for a specific form of corrupted audio, and the training may make the model unsuitable for handling other forms of corrupted audio. Summary of the Invention
[0003] According to one aspect, a method is provided, comprising: obtaining audiovisual data including visual data associated with a person and audio data associated with the person; determining pronunciation data associated with the person's speech based on the visual data; converting the speech into encoded data; and synthesizing the speech based on the encoded data to obtain synthesized speech.
[0004] The method may further include outputting synthesized speech while playing or rendering the visual data, wherein the synthesized speech is synchronized with movement associated with the person.
[0005] The method may further include, in response to determining the damaged portion of the audio data: determining a duration during which the damaged portion occurs; and outputting the synthesized speech during the duration.
[0006] Determining pronunciation data associated with the speech may include determining, by the first model, visual cues of the person using the visual data.
[0007] Converting the speech into encoded data may include converting the visual cues into pronunciation data by the first model.
[0008] Synthesizing the speech may include generating synthesized speech based on pronunciation data determined according to the visual prompt.
[0009] Converting the speech into the encoded data may include converting the speech into the encoded data via a second model. The second model may be trained to encode the visual cues by assigning codes to the visual cues.
[0010] The method may further include removing background noise from the audio data, wherein determining the pronunciation data is based on visual data including visual cues associated with the person.
[0011] The visual cue may include one or more mouth movements associated with the person.
[0012] According to another aspect, a device is provided, comprising: one or more processors; and at least one memory storing instructions, which, when executed by the one or more processors, cause the device to: obtain audio-visual data, the audio-visual data comprising visual data associated with a person and audio data associated with the person; determine pronunciation data associated with the person's speech based on the visual data by utilizing a first model; convert the speech into encoded data by utilizing a second model; and synthesize the speech based on the encoded data by utilizing the second model to obtain synthesized speech.
[0013] When executed by one or more processors, the instructions may further cause the device to: present the audiovisual data and the synthesized speech through a display and a speaker; and synchronize the synthesized speech with the person's movements while presenting the audiovisual data.
[0014] When executed by the one or more processors, the instructions further cause the device, in response to determining the corrupted portion of the audio data: determining a duration during which the corrupted portion occurs; and presenting the synthesized speech during the duration.
[0015] When executed by the one or more processors, the instructions further cause the apparatus to determine pronunciation data associated with the speech based on visual cues associated with the person determined by the first model using the visual data.
[0016] When executed by one or more processors, the instructions further cause the device to: convert speech into encoded data based on converting the visual cues into pronunciation data by the first model, and optionally generate synthesized speech based on the pronunciation data determined based on the visual cues, and further optionally, wherein the second model is trained to encode the visual cues by assigning codes to the visual cues.
[0017] When executed by the one or more processors, the instructions further cause the device to: remove background noise from the audio data; and determine pronunciation data based on the visual data including visual cues of the person.
[0018] According to another aspect, a non-transitory computer-readable medium storing instructions is provided, which, when executed, causes: obtaining audiovisual data, the audiovisual data including visual data associated with a person and audio data associated with the person; determining pronunciation data associated with the person's speech based on the visual data by utilizing a first model; converting the speech into encoded data by utilizing a second model; and synthesizing the speech based on the encoded data by utilizing the second model to obtain synthesized speech.
[0019] The instructions, when executed, may also cause: outputting the synthesized speech as computer-generated synthesized speech by utilizing a second model; and outputting the computer-generated synthesized speech while playing or rendering the audio-visual data, wherein the computer-generated synthesized speech is synchronized with movements associated with the person.
[0020] The instructions, when executed, may also be responsive to determining a damaged portion of the audio data: determining a duration during which the damaged portion occurs; and outputting the synthesized speech during the duration.
[0021] Some examples of the present disclosure relate to devices (e.g., head-mounted displays, communication devices) that include one or more machine learning models designed to generate synthesized speech and / or text and replace corrupted audio in multimedia content and / or multimedia files with the synthesized speech and / or text. The one or more machine learning models can rely on visual cues from a person captured in the multimedia content and / or multimedia files and predicted pronunciation tokens corresponding to the visual cues.
[0022] In some aspects of the present disclosure, a speech recognition system may be provided that can determine speech in the presence of various forms of corrupted audio. The speech recognition system may include a pseudo-audio-visual model (P-AVSR) that can receive data that may include video (of a person / speaker) and audio. The audio may include corrupted audio (e.g., background noise, speech from another person), and pronunciation data may be determined from visual content (e.g., a person's mouth movements). Thus, the audio-visual model can rely on visual cues in the form of speech units to determine textual data from the speech. Additionally, the speech recognition system may include a self-supervised learning (SSL) tokenizer that can receive the determined textual data and convert the textual data into clean speech. The SSL tokenizer can determine the sound emitted by a letter or sequence of letters, thereby determining a sound emitted based on the textual data. The sound recognized / determined by the SSL tokenizer may include human speech. The determined speech (e.g., human speech) may be synthesized to match the received textual data and presented as computer-generated synthesized speech and / or text.
[0023] In one exemplary aspect of the present disclosure, a method is provided for enabling (one or more) devices to reconstruct speech based on multimedia files and / or multimedia content. The method may include obtaining audiovisual data, the audiovisual data including i) visual data associated with a person and ii) audio data associated with the person. The method may also include determining text data associated with the person's speech based on the visual data. The method may also include converting the speech into coded data. The method may also include synthesizing the speech based on the coded data to obtain synthesized speech.
[0024] In another exemplary aspect of the present disclosure, a device for reconstructing speech based on multimedia files and / or multimedia content is provided. The device may include one or more processors and a memory including computer program code instructions. The memory and the computer program code instructions are configured to, together with at least one of the multiple processors, cause the device to obtain audiovisual data, the audiovisual data including i) visual data associated with a person and ii) audio data associated with the person. The memory and the computer program code instructions are further configured to, together with the processor, cause the device to determine textual data associated with the person's speech based on the visual data using a first model. The memory and the computer program code are further configured to, together with the processor, cause the device to convert the speech into encoded data using a second model. The memory and the computer program code are further configured to, together with the processor, cause the device to synthesize speech based on the encoded data using the second model to obtain synthesized speech. Additionally, optionally or alternatively, in some examples of the present disclosure, the memory and the computer program code may also be configured to, together with the processor, cause the device to present the audiovisual data and the synthesized speech. In some aspects of the present disclosure, the device may present the audiovisual data and the synthesized speech through a display and a speaker. For example, the display may play or render the visual content of the audiovisual data, and the speaker may play or output the audio content and synthesized speech of the audiovisual data.
[0025] In another example aspect of the present disclosure, a computer program product is provided that enables (one or more) devices to reconstruct speech based on multimedia files and / or multimedia content. The computer program product includes at least one computer-readable storage medium having computer-executable program code instructions stored therein. The computer-executable program code instructions may include program code instructions configured to obtain audio-visual data, the audio-visual data including i) visual data associated with a person and ii) audio data associated with the person. The computer program product may also include program code instructions configured to determine text data associated with the person's speech based on the visual data using a first model. The computer program product may also include program code instructions configured to convert the speech into encoded data using a second model. The computer program product may also include program code instructions configured to synthesize speech based on the encoded data using the second model to obtain synthesized speech.
[0026] Additional advantages will be set forth in part in the following description or may be learned through practice. These advantages will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims. It should be understood that both the foregoing general description and the following detailed description are exemplary and illustrative only and are not restrictive as claimed.
[0027] It should be understood that any feature described herein as suitable for incorporation into one or more aspects or embodiments of the present disclosure is intended to be generalizable to any and all aspects and embodiments of the present disclosure. Other aspects of the present disclosure will be apparent to those skilled in the art from the specification, claims, and drawings of the present disclosure. The foregoing general description and the following detailed description are exemplary and illustrative only and are not limiting of the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Certain features of the subject technology are set forth in the appended claims.For purposes of explanation, however, several examples of the subject technology are set forth in the following figures.
[0029] Figure 1 Displays presenting media files or media content are shown in accordance with aspects of the present disclosure.
[0030] Figure 2 A block diagram of a system for reconstructing speech from a media file or media content is shown, in accordance with aspects of the present disclosure.
[0031] Figure 3 Shown according to various aspects of the present disclosure Figure 2 Additional block diagrams of the system are shown, illustrating other features and functions of the system.
[0032] Figure 4A 、 Figure 4B and Figure 4C Exemplary movements of a person according to aspects of the present disclosure are shown.
[0033] Figure 5 An example of a flow chart illustrating the operation of a device that can reconstruct speech from a multimedia file or multimedia content according to aspects of the present disclosure is shown.
[0034] Figure 6 An example of a flow chart illustrating alternative operations for a device that can reconstruct speech from multimedia files or multimedia content is shown, according to aspects of the present disclosure.
[0035] Figure 7 A diagram illustrating an example of an artificial reality system in accordance with aspects of the present disclosure.
[0036] Figure 8 An example of a machine learning framework including a machine learning model and a training database according to aspects of the present disclosure is shown. DETAILED DESCRIPTION
[0037] Some embodiments of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments of the present disclosure are shown. Indeed, the various embodiments of the present disclosure may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein.
[0038] The present disclosure relates to a speech enhancement system designed to determine and synthesize speech from a multimedia file, including human speech present in the multimedia file. Specifically, the system can determine and synthesize human speech even in the presence of corrupted audio within the multimedia file. Furthermore, the speech enhancement system can utilize multiple models pre-trained on different datasets, allowing the speech enhancement system to manage various forms of corrupted audio, such as partial audio loss, complete audio loss, overlap with speech from other speakers, and overlap with ambient noise. This speech enhancement system can provide a more efficient speech synthesis approach compared to using multiple discrete systems.
[0039] In one or more embodiments, the system includes a pseudo-audio-visual speech recognition (P-AVSR) model that receives data in the form of a multimedia file including audio and video of a speaker, the speaker's audio including corrupted audio (e.g., background noise, speech from other people, poor microphone performance). The P-AVSR model may include a speech recognition model that is designed to determine what the speaker is saying based on both audio and visual cues, and output the determined speech as pronunciation data. For example, the P-AVSR model can be pre-trained to rely on a person's mouth movements. Typically, people tend to have similar mouth movements for specific sounds (e.g., an "ahh" sound, an "eee" sound, an "ohh" sound). In addition, the P-AVSR model can rely on visual cues independently of the recorded audio and output pronunciation data based on visual cues (e.g., a person's mouth movements), including in some cases using only visual cues. Advantageously, the P-AVSR model may not be negatively affected by corrupted audio in the multimedia file. Additionally, the use of visual cues can allow the system to manage multiple forms of corrupted audio, thereby allowing the system to function as a general voice enhancement system that manages different forms of corrupted audio in multimedia files by determining voice regardless of the different forms of corrupted audio.
[0040] Additionally, the speech enhancement system can include a self-supervised language (SSL) tokenizer model designed to encode clean speech audio (e.g., speech without corruption) into pronunciation symbols. Prior to being utilized, the SSL tokenizer model can be pre-trained with one or more datasets having speech (e.g., clean speech). Through training, the SSL tokenizer model can learn specific sounds in words and assign encoding data (e.g., symbols in the form of encoded numbers or alphanumeric phrases) to each specific sound. For example, the "ahh" sound can be encoded or labeled as 10, the "eee" sound can be encoded as 20, and the "ohh" sound can be encoded as 30.
[0041] When trained, the SSL tokenizer can create training data for training the P-AVSR model. Given a multimedia file containing both video and audio, the P-AVSR model, once trained, can be designed to predict pronunciation symbols. The P-AVSR model can generate symbols for every perceived sound that can be generated based on the video and audio. Furthermore, the audio may include corrupted audio.
[0042] The speech enhancement system may also include a symbol-to-speech (TTS) model. Similar to the P-AVSR model, the TTS model may also be trained on data generated by the SSL word segmenter. Using the symbol sequence generated by the P-AVSR model, the TTS model may output synthesized speech in the form of computer-generated spoken language or computer-generated text. Systems that incorporate synthesized speech in this manner provide several advantages. For example, when speech is determined from damaged audio and the damaged audio is removed, the computer-generated synthesized speech and / or text may be played (e.g., rendered, output, etc.) over the damaged audio in a multimedia file and synchronized with the mouth movements of the person in the multimedia file. In this regard, the computer-generated synthesized speech and / or text takes into account any missing speech from the person in the multimedia file due to damage.
[0043] Reference below Figures 1 to 7 These and other embodiments are discussed. However, those skilled in the art will readily appreciate that the detailed description given herein with respect to these figures is for explanatory purposes only and should not be construed as limiting.
[0044] Figure 1 A display 100 is shown presenting a media file 102 according to various aspects of the present disclosure. The media file 102, which represents other media files shown and / or described herein, may include a multimedia file having an audio component and a video component. In this regard, the media file 102 may take the form of a recorded audiovisual file that is stored on a storage device such as a server or a memory on another computing system and can be played back, for example, via the display 100. As shown, a person 104 (e.g., a real person, an animated character) is speaking, as indicated by a bubble 106. The media file 102 may utilize a microphone 108 to record the person 104. In addition to the person 104 speaking, the media file 102 may also include other forms of sound detected by the microphone 108. For example, the media file 102 may include a person "off-camera" or speaking in the background, as indicated by a bubble 110. Additionally, a media player 112 is playing (e.g., rendering or outputting) music 114.
[0045] In some cases, other forms of sound detected by microphone 108 may corrupt the speech of person 104. Additionally, microphone 108 may experience a malfunction, and during the malfunction, audio of person 104 may not be available. Other problems may arise when recording media file 102, such as internet connectivity issues, thereby reducing the overall quality of the audio in media file 102. Any one or more of the above-mentioned problems may be present in media file 102, and the speech from person 104 may not be fully intelligible. In some examples of the present disclosure, media file 102 may be referred to herein as media content 102.
[0046] Figure 2 A block diagram of a system 220 for reconstructing speech from a media file according to aspects of the present disclosure is shown. In some examples of the present disclosure, the system 220 may be referred to herein as a communication device. In some examples of the present disclosure, the system 220 may be a single integrated standalone device. In other examples of the present disclosure, one or more components of the system 220 may be remote from the system 220. In some examples, the system 220 may, but need not, be a head mounted display (HMD) (e.g., Figure 7 HMD 610), smart glasses, augmented / virtual reality devices, etc. System 220 can be used to recognize relevant speech from media files and / or media content (e.g., speech from / associated with person 104, where the speech is in Figure 1 102 ), converting relevant speech (e.g., speech from a person / associated with a person) into pronunciation symbols, and converting the pronunciation symbols into computer-generated synthesized speech. In this respect, system 220 can take the form of a speech recognition system.
[0047] As shown, system 220 may include one or more processors 222. As non-limiting examples, processor(s) 222 may include circuitry such as a central processing unit (CPU), a graphics processing unit (GPU), one or more microcontrollers, one or more microelectromechanical systems (MEMS) controllers, an application specific integrated circuit (ASIC), or a combination thereof.
[0048] The system 220 may also include a memory circuit 224 in communication with one or more processors 222. The memory circuit 224 may include read-only memory (ROM), random access memory (RAM), or a combination thereof. The one or more processors 222 may execute instructions stored on the memory circuit 224, such as recognizing speech recognition instructions and speech synthesis instructions.
[0049] The system 220 may also include an audio processor 226 in communication with the one or more processors 222. The audio processor 226 is designed to analyze the audio data from the media file and determine various types of audio, including whether the audio is attributed to the person speaking on the media file, other people who are speaking, and background noise. The audio processor 226 may be implemented as software stored on the memory circuit 224, or as hardware via the one or more processors 222. Alternatively, the audio processor 226 may be a microphone ( Figure 2 A portion of the ).
[0050] The system 220 may also include an image processor 228 in communication with the one or more processors 222. The image processor 228 is designed to process images of a person (e.g., Figure 1The image processor 228 may analyze the image data based on one or more visual cues (e.g., mouth movements) of the person 104 in the video. For example, the image processor 228 may process an image of a person speaking to determine the speech units spoken by the person with his or her mouth while speaking. As a result, the image processor 228 may determine the type of sound made by the person while speaking based on the visual cues. The image processor 228 may be implemented as software stored on the memory circuit 224 or as hardware via the one or more processors 222. Furthermore, the image processor 228 may recognize speech independently of the audio processor 226.
[0051] The system 220 may also include one or more input-output devices 229 (I / O devices) in communication with the one or more processors 222. As non-limiting examples, the input-output devices 229 may include a display (representing one or more displays), a microphone (representing one or more microphones), and a speaker (representing one or more audio speakers). As a non-limiting example, the system 220 may take the form of a head-mounted device (HMD).
[0052] The system 220 may also include one or more models 230 in communication with the one or more processors 222. As non-limiting examples, the one or more models 230 may include a speech recognition model, a self-supervised learning model, and a symbol-to-speech model. The one or more models 230 may be implemented as software stored on the memory circuit 224 or as hardware via the one or more processors 222.
[0053] Figure 3 Shown according to various aspects of the present disclosure Figure 2 An additional block diagram of system 220 is shown, illustrating other features and functions of system 220. As shown, system 220 can receive media file 102 (e.g., media content) that includes video portion 116 and audio portion 118. Video portion 116 and audio portion 118 can include digitally stored representations of images and sounds, respectively, of media file 102. Furthermore, in some cases, audio portion 118 includes corrupted audio in which human speech in media file 102 is difficult to understand.
[0054] The one or more models 230 may include an audio-visual model 232. In one or more embodiments, the audio-visual model 232 takes the form of a P-AVSR model. In this regard, the audio-visual model 232 may take the form of a speech recognition model designed to recognize a person (e.g., Figure 1104) and converts the recognized speech into pronunciation symbols, where the pronunciation symbols correspond to the speech of the person determined by the audio-visual model 232. In particular, the audio-visual model 232 can analyze (e.g., sample) the video portion 116 of the media file 102 to determine visual cues (e.g., mouth movements) of the person in the media file 102. In this regard, the audio-visual model 232 can determine speech units based on the determined mouth movements of the person in the media file 102. The audio-visual model 232 can be trained with data (e.g., video data, audio data, pronunciation data) of various images with visual cues (such as mouth movements) and individual sounds made with specific visual cues (e.g., specific mouth movements). Thus, the audio-visual model 232 can learn to recognize pronunciations associated with specific mouth movement sequences. Furthermore, the audio-visual model 232 can effectively ignore the audio portion 118 of the media file 102 and use only visual cues. Advantageously, any corrupted portion (or portions) of the audio portion 118 in the form of background noise or other human speech are removed and do not affect the pronunciation symbols determined by the audio-visual model 232 .
[0055] The audio-visual model 232 may include a video encoder 234, an audio encoder 236, or an audio-visual fusion block 238. The video encoder 234 and the audio encoder 236 may compress the video portion 116 and the audio portion 118, respectively, for subsequent playback (e.g., when generating computer-generated synthesized speech and / or text, as described below). The audio-visual fusion block 238 may merge the video and audio from the video encoder 234 and the audio encoder 236, respectively. Additionally, the audio-visual model 232 may determine one or more segments of the audio portion 118 that are damaged. The computer-generated synthesized speech provided by the system 220 (discussed below) may be played back (e.g., rendered, output, etc.) for the duration of the damaged audio segment. In this regard, the audio-visual fusion block 238 may merge the video from the video encoder 234 with the audio from the audio encoder 236, as well as with the computer-generated synthesized speech.
[0056] The one or more models 230 may also include an SSL model 240 designed to encode clean speech audio (e.g., speech without corruption) into pronunciation symbols. The SSL model 240 can be trained on audio data. Thus, the SSL model 240 can learn to group sounds into symbols representing different pronunciations. In this regard, the SSL model 240 can include a tokenizer 242 designed to generate encoded data in the form of symbols (e.g., encoded numbers, alphanumeric phrases) and assign them to the learned sounds, thereby effectively assigning symbols to letters or sequences of letters to encode the audio data. When trained, the SSL model 240 can create training data for training the audio-visual model 232. In this regard, given a multimedia file that includes both video and audio, the audio-visual model 232, once trained, is designed to predict pronunciation symbols. The audio-visual model 232 can generate symbols for each perceived sound that can be made based on the pronunciation data.
[0057] The one or more models 230 may also include a TTS model 244 (a symbol-to-speech model). Similar to the audio-visual model 232, the TTS model 244 may also be trained using data generated by the SSL model 240. In one or more embodiments, the TTS model 244 comprises a pseudo-TTS model. When trained, the TTS model 244 may generate synthesized speech (e.g., computer-generated synthesized speech and / or computer-generated text) representing the pronunciation of the speech predicted by the audio-visual model 232. In this regard, the system 220 may generate an audio output 250 in the form of synthesized speech, wherein the synthesized speech represents clean speech. In one or more embodiments, the synthesized speech from the TTS model 244 (e.g., audio output 250) may be merged with the video from the video portion 116 of the media file 102. Furthermore, in one or more embodiments, corrupted audio in the audio portion 118 of the media file 102 is replaced or substituted by the computer-generated synthesized speech from the TTS model 244 (e.g., audio output 250). Advantageously, media file 102 may be played back with computer-generated synthesized speech and / or synthesized text in place of corrupted audio, thereby providing a viewer of media file 102 with an enhanced version of media file 102, particularly potentially increasing comprehension of media file 102 due in part to the computer-generated synthesized speech.
[0058] System 220 can also combine computer generated synthesized speech with human speech (e.g. Figure 1The computer-generated synthesized speech can be synchronized with the mouth movements of the person 104 shown on the display. In this regard, the computer-generated synthesized speech can be presented and played back to the viewer in a manner that is synchronized (e.g., consistent and aligned) with the person's mouth movements. Thus, the computer-generated synthesized speech and / or text can be heard simultaneously with the person's mouth movements shown on the display, without delay or being played before the person's mouth movements. Advantageously, corrupted audio can be replaced with the computer-generated synthesized speech.
[0059] Furthermore, the system 220 can determine corrupted portions of the audio portion 118, including respective durations of the corrupted portions. In this regard, the system 220 can remove (e.g., mute) the audio portion 118 for each corrupted portion for the duration of the respective corrupted audio portion of the audio portion 118. Furthermore, the system 220 can play (e.g., render, output) the computer-generated synthesized speech for the duration of each corrupted portion of the audio portion 118 while muting one or more corrupted portions of the audio portion 118.
[0060] Figure 4A 、 Figure 4B and Figure 4C illustrative motions of a person according to aspects of the present disclosure are shown. Figure 4A , person 304a is making an "ahh" sound. Figure 4B , person 304b is making the sound of "eee". Figure 4C , person 304c is making the sound "ohh". Each mouth movement of persons 304a, 304b, and 304c represents a respective speech unit. The sounds made by persons 304a, 304b, and 304c can be used as the basis for the SSL model 240 ( Figure 3 4 (shown). For example, SSL model 240 can be trained by receiving image data of each of persons 304a, 304b, and 304c and associated sounds uttered by persons 304a, 304b, and 304c based on respective mouth movements and / or mouth positioning of persons 304a, 304b, and 304c. Additionally, SSL model 240 can be trained with text data corresponding to each sound uttered by persons 304a, 304b, and 304c. The sounds uttered by persons 304a, 304b, and 304c are exemplary, and several additional mouth movements, associated sounds, and associated text data can be used as training data.
[0061] Figure 5 and Figure 6 Examples of flowcharts illustrating the operation of a device that can reconstruct speech from multimedia files and / or multimedia content (e.g., media files and / or media content) according to aspects of the present disclosure are shown. Each flowchart shown and / or described includes a flowchart that can be used by a system (e.g., Figure 2 and Figure 3 Thus, as a non-limiting example, the steps of the flowchart may be implemented in part by a display, a speaker, and one or more processors, and / or by a display, and / or by a speaker, etc. As another example, the steps of the flowchart may be implemented in part by a medium (e.g., a non-transitory computer-readable medium storing instructions executable by a device (e.g., one or more processors)).
[0062] Figure 5 An example of a flow chart 400 illustrating the operation of a device for reconstructing speech from multimedia files and / or multimedia content according to various aspects of the present disclosure is shown. At operation 402, audiovisual data is obtained. The audiovisual data may include i) visual data associated with a person and ii) audio data associated with the person. The audio data may include damaged audio data. At operation 404, pronunciation data is determined based on the visual data. The pronunciation data may also be based on damaged audio data (when present). The pronunciation data may be associated with the person's speech. At operation 406, the speech is converted into coded data. At operation 408, speech is synthesized based on the coded data to obtain synthesized speech.
[0063] Figure 6 An example of a flow chart 500 illustrating alternative operations for a device that can reconstruct speech from multimedia files and / or multimedia content according to various aspects of the present disclosure is shown. At operation 502, audio-visual data is obtained. The audio-visual data may include i) visual data associated with a person and ii) audio data associated with the person. The audio data may include damaged audio data. At operation 504, pronunciation data is determined using a first model. The pronunciation data may be associated with the person's speech based on the visual data. The pronunciation data may also be based on damaged audio data (when present). At operation 506, the speech is converted into coded data using a second model. At operation 508, speech is synthesized based on the coded data to obtain synthesized speech using the second model.
[0064] Figure 7An example of an artificial reality system 600 is shown. The subject matter disclosed herein (e.g., one or more portions of system 220) can be incorporated into artificial reality system 600. Artificial reality system 600 can include a head-mounted display (HMD) 610 (e.g., smart glasses and / or augmented / virtual reality devices) comprising a frame 612, one or more displays 614, a computing device 608 (also referred to herein as a computer), and a controller 604. In some examples, HMD 610 can capture one or more text items from one or more images / videos associated with a real-world environment in the field of view of one or more cameras (e.g., cameras 616, 618) of artificial reality system 600. HMD 610 can utilize the captured text from the one or more images / videos to trigger one or more actions / functions of artificial reality system 600. Display 614 can be transparent or translucent, allowing a user wearing HMD 610 to see through display 614 to see the real world (e.g., the real-world environment) while visual artificial reality content is displayed to the user. The HMD 610 may include an audio device 606 (e.g., a speaker / microphone) that may provide audio artificial reality content to the user. The HMD 610 may include one or more cameras 616, 618 that may capture images and / or video of the environment. In one exemplary embodiment, the HMD 610 may include one or more cameras 618, which may be rear-facing cameras that track the user's eye movements and / or gaze.
[0065] One of the cameras 616 may be a forward-facing camera that captures images and / or video of the environment that a user wearing the HMD 610 can view. The one or more cameras 616 may also be referred to herein as a front camera. The HMD 610 may include an eye-tracking system to track vergence movements of the user wearing the HMD 610. In one exemplary embodiment, the one or more cameras 618 may be an eye-tracking system. In some exemplary embodiments, the one or more cameras 618 may be a camera configured to view at least one eye of the user to capture glint images (e.g., and / or glint signals). The one or more cameras 618 may also be referred to herein as a rear camera. The HMD 610 may include a microphone for the audio device 606 to capture voice input from the user. The artificial reality system 600 may also include a controller 604 that includes a touchpad and one or more buttons. The controller 604 may receive input from the user and relay the input to the computing device 608. The controller 604 may also provide haptic feedback to the user or users. The computing device 608 may be connected to the HMD 610 and the controller 604 via a cable or a wireless connection. Computing device 608 can control HMD 610 and controller 604 to provide augmented reality content to one or more users and receive input from one or more users. In some example embodiments, controller 604 can be a standalone controller or integrated within HMD 610. Computing device 608 can be a standalone host computer device, an onboard computer device integrated with HMD 610, a mobile device, or any other hardware platform capable of providing artificial reality content to a user and receiving input from the user. In some example embodiments, HMD 610 can include an artificial reality system / virtual reality system.
[0066] Figure 8 An example of a machine learning framework 700 including a machine learning model 750 and a training database 760 according to aspects of the present disclosure is shown. The machine learning framework 700 can be hosted locally in a computing device or remotely. For purposes of illustration and not limitation, for example, the machine learning framework 700 can be hosted locally on a Figure 3 In the system 220 shown. The training database 760 can include several tasks (e.g., speech recognition tasks, self-supervised speech learning tasks). Using the training database 760, the machine learning framework 700 can train the machine learning model 750 to translate received text from one language to another language, and vice versa. In some aspects, for example, the machine learning model 750 can be stored by a computing device. In other aspects, for example, the machine learning model 750 can reside within a computing system, such as a portable electronic device, an HMD, a server, etc.
[0067] The training database 760 may include multiple training data sets, which may include one or more word sequences in the form of phrases and / or sentences. The one or more word sequences may include labeled data and / or unlabeled data. The word sequences may be labeled as including mouth movements. In addition, the word sequences may be labeled as including specific sounds. The labeled training data sets may be used, for example, to train a machine translation model, such as the machine learning model 750. The unlabeled training data sets may be used, for example, to validate training. The training database 760 employed by the machine learning framework 700 may be fixed or periodically updated. One or more models 230 may be used in a manner similar to or identical to the machine learning model 750.
[0068] Methods, systems, or devices, etc., as described herein, can provide speech reconstruction. A method, system, computer-readable storage medium, or device can determine speech in the presence of various forms of corrupted audio. A speech recognition system can include a pseudo-audio-visual model (P-AVSR) that can receive data that can include video (of a person / speaker) and audio. The audio can include corrupted audio (e.g., background noise, speech from another person), and textual data can be determined from visual content (e.g., a person's mouth movements). Thus, the audio-visual model can rely on visual cues in the form of speech units to determine textual data from speech. The speech recognition system can include a self-supervised learning (SSL) tokenizer that can receive determined textual data and convert the textual data into clean speech. The SSL tokenizer can determine the sound emitted by a letter or letter sequence, thereby determining a sound emitted based on the textual data. The sound recognized / determined by the SSL tokenizer can include human speech. The determined speech (e.g., human speech) can be synthesized to match the received textual data and can be presented as computer-generated synthesized speech and / or text. All combinations in this paragraph (including removal or addition of steps) are designed in a manner consistent with the rest of the detailed description.
[0069] In one exemplary aspect of the present disclosure, a method enables (one or more) devices to reconstruct speech based on provided multimedia files and / or multimedia content. The method may include obtaining audiovisual data, the audiovisual data including i) visual data associated with a person and ii) audio data associated with the person. The method may also include determining text data associated with the person's speech based on the visual data. The method may also include converting the speech into coded data. The method may also include synthesizing speech based on the coded data to obtain synthesized speech. All combinations of this paragraph and previous paragraphs (including removal or addition of steps) are designed in a manner consistent with other parts of the detailed description.
[0070] In an exemplary aspect of the present disclosure, a device for reconstructing speech based on multimedia files and / or multimedia content is provided. The device may include one or more processors and a memory including computer program code instructions. The memory and the computer program code instructions are configured to, using at least one of the processors, cause the device to obtain audiovisual data, the audiovisual data including i) visual data associated with a person and ii) audio data associated with the person. The memory and the computer program code instructions may be configured to, together with the processor, cause the device to determine textual data associated with the person's speech based on the visual data using a first model. The memory and the computer program code may be configured to, together with the processor, cause the device to convert the speech into encoded data using a second model. The memory and the computer program code may also be configured to, together with the processor, cause the device to synthesize speech based on the encoded data using the second model to obtain synthesized speech. In some examples of the present disclosure, the memory and the computer program code may also be configured to, together with the processor, cause the device to present the audiovisual data and the synthesized speech. In some aspects of the present disclosure, the device may present the audiovisual data or the synthesized speech via a display or a speaker. For example, the display may play or render the visual content of the audiovisual data, and the speaker may play or output the audio content of the audiovisual data or the synthesized speech. All combinations of this and previous paragraphs (including removal or addition of steps) are designed in a manner consistent with the rest of the detailed description.
[0071] In an exemplary aspect of the present disclosure, a computer program product can enable (one or more) devices to reconstruct speech based on provided multimedia files and / or multimedia content. The computer program product may include at least one computer-readable storage medium having computer-executable program code instructions stored therein. The computer-executable program code instructions may include program code instructions configured to obtain audio-visual data, the audio-visual data including i) visual data associated with a person and ii) audio data associated with the person. The computer program product may also include program code instructions configured to determine text data associated with the person's speech based on the visual data using a first model. The computer program product may also include program code instructions configured to convert the speech into coded data using a second model. The computer program product may also include program code instructions configured to synthesize speech based on the coded data using the second model to obtain synthesized speech. All combinations of this paragraph and the previous paragraph (including the removal or addition of steps) are designed in a manner consistent with other parts of the detailed description.
[0072] The foregoing description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the various aspects shown herein, but should be given the full scope consistent with the claim language, wherein unless clearly stated otherwise, reference to an element in the singular does not intend to mean "one and only one", but "one or more". Unless otherwise expressly stated, the term "some" refers to one or more. Male pronouns (e.g., he) include female and neutral genders (e.g., she and it), and vice versa. Titles and subtitles (if any) are used only for convenience and do not limit this subject disclosure.
[0073] Like reference numerals refer to like elements throughout. As used herein, the terms "data," "content," "information," and similar terms may be used interchangeably to refer to data that can be sent, received, and / or stored according to embodiments of the present disclosure. Furthermore, the term "exemplary," as used herein, is not provided to convey any qualitative assessment, but rather merely to convey an illustration of an example.
[0074] As defined herein, "computer-readable storage media," which refer to non-transitory, physical, or tangible storage media (eg, volatile or non-volatile memory devices), may be distinguished from "computer-readable transmission media," which refer to electromagnetic signals.
[0075] As referred to herein, a metaverse may represent an immersive virtual space or world in which devices may be utilized in a network, wherein one or more social connections may, but need not, exist in the network or with environments in the metaverse or world. A metaverse or metaverse network may be associated with a three-dimensional (3D) virtual world, an online game (e.g., a video game), one or more content items (such as, for example, images, videos, non-fungible tokens (NFTs)), and wherein the content items may be purchased, for example, with digital currency (e.g., cryptocurrency) and other suitable currencies. In some examples, a metaverse or metaverse network may enable the generation and provision of an immersive virtual space in which remote users may socialize, collaborate, learn, shop, and / or engage in various other activities within the virtual space, including through the use of augmented reality (AR) / virtual reality (VR) / mixed reality (MR).
[0076] In addition, as used in the specification, including the appended claims, the singular forms "a", "an", and "the" include plural forms, and references to a particular numerical value include at least that particular value, unless the context clearly dictates otherwise. As used herein, the term "plurality" means more than one. When a range of values is expressed, another embodiment includes from one particular value and / or to another particular value. Similarly, when values are expressed as approximations, by using the antecedent "about", it will be understood that the particular value forms another embodiment. All ranges are inclusive and combinable. It should be understood that the terms used herein are for the purpose of describing particular aspects only and are not intended to be limiting.
[0077] It should be understood that certain features of the disclosed subject matter described herein in the context of separate embodiments for the sake of clarity may also be provided in combination in a single embodiment. Conversely, various features of the disclosed subject matter described herein in the context of a single embodiment for the sake of brevity may also be provided individually or in any subcombination. Furthermore, any reference to a value stated in a range includes every value within that range.
[0078] It should be understood that the methods and systems described herein are not limited to specific methods, specific components or specific embodiments. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting.
[0079] As used herein, the phrase "at least one of" preceding a list of items, where the terms "and" or "or" are used to separate any item, modifies the entire list, rather than each member of the list (i.e., each item). The phrase "at least one of" does not require selection of at least one of each item listed; rather, the phrase allows for a meaning that includes at least one of any one item, and / or at least one of any combination of items, and / or at least one of each item. For example, the phrase "at least one of A, B, and C" or "at least one of A, B, or C" each refers to only A, only B, or only C; any combination of A, B, and C; and / or at least one of each of A, B, and C.
[0080] The predicate terms "configured to," "operable to," and "programmed to" do not imply any specific tangible or intangible modification of the subject matter, but are intended to be used interchangeably. In one or more embodiments, a processor configured to monitor and control an operation or component may also refer to a processor programmed to monitor and control an operation or a processor operable to monitor and control an operation. Similarly, a processor configured to execute code may be interpreted as a processor programmed to execute code or operable to execute code.
[0081] Phrases such as an aspect, this aspect, another aspect, some aspects, one or more aspects, one embodiment, this embodiment, another embodiment, some embodiments, one or more embodiments, an example, this example, another example, some examples, one or more examples, a configuration, this configuration, another configuration, some configurations, one or more configurations, the subject technology, the disclosure, the present disclosure, other variations thereof, etc., are used for convenience and do not imply that the disclosure associated with such phrases is essential to the subject technology or that such disclosure applies to all configurations of the subject technology. The disclosure associated with such phrases may apply to all configurations or one or more configurations. The disclosure associated with such phrases may provide one or more examples. Phrases such as an aspect or some aspects may refer to one or more aspects and vice versa, and similarly applies to the other aforementioned phrases.
[0082] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" or "example" is not necessarily to be construed as preferred or advantageous over other embodiments. Furthermore, to the extent that the terms "including," "having," and the like are used in the specification or claims, such terms are intended to be inclusive in a manner similar to the term "comprising" as "comprising" is interpreted when used as a transitional word in a claim. References in this specification to "an example," "an example," and the like may mean that the particular feature, function, or characteristic being described is included in at least one example of the embodiment. Occurrences of such phrases in this specification are not necessarily all referring to the same example and are not necessarily mutually exclusive.
[0083] When an element is referred to herein as being "connected" or "coupled" to another element, it should be understood that the element can be directly connected to the other element or that intervening elements may exist between the elements. Conversely, when an element is referred to as being "directly connected" or "directly coupled" to another element, it should be understood that there are no intervening elements in the "direct" connection between the elements. However, the presence of a direct connection does not preclude other connections that may have intervening elements.
[0084] Alternative Embodiments
[0085] The foregoing description of the embodiments has been presented for purposes of illustration; it is not intended to be exhaustive or to limit the patent rights to the precise forms disclosed. Those skilled in the relevant art will appreciate that many modifications and variations are possible in light of the above disclosure.
[0086] Some parts of this specification describe embodiments in terms of applications and symbolic representations of operations on information. Those skilled in the art of data processing typically use these application descriptions and representations to effectively convey the essence of their work to others skilled in the art. Although described functionally, computationally, or logically, these operations are understood to be implemented by computer programs or equivalent circuits, microcode, etc. In addition, without loss of generality, it sometimes proves convenient to refer to these operational configurations as components. The described operations and their associated components may be embodied in software, firmware, hardware, or any combination thereof.
[0087] Any steps, operations, or processes described herein may be performed or implemented using one or more hardware or software components, alone or in combination with other devices. In one embodiment, the software components are implemented using a computer program product comprising a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the steps, operations, or processes described.
[0088] The embodiments may also relate to an apparatus for performing the operations described herein. The apparatus may be specially constructed for the desired purpose, and / or it may include a computing device selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a non-transitory tangible computer-readable storage medium or any type of medium suitable for storing electronic instructions, which may be coupled to a computer system bus. Furthermore, any computing system mentioned in the specification may include a single processor, or may be an architecture that employs multiple processor designs to increase computing power.
[0089] Embodiments may also relate to products produced by the computing processes described herein. Such products may include information produced by the computing processes, wherein the information is stored on a non-transitory tangible computer-readable storage medium, and may include any embodiment of a computer program product or other data combination described herein.
[0090] The language used in the specification is selected primarily for readability and instructional purposes and may not be selected to delineate or limit the subject matter of the invention. Accordingly, it is intended that the scope of the present patent rights be limited not by this detailed description, but rather by any claims of an application based hereon. Accordingly, the disclosure of the embodiments is intended to illustrate, not to limit, the scope of the patent rights set forth in the appended claims.
Claims
1. A method comprising: obtaining audiovisual data, the audiovisual data comprising visual data associated with a person and audio data associated with the person; determining pronunciation data associated with the person's speech based on the visual data; Converting the speech into coded data; as well as The speech is synthesized based on the encoded data to obtain synthesized speech.
2. The method according to claim 1, further comprising: The synthesized speech is outputted while playing or rendering the visual data, wherein the synthesized speech is synchronized with movement associated with the person.
3. The method of claim 1 or 2, further comprising, in response to determining the damaged portion of the audio data: determining the duration of the presence of the damaged portion; and The synthesized speech is output during the duration.
4. A method according to any preceding claim, wherein: Determining the pronunciation data associated with the speech includes determining, by a first model, visual cues of the person using the visual data.
5. The method according to claim 4, wherein Converting the speech into the encoded data comprises converting the visual cues into the pronunciation data by the first model, and optionally, wherein synthesizing the speech comprises generating the synthesized speech based on the pronunciation data determined from the visual cues.
6. The method according to claim 5, wherein: Converting the speech into the encoded data includes converting the speech into the encoded data by a second model, wherein the second model is trained to encode the visual cues by assigning codes to the visual cues.
7. The method according to any one of the preceding claims, further comprising: Background noise is removed from the audio data, wherein determining the pronunciation data is based on the visual data comprising visual cues associated with the person, and optionally wherein the visual cues include one or more mouth movements associated with the person.
8. A device comprising: one or more processors; as well as at least one memory storing instructions that, when executed by the one or more processors, cause the apparatus to: obtaining audiovisual data, the audiovisual data comprising visual data associated with a person and audio data associated with the person; determining pronunciation data associated with the person's speech based on the visual data by utilizing a first model; converting the speech into encoded data by utilizing a second model; as well as The speech is synthesized based on the encoded data by using the second model to obtain synthesized speech.
9. The apparatus according to claim 8, wherein When executed by the one or more processors, the instructions further cause the device to: presenting the audiovisual data and the synthesized speech via a display and a speaker; and The synthesized speech is synchronized with the movement of the person while presenting the audiovisual data.
10. The apparatus according to claim 8 or claim 9, wherein When executed by the one or more processors, the instructions further cause the device, in response to determining the corrupted portion of the audio data: determining the duration of the presence of the damaged portion; and The synthesized speech is presented during the duration.
11. The apparatus according to any one of claims 8 to 10, wherein When executed by the one or more processors, the instructions further cause the device to: The pronunciation data associated with the speech is determined based on visual cues associated with the person determined by the first model using the visual data.
12. The apparatus according to claim 11, wherein When executed by the one or more processors, the instructions further cause the device to: Based on converting the visual cue into the pronunciation data by the first model, converting the speech into the encoded data, and optionally, generating the synthesized speech based on the pronunciation data determined according to the visual prompt, and further optionally, Wherein, the second model is trained to encode the visual cues by assigning codes to the visual cues.
13. The apparatus according to any one of claims 8 to 12, wherein When executed by the one or more processors, the instructions further cause the device to: removing background noise from the audio data; and The pronunciation data is determined based on the visual data including visual cues of the person.
14. A non-transitory computer-readable medium storing instructions that, when executed, cause: obtaining audiovisual data, the audiovisual data comprising visual data associated with a person and audio data associated with the person; determining pronunciation data associated with the person's speech based on the visual data by utilizing a first model; converting the speech into encoded data by utilizing a second model; as well as The speech is synthesized based on the encoded data by using the second model to obtain synthesized speech.
15. The non-transitory computer readable medium of claim 14, wherein: The instructions, when executed, further cause: outputting the synthesized speech as computer-generated synthesized speech by utilizing the second model; as well as outputting the computer-generated synthesized speech while playing or rendering the audiovisual data, wherein the computer-generated synthesized speech is synchronized with movements associated with the person, and optionally, wherein the instructions, when executed, further cause, in response to determining a corrupted portion of the audio data: determining the duration of the presence of the damaged portion; and The synthesized speech is output during the duration.