Techniques for secure synthesis of voice using the natural voice of a spoker during a
By updating voice profiles in real-time in communication sessions, the risk of deep forgery is solved, communication security is improved, fraudsters are prevented from impersonating others, and secure language translation and voice synthesis are achieved.
Patent Information
- Application Number
- CN202380080849.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-08
- Filing Date
- 2023-10-19
- Publication Date
- 2025-07-04
AI Technical Summary
Existing language translation and voice synthesis technologies pose a risk of deep forgery in real-time communication sessions, and fraudsters can impersonate others by replaying other people's voice records, resulting in communication security issues.
The speaker's voice profile is generated and refreshed at regular intervals during a communication session, and the voice profile is continuously updated by processing fixed-length audio data samples, and overwriting old profiles in volatile memory, reducing the possibility of fraudsters exploiting other people's voice profiles.
Significantly reduce the risk of deep forgery, preventing fraudsters from impersonating others by updating voice profiles in real time, improving the security and reliability of communication sessions.
Smart Images

Figure CN120266201A_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to a technique for securely synthesizing a speaker's natural voice when translating the speaker's speech from a first language to a second language during a real-time communication session, such as a one-on-one voice call, an audio- or video-based conference, or a similar online event. More specifically, the present application describes a technique for creating and then continuously recreating a speaker's voice profile for the duration of a communication session, thereby reducing the opportunity to mimic another person with a deepfake voice clone. Background Art
[0002] The use of artificial intelligence with language translation and voice synthesis technologies enables the translation of speech from a first language to a second language with high accuracy and minimal processing latency, near real-time. Language translation and voice synthesis technologies have been deployed in various applications and for various different use cases. For example, some personal communication applications use these technologies to enable voice calls and video calls between two or more call participants who would otherwise speak different languages. As an example, during a voice call utilizing these technologies, a first person who speaks a first language (e.g., English) can converse with a second person who speaks a second language (e.g., French). Similarly, many audio- and video-based conference applications use these technologies to allow one-to-many broadcasts, where a presenter speaks in a language that may be different from the language understood by other conference participants. As an example, an application facilitating a broadcast event, such as a conference application, can utilize these technologies to translate the presenter's speech into a synthetic voice in one or more alternative languages, suitable for other conference participants. Brief Description of the Drawings
[0003] Embodiments of the present invention are illustrated by way of example and not limitation in the accompanying drawings, in which:
[0004] Figure 1 is a diagram illustrating an example of an application having a voice profile service for continuously updating a speaker's voice profile when generating a language translation synthetic voice, according to some embodiments.
[0005] Figure 2 is a diagram illustrating a timing diagram of a system using a voice profile service such as that described in Figure 1 according to an embodiment of the present invention.
[0006] Figure 3 is a user interface diagram illustrating an example of a user interface of an application for facilitating the generation of a visual notification when a significant change to a speaker's voice profile for generating synthetic voice is detected during a communication session, according to an embodiment of the present invention.
[0007] Figure 4 is a flowchart illustrating an example of a technique for translating speech from a first language to a second language according to an embodiment of the present invention, wherein the synthesized speech in the second language has the audible characteristics of a speaker's natural voice.
[0008] Figure 5 is a block diagram illustrating a software architecture that can be installed on any computing device among various computing devices to execute the methods described herein.
[0009] Figure 6 illustrates a pictorial representation of a machine in the form of an exemplary computer system (e.g., a server computer), within which a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein. Detailed Description
[0010] Methods and systems for securely synthesizing a speaker's natural voice when translating the speaker's speech from a first language to a second language are described herein, as can occur during a voice call, an audio - or video - based conference call, or a similar online event. In the following description, for purposes of explanation, numerous specific details and features are set forth in order to provide a thorough understanding of various aspects of different embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention can be practiced and / or implemented using different combinations of the many details and features presented herein.
[0011] Technological advancements in speech synthesis have enabled the generation of synthesized speech with audible characteristics that mimic a person's natural voice - a concept often referred to as voice cloning. Many conventional techniques for voice cloning when synthesizing speech involve generating what is referred to herein as a voice profile. The term "voice profile" refers to voice data that encapsulates the unique voice characteristics of a real person, and which can be used by a speech synthesizer to generate synthesized speech with audible characteristics consistent with the speaker's natural or real voice. For example, when a speech synthesizer is generating synthesized speech, the voice profile is used as an input to the speech synthesizer to produce synthesized speech that mimics the voice of the person associated with the voice profile. Thus, a voice profile can be considered a digital copy of a person's natural voice.
[0012] For many conventional communication applications and services that use a voice synthesizer to provide language translation services, a voice profile for a person is created by prompting the person to speak and then capturing an audio recording of the person. For example, a speaker may be prompted to read back a specific sentence or word arrangement. The captured audio recording is then processed to identify the unique voice characteristics of the person, which are then stored in a voice profile for the person. When the person initiates a voice call or other communication session using the communication application, the person's voice profile is retrieved from storage and used as an input to the voice synthesizer when creating a language translation synthetic voice during a real-time communication session, such as a voice call or conference event.
[0013] The techniques described above are inherently risky due to the voice profiles stored by communication applications or services. For example, if the voice profile becomes inadvertently accessible to others, the voice profile may be used for malicious activities. For example, if a fraudster gains access to the voice profile, the fraudster may use another person's voice profile to deceive others by pretending to be the person to whom the voice profile belongs. Because the synthetic voice generated using the voice profile sounds to others as if it is being spoken in real-time or near real-time by a known and trusted person, the target participant may trust the fraudster's message and take some action contrary to his or her own best interests. These types of illegal activities or schemes - commonly referred to as deepfakes - have received widespread attention for their potential use in creating fake news, catfishing, bullying, and financial fraud.
[0014] Recently, communication applications and services have addressed the problems set forth above by developing techniques for capturing a speaker's voice recording and generating a voice profile for the speaker in real-time or near real-time, such as by sampling the speaker's voice at the beginning of a communication session. For example, when a communication session is initiated and the participant first speaks, the system detects the voice and captures a fixed-length sample of the voice, and generates a voice profile based on the fixed-length sample of the voice. Then, for the remainder of the communication session, the voice profile is used when generating a language translation synthetic voice in the speaker's natural voice. Thus, when the voice profile is created during an actual communication session in which the voice profile will be used, the communication application or service does not need to store a voice profile for the speaker. The voice profile derived for the speaker at the beginning of the communication session is used throughout the duration of the communication session to generate a synthetic voice in the speaker's natural voice.
[0015] However, the technique of using a voice profile to generate synthetic speech is still vulnerable to deepfake scenarios. For example, if a fraudster has previously captured a recording of another person, the fraudster can play back the recording of that person's voice (and speech) at the start of a communication session. Then, the voice profile service of the system will create a voice profile with the voice characteristics of the other person for the benefit of the fraudster, thereby allowing the fraudster to impersonate the other person for the remainder of the communication session. For example, after initially playing back the recording of the other person at the start of the communication session, the fraudster will be able to speak in his or her own voice, and the language translation synthetic speech will be generated using the voice profile of the other person, thus allowing the fraudster to impersonate the other person.
[0016] According to an embodiment of the present invention, a communication application or service has a voice profile service that generates and refreshes a voice profile for a speaker at regular intervals throughout the duration of a communication session. As an example, when a communication session is first initiated and the first participant in the communication session starts speaking, a fixed-length (e.g., 8 seconds) sample of the audio data representing the voice of the first participant is processed by the voice profile service to generate a voice profile for the first participant. Then, at some regular interval (e.g., every 30 seconds), the voice profile service will sample the voice of the first participant again to generate a refreshed or updated voice profile. In some cases, when generating an updated voice profile, the updated voice profile is written to a volatile memory storage device such that it overwrites the previously generated voice profile. This process continues continuously during the duration of the communication session. By iteratively refreshing or updating the voice profile of each participant in the communication session, the risk of deepfake scenarios is significantly reduced, where one person uses the voice profile of another person to achieve some fraudulent goal. For example, if a fraudster starts a call by playing back a recording of another person's voice, the initial voice profile generated for synthesizing translated speech will be of the other person's voice. However, when subsequent voice profiles are generated, the subsequent voice profiles will be based on the fraudster's voice, and thus, other participants in the communication session will observe a significant change in the voice of the synthesized speech. In fact, the change in voice will also be a signal to the participants that some malicious activity may be taking place. Other aspects and advantages of the present invention will be apparent from the following description of several drawings. The disclosed technical solution of resampling the voice of the first participant (e.g., the speaker) to generate a refreshed or updated voice profile thus solves the technical problem of preventing the abuse of previously generated voice profiles.
[0017] Figure 1FIG. 0 is a diagram illustrating an example of a communication service 100 having a voice profile service 102 for continuously updating a speaker's voice profile 104 when generating a language translation synthesized voice 116. The communication service 100 can facilitate one-to-one voice calls or voice calls with more than two call participants, such as may be the case with audio conferencing or video conferencing applications. In other cases, the communication service 100 can facilitate one-to-many broadcast-like events, such as online meetings or online seminars—sometimes referred to as webinars—where one or more dedicated presenters (e.g., speakers) can present to a large audience of participants. Thus, in some embodiments, the communication service 100 can facilitate the communication of both audio and visual information, including both video-based conferencing and web-based data presentation. However, in other cases, the communication service 100 can be a purely audio-based service. Additionally, the techniques described herein can be implemented using local digital communication services—e.g., such as Voice over Internet Protocol (VoIP) systems, Integrated Services Digital Network (ISDN), or other proprietary digital-based systems, but can also use analog-based systems, such as Plain Old Telephone Service (POTS).
[0018] As shown in Figure 1 FIG. 0, in addition to the voice profile service 102, the communication service 100 also includes a translation service 106 and a voice synthesizer service 108. When a communication session is first initiated, voice 112 in a first language of a first speaker is received at the communication service 100 and processed in parallel by the translation service 106 and the voice profile service 102. Although shown as separate services in Figure 1 FIG. 0, in some embodiments, the voice profile service 102 can be a sub-component of the translation service 106. In any case, the translation service 106 processes the received audio data representing the voice 112 by translating the speaker's voice 112 from the first language to a second language, while at the same time, the voice profile service processes a fixed-length sample of the audio data to derive a first voice profile for the speaker.
[0019] Translating the received speech 112 from a first language to a second language is typically accomplished in two steps. The translation service 106 first processes the audio data representing the received speech 112 to identify the component parts of the speech in the first language. Next, these identified component parts of the speech in the first language are translated or converted into component parts of the speech in the second language. In various embodiments, the component parts of the speech may take different forms - referred to as symbolic linguistic representations. For example, in some embodiments, the component parts of the speech may be text (e.g., words). However, in other embodiments, the component parts of the speech may be speech transcript - symbols that provide a visual representation of the speech sounds (or phonemes). In either case, after performing speech recognition to identify or recognize the component parts of the speech, the recognized speech is then converted or translated into component parts of the speech in the second (target) language. Thus, the output of the translation service 106 is a symbolic linguistic representation of the recognized speech 114 in the target or translated language.
[0020] When the translation service 106 is translating the speech received from a speaker, the voice profile service 102 obtains and processes a fixed - length sample of the audio data representing the speech 112 to generate a voice profile 104 for the speaker. Specifically, the voice profile service processes a fixed - length sample of the audio data representing the speech from the speaker to identify various acoustic characteristics. These acoustic characteristics are then used to generate the speaker's voice profile 104. As shown in Figure 1 the speaker's voice profile is an input to the speech synthesizer service 108, which uses the voice profile 104 as an input to generate a synthesized speech 116 in the target or translated language, to give the synthesized speech the audible characteristics of the speaker's natural voice.
[0021] Now referring to Figure 2 a timing diagram is shown to illustrate how the processes described above are iteratively performed at fixed intervals such that when a speaker speaks continuously during a communication session, the voice profile is continuously refreshed or updated. As shown in Figure 2As illustrated, the line with reference numeral 200 represents the timeline during which the speaker is speaking in a first language (e.g., English). The point designated as "T = 0" in the timeline represents the start of the communication session and thus the time when the speaker starts speaking in English. As shown by reference numeral 202, an eight - second sample of the audio data representing the first eight seconds of voice from the first thirty - second interval 204 is processed to generate a first voice profile 206 (e.g., "Voice Profile #1") for the speaker. Then, the speech synthesizer service 108 uses this first voice profile 206 to generate a portion of the synthetic voice 208 in a target or translated language (e.g., French), giving the synthetic voice audible characteristics according to the natural speech of the speaker.
[0022] As the speaker continues speaking, a second eight - second sample of the audio data 210 is captured during a second interval 212 and processed by the voice profile service to generate a second instance of the speaker's voice profile 214. Then, the speaker's new voice profile 214 is used as an input to the speech synthesizer service 108 to generate a second portion of the synthetic voice 216, giving the synthetic voice audible characteristics consistent with the natural speech of the speaker. This process is carried out continuously during the duration of the communication session, making it extremely difficult (if not impossible) for a fraudster to deceive the system by playing a recording of another person's voice.
[0023] Although each interval is shown in Figure 2 as being thirty seconds in length, in various alternative embodiments, the fixed interval can be of a different length, e.g., such as being between ten and forty seconds in length. Similarly, the sample of the audio data from which the voice profile is generated is shown in Figure 2 as being eight seconds. However, in alternative embodiments, the length of the sample can be greater than or less than eight seconds, such as being between two and ten seconds. Further, in the Figure 2 timing diagram, the sample of the audio data obtained during the fixed interval and from which the voice profile is derived is shown as being obtained from the first part (e.g., the first eight seconds) of the fixed interval (e.g., 30 seconds of audio data). However, in some embodiments, the audio data from the entire fixed interval is first processed and the sample is obtained from the portion (e.g., eight seconds) of the audio data having the best energy (e.g., the most significant voice characteristics). By selecting the portion of the audio data obtained during the fixed interval having the best energy, the quality of the resulting voice profile is improved and the situation where the audio sample has only silence or poor audio characteristics is avoided.
[0024] Referring again to Figure 1, according to some embodiments, whenever a voice profile is generated for a speaker, a new or updated voice profile for the speaker can be written to a volatile memory storage device by overwriting a previously generated voice profile for the speaker, thereby ensuring for security purposes that no more than one profile is temporarily stored in the volatile memory at any given time. Similarly, in some embodiments, when it is detected that a participant has terminated his or her participation in the communication session, the voice profile of the participant will be removed or deleted from the memory. When the communication session ends, any voice profiles generated for the participants in the communication session are deleted.
[0025] According to some embodiments of the present invention, the communication service 100 can facilitate the presentation of visual information within the user interface of an end user's computing device, and the end user establishes a communication session via the user interface. Thus, as Figure 1 illustrated by the dashed box with reference numeral 110 in the figure, in some embodiments, the optional user interface component 110 can be part of the communication service.
[0026] According to some embodiments, the user interface component 110 can generate a visual indicator for presentation via a user interface presented on the display of a participant's device, where the visual indicator serves as a warning that an updated voice profile for a particular speaker has changed significantly from a previously generated voice profile for the same speaker. For example, in some embodiments, the voice profile service can include a voice profile verification component or service (not shown in Figure 1 ). During a communication session, when the voice profile service 102 generates a voice profile for a speaker, before overwriting the previously generated voice profile for the speaker, the voice profile verification component can analyze the new voice profile to determine whether the new voice profile is significantly different from the previously generated voice profile. When the determined difference between some aspects of the new and old voice profiles exceeds a predetermined threshold, the voice profile verification service can cause the user interface component 110 to generate a visual indicator for presentation via a user interface shown on the display of a participant's device. The visual indicator will serve as a warning to the participant that the voice profile of the "speaker" has changed significantly, thus indicating the possibility of a fraudster. In Figure 3 the icon with reference numeral 300 and the "Warning" label in the figure are examples of visual indicators that can be presented to the participants of a communication session when a significant change has been detected in the continuously generated voice profiles of a speaker.
[0027] Figure 4is a flowchart illustrating an example of a technique (e.g., method 400) for translating speech received from a speaker's device from a first language to a second language and then generating synthetic speech for transmission to another participant. At method operation 402, during a real-time language translation communication session, the speech of a first participant (e.g., the speaker) is translated from the first language to the second language and then replayed as synthetic speech in the second language, obtaining a fixed-length sample of audio data representing the speaker's speech in the first language. The sample of audio data is processed by a voice profile service to generate a voice profile for the speaker. Generating the voice profile by the voice profile service is performed iteratively at regular intervals using audio data captured within a relevant interval. For example, in some embodiments, the interval length may be thirty seconds. Thus, every thirty seconds, the voice profile service obtains a fixed-length sample of audio data representing the speaker's speech and uses the sample to generate a voice profile for the speaker. In some embodiments, as each new or updated voice profile is generated, the new or updated voice profile is written to the same space in memory as the previously generated voice profile for the speaker, such that the new or updated voice profile overwrites the previously generated voice profile for the speaker. In this context, overwriting one voice profile with another may involve directly overwriting the voice profile - i.e., writing the new voice profile to the same memory address as the previous voice profile without first deleting or removing the existing voice profile from memory. Alternatively, overwriting the existing voice profile may involve first deleting or removing the existing voice profile from memory - to ensure that the existing voice profile is no longer available in the application domain before writing the new voice profile to memory.
[0028] At method operation 404, simultaneously - i.e., during the language translation communication session - the speaker's speech is translated to generate a text or phonetic transcription in the second language. For example, a speech translation service processes the speech received from the speaker's device and generates a symbolic language representation of the recognized speech as output. The representation may be text, or alternatively a phonetic transcription. In either case, the output of the translation service is provided as input to a speech synthesizer. The speech synthesizer uses the voice profile currently stored in memory - i.e., the most recently generated voice profile that has been written to memory - to generate synthetic speech based on the output of the translation service, such that the synthetic speech is based on the speaker's speech, is in the second language, and has audible characteristics consistent with the natural speech of the speaker associated with the voice profile used in generating the synthetic speech.
[0029] Finally, at method operation 406, when the communication session ends, or when one or more participants exit or leave the communication session, any and all associated voice profiles generated during the communication session are deleted from memory or otherwise made inaccessible.
[0030] By continuously and iteratively generating a speaker's voice profile during a communication session, a person with malicious intent cannot easily use a recording of another person's voice to generate a voice profile based on that other person's voice and thus cannot easily impersonate that person during the communication session.
[0031] Machine and software architecture
[0032] Figure 5 is a block diagram 800 that illustrates a software architecture 802 that can be installed on any of a variety of computing devices to perform the methods described herein. Figure 6 This is merely a non-limiting example of a software architecture, and it will be appreciated that many other architectures can be implemented to facilitate the functionality described herein. In various embodiments, the software architecture 802 is implemented by hardware of a machine 900 such as Figure 6 The machine 900 includes a processor 910, a memory 930, and input / output (I / O) components 950. In this example architecture, the software architecture 802 can be conceptualized as a stack of layers where each layer can provide a specific function. For example, the software architecture 802 includes layers such as an operating system 804, libraries 806, frameworks 808, and applications 810. In operation, according to some embodiments, the application 810 makes API calls 812 through the software stack and receives messages 814 in response to the API calls 812.
[0033] In various implementations, the operating system 804 manages hardware resources and provides common services. The operating system 804 includes, for example, a kernel 820, services 822, and drivers 824. According to some embodiments, the kernel 820 acts as an abstraction layer between the hardware and other software layers. For example, the kernel 820 provides memory management, processor management (e.g., scheduling), component management, networking, and security settings, among other functions. The services 822 can provide other common services to other software layers. According to some embodiments, the drivers 824 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 824 can include a display driver, a camera driver, or a low-power driver, a flash driver, a serial communication driver (e.g., a Universal Serial Bus (USB) driver), drivers, an audio driver, a power management driver, and the like.
[0034] In some embodiments, library 806 provides a low-level common infrastructure utilized by application 810. Library 806 can include system libraries 830 (e.g., C standard library) capable of providing functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, library 806 can include API libraries 832, such as media libraries (e.g., libraries for supporting the presentation and manipulation of various media formats such as Moving Picture Experts Group-4 (MPEG4), High Efficiency Video Coding (H.264 or AVC), Moving Picture Experts Group Layer 3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., OpenGL framework for rendering in two dimensions (2D) and three dimensions (3D) in a graphics context on a display), libraries for databases (e.g., SQLite for providing various relational database functions), web libraries (e.g., WebKit for providing web browsing functions), etc. Library 806 can also include a variety of other libraries 834 to provide many other APIs to application 810.
[0035] According to some embodiments, framework 808 provides a high-level common infrastructure that can be utilized by application 810. For example, framework 608 provides various GUI functions, advanced resource management, advanced location services, etc. Framework 808 can provide a wide range of other APIs that can be utilized by application 810, some of which may be specific to a particular operating system 804 or platform. In an example embodiment, application 810 includes home application 850, contacts application 852, browser application 854, book reader application 856, location application 858, media application 860, messaging application 862, game application 864, and a variety of other applications, such as third-party applications 866. According to some embodiments, application 810 is a program that executes functions defined in the program. One or more of the applications 810 can be created using various programming languages structured in various ways, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a specific example, third-party application 866 (e.g., an application developed using an ANDROID TM or IOS TM software development kit (SDK) by an entity other than the vendor of a particular platform) can be on a mobile operating system such as IOS TM 、ANDROID TM 、 Mobile software running on a mobile operating system (e.g., iOS or another mobile operating system). In this example, the third-party application 866 can call the API calls 812 provided by the operating system 804 to facilitate the functions described herein.
[0036] Figure 6 A graphical representation of a machine 900 in the form of a computer system according to an example embodiment is illustrated, within which a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein. Specifically, Figure 6 A diagrammatic representation of an example form of a machine 900 is shown, within which instructions 916 (e.g., software, program, application, applet, application, or other executable code) can be executed to cause the machine 900 to perform any one or more of the methods discussed herein. For example, the instructions 916 can cause the machine 900 to perform any one of the methods or algorithmic techniques described herein. Additionally or alternatively, the instructions 916 can implement any one of the systems described herein. The instructions 916 transform the general non-programmed machine 900 into a particular machine 900 programmed to perform the described and illustrated functions in the described manner. In an alternative embodiment, the machine 900 operates as a stand-alone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 900 can operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 900 can include, but is not limited to: server computers, client computers, PCs, tablet computers, laptop computers, netbooks, set-top boxes (STBs), PDAs, entertainment media systems, cellular phones, smart phones, mobile devices, wearable devices (e.g., smart watches), smart home devices (e.g., smart appliances), other smart devices, network devices, network routers, network switches, bridges, or any machine capable of sequentially or otherwise executing the instructions 916 to specify actions to be taken by the machine 900. Further, although only a single machine 900 is illustrated, the term "machine" should also be regarded as including a collection of machines 900 that individually or jointly execute the instructions 916 to perform any one or more of the methods discussed herein.
[0037] Machine 900 may include a processor 910, a memory 930, and I / O components 950, which may be configured to communicate with each other, such as via a bus 902. In an example embodiment, the processor 910 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 912 and a processor 914 that may execute instructions 916. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that may execute instructions simultaneously. Although Figure 6 multiple processors 910 are shown, machine 900 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0038] The memory 930 may include a main memory 932, a static memory 934, and a storage unit 936, all of which may be accessed by the processor 910, such as via the bus 902. The main memory 930, the static memory 934, and the storage unit 936 store instructions 916 embodying any one or more of the methods or functions described herein. The instructions 916 may also reside, in whole or in part, within at least one of the main memory 932, the static memory 934, the storage unit 936, the processor 910 (e.g., within the processor), or any suitable combination thereof, during execution by the machine 900.
[0039] The I / O components 950 may include a variety of components to receive input, provide output, produce output, send information, exchange information, capture measurements, and so forth. The specific I / O components 950 included in a particular machine will depend on the type of machine. For example, a portable machine, such as a mobile phone, will likely include a touch input device or other such input mechanism, while a headless server machine will likely not include such a touch input device. It will be appreciated that the I / O components 950 may include Figure 6Many other components not shown. The I / O components 950 are grouped according to functionality for simplicity of the following discussion only, and such grouping is in no way restrictive. In various example embodiments, the I / O components 950 may include an output component 952 and an input component 954. The output component 952 may include visual components (e.g., a display, such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibration motor, a resistance mechanism), other signal generators, etc. The input component 954 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical - optical keyboard, or other alphanumeric input components), point - based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or another pointing instrument), haptic input components (e.g., a physical button, a touch screen that provides the location and / or force of a touch or touch gesture, or other haptic input components), audio input components (e.g., a microphone), etc.
[0040] In additional example embodiments, the I / O components 950 may include a biometric component 956, a motion component 958, an environmental component 960, or a location component 962, as well as various other components. For example, the biometric component 956 may include components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying a person (e.g., voice recognition, retina recognition, face recognition, fingerprint recognition, or electroencephalogram - based recognition), etc. The motion component 958 may include an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. The environmental component 960 may include, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect the ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor for detecting the concentration of a hazardous gas for safety or measuring pollutants in the atmosphere), or other components that may provide an indication, measurement, or signal corresponding to the surrounding physical environment. The location component 962 may include a location sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects the air pressure from which altitude is derived), a direction sensor component (e.g., a magnetometer), etc.
[0041] A variety of techniques can be used to implement communication. The I / O component 950 can include a communication component 964 that is operable to couple the machine 900 to the network 980 or the device 970 via a coupling 982 and a coupling 972, respectively. For example, the communication component 964 can include a network interface component or another suitable device to interface with the network 980. In another example, the communication component 964 can include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, components (e.g., low power), components, and other communication components to provide communication via other modalities. The device 970 can be another machine or any of a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0042] In addition, the communication component 964 can detect identifiers or include components operable to detect identifiers. For example, the communication component 964 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor that detects one-dimensional barcodes such as universal product code (UPC) barcodes, and multi-dimensional barcodes such as quick response (QR) codes, Aztec codes, data matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying tagged audio signals). Additionally, various information can be derived via the communication component 964, such as a location via Internet protocol (IP) geolocation, a location via signal triangulation, a location via detecting an NFC beacon signal that can indicate a specific location, etc.
[0043] Executable instructions and machine storage media
[0044] Various memories (i.e., 930, 932, 934, and / or the memory of the (one or more) processors 910) and / or the storage unit 936 can store one or more sets of instructions and data structures (e.g., software) that embody or are used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 916), when run by the (one or more) processors 910, cause various operations to implement the disclosed embodiments.
[0045] As used herein, the terms "machine storage medium", "device storage medium", and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms refer to one or more storage devices and / or media that store executable instructions and / or data (e.g., a centralized or distributed database and / or associated caches and servers). Thus, the terms should be considered to include, but not be limited to: solid-state memories, and optical and magnetic media, including memories internal or external to a processor. Specific examples of machine storage media, computer storage media, and / or device storage media include: non-volatile memories, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGA, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium", "computer storage medium", and "device storage medium" specifically exclude carrier waves, modulated data signals, and other such media, at least some of which are covered under the term "signal medium" discussed below.
[0046] transmission medium
[0047] In various example embodiments, one or more portions of network 980 can be an ad hoc network, an intranet, an extranet, a VPN, a LAN, a WLAN, a WAN, a WWAN, a MAN, the Internet, a portion of the Internet, a portion of the PSTN, a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, a network, another type of network, or a combination of two or more such networks. For example, network 980 or a portion of network 980 can include a wireless or cellular network, and coupling 982 can be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or another type of cellular or wireless coupling. In this example, coupling 982 can implement any one of a variety of types of data transfer technologies, such as single carrier radio transmission technology (Ix RTT), evolved data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rate GSM evolution (EDGE) technology, third generation partnership project (3GPP) including 3G, fourth generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), worldwide interoperability for microwave access (Wi MAX), long term evolution (LTE) standards, other standards defined by various standards setting organizations, other long-range protocols, or other data transfer technologies.
[0048] Instructions 916 may be sent or received over network 980 using a transmission medium via a network interface device (e.g., a network interface component included in communication component 964) and utilizing any one of a variety of well-known transmission protocols (e.g., HTTP). Similarly, instructions 916 may be sent or received to / from device 070 via coupling 972 (e.g., a peer-to-peer coupling). The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure. The terms "transmission medium" and "signal medium" should be regarded as including any non-transitory medium that is capable of storing, encoding, or carrying instructions 916 for execution by machine 900, and including digital or analog communication signals or other non-transitory media to facilitate the communication of such software. Thus, the terms "transmission medium" and "signal medium" should be regarded as including any form of modulated data signal, carrier wave, etc. A "modulated data signal" refers to a signal having one or more of its characteristics set or changed to encode information in the signal.
[0049] Computer-readable medium
[0050] The terms "machine-readable medium", "computer-readable medium", and "device-readable medium" mean the same thing and may be used interchangeably in this disclosure. The terms are defined to include both machine storage media and transmission media. Thus, the terms include storage devices / media and carrier waves / modulated data signals.
Claims
1. A system for securely synthesizing speech in a speaker's natural voice during a language translation communication session, the system comprising: A hardware processor; And One or more memory storage devices storing instructions which, when run by the hardware processor, cause the system to perform operations including the following: During the language translation communication session: Periodically replace the speaker's voice profile at a fixed interval by: i) obtaining a sample of audio data from a stream of audio data received from the speaker's device during the fixed interval, the stream of audio data representing the speaker's speech in a first language, ii) generating a voice profile for the speaker based on the sample of audio data, and iii) writing the speaker's voice profile to the memory storage device; Process the stream of audio data received from the speaker's device by generating a stream of audio data for transmission to a participant in the communication session, the stream of audio data for transmission to the participant representing synthesized speech that has been translated from a first language to a second language, the synthesized speech being derived using the speaker's most recently generated voice profile stored in the memory storage device; And Transmit the stream of audio data representing the synthesized speech in the second language to the participant's device for playback.
2. The system according to claim 1, wherein, The operations further include: In response to detecting termination of the language translation communication session, delete the speaker's voice profile from the memory storage device.
3. The system according to claim 1, wherein Writing the speaker's voice profile to the memory storage device includes overwriting an instance of the speaker's voice profile previously written to the memory storage device.
4. The system according to claim 3, wherein, The language translation communication session is facilitated by an application running on the participant's device, the application having a user interface for presentation on a display of the device, and the operations further include: Before overwriting an instance of the speaker's voice profile previously written to the memory storage device: Compare the speaker's voice profile with the instance of the speaker's voice profile previously written to the memory storage device; and Based on the comparison, determine that the difference between the speaker's voice profile and the instance of the speaker's voice profile previously written to the memory storage device exceeds a threshold; and In response to the determination, cause a visual indicator to be presented via the user interface of the application running on the participant's device, wherein the visual indicator indicates detection of a significant change in the speaker's voice profile.
5. The system according to claim 1, wherein, Processing the stream of audio data received from the speaker's device by generating a stream of audio data for transmission to the participant's device includes: Perform speech recognition on a stream of the audio data received from the speaker's device to generate a signed language representation of the speech in the second language for the speech in the first language; and Utilize a speech synthesizer to generate a stream of the audio data for transmission to the device of the participant in the communication session, using the signed language representation of the speech in the second language as input.
6. The system according to claim 5, wherein, The signed language representation of the speech in the second language includes:[[]] Text in the second language; or A speech transcription in the second language.
7. The system according to claim 1, wherein, Obtaining a sample of the audio data from the stream of the audio data received from the speaker's device includes:[[]] Obtaining a sample of the audio data having a duration in the range of 2 - 10 seconds.
8. A computer-implemented method for securely synthesizing speech in a speaker's natural voice during a language translation communication session, the method comprising:[[]] During the language translation communication session:[[]] Periodically replace, at a fixed interval, the voice profile for the speaker by: i) obtaining a sample of the audio data from a stream of the audio data received from the speaker's device during the fixed interval, the stream of the audio data representing speech in the first language from the speaker, ii) generating a voice profile for the speaker based on the sample of the audio data, and iii) writing the voice profile of the speaker to a memory storage device; Process the stream of the audio data received from the speaker's device by generating a stream of the audio data for transmission to the participant in the communication session, the stream of the audio data for transmission to the participant representing synthesized speech that has been translated from the first language to the second language, the synthesized speech being derived using the most recently generated voice profile of the speaker stored in the memory storage device; And Transmit the stream of the audio data representing the synthesized speech in the second language to the device of the participant for playback.
9. The computer-implemented method according to claim 8, further comprising:[[]] In response to detecting the termination of the language translation communication session, delete the voice profile of the speaker from the memory storage device.
10. The computer-implemented method according to claim 8, wherein, Writing the voice profile of the speaker to the memory storage device includes overwriting an instance of the voice profile of the speaker previously written to the memory storage device.
11. The computer-implemented method according to claim 10, wherein, The language translation communication session is facilitated by an application running on the device of the participant, the application having a user interface for presentation on a display of the device, the method further comprising:[[]] Before overwriting an instance of the voice profile of the speaker previously written to the memory storage device:[[]] Compare the voice profile of the speaker with the instance of the voice profile of the speaker previously written to the memory storage device; and Determine that a difference between the voice profile of the speaker and an instance of the voice profile of the speaker previously written to the memory storage device exceeds a threshold based on the comparison; and In response to the determination, cause a visual indicator to be presented via the user interface of the application running on the device of the participant, wherein the visual indicator indicates detection of a significant change to the voice profile of the speaker.
12. The computer-implemented method according to claim 9, wherein, Processing the stream of audio data received from the device of the speaker by generating a stream of audio data for transmission to a device of a participant in the communication session includes: Performing speech recognition on the stream of audio data received from the device of the speaker to generate a symbolic language representation of the speech in the second language for the speech in the first language; and Using a speech synthesizer, and using the symbolic language representation of the speech in the second language as input, generate a stream of the audio data for transmission to the device of the participant in the communication session.
13. The computer-implemented method according to claim 12, wherein, The symbolic language representation of the speech in the second language includes: Text in the second language; or A speech transcription in the second language.
14. The computer-implemented method according to claim 8, wherein, Obtaining a sample of audio data from the stream of audio data received from the device of the speaker includes: Obtaining a sample of audio data having a duration in the range of 2 - 10 seconds.
15. A computer-implemented system for securely synthesizing speech in a speaker's natural voice during a language translation communication session, the system comprising: During the language translation communication session: A unit for periodically replacing, at a fixed interval, a unit of the voice profile of the speaker by: i) obtaining a sample of audio data from a stream of audio data received from the device of the speaker during the fixed interval, the stream of audio data representing speech in a first language from the speaker, ii) generating a voice profile for the speaker based on the sample of audio data, and iii) writing the voice profile of the speaker to a memory storage device; A unit for processing the stream of audio data received from the device of the speaker by generating a stream of audio data for transmission to a participant in the communication session, the stream of audio data for transmission to the participant representing synthesized speech that has been translated from a first language to a second language and that is derived using the most recently generated voice profile of the speaker stored in the memory storage device; And A unit for transmitting the stream of audio data representing the synthesized speech in the second language to the device of the participant for playback.