Automatic Voice-Over Generation
The method automates voice-over generation for advertising campaigns by using ad campaign attributes to create tailored synthesized speech, addressing the inefficiencies and costs of human-driven approaches, enhancing ad effectiveness.
Patent Information
- Application Number
- JP2024507124
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-08-07
- Filing Date
- 2022-07-20
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-07-20
AI Technical Summary
Current voice-over generation methods for advertising campaigns are time-consuming and costly due to the need for human voice actors and lack of automated systems that can generate synthesized speech with characteristics tailored to campaign attributes.
A computer-implemented method and system that automatically generates voice-over scripts and synthesized speech based on ad campaign attributes, using a text-to-speech system to overlay the speech onto advertisements, with features like prosody and speaker embeddings to match target demographics and regions.
Enables efficient and cost-effective voice-over generation that accurately represents campaign attributes, reducing the need for human actors and improving ad effectiveness by aligning speech characteristics with target audiences.
Smart Images

Figure 0007763327000001 
Figure 0007763327000002 
Figure 0007763327000003
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates to automatic voice-over generation. [Background technology]
[0002] Voice-over generation is the process of generating audible audio for an audio or video advertising campaign that explains and / or provides additional context to the advertising campaign's viewers. Voice-over generation has grown in popularity in recent years because adding voice-over to an advertising campaign has been proven to significantly improve the effectiveness of the campaign. A key aspect of voice-over generation is deciding what to say during the voice-over and how it should sound to appeal to the target customers viewing the advertising campaign. However, deciding what to say and how to say it is a significant task for many companies and advertising agencies, as it can be time-consuming and expensive to hire the right voice actors to speak the voice-over audio used in the advertising campaign. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] U.S. Patent Application No. 16 / 867,427 [Non-patent literature]
[0004] [Non-Patent Document 1] van den Oord, Parallel WaveNet: Fast High-Fidelity Speech Synthesis, available at https: / / arxiv.org / pdf / 1711.10433.pdf Summary of the Invention [Means for solving the problem]
[0005] One aspect of the present disclosure provides a computer-implemented method that, when executed on data processing hardware, causes the data processing hardware to perform operations. The operations include receiving a voice-over request to generate synthesized voice-over speech for a targeted advertisement having one or more ad campaign attributes. The operations also include generating a voice-over script including a sequence of text for the synthesized voice-over speech based on the one or more ad campaign attributes. The operations also include generating the synthesized voice-over speech using a text-to-speech (TTS) system. The TTS system is configured to receive as input the sequence of text for the voice-over script and generate as output synthesized voice-over speech having speech characteristics specified by the target TTS vertical. The operations also include overlaying the synthesized voice-over speech on the targeted advertisement.
[0006] Implementations of the present disclosure may include one or more of any of the following features: In some implementations, the operations further include selecting a target TTS vertical based on one or more ad campaign attributes. The speech characteristics specified by the target TTS vertical may include at least one of an utterance embedding that specifies prosody / style information to be conveyed by the synthesized voice-over speech and a speaker embedding that specifies phonetic characteristics of the synthesized voice-over speech.
[0007] Optionally, the ad campaign attributes may include at least one of a headline, a call to action, a geographic region, a language, or an audience demographic. In some examples, the sequence of text of the voice-over script includes one or more words, and overlaying the synthesized voice-over speech on the targeted advertisement includes, if the target advertisement has a playback time including the respective timestamps, determining respective timestamps at which the one or more words of the voice-over script should be spoken by the synthesized voice-over speech, and aligning the synthesized voice-over speech with the targeted advertisement such that segments of the synthesized voice-over speech corresponding to the one or more words of the voice-over script occur at the respective timestamps of the target advertisement.
[0008] In some implementations, generating a voice-over script for the synthesized voice-over speech may include identifying one or more words associated with the ad campaign having one or more ad campaign attributes by identifying phrases from a uniform resource locator (URL) of a landing page associated with the ad campaign and ranking each of the identified phrases from the landing page URL. The rank of each of the phrases corresponds to the likelihood that the respective phrase is associated with one or more ad campaign attributes of the ad campaign. Here, the operations may further include determining whether the rank of the identified phrase meets a threshold. Generating the voice-over script may occur if the rank of one of the identified phrases meets the threshold and the sequence of text in the voice-over script represents the identified phrase that meets the threshold.
[0009] In these implementations, in response to determining that the rank of the identified phrase does not meet the threshold, the operations further include accessing a corpus of advertisements associated with different advertisement campaigns, each advertisement associated with a respective advertisement campaign having a respective voice-over script and a set of advertisement campaign attributes; identifying one or more advertisements from the corpus of advertisements having advertisement campaign attributes similar to one or more advertisement campaign attributes of the voice-over request; and generating a voice-over script for the synthesized voice-over speech based on the respective voice-over scripts of the identified one or more advertisements having advertisement campaign attributes similar to the one or more advertisement campaign attributes of the voice-over request.
[0010] In some examples, the TTS system includes a TTS model configured to convert a sequence of text of a voice-over script into a corresponding synthesized speech representation of the voice-over script, and a TTS synthesizer configured to generate synthesized voice-over speech from the synthesized speech representation output from the TTS model. Optionally, one or more ad campaign attributes may be associated with the human-created ad campaign.
[0011] Another aspect of the present disclosure provides a system including data processing hardware and memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving a voice-over request to generate synthesized voice-over speech for a targeted advertisement having one or more ad campaign attributes. The operations also include generating a voice-over script including a sequence of text for the synthesized voice-over speech based on the one or more ad campaign attributes. The operations also include generating the synthesized voice-over speech using a text-to-speech (TTS) system. The TTS system is configured to receive as input the sequence of text for the voice-over script and generate as output synthesized voice-over speech having speech characteristics specified by the target TTS vertical. The operations also include overlaying the synthesized voice-over speech on the targeted advertisement.
[0012] Implementations of the present disclosure may include one or more of any of the following features: In some implementations, the operations further include selecting a target TTS vertical based on one or more ad campaign attributes. The speech characteristics specified by the target TTS vertical may include at least one of an utterance embedding that specifies prosody / style information to be conveyed by the synthesized voice-over speech and a speaker embedding that specifies phonetic characteristics of the synthesized voice-over speech.
[0013] Optionally, the ad campaign attributes may include at least one of a headline, a call to action, a geographic region, a language, or an audience demographic. In some examples, the sequence of text of the voice-over script includes one or more words, and overlaying the synthesized voice-over speech on the targeted advertisement includes, where the target advertisement has a playback time including the respective timestamps, determining respective timestamps at which the one or more words of the voice-over script should be spoken by the synthesized voice-over speech, and aligning the synthesized voice-over speech with the targeted advertisement such that segments of the synthesized voice-over speech corresponding to the one or more words of the voice-over script occur at the respective timestamps of the target advertisement.
[0014] In some implementations, generating a voice-over script for the synthesized voice-over speech may include identifying one or more words associated with the ad campaign having one or more ad campaign attributes by identifying phrases from a uniform resource locator (URL) of a landing page associated with the ad campaign and ranking each of the identified phrases from the landing page URL. The rank of each of the phrases corresponds to the likelihood that the respective phrase is associated with one or more ad campaign attributes of the ad campaign. Here, the operation may further include determining whether the rank of the identified phrase meets a threshold. Generating the voice-over script may occur if the rank of one of the identified phrases meets the threshold and the sequence of text in the voice-over script represents the identified phrase that meets the threshold.
[0015] In these implementations, in response to determining that the rank of the identified phrase does not meet the threshold, the operations further include accessing a corpus of advertisements associated with different advertisement campaigns, each advertisement associated with a respective advertisement campaign having a respective voice-over script and a set of advertisement campaign attributes; identifying one or more advertisements from the corpus of advertisements having advertisement campaign attributes similar to the one or more advertisement campaign attributes of the voice-over request; and generating a voice-over script for the synthesized voice-over speech based on the respective voice-over scripts of the identified one or more advertisements having advertisement campaign attributes similar to the one or more advertisement campaign attributes of the voice-over request.
[0016] In some examples, the TTS system includes a TTS model configured to convert a sequence of text of a voice-over script into a corresponding synthesized speech representation of the voice-over script, and a TTS synthesizer configured to generate synthesized voice-over speech from the synthesized speech representation output from the TTS model. Optionally, one or more ad campaign attributes may be associated with the human-created ad campaign.
[0017] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a schematic diagram of an example system for automatic voice-over generation. [Figure 2A] FIG. 1 is a schematic diagram of an exemplary script generator. [Figure 2B] FIG. 1 is a schematic diagram of an exemplary script generator. [Figure 3] 1 is a schematic diagram of an exemplary text-to-speech system. [Figure 4A] FIG. 2 is a schematic diagram of an exemplary audio overlay module. [Figure 4B] FIG. 2 is a schematic diagram of an exemplary audio overlay module. [Figure 5] 1 is a flowchart of an exemplary arrangement of method operations for performing automatic voice-over generation. [Figure 6] FIG. 1 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0019] Like reference numbers in the various drawings indicate like elements.
[0020] An advertising campaign generally refers to an advertising strategy designed to promote brand awareness, increase sales, and / or improve communication within one or more markets. Advertising campaigns often include goals or objectives centered around a brand or product. Some goals include acquiring or expanding customers, promoting current products, and / or launching new products. The design or strategy of an advertising campaign may also seek to associate a particular emotion or feeling with the brand or product. For example, an advertising campaign may market a new toy as fun, exciting, and playful, while a new pair of women's work boots may be marketed as rugged, outdoorsy, practical, and generally strong. In this sense, an advertising campaign includes one or more attributes that characterize the advertising campaign's strategy. These attributes may specify properties of the ad campaign such as the target audience (e.g., demographic details such as age, gender, social class, marital status, education level, interests, habits, and / or hobbies), the type of ad content (e.g., static ads such as photos or images, or dynamic ads such as videos), the content of one or more ads associated with the ad campaign (i.e., the content of the ads), the form factor of the ads (e.g., videos embedded around ads on a web page vs. commercials within the primary content), metrics of the ad campaign, and / or information about where the ad content should be placed / hosted.
[0021] As digital marketing continues to expand, advertising campaigns have become more sophisticated to understand the social engineering of advertising. Based on this understanding, it has been observed that advertising campaigns that include well-crafted (i.e., scripted) audio in their advertisements generally have a greater effect on their target audiences when compared to advertisements without audio content. Accordingly, advertising campaigns often seek to include voice-over production techniques to associate a voice or narrative with media content, particularly media content without audible sound. In this regard, advertising agencies and entities running advertising campaigns are increasingly seeking to add voice-overs to advertisements associated with their advertising campaigns.
[0022] Unfortunately, for a voice-over to be effective for an advertising campaign, it must adequately describe the advertising campaign's product, service, or company in a voice that is representative of the target consumer. In other words, the voice-over must reflect the objectives, goals, goals, and / or attributes associated with a particular product or brand advertising campaign. Therefore, the voice-over production process typically involves multiple iterations to generate a voice-over script that is carefully tailored (or adequately described) for the advertising campaign. With a carefully selected voice-over script, the voice for the voice-over is used to represent key characteristics of the advertising campaign. That is, the voice for the voice-over is selected to have a prosody / speech style (e.g., intonation, pitch, rhythm, etc.) that corresponds to one or more characteristics of the advertising campaign. For example, returning to the work boot advertisement, a female voice actress may be selected to speak with a slow, deliberate, and confident speech rhythm to create intensity.
[0023] Furthermore, if a product's advertising campaign spans multiple regions or countries, different voice actors may be required to have voice characteristics representative of the advertising campaign. An advertisement for work boots airing in the United States may include a voice-over script spoken with voice characteristics representative of an American (i.e., American English), while an advertising campaign for the same product airing in the United Kingdom may include a voice-over script spoken with voice characteristics representative of a British (i.e., British English). This means that the voice-over script may be spoken with voice characteristics representative of different languages, different genders, and / or different accents / dialects, as needed, to target consumers of the advertising campaign. Due to these varying requirements of advertising campaigns, generating voice-overs to effectively target audiences can quickly become complex and costly in terms of voice-over script generation and voice-over generation from the voice-over scripts. For example, an advertising campaign spanning multiple countries may require voice-over scripts in multiple languages and multiple voice-over actors / actresses who speak the particular languages.
[0024] Some current approaches attempt to address these voice-over generation issues by synthesizing voice-overs for advertising campaigns. Using synthetic speech has the advantage of not relying on human voice actors to provide voice-over speech. However, these current approaches do not base the synthesized speech's voice characteristics on advertising campaign attributes. That is, current voice-over approaches do not use speech synthesizers to generate synthesized speech according to speech characteristics that specifically represent one or more attributes of an advertising campaign. Furthermore, even when current approaches generate synthesized speech without speech characteristics specific to an advertising campaign, these implementations generally rely on receiving a voice-over script or speech transcript to generate the synthesized speech. In other words, voice-over scripts are not automatically generated (i.e., machine / computer-generated) but are instead created by advertising professionals or entities associated with a brand or product (i.e., human-generated). This means that even when using synthesized speech, voice-over script generation can still cause a bottleneck in the overall voice-over generation process.
[0025] Implementations herein are directed to a method for automatic voice-over generation. The method executes a voice-over generation model that receives a voice-over request to generate synthesized voice-over speech for a targeted advertisement having one or more ad campaign attributes. The ad campaign attributes may be computer-generated or provided by a user of the voice-over request. These campaign attributes provide context to the voice-over generation model in terms of what to say in the voice-over script and how to say it. The voice-over generation model generates a voice-over script to convert into synthesized voice-over speech having the speech characteristics of the one or more ad campaign attributes. That is, the voice-over generation model generates the voice-over script and determines the speech characteristics that the resulting synthesized voice-over speech should convey to specifically target an intended group of consumers based on the ad campaign attributes. The voice-over generation model then overlays the synthesized voice-over speech onto the targeted advertisement. An entity managing the ad campaign (e.g., an entity generating the voice-over request) can then deploy the targeted advertisement with the synthesized voice-over speech overlay to the target audience. As used herein, a voice-over request may include an explicit request from a user to generate a voice-over script to convert into synthesized voice-over speech for inclusion in a targeted advertisement, or a computing device may automatically generate a voice-over request upon detecting that a particular targeted advertisement does not include voice-over speech.
[0026] 1 , in some implementations, an exemplary system 100 includes one or more user devices 110 that communicate with a remote system 130 via a network 130. The user devices 110 may correspond to any computing device, such as a desktop workstation, a laptop workstation, or a mobile device (i.e., a smartphone). The user devices 110 include computing resources 112 (e.g., data processing hardware) and / or storage resources 114 (e.g., memory hardware). The remote system 130 is configured to receive voice-over requests 102 from the user devices 110 associated with respective users 10 via the network 120. The remote system 130 may be multiple computers or a distributed system (e.g., a cloud environment) having scalable / elastic resources 134, including computing resources 134 (e.g., data processing hardware) and / or storage resources 136 (e.g., memory hardware).
[0027] The voice-over request 102 requests that the voice-over generator 140 generate synthesized voice-over speech 352 for the targeted advertisement 104. Here, synthesized voice-over speech 352 refers to machine-generated speech generated from a voice-over script overlaid as audio onto media content (e.g., the targeted advertisement 104). The targeted advertisement 104 may be an audio or video advertisement that does not include any voice-over or includes only a portion of the voice-over of the targeted advertisement 104. While the examples herein are directed to automatically generating synthesized voice-over speech 352 to be overlaid on a targeted advertisement, implementations herein are equally applicable to automatically generating synthesized voice-over speech 352 for other types of media content, such as, but not limited to, documentaries, musical performances, and educational videos, to name a few. The targeted advertisement 104 is associated with an ad campaign having one or more ad campaign attributes 106, which are assigned to the targeted advertisement 104. The ad campaign attributes 106 may provide context for the targeted advertisement 104 and the target consumers (i.e., target audience) of the targeted advertisement 104. Accordingly, the voice-over generation model 140 generates synthesized voice-over speech 352 based on one or more ad campaign attributes 106. In some examples, the user device 110 or the remote system 130 executes a voice-over detector 180 configured to detect whether the targeted advertisement 104 (i.e., audio-video data) includes voice-over speech. In these examples, the voice-over detector 180 may output an indication when voice-over speech / content is not detected (e.g., not present) from the targeted advertisement 104. The output indication may serve as a suggestion prompting the user 10 to provide a voice-over request 102 to the voice-over generation model 140.Alternatively, the voice-over detector 180 may automatically generate and provide a voice-over request 102 to the voice-over generator 140 to request the voice-over generator 140 to generate synthesized voice-over speech 352 for the targeted advertisement 104.
[0028] An ad campaign may be configured by an advertiser or some other ad management entity (e.g., user 10 as shown). The advertiser or ad management entity may provide ad campaign attributes 106 (also referred to as attributes 106) for the ad campaign when the campaign is configured, or the ad campaign system may infer or automatically generate one or more ad campaign attributes 106 based on ad information provided to the ad campaign system by the advertiser / ad management entity. That is, ad campaign attributes 106 may be associated with a human-created or computer-generated ad campaign. In some examples, the user (e.g., user 10) that generates the voice-over request 102 is the same entity that coordinates (e.g., configures) the ad campaign (and attributes 106). In other examples, the ad campaign system may automatically generate the voice-over request. For example, the ad campaign system (e.g., in conjunction with voice-over generator 140) is configured to detect when an advertisement associated with an ad campaign lacks voice-over content and to provide the entity responsible for the ad campaign with the option of generating synthesized voice-over speech. In some implementations, the advertising campaign system (e.g., voice-over generator 140) automatically generates synthesized voice-over speech 352 for a particular advertisement (e.g., an advertisement lacking voice-over content) and recommends the automatically generated synthesized voice-over speech to the entity responsible for the advertising campaign (e.g., user 10).
[0029] The ad campaign attributes 106 may include, but are not limited to, a headline, a call-to-action, a geographic region, a language, or an audience demographic. A headline may include a slogan or saying related to the brand (e.g., company) or product of the targeted ad 104, such as "Visit ABC123.com to get coupons" or "Everyday Performance Apparel." A call-to-action may include an action that the target consumer of the ad should take. Examples of calls-to-action include "Shop Now," "Download the app now," or "Click the link to learn more." A geographic region may include a target area of the ad campaign, such as a particular country, state, city, or region. A language may include the intended language of the targeted ad 104. Audience demographics may include the target consumer (i.e., the target audience) of the targeted ad 104. For example, an audience demographic may be men aged 18-30 or women aged 40-62. Audience demographics may provide key characteristics about the target consumer that the advertising entity of the targeted ad 104 is attempting to capture with the targeted ad 104. The ad campaign attributes 106 may also include a landing page uniform resource locator (URL), product type, and / or industry associated with the content (e.g., brand or product) of the targeted ad 104 .
[0030] In some implementations, the voice-over request 102 requests synthesized voice-over speech 352 having speech characteristics 304 that represent one or more ad campaign attributes 106 of the targeted advertisement 104. The voice-over generator 140 may be configured to generate the synthesized voice-over speech 352 for the targeted advertisement 104 of the voice-over request 102 by executing on the remote system 130, the user device 110, or some combination thereof. More specifically, the voice-over generator 140 may include a script generator 200, a text-to-speech (TTS) system 300, and a speech overlay module 400. The script generator 200 is configured to generate a voice-over script 252 (i.e., a computer / machine-generated voice-over script 252) for the targeted advertisement 104. Here, when the script generator 200 generates the voice-over script 252, the voice-over script 252 may be completely machine-generated without human input during script generation. The voice-over script 252 includes a sequence of text that represents content to be spoken as synthesized voice-over speech during the targeted advertisement 104. In particular, the voice-over script 252 includes a text representation of one or more words to be spoken as synthesized voice-over speech during the targeted advertisement 104. To automatically generate the voice-over script 252 associated with the targeted advertisement 104, the script generator 200 generates a sequence of text that represents (i.e., characterizes) one or more ad campaign attributes 106. That is, the script generator 200 generates the voice-over script 252 based on the one or more ad campaign attributes 106 such that the voice-over script 252 is associated with the targeted advertisement 104. Once the script generator 200 generates the voice-over script 252, the script generator 200 communicates the voice-over script 252 to the TTS system 300.
[0031] The script generator 200 may implement one or more language models for automatically generating a voice-over script 252 based on one or more ad campaign attributes 106. In some implementations, the script generator 200 includes one or more language models trained with captions of training voice-over speech extracted from a corpus of existing advertisements (e.g., training advertisements) 208, 208a-n (FIG. 2B). In particular, the captions serve as reference voice-over scripts 252R (FIG. 2B). In these implementations, advertisements 208 in the advertisement corpus may be associated with corresponding reference campaign attributes 106R (FIG. 2B), which may further be used as labels for tuning the language model during training.
[0032] The TTS system 300 is configured to convert the voice-over script 252 into a corresponding synthesized voice-over speech 352 having speech characteristics 304 specified by a target TTS vertical 312 representing the ad campaign attributes 106. That is, the TTS system 300 determines how to say the voice-over script 252 based on the ad campaign attributes 106 and / or the voice-over script 252. The target TTS vertical 312 may convey a particular “character” of the voice-over speech 352 that best suits the target advertisement 104. Thus, the TTS system 300 may select the target TTS vertical 312 based on the ad type / vertical associated with the target advertisement 104. The TTS system 300 may use the ad campaign attributes 106 to identify the ad type / vertical, thereby selecting an appropriate target TTS vertical 312 associated therewith and its type corresponding to the speech characteristics 304. The speech characteristics 304 specified by the target TTS vertical 312 may include many linguistic elements not provided by the text input for generating synthesized speech. A subset of these linguistic elements, collectively referred to as prosody, may include intonation (variations in pitch), stress (stressed and unstressed syllables), duration, volume, tone, rhythm, and speaking style. Prosody may indicate the emotional state of the speech, the form of the speech (e.g., statement, question, command, etc.), the presence of sarcasm or sarcasm in the speech, uncertainty in the knowledge of the speech, or other linguistic elements that cannot be encoded by the grammar or lexical choices of the input text. Linguistic elements may also include the accent, dialect, and / or language of a particular speaker in a geographic region. The TTS system 300 sends the synthesized voice-over speech 352 to the speech overlay module 400.
[0033] The speech overlay module 400 is configured to overlay the synthesized voice-over speech 352 generated by the TTS system 300 onto the targeted advertisement 104 to generate a voice-over advertisement 450. Here, the voice-over advertisement 450 includes a targeted advertisement 104 (i.e., an audio advertisement or a video advertisement) with synthesized voice-over speech 352 in a targeted TTS vertical 312 that represents an advertising campaign attribute 106. When the speech overlay module 400 overlays the synthesized voice-over speech 352 onto the targeted advertisement 104, the speech overlay module 400 may be configured to align the synthesized voice-over speech 352 with a particular portion or portions of the targeted advertisement 104. For example, the synthesized voice-over speech 352 may include 10 seconds of speech, and the targeted advertisement 104 may be 20 seconds long. Here, the speech overlay module 400 determines when 10 seconds of synthesized voice-over speech 352 will be spoken during the 20 seconds of the targeted advertisement 104. The voice-over generator 140 provides the voice-over advertisement 450 to the entity or system responsible for implementing the advertising campaign. For example, as shown in FIG. 1 , the voice-over generator 140 communicates the voice-over advertisement 450 to a user 10 associated with the user device 110.
[0034] In some examples, the voice-over request 102 includes only the targeted advertisement 104 and one or more ad campaign attributes 106. Accordingly, the script generator 200 is configured to determine / generate a voice-over script 252 for converting the voice-over advertisement 450 into corresponding synthesized voice-over speech 352 based on the ad campaign attributes 106. Referring now to FIG. 2A , in some implementations, the exemplary script generator 200, 200a includes a scraper 210, a classifier 220, and a text generator 250. In some cases, the ad campaign attributes 106 of the targeted advertisement 104 include a landing page uniform resource locator (URL) 204. The landing page URL 204 may be any web page that includes content associated with the targeted advertisement 104 (e.g., content associated with the company, brand, or product of the targeted advertisement 104). For example, the targeted advertisement 104 may be a video advertisement that includes a landing page URL 204 linked to the homepage of the company of the targeted advertisement 104, a webpage containing detailed information about the product of the targeted advertisement 104, or any other webpage associated with the source, brand, and / or product of the targeted advertisement 104.
[0035] The script generator 200 may communicate with an online database 202 to access landing page URLs 204 of the targeted advertisements 104. In particular, the scraper 210 receives the targeted advertisements 104 and one or more ad campaign attributes 106 and obtains the landing page URLs 204 of the targeted advertisements 104 by accessing the online database 202. Once the scraper 210 obtains the landing page URLs 204, the scraper 210 is configured to parse the content of the landing page URLs 204 to identify phrases 212. That is, the landing page URLs 204 include various content, such as phrases, graphics, videos, links, etc., and the scraper 210 parses the content to identify phrases 212 from among other content included in the landing page URLs 204. The identified phrases 212 may include a single word, one or more words, punctuation marks, symbols, and / or numbers. In some cases, the landing page URL 204 is associated with the company, brand, and / or product of the targeted advertisement 104, and therefore the landing page URL 204 includes phrases that may be included in the voice-over script 252. In other words, phrases from the landing page URL 204 may be candidate phrases for potential inclusion in the voice-over script 252.
[0036] For example, a targeted advertisement 104 for a sportswear company may include a landing page URL 204 linked to the sportswear company's home page. A scraper 210 may access an online database 202 to obtain the landing page URL 204 for the sportswear company. The scraper 210 then parses the content of the landing page URL 204 and identifies one or more phrases 212 from the landing page URL 204, including "Buy Now," "Terms of Use," "20% Off," "The Style You Need Now," and "Shipping Information." The scraper 210 sends each of the identified phrases 212 to a classifier 220.
[0037] While one or more of the identified phrases 212 identified by the scraper 210 may be related to the targeted advertisement 104, other phrases 212 identified by the scraper 210 are not related to the targeted advertisement 104. Thus, the script generator 200a uses only the identified keywords related to the targeted advertisement 104 to generate the voice-over script 252. Thus, the classifier 220 is configured to classify which of the identified phrases 212 are key phrases 212, 212K of the targeted advertisement 104 based on the ad campaign attributes 106 of the targeted advertisement 104. The classifier 220 determines whether the identified phrase 212 is a key phrase 212K by ranking each of the identified phrases 212 from the landing page URL 204. Here, the rank of each of the identified phrases 212 corresponds to the likelihood that the identified phrase 212 is related to the ad campaign attributes 106 of the targeted advertisement 104 (e.g., the likelihood that the identified phrase 212 is a key phrase 212K).
[0038] Continuing with the above example, the classifier 220 ranks each of the identified phrases 212 received from the scraper 210 using the ad campaign attributes 106 of the targeted advertisement 104: "Buy Now," "Terms and Conditions," "20% Off," "Styles You Need Now," and "Shipping Information." Here, the ad campaign attributes 106 include apparel companies, athletics, and audience demographics of people ages 12 to 40. In this example, the classifier 220 may rank each of the identified phrases 212 from 0, indicating that the identified phrase 212 is least likely to be associated with the targeted advertisement 104, to 1, indicating that the identified phrase 212 is most likely to be associated with the targeted advertisement 104. The classifier 220 ranks "Buy Now" with a probability of 0.85, "Terms and Conditions" with a probability of 0.3, "20% Off" with a probability of 0.75, "Styles You Need Now" with a probability of 0.9, and "Shipping Information" with a probability of 0.35. The classifier 220 determines from the ad campaign attributes 106 that the targeted ad 104 is related to an advertisement for sportswear and that the identified phrases 212 "Buy Now," "The Style You Need Now," and "20% Off" are more likely to be related to the targeted ad 104 than the identified phrases 212 "Terms and Conditions" and "Shipping Information."
[0039] In some implementations, the classifier 220 classifies whether the identified phrases 212 are key phrases 212K by determining whether the rank associated with each identified phrase 212 satisfies a threshold. That is, the threshold indicates the minimum rank of the identified phrases 212 (e.g., the likelihood that the identified phrase 212 is relevant to the targeted advertisement 104) for the classifier 220 to classify the identified phrase 212 as a key phrase 212K. Thus, the classifier 220 determines whether each of the identified phrases 212 is a key phrase 212K and sends each of the key phrases 212K to the text generator 250. In this example, the classifier 220 has a threshold of 0.7 and determines that "Buy Now," "20% Off," and "The Style You Need Now" are key phrases 212K. The classifier 220 then sends the key phrases 212K to the text generator 250.
[0040] The text generator 250 is configured to generate the voice-over script 252 using one or more key phrases 212K received from the classifier 220. The text generator 250 may implement one or more language models to generate the voice-over script 252 using the one or more key phrases 212K. The one or more key phrases 212K may be “seed phrases” that the text generator 250 uses to generate the voice-over script 252. Here, the voice-over script 252 includes a sequence of text that represents one or more ad campaign attributes 106. The voice-over script 252 may include all words from the key phrases 212K, only some of the words from the key phrases 212K, or no words from the key phrases 212K. The text generator 250 generates the voice-over script 252 using the key phrases 212K and by generating additional words related to the key phrases 212K and / or the ad campaign attributes 106. In particular, if the text generator 250 only used the key phrases 212K to generate the voice-over script 252, the voice-over script 252 may sound incomplete and choppy. Therefore, the text generator 250 generates additional words related to the key phrases 212K and the advertising campaign attributes 106 to generate a complete and coherent voice-over script 252.
[0041] Continuing with this example, text generator 250 receives key phrases 212K "Buy now," "20% off," and "The styles you need now" and generates voice-over script 252 "Buy all your sportswear styles now and get an additional 20% off." Now, if text generator 250 simply used key phrases 212K, voice-over script 252 would be "Buy now the styles you need now and get 20% off," which would not be a coherent description of targeted advertisement 104. Therefore, text generator 250 uses key phrases 212K and ad campaign attributes 106 to generate additional words for the complete voice-over script 252.
[0042] In some implementations, the classifier 220 determines that none of the identified phrases 212 ranks meet the threshold. Here, none of the identified phrases 212 may meet the threshold because the targeted advertisement 104 does not include a landing page URL 204, the landing page URL does not include much text, and / or the landing page URL does not include text that is sufficiently relevant to the targeted advertisement 104 (e.g., related to the attributes 106 of the targeted advertisement 104). Here, the classifier 220 cannot send the key phrases 212K to the text generator 250 to generate the voice-over script 252. Notably, in these implementations, the script generator 200 must generate the entire voice-over script 252 using classification-free generation from the “seed values” (e.g., key phrases 212K) classified by the classifier 220.
[0043] Thus, in some cases, the script generator 200 needs to generate the voice-over script 252 without using the key phrase 212K from the landing page URL 204. Referring now to FIG. 2B , in some implementations, the exemplary script generator 200, 200b includes an advertisement database 206, an advertisement identifier 230, and a text generator 250. The advertisement database 206 includes a corpus of advertisements 208, 208a-n, each associated with a respective advertisement campaign including a reference voice-over script 252, 252R and a set of reference ad campaign attributes 106, 106R. For example, the advertisement database 206 corresponds to a YouTube® advertisement database, where each of the numerous advertisements in the advertisement database has a reference voice-over script 252R and a set of reference ad campaign attributes 106R. The reference voice-over script 252R may correspond to a caption for the corresponding voice-over speech in each advertisement in the corpus of advertisements 208. In some examples, an automatic speech recognition (ASR) system performs speech recognition on the voice-over speech to generate captions that correspond to the reference voice-over script 252R.
[0044] The script generator 200b is configured to determine a voice-over script 252 for the targeted advertisement 104 using a corpus of advertisements 208 retrieved from the advertisement database 206. In particular, the advertisement identifier 230 identifies one or more advertisements 208 that have reference ad campaign attributes 106R similar to the ad campaign attributes 106 of the target advertisement 104. The advertisement identifier 230 determines that advertisements 208 that include reference ad campaign attributes 106R similar to the ad campaign attributes 106 of the target advertisement 104 are likely to have a reference voice-over script 252R that represents the target advertisement 104. The advertisement identifier 230 uses the similarity score to identify advertisements 208 from the corpus of advertisements 208 as having reference ad campaign attributes 106R similar to one or more ad campaign attributes of the target advertisement 104. That is, the ad identifier 230 may assign each of the ads 208 a similarity score indicating the similarity between the ad campaign attributes 106 of the target ad 104 and the reference ad campaign attributes 106R of each ad 208 from the corpus of ads 208.
[0045] The advertisement identifier 230 may determine whether the similarity score of each advertisement 208 meets a similarity threshold. The similarity threshold may represent a minimum similarity required between the advertisement campaign attributes 106 and the reference advertisement campaign attributes 106R to generate a voice-over script 252 for the target advertisement 104 using the reference voice-over script 252R. If the similarity score of the advertisement 208 meets the similarity threshold, the advertisement identifier 230 sends the reference voice-over script 252R to the text generator 250. If the similarity score of the advertisement 208 does not meet the similarity threshold, the advertisement identifier 230 does not send the reference voice-over script 252R to the text generator 250. The advertisement identifier 230 may send multiple reference voice-over scripts 252R, 252Ra-n if multiple similarity scores meet the similarity threshold.
[0046] Using one or more reference voice-over scripts 252R, the text generator 250 generates a voice-over script 252 for the targeted advertisement 104. That is, the text generator 250 uses reference voice-over scripts 252R from existing advertisements 208 that have similar reference ad campaign attributes 106R as the ad campaign attributes 106 of the targeted advertisement 104 to generate a voice-over script 252 specific to the targeted advertisement 104.
[0047] In another additional implementation, as described above, the text generator 250 includes a language model trained on captions of training voice-over speech extracted from a corpus of advertisements 208, where each caption corresponds to a corresponding reference voice-over script 252R. Similarly, the reference ad campaign attributes 106R associated with each advertisement can be used as labels to adjust the language model during training. Thus, the text generator 250 implementing the trained language model can be configured to receive the ad campaign attributes 106 as input and generate the voice-over script 252 as output.
[0048] 3 , in some implementations, a TTS system 300 includes a TTS vertical selector 310, a TTS model 320, and a synthesizer 350 for outputting respective synthesized speech 352 having an intended prosody / style specified by a unique set of speech characteristics 304. The TTS vertical selector 310 is configured to select a target TTS vertical 312 that specifies the set of speech characteristics 304 for the resulting synthesized voice-over speech 352 based on one or more ad campaign attributes 106 associated with the target advertisement 104. The selection of the target TTS vertical 312 by the TTS vertical selector 310 may be further based on the voice-over script 252 output by the script generator 200.
[0049] As previously mentioned, the target TTS vertical 312 may convey a particular "character" of voice-over-speech 352 that best suits the target advertisement 104. In other words, the target TTS vertical 312 conveys a virtual voice actor that speaks in a speaking style / prosody typically associated with the advertisement type / vertical associated with the target advertisement. Thus, the TTS vertical selector 310 may select the target TTS vertical 312 based on the advertisement type / vertical associated with the target advertisement 104. The TTS system 300 may use the ad campaign attributes 106 to identify the advertisement type / vertical, thereby selecting an appropriate target TTS vertical 312 to associate therewith and its type corresponding to the speech characteristics 304. For example, advertisements in verticals related to technology, retail, and consumer packaged goods may be associated with the "Creator" TTS vertical 312, which specifies speech characteristics 304 with a youthful voice and an energetic, upbeat speech style / prosody, while advertisements in verticals related to healthcare and finance may be associated with the "Expert" TTS vertical 312, which specifies speech characteristics 304 with an adult voice and an informative, direct, confident, and deliberate speech style / prosody. As another example, advertisements in verticals related to automotive, consumer packaged goods, education and government, and media entertainment advertisements may be associated with the "Announcer" TTS vertical 312, which specifies speech characteristics 304 with a low-pitched adult voice and a speech style / prosody indicative of a direct, hard seller. Advertisements in beauty, fashion, travel, and wellness may further be associated with the premium TTS vertical 312, which specifies speech characteristics 304 with a relaxed, smooth, and soft speech style / prosody.
[0050] The TTS vertical selector 310 may be a heuristic-based or neural network-based model that selects the target TTS vertical 312 based on the ad campaign attributes 106. That is, the TTS vertical selector 310 may learn from correlations between the speech characteristics conveyed by the voice actor who spoke the voice-over speech in the reference advertisement 208, the corresponding reference voice-over script 252R (e.g., captions for the voice-over speech), and the ad types / verticals associated with the advertisements 208 in the corpus of reference advertisements 208.
[0051] As previously mentioned, the speech characteristics 304 specified by the target TTS vertical 312 may include many linguistic elements not provided or conveyed by the voice-over script 252 (i.e., the text input). A subset of these linguistic elements is collectively referred to as prosody and may include intonation (variations in pitch), stress (stressed and unstressed syllables), duration, volume, tone, rhythm, and speaking style. Prosody may indicate the emotional state of the speech, the form of the speech (e.g., statement, question, command, etc.), the presence of sarcasm or sarcasm in the speech, uncertainty in the knowledge of the speech, or other linguistic elements that cannot be encoded by the grammar or lexical choices of the input text. Linguistic elements may also include the accent, dialect, and / or language of a particular speaker of a geographic region.
[0052] The speech characteristics 304 specified by the target TTS vertical 312 may include at least one of an utterance embedding 304 a, an accent / dialect identifier 304 b, or a speaker embedding 304 c. The utterance embedding 304 a may include latent variables specifying the intended prosody / style, so that the TTS model 320 predicts a synthesized speech expression 322 that conveys the intended prosody / style specified by the utterance embedding 304 a. That is, for example, the utterance embedding 304 a may represent prosody / style information and / or accent / dialect information associated with the synthesized speech expression 322 that the TTS model 320 aims to replicate. For example, the speech embeddings 304a may represent an energetic and upbeat speaking style / prosody for the "creator" TTS vertical 312, an informative, direct, confident, and deliberate speaking style / prosody for the "expert" TTS vertical 312, speaking style / prosody information that conveys a direct, hard sell for the "announcer" category, and a relaxed, smooth, and soft style / prosody for the "luxury" TTS vertical 312. Other TTS verticals 312 that map to different speaking styles / prosody are also envisioned.
[0053] The accent / dialect identifier 304b indicates the target accent / dialect of the resulting synthesized voice-over speech 352. For example, the accent / dialect identifier 304b may identify a target British English or American English accent / dialect. In some examples, the accent / dialect identifier 304b identifies a fine-grained dialect, such as a Texas accent of American English, a Midwestern accent of American English, a South London accent of British English, or a Manchester accent of British English. The accent / dialect identifier 304b may further function as a language identifier if the TTS model 320 is multilingual, thereby enabling the TTS model 320 to be tuned to generate synthesized speech expressions 322 in multiple languages different from the voice-over script 252.
[0054] The speaker embedding 304c may indicate the phonetic characteristics of the target voice of the resulting synthesized voice-over speech 352. For example, the speaker embedding 304c may indicate whether the target voice is male / female, child / adult, low / high-pitched, etc. The speaker embedding 304c may convey the speaker identifier of the particular voice actor who spoke the reference utterances used to train the TTS system 300. Thus, the TTS system 300 may use the utterance embedding, accent / dialect identifier 304b, and speaker embedding 304c to clone the target speaker's voice in the synthesized voice-over speech 352 across different accents / dialects and speaking styles / prosody.
[0055] The TTS model 320 is configured to receive the speech characteristics 304 specified by the target TTS vertical 312 and convert the corresponding text of the voice-over script 252 into a synthesized speech representation 322. The synthesized speech representation 322 therefore conveys the speaking style / prosody associated with the “character” represented by the TTS vertical 312. The speaker embedding 304c may tune the TTS model 320 to replicate a vocal clone of any particular target voice with the same speaking style / prosody associated with the “character” represented by the TTS vertical 312. Similarly, the accent / dialect identifier 304b may tune the TTS model 320 to generate synthesized speech representations 322 in a variety of different accents / dialects and with the same speaking style / prosody. This scenario is particularly advantageous because it allows voice-over speech to be generated across different accents / dialects associated with the geographic region in which the targeted advertisement 104 is served. For example, a voice-over script 252 for a new car leasing advertising campaign can be used to generate synthesized voice-over speech 352 in a Midwestern accent for consumers viewing / listening to the targeted advertisement 104 in Michigan, and to generate synthesized voice-over speech 352 in a Texas accent for consumers viewing / listening to the targeted advertisement 104 in Texas.
[0056] The synthesized speech representation 322 output by the TTS model 320 may include a sequence of mel-frequency spectrograms. In some examples, the TTS model 320 includes a variational autoencoder-based (VAE-based) TTS model having a decoder portion configured to decode the voice-over script 252 into a corresponding synthesized speech representation 322 including speech units of pitch, energy, and phoneme duration (e.g., fixed-length frames (e.g., 5 milliseconds)) that convey prosodic / style information associated with the target TTS vertical 312 selected by the TTS vertical selector 310. Further details of the VAE-based TTS model are described with reference to U.S. Patent Application No. 16 / 867,427, filed May 5, 2020, the entire contents of which are incorporated by reference. The synthesized speech representation 322 may additionally or alternatively include vocoder parameters including mel-cepstral coefficients (MCEPs), aperiodic components, and phonetic components of each speech unit.
[0057] In the illustrated example, the TTS system 300 includes a single TTS model 320, which may be trained on existing advertisements in the corpus of advertisements 208, which may span multiple ad types / verticals such that the voice-over speech spans the different speaking styles / prosody associated with these verticals. Additionally or alternatively, the TTS model 320 may be trained to learn how to synthesize speech that matches reference utterances of human speech spoken by different voice actors. For example, a set of one or more voice actors may speak reference utterances from a reference voice-over script 252R having a speaking style / prosody associated with the “Announcer” TTS vertical, and the TTS model 320 and TTS synthesizer 350 may learn to generate synthesized voice-over speech 352 that matches the reference utterances. These reference utterances may be labeled with the associated TTS vertical. This process can be repeated with the same and / or different sets of voice actors for speaker-referenced utterances having speaking styles / prosody associated with other TTS verticals, e.g., the "Expert," "Advanced," and / or "Creator" verticals.
[0058] In additional implementations, the TTS system 300 includes multiple TTS models 320, each trained to generate synthesized speech expressions with a different speaking style / prosody. For example, the TTS system 300 may include a respective TTS model 320 for each target TTS vertical 312. An appropriate TTS model 320 can then be selected to convert the voice-over script 252 based on the target TTS vertical 312 selected by the TTS vertical selector 310. Similarly, the TTS system 300 may include multiple TTS models 320, each trained to generate synthesized speech expressions in a different voice and / or accent / dialect. In one example, the voice-over script 252 in a first language can be translated / transliterated into a second language and provided to a TTS model 320 trained to generate synthesized speech in the second language.
[0059] The TTS synthesizer 350 is configured to receive as input the synthesized speech representation 322 output by the TTS model 320 and to generate as output synthesized voice-over speech 352 that conveys the unique set of speech characteristics 304 specified by the target TTS vertical 312. The TTS synthesizer 350 may include a vocoder network for converting the Mel-frequency spectrogram sequence into a time-domain audio waveform. The time-domain audio waveform includes an audio waveform that defines the amplitude of an audio signal over time. The vocoder network may be any network configured to receive the Mel-frequency spectrogram and generate audio output samples based on the Mel-frequency spectrogram. For example, the vocoder network may be or be based on the parallel feed-forward neural network described in van den Oord, Parallel WaveNet: Fast High-Fidelity Speech Synthesis, available at https: / / arxiv.org / pdf / 1711.10433.pdf, incorporated herein by reference. Alternatively, TTS synthesizer 350 may be an autoregressive neural network. In some examples, TTS synthesizer 350 transforms fixed-length frames of pitch, energy, and phoneme duration represented by synthesized speech representation 322 to generate synthesized voice-over speech 352. For example, a unit selection module or a WaveNet module may use the frames to generate synthesized voice-over speech 352.
[0060] 4A and 4B, the speech overlay module 400 is configured to overlay synthesized voice-over speech 352 onto the targeted advertisement 104 to generate a voice-over advertisement 450. That is, the speech overlay module 400 determines when one or more words of the voice-over script 252 should be spoken by the synthesized voice-over speech 352. In this regard, the speech overlay module 400 may tailor the synthesized voice-over speech 352 to a particular playback time during the duration of the targeted advertisement 104.
[0061] In some configurations, the speech overlay module 400 may include a time stamper 410 and an aligner 420. The time stamper 410 is configured to determine respective time stamps T for a set of one or more words of the voice-over script 252. The time stamps T may represent a fixed time unit (e.g., 1 second, 0.5 seconds, 5 seconds, etc.). The time stamps T of each of the one or more words determine the sequence (e.g., order) in which the one or more words are spoken and / or the length for which the one or more words are spoken. The aligner 420 is configured to align the time stamps T of the one or more words of the voice-over script 252 with the play time time stamps P of the targeted advertisement 104. That is, the targeted advertisement 104 may include 9 seconds of play time with each play time time stamp P representing 1 second (i.e., P=9), and the time stamps T of the voice-over script 252 may include 5 seconds of speech with each time stamp representing 1 second (i.e., T=5). Here, the aligner 420 aligns the 5-second voice-over script 252 with the 9-second play time of the targeted advertisement 104. For example, the voice-over script 252 begins at the 3rd second of the play time of the targeted advertisement 104 and therefore ends at the 7th second of the play time of the targeted advertisement 104.
[0062] 4A , in some implementations, a time stamper 410 determines respective time stamps T for a set of one or more words of a voice-over script 252. That is, the time stamper 410 determines respective time stamps T for the start of the one or more words and the duration for which the one or more words are spoken. Here, the set of one or more words does not include pauses or silences between respective time steps T. Thus, the time stamper 410 determines only respective time stamps T for the start of the set of one or more words of the voice-over script 252 and the length for which the one or more words are spoken. For example, as shown in FIG. 4A , the time stamper 410 receives the set of one or more words of the voice-over script 252 and the synthesized voice-over speech 352 corresponding to "Download our new app now!" Here, the time stamper 410 determines that the set of one or more words begins at time stamp T=1 and its duration is 5 time stamps T (e.g., 5 seconds). Thus, the set of one or more words begins at timestamp T=1, there are no silences or pauses at any of the timestamps T between T=1 and T=5, and the set ends at timestamp T=5. Notably, the timestamp 410 determines only one respective timestamp T for the set of one or more words, rather than determining a timestamp T for each word of the phrase. In other words, instead of having to generate a timestamp T for each word of the phrase, the timestamp 410 may generate a single timestamp T that can be used as a key timestamp for overlaying the synthesized voice-over speech 352 onto the targeted advertisement 104. Here, the key timestamp may be the beginning, midpoint, or end of a segment of the synthesized voice-over speech 352, and the aligner 420 uses only the key timestamp to overlay the synthesized voice-over speech 352 onto the targeted advertisement 104 at the desired time.For example, the time stamper 410 determines the timestamp T of the midpoint, the word "new", and the aligner 420 aligns the word "new" to the midpoint (eg, 5 seconds) of the play time of the targeted advertisement 104.
[0063] The aligner 420 receives the synthesized voice-over speech 352 and the associated timestamp T from the time stamper 410. As shown in FIG. 4A , the playback time of the target advertisement 104 is 9 seconds, and each playback time step P is equal to 1 second (i.e., P=9). The aligner 420 uses each time step T to align the synthesized voice-over speech 352 to the playback time step P. The aligner 420 determines that the synthesized voice-over speech 352 begins at P=3 and ends at P=7. Therefore, the aligner 420 aligns each timestamp T=1 through T=5 of the synthesized voice-over speech 352 to the playback time steps P=3 through P=7. After the aligner 420 aligns the synthesized voice-over speech 352 to the target advertisement 104, the speech overlay module 400 generates the voice-over advertisement 450.
[0064] In some cases, the speech overlay module 400 can independently control the rhythm (e.g., timing) of each spoken word of the synthesized voice-over speech 352. That is, the synthesized voice-over speech 352 may not be spoken continuously and may include one or more pauses or silences between words. Referring now to FIG. 4B , in some implementations, the time stamper 410 independently determines a respective timestamp T for each of one or more words of the voice-over script 252. That is, there may be blank spaces between one or more words from the voice-over script 252. As shown in FIG. 4B , the time stamper 410 receives the synthesized voice-over speech 352 and the voice-over script 252 corresponding to "Buy one world-class luxury car now." The time stamper 410 independently determines a respective timestamp T for each of the one or more words. For example, the time stamper may determine that a pause is needed between the words "world class," "luxury," and "car" and the phrase "Buy one now." Thus, time stamper 410 determines the following timestamps: T=1 for "world," T=2 for "class," T=4 for "luxury," and T=6 for "car," T=8 for "buy," T=9 for "one," and T=10 for "right now." Time stamper 410 then also determines that there should be a pause or silence at timestamps T=3 and T=6.
[0065] The time stamper 410 sends the synthesized voice-over speech 352 and the corresponding timestamp T to the aligner 420. The aligner 420 is configured to align the timestamp T of one or more words of the synthesized voice-over speech 352 with the playback time P of the target advertisement 104. That is, the aligner 420 determines when the synthesized voice-over speech 352 is spoken during the playback time of the target advertisement 104. In some examples, the aligner 420 aligns with the start and end of the synthesized voice-over speech 352, but does not add or remove silences or pauses between one or more words of the synthesized voice-over speech 352 other than those the aligner 420 received as a communication from the time stamper 410. For example, the time stamper 410 determined that there is a timestamp of silence between “class” and “luxury” at timestamp T=3. Here, the aligner 420 cannot add or remove silence between “class” and “luxury.” Thus, the aligner 420 determines where the synthesized voice-over speech 352 is spoken, but does not affect the rhythm (eg, timing) of the synthesized speech as determined by the time stamper.
[0066] For example, the aligner 420 aligns nine timestamps T from the time stamper to twelve playtime timestamps P of the target advertisement 104. The aligner 420 determines that the first timestamp T=1 matches the second playtime timestamp P=2, and the last timestamp T=9 matches the tenth playtime timestamp P=10. Here, the aligner 420 aligns where the synthesized voice-over speech 352 is spoken during the playtime of the target advertisement 104 (e.g., starting at playtime timestamp P=2 and ending at playtime timestamp P=10) without affecting the rhythm of the synthesized voice-over speech 352 set by the time stamper 410. After the aligner 420 aligns the synthesized voice-over speech 352 to the playtime of the target advertisement 104, the speech overlay module 400 generates the voice-over advertisement 450 in response to the voice-over request 102. For example, the speech overlay module 400 or the voice-over generator 140 communicates a voice-over advertisement 450 to the user 10 associated with the voice-over request 102 .
[0067] 5 is a flowchart of an example configuration of operations for a method 500 for performing automatic voice-over generation. At operation 502, the method 500 includes receiving a voice-over request 102 for generating synthesized voice-over speech 352 for a targeted advertisement 104 having one or more ad campaign attributes 106. At operation 504, the method 500 includes generating a voice-over script 252 for the synthesized voice-over speech 352, the voice-over script 252 including a sequence of text, based on the one or more ad campaign attributes. At operation 506, the method 500 includes generating the synthesized voice-over speech 352 using a text-to-speech (TTS) system 300. The TTS system 300 is configured to receive as input the sequence of text for the voice-over script 252 and generate as output synthesized voice-over speech having speech characteristics specified by a target TTS vertical 312. At operation 508 , the method 500 includes overlaying the synthesized voice-over-speech 352 onto the targeted advertisement 104 .
[0068] 6 is a schematic diagram of an exemplary computing device 600 that may be used to implement the systems and methods described herein. Computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown here, their connections and relationships, and their functions are for illustrative purposes only and are not intended to limit the implementation of the invention(s) described and / or claimed herein.
[0069] Computing device 600 includes a processor 610 (e.g., data processing hardware 112, 134), memory 620 (e.g., memory hardware 114, 136), a storage device 630, a high-speed interface / controller 640 that connects to memory 620 and a high-speed expansion port 650, and a low-speed interface / controller 660 that connects to a low-speed bus 670 and storage device 630. Each of the components 610, 620, 630, 640, 650, and 660 are interconnected using various buses and may be mounted on a common motherboard or in other manners as needed. Processor 610 can process instructions for execution within computing device 600, including instructions stored in memory 620 or storage device 630, to display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 680, coupled to high-speed interface 640. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and types of memory, as needed. Also, multiple computing devices 600 may be connected, each providing a portion of the required operations (eg, as a bank of servers, a group of blade servers, or a multi-processor system).
[0070] The memory 620 stores information non-transiently within the computing device 600. The memory 620 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. The non-transient memory 620 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) used by the computing device 600. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks and tapes.
[0071] The storage device 630 can provide mass storage for the computing device 600. In some implementations, the storage device 630 is a computer-readable medium. In various different implementations, the storage device 630 may be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or devices in a storage area network or other configuration. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as memory 620, the storage device 630, or memory on the processor 610.
[0072] The high-speed controller 640 manages bandwidth-intensive operations of the computing device 600, while the low-speed controller 660 manages less bandwidth-intensive operations. Such assignments are merely examples. In some implementations, the high-speed controller 640 is coupled to memory 620, a display 680 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 650 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 660 is coupled to a storage device 630 and a low-speed expansion port 690. The low-speed expansion port 690, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled, for example, through a network adapter, to one or more input / output devices such as a keyboard, pointing device, scanner, or a networking device such as a switch or router.
[0073] The computing device 600, as shown in the drawing, can be implemented in many different forms. For example, it can be implemented as a standard server 600a, or multiple times within a group of such servers 600a, as a laptop computer 600b, or as part of a rack server system 600c.
[0074] Various implementations of the systems and techniques described herein can be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementations in one or more computer programs executable and / or interpretable on a programmable system that includes at least one programmable processor, which can be special-purpose or general-purpose, coupled to receive data and instructions from and send data and instructions to a storage system, at least one input device, and at least one output device.
[0075] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application," "app," or "program." Examples of applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0076] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0077] The processes and logic flows described herein can be executed by one or more programmable processors, also referred to as data processing hardware, which execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be executed by special-purpose logic circuitry, such as, for example, an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operatively coupled to receive data from or transfer data to them, or both. However, such devices are not required for a computer. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all types of non-volatile memory, media, and memory devices, including semiconductor memory devices such as EPROM, EEPROM, flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0078] To provide for user interaction, one or more aspects of the present disclosure can be implemented on a computer having a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and a pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, tactile feedback, etc., and input from the user can be received in any form, such as acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending and receiving documents to and from devices used by the user, for example, by sending web pages to a web browser on the user's client device in response to a request received from the web browser.
[0079] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]
[0080] 10 users 100 systems 102 Voice Over Request 104 Targeted Advertising 106 Ad Campaign Attributes 106R Reference Campaign Attributes 110 User Devices 112 computing resources 112 Data Processing Hardware 114 Storage Resources 114 Memory Hardware 120 Network 130 Remote Systems 134 Data Processing Hardware 134 computing resources 134 Scalable / Elastic Resources 136 Memory Hardware 136 Storage Resources 140 Voice-Over Generator 140 Voice-over-bass generation model 180 Voice Over Detector 200 Script Generator 200a Script Generator 200b Script Generator 202 Online Database 204 Landing Page Uniform Resource Locator (URL) 206 Advertising Database 208 Advertisement 208a~n Advertisement 210 Scraper 212 phrases 212 Key Phrases 212K Keyphrases 220 Classifier 230 Advertising Identifier 250 Text Generator 252 Voice-over scripts 252R Reference Voice-Over Script 252Ra~n Reference Voice-over Script 300 Text-to-Speech (TTS) Systems 304 Speech Characteristics 304a Utterance Embedding 304b Accent / Dialect Identifier 304c Speaker Embedding 310 Vertical Selector 312 Target TTS Vertical 320 TTS model 322 Synthesized Speech Expressions 350 TTS Synthesizer 352 Voice Over Speech 400 Speech Overlay Module 410 Time Stamper 420 Alaina 450 Voice-over Ads 500 ways 600 computing devices 600a Standard Server 600b laptop computer 600c Rack Server System 610 processor 620 memory 630 Storage Devices 640 High-Speed Interface / Controller 650 High-Speed Expansion Port 660 Low-Speed Interface / Controller 670 Slow Bus 680 display 690 Low-Speed Expansion Port
Claims
1. When executed on data processing hardware (134), the data processing hardware (134) receiving a voice-over request (102) for generating synthesized voice-over speech (352) for a targeted advertisement (104) having one or more ad campaign attributes (106); generating a voice-over script (252) for the synthesized voice-over speech (352) based on the one or more ad campaign attributes (106), the voice-over script (252) comprising a sequence of text; generating the synthesized voice-over speech (352) using a text-to-speech (TTS) system (300), the TTS system (300) comprising: receiving as input the sequence of text of the voice-over script (252); generating as output the synthesized voice-over speech (352), the synthesized voice-over speech (352) having speech characteristics (304) specified by a target TTS vertical (312); and overlaying the synthesized voice-over-speech (352) onto the targeted advertisement (104); A computer-implemented method (500) for performing operations comprising: the sequence of text of the voice-over script (252) comprises one or more words, and overlaying the synthesized voice-over speech (352) onto the targeted advertisement (104) comprises: determining respective timestamps at which the one or more words of the voice-over script (252) should be spoken by the synthesized voice-over speech (352), the targeted advertisement (104) having a playback time comprising the respective timestamps; aligning the synthesized voice-over speech (352) with the targeted advertisement (104) such that segments of the synthesized voice-over speech (352) corresponding to the one or more words of the voice-over script (252) occur at the respective timestamps of the targeted advertisement (104); A computer-implemented method (500) comprising:
2. 10. The computer-implemented method of claim 1, wherein the operations further comprise selecting the target TTS vertical based on the one or more ad campaign attributes.
3. 2. The computer-implemented method of claim 1, wherein the speech characteristics specified by the target TTS vertical comprise at least one of an utterance embedding that specifies prosodic / style information to be conveyed by the synthesized voice-over speech, an accent / dialect identifier that specifies the accent / dialect to be conveyed by the synthesized voice-over speech, and a speaker embedding that specifies voice characteristics of the synthesized voice-over speech.
4. The advertising campaign attributes (106) Headline, call to action, geographical region, language, or 10. The computer-implemented method of claim 1, further comprising at least one of: an audience demographic;
5. generating the voice-over script (252) for the synthesized voice-over speech (352), identifying a phrase (212) from a uniform resource locator (URL) (204) of a landing page associated with the advertising campaign; ranking each of the phrases (212) identified from the landing page URL (204), the rank of each of the phrases (212) corresponding to the likelihood that the respective phrase (212) is associated with the one or more ad campaign attributes of the ad campaign; 10. The computer-implemented method of claim 1, further comprising identifying one or more words associated with an ad campaign having the one or more ad campaign attributes by:
6. 6. The computer-implemented method (500) of claim 5, wherein the operations further comprise determining whether the rank of the identified phrase (212) meets a threshold.
7. generating the voice-over script (252) occurs if the rank of one of the identified phrases (212) meets the threshold; 7. The computer-implemented method of claim 6, wherein the sequence of text in the voice-over script represents the identified phrases that satisfy the threshold.
8. In response to determining that the rank of the identified phrase (212) does not meet the threshold, the operation: accessing a corpus of advertisements (208) associated with different advertisement campaigns, each advertisement (208) associated with a respective advertisement campaign having a respective voice-over script (252R) and a set of advertisement campaign attributes (106R); identifying one or more advertisements (208) from the corpus of advertisements (208) having similar ad campaign attributes (106R) as the one or more ad campaign attributes (106) of the voice-over request (102); generating the voice-over script (252) for the synthesized voice-over speech (352) based on the respective voice-over scripts (252R) of the identified one or more advertisements (208) having advertising campaign attributes (106R) similar to the one or more advertising campaign attributes (106) of the voice-over request (102); The computer-implemented method (500) of claim 6 further comprising:
9. The TTS system (300) a TTS model (320) configured to convert the sequence of text of the voice-over script (252) into a corresponding synthesized speech representation (322) of the voice-over script (252); a TTS synthesizer (350) configured to generate the synthesized voice-over speech (352) from the synthesized speech expression (322) output from the TTS model (320); The computer-implemented method (500) of claim 1, comprising:
10. 10. The computer-implemented method (500) of any one of claims 1 to 9, wherein the one or more ad campaign attributes (106) are associated with a human-created ad campaign.
11. data processing hardware (134); memory hardware (136) in communication with the data processing hardware (134), the memory hardware (136) being configured to, when executed by the data processing hardware (134), cause the data processing hardware (134) to: receiving a voice-over request (102) to generate synthesized voice-over speech (352) for a targeted advertisement (104) having one or more ad campaign attributes (106); generating a voice-over script (252) for the synthesized voice-over speech (352) based on the one or more ad campaign attributes (106), the voice-over script (252) comprising a sequence of text; generating the synthesized voice-over speech (352) using a text-to-speech (TTS) system (300), the TTS system (300) comprising: receiving as input the sequence of text of the voice-over script (252); generating as output the synthesized voice-over speech (352), the synthesized voice-over speech (352) having speech characteristics (304) specified by a target TTS vertical (312); and overlaying the synthesized voice-over-speech (352) onto the targeted advertisement (104); said memory hardware (136) storing instructions for performing operations comprising: A system (100) comprising: the sequence of text of the voice-over script (252) comprises one or more words, and overlaying the synthesized voice-over speech (352) onto the targeted advertisement (104); determining respective timestamps at which the one or more words of the voice-over script (252) should be spoken by the synthesized voice-over speech (352), wherein the targeted advertisement (104) has a playback time comprising the respective timestamps; aligning the synthesized voice-over speech (352) with the targeted advertisement (104) such that segments of the synthesized voice-over speech (352) corresponding to the one or more words of the voice-over script (252) occur at the respective timestamps of the targeted advertisement (104); A system (100) comprising:
12. 12. The system (100) of claim 11, wherein the operations further comprise selecting the target TTS vertical (312) based on the one or more ad campaign attributes (106).
13. 12. The system of claim 11, wherein the speech characteristics specified by the target TTS vertical comprise at least one of an utterance embedding that specifies prosodic / style information to be conveyed by the synthesized voice-over speech, an accent / dialect identifier that specifies the accent / dialect to be conveyed by the synthesized voice-over speech, and a speaker embedding that specifies voice characteristics of the synthesized voice-over speech.
14. The advertising campaign attributes (106) Headline, call to action, geographical region, language, or 12. The system (100) of claim 11, comprising at least one of an audience demographic.
15. generating the voice-over script (252) for the synthesized voice-over speech (352), Identifying a phrase (212) from a uniform resource locator (URL) (204) of a landing page associated with the advertising campaign; ranking each of the phrases (212) identified from the landing page URL (204), the rank of each of the phrases (212) corresponding to the likelihood that the respective phrase (212) is associated with the one or more ad campaign attributes of the ad campaign; 12. The system (100) of claim 11, comprising identifying one or more words associated with an ad campaign having the one or more ad campaign attributes (106) by:
16. 16. The system (100) of claim 15, wherein the operation further comprises determining whether the rank of the identified phrase (212) meets a threshold.
17. generating the voice-over script (252) occurs if the rank of one of the identified phrases (212) meets the threshold; 17. The system (100) of claim 16, wherein the sequence of text in the voice-over script (252) represents the identified phrase (212) that satisfies the threshold.
18. In response to determining that the rank of the identified phrase (212) does not meet the threshold, the operation: accessing a corpus of advertisements (208) associated with different advertisement campaigns, each advertisement (208) associated with a respective advertisement campaign having a respective voice-over script (252R) and a set of advertisement campaign attributes (106R); identifying one or more advertisements (208) from the corpus of advertisements (208) having similar ad campaign attributes (106R) as the one or more ad campaign attributes (106) of the voice-over request (102); generating the voice-over script (252) for the synthesized voice-over speech (352) based on the respective voice-over scripts (252R) for the identified one or more advertisements (208) having advertising campaign attributes (106R) similar to the one or more advertising campaign attributes (106) of the voice-over request (102); 17. The system (100) of claim 16, further comprising:
19. The TTS system (300) a TTS model (320) configured to convert the sequence of text of the voice-over script (252) into a corresponding synthesized speech representation (322) of the voice-over script (252); a TTS synthesizer (350) configured to generate the synthesized voice-over speech (352) from the synthesized speech expression (322) output from the TTS model (320); The system (100) of claim 11, comprising:
20. 20. The system (100) of any one of claims 11 to 19, wherein the one or more ad campaign attributes (106) are associated with a human-created ad campaign.
Citation Information
Patent Citations
Text-based image generation
JP2014519082A
Speech translation method and system using a multilingual text-to-speech synthesis model
JP2021511534A
Speech synthesis prosody using a BERT model
US11881210B2
Systems, methods and computer program products for generating script elements and call to action components therefor
US20190355024A1