Advertisement assistance learning method, advertisement assistance method, advertisement assistance learning device, advertisement assistance device, and program
By using a promotion word estimation model to analyze voice features and customer attributes, the method generates tailored promotional texts that align with the speaker's voice characteristics, enhancing advertising effectiveness.
Patent Information
- Application Number
- PCT/JP2023/042170
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-30
AI Technical Summary
Existing advertising methods lack effectiveness in tailoring promotional messages to the voice characteristics of the speaker, which can impact customer perception and purchasing willingness.
A computer-based method that uses a promotion word estimation model to learn the relationship between voice features, customer attributes, and product types, generating tailored promotional words and texts that match the speaker's voice characteristics.
The method supports effective advertising by generating promotional texts that enhance purchasing desire based on the speaker's voice characteristics, improving advertising outcomes.
Smart Images

Figure 00000014_0000 
Figure 00000015_0000 
Figure 00000017_0000
Abstract
Description
Advertising support learning method, advertising support method, advertising support learning device, advertising support device and program
[0001] The present invention relates to an advertisement-assisted learning method, an advertisement-assisted learning device, an advertisement-assisted device, and a program.
[0002] When promoting products in stores, salespeople use promotional texts written in a manual or ones they have thought up themselves, but at this point it is not clear whether the promotion is effective for customers.
[0003] Research on the impression of speech has shown that listeners' impressions differ depending on the characteristics of the speech. There are also reports that the perceived sense of security differs depending on the gender of the speaker.
[0004] Jiayuan Dong et al. (2020), "Female Voice Agents in Fully Autonomous Vehicles Are Not Only More Likeable and Comfortable, But Also More Competent", Proceedings of the Human Factors and Ergonomics Society Annual Meeting 64(1):1033-1037Kobayashi, M., Hamada, Y., and Akagi, M. (2022), "Acoustic features correlated to perceived urgency in evacuation announcements", Speech Communication, 22-34. doi:https: / / doi.org / 10.1016 / j.specom.361 2022.03.001
[0005] For example, when advertising insurance using advertising text containing the word "peace of mind," a salesperson with a reassuring voice may be able to more effectively stimulate purchasing desire. On the other hand, if a salesperson has a tense voice, advertising text containing a word of urgency, such as "now," may be able to more effectively stimulate purchasing desire.
[0006] The present invention has been made in view of the above points, and has as its object to support effective advertising in accordance with the characteristics of the speaker's voice.
[0007] To solve the above problem, a computer executes an advertising word estimation model training procedure, which uses training data including speech features of speech data containing product advertisements, attributes of customers who are the targets of the advertisements, the type of product, and words contained in the advertisements, and causes a machine learning model to learn the relationship between input and output when the speech features, customer attributes, and type of product are input and the words are output.
[0008] This can support effective advertising that is tailored to the characteristics of the speaker's voice.
[0009] Fig. 1 is a diagram showing an example of the hardware configuration of the advertising support device 10 in a first embodiment. Fig. 2 is a diagram showing an example of the functional configuration of the advertising support device 10 (advertising support learning device) when learning the advertising word estimation model 13 of the first embodiment. Fig. 3 is a diagram showing an example of learning data for the advertising word estimation model 13. Fig. 4 is a diagram showing an example of the functional configuration of the advertising support device 10 when making an inference using the advertising word estimation model 13 of the first embodiment. Fig. 5 is a diagram showing an example of the functional configuration of the advertising support device 10 when learning the advertising word estimation model 13 of a second embodiment.
[0010] This embodiment discloses an advertising support device 10 that generates advertising text that matches the characteristics of a speaker's voice and increases listeners' purchasing motivation. Specifically, the advertising support device 10 (advertising support learning device) learns a model (hereinafter referred to as an "advertising word estimation model 13") that outputs advertising words that match the characteristics of the speaker's voice, based on the speaker's voice, product type, and customer attributes. The advertising support device 10 uses an existing large-scale language model or the like to generate advertising text using the advertising words output by the model.
[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0012] 1 is a diagram showing an example of the hardware configuration of an advertising support device 10 according to the first embodiment. The advertising support device 10 in FIG. 1 includes a drive device 100, an auxiliary storage device 102, a memory device 103, a processor 104, and an interface device 105, all of which are interconnected via a bus B.
[0013] The program that realizes the processing in the advertising support device 10 is provided by a recording medium 101 such as a CD-ROM. When the recording medium 101 storing the program is set in the drive device 100, the program is installed from the recording medium 101 to the auxiliary storage device 102 via the drive device 100. However, the program does not necessarily have to be installed from the recording medium 101, but may be downloaded from another computer via a network. The auxiliary storage device 102 stores the installed program as well as necessary files, data, etc.
[0014] When an instruction to start a program is received, the memory device 103 reads the program from the auxiliary storage device 102 and stores it. The processor 104 is a CPU or a GPU (Graphics Processing Unit), or a CPU and a GPU, and executes functions related to the advertising support device 10 in accordance with the program stored in the memory device 103. The interface device 105 is used as an interface for connecting to a network.
[0015] 2 is a diagram showing an example of the functional configuration of the advertising support device 10 (advertising support learning device) during learning of the advertising word estimation model 13 according to the first embodiment. In Fig. 2, the advertising support device 10 includes a speech feature extraction unit 11, an advertising word estimation model learning unit 12, and an advertising word estimation model 13. Each of these units is realized by processing executed by a processor 104 of one or more programs installed in the advertising support device 10.
[0016] The speech feature extraction unit 11 performs signal processing such as Fourier transform on the input experimental speech data for each frame, or extracts and outputs a speech feature vector by using a speech feature extraction tool such as OpenSMILE [Reference 3]. During training, the speech feature extraction unit 11 extracts a vector as a speech feature (speech feature vector) from the training speech data.
[0017] The advertising word estimation model training unit 12 uses training data including training speech feature vectors, training customer attribute data, training product type data, and training advertising word data to train the advertising word estimation model 13 on the relationship between input and output when the training speech feature vectors, training customer attribute data, and training product type data are input and the training advertising word data is output. In other words, the advertising word estimation model 13 is a machine learning model that receives input of the speech feature vectors, customer attribute data, and product type data and outputs advertising word data. A conventional DNN such as that described in [Reference 4] can be used to train the advertising word estimation model 13.
[0018] The training audio data is, for example, existing advertising audio data (audio data including advertising). Training audio data can be prepared by extracting commercial audio from video distribution services, web radio, etc. However, background music must be removed as much as possible. The audio should be about one to two sentences long. There should be one speaker, and audio in which multiple people speak simultaneously within a sentence should not be used. Note that multiple pieces of training audio data are prepared.
[0019] The training speech feature vector is a vector representing the pitch, intonation, etc. of a voice output for each training speech data by the speech feature extraction unit 11. There is no restriction on the number of dimensions as long as the number of dimensions of all speech feature vectors is uniform.
[0020] The training customer attribute data is data that numerically represents the attributes (age, gender, place of residence, etc.) of customers who are the target of the advertisement related to the training audio data (advertising audio), and is prepared for each training audio data. The training customer attribute data is obtained by listening to the training audio data or by manually annotating the transcript of the training audio data. In the case of an advertisement audio that says, "If you are under 35, you can save on your rent for three years!", the customer attribute is annotated as, for example, "male, in his 20s." If the attributes of the target customer are not clearly stated in the training audio data, they are obtained from the overall content of the original commercial used to extract the training audio data, or are inferred from the content of the product being advertised. When the annotation is quantified, it is expressed as a real value, 1-hot data, etc.
[0021] The training product type data is data that numerically represents the category of the product (product type such as housing or insurance) being advertised in the training audio data (advertising audio), and is prepared for each training audio data. The training product type data is obtained by listening to the training audio data or by manually annotating the transcript of the training audio data. In the case of an advertising audio that says, "If you're under 35, you'll save on your rent for three years!", the product type is annotated as, for example, "housing." If the target product is not clearly stated in the training audio data, it is obtained from the overall content of the original commercial used to extract the training audio data. When the annotation is quantified, it is expressed as 1-hot data or the like.
[0022] The training advertising word data is data that numerically represents advertising words contained in the training audio data (advertising audio), and is prepared for each training audio data. An advertising word is a single word that corresponds to an expression or phrase that particularly stimulates purchasing desire among the group of words contained in the advertisement, such as "highly praised" or "great deal," as shown in [Reference 1] (Table 9). The advertising word is also obtained by listening to the training audio data or by manually annotating the transcript of the training audio data. For example, in the case of an advertising audio that says, "If you're under 35, you can save three years on your rent!", the advertising word is annotated as "great deal." When quantifying the annotation, it is expressed as 1-hot data or a vector obtained using a method such as [Reference 2].
[0023] The advertising word estimation model 13 is a model that learns the relationship between the speech feature vector, customer attribute data, product type data, and advertising words. The advertising word estimation model 13 receives the speech feature vector, customer attribute data, and product type data, and outputs advertising words that match the characteristics of the speaker's voice.
[0024] In training the advertising word estimation model 13, the advertising word estimation model training unit 12 generates training data consisting of a set of a speech feature vector, customer attribute data, product type data, and advertising words for each training speech data.
[0025] Fig. 3 is a diagram showing an example of training data for the advertising word estimation model 13. One row in Fig. 3 corresponds to one piece of training data based on one piece of training speech data.
[0026] The advertising word estimation model training unit 12 processes each of these training data one by one, and inputs the input data (speech feature vector, customer attribute data, and product type data) of the training data to be processed (hereinafter referred to as "target training data") to the advertising word estimation model 13. The advertising word estimation model training unit 12 updates the parameters of the advertising word estimation model 13 so that the output from the advertising word estimation model 13 approaches (matches) the output data of the target training data (training advertising word data). More specifically, the advertising word estimation model 13 assigns each of the many words that are candidate advertising words to each dimension, and outputs a vector (probability distribution of candidate advertising words) in which the value of each dimension is the probability that the word is a correct answer for the input data. The advertising word estimation model training unit 12 calculates a loss using a known method based on the correct answer vector (the correct answer to the probability distribution of candidate advertising words) in which the value of the dimension corresponding to the advertising word data of the target training data is 1 and the value of the dimensions corresponding to other words is 0, and the vector output from the advertising word estimation model 13. The advertising word estimation model learning unit 12 updates the parameters of the advertising word estimation model 13 based on the loss using a known technique such as backpropagation.
[0027] Next, the inference process will be described. During the inference process, the advertising support device 10 estimates advertising words that match the voice of a certain speaker using the advertising word estimation model 13, and generates advertising text using these advertising words, thereby generating advertising text that matches the speaker's voice and increases purchasing motivation. Note that the advertising support device 10 during learning (advertising support learning device) and the advertising support device 10 during inference may be realized using different computers.
[0028] 4 is a diagram showing an example of the functional configuration of the advertising support device 10 during inference using the advertising word estimation model 13 of the first embodiment. In FIG. 4, parts that are the same as or correspond to those in FIG. 2 are given the same reference numerals.
[0029] 4, the advertising support device 10 includes a speech feature extraction unit 11, an advertising word estimation unit 14, an advertising text generation unit 15, an advertising word estimation model 13, and a large-scale language model 16. Each of these units is realized by a process executed by a processor 104 of one or more programs installed in the advertising support device 10.
[0030] During inference, the speech feature extraction unit 11 generates a speech feature vector of the input speech. The input speech is speech data that records the speech of a single speaker (e.g., a store clerk) who will be advertising using the promotional text that will ultimately be generated. Since it is sufficient to be used to acquire the features of the speaker's voice, the spoken text does not have to be a promotional text and can be any spoken text.
[0031] The advertising word estimation unit 14 estimates advertising words using the advertising word estimation model 13 based on the speech feature vector, customer attribute data, and product type data output by the speech feature extraction unit 11 for the input speech. That is, the advertising word estimation unit 14 inputs these speech feature vector, customer attribute data, and product type data to the advertising word estimation model 13. The advertising word estimation unit 14 estimates the candidate with the highest probability as the advertising word from the probabilities output by the advertising word estimation model 13 for each advertising word candidate. The customer attribute data at the time of estimation is data that numerically represents the attributes (age, gender, place of residence, etc.) of customers targeted by the product that the user wants to advertise. The product type data at the time of estimation is data that numerically represents the category (housing, insurance, etc.) of the product in question.
[0032] The large-scale language model 16 is an existing large-scale language model, such as ChatGPT, that can generate sentences from prompts.
[0033] The advertising text generation unit 15 generates an input prompt that instructs the generation of text to advertise a product related to the product type data to a customer related to the customer attribute data using advertising words. The advertising text generation unit 15 inputs the input prompt into the large-scale language model 16, causing the large-scale language model 16 to generate advertising text. For example, if the product type data is "insurance," the customer attribute data is "male in his 60s," and the advertising word is "peace of mind," the advertising text generation unit 15 generates an input prompt such as "Please create a text advertising insurance to males in their 60s using the word 'peace of mind'." The large-scale language model 16 that has received the input prompt outputs advertising text such as "Advertising text: 'This insurance will give you peace of mind in your old age!'"
[0034] As described above, according to the first embodiment, it is possible to generate advertising words that match the characteristics of a speaker's voice. Furthermore, the advertising word estimation model 13, which is trained using the speaker's voice, product type, and customer attribute data, can generate advertising text that uses advertising words that match the characteristics of the speaker's voice and that enhances the listener's purchasing motivation. Therefore, it is possible to support effective advertising that matches the characteristics of the speaker's voice.
[0035] Next, a second embodiment will be described. In the second embodiment, differences from the first embodiment will be described. Points not specifically mentioned in the second embodiment may be the same as those in the first embodiment.
[0036] In the first embodiment, when training the advertising word estimation model 13, it was necessary to prepare training advertising word data by manually annotating it. However, there is a possibility that the words evaluated as advertising words may vary depending on the annotator.
[0037] In advertising speech, when promoting words that stimulate purchasing desire, such as "highly praised" or "great value," there is a tendency to emphasize those words. Therefore, in the second embodiment, advertising words are automatically acquired using a model (emphasis estimation model 17) trained using a method for acquiring emphasized parts as described in [Reference 5] (FIG. 5).
[0038] 5 is a diagram showing an example of the functional configuration of the advertising support device 10 during learning of the advertising word estimation model 13 according to the second embodiment. In FIG. 5, the same components as those in FIG. 2 are designated by the same reference numerals, and their description will be omitted.
[0039] 5, the advertising support device 10 further includes an emphasis estimation model 17 and an advertising word extraction unit 18. These units are realized by the processor 104 executing one or more programs installed in the advertising support device 10.
[0040] The emphasis estimation model 17 is an existing model that estimates emphasized parts of speech. In other words, the emphasis estimation model 17 is a model that has already learned the relationship between speech and emphasized parts. The emphasis estimation model 17 can be realized using publicly known technology. The emphasized parts can be estimated using a method such as [Reference 5], and a conventional DNN such as [Reference 4] can also be used for learning.
[0041] The advertising word extraction unit 18 receives training speech data as input and estimates words corresponding to emphasized portions of the training speech data using the emphasis estimation model 17. For example, when training speech data such as "Get a discount on your rent for three years if you're under 35!" is input, the advertising word extraction unit 18 outputs the word "bargain." In the second embodiment, the advertising word estimation model training unit 12 uses the words output by the advertising word extraction unit 18 for certain training speech data as training advertising words corresponding to the training speech data to train the advertising word estimation model 13.
[0042] [References] [Reference 1] Teppei Otani, "Commercially Active Words in Headlines: The Case of Magazine Article Headlines," Bulletin of Hokuriku University, No. 50, pp. 117-132, 2020. [Reference 2] Tomas Mikolov, Kai Chen, G. Corrado, and J. Dean, "Efficient Estimation of Word Representations in Vector Space," In Proceedings of the International Conference on Learning Representations, 2013. [Reference 3] F. Eyben, M. Woellmer, and B. Schuller, "OpenSMILE: The Munich Versatile and Fast Open-Source Audio Feature Extractor," in ACM International Conference on Multimedia (MM 2010), Florence, Italy, pp. 1459-1462, 2010. [Reference 4] Han, K., Yu, D. and Tashev, I., "Speech Emotion "Recognition Using Deep Neural Network and Extreme Learning Machine," Proc. of INTERSPEECH, pp. 223-227, 2014. [Reference 5] Nakajima Shuji et al., "Prediction of Emphasized Accent Phrases from Advertising Text for Expressive Text-to-Speech Synthesis," Transactions of Information Processing Society of Japan 56 (12), 2384-2394, 2015. The above describes in detail the embodiments of the present invention, but the present invention is not limited to such specific embodiments, and various modifications and variations are possible within the scope of the gist of the present invention as defined in the claims.
[0043] REFERENCE SIGNS LIST 10 Advertising support device 11 Speech feature extraction unit 12 Advertising word estimation model learning unit 13 Advertising word estimation model 14 Advertising word estimation unit 15 Advertising text generation unit 16 Large-scale language model 17 Emphasis estimation model 18 Advertising word extraction unit 100 Drive device 101 Recording medium 102 Auxiliary storage device 103 Memory device 104 Processor 105 Interface device B Bus
Claims
1. A learning method for promoting advertising support, characterized in that a computer executes a learning procedure for an advertising word estimation model that uses learning data including voice feature amounts of voice data including product advertising, attributes of customers targeted by the advertising, types of the products, and words included in the advertising, and trains a machine learning model on the relationship between the input and output when the voice feature amounts, the customer attributes, and the product types are input and the words are output.
2. The learning method for promoting advertising support according to claim 1, characterized in that a computer executes an advertising word extraction procedure for estimating a word corresponding to a emphasized part in the voice data using a model for estimating a emphasized part in voice, and the learning procedure for the advertising word estimation model uses the word estimated by the advertising word extraction procedure as the word included in the advertising in the learning data.
3. A method for promoting advertising support, characterized in that a computer executes an advertising word estimation procedure for estimating a word to be included in an advertising text of a product being advertised by an input voice based on an output from a machine learning model when an input voice feature amount of the input voice, attributes of a customer targeted by the input voice for advertising, and a type of the product being advertised by the input voice are input to the machine learning model that has learned the relationship between the input and output when the voice feature amounts of voice data including product advertising, attributes of customers targeted by the advertising, types of the products, and words included in the advertising are input and the words are output.
4. The method for promoting advertising support according to claim 3, characterized in that a computer executes an advertising text generation procedure for generating, using a large language model, an advertising text that includes the word estimated by the advertising word estimation procedure and advertises the product being advertised by the input voice to the customer targeted by the input voice for advertising.
5. An apparatus for promoting advertising support, characterized by comprising an advertising word estimation model learning unit configured to train a machine learning model on the relationship between the input and output when voice feature amounts of voice data including product advertising, attributes of customers targeted by the advertising, types of the products, and words included in the advertising are input and the words are output.
6. Using learning data including the acoustic feature amount of voice data including product promotion, the attributes of customers targeted by the promotion, the type of the product, and the words included in the promotion, for a machine learning model that has learned the relationship between the input and output when the acoustic feature amount, the attributes of the customer, and the type of the product are input and the words are output, based on the output from the machine learning model when the acoustic feature amount of the input voice, the attributes of the customer targeted by the input voice for promotion, and the type of the product being promoted by the input voice are input, a promotion word estimation unit configured to estimate words to be included in the promotion text of the product being promoted by the input voice. A promotion support device characterized by having.
7. A program characterized by causing a computer to execute the promotion support learning method according to claim 1 or 2.
8. A program characterized by causing a computer to execute the promotion support method according to claim 3 or 4.
Citation Information
Patent Citations
Information processor, information processing method, and program
JP2021033215A
Program, device and method for selecting items based on compensated effects, and item effect estimation program
JP2021068238A
Voice generation method, voice generation device, and voice generation program
WO2023017582A1