A method, device and system for audio advertisement delivery

By dynamically matching audio ads in the cloud, combining audio program information and user characteristics, the problem of low matching rate of audio ads in podcasts is solved, resulting in a better user experience and better advertising performance.

CN119923660BActive Publication Date: 2026-04-24HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2022-09-30
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing audio advertising methods have poor matching with audio programs in podcasts, which affects the user listening experience and results in poor advertising performance.

Method used

By receiving advertising requests from clients through the cloud, and combining audio program information and user characteristics, the system dynamically matches audio ads that match the target ad slots. It also optimizes ad selection using vector representation and ad ranking models to ensure that audio ads match user preferences and program content.

Benefits of technology

It improves the matching accuracy and delivery effectiveness of audio ads, meets users' personalized needs, and enhances user experience and ad completion rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119923660B_ABST
    Figure CN119923660B_ABST
Patent Text Reader

Abstract

A method for audio advertisement delivery, comprising: sending an advertisement request (301) to a cloud device (20, 90, 110) by a client (100) when playing an audio program, the advertisement request including information of the audio program, identification of a target advertisement slot, and user features; determining a vector representation (302) of the target advertisement slot by the cloud device (20, 90, 110) according to the information of the audio program and the identification of the target advertisement slot, the vector representation of the target advertisement slot being used to describe the content involved in the audio program within a period of time before the target advertisement slot; determining an audio advertisement (303) matching the target advertisement slot by the cloud device (20, 90, 110) according to the user features and the vector representation of the target advertisement slot; sending the audio advertisement (304) to the client (100) by the cloud device (20, 90, 110), and playing the audio advertisement (305) by the client (100) when playing the audio program to the target advertisement slot. The matching degree of the audio advertisement and the audio program is higher, and the user features are combined, so that the personalized needs of the user can be better met, and the delivery effect of the audio advertisement can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a method, device, and system for audio advertising. Background Technology

[0002] Podcasts are recorded online radio or audio programs, such as audiobooks, stand-up comedy, and current affairs news. The podcast market is booming, with hundreds of millions of users. Along with the rapid development of podcasts, the corresponding advertising market share is also growing. With the development of podcasts both domestically and internationally, audio advertising has become an important form of advertising.

[0003] Inserting audio ads into audio programs requires first identifying the ad slots within the audio program offline—that is, the positions where the audio ads will be inserted. Then, audio ads are configured for each ad slot. This way, when the audio program reaches that ad slot, the audio ad configured for that slot will play.

[0004] This type of offline audio advertising has a low degree of relevance to audio programs, often affecting the continuity of users' listening to audio programs, resulting in poor advertising performance. Summary of the Invention

[0005] This application provides a method for delivering audio advertisements, used to deliver audio advertisements that meet the personalized needs of users within audio programs. This application also provides corresponding devices, systems, computer-readable storage media, and computer program products.

[0006] The first aspect of this application provides a method for delivering audio advertisements, comprising: receiving an advertisement request from a client via a cloud platform, the advertisement request including information about an audio program, an identifier of a target advertisement slot, and user characteristics, wherein the target advertisement slot is one of at least one advertisement slots extracted from the audio program, and the advertisement request is triggered when the client plays the audio program; determining a vector representation of the target advertisement slot based on the information about the audio program and the identifier of the target advertisement slot, the vector representation of the target advertisement slot being used to describe the content involved in the audio program within a certain period before the target advertisement slot; obtaining an audio advertisement matching the target advertisement slot based on the user characteristics and the vector representation of the target advertisement slot; and sending the audio advertisement to the client, the audio advertisement being played by the client when the audio program reaches the target advertisement slot.

[0007] In this application, the cloud can be software or services of a cloud platform, or software or services deployed on nodes in a network, such as edge nodes. The cloud can run on a standalone physical machine or on virtualized resources.

[0008] In this application, the client can be a terminal device or an application, for example, the application runs on a terminal device for user use.

[0009] In this application, when the client plays an audio program, it usually means when the client is about to reach the target ad slot. Typically, an ad request is triggered at a preset time point before reaching the target ad slot. This preset time point can be 5 seconds away from the target ad slot or other time points representing duration.

[0010] In this application, audio advertising refers to advertising played via audio. An advertising request is used to request audio advertising from the cloud.

[0011] In this application, the information of the audio program may be an identifier or index of the audio program. The audio program is one that the client is about to play, is currently playing, or has just finished playing. The audio program may be an audiobook, a song in audio format, a crosstalk performance, or current affairs news, etc.

[0012] In this application, during the ad slot mining stage, one or more ad slots can be mined from an audio program. Each ad slot in an audio program has a unique identifier. Each ad slot also has a vector representation. The identifier of the same ad slot is associated with its vector representation. Furthermore, the representation and vector representation of at least one ad slot in each audio program are stored in association with that audio program. This information can all be stored in the audio content library of a cloud platform. The vector representation of an ad slot refers to the vector obtained by encoding the content involved in that ad slot within a certain period of time. In this application, "a period of time" can be a duration, such as one minute or other numerical values ​​representing duration. For example, in one implementation, the specific numerical value can be preset; in another implementation, the duration can be randomly selected within a range.

[0013] In this application, user characteristics may include user profiles and user behavior characteristics. User profiles may include basic user information, such as gender, age, and hobbies. User behavior characteristics may include user behavior information such as clicking, saving, and commenting on historical audio programs.

[0014] In this application, based on the information of the audio program, the identifiers of all advertising slots associated with the audio program and the vector representations of the advertising slots can be found. Furthermore, based on the identifier of the target advertising slot, the vector representation of the target advertising slot can be determined.

[0015] In this application, one or more audio advertisements strongly correlated with the audio program being listened to by the user in front of the target ad slot can be identified based on the vector representation of the target ad slot. Furthermore, the audio advertisements can be further filtered or processed based on user characteristics to obtain audio advertisements matching the target ad slot. Because the audio advertisements identified in this application have a higher degree of matching with the audio program and incorporate user characteristics, they better meet the personalized needs of users and can improve the effectiveness of audio ad delivery.

[0016] In one possible implementation, the above steps include: the cloud determining the audio ad that matches the target ad slot based on user characteristics and the vector representation of the target ad slot, including: the cloud recalling multiple audio ads from the audio ad library based on the vector representation of the target ad slot; and the cloud obtaining the audio ad that matches the target ad slot from the multiple audio ads based on user characteristics.

[0017] In this possible implementation, the cloud recalls multiple audio ads that are strongly related to the audio program that the user is listening to in front of the target ad slot based on the vector representation of the target ad slot, and then selects the ad with the highest matching degree with the user characteristics. This can improve the matching degree between audio ads and audio programs playing on the client.

[0018] In one possible implementation, the above steps include: the cloud obtaining audio ads that match the target ad slot from multiple audio ads based on user characteristics, including: the cloud predicting the completion rate of multiple audio ads based on user characteristics and an ad ranking model, wherein the audio ad with the highest completion rate is the audio ad that matches the target ad slot, or the audio ad with the highest completion rate is the source ad of the audio ad that matches the target ad slot, and the ad ranking model is a model that takes user characteristics as input and completion rate as output.

[0019] In this possible implementation, the completion rate refers to the predicted probability that an audio ad will be played in its entirety. The closer the content and style of an audio ad are to the user's preferences, the higher the probability of it being played in its entirety, and the better the performance after it is delivered. Therefore, the completion rate of multiple recalled audio ads can be predicted based on user characteristics, and the audio ad with the highest completion rate can be selected as the main audio ad or as the source ad for an audio ad campaign, thus improving the effectiveness of audio ad delivery.

[0020] In one possible implementation, when the audio ad with the highest completion rate is the source ad of the audio ad that matches the target ad slot, the method further includes: the cloud adjusts the style of the audio ad with the highest completion rate according to the style of the audio program and user characteristics to obtain an audio ad that matches the target ad slot.

[0021] In this possible implementation, after the cloud determines the audio ad with the highest completion rate through the ad ranking model, it can further adjust the style of the audio ad with the highest completion rate according to the style of the audio program that the user wants to play or is currently playing and the user characteristics. This can improve the user's acceptance of the audio ad, thereby improving the effectiveness of the audio ad delivery.

[0022] In one possible implementation, the above steps include: the cloud adjusting the style of the audio ad with the highest completion rate based on the style of the audio program and user characteristics to obtain an audio ad matching the target ad slot, including: the cloud adjusting the object sound in the audio ad with the highest completion rate based on the style vector of the object sound in the audio program and the style vector of user preference, wherein the style vector of the object sound in the audio program is obtained by encoding the object sound in the audio program, and the style vector of user preference is obtained by encoding user characteristics; the cloud adjusting the background music in the audio ad with the highest completion rate based on the style vector of the background music in the audio program and the style vector of user preference, wherein the style vector of the background music in the audio program is obtained by encoding the background music in the audio program; and the cloud fusing the adjusted object sound and the adjusted background music of the audio ad with the highest completion rate to obtain an audio ad matching the target ad slot.

[0023] In this possible implementation, if the audio program includes both subject sound and background music, the subject sound and background music can be separated and encoded separately to obtain style vectors for the subject sound and background music. These style vectors are then combined with the user's preferred style vectors to adjust the subject sound and background music in the highest-scoring audio ad, resulting in the final audio ad. The adjusted style of the audio ad matches the style of the audio program, satisfying the user's style preferences, improving user experience, and thus enhancing the effectiveness of the audio ad campaign.

[0024] In one possible implementation, before receiving the advertising request sent by the client in the cloud, the method further includes: the cloud determining at least one advertising slot based on the temporal information of the audio program in the speech state and the text content after the audio program is converted into text; the cloud encoding the text content of each advertising slot in the at least one advertising slot within a previous period to obtain a vector representation of each advertising slot.

[0025] In this possible implementation, the time-domain information may include amplitude (amplitude can also be described as sound intensity), the change of amplitude over time, etc.

[0026] The cloud can also perform ad slot mining before receiving ad requests. The process of mining ad slots can involve determining the ad slots for an audio program based on its temporal information in speech mode and the text content after the audio program is converted to text. In this application, because ad slots can be mined based on both the temporal information and text content of the audio program, the quality of the mined ad slots is high. Inserting an ad into this slot typically does not affect the continuity of the audio program, thereby improving user experience and enhancing the effectiveness of audio ad delivery.

[0027] In one possible implementation, the method further includes: storing the audio program, the identifier of each ad slot, and the vector representation of each ad slot in association in the cloud.

[0028] In this possible implementation, the audio program, the identifier of each ad slot in the audio program, and the vector representation of each ad slot are stored together. This makes it easier to quickly determine the audio ad that matches the requested target ad slot when the client sends an ad request, thereby improving the efficiency of audio ad delivery.

[0029] In one possible implementation, the above steps are as follows: The cloud determines at least one advertising slot based on the temporal information of the audio program in speech mode and the text content after the audio program is converted into text. This includes: when the temporal information is amplitude, if the duration for which the amplitude of the audio program in speech mode is continuously lower than the amplitude threshold exceeds a first threshold, then the cloud determines the duration for which the amplitude is continuously lower than the amplitude threshold as the first basic advertising slot; if the time interval between two adjacent words in the text content after the audio program is converted is greater than a second threshold, then the cloud determines the time interval between two adjacent words as the second basic advertising slot. The time interval between two adjacent words is determined by the timestamp of each word during text conversion; the cloud determines at least one advertising slot from the union of the first basic advertising slot and the second basic advertising slot.

[0030] In this possible implementation, selecting ad slots from the union of the first and second basic ad slots can expand the range of ad slot selection.

[0031] In one possible implementation, the above steps include: the cloud determining at least one advertising slot from the union of the first basic advertising slot and the second basic advertising slot, which includes: the cloud selecting at least one advertising slot with the largest weight from the union of the first basic advertising slot and the second basic advertising slot as at least one advertising slot for the audio program, wherein the weight of each advertising slot is determined by the punctuation mark and / or the segmentation position of the text segment corresponding to each advertising slot.

[0032] In one possible implementation, the weight of the corresponding basic advertising slots can be increased by using punctuation marks, text segmentation, etc., and then at least one basic advertising slot with the highest weight can be selected as at least one advertising slot for the audio program. This can improve the quality of the selected advertising slots.

[0033] In one possible implementation, user characteristics include user profiles and user behavior characteristics related to historical audio programs.

[0034] A second aspect of this application provides a method for delivering audio advertisements, comprising: when a client plays an audio program, sending an advertisement request to the cloud, the advertisement request including information about the audio program, an identifier of a target advertisement slot, and user characteristics, wherein the target advertisement slot is one of at least one advertisement slot extracted from the audio program; the client receiving an audio advertisement sent by the cloud that matches the target advertisement slot; and the client playing the audio advertisement when playing the audio program to the target advertisement slot.

[0035] In this application, when the client plays an audio program, it usually means when the client is about to reach the target ad slot. Typically, an ad request is triggered at a preset time point before reaching the target ad slot. This preset time point can be 5 seconds away from the target ad slot or other time points representing duration.

[0036] In one possible implementation, user characteristics include user features and user preferences for audio programs.

[0037] A third aspect of this application provides a method for mining advertising slots, comprising: acquiring an audio program of the advertising slot to be mined in the cloud; determining at least one advertising slot in the cloud based on the temporal information of the audio program in speech mode and the text content after the audio program is converted into text; and encoding the text content of each advertising slot in the at least one advertising slot in the cloud within a previous period of time to obtain a vector representation of each advertising slot.

[0038] In one possible implementation, the method further includes: storing the audio program, the identifier of each ad slot, and the vector representation of each ad slot in association in the cloud.

[0039] In one possible implementation, the above steps are as follows: The cloud determines at least one advertising slot based on the temporal information of the audio program in speech mode and the text content after the audio program is converted into text. This includes: when the temporal information is amplitude, if the duration for which the amplitude of the audio program in speech mode is continuously lower than the amplitude threshold exceeds a first threshold, then the cloud determines the duration for which the amplitude is continuously lower than the amplitude threshold as the first basic advertising slot; if the time interval between two adjacent words in the text content after the audio program is converted is greater than a second threshold, then the cloud determines the time interval between two adjacent words as the second basic advertising slot. The time interval between two adjacent words is determined by the timestamp of each word during text conversion; the cloud determines at least one advertising slot from the union of the first basic advertising slot and the second basic advertising slot.

[0040] In one possible implementation, the above steps include: the cloud determining at least one advertising slot from the union of the first basic advertising slot and the second basic advertising slot, which includes: the cloud selecting at least one advertising slot with the largest weight from the union of the first basic advertising slot and the second basic advertising slot as at least one advertising slot for the audio program, wherein the weight of each advertising slot is determined by the punctuation mark and / or the segmentation position of the text segment corresponding to each advertising slot.

[0041] In a fourth aspect, this application provides a cloud-based device for executing the method described in the first aspect or any possible implementation thereof. Specifically, the cloud-based device includes modules or units for executing the method described in the first aspect or any possible implementation thereof, such as a processing unit, a sending unit, and a receiving unit.

[0042] In a fifth aspect, this application provides a client for executing the method of the second aspect described above. Specifically, the client includes modules or units for executing the method of the second aspect or any possible implementation thereof, such as a receiving unit, a display unit, and a sending unit.

[0043] In a sixth aspect, this application provides a cloud-based device for executing the methods described in the first aspect or any possible implementation thereof. Specifically, the cloud-based device includes modules or units for executing the methods described in the third aspect or any possible implementation thereof, such as a processing unit, a sending unit, and a receiving unit.

[0044] A seventh aspect of this application provides a cloud-based device. The cloud-based device may include at least one processor, a memory, and a communication interface. The processor is coupled to the memory and the communication interface. The memory is used to store instructions, the processor is used to execute the instructions, and the communication interface is used to communicate with other network elements under the control of the processor. When executed by the processor, the instructions cause the processor to perform the method of the first aspect or any possible implementation thereof.

[0045] The eighth aspect of this application provides a client including a transceiver, a processor, and a memory, wherein the transceiver and the processor are coupled to the memory, and the memory is used to store a program or instructions that, when executed by the processor, cause a cloud device to perform the methods of the second aspect or any possible implementation thereof.

[0046] A ninth aspect of this application provides a cloud-based device. The cloud-based device may include at least one processor, a memory, and a communication interface. The processor is coupled to the memory and the communication interface. The memory is used to store instructions, the processor is used to execute the instructions, and the communication interface is used to communicate with other network elements under the control of the processor. When executed by the processor, the instructions cause the processor to perform a method of the third aspect or any possible implementation thereof.

[0047] The tenth aspect of this application provides a chip system including one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the interface circuits are used to receive signals from the memory of a cloud device and send signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the cloud device executes the method in the first aspect or any possible implementation of the first aspect.

[0048] The eleventh aspect of this application provides a chip system including one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the interface circuits are used to receive signals from the client's memory and send signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the client executes the method in the second aspect or any possible implementation of the second aspect.

[0049] The twelfth aspect of this application provides a chip system including one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the interface circuits are used to receive signals from the memory of a cloud device and send signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the cloud device executes the method in the aforementioned third aspect or any possible implementation of the third aspect.

[0050] The thirteenth aspect of this application provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed on a computer device, cause the computer device to perform the method described in the first aspect or any possible implementation thereof.

[0051] The fourteenth aspect of this application provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed on a computer device, cause the computer device to perform the methods of the second aspect or any possible implementation thereof.

[0052] The fifteenth aspect of this application provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed on a computer device, cause the computer device to perform the methods of the aforementioned third aspect or any possible implementation thereof.

[0053] The sixteenth aspect of this application provides a computer device program product including computer device program code, which, when executed on a computer device, causes the computer device to perform the method in the first aspect or any possible implementation thereof.

[0054] The seventeenth aspect of this application provides a computer device program product including computer device program code, which, when executed on a computer device, causes the computer device to perform the methods in the aforementioned second aspect or any possible implementation thereof.

[0055] The eighteenth aspect of this application provides a computer device program product including computer device program code, which, when executed on a computer device, causes the computer device to perform the methods in the aforementioned third aspect or any possible implementation thereof.

[0056] The nineteenth aspect of this application provides an audio advertising system, which includes a cloud device and a client. The cloud device is used to execute the methods in the first aspect or any possible implementation thereof, and the client is used to execute the methods in the second aspect or any possible implementation thereof.

[0057] The twentieth aspect of this application provides an audio advertising system, which includes a cloud device and an audio content library. The cloud device retrieves audio programs from the audio content library and performs the methods described in the third aspect or any possible implementation thereof.

[0058] The technical effects of the second to twentieth aspects or any of their possible implementations can be found in the first aspect or the technical effects of different possible implementations of the first aspect, and will not be repeated here. Attached Figure Description

[0059] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] Figure 1A This is a schematic diagram of the architecture of the audio advertising system provided in the embodiments of this application;

[0061] Figure 1B This is another schematic diagram of the architecture of the audio advertising system provided in the embodiments of this application;

[0062] Figure 2A This is a schematic diagram of the structure of a client provided in an embodiment of this application;

[0063] Figure 2B This is a schematic diagram of the structure of the cloud device provided in an embodiment of this application;

[0064] Figure 3 This is a schematic diagram of an embodiment of the audio advertising delivery method provided in this application.

[0065] Figure 4 This is a schematic diagram of the structure of an advertising ranking model provided in an embodiment of this application;

[0066] Figure 5 This is a schematic diagram of another embodiment of the audio advertising delivery method provided in this application;

[0067] Figure 6 This is a schematic diagram of an embodiment of the method for mining advertising slots provided in this application;

[0068] Figure 7 This is a schematic diagram illustrating a scenario provided in an embodiment of this application;

[0069] Figure 8 This is a schematic diagram of an embodiment of the mining of advertising slots and audio advertising placement provided in this application;

[0070] Figure 9 This is another structural schematic diagram of the cloud device provided in the embodiments of this application;

[0071] Figure 10 This is another structural schematic diagram of the client provided in the embodiments of this application;

[0072] Figure 11 This is another structural schematic diagram of the cloud device provided in the embodiments of this application. Detailed Implementation

[0073] The embodiments of this application are described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. As those skilled in the art will recognize, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0074] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0075] This application provides a method for delivering audio advertisements, which is used to deliver audio advertisements that meet the personalized needs of users within audio programs. This application also provides corresponding devices, systems, computer-readable storage media, and computer program products, etc., which are described in detail below.

[0076] Figure 1A This is a schematic diagram of the architecture of the audio advertising system provided in the embodiments of this application.

[0077] like Figure 1A As shown in the illustration, the audio advertising system provided in this application embodiment includes a cloud platform and multiple clients, which can communicate with the multiple clients via a network. The audio advertising system provided in this application embodiment may also include an audio content library and an audio advertising library; of course, the audio content library and / or audio advertising library can also be integrated into the cloud platform.

[0078] In this embodiment, the cloud can be software or services of a cloud platform, or software or services deployed on nodes in a network, such as edge nodes. The client can be a terminal device or an application, for example, the application runs on a terminal device for user use.

[0079] When a client uses a podcast-type application, it can retrieve audio programs from an audio content library. While playing an audio program, it can send an advertising request to the cloud. The cloud can then determine an audio ad from the audio ad library that meets the user's personalized needs based on the information related to the audio program carried in the ad request and the user's characteristics. This ad is then sent to the client for the client to display during audio program playback.

[0080] In this application, audio advertising refers to advertising played via audio. An advertising request is used to request audio advertising from the cloud.

[0081] In this application, user characteristics may include user profiles and user behavior characteristics. User profiles may include basic user information, such as gender, age, and hobbies. User behavior characteristics may include user behavior information such as clicking, saving, and commenting on historical audio programs.

[0082] The audio advertising system provided in this application embodiment can determine audio advertisements by combining user characteristics. The audio advertisements determined in this way can better meet the personalized needs of users and improve the delivery effect of audio advertisements.

[0083] In this embodiment, the information related to the audio program may include the identifier of the audio program, and may also include the identifier of the advertising slot pre-mined from the audio program. This allows the cloud to determine the audio advertisement for the advertising slot specified in the advertising request, further improving the matching degree between the audio advertisement and the audio program. An advertising slot refers to the time period within the audio program used to play an audio advertisement.

[0084] In this embodiment of the application, the process of mining advertising slots is usually offline, but online mining is also possible. The following describes the process in conjunction with... Figure 1B This paper introduces an audio advertising system used to identify ad slots.

[0085] like Figure 1B As shown, the audio advertising system may include a cloud platform and an audio content library. This audio content library can be integrated into the cloud, which can interact with... Figure 1A The cloud in the context can be the same device or different devices.

[0086] When mining ad slots, the cloud can obtain the audio program of the ad slot to be mined from the audio content library. Then, based on the temporal information of the audio program in speech mode and the text content after the audio program is converted into text, the cloud determines at least one ad slot. The cloud encodes the text content of each ad slot in the at least one ad slot in the previous period to obtain the vector representation of each ad slot.

[0087] In this embodiment of the application, the time-domain information may include amplitude (amplitude can also be described as sound intensity), the change of amplitude over time, etc.

[0088] The cloud will store the identifier and vector representation of each ad slot for the same audio program, along with the audio program itself. If the audio program is stored in an audio content library, the identifier and vector representation of each ad slot for the same audio program can be returned to the audio content library. The audio content library stores the audio program, along with the identifier and vector representation of each ad slot for that audio program. Figure 1B In this context, the audio content library can store multiple audio programs and their corresponding advertising slots, along with their identifiers and vector representations. For example, audio program 1 has x corresponding advertising slots, with the identifier and vector representation of each advertising slot corresponding to audio program 1 being advertising slot 1, vector representation 1, ..., advertising slot x, and vector representation x, respectively. Audio program M has y corresponding advertising slots, with the identifier and vector representation of each advertising slot corresponding to audio program 1 being advertising slot 1, vector representation 1, ..., advertising slot y, and vector representation y, respectively. Here, x, y, and M are all positive integers.

[0089] The ad slot mining scheme provided in this application can mine ad slots based on the temporal information of the audio program in voice mode and the text content after the audio program is converted into text format. The quality of the mined ad slots is high, and inserting audio ads into these slots usually does not affect the continuity of the audio program, thereby improving user experience and the effectiveness of audio ad delivery. Moreover, by associating and storing the audio program, the identifier of each ad slot in the audio program, and the vector representation of each ad slot in the audio content library, it is easy to quickly determine the audio ad that matches the requested target ad slot when the client sends an ad request, thus improving the efficiency of audio ad delivery.

[0090] In this embodiment of the application, the cloud can be a physical machine, a virtual machine (VM), or a container, etc. It can also be understood as an advertising system, or as a device in an advertising system.

[0091] When the client is a terminal device, this terminal device (also called user equipment, UE) is a device with wireless transceiver capabilities. It can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; it can also be deployed on water (such as ships); and it can also be deployed in the air (such as airplanes, balloons, and satellites). Terminals can be mobile phones, tablets, computers with wireless transceiver capabilities, virtual reality (VR) terminals, augmented reality (AR) terminals, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical care, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, and wireless terminals in the Internet of Things (IoT), etc.

[0092] The structure of the terminal device provided in this application embodiment can be referred to as follows: Figure 2A For a better understanding, the structure of a cloud-based device can be found in the following reference. Figure 2B To understand.

[0093] Please refer to Figure 2A This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 2A As shown, the terminal device may include a processor 101, a transceiver 102, a memory 103, and a bus 104. The processor 101, transceiver 102, and memory 103 are interconnected via the bus 104. In embodiments of this application, the processor 101 is used to control and manage the operations of the terminal device 10; for example, the processor 101 is used to control the playback of audio programs and audio advertisements. The transceiver 102 is used to support communication by the terminal device 10; for example, the transceiver 102 can execute steps of sending advertisement requests and receiving audio advertisements. The memory 103 is used to store the program code and data of the terminal device 10.

[0094] The processor 101 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, etc. The bus 104 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 2A The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0095] above Figure 2A The structure of the terminal device has been introduced. The following section will combine... Figure 2B Introduce the structure of cloud-based devices.

[0096] Figure 2B A schematic diagram of a possible logical structure of a cloud device provided for embodiments of this application. For example... Figure 2B As shown in the embodiments of this application, the cloud device 20 includes a processor 201, a communication interface 202, a memory 203, and a bus 204. The processor 201, communication interface 202, and memory 203 are interconnected via the bus 204. In the embodiments of this application, the processor 201 is used to control and manage the operations of the cloud device 20; for example, the processor 201 is used to execute the process of determining audio advertisements. The communication interface 202 is used to support communication by the cloud device 20; for example, the communication interface 202 can execute the steps of receiving advertisement requests and sending audio advertisements. The memory 203 is used to store the program code and data of the cloud device 20.

[0097] The processor 201 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, etc. The bus 204 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 2B The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0098] The method for audio advertising delivery provided in this application embodiment is described below. The cloud-based execution involved in this method can be performed by the cloud itself, or by cloud components (such as processors, chips, or chip systems).

[0099] Figure 3 This is a schematic diagram of an embodiment of the audio advertising delivery method provided in this application.

[0100] like Figure 3 As shown, one embodiment of the audio advertising delivery method provided in this application includes:

[0101] 301. The client sends an advertising request to the cloud. Correspondingly, the cloud receives the advertising request from the client.

[0102] The ad request is triggered when the client plays an audio program.

[0103] In this embodiment of the application, when the client plays an audio program, it usually means when the client is about to reach the target ad slot. Typically, an ad request is triggered at a preset time point before reaching the target ad slot. This preset time point can be 5 seconds away from the target ad slot or other time points that represent the duration.

[0104] In this embodiment of the application, the advertising request includes information about the audio program, the identifier of the target advertising slot, and user characteristics.

[0105] In this embodiment of the application, the target advertising slot is one of at least one advertising slot extracted from the audio program, such as: Figure 1BThe ad slot 1 corresponding to the audio program 1 can also be any other ad slot.

[0106] 302. The cloud determines the vector representation of the target ad slot based on the information of the audio program and the identifier of the target ad slot.

[0107] The vector representation of the target ad slot is used to describe the content in an audio program that occurs within a certain period of time before the target ad slot.

[0108] In this embodiment, the information of the audio program may be an identifier or index of the audio program. The audio program is one that the client is about to play, is currently playing, or has just finished playing. The audio program may be an audiobook, a song in audio format, a crosstalk performance, or current affairs news, etc.

[0109] In this embodiment, during the ad slot mining stage, one or more ad slots can be mined from an audio program. Each ad slot in an audio program has a unique identifier. Each ad slot has a vector representation, and the identifier of the same ad slot is associated with its vector representation. Furthermore, the representation and vector representation of at least one ad slot in each audio program are stored in association with that audio program. This information can all be stored in the audio content library of a cloud platform. The vector representation of an ad slot refers to the vector obtained by encoding the content involved in that ad slot within a certain period of time. In this application, "a period of time" can be a duration, such as 1 minute or other numerical values ​​representing duration. For example, in one implementation, the specific numerical value can be preset; in another implementation, the duration can be randomly selected within a range.

[0110] like Figure 1B As illustrated in the diagram, if the information of the audio program is audio program 1 and the identifier of the target advertising slot is advertising slot 1, then the vector representation 1 of advertising slot 1 can be determined based on audio program 1 and advertising slot 1.

[0111] 303. The cloud determines the audio ad that matches the target ad slot based on user characteristics and the vector representation of the target ad slot.

[0112] In this embodiment of the application, user characteristics can reflect the user's preference for audio program types or styles, such as the types and content of audio programs that the user likes to listen to, as well as the audio program narrators they like.

[0113] In this embodiment, because the vector representation of the ad slot can reflect the content of the audio program, when matching audio ads, audio ads that are related to the content of the audio program can be selected. Such audio ads have a high degree of integration with the audio program and will not affect the continuity of the user's listening to the audio program. The cloud then combines the user's preferences to further determine the audio ads, which can obtain the audio ads with the best matching degree to the target ad slot.

[0114] 304. The cloud sends audio ads to the client. Correspondingly, the client receives audio ads from the cloud.

[0115] Audio ads are played on the client side when the audio program is played to the target ad slot.

[0116] 305. The client plays an audio ad in the target ad slot.

[0117] In this embodiment, one or more audio advertisements strongly correlated with the audio program being listened to by the user in front of the target ad slot can be determined based on the vector representation of the target ad slot. Furthermore, the audio advertisements can be further filtered or processed based on user characteristics to obtain audio advertisements matching the target ad slot. Because the audio advertisements determined in this application have a higher degree of matching with the audio program and incorporate user characteristics, they better meet the personalized needs of users and can improve the effectiveness of audio ad delivery.

[0118] Step 303 above may include: the cloud recalls multiple audio ads from the audio ad library based on the vector representation of the target ad slot; the cloud obtains the audio ad that matches the target ad slot from the multiple audio ads based on user characteristics.

[0119] Further possibilities include: the cloud predicts the completion rate of multiple audio ads based on user characteristics and an ad ranking model, wherein the audio ad with the highest completion rate is the audio ad that matches the target ad slot, or the audio ad with the highest completion rate is the source ad of the audio ad that matches the target ad slot, and the ad ranking model is a model that takes user characteristics as input and takes completion rate as output.

[0120] In other words, in this embodiment of the application, the process of determining audio advertisements may include several parts, namely, advertisement recall, advertisement sorting and style transfer, which will be described below.

[0121] 1. Advertising recall.

[0122] Ad recall refers to retrieving multiple audio ads from an audio ad library that are related to the content described by the vector representation of the target ad slot.

[0123] 2. Advertisement sorting.

[0124] The cloud can use an ad ranking model to score each recalled audio ad, or it can sort the ads based on the scores and select the highest-scoring audio ad.

[0125] In this embodiment, the ad ranking model is a model that takes user characteristics as input and outputs completion rate. This ad ranking model can be a machine learning model; for more information on ad ranking models, please refer to [link to relevant documentation]. Figure 4 To understand. For example Figure 4 As shown, the input to this ad ranking model can include audio programs, audio ads, audio ad text (text-based audio ads), slot text (text-based ad slots), and can also include slot weights, audio ad features, audio program features, context features, etc. Additionally, the input to the ad ranking model provided in this embodiment also includes user features. Thus, after the embedding layer processes the slot weights, audio ad features, audio program features, context features, and user feature information, and performs audio encoding on the audio programs and audio ads, and text encoding on the audio ad text and slot text, all of these are input to the connection and flattening layers for processing. Finally, the completion rate of the audio ad for the user is input through the neural network. The neural network can include a convolutional neural network (DNN), a deep interest network (DIN), or a deep factorization machine (DeepFM).

[0126] Multiple recalled audio ads can be input into the ad ranking model individually, or they can be input all at once or in batches to obtain the completion rate for each audio ad. The completion rate can be understood as a score for the audio ad, with the audio ad with the highest score having the highest completion rate.

[0127] Completion rate refers to the predicted probability that an audio ad will be played in its entirety. The closer the content and style of an audio ad are to user preferences, the higher the probability of it being played in its entirety, and the better the campaign performance will be. Therefore, by predicting the completion rate of multiple recalled audio ads based on user characteristics, the audio ad with the highest completion rate can be selected as the primary audio ad, thus improving the effectiveness of audio ad campaigns.

[0128] 3. Style transfer.

[0129] In this embodiment, after the cloud determines the audio ad with the highest completion rate using an ad ranking model, it can further adjust the style of the audio ad with the highest completion rate based on the style of the audio program the user wants to play or is currently playing, and the user's characteristics, to obtain an audio ad that matches the target ad slot. This can improve the user's acceptance of audio ads, thereby improving the effectiveness of audio ad delivery.

[0130] It should be noted that after determining the audio ad with the highest completion rate, it can be directly identified as the audio ad that matches the target ad slot, or the audio ad with the highest completion rate can be used as the source ad for the aforementioned style transfer to obtain the audio ad that matches the target ad slot. This application does not limit this.

[0131] The style transfer process provided in this application's embodiments can be found in [reference needed]. Figure 5 To understand.

[0132] like Figure 5 As shown, the process may include:

[0133] 501. Separate object sounds and background music in audio ads in the cloud.

[0134] The subject voice is usually the main voice in an audio advertisement, such as the narrator in a narrated advertisement.

[0135] 502. Separate object sounds and background music from audio programs in the cloud.

[0136] 503. Obtain the style vector of the object sound in the cloud-encoded audio program.

[0137] 504. Obtain the style vector of the background music in the cloud-encoded audio program.

[0138] 505. Cloud-based encoding of user features yields style vectors representing user preferences.

[0139] 506. The cloud adjusts the object's voice in the highest-scoring audio ad based on the object's voice style vector and the user's preferred style vector.

[0140] Step 506 involves migrating the style of the object's voice in the audio ad.

[0141] 507. The cloud platform adjusts the background music in the highest-scoring audio ads based on the style vector of the background music and the style vector of user preferences.

[0142] Step 507 involves migrating the style of the background sound in the audio ad.

[0143] Style transfer can be the replacement or partial adjustment of the style of an object's sound or background music.

[0144] 508. Combine the object sound and background music after style transfer in steps 506 and 507 to obtain an audio advertisement.

[0145] It should be noted that if there is no background music in the audio advertisement or audio program, the steps related to background music processing, such as one or more of steps 501, 502, 504, 507 or 508, may not be performed.

[0146] In this embodiment, the style of the audio advertisement is adjusted to match the style of the audio advertisement, which can meet the user's style preference, improve the user experience, and thus improve the delivery effect of the audio advertisement.

[0147] The above describes the process of online audio advertising. Below, we will discuss... Figure 1B The process of mining ad slots, whether offline or online, will be described in further detail.

[0148] like Figure 6 As shown, one embodiment of the method for mining advertising slots provided in this application includes:

[0149] 601. Obtain audio programs for ad slots to be explored from the cloud.

[0150] The cloud can retrieve audio programs from the audio content library that are available for ad slots.

[0151] 602. The cloud performs detection on audio programs in voice mode.

[0152] 603. When the time domain information is amplitude, if the duration for which the amplitude of the audio program in the voice state is continuously lower than the amplitude threshold exceeds the first threshold, the cloud will determine the duration for which the amplitude is continuously lower than the amplitude threshold as the first basic advertising slot.

[0153] 604. The cloud converts audio programs into text content and records the timestamps of words in the text content, and determines the time interval between two adjacent words based on the timestamps of two adjacent words.

[0154] The first threshold and the second threshold can be the same or different.

[0155] 605. If the time interval between two adjacent words in the text content after the audio program is converted is greater than the second threshold, the cloud will determine the time interval between the two adjacent words as the second basic advertising slot.

[0156] 606. Restore punctuation marks in text content via cloud.

[0157] 607. The cloud adds weight to the basic ad slots corresponding to the ending punctuation marks.

[0158] In this embodiment of the application, the basic advertising slot refers to the advertising slot that is the union of the first basic advertising slot and the second basic advertising slot.

[0159] Of course, the first basic ad slot and the second basic ad slot can overlap.

[0160] 608. The cloud divides the text content of the recovery symbol into text segments and increases the weight of the basic advertising slots between two text segments.

[0161] The process of steps 602 to 608 above can be referred to Figure 7 To understand this, we need to look at examples.

[0162] like Figure 7 As shown, the cloud can first process audio programs ( Figure 7 The example shown can be a segment extracted from an audio program. If a segment of audio is detected with a very small amplitude (less than the amplitude threshold) and a duration exceeding the first threshold, it can be determined that the object in the audio program is not making a sound during this duration, that is, it is in a paused state. This duration can also be determined as a first basic advertising slot.

[0163] Then, the cloud can Figure 7 The audio clip shown has been converted into text using speech recognition. Figure 7 The converted text includes the phrase: "The weather is so nice today, let's go to the Summer Palace in Haidian District." The timestamps are as follows: "Today" (1), "weather" (2), "nice" (3), "we" (6), "where" (7), "play" (8), "Summer Palace" (12), "in" (13), and "Haidian District" (14). From the timestamps, we can see that timestamps 4 and 5 are missing between "nice" and "we," indicating a pause between them. If the second threshold is 1 (or any other value between 0 and 2), this can be identified as a second basic ad slot. Similarly, timestamps 9, 10, and 11 are missing between "play" and "Summer Palace," making this also a second basic ad slot. "Haidian District" can also be identified as a second basic ad slot. Figure 7 As can be seen, the first basic advertising slot largely overlaps with the second basic advertising slot between "Play" and "Summer Palace". For ease of description, the following will... Figure 7 The first and second basic advertising slots shown in the diagram are collectively referred to as basic advertising slots.

[0164] Next, we can restore the punctuation marks in the text, such as... Figure 7As shown, the text content after restoring punctuation includes: "The weather is so nice today, where should we go? The Summer Palace is in Haidian District." Because the pauses of the question mark "?" and the period "." are usually longer than those of the comma ",", the two basic ad slots corresponding to "?" and "." can be weighted up, that is, the weight of these two basic ad slots can be increased.

[0165] Furthermore, it can also be used for Figure 7 The illustrated text content is segmented. "The weather is so nice today, where shall we go?" can be divided into text segment 1, and "The Summer Palace is in Haidian District." can be divided into text segment 2. The pause at the division point between the two text segments will be longer, so the weight of the basic ad slot at the division point between the two text segments can be further increased, that is, the weight of the basic ad slot between "play" and "Summer Palace" can be further increased.

[0166] 609. The cloud selects at least one basic ad slot with the highest weight to determine as at least one ad slot for the audio program.

[0167] After performing the above operations, the basic ad slots between "Play" and "Summer Palace" will have the highest weight. If in Figure 7 If you select an ad slot from the audio clip shown, you can choose the basic ad slot between "Play" and "Summer Palace" as the ad slot for that audio clip.

[0168] 610. The cloud encodes the text content of a preset length before each ad slot in at least one ad slot to obtain a vector representation of each ad slot.

[0169] The cloud can encode the text "The weather is so nice today, where shall we go?" into a vector representation of the ad slot between "go play" and "Summer Palace".

[0170] 611. The cloud can send the identifier and vector representation of each ad slot to the audio content library for associated storage with the audio content.

[0171] In this way, when a user plays this audio clip and it reaches the ad slot between "Play" and "Summer Palace," a travel-related audio ad can be selected for delivery. The delivery process can also consider the user's city, the types of attractions the user has previously visited, and the audio ad that matches the prompt "The weather is so nice today, where should we go?" This ensures a continuous flow between the audio program and the audio ad, without disrupting the user's listening experience and increasing the effectiveness of the audio ad delivery.

[0172] To better understand the relationship between the ad mining process and the ad delivery process provided in the embodiments of this application, the following will be combined with... Figure 8The two processes are described in a combined manner.

[0173] 801. Ad slot mining based on audio content.

[0174] This process may include ad slot identification and ad slot vector representation.

[0175] The process of step 801 can be understood by referring to steps 601 to 611 above, and will not be repeated here.

[0176] 802. During the audio ad delivery process, when the client plays the audio program, it determines whether the ad slot has been reached. If so, proceed to step 803; otherwise, continue playing the audio program.

[0177] 803. When playing to the ad slot, determine whether to send an ad request. If yes, proceed to step 804; otherwise, continue playing the audio program.

[0178] If enough ads have already been played, meaning the duration of the music ads has exceeded the preset value, then subsequent ad slots will not play ads, and there is no need to send ad requests.

[0179] 804. Recall multiple audio ads from the audio ad library.

[0180] 805. Personalize the audio ad sorting for multiple audio ads.

[0181] Steps 804 and 805 can be understood by referring to the previous introductions to ad recall and ad scoring.

[0182] 806. Determine whether to run audio ads.

[0183] The decision to run an audio ad is based on the scoring results. If all audio ads score below the threshold, no audio ads will be run. If the highest-scoring audio ad scores above the threshold, then the highest-scoring audio ad will be run, and step 807 will proceed.

[0184] 807. Audio advertising style migration.

[0185] This step 807 can be understood by referring to the introduction in the previous section on style transfer.

[0186] 808. Play the audio ad after style migration, and then continue playing the audio program.

[0187] The ad slot mining process described above combines text content to generate vector representations. This allows for the identification of audio ads that better match the audio content during ad delivery. Furthermore, user characteristics are used when delivering audio ads, which better meets users' personalized needs and improves the effectiveness of audio ad delivery.

[0188] The above describes the methods for finding advertising slots and for placing audio ads. The following describes the cloud device and client in the embodiments of this application with reference to the accompanying drawings.

[0189] like Figure 9 As shown, a structure of the cloud device 90 provided in this application embodiment includes:

[0190] The receiving unit 901 is configured to receive an advertising request from a client. The advertising request includes information about the audio program, an identifier of the target advertising slot, and user characteristics. The target advertising slot is one of at least one advertising slot extracted from the audio program. The advertising request is triggered when the client plays the audio program. The receiving unit 901 can perform step 301 in the above method embodiment.

[0191] The first processing unit 902 is configured to determine the vector representation of the target advertising slot based on the information of the audio program and the identifier of the target advertising slot. The vector representation of the target advertising slot describes the content of the audio program during a period of time preceding the target advertising slot. This first processing unit 902 can execute step 302 in the above method embodiment.

[0192] The second processing unit 903 is used to obtain an audio advertisement matching the target ad slot based on user characteristics and the vector representation of the target ad slot. This second processing unit 903 can execute step 303 in the above method embodiment.

[0193] The sending unit 904 is used to send an audio advertisement to the client, which is played when the client plays an audio program and reaches the target advertisement slot. The sending unit 904 can perform step 304 in the above method embodiment.

[0194] In this embodiment, one or more audio advertisements strongly correlated with the audio program being listened to by the user in front of the target ad slot can be determined based on the vector representation of the target ad slot. Furthermore, the audio advertisements can be further filtered or processed based on user characteristics to obtain audio advertisements matching the target ad slot. Because the audio advertisements determined in this application have a higher degree of matching with the audio program and incorporate user characteristics, they better meet the personalized needs of users and can improve the effectiveness of audio ad delivery.

[0195] Optionally, the second processing unit 903 is specifically used to recall multiple audio advertisements from the audio advertisement library based on the vector representation of the target advertisement slot; and to obtain an audio advertisement that matches the target advertisement slot from the multiple audio advertisements based on user characteristics.

[0196] Optionally, the second processing unit 903 is specifically used to predict the completion rate of multiple audio ads based on user characteristics and an ad ranking model, wherein the audio ad with the highest completion rate is the audio ad that matches the target ad slot, or the audio ad with the highest completion rate is the source ad of the audio ad that matches the target ad slot, and the ad ranking model is a model that takes user characteristics as input and completion rate as output.

[0197] Optionally, the second processing unit 903 is specifically used to adjust the style of the audio ad with the highest completion rate to obtain an audio ad that matches the target ad slot when the audio ad with the highest completion rate is the source ad of the audio ad that matches the target ad slot, based on the style of the audio program and user characteristics.

[0198] Optionally, the second processing unit 903 is specifically used to adjust the object sound in the audio advertisement with the highest completion rate based on the style vector of the object sound in the audio program and the style vector of user preference. The style vector of the object sound in the audio program is obtained by encoding the object sound in the audio program, and the style vector of user preference is obtained by encoding user features. Based on the style vector of the background music in the audio program and the style vector of user preference, the background music in the audio program is adjusted based on the style vector of the background music in the audio program and the style vector of user preference. The style vector of the background music in the audio program is obtained by encoding the background music in the audio program. The adjusted object sound in the audio advertisement with the highest completion rate and the adjusted background music in the audio advertisement with the highest completion rate are fused to obtain an audio advertisement that matches the target advertisement slot.

[0199] Optionally, the first processing unit 902 is further configured to determine at least one advertising slot based on the temporal information of the audio program in the speech state and the text content after the audio program is converted into text; and to encode the text content of each advertising slot in the at least one advertising slot in the previous period of time to obtain a vector representation of each advertising slot.

[0200] Optionally, the first processing unit 902 is specifically used to determine the duration of the amplitude continuously below the amplitude threshold as a first basic advertising slot when the time domain information is amplitude. If the duration of the amplitude of the audio program in the speech state being continuously below the amplitude threshold exceeds a first threshold, then the duration of the amplitude continuously below the amplitude threshold is determined as a first basic advertising slot. If the time interval between two adjacent words in the text content after the audio program is converted is greater than a second threshold, then the time interval between two adjacent words is determined as a second basic advertising slot. The time interval between two adjacent words is determined by the timestamp of each word during text conversion. At least one advertising slot is determined from the union of the first basic advertising slot and the second basic advertising slot.

[0201] Optionally, the first processing unit 902 is specifically used to select at least one advertising slot with the largest weight from the union of the first basic advertising slot and the second basic advertising slot to determine at least one advertising slot for the audio program, wherein the weight of each advertising slot is determined by the punctuation mark and / or the segmentation position of the text segment corresponding to each advertising slot.

[0202] Optionally, user characteristics include user profiles and user behavior characteristics related to historical audio programs.

[0203] In this embodiment, the operations performed by each unit in the cloud device 90 are the same as those described above. Figures 3 to 8 The embodiments shown are similar and will not be repeated here.

[0204] like Figure 10 As shown, a structure of the client 100 provided in this application embodiment includes:

[0205] The sending unit 1001 is used to send an advertising request to the cloud when playing an audio program. The advertising request includes information about the audio program, the identifier of the target advertising slot, and user characteristics. The target advertising slot is one of at least one advertising slot extracted from the audio program.

[0206] The receiving unit 1002 is used to receive audio advertisements sent from the cloud that match the target advertisement slot.

[0207] The processing unit 1003 is used to play audio advertisements when playing audio programs to the target advertisement slot.

[0208] Optionally, user characteristics include user profiles and user behavior characteristics related to historical audio programs.

[0209] In this embodiment, the operations performed by each unit in the client 100 are the same as those described above. Figures 3 to 8 The embodiments shown are similar and will not be repeated here.

[0210] like Figure 11As shown in the embodiments of this application, another structure of the cloud device 110 is also provided, including:

[0211] The acquisition unit 1101 is used to acquire the audio program of the advertising slot to be mined.

[0212] The first processing unit 1102 is used to determine at least one advertising slot based on the temporal information of the audio program in the speech state and the text content of the audio program after it has been converted into text.

[0213] The second processing unit 1103 is used to encode the text content of each advertising slot in the at least one advertising slot within a certain period of time in order to obtain a vector representation of each advertising slot.

[0214] Optionally, the first processing unit 1102 is specifically used to determine the duration of the amplitude continuously below the amplitude threshold as a first basic advertising slot when the time domain information is amplitude. If the duration of the amplitude of the audio program in the speech state being continuously below the amplitude threshold exceeds a first threshold, then the duration of the amplitude continuously below the amplitude threshold is determined as a first basic advertising slot. If the time interval between two adjacent words in the text content after the audio program is converted is greater than a second threshold, then the time interval between two adjacent words is determined as a second basic advertising slot. The time interval between two adjacent words is determined by the timestamp of each word during text conversion. At least one advertising slot is determined from the union of the first basic advertising slot and the second basic advertising slot.

[0215] Optionally, the first processing unit 1102 is specifically used to select at least one advertising slot with the largest weight from the union of the first basic advertising slot and the second basic advertising slot to determine at least one advertising slot of the audio program. The weight of each advertising slot in the at least one advertising slot is determined by the punctuation mark and / or the segmentation position of the text segment corresponding to each advertising slot.

[0216] In this embodiment, the operations performed by each unit in the cloud device 110 are the same as those described above. Figures 3 to 8 The embodiments shown are similar and will not be repeated here.

[0217] In another embodiment of this application, a computer-readable storage medium is also provided, which stores computer-executable instructions. When the processor of the cloud device executes the computer-executable instructions, the cloud device performs the aforementioned... Figures 3 to 8 The steps performed by the cloud-based device.

[0218] In another embodiment of this application, a computer-readable storage medium is also provided, which stores computer-executable instructions. When the client's processor executes the computer-executable instructions, the client performs the aforementioned... Figures 3 to 8 The steps performed by the client.

[0219] In another embodiment of this application, a computer program product is also provided, which includes computer program code. When the computer program code is executed on a computer, the computer device performs the above-described... Figures 3 to 8 The steps performed by the cloud device or client.

[0220] In another embodiment of this application, a chip system is also provided, the chip system including one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the interface circuits are used to receive signals from the terminal's memory and send signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the terminal performs the aforementioned actions. Figures 3 to 8 The steps performed by the cloud device or client. In one possible design, the chip system may also include a memory for storing program instructions and data necessary for controlling the device. The chip system may consist of chips or may include chips and other discrete components.

[0221] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0222] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0223] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented wholly or partially through software, hardware, firmware, or any combination thereof.

[0224] When the integrated unit is implemented using software, it can be implemented in whole or in part as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

Claims

1. A method for delivering audio advertisements, characterized in that, include: The cloud receives an advertising request from the client. The advertising request includes information about the audio program, the identifier of the target advertising slot, and user characteristics. The target advertising slot is one of at least one advertising slot extracted from the audio program. The advertising request is triggered when the client plays the audio program. The cloud determines the vector representation of the target advertising slot based on the information of the audio program and the identifier of the target advertising slot. The vector representation of the target advertising slot is used to describe the content involved in the audio program during a period of time before the target advertising slot. The cloud platform obtains an audio ad that matches the target ad slot based on the user characteristics and the vector representation of the target ad slot. The cloud sends the audio advertisement to the client, and the audio advertisement is played by the client when the audio program is played to the target advertisement slot.

2. The method according to claim 1, characterized in that, The cloud platform determines audio ads that match the target ad slot based on the user characteristics and the vector representation of the target ad slot, including: The cloud platform recalls multiple audio ads from the audio ad library based on the vector representation of the target ad slot; The cloud platform obtains audio ads that match the target ad slot from the multiple audio ads based on the user characteristics.

3. The method according to claim 2, characterized in that, The cloud-based system obtains audio ads that match the target ad slot from the plurality of audio ads based on the user characteristics, including: The cloud platform predicts the completion rate of the multiple audio ads based on the user characteristics and the ad ranking model. The audio ad with the highest completion rate is the audio ad that matches the target ad slot, or the audio ad with the highest completion rate is the source ad of the audio ad that matches the target ad slot. The ad ranking model is a model that takes user characteristics as input and completion rate as output.

4. The method according to claim 3, characterized in that, When the audio ad with the highest completion rate is the source ad of the audio ad that matches the target ad slot, the method further includes: The cloud platform adjusts the style of the audio ad with the highest completion rate based on the style of the audio program and the user characteristics to obtain an audio ad that matches the target ad slot.

5. The method according to claim 4, characterized in that, The cloud platform adjusts the style of the audio ad with the highest completion rate based on the style of the audio program and the user characteristics to obtain an audio ad that matches the target ad slot, including: The cloud platform adjusts the object sound in the audio advertisement with the highest completion rate based on the style vector of the object sound in the audio program and the style vector of user preference. The style vector of the object sound in the audio program is obtained by encoding the object sound in the audio program, and the style vector of user preference is obtained by encoding the user features. The cloud platform adjusts the background music in the audio advertisement with the highest completion rate based on the style vector of the background music in the audio program and the style vector of the user's preference. The style vector of the background music in the audio program is obtained by encoding the background music in the audio program. The cloud-integrated audio of the target ad with the highest completion rate, along with the background music, is used to obtain an audio ad that matches the target ad slot.

6. The method according to any one of claims 1-5, characterized in that, Before the cloud receives the advertising request sent by the client, the method further includes: The cloud determines at least one advertising slot based on the temporal information of the audio program in voice mode and the text content after the audio program is converted into text. The cloud platform encodes the text content of each ad slot within a certain period of time in the at least one ad slot to obtain a vector representation of each ad slot.

7. The method according to claim 6, characterized in that, The method further includes: The cloud storage associates and stores the audio program, the identifier of each advertising slot, and the vector representation of each advertising slot.

8. The method according to claim 6, characterized in that, The cloud platform determines at least one advertising slot based on the temporal information of the audio program in speech mode and the text content after the audio program is converted into text, including: When the time domain information is amplitude, if the duration for which the amplitude of the audio program in voice state is continuously lower than the amplitude threshold exceeds the first threshold, then the cloud determines the duration for which the amplitude is continuously lower than the amplitude threshold as the first basic advertising slot. If the time interval between two adjacent words in the text content after the audio program is converted is greater than the second threshold, the cloud will determine the time interval between the two adjacent words as the second basic advertising slot. The time interval between the two adjacent words is determined by the timestamp of each word during text conversion. The cloud determines the at least one advertising slot from the union of the first basic advertising slot and the second basic advertising slot.

9. The method according to claim 8, characterized in that, The cloud determines the at least one ad slot from the union of the first basic ad slot and the second basic ad slot, including: The cloud platform selects at least one advertising slot with the largest weight from the union of the first basic advertising slot and the second basic advertising slot to determine at least one advertising slot for the audio program. The weight of each advertising slot is determined by the punctuation mark and / or the segmentation position of the text segment corresponding to each advertising slot.

10. The method according to any one of claims 1-5, characterized in that, The user characteristics include user profiles and user behavior characteristics related to historical audio programs.

11. A method for delivering audio advertisements, characterized in that, include: When the client plays an audio program, it sends an advertising request to the cloud. The advertising request includes information about the audio program, the identifier of the target advertising slot, and user characteristics. The target advertising slot is one of at least one advertising slot extracted from the audio program. The client receives an audio advertisement sent from the cloud that matches the target ad slot. The audio advertisement that matches the target ad slot is obtained based on user characteristics and the vector representation of the target ad slot. The vector representation of the target ad slot is determined based on the information of the audio program and the identifier of the target ad slot. The vector representation of the target ad slot is used to describe the content involved in the audio program during a period of time before the target ad slot. The client plays the audio advertisement when the audio program is played to the target advertisement slot.

12. The method according to claim 11, characterized in that, The user characteristics include user profiles and user behavior characteristics related to historical audio programs.

13. A method for identifying advertising slots, characterized in that, include: Retrieve audio programs for ad slots to be explored from the cloud; The cloud determines at least one advertising slot based on the temporal information of the audio program in voice mode and the text content after the audio program is converted into text. The cloud platform encodes the text content of each ad slot within a certain period of time in the at least one ad slot to obtain a vector representation of each ad slot.

14. The method according to claim 13, characterized in that, The method further includes: The cloud storage associates and stores the audio program, the identifier of each advertising slot, and the vector representation of each advertising slot.

15. The method according to claim 13 or 14, characterized in that, The cloud platform determines at least one advertising slot based on the temporal information of the audio program in speech mode and the text content after the audio program is converted into text, including: When the time domain information is amplitude, if the duration for which the amplitude of the audio program in voice state is continuously lower than the amplitude threshold exceeds the first threshold, then the cloud determines the duration for which the amplitude is continuously lower than the amplitude threshold as the first basic advertising slot. If the time interval between two adjacent words in the text content after the audio program is converted is greater than the second threshold, the cloud will determine the time interval between the two adjacent words as the second basic advertising slot. The time interval between the two adjacent words is determined by the timestamp of each word during text conversion. The cloud determines the at least one advertising slot from the union of the first basic advertising slot and the second basic advertising slot.

16. The method according to claim 15, characterized in that, The cloud determines the at least one ad slot from the union of the first basic ad slot and the second basic ad slot, including: The cloud platform selects at least one advertising slot with the largest weight from the union of the first basic advertising slot and the second basic advertising slot to determine at least one advertising slot for the audio program. The weight of each advertising slot is determined by the punctuation mark and / or the segmentation position of the text segment corresponding to each advertising slot.

17. A cloud-based device, characterized in that, include: A communication interface, a processor, and a memory, wherein the communication interface and the processor are coupled to the memory, the memory being used to store programs or instructions that, when executed by the processor, cause the cloud device to perform the method as described in any one of claims 1 to 10, or to perform the method as described in any one of claims 13 to 16.

18. A client application, characterized in that, include: A transceiver, a processor, and a memory, wherein the transceiver and the processor are coupled to the memory, the memory being used to store a program or instructions that, when executed by the processor, cause the client to perform the method as described in claim 11 or 12.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computer device, cause the computer device to perform the method as described in any one of claims 1 to 16.

20. A computer program product, characterized in that, The computer program product includes computer program code that, when run on a computer device, causes the computer device to perform the method as described in any one of claims 1 to 16.

21. A chip system, characterized in that, The chip system includes one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the interface circuits are used to receive signals from the memory of the cloud device and send signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the cloud device performs the method as described in any one of claims 1 to 10, or performs the method as described in any one of claims 13 to 16.

22. A chip system, characterized in that, The chip system includes one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the interface circuits are used to receive signals from the client's memory and send signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the client performs the method as described in claim 11 or 12.

23. An audio advertising system, characterized in that, include: A client and a cloud, wherein the cloud is used to perform the method according to any one of claims 1-10, and the client is used to perform the method according to claim 11 or 12.

Citation Information

Patent Citations

  • Advertisement data pushing method and device

    CN113159836A

  • System and method for advertisement augmentation via a called voice connection

    US8995427B1