Advertisement intelligent recognition method and system based on audio and video data feature analysis

By analyzing audio and video features, generating advertising tags and dynamically adjusting the bit rate, the problem of poor user experience in traditional control strategies is solved, and more efficient resource utilization and playback effects are achieved.

CN119809720BActive Publication Date: 2025-09-23HUBEI RONGHUI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411858329.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-09-23
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

Traditional advertising bit rate control strategies cannot reasonably allocate advertising content according to its characteristics, resulting in a poor user experience, possible waste of bandwidth resources, or decreased playback quality.

Method used

By performing feature analysis on advertising audio and video data, extracting audio and video feature information, generating advertising labels and determining control strategies, including bit rate control, determining advertising complexity based on audio text and spectrogram analysis, and dynamically adjusting the bit rate based on user preferences and advertising quality scores.

Benefits of technology

It achieves dynamic bit rate adjustment based on the complexity of the advertisement, improves user experience and acceptance, reduces bandwidth waste, and ensures that advertisements are played in the best state in different network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119809720B_ABST
    Figure CN119809720B_ABST
Patent Text Reader

Abstract

The present application discloses an intelligent advertisement recognition method and system based on audio and video data feature analysis, relating to the field of audio and video processing technology. The method comprises: obtaining an audio text of the advertisement audio by performing speech recognition on the advertisement audio; obtaining a first audio feature from the audio text; obtaining a spectrogram of the advertisement audio based on the acoustic signal based on a short-time Fourier transform; determining multiple key point positions of the spectrogram based on the frequency peaks of the spectrogram; extracting a second audio feature of the spectrogram based on the multiple key point positions; determining first feature information based on the first audio feature and the second audio feature; determining second feature information from the advertisement video; generating an advertisement tag based on the first feature information and the second feature information, and determining key information of the advertisement; and determining an advertisement control strategy based on the key information and the second feature information, wherein the advertisement control strategy includes a bit rate control strategy. The present application can effectively improve the user experience of watching advertisement videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of audio and video processing technology, and in particular to a method and system for intelligently identifying advertisements based on feature analysis of audio and video data. Background Art

[0002] With the rapid development of audio and video technologies, advertising formats have evolved from simple text or images to multimedia content that incorporates rich visual and auditory elements. This shift not only increases the appeal of ads but also places higher demands on advertising communication strategies, particularly in terms of network transmission and playback stability. Bitrate, a key metric for measuring the speed of audio and video data transmission, directly determines the smoothness, clarity, and user experience of ad playback. Therefore, how to rationally adjust the bitrate of ads to meet the needs of different users has become a pressing issue in the advertising communication field.

[0003] Traditional ad bitrate control strategies primarily rely on fixed bitrate settings. However, these strategies often overlook the characteristics of ad content, such as audio complexity and video dynamic range. For example, ads with high audio complexity require more bitrate to maintain sound quality, while ads with a large video dynamic range require a higher bitrate to maintain clarity and smoothness. However, fixed bitrate settings fail to appropriately allocate bitrate based on the characteristics of the ad content. For these reasons, in high-bitrate environments, excessive bitrate allocation for simple ad content can waste bandwidth resources. In low-bitrate environments, insufficient bitrate allocation for complex ad content can lead to poor playback quality, such as audio distortion and video stuttering. Therefore, traditional fixed bitrate control strategies can result in a poor user experience and severely impact user acceptance of the ad. Summary of the Invention

[0004] The embodiments of the present application provide a method and system for intelligently identifying advertisements based on feature analysis of audio and video data, which are used to improve the user experience of watching advertisement videos.

[0005] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, a method for intelligently identifying advertisements based on feature analysis of audio and video data is provided, comprising:

[0007] Get the ad video and corresponding ad audio of the ad to be tested;

[0008] Performing voice recognition on the advertisement audio to obtain an audio text of the advertisement audio;

[0009] Obtaining a first audio feature through the audio text;

[0010] Acquiring an acoustic signal of the advertising audio, and obtaining a spectrogram of the advertising audio according to the acoustic signal based on a short-time Fourier transform;

[0011] Determining a plurality of key point positions of the spectrogram according to frequency peaks of the spectrogram;

[0012] Extracting a second audio feature of the spectrogram based on the plurality of key point positions;

[0013] determining first feature information according to the first audio feature and the second audio feature;

[0014] Determining second feature information of the advertising video based on the key point positions;

[0015] extracting a first text content of the first feature information and a second text content of the second feature information, and calculating a similarity between the first text content and the second text content;

[0016] If the similarity is greater than a preset threshold, generating an advertisement tag according to the first feature information and the second feature information;

[0017] Determining key information of the advertisement according to the advertisement tag;

[0018] An advertisement regulation strategy is determined based on the key information and the second characteristic information, the advertisement regulation strategy including a bit rate regulation strategy, wherein the second characteristic information and the key information are used to characterize the advertisement complexity of the advertisement to be tested.

[0019] In a possible implementation of the first aspect, obtaining the first audio feature through the audio text includes:

[0020] Obtaining the domain type of the advertisement to be tested;

[0021] Dividing the audio text according to preset word categories to obtain multiple audio words;

[0022] determining an audio keyword from the plurality of audio words according to the field type;

[0023] Determining a corresponding preset database according to the audio keyword;

[0024] Determining the number of times the audio keyword appears in the audio text;

[0025] Obtaining a first repetition degree of the audio keyword according to the number of times the audio keyword appears in the audio text;

[0026] Determining the number of times the audio keyword appears in a preset database, wherein the preset database stores historical advertisement text keywords;

[0027] Obtaining a second repetition degree of the audio keyword according to the number of times the audio keyword appears in a preset database;

[0028] A weight value of the first repetition degree and the second repetition degree is calculated, and if the weight value is greater than a preset threshold, a feature keyword is determined, where the feature keyword is used to represent the first audio feature.

[0029] In a possible implementation of the first aspect, extracting features of the spectrogram based on the multiple key point positions to obtain the second audio feature includes:

[0030] Dividing the spectrogram according to the key point positions to obtain a plurality of spectrogram segments;

[0031] determining the frequency, amplitude, and intensity of each of the spectrogram segments;

[0032] Based on the frequency, amplitude and intensity of each of the spectrogram segments, basic feature information of the speech segment corresponding to each of the spectrogram segments is obtained;

[0033] A preset portrait generation algorithm is used to construct a corresponding voice portrait through each of the basic feature information, wherein the voice portrait includes an audio tag, and the audio tag is used to characterize the second audio feature.

[0034] In a possible implementation of the first aspect, the basic feature information of the speech segment corresponding to each spectrogram segment includes age feature, gender feature and emotion feature, wherein the age feature is determined based on the frequency and intensity of the spectrogram segment, the gender feature is determined based on the frequency and amplitude of the spectrogram segment, and the emotion feature is determined based on the intensity and amplitude of the spectrogram segment.

[0035] In a possible implementation of the first aspect, determining the first feature information according to the first audio feature and the second audio feature includes:

[0036] The semantic similarity between the first audio feature and the second audio feature is calculated, and a keyword with the semantic similarity greater than a preset similarity is used as the first feature information.

[0037] In a possible implementation manner of the first aspect, determining the second feature information based on the key point position includes:

[0038] Divide the advertisement video based on the key point positions to obtain multiple video slices;

[0039] Feature extraction is performed on each of the video slices using a preset recognition model to obtain video features, wherein the video features are used to represent the second feature information.

[0040] In a possible implementation of the first aspect, the key information includes a brand, and the video features include video resolution, clarity, color, contrast, duration, loading time, and inter-frame difference. Determining an advertising regulation strategy based on the key information and the second feature information includes:

[0041] Clustering the key information to obtain user preference characteristics for the advertisement to be tested;

[0042] According to the preference characteristics, a corresponding search algorithm is retrieved from a preset database;

[0043] Using the retrieval algorithm to retrieve a target advertising video in a user terminal, and determining the target audience category of the user based on the target advertising video;

[0044] Determining the advertisement quality score of the advertisement to be tested by using the video features of the advertisement to be tested based on a preset quality scoring algorithm;

[0045] The bit rate regulation strategy is determined according to the advertisement quality score, the brand, and the target audience category.

[0046] In a possible implementation of the first aspect, determining the bitrate control strategy based on the advertisement quality score, the brand, and the target audience includes:

[0047] Searching the brand in a preset database to obtain the corresponding brand value;

[0048] When any two of the following conditions are met: the target audience category is a first preset type, the advertisement quality score is greater than a first preset threshold, and the brand value is greater than a second preset threshold, determining the bitrate control strategy to be a first strategy, and the first strategy is used to increase the bitrate;

[0049] When any two of the following conditions are met: the target audience category is a second preset type, the advertising quality score is not greater than the first preset threshold, and the brand value is not greater than the second preset threshold, the bit rate control strategy is determined to be the second strategy, and the second strategy is used to lower the bit rate.

[0050] In a second aspect, the present application provides a machine-readable storage medium having stored thereon instructions for enabling a machine to execute the above-mentioned method for intelligent advertisement recognition based on feature analysis of audio and video data.

[0051] In a third aspect, the present application provides an intelligent advertising recognition system based on audio and video data feature analysis, comprising:

[0052] a memory configured to store instructions; and

[0053] The processor is configured to call the instructions from the memory and implement the above-mentioned advertisement intelligent recognition method based on audio and video data feature analysis when executing the instructions.

[0054] Through the above technical solution, by performing speech recognition on the advertising audio to obtain audio text, and extracting keywords from the audio text to determine the first audio feature, it is possible to help obtain the theme and core content of the advertising audio, saving time and energy. By obtaining the acoustic signal of the advertising audio and obtaining the spectrogram of the advertising audio, the spectral distribution of the advertising audio can be intuitively displayed, which helps to further analyze the characteristics of the advertising audio and determine the second audio feature of the advertising audio. Determining the first feature information through the first audio feature and the second audio feature helps to more accurately determine the feature information of the advertising audio and ensure the accuracy of the advertising audio analysis. By performing feature extraction on the advertising video to obtain the second feature information of the advertising video, the information of the advertisement can be accurately grasped through the feature information. By calculating the similarity between the first feature information and the second feature information, and generating an advertising label based on the feature information whose similarity is greater than a preset threshold, the key information of the advertisement to be tested is determined, which can better clarify the key information of the advertisement and help improve the advertising effect. Based on the ad's key information and secondary feature information, an ad control strategy is determined and the ad's bitrate is adjusted to ensure that the ad video is played at a reasonable bitrate. This not only reduces the waste of bandwidth resources caused by excessive bitrate allocation, but also ensures that the ad is presented to users in an optimal state under different network environments and device performance, thereby improving user acceptance and satisfaction with the ad. In summary, this technical solution can dynamically adjust the ad video bitrate based on the ad's complexity, effectively improving the user experience and ad acceptance.

[0055] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 A schematic diagram of the structure of an intelligent advertisement recognition method based on audio and video data feature analysis provided in an embodiment of the present application;

[0057] Figure 2 This is one of the interface display diagrams of a management platform for an intelligent advertising identification method based on audio and video data feature analysis provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation methods described herein are only used to illustrate and explain the embodiments of the present application and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0059] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0060] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0061] Figure 1 The following schematically shows a flow chart of an advertisement intelligent recognition method based on audio and video data feature analysis according to an embodiment of the present application. Figure 1 As shown, an embodiment of the present application provides an advertisement intelligent recognition method based on audio and video data feature analysis, which may include the following steps.

[0062] S101, obtaining an advertisement video and corresponding advertisement audio of an advertisement to be tested;

[0063] S102, performing voice recognition on the advertisement audio to obtain an audio text of the advertisement audio;

[0064] S103, obtaining a first audio feature through the audio text;

[0065] S104: Acquire an acoustic signal of the advertising audio, and obtain a spectrogram of the advertising audio based on the acoustic signal based on a short-time Fourier transform;

[0066] S105, determining multiple key point positions of the spectrogram according to the frequency peaks of the spectrogram;

[0067] S106. Extracting a second audio feature of the spectrogram based on the multiple key point positions;

[0068] S107. Determine first feature information based on the first audio feature and the second audio feature;

[0069] S108: Determine second feature information of the advertising video based on the key point positions;

[0070] S109, extracting the first text content of the first feature information and the second text content of the second feature information, and calculating the similarity between the first text content and the second text content;

[0071] S110: If the similarity is greater than a preset threshold, generate an advertisement tag based on the first feature information and the second feature information;

[0072] S111. Determine key information of the advertisement based on the advertisement tag;

[0073] S112. Determine an advertisement control strategy based on the key information and the second characteristic information, where the advertisement control strategy includes a bit rate control strategy. The second characteristic information and the key information are used to characterize the advertisement complexity of the advertisement to be tested.

[0074] Figure 2 The embodiment of the present invention provides an advertisement intelligent identification method management platform based on audio and video data feature analysis, wherein Figure 2 As shown in the interface display diagram, it can be seen that the spectrogram of the advertising audio and the video slices of the advertising video are analyzed to determine information such as the name, brand and category of the advertisement.

[0075] First, obtain the advertising video and the corresponding audio of the advertisement to be tested, and perform voice recognition on the advertising audio to obtain the audio text of the advertising audio. The advertisement to be tested can be an advertisement of any type of field, such as automobile field advertisement, home appliance field advertisement, etc., and this application does not limit this. It should be noted that in this embodiment, the advertisement to be tested is an advertisement that contains both video and audio. Specifically, in the advertisement to be tested, the voice content in the advertising audio can be converted into text form through a voice recognition algorithm. The voice recognition algorithm can be an algorithm based on dynamic time regularization. The algorithm based on dynamic time regularization is a method based on template matching, which performs recognition by calculating the similarity between the input voice and the template.

[0076] In one embodiment, a first audio feature can be obtained from an audio text. Specifically, the domain type of the advertisement to be tested can be determined from the audio text. For example, if one or more keywords in the advertisement are determined to be juice, the domain type of the advertisement to be tested is the beverage domain. After determining the domain type of the advertisement to be tested, the audio text can be divided according to preset word categories. In this embodiment, the preset word categories can be determined based on actual conditions. For example, the preset word categories can be two-character words or four-character words. The word categories can be set according to the domain of the advertisement. For example, if the domain of the advertisement is agriculture, the two-character words can be set to words such as fertilizer, high yield, and corn. After dividing the audio text according to the preset word categories, audio words are extracted from the audio text according to the preset two-character words or four-character words to obtain multiple audio words. After obtaining the multiple audio words, audio keywords are determined from the multiple audio words based on the domain type. In this embodiment, audio keywords can be determined based on the domain type of the advertisement. The number of times the audio keywords appear in the audio text can be first determined, and the first repetition degree of the audio keywords can be obtained based on the number of times the audio keywords appear in the audio text. In addition, the number of times the audio keywords appear in a preset database can be counted. The preset database in this embodiment stores historical advertisement text keywords. The second repetition degree of the audio keyword can be obtained by the number of times the audio keyword appears in the preset database. After obtaining the first repetition degree and the second repetition degree, the weight value of the first repetition degree and the second repetition degree can be calculated. If the weight value is greater than the preset threshold, the characteristic keyword is determined, and the characteristic keyword is used to characterize the first audio feature. In this embodiment, the first audio feature refers to the audio feature obtained by the keyword of the advertising audio. For example, if the keyword of the advertising audio is mobile phone, high-definition sound quality, ultra-long battery life and young people, the number of times "mobile phone" appears in the preset data is determined, and the second repetition degree of "mobile phone" can be calculated through the formula of the second repetition degree.

[0077] In another embodiment, audio features can also be obtained from the acoustic signal of the advertising audio. In this embodiment, the acoustic signal of the advertising audio can be converted into a spectrogram of the advertising audio using a short-time Fourier transform. The short-time Fourier transform is a mathematical transformation related to the Fourier transform that is used to determine the frequency and phase of a sine wave in a local region of a time-varying signal. The spectrogram refers to a speech spectrum graph and can be obtained by processing the received time-domain signal. Since the technique of obtaining a spectrogram using a short-time Fourier transform is widely used, this application will not further elaborate on it.

[0078] The positions of the frequency peaks in the spectrogram are determined using the spectrogram of the advertising audio, and the positions of the frequency peaks in the spectrogram are used as the multiple key point positions of the spectrogram. Based on the multiple key point positions, second audio features of the spectrogram are extracted. In this embodiment, the second audio features refer to characteristic information of the advertising audio obtained by analyzing the spectrogram. In a specific implementation, the spectrogram can be divided according to the key point positions to obtain multiple spectrogram segments, and the frequency, amplitude, and intensity of each spectrogram segment are determined. Based on the frequency, amplitude, and intensity of each spectrogram segment, basic feature information of the speech segment corresponding to each spectrogram segment is obtained. The basic feature information of the speech segment corresponding to each spectrogram segment may include age features, gender features, and emotional features. The age feature is determined based on the frequency and intensity of the spectrogram segment, the gender feature is determined based on the frequency and amplitude of the spectrogram segment, and the emotional feature is determined based on the intensity and amplitude of the spectrogram segment.

[0079] After obtaining the first audio feature and the second audio feature, the semantic similarity between the first audio feature and the second audio feature is calculated. In this embodiment, the semantic similarity between the first audio feature and the second audio feature can be calculated using a word embedding model, such as Word2Vec and GloVe. The semantic similarity is measured by mapping words to a high-dimensional vector space and calculating the cosine similarity between the words. After obtaining the semantic similarity, keywords with a semantic similarity greater than a preset similarity are used as first feature information. In this embodiment, the preset similarity can be determined based on actual conditions. The first feature information refers to feature information whose semantic similarity between the first audio feature and the second audio feature exceeds a preset threshold.

[0080] Based on the key point positions, the second feature information of the advertising video is determined. In this embodiment, the second feature information refers to the feature information obtained after feature extraction of the advertising video, and the key point position refers to the position of the frequency peak of the spectrogram. First, the position of the frequency peak in the spectrogram is determined, and the advertising video is divided based on the position of the frequency peak to obtain multiple video slices. After obtaining the video slices, feature extraction is performed on each video slice through a preset recognition model. In this embodiment, the preset recognition model can be a 3D convolutional neural network model. The 3D convolutional neural network model is a spatiotemporal extension of the 2D convolutional neural network model. It not only slides the convolution kernel in the spatial domain, but also slides in the temporal domain. It can simultaneously mine the spatial information and temporal information of the video frame, mine the features in the advertising video, and obtain video features, wherein the video features are used to characterize the second feature information. In this embodiment, the video features include video resolution, clarity, color, contrast, duration, loading time, and inter-frame differences.

[0081] After determining the first feature information of the advertising audio and the second feature information of the advertising video, the first text content of the first feature information and the second text content of the second feature information can be extracted, and the similarity between the first text content and the second text content can be calculated to facilitate the subsequent determination of the advertising label. In this embodiment, the similarity between the first text content and the second text content can be calculated through word embedding models, such as Word2Vec and GloVe, by mapping words to high-dimensional vector space and calculating the cosine similarity between words to measure semantic similarity. If the similarity is greater than a preset threshold, an advertising label is generated based on the first feature information and the second feature information. In this embodiment, the advertising label can include the type label, style label, etc. of the advertisement. After obtaining the advertising label, the key information of the advertisement is determined based on the advertising label. In this embodiment, the key information includes brand, advertising category, advertising theme, etc.

[0082] The advertising control strategy is determined based on the key information and the second feature information. In this embodiment, the advertising control strategy includes a bitrate control strategy and a scheme control strategy. The second feature information and the key information are used to characterize the complexity of the advertisement under test. In this embodiment, advertisement complexity refers to the complexity of the advertisement video content. Specifically, the key information is first clustered to obtain the user's preference characteristics for the advertisement under test. In this embodiment, the user's preference characteristics for the advertisement under test can include the user's preference for the type of advertisement, the theme of the advertisement, and the style of the advertisement. After obtaining the user's preference characteristics, a corresponding search algorithm is retrieved from a preset database. The search algorithm is then used to retrieve the target advertisement video from the user's terminal. The target audience category of the user is determined based on the target advertisement video. In this embodiment, the target audience category can be young people, middle-aged people, and the elderly. After determining the target audience category, the advertisement quality score of the advertisement under test is determined using a preset quality scoring algorithm and the video features of the advertisement under test. The bitrate control strategy is then determined based on the advertisement quality score, brand, and target audience category.

[0083] Through the above technical solution, by performing speech recognition on the advertising audio to obtain audio text, and extracting keywords from the audio text to determine the first audio feature, it is possible to help obtain the theme and core content of the advertising audio, saving time and energy. By obtaining the acoustic signal of the advertising audio and obtaining the spectrogram of the advertising audio, the spectrum distribution of the advertising audio can be intuitively displayed through the spectrogram of the advertising audio, which helps to further analyze the characteristics of the advertising audio and determine the second audio feature of the advertising audio. By determining the first feature information through the first audio feature and the second audio feature, it is possible to more accurately determine the feature information of the advertising audio and ensure the accuracy of the advertising audio analysis. By performing feature extraction on the advertising video to obtain the second feature information of the advertising video, the information of the advertisement can be accurately grasped through the feature information. By calculating the similarity between the first feature information and the second feature information, and generating an advertising label based on the feature information whose similarity is greater than a preset threshold, the key information of the advertisement to be tested is determined, which can better clarify the key information of the advertisement and help improve the advertising effect. Based on the key information and second characteristic information of the advertisement, the advertisement regulation strategy is determined and the bit rate of the advertisement is adjusted to ensure that the advertisement video is played at a reasonable bit rate. This not only reduces the waste of bandwidth resources caused by excessive bit rate allocation, but also ensures that the advertisement can be presented to users in the best state under different network environments and device performance, thereby improving user acceptance and satisfaction with the advertisement.

[0084] In one implementation of this embodiment, obtaining the first audio feature through the audio text includes:

[0085] S210, obtaining the domain type of the advertisement to be tested;

[0086] S220: Divide the audio text according to preset word categories to obtain multiple audio words;

[0087] S230, determining audio keywords from multiple audio words according to the field type;

[0088] S240, determining a corresponding preset database according to the audio keyword;

[0089] S250, determining the number of times the audio keyword appears in the audio text;

[0090] S260: Obtain a first repetition degree of the audio keyword based on the number of times the audio keyword appears in the audio text;

[0091] S270: Determine the number of times the audio keyword appears in a preset database, where historical advertisement text keywords are stored.

[0092] S280, obtaining a second repetition degree of the audio keyword according to the number of times the audio keyword appears in the preset database;

[0093] S290: Calculate weight values ​​of the first repetition degree and the second repetition degree. If the weight values ​​are greater than a preset threshold, determine a feature keyword, where the feature keyword is used to represent the first audio feature.

[0094] First, the field type of the advertisement to be tested is obtained. In this embodiment, the field type of the advertisement to be tested can be obtained through the audio text of the advertisement. For example, through the audio text of the advertisement, the keywords of the advertisement are determined to be plants and fertilizers, and the field type of the advertisement to be tested is determined to be the agricultural field. After determining the field type of the advertisement to be tested, the audio text is divided according to preset word categories. In this embodiment, the preset word categories can be determined according to actual conditions. For example, the preset word categories can be two-character words or four-character words. The word categories can be set according to the field of the advertisement. For example, if the field of the advertisement is the agricultural field, the preset two-character words can be set to words such as fertilizer, high yield, and corn, and the preset four-character words can be set to four-character words such as farmland promotion and green advertising. The audio text is divided according to the preset word categories to obtain multiple audio words. That is to say, after obtaining the audio text, the audio text is divided according to the preset word types, such as two-character words or four-character words. For example, the audio text of the advertisement is "Health is the cornerstone of life, and our health products are your right-hand man to protect your health. With carefully selected natural ingredients and scientific proportions, our health products will create a healthy new life for you. Take care of yourself and start by choosing our health products, so that health is no longer a luxury. According to the field type of the audio text, the preset word categories can be health, health products, etc. The audio text is divided according to the preset word categories, and the audio words obtained can be health, life, health products, protection, and scientific proportions.

[0095] According to the determined audio words and field types, audio keywords are determined among multiple audio words. Specifically, audio words with repeated semantics are determined among the audio words, and the audio words with repeated semantics are determined as audio keywords. For example, health care products and health appear repeatedly in the audio words. Therefore, health care products and health are audio keywords among multiple audio words.

[0096] To more accurately determine the first audio feature of an advertisement, after determining the audio keyword, the corresponding preset database can be determined based on the audio keyword. In this embodiment, the preset database can be determined based on the field type of the audio keyword. For example, if the audio keyword is medicine, the preset database can be determined to be a text database in the health field. After obtaining the corresponding preset database, the number of times the audio keyword appears in the audio text is determined. The more times the audio keyword appears in the audio text, the more important the audio keyword is. Based on the number of times the audio keyword appears in the audio text, the first repetition degree of the audio keyword is determined. The formula for determining the first repetition degree of the audio keyword is as follows:

[0097]

[0098] Among them, n i,j Indicates the number of times the audio keyword appears in the audio text, TF i,j Indicates the first repetition degree of the audio keyword;

[0099] After determining the first repetition degree of the audio keyword, the number of times the audio keyword appears in a preset database is determined. The preset database stores historical advertisement text keywords. That is, the number of times the audio keyword appears in the preset database is determined, and the second repetition degree of the audio keyword is obtained based on the number of times the audio keyword appears in the preset database. The formula for determining the second repetition degree of the audio keyword is as follows:

[0100]

[0101] Where N represents the number of historical advertising text keywords stored in the preset database, n i Indicates the number of times the audio keyword appears in the historical advertising text keywords stored in the preset database, IDF i,j Indicates the second repetition degree of the audio keyword;

[0102] After obtaining the first repetition degree and the second repetition degree of the audio keyword, the weight value of the first repetition degree and the second repetition degree is calculated, and the weight value of the audio keyword is obtained by multiplying the first repetition degree and the second repetition degree. The formula is:

[0103] W i,j =TF i,j ×IDF i,j

[0104] Among them, W i,j Indicates the weight value of the audio keyword;

[0105] When the weight value of the audio keyword is greater than a preset threshold, a feature keyword can be determined. The feature keyword is used to characterize the first audio feature. In this embodiment, the preset threshold can be determined according to actual conditions. The feature keyword is determined by the audio keyword, and the audio keyword greater than the preset threshold is used as the feature keyword.

[0106] By determining the audio keywords of the audio text and then determining the first audio feature, the keyword features of the advertising audio can be accurately obtained, which helps to accurately grasp the information of the advertising audio, and can more accurately target the target users, ensuring that the advertising information can better meet the needs of users and improve the user experience.

[0107] In one implementation of this embodiment, extracting features of the spectrogram based on multiple key point positions to obtain the second audio features includes:

[0108] S310, dividing the spectrogram according to the key point positions to obtain a plurality of spectrogram segments;

[0109] S320, determining the frequency, amplitude, and intensity of each spectrogram segment;

[0110] S330, based on the frequency, amplitude and intensity of each spectrogram segment, obtaining basic feature information of the speech segment corresponding to each spectrogram segment;

[0111] S340. Use a preset portrait generation algorithm to construct a corresponding voice portrait through each basic feature information, wherein the voice portrait includes an audio tag, and the audio tag is used to represent the second audio feature.

[0112] Based on the multiple key point positions, features of the spectrogram are extracted to obtain the second audio features. Specifically, the spectrogram is first divided according to the key point positions. The spectrogram refers to a speech spectrum graph, which is a spectrum graph obtained by processing the received time domain signal. In this embodiment, the multiple key point positions refer to the positions of frequency peaks in the spectrogram. By dividing the spectrogram according to the positions of the frequency peaks, a plurality of spectrogram segments can be obtained.

[0113] After obtaining multiple spectrogram segments, the frequency, amplitude, and intensity of each spectrogram segment are determined. In a spectrogram, frequency is typically represented on the horizontal axis, demonstrating the presence of different frequency components in the speech signal. Amplitude refers to the amplitude of the signal at different frequencies, reflecting the signal's strength. Intensity, closely related to amplitude, is typically used to describe the energy or power of a signal. In a spectrogram, intensity can be intuitively represented by amplitude; a larger amplitude indicates a higher signal strength at that frequency.

[0114] The frequency, amplitude, and intensity of each spectrogram segment are used to obtain the basic characteristic information of the speech segment corresponding to each spectrogram segment. That is, the frequency, amplitude, and intensity of each spectrogram segment can be determined by analyzing the spectrogram segments, and the basic characteristic information is determined based on the frequency, amplitude, and intensity. In this embodiment, the basic characteristic information includes age, gender, and emotional characteristics. Age characteristics are determined by analyzing the frequency and intensity of the spectrogram segments. For example, young people have tight vocal cords, producing high-pitched voices and hearing higher-frequency sound waves, while elderly people have loose vocal cords, producing low-pitched voices and typically only hearing lower-frequency sound waves. Therefore, a higher frequency in the spectrogram segment likely indicates a younger age. Generally speaking, with age, speech may become hoarser and its intensity may decrease. Therefore, a higher intensity in the spectrogram segment likely indicates a younger age. Gender characteristics can be determined by the frequency and amplitude of the spectrogram segments. For example, female voices generally have higher frequencies, while male voices generally have lower frequencies. This is due to differences in the structure of male and female vocal cords. Women's vocal cords are shorter and thinner, vibrating at a higher frequency; men's vocal cords are longer, wider, and thicker, vibrating at a lower frequency. Amplitude represents the strength of a signal. Generally speaking, male voices are stronger than female voices. Therefore, spectrogram segments with higher frequencies and lower intensities are female, while spectrogram segments with lower frequencies and higher intensities are male. Emotional characteristics can be determined by the intensity and amplitude of spectrogram segments. In other words, intensity generally reflects the energy level of the speech signal and is related to the speaker's emotional expression. High intensity may indicate strong emotions, such as anger or excitement, while low intensity may indicate calmness or sadness. Amplitude represents the amplitude of the signal's vibrations and is related to the loudness and clarity of the speech. Different amplitudes may correspond to different emotional states. For example, a high pitch may indicate surprise or happiness, while a low pitch may indicate sadness or dissatisfaction. Therefore, emotional characteristics can be determined by the intensity and amplitude of spectrogram segments.

[0115] After obtaining the basic feature information of the speech segment corresponding to each spectrogram segment, a preset portrait generation algorithm is used to construct a corresponding speech portrait based on each basic feature information. The preset portrait generation algorithm can be a generative adversarial network algorithm. The generative adversarial network algorithm consists of a generator and a discriminator. Through adversarial training, the generator gradually learns to generate high-quality images. Using the generative adversarial network algorithm, a corresponding speech portrait is constructed based on each basic feature information. Using artificial intelligence technology, by analyzing an individual's speech characteristics, various information about the speaker can be identified, such as emotion, social status, upbringing, age, weight, height, and facial features. By constructing a speech portrait, which includes an audio label, the label of each spectrogram segment is determined, thereby determining the second audio feature. In this embodiment, the second audio feature refers to the characteristic information of the advertising audio obtained by analyzing the spectrogram.

[0116] By dividing the key point positions of the spectrogram, we obtain spectrogram segments and determine the age, gender and emotional characteristics of each spectrogram segment. Based on the basic feature information, we generate a voice portrait to obtain the second audio feature, which can accurately determine the information of the advertising audio and help to more accurately determine the audio label of the advertisement.

[0117] In one implementation manner of this embodiment, determining the first feature information according to the first audio feature and the second audio feature includes:

[0118] S410: Calculate the semantic similarity between the first audio feature and the second audio feature, and use keywords with a semantic similarity greater than a preset similarity as first feature information.

[0119] The first feature information is determined based on the first audio feature and the second audio feature. Specifically, by calculating the semantic similarity between the first audio feature and the second audio feature, the keywords with semantic similarity greater than the preset similarity are used as the first feature information. In this embodiment, the first audio feature refers to the audio feature obtained through the keyword of the advertising audio; the second audio feature refers to the feature information of the advertising audio obtained by analyzing the spectrogram; the first feature information refers to the feature information whose semantic similarity in the first audio feature and the second audio feature exceeds the preset threshold; the preset similarity can be determined according to the actual situation. The keywords with semantic similarity greater than the preset similarity in the first audio feature and the second audio feature are retained as the first feature information. The semantic similarity between the first audio feature and the second audio feature can be calculated by using a word embedding model, such as Word2Vec and GloVe, by mapping words to a high-dimensional vector space and calculating the cosine similarity between words to measure the semantic similarity.

[0120] By calculating the semantic similarity between the first audio feature and the second audio feature, and using keywords with a semantic similarity greater than a preset similarity as the first feature information, the feature information of the advertising audio can be accurately obtained, which helps to better grasp the information of the advertisement, improve user acceptance and satisfaction with the advertisement, and optimize the user experience.

[0121] In one implementation of this embodiment, determining the second feature information based on the key point position includes:

[0122] S510, dividing the advertisement video based on key point positions to obtain multiple video slices;

[0123] S520: Perform feature extraction on each video slice using a preset recognition model to obtain video features, where the video features are used to represent the second feature information.

[0124] Based on the key point locations, the second feature information is determined. Specifically, the advertisement video is first divided based on the key point locations to obtain multiple video slices. In this embodiment, the key point locations are the locations of frequency peaks in the spectrogram. The advertisement video is divided based on the locations of frequency peaks in the spectrogram to obtain multiple video slices. Video slicing is the process of dividing a complete video file into multiple small segments that can be independently processed, transmitted, and played.

[0125] After obtaining multiple video slices, feature extraction is performed on each video slice using a preset recognition model. In this embodiment, the preset recognition model is a 3D convolutional neural network model. The 3D convolutional neural network model is a spatiotemporal extension of the 2D convolutional neural network model. It not only slides the convolution kernel in the spatial domain, but also in the temporal domain. It can simultaneously mine the spatial information and temporal information of the video frame, mine the features in the advertising video, and obtain video features. The 3D convolutional neural network model is used to continuously mine and extract video features from the advertising video to determine the video features of each video slice. In this embodiment, the video features include video resolution, clarity, color, contrast, duration, loading time, and inter-frame difference. That is, the resolution, clarity, color, contrast, duration, loading time, and inter-frame difference in each video slice are separately confirmed. The determined video features are used to characterize the second feature information. In this embodiment, the second feature information refers to the feature information obtained after feature extraction of the advertising video.

[0126] By dividing the advertising video, video slices are obtained, and feature extraction is performed on each video slice to obtain video features. By analyzing the advertising video, the second feature information of the advertisement can be better determined, which is helpful for subsequent analysis of the advertisement and improves the processing efficiency of advertising video analysis.

[0127] In one implementation of this embodiment, the key information includes the brand, and the video features include video resolution, clarity, color, contrast, duration, loading time, and inter-frame difference. Based on the key information and the second feature information, an advertising control strategy is determined, including:

[0128] S610: Clustering key information to obtain user preference characteristics for the advertisement to be tested;

[0129] S620: searching a preset database for a corresponding search algorithm based on the preference characteristics;

[0130] S630: Using a retrieval algorithm to retrieve a target advertising video from the user's user terminal, and determining the target audience category of the user based on the target advertising video;

[0131] S640: Determine an advertisement quality score of the advertisement to be tested based on a preset quality scoring algorithm and the video features of the advertisement to be tested;

[0132] S650: Determine a bitrate control strategy based on the advertisement quality score, brand, and target audience category.

[0133] Based on the key information and the second feature information, an advertising control strategy is determined. In this embodiment, the key information may include brand, ad category, ad theme, etc. In this embodiment, the advertising control strategy includes a bitrate control strategy and an advertising scheme control strategy. The bitrate control strategy refers to a strategy for adjusting the data traffic used by the advertising video per unit time. The scheme control strategy involves gaining a deep understanding of the target audience's age, gender, location, interests, and hobbies to develop a more effective advertising strategy. First, the key information is clustered to obtain the user's preference characteristics for the tested ads. In this embodiment, the user's preference characteristics for the tested ads are mainly reflected in content preferences, personal characteristics, advertising effectiveness, and scene associations. Cluster analysis is an analysis that groups a collection of physical or abstract objects into multiple clusters composed of similar objects. Clustering the key information obtains the user's preference characteristics for the tested ads. For example, if the key information of the advertising video is middle-aged and elderly, health products, and health, it can be determined that the user's preference characteristics for the tested ads are that they tend to watch advertising videos about health products for middle-aged and elderly people, and advertising videos that help maintain good health.

[0134] After determining the user's preference characteristics for the tested advertisement, a corresponding search algorithm is retrieved from a preset database based on the preference characteristics. Users have different requirements for advertising videos when viewing them. For example, older people tend to watch health-related advertisements and have lower requirements for clarity. Young people tend to watch higher-quality advertisements and have higher requirements for clarity, preferring high-definition advertisements. High-definition advertisements can increase young people's desire to purchase the products featured in the advertisements, thereby increasing product sales. Therefore, users have inconsistent preferences for advertisements when viewing them. Based on the user's preference characteristics, a corresponding search algorithm is retrieved from a preset database. A user-specific search algorithm is then used to determine the user's target advertisement video. For example, young people are more concerned with advertisement clarity, the type of product featured in the advertisement, and the duration of the advertisement. Therefore, a search algorithm is determined from a preset database based on the clarity, type of product featured, and duration of the advertisement. The user's target advertisement video is then retrieved from the user terminal using the specific algorithm.

[0135] A retrieval algorithm is used to retrieve a target advertising video from the user's terminal, and the user's target audience category is determined based on the target advertising video. Specifically, the user's specific algorithm is used to retrieve the target advertising video from the user's terminal. For example, if the user's preference is the type of product in the advertising video, the user's specific retrieval algorithm is used to determine the advertising video for the user's preferred brand of product and play it for the user. After determining the user's target advertising video, the user's target audience category is determined based on the target advertising video. Specifically, if the user's target advertising video is for fast-moving consumer goods (FMCG), the user's target audience category can be determined to be FMCG users, indicating that the user prefers to watch and purchase FMCG products.

[0136] The quality score of the ad under test is determined based on the video features of the ad under test using a preset quality scoring algorithm. In this embodiment, the quality scoring algorithm is preset and used to evaluate the ad quality score, which can reflect user satisfaction with the ad video. The criteria for judging the ad video in the preset quality scoring algorithm can be determined based on actual conditions. The quality score of the ad under test is determined using the preset quality scoring algorithm.

[0137] After determining the quality score of the advertisement to be tested, the bit rate control strategy is determined based on the advertisement quality score, brand and target audience category. In this embodiment, key information includes brand, advertisement category, advertisement theme, etc. The bit rate control strategy of the advertisement video is determined based on the quality score of the advertisement to be tested, the value of the brand and the target audience category. For example, young people in the target audience category have relatively high requirements for the bit rate of the advertisement video, and the brand value is also relatively high for high-end luxury goods. When the advertisement quality score is relatively high, the bit rate of the advertisement video needs to be increased to ensure the viewing effect of the user.

[0138] By capturing user preferences for the ads being tested, we can accurately identify the ad videos they prefer. Based on these preferences, we can retrieve the corresponding search algorithm from a pre-set database to accurately obtain the user's target ad video. By determining the bitrate control strategy based on ad quality rating, brand, and target audience category, we can precisely determine the bitrate adjustment value for the user's viewing experience, ensuring a high level of user satisfaction.

[0139] In one implementation of this embodiment, a rate control strategy is determined based on the advertisement quality score, brand, and target audience, including:

[0140] S710, searching the brand in a preset database to obtain the corresponding brand value;

[0141] S720: When any two of the following conditions are met: the target audience category is a first preset type, the advertisement quality score is greater than a first preset threshold, and the brand value is greater than a second preset threshold, determining the bitrate control strategy to be a first strategy for increasing the bitrate;

[0142] S730: When any two of the following conditions are met: the target audience category is the second preset type, the advertisement quality score is not greater than the first preset threshold, and the brand value is not greater than the second preset threshold, the bit rate control strategy is determined to be the second strategy, and the second strategy is used to lower the bit rate.

[0143] The bitrate control strategy is determined based on the ad quality score, brand, and target audience. Specifically, the brand is searched in a preset database to obtain the corresponding brand value. Brand value refers to the consumer's comprehensive evaluation and recognition of the brand value, reflecting the brand's comprehensive image and appeal in the minds of consumers.

[0144] After obtaining the brand value, first, determine the target audience category. When the target audience category is the first preset type, if any two of the following conditions are met: the advertising quality score is greater than the first preset threshold and the brand value is greater than the second preset threshold, determine the bitrate control strategy as the first strategy, which is used to increase the bitrate.

[0145] In this embodiment, the first preset type can be young people or middle-aged people, and the first preset threshold and the second preset threshold can be determined according to actual conditions. For example, when it is determined that the target audience category is 20-year-old young people, the advertising quality score is greater than the first preset threshold and the brand value is greater than the second preset threshold, the first strategy is executed to increase the bit rate, improve the user's perception of the advertisement, and maximize the user's interest in the advertisement. In this embodiment, the first strategy is to increase the bit rate of the advertising video. The bit rate of the advertising video can be a fixed bit rate. For example, if the current bit rate is 5Mbps, if the target audience category is the first preset type, the advertising quality score is greater than the first preset threshold and the brand value is greater than the second preset threshold, the current bit rate is increased to 15Mbps. At this time, 15Mbps is a fixed bit rate, which is suitable for high-quality video playback scenarios. Similarly, if the current bit rate is 15Mbps, the target audience category is the first preset type, the advertising quality score is greater than the first preset threshold, and the brand value is greater than the second preset threshold, the first strategy is still to increase the current bit rate to 15Mbps. However, since the current bit rate is the same as the fixed bit rate, in actual implementation, the first strategy can be not implemented.

[0146] It should be noted that for the following three conditions, only at least two of them need to be met. The three conditions are: (1) the target audience category is the first preset type; (2) the advertising quality score is greater than the first preset threshold; (3) the brand value is greater than the second preset threshold.

[0147] In another embodiment, the bit rate of the advertising video can be dynamically adjusted based on the above three conditions. The population corresponding to the target audience category can be divided into m age intervals according to age, the advertising quality score can be divided into n intervals, and the brand value can be divided into p intervals. Taking Table 1 below as an example, it can be seen that the advertising video bit rate can be dynamically adjusted based on the target audience category age interval, the advertising quality score, and the brand value. Table 1 is as follows:

[0148] Target audience category and age range Ad quality score > 75 Brand value > 100 Bitrate setting Youth [18-30] 88 120 15Mbps Youth [30-44] 80 150 20Mbps Middle-aged [45-50] 86 200 15Mbps Middle-aged [50-59] 95 120 10Mbps

[0149] It should be noted that Table 1 is only a partial illustrative description and is not intended to limit the present application.

[0150] In a specific implementation, the bit rate corresponding to the target audience category age range, advertisement quality score and brand value can be obtained by querying the preset bit rate database.

[0151] Secondly, when the target audience category is the second preset type, if any two of the following conditions are met: the advertising quality score is not greater than the first preset threshold, and the brand value is not greater than the second preset threshold, the bit rate control strategy is determined to be the second strategy, which is used to lower the bit rate.

[0152] In this embodiment, the second preset type can be elderly people over 59 years old, and the first preset threshold and the second preset threshold can be determined according to actual conditions. For example, when the target audience category is determined to be middle-aged people aged 60, and the advertising quality score is less than the first preset threshold and the brand value is less than the second preset threshold, then the second strategy is executed to lower the bit rate, and optimize the transmission efficiency of the advertising video and the various storage requirements of the user for storing the advertising video while ensuring the quality of the advertising video. In this embodiment, the second strategy is used to lower the bit rate, and the bit rate of the advertising video can be a fixed bit rate. For example, if the current bit rate is 15Mbps, if the target audience category is the second preset type, the advertising quality score is less than the first preset threshold and the brand value is less than the second preset threshold, the current bit rate is lowered to 10Mbps. At this time, 10Mbps is a fixed bit rate, which is suitable for playback scenarios with limited network speed. Similarly, if the current bit rate is 10Mbps, and the target audience category is the second preset type, the advertising quality score is less than the first preset threshold, and the brand value is less than the second preset threshold, the second strategy is still to lower the current bit rate to 10Mbps. However, since the current bit rate is the same as the fixed bit rate, in actual implementation, the second strategy can be not implemented.

[0153] It should be noted that for the following three conditions, only at least two of them need to be met. The three conditions are: (1) the target audience category is the second preset type; (2) the advertising quality score is less than the first preset threshold; (3) the brand value is less than the second preset threshold.

[0154] In another embodiment, the bit rate of the advertising video can be dynamically adjusted based on the above three conditions. The population corresponding to the target audience category can be divided into m age intervals according to age, the advertising quality score can be divided into n intervals, and the brand value can be divided into p intervals. Taking Table 2 below as an example, it can be seen that the advertising video bit rate can be dynamically adjusted based on the target audience category age interval, the advertising quality score, and the brand value. Table 2 is as follows:

[0155] Target audience category and age range Ad quality score <75 Brand value <100 Bitrate setting Old age [60-64] 70 99 10Mbps Elderly [65-70] 72 80 15Mbps Old age [71-75] 69 97 13Mbps Old age [76-80] 74 95 18Mbps

[0156] It should be noted that Table 2 is only a partial illustrative description and is not intended to limit the present application solution.

[0157] In a specific implementation, the bit rate corresponding to the target audience category age range, advertisement quality score and brand value can be obtained by querying the preset bit rate database.

[0158] Determining the adjusted bit rate of an ad based on brand value, target audience type, and ad quality ratings not only ensures ad video quality but also saves storage space and bandwidth costs, improves ad delivery efficiency, enhances the user viewing experience, adapts to different scenario needs, and reduces costs. This is of great significance.

[0159] An embodiment of the present application also provides a machine-readable storage medium having stored thereon instructions for enabling a machine to execute the above-mentioned method for intelligent advertisement recognition based on feature analysis of audio and video data.

[0160] An embodiment of the present application further provides an electronic device, including:

[0161] a memory configured to store instructions; and

[0162] The processor is configured to call instructions from the memory and implement the above-mentioned advertisement intelligent recognition method based on audio and video data feature analysis when executing the instructions.

[0163] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0164] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0165] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The function specified in one or more boxes.

[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0167] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0168] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0169] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0170] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0171] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. An intelligent advertisement recognition method based on audio and video data feature analysis, characterized in that: include: Get the ad video and corresponding ad audio of the ad to be tested; Performing voice recognition on the advertisement audio to obtain an audio text of the advertisement audio; Obtaining a first audio feature through the audio text; Acquiring an acoustic signal of the advertising audio, and obtaining a spectrogram of the advertising audio according to the acoustic signal based on a short-time Fourier transform; Determining a plurality of key point positions of the spectrogram according to frequency peaks of the spectrogram; Extracting a second audio feature of the spectrogram based on the plurality of key point positions; determining first feature information according to the first audio feature and the second audio feature; Determining second feature information of the advertising video based on the key point positions; extracting a first text content of the first feature information and a second text content of the second feature information, and calculating a similarity between the first text content and the second text content; If the similarity is greater than a preset threshold, generating an advertisement tag according to the first feature information and the second feature information; Determining key information of the advertisement according to the advertisement tag; An advertisement regulation strategy is determined based on the key information and the second characteristic information, the advertisement regulation strategy including a bit rate regulation strategy, wherein the second characteristic information and the key information are used to characterize the advertisement complexity of the advertisement to be tested.

2. The method according to claim 1, characterized in that Obtaining a first audio feature through the audio text includes: Obtaining the domain type of the advertisement to be tested; Dividing the audio text according to preset word categories to obtain multiple audio words; determining audio keywords from the plurality of audio words according to the field type; Determining a corresponding preset database according to the audio keyword; Determining the number of times the audio keyword appears in the audio text; Obtaining a first repetition degree of the audio keyword according to the number of times the audio keyword appears in the audio text; Determining the number of times the audio keyword appears in a preset database, wherein the preset database stores historical advertisement text keywords; Obtaining a second repetition degree of the audio keyword according to the number of times the audio keyword appears in a preset database; A weight value of the first repetition degree and the second repetition degree is calculated, and if the weight value is greater than a preset threshold, a feature keyword is determined, where the feature keyword is used to represent the first audio feature.

3. The method according to claim 2, characterized in that The extracting the features of the spectrogram based on the multiple key point positions to obtain the second audio features includes: Dividing the spectrogram according to the key point positions to obtain a plurality of spectrogram segments; determining the frequency, amplitude, and intensity of each of the spectrogram segments; Based on the frequency, amplitude and intensity of each of the spectrogram segments, basic feature information of the speech segment corresponding to each of the spectrogram segments is obtained; A preset portrait generation algorithm is used to construct a corresponding voice portrait through each of the basic feature information, wherein the voice portrait includes an audio tag, and the audio tag is used to characterize the second audio feature.

4. The method according to claim 3, characterized in that The basic feature information of the speech segment corresponding to each spectrogram segment includes age feature, gender feature and emotion feature, wherein the age feature is determined based on the frequency and intensity of the spectrogram segment, the gender feature is determined based on the frequency and amplitude of the spectrogram segment, and the emotion feature is determined based on the intensity and amplitude of the spectrogram segment.

5. The method according to claim 3, characterized in that The determining first feature information according to the first audio feature and the second audio feature includes: The semantic similarity between the first audio feature and the second audio feature is calculated, and a keyword with the semantic similarity greater than a preset similarity is used as the first feature information.

6. The method according to claim 3, characterized in that The determining of the second feature information based on the key point position includes: Divide the advertisement video based on the key point positions to obtain multiple video slices; Feature extraction is performed on each of the video slices using a preset recognition model to obtain video features, wherein the video features are used to represent the second feature information.

7. The method according to claim 6, characterized in that The key information includes brand, the video features include video resolution, clarity, color, contrast, duration, loading time, and inter-frame difference, and determining an advertising regulation strategy based on the key information and the second feature information includes: Clustering the key information to obtain user preference characteristics for the advertisement to be tested; According to the preference characteristics, a corresponding search algorithm is retrieved from a preset database; Retrieving a target advertising video from a user terminal using the retrieval algorithm, and determining a target audience category of the user based on the target advertising video; Determining an advertisement quality score of the advertisement to be tested based on the video features of the advertisement to be tested based on a preset quality scoring algorithm; The bit rate regulation strategy is determined according to the advertisement quality score, the brand, and the target audience category.

8. The method according to claim 7, characterized in that Determining the bitrate control strategy according to the advertisement quality score, the brand, and the target audience includes: Searching the brand in a preset database to obtain the corresponding brand value; When any two of the following conditions are met: the target audience category is a first preset type, the advertisement quality score is greater than a first preset threshold, and the brand value is greater than a second preset threshold, determining the bitrate control strategy to be a first strategy, and the first strategy is used to increase the bitrate; When any two of the following conditions are met: the target audience category is a second preset type, the advertising quality score is not greater than the first preset threshold, and the brand value is not greater than the second preset threshold, the bit rate control strategy is determined to be the second strategy, and the second strategy is used to lower the bit rate.

9. A machine-readable storage medium, characterized in that The machine-readable storage medium stores instructions for enabling a machine to execute the advertisement intelligent recognition method based on audio and video data feature analysis according to any one of claims 1 to 8.

10. An intelligent advertising recognition system based on audio and video data feature analysis, characterized in that: include: a memory configured to store instructions; as well as The processor is configured to call the instruction from the memory and implement the advertisement intelligent recognition method based on audio and video data feature analysis according to any one of claims 1 to 8 when executing the instruction.

Citation Information

Patent Citations

  • Advertisement generation method, delivery method, advertisement generation device and delivery device

    CN112561549A

  • Advertisement management and distribution system and method based on Internet of Things

    CN118608209A