An advertisement video classification method, device and electronic equipment

CN116052052BActive Publication Date: 2026-08-21FEISHU YITU (SHANGHAI) NETWORK TECHNOLOGY CO LTD +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310089049.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-03
Publication Date
2026-08-21
Estimated Expiration
2043-02-03

AI Technical Summary

Technical Problem

[0003]现有的解决方案是在广告账户层级进行人工分类,需要人工设定广告账户层级的一级、二级、三级分类,耗费人力物力,而且人工设置的分类数据颗粒度不准确

Benefits of technology

[0024] The method, apparatus, electronic device, and computer-readable storage medium provided in this invention, compared to manual classification at the advertising account level, achieve automatic classification of e-commerce advertising video materials with high accuracy, no need for manual intervention, and saves labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052052B_ABST
    Figure CN116052052B_ABST
Patent Text Reader

Abstract

The application provides an advertisement video classification method and device, electronic equipment and computer readable storage medium, wherein the method comprises: using labeled advertisement video data as a benchmark test set; using an advertisement picture classification method to classify key frames of extracted local advertisement video materials; using advertisement video data with consistent copywriting classification data and local advertisement video classification data as video classification training sample data; training a video classification deep learning model using the video classification training sample data; performing parameter tuning on the benchmark test set to obtain an optimal video classification deep learning model; and obtaining final classification data of the advertisement video. Compared with manual classification at the advertisement account level, the method, device, electronic equipment and computer readable storage medium provided by the application realize automatic classification of e-commerce advertisement video materials, and have high classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of e-commerce technology, and more specifically, to a method, apparatus, electronic device, and computer-readable storage medium for classifying advertising videos. Background Technology

[0002] E-commerce advertising creative elements include images, videos, and copy. Before promoting products in a particular country or region, advertisers in the e-commerce industry refer to Benchark data for that area to conduct pre-campaign performance analysis, product selection, and budget allocation. Therefore, e-commerce Benchark data is crucial. E-commerce Benchark data includes performance data for advertising across different countries and product categories. Performance data includes metrics such as spend, impressions, clicks, purchases, click-through rate, conversion rate, CPS, and ROI. Product categories include primary, secondary, and tertiary product classification data. Primary categories include apparel; secondary categories include menswear, womenswear, and wedding dresses; and tertiary categories include shirts, T-shirts, suits, and trousers. Specific classification examples are shown in Table 1. Therefore, to generate e-commerce Benchark data, e-commerce advertisements need to be categorized.

[0003] The existing solution involves manual categorization at the ad account level. This requires manually setting first-, second-, and third-level categories for ad accounts, which is labor-intensive and resource-intensive. Furthermore, manually set categorization data lacks granularity. For example, if an ad account contains both ads promoting clothing and ads promoting shoes, manually setting the account level only to "clothing" will result in inaccurate categorization data and inaccurate e-commerce industry Benchark data, thus limiting the reference value of Benchark data.

[0004] More than 50% of e-commerce ads contain video content. If the video content in the ads is manually categorized, it will face problems of inefficiency and inaccurate categorization.

[0005] Table 1

[0006] Summary of the Invention

[0007] To address the existing technical problems, embodiments of the present invention provide a method, apparatus, electronic device, and computer-readable storage medium for classifying advertising videos.

[0008] In a first aspect, embodiments of the present invention provide a method for classifying advertising videos, including:

[0009] The randomly selected advertising video data was manually classified and labeled, and the labeled advertising video data was used as a benchmark test set.

[0010] The keyframes of the extracted local advertising video material are classified using an advertising image classification method to obtain the classification data of the local advertising video.

[0011] Obtain the copywriting classification data of all advertisements, and use the advertisement video data whose copywriting classification data is consistent with the classification data of the local advertisement video as the training sample data for video classification;

[0012] A deep learning model for video classification is trained using the training sample data for video classification.

[0013] Parameters were fine-tuned on the benchmark set to obtain the optimal deep learning model for video classification.

[0014] The video is classified and predicted using the optimal video classification deep learning model to obtain the final classification data for the advertising video.

[0015] Secondly, embodiments of the present invention provide a device for classifying advertising videos, comprising:

[0016] The classification and labeling module is used to manually classify and label randomly selected advertising video data, and use the labeled advertising video data as a benchmark test set.

[0017] The local classification module is used to classify the keyframes of the extracted local advertising video materials using advertising image classification methods, so as to obtain the classification data of the local advertising video.

[0018] The video sample module is used to acquire the text classification data of all advertisements, and to use the advertisement video data whose text classification data is consistent with the classification data of the local advertisement video as the training sample data for video classification.

[0019] The model training module is used to train a deep learning model for video classification using the training sample data of the video classification.

[0020] The parameter tuning module is used to perform parameter tuning on the benchmark test set to obtain the best video classification deep learning model.

[0021] The final classification module is used to classify and predict the video based on the best video classification deep learning model to obtain the final classification data of the advertising video.

[0022] Thirdly, embodiments of the present invention provide an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the transceiver, the memory, and the processor are connected via the bus, characterized in that the computer program, when executed by the processor, implements the steps in the advertising video classification method as described above.

[0023] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps in the advertising video classification method described above.

[0024] The method, apparatus, electronic device, and computer-readable storage medium provided in this invention, compared to manual classification at the advertising account level, achieve automatic classification of e-commerce advertising video materials with high accuracy, no need for manual intervention, and saves labor costs. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the background art, the accompanying drawings used in the embodiments of the present invention or the background art will be described below.

[0026] Figure 1 A flowchart illustrating a method for classifying advertising videos provided by an embodiment of the present invention is shown;

[0027] Figure 2 A schematic diagram of the deep learning model structure in step S111 of an embodiment of the present invention is shown;

[0028] Figure 3 A flowchart of the advertising image classification method provided in an embodiment of the present invention is shown;

[0029] Figure 4 A flowchart illustrating the method for obtaining copy classification data of all advertisements provided in an embodiment of the present invention is shown;

[0030] Figure 5 A schematic diagram of the structure of an advertising video classification device provided in an embodiment of the present invention is shown;

[0031] Figure 6 A schematic diagram of the structure of an electronic device for classifying advertising videos provided in an embodiment of the present invention is shown. Detailed Implementation

[0032] Those skilled in the art will understand that embodiments of the present invention can be implemented as methods, apparatuses, electronic devices, and computer-readable storage media. Therefore, embodiments of the present invention can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. Furthermore, in some embodiments, embodiments of the present invention can also be implemented as a computer program product contained in one or more computer-readable storage media, the computer-readable storage media containing computer program code.

[0033] The aforementioned computer-readable storage medium may be any combination of one or more computer-readable storage media. Computer-readable storage media include: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any combination thereof. In embodiments of the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0034] The computer program code contained in the aforementioned computer-readable storage medium may be transmitted using any suitable medium, including wireless, wire, optical fiber, radio frequency (RF), or any suitable combination thereof.

[0035] Computer program code for performing the operations of embodiments of the present invention can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The computer program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer or an external computer via any type of network, including a local area network (LAN) or a wide area network (WAN).

[0036] The embodiments of the present invention will now be described with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer-readable storage media.

[0037] It should be understood that each block of a flowchart and / or block diagram, as well as combinations of blocks in a flowchart and / or block diagram, can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine that, when executed by a computer or other programmable data processing apparatus, creates means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.

[0038] These computer-readable program instructions may also be stored in a computer-readable storage medium that enables a computer or other programmable data processing device to function in a particular manner. In this way, the instructions stored in the computer-readable storage medium produce an instruction apparatus product that includes the functions / operations specified in the blocks of a flowchart and / or block diagram.

[0039] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer-implemented process, such that the instructions that execute on the computer or other programmable data processing apparatus provide a process for implementing the functions / operations specified in the blocks of a flowchart and / or block diagram.

[0040] The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0041] Figure 1 A flowchart illustrating a method for classifying advertising videos according to an embodiment of the present invention is shown. Figure 1 As shown, the method includes:

[0042] Step S101: Manually label the randomly selected advertising video data, and use the labeled advertising video data as a benchmark test set;

[0043] 20,000 video data points were randomly sampled from the sample space of all advertising video data. These video data were then manually classified and labeled. The labeled and classified advertising video data was used as a benchmark set to evaluate the classification performance. The formula for calculating the classification performance index is: Accuracy = Number of correctly classified entries / Total number of entries in the benchmark set.

[0044] The format of the benchmark test set for manual classification and labeling is shown in Table 2.

[0045] Table 2

[0046] http: / / media.meetsocial.com / aaa.mp4 apparel Women's clothing Pants http: / / media.meetsocial.com / bbb.mp4 Shoes, boots, bags Footwear and accessories Women's shoes …… …… …… ……

[0047] Step S103: Extract multiple keyframes from the local advertisement video material. The keyframe extraction algorithm uses the perceptual hash algorithm, and the specific steps include:

[0048] Step S1031: Reduce size: Reduce the image size to 8*8 pixels, a total of 64 pixels;

[0049] Step S1032: Simplify colors: Convert the reduced image to 64 grayscale levels, meaning that all pixels have a total of 64 colors;

[0050] Step S1033: Calculate the average grayscale value: Calculate the average grayscale value of all 64 pixels;

[0051] Step S1034: Compare the grayscale of pixels: Compare the grayscale of each pixel with the average grayscale value. If the grayscale value is greater than or equal to the average grayscale value, it is recorded as 1; if the grayscale value is less than the average grayscale value, it is recorded as 0.

[0052] Step S1035: Combine the comparison results from step S1034 to form a 64-bit integer, which is the fingerprint of the image.

[0053] Step S1036: Compare the fingerprints of the preceding and following frames to see how many bits in the 64-bit fingerprint are different, i.e., calculate the Hamming distance (the Hamming distance between two strings of equal length is the number of different characters at corresponding positions in the two strings). If the Hamming distance is greater than 10, it indicates that the second image has changed significantly, meaning the next frame is a keyframe.

[0054] Step S105: Use the advertising image classification method to classify the keyframes extracted in step S103 to obtain the classification data of the local advertising video;

[0055] For example, in the first-level category (as shown in Table 1), the category score data can be represented as: {"Clothing": 0.8, "Shoes and Bags": 0.1, "Jewelry / Accessories": 0.1, "Digital Appliances": 0.0, ..., "Digital Appliances": 0.0}, where the sum of the scores is 1.0. The category with the highest score is the category data of the keyframe, so the category of this keyframe is "Clothing", and the score of "Clothing" is 0.8.

[0056] like Figure 3 As shown, in this embodiment of the invention, the advertising image classification method specifically includes:

[0057] Step S1051: Determine the classification comparison table between the advertising image classification system data and the e-commerce website classification system, and crawl the advertising image data of the corresponding categories on the e-commerce website based on the classification data in the classification comparison table;

[0058] Step S1052: Divide the set of advertising image data into a training set and a validation set according to the proportion;

[0059] Step S1053: Use the advertising image data in the training set as training sample data to train a deep convolutional neural network classification model;

[0060] Step S1054: Evaluate the performance of the deep convolutional neural network classification model based on the advertising image data in the validation set;

[0061] Step S1055: Select the deep convolutional neural network classification model with the best model performance for the classification prediction of advertising images, and obtain the advertising image classification system data.

[0062] Step S107: Use a voting method to vote on the classification data of the keyframes of the local advertising video material, and obtain the classification with the highest number of votes as the voting classification data of the advertising video;

[0063] If multiple categories have the same number of votes, the category with the highest average category score is taken as the category result of this advertisement video. The average category score is taken as the category score of this advertisement video, as shown in Table 3.

[0064] Table 3

[0065]

[0066] As shown in Table 3, Video 1 received 2 votes in the Clothing category and 2 votes in the Footwear and Bags category. The average score for Clothing was (0.8 + 0.9) / 2 = 0.85, and the average score for Footwear and Bags was (0.7 + 0.6) / 2 = 0.65. Therefore, the category of this video is Clothing. The score for this video category is the average of the scores of all categories, which is: {"Clothing": 0.85, "Footwear and Bags": 0.15, "Accessories": 0.0, "Digital Appliances": 0.0, ..., "Digital Appliances": 0.0};

[0067] Step S109: Obtain the copy classification data of all advertisements, and use the advertisement video data whose copy classification data is consistent with the classification data of the local advertisement video as the training sample data for video classification. The classification and score data format is similar to the classification data in step S105.

[0068] like Figure 4 As shown, in this embodiment of the invention, the step of obtaining the copywriting classification data of all advertisements specifically includes:

[0069] Step S1091: Obtain the content information of the advertising copy, and concatenate the content information to form the first copy information;

[0070] Step S1092: Obtain the content information of the advertising landing page, and concatenate the content information to form the second copy information;

[0071] Step S1093: Clean the URL link information of the advertising landing page and use the obtained relevant information as third-party copy information;

[0072] Step S1094: Translate the first, second, and third copy information into English, and perform binary classification prediction on the translated advertising copy;

[0073] Step S1095: Call the BERT multi-classification model to classify the effective advertising copy at each level to obtain the required advertising classification data.

[0074] Step S111: Using the training sample data obtained in step S109, train a deep learning model for video classification, and perform parameter tuning on the benchmark test set shown in Table 2 to obtain the best deep learning model for video classification.

[0075] The advertising video data is classified and predicted using the best video classification deep learning model mentioned above, and the classification and score data of all advertising videos are obtained. The format of the classification and score data is similar to that of the classification data in step S105.

[0076] The aforementioned deep learning model employs a pre-trained CNN+LSTM architecture. The pre-trained CNN model extracts feature vectors from the first 72 frames of the advertisement video, while the LSTM is used for video frame sequence classification, ultimately predicting the score for each category. The model structure is as follows: Figure 2 As shown.

[0077] The parameters of the deep learning model were tuned using a grid search approach. The hyperparameters tuned were the learning rate and the batch size. Specifically, the tuning method involved iterating through parameter pairs in the learning rate set {0.00005, 0.00008, 0.0001, 0.0002} and the batch size set {16, 32, 64}. Under these hyperparameter combinations, a CNN+LSTM deep learning model was trained. After the model training converged, its classification performance metric, Accuracy, was evaluated on a benchmark set. The model with the highest Accuracy was selected as the optimal deep learning model for video classification and used for the final classification prediction of advertising videos.

[0078] Step S113: Combining the two sets of advertising video classification data from Step S107 and Step S111, the classification scores for each category are obtained by weighting and normalizing. The category with the highest classification score is taken as the final classification data for the advertising video.

[0079] The following is an example of weighted and normalized classification: Suppose that for video 1, the classification score data obtained through step S107 is {"Clothing": 0.8, "Shoes and Bags": 0.1, "Jewelry / Accessories": 0.1, "Digital Appliances": 0.0, ..., "Digital Appliances": 0.0}; and the classification score data obtained through step S111 is {"Clothing": 0.1, "Shoes and Bags": 0.9, "Jewelry / Accessories": 0.0, "Digital Appliances": 0.0, ..., "Digital Appliances": 0.0}.

[0080] Assuming the classification result weight in step S107 is 0.9 and the classification result weight in step S111 is 0.7, the weighted sum and normalization calculation formula is as follows:

[0081] For the "Clothing" category, the final score is calculated as follows: (0.9*0.8+0.7*0.1) / (0.9+0.7)=0.494;

[0082] For the "Shoes, Boots and Bags" category, the final score is calculated as follows: (0.9*0.1+0.7*0.9) / (0.9+0.7)=0.45;

[0083] ...

[0084] Other categories can be calculated in the same way.

[0085] Since the "Clothing" category had the highest final score, Video 1 was ultimately classified as "Clothing" with a score of 0.494.

[0086] The above weighted and normalized calculation methods are expressed by the following unified formula:

[0087]

[0088] Among them, S i S is the final score for the i-th ad video category. i1 and S i2 W1 and W2 are the scores of the i-th category obtained in steps S107 and S111, respectively, and the weights of steps S107 and S111 are the scores of the i-th category.

[0089] In addition, the weights W1 and W2 are selected by enumeration parameter tuning. Different weight combinations of W1 weight set {0.6,0.7,0.8,0.9,1.0} and W2 weight set {0.6,0.7,0.8,0.9,1.0} are selected, and W1 != W2 or W1=W2=1.0. The final classification data of the benchmark test set is calculated using this weight combination, and the classification performance index Accuracy of the benchmark test set is evaluated. The weight combination with the highest classification performance index Accuracy is selected.

[0090] Step S115: For the first-level, second-level, and third-level categories, repeat steps S103-S113 above to obtain the final first-level, second-level, and third-level category data of the advertising video.

[0091] The advertising video classification method of this invention, compared with the manual classification at the advertising account level, achieves automatic classification of e-commerce advertising video materials with high classification accuracy, no need for manual intervention, and saves labor costs.

[0092] The advertising video classification method of this invention has a high accuracy rate in classifying e-commerce advertising video materials, and the obtained Benchark classification data has significant reference value.

[0093] The above text combined Figures 1 to 4 The method for classifying advertising videos according to embodiments of the present invention is described in detail below, in conjunction with... Figure 5 The present invention describes in detail an advertising video classification device according to an embodiment of the present invention.

[0094] Figure 5 A schematic diagram of the structure of an advertising video classification device provided in an embodiment of the present invention is shown. Figure 5 As shown, the classification device for the advertising video includes:

[0095] The classification and labeling module 10 is used to manually classify and label randomly selected advertising video data, and use the labeled advertising video data as a benchmark test set.

[0096] The local classification module 20 is used to classify the keyframes of the extracted local advertising video material using an advertising image classification method to obtain the classification data of the local advertising video.

[0097] The video sample module 30 is used to acquire the text classification data of all advertisements, and use the advertisement video data whose text classification data is consistent with the classification data of the local advertisement video as the training sample data for video classification.

[0098] The model training module 40 is used to train a deep learning model for video classification using the training sample data of the video classification.

[0099] The parameter tuning module 50 is used to perform parameter tuning on the benchmark test set to obtain the best video classification deep learning model.

[0100] The final classification module 60 is used to perform classification prediction on the video based on the best video classification deep learning model to obtain the final classification data of the advertising video.

[0101] Optionally, in an embodiment of the present invention, the apparatus further includes:

[0102] The data voting module 70 is used to vote on the classification data of the key frames of the local advertising video material using a voting method, and obtain the classification with the highest number of votes as the voting classification data of the advertising video;

[0103] The learning classification module 80 is used to perform classification prediction on the video based on the best video classification deep learning model to obtain the learning classification data of the advertising video.

[0104] The final classification module 60 is used to obtain classification scores for each category based on the voting classification data and the learning classification data using a weighted and normalized method, and to use the category with the highest classification score as the final classification data for the advertising video.

[0105] In an embodiment of the present invention, optionally, the local classification module 20 includes:

[0106] The crawling submodule 21 is used to determine the classification comparison table between the advertising image classification system data and the e-commerce website classification system, and crawl the advertising image data of the corresponding category of the e-commerce website according to the classification data in the classification comparison table;

[0107] Submodule 22 is used to divide the set of advertising image data into a training set and a validation set according to a ratio;

[0108] Training submodule 23 is used to train a deep convolutional neural network classification model by using the advertising image data in the training set as training sample data.

[0109] Evaluation submodule 24 is used to evaluate the model performance of the deep convolutional neural network classification model based on the advertising image data in the validation set;

[0110] The classification submodule 25 is used to select the deep convolutional neural network classification model with the best model performance for the classification prediction of the advertising image, and obtain the advertising image classification system data.

[0111] In an embodiment of the present invention, optionally, the video sample module 30 includes:

[0112] The first copywriting submodule 31 is used to obtain the content information of the advertising copy and to splice the content information as the first copywriting information.

[0113] The second copywriting submodule 32 is used to obtain the content information of the advertising landing page and to splice the content information as the second copywriting information;

[0114] The third copywriting submodule 33 is used to clean the URL link information of the advertising landing page and use the obtained relevant information as the third copywriting information.

[0115] The translation submodule 34 is used to translate the first copy information, the second copy information and the third copy information into English, and to perform binary classification prediction on the translated advertising copy;

[0116] The classification submodule 35 is used to call the BERT multi-classification model to classify effective advertising copy at various levels and obtain the required advertising classification data.

[0117] The video classification device of this invention, compared with the manual classification at the advertising account level, realizes the automatic classification of e-commerce advertising video materials with high classification accuracy, no manual intervention required, and saves labor costs.

[0118] The advertising video classification device of this invention has a high accuracy rate in classifying e-commerce advertising video materials, and the obtained Benchark classification data has significant reference value.

[0119] In addition, embodiments of the present invention also provide an electronic device, including a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor. The transceiver, the memory, and the processor are respectively connected via the bus. When the computer program is executed by the processor, it implements the various processes of the above-described advertising video classification method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0120] For details, see Figure 6 As shown, this embodiment of the invention also provides an electronic device, which includes a bus 61, a processor 62, a transceiver 63, a bus interface 64, a memory 65, and a user interface 66.

[0121] In this embodiment of the invention, the electronic device further includes: a computer program stored in a memory 65 and executable on a processor 62, wherein the computer program, when executed by the processor 62, performs the following steps:

[0122] The randomly selected advertising video data was manually classified and labeled, and the labeled advertising video data was used as a benchmark test set.

[0123] The keyframes of the extracted local advertising video material are classified using an advertising image classification method to obtain the classification data of the local advertising video.

[0124] Obtain the copywriting classification data of all advertisements, and use the advertisement video data whose copywriting classification data is consistent with the classification data of the local advertisement video as the training sample data for video classification;

[0125] A deep learning model for video classification is trained using the training sample data for video classification.

[0126] Parameters were fine-tuned on the benchmark set to obtain the optimal deep learning model for video classification.

[0127] The video is classified and predicted using the optimal video classification deep learning model to obtain the final classification data for the advertising video.

[0128] Optionally, when the computer program is executed by the processor 62, it may also perform the following steps:

[0129] The classification data of the keyframes in the local advertising video material are voted on using a voting method, and the classification with the highest number of votes is used as the voting classification data of the advertising video;

[0130] Based on the optimal video classification deep learning model, the video is classified and predicted to obtain the learning classification data of the advertising video;

[0131] Based on the voting classification data and the learning classification data, a weighted and normalized method is used to obtain the classification scores for each category, and the category with the highest classification score is used as the final classification data for the advertising video.

[0132] Alternatively, when the computer program is executed by the processor 62, it may also perform the following steps:

[0133] The advertising image classification method includes:

[0134] A classification comparison table is established between the advertising image classification system data and the e-commerce website classification system. Based on the classification data in the classification comparison table, advertising image data of the corresponding categories on the e-commerce website is crawled.

[0135] The set of advertising image data is divided into a training set and a validation set according to the proportions.

[0136] The advertising image data in the training set is used as training sample data to train a deep convolutional neural network classification model.

[0137] The performance of the deep convolutional neural network classification model is evaluated based on the advertising image data in the validation set.

[0138] The deep convolutional neural network classification model with the best performance is selected for classification prediction of the advertising images to obtain the advertising image classification system data.

[0139] Alternatively, when the computer program is executed by the processor 62, it may also perform the following steps:

[0140] The steps for obtaining the copywriting classification data of all advertisements include:

[0141] Obtain the content information of the advertising copy, and then combine and process this content information to form the first copy information;

[0142] Obtain the content information from the ad landing page, and then combine and process this content information to create the second copy.

[0143] The URL link information of the advertising landing page is cleaned, and the obtained relevant information is used as third-party copy information;

[0144] The first, second, and third copy information are translated into English, and the translated advertising copy is then subjected to binary classification prediction.

[0145] The BERT multi-classification model is used to classify effective advertising copy at various levels to obtain the required advertising classification data.

[0146] Transceiver 63 is used to receive and send data under the control of processor 62.

[0147] exist Figure 6 In the bus architecture (represented by bus 61), bus 61 may include any number of interconnected buses and bridges, and bus 61 will connect various circuits including one or more processors represented by processor 62 and memory represented by memory 65.

[0148] Bus 61 represents one or more of several types of bus architectures, including memory buses and memory controllers, peripheral buses, Accelerated Graphics Port (AGP), processors, or local buses using any bus architecture from various bus architectures. By way of example and not limitation, such architectures include: Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MCA) buses, Enhanced ISA (EISA) buses, Video Electronics Standards Association (VESA) buses, and Peripheral Component Interconnect (PCI) buses.

[0149] Processor 62 can be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processors mentioned above include: general-purpose processors, central processing units (CPUs), network processors (NPs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), complex programmable logic devices (CPLDs), programmable logic arrays (PLAs), microcontroller units (MCUs) or other programmable logic devices, discrete gates, transistor logic devices, and discrete hardware components. They can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. For example, the processor can be a single-core processor or a multi-core processor, and the processor can be integrated on a single chip or located on multiple different chips.

[0150] Processor 62 can be a microprocessor or any conventional processor. The method steps disclosed in the embodiments of the present invention can be directly executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in readable storage media known in the art, such as Random Access Memory (RAM), Flash Memory, Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), registers, etc. The readable storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0151] Bus 61 can also connect various other circuits, such as peripheral devices, voltage regulators, or power management circuits. Bus interface 64 provides an interface between bus 61 and transceiver 63, all of which are well known in the art. Therefore, the embodiments of the present invention will not be described further.

[0152] Transceiver 63 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. For example, transceiver 63 receives external data from other devices, and transceiver 63 is used to send data processed by processor 62 to other devices. Depending on the nature of the computer system, a user interface 66 may also be provided, such as a touchscreen, physical keyboard, monitor, mouse, speaker, microphone, trackball, joystick, or stylus.

[0153] It should be understood that, in embodiments of the present invention, memory 65 may further include memory remotely configured relative to processor 62, and such remotely configured memory may be connected to a server via a network. One or more portions of the aforementioned network may be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless local area network (WLAN), wide area network (WAN), wireless wide area network (WWAN), metropolitan area network (MAN), Internet, public switched telephone network (PSTN), ordinary old-style telephone service (POTS), cellular telephone network, wireless network, Wi-Fi network, and combinations of two or more of the aforementioned networks. For example, cellular telephone networks and wireless networks can be Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), WiMAX, General Packet Radio Service (GPRS), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), LTE Frequency Division Duplex (FDD), LTE Time Division Duplex (TDD), Advanced Long Term Evolution (LTE-A), Universal Mobile Telecommunications System (UMTS), Enhanced Mobile Broadband (eMBB), Massive Machine Type Communication (mMTC), Ultra Reliable Low Latency Communications (uRLLC), etc.

[0154] It should be understood that the memory 65 in the embodiments of the present invention may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory includes: read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory.

[0155] Volatile memory includes random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 65 of the electronic device described in this embodiment includes, but is not limited to, the above and any other suitable types of memory.

[0156] In this embodiment of the invention, the memory 65 stores the following elements of the operating system 651 and the application 652: executable modules, data structures, or subsets thereof, or extended sets thereof.

[0157] Specifically, the operating system 651 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 652 includes various applications, such as a media player and a browser, used to implement various application functions. Programs implementing the methods of this embodiment of the invention can be included in the application program 652. The application program 652 includes applets, objects, components, logic, data structures, and other computer system executable instructions that perform specific tasks or implement specific abstract data types.

[0158] Furthermore, this embodiment of the invention also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the various processes of the above-described advertising video classification method embodiment and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0159] Specifically, when a computer program is executed by a processor, it can perform the following steps:

[0160] The randomly selected advertising video data was manually classified and labeled, and the labeled advertising video data was used as a benchmark test set.

[0161] The keyframes of the extracted local advertising video material are classified using an advertising image classification method to obtain the classification data of the local advertising video.

[0162] Obtain the copywriting classification data of all advertisements, and use the advertisement video data whose copywriting classification data is consistent with the classification data of the local advertisement video as the training sample data for video classification;

[0163] A deep learning model for video classification is trained using the training sample data for video classification.

[0164] Parameters were fine-tuned on the benchmark set to obtain the optimal deep learning model for video classification.

[0165] The video is classified and predicted using the optimal video classification deep learning model to obtain the final classification data for the advertising video.

[0166] Alternatively, when a computer program is executed by a processor, it may also perform the following steps:

[0167] The classification data of the keyframes in the local advertising video material are voted on using a voting method, and the classification with the highest number of votes is used as the voting classification data of the advertising video;

[0168] Based on the optimal video classification deep learning model, the video is classified and predicted to obtain the learning classification data of the advertising video;

[0169] Based on the voting classification data and the learning classification data, a weighted and normalized method is used to obtain the classification scores for each category, and the category with the highest classification score is used as the final classification data for the advertising video.

[0170] Alternatively, when a computer program is executed by a processor, it may also perform the following steps:

[0171] The advertising image classification method includes:

[0172] A classification comparison table is established between the advertising image classification system data and the e-commerce website classification system. Based on the classification data in the classification comparison table, advertising image data of the corresponding categories on the e-commerce website is crawled.

[0173] The set of advertising image data is divided into a training set and a validation set according to the proportions.

[0174] The advertising image data in the training set is used as training sample data to train a deep convolutional neural network classification model.

[0175] The performance of the deep convolutional neural network classification model is evaluated based on the advertising image data in the validation set.

[0176] The deep convolutional neural network classification model with the best performance is selected for classification prediction of the advertising images to obtain the advertising image classification system data.

[0177] Alternatively, when a computer program is executed by a processor, it may also perform the following steps:

[0178] The steps for obtaining the copywriting classification data of all advertisements include:

[0179] Obtain the content information of the advertising copy, and then combine and process this content information to form the first copy information;

[0180] Obtain the content information from the ad landing page, and then combine and process this content information to create the second copy.

[0181] The URL link information of the advertising landing page is cleaned, and the obtained relevant information is used as third-party copy information;

[0182] The first, second, and third copy information are translated into English, and the translated advertising copy is then subjected to binary classification prediction.

[0183] The BERT multi-classification model is used to classify effective advertising copy at various levels to obtain the required advertising classification data.

[0184] Computer-readable storage media include: permanent and non-permanent, removable and non-removable media, which are tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media include: electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, and any suitable combination thereof. Computer-readable storage media include: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, optical disc read-only memory (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape storage, magnetic disk storage or other magnetic storage devices, memory sticks, mechanical encoding devices (e.g., punched cards or raised structures in grooves on which instructions are recorded), or any other non-transfer medium that can be used to store information accessible by a computing device. As defined in the embodiments of the present invention, computer-readable storage media do not include temporary signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.

[0185] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0186] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. When implemented in software, it can be implemented wholly or partially in the form of a computer program product. The computer program product includes one or more computer program instructions. The computer program instructions include: assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk and C++, and procedural programming languages ​​such as C or similar programming languages.

[0187] When the computer program instructions are loaded and executed on a computer, all or part of the process or function described in the embodiments of the present invention is generated. The computer may be a computer, a dedicated computer, a computer network, or other editable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, twisted pair, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave) means. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, magnetic disk, magnetic tape), an optical medium (e.g., optical disc), or a semiconductor medium (e.g., solid state drive (SSD)). Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of the present invention.

[0188] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing embodiments of the present invention, and will not be repeated here.

[0189] In the several embodiments provided in this application, it should be understood that the disclosed apparatus, electronic devices, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, or it may be an electrical, mechanical, or other form of connection.

[0190] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to solve the problems addressed by the embodiments of the present invention, depending on actual needs.

[0191] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0192] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (including: a personal computer, a server, a data center, or other network device) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media listed above that can store program code.

[0193] The above description is merely a specific implementation of the embodiments of the present invention, but the protection scope of the embodiments of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of the present invention should be included within the protection scope of the embodiments of the present invention. Therefore, the protection scope of the embodiments of the present invention should be determined by the protection scope of the claims.

Claims

1. A method for classifying advertising videos, characterized in that, include: The randomly selected advertising video data was manually classified and labeled, and the labeled advertising video data was used as a benchmark test set. The keyframes of the extracted local advertising video material are classified using an advertising image classification method to obtain the classification data of the local advertising video. Obtain the copywriting classification data of all advertisements, and use the advertisement video data whose copywriting classification data is consistent with the classification data of the local advertisement video as the training sample data for video classification; A deep learning model for video classification is trained using the training sample data for video classification. Parameters were fine-tuned on the benchmark set to obtain the optimal deep learning model for video classification. The video is classified and predicted based on the optimal video classification deep learning model to obtain the final classification data of the advertising video; Keyframes extracted from local advertisement video footage include: Reduce the image size to 64 pixels; Convert the reduced image to 64 levels of grayscale; Calculate the average grayscale value of all 64 pixels; The grayscale value of each pixel is compared with the average grayscale value to obtain a comparison result; a value greater than or equal to the average grayscale value is recorded as 1, and a value less than the average grayscale value is recorded as 0. The comparison results are combined to form the fingerprint of the image; Calculate the Hamming distance between the fingerprints of the preceding and following frames. If the Hamming distance is greater than 10, the following frame is determined to be a keyframe. The method further includes: The classification data of the keyframes in the local advertising video material are voted on using a voting method, and the classification with the highest number of votes is used as the voting classification data of the advertising video; Based on the optimal video classification deep learning model, the video is classified and predicted to obtain the learning classification data of the advertising video; Based on the voting classification data and the learning classification data, a weighted and normalized method is used to obtain the classification scores for each category, and the category with the highest classification score is used as the final classification data for the advertising video.

2. The method according to claim 1, characterized in that, The advertising image classification method includes: A classification comparison table is established between the advertising image classification system data and the e-commerce website classification system. Based on the classification data in the classification comparison table, advertising image data of the corresponding categories on the e-commerce website is crawled. The set of advertising image data is divided into a training set and a validation set according to the proportions. The advertising image data in the training set is used as training sample data to train a deep convolutional neural network classification model. The performance of the deep convolutional neural network classification model is evaluated based on the advertising image data in the validation set. The deep convolutional neural network classification model with the best performance is selected for classification prediction of the advertising images to obtain the advertising image classification system data.

3. The method according to claim 1, characterized in that, The steps for obtaining the copywriting classification data of all advertisements include: Obtain the content information of the advertising copy, and then combine and process this content information to form the first copy information; Obtain the content information from the ad landing page, and then combine and process this content information to create the second copy. The URL link information of the advertising landing page is cleaned, and the obtained relevant information is used as third-party copy information; The first, second, and third copy information are translated into English, and the translated advertising copy is then subjected to binary classification prediction. The BERT multi-classification model is used to classify effective advertising copy at various levels to obtain the required advertising classification data.

4. A device for classifying advertising videos, characterized in that, include: The classification and labeling module is used to manually classify and label randomly selected advertising video data, and use the labeled advertising video data as a benchmark test set. The local classification module is used to classify the keyframes of the extracted local advertising video materials using advertising image classification methods, so as to obtain the classification data of the local advertising video. The video sample module is used to acquire the text classification data of all advertisements, and to use the advertisement video data whose text classification data is consistent with the classification data of the local advertisement video as the training sample data for video classification. The model training module is used to train a deep learning model for video classification using the training sample data of the video classification. The parameter tuning module is used to perform parameter tuning on the benchmark test set to obtain the best video classification deep learning model. The final classification module is used to predict the classification of the video based on the best video classification deep learning model to obtain the final classification data of the advertising video. Extracting keyframes from local advertising video footage includes: reducing the image to 64 pixels; converting the reduced image to 64 levels of grayscale; calculating the average grayscale value of all 64 pixels; comparing the grayscale value of each pixel with the average grayscale value to obtain a comparison result; values ​​greater than or equal to the average grayscale value are recorded as 1, and values ​​less than the average grayscale value are recorded as 0; combining the comparison results to form the fingerprint of the image; calculating the Hamming distance between the fingerprints of consecutive frames; if the Hamming distance is greater than 10, the subsequent frame is determined to be a keyframe. The device further includes: The data voting module is used to vote on the classification data of the key frames of the local advertising video material using a voting method, and the classification with the highest number of votes is used as the voting classification data of the advertising video. The learning classification module is used to classify and predict the video based on the best video classification deep learning model to obtain the learning classification data of the advertising video. The final classification module is used to obtain classification scores for each category based on the voting classification data and the learning classification data using a weighted and normalized method, and to use the category with the highest classification score as the final classification data for the advertising video.

5. The apparatus according to claim 4, characterized in that, The local classification module includes: The crawling submodule is used to determine the classification comparison table between the advertising image classification system data and the e-commerce website classification system, and crawl the advertising image data of the corresponding category of the e-commerce website according to the classification data in the classification comparison table; The partitioning submodule is used to divide the set of advertising image data into a training set and a validation set according to a ratio. The training submodule is used to train a deep convolutional neural network classification model by using the advertising image data in the training set as training sample data. The evaluation submodule is used to evaluate the model performance of the deep convolutional neural network classification model based on the advertising image data in the validation set. The classification submodule is used to select the deep convolutional neural network classification model with the best model performance for the classification prediction of the advertisement image, thereby obtaining the advertisement image classification system data.

6. The apparatus according to claim 4, characterized in that, The video sample module includes: The first copywriting submodule is used to obtain the content information of the advertising copy and to splice the content information as the first copywriting information. The second copywriting submodule is used to obtain the content information of the advertising landing page and to concatenate the content information as the second copywriting information. The third copywriting submodule is used to clean the URL link information of the advertising landing page and use the obtained relevant information as the third copywriting information. The translation submodule is used to translate the first copy information, the second copy information, and the third copy information into English, and to perform binary classification prediction on the translated advertising copy. The classification submodule is used to call the BERT multi-classification model to classify effective advertising copy at various levels and obtain the required advertising classification data.

7. An electronic device comprising a bus, a transceiver, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the transceiver, the memory, and the processor are connected via the bus, characterized in that, When the computer program is executed by the processor, it implements the steps in the method for classifying advertising videos as described in any one of claims 1 to 3.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the method for classifying advertising videos as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Convolutional neutral network-based deceptive advertisement page identification method and apparatus

    CN107886344A

  • Video classification method and device, equipment and storage medium

    CN113159010A

  • Advertisement classification method and device and electronic equipment

    CN115186117A