Content detection model training method, content detection method and device
By combining the content feature clustering and user characteristics of multimedia data, the content detection model is trained, and the problem of high cost of multimedia data delivery is solved, and efficient analysis of user preference is achieved.
Patent Information
- Application Number
- CN202210265805.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-03-17
AI Technical Summary
The prior art has high cost of delivery during the multimedia data delivery process, making it difficult to effectively evaluate the user's preference for multimedia data.
By extracting the content features of multimedia data for clustering, obtaining the clustering center, and training the content detection model with user feature vectors and behavior category labels, predicting the user's behavior categories for multimedia data.
Without serving multimedia data, it can predict the user's behavioral categories of multimedia data, analyze the user's preferences, and reduce the delivery cost.
Smart Images

Figure CN114595346B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technology, and in particular to a training method for a content detection model, a content detection method, a device and equipment. Background Art
[0002] After users upload multimedia materials, they can generate and publish large amounts of multimedia data by combining them in various ways. For example, when the multimedia materials are advertising multimedia materials and the multimedia data is video data, the multimedia data specifically refers to the advertising video data. However, not all published multimedia data is popular with users. Therefore, it is necessary to identify user-friendly content from a large amount of multimedia data and analyze it to subsequently generate high-quality multimedia data that is more popular with users.
[0003] Currently, large amounts of multimedia data can be distributed first, and then user behavior information, such as clicks, likes, and completions, can be collected. This behavior information can then be used to assess user preferences for the multimedia data. However, distributing large amounts of multimedia data can result in high costs. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a content detection model training method, a content detection method, an apparatus and a device, which can effectively detect the user's behavior category on multimedia data on the basis of reducing the delivery cost, so as to predict the user's preference for the content in the multimedia data.
[0005] To solve the above problems, the technical solutions provided in the embodiments of the present application are as follows:
[0006] In a first aspect, embodiments of the present application provide a method for training a content detection model, which extracts content features of at least one category of first multimedia data, clusters the content features of each category of the first multimedia data, and obtains multiple cluster centers of the content features of each category. The method includes:
[0007] extracting content features of at least one category of the second multimedia data, comparing the content features of each category of the second multimedia data with respective cluster centers of the content features of the corresponding category, and obtaining the cluster center to which the content features of each category of the second multimedia data belong;
[0008] obtaining a content feature vector of the second multimedia data according to the cluster center to which the content feature of each category of the second multimedia data belongs;
[0009] Get the user feature vector of the user account;
[0010] A content detection model is trained using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data. The content detection model is used to output a prediction result of the behavior category of the target user account for the target multimedia data.
[0011] A second aspect of an embodiment of the present application provides a content detection method, the method comprising:
[0012] extracting content features of at least one category of the target multimedia data, comparing the content features of each category of the target multimedia data with respective cluster centers of content features of the corresponding category, and obtaining the cluster center to which the content features of each category of the target multimedia data belong;
[0013] Obtaining a content feature vector of the target multimedia data according to the cluster center to which the content feature of each category of the target multimedia data belongs;
[0014] Obtain the user feature vector corresponding to the target user account;
[0015] The content feature vector of the target multimedia data and the user feature vector of the target user account are input into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data. The content detection model is trained according to the above-mentioned content detection model training method.
[0016] A third aspect of an embodiment of the present application provides a training device for a content detection model, the device comprising:
[0017] a first extraction unit, configured to extract content features of at least one category of the first multimedia data, and cluster the content features of each category of the first multimedia data to obtain a plurality of cluster centers of the content features of each category;
[0018] a second extraction unit, configured to extract content features of at least one category of the second multimedia data, compare the content features of each category of the second multimedia data with the respective cluster centers of the content features of the corresponding category, and obtain the cluster center to which the content features of each category of the second multimedia data belong;
[0019] a first acquiring unit, configured to obtain a content feature vector of the second multimedia data according to a cluster center to which the content feature of each category of the second multimedia data belongs;
[0020] A second acquisition unit is used to obtain a user feature vector of a user account;
[0021] A training unit is used to train a content detection model using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data. The content detection model is used to output a prediction result of the behavior category of the target user account for the target multimedia data.
[0022] A fourth aspect of an embodiment of the present application provides a content detection device, the device comprising:
[0023] an extraction unit, configured to extract content features of at least one category of the target multimedia data, compare the content features of each category of the target multimedia data with respective cluster centers of content features of the corresponding category, and obtain the cluster center to which the content features of each category of the target multimedia data belong;
[0024] a first acquiring unit, configured to obtain a content feature vector of the target multimedia data according to a cluster center to which the content feature of each category of the target multimedia data belongs;
[0025] A second acquisition unit is used to obtain a user feature vector corresponding to the target user account;
[0026] The first input unit is used to input the content feature vector of the target multimedia data and the user feature vector of the target user account into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data. The content detection model is trained according to the above-mentioned content detection model training method.
[0027] According to a fifth aspect of the present application, an electronic device is provided, including:
[0028] one or more processors;
[0029] a storage device having one or more programs stored thereon,
[0030] When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned content detection model training method or the above-mentioned content detection method.
[0031] A sixth aspect of an embodiment of the present application provides a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, it implements the training method of the content detection model as described above, or the content detection method as described above.
[0032] It can be seen that the embodiments of the present application have the following beneficial effects:
[0033] The embodiment of the present application provides a training method for a content detection model, a content detection method, an apparatus and a device, which first extracts the content features of at least one category of the first multimedia data, clusters the content features of each category of the first multimedia data, and obtains multiple cluster centers of the content features of each category. Then, after extracting the content features of at least one category of the second multimedia data, the content features of each category of the second multimedia data are compared with the respective cluster centers of the content features of the corresponding category to obtain the cluster center to which the content features of each category of the second multimedia data belong. According to the cluster center to which the content features of each category of the second multimedia data belong, the content feature vector of the second multimedia data is obtained. The content detection model is trained using the obtained content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data. The content detection model after training is able to output the prediction result of the behavior category of the target user account for the target multimedia data. In this way, the content detection model can be used to predict the user's behavior category for the multimedia data without delivering the multimedia data, and then the user's liking for the multimedia data can be analyzed. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A schematic diagram of a framework of an exemplary application scenario provided in an embodiment of the present application;
[0035] Figure 2 A flowchart of a method for training a content detection model provided in an embodiment of the present application;
[0036] Figure 3a A schematic diagram of a first multimedia data clustering provided in an embodiment of the present application;
[0037] Figure 3b A schematic diagram of a second multimedia data clustering provided in an embodiment of the present application;
[0038] Figure 4a A schematic diagram of a content detection model provided in an embodiment of the present application;
[0039] Figure 4b A schematic diagram of another content detection model provided in an embodiment of the present application;
[0040] Figure 5a A schematic diagram of another content detection model provided in an embodiment of the present application;
[0041] Figure 5b A schematic diagram of another content detection model provided in an embodiment of the present application;
[0042] Figure 6A schematic diagram of another exemplary application scenario provided in an embodiment of the present application;
[0043] Figure 7 A flowchart of a content detection method provided in an embodiment of the present application;
[0044] Figure 8 A schematic diagram of a training method for a user account recall model provided in an embodiment of the present application;
[0045] Figure 9 A schematic diagram of the structure of a content detection model training device provided in an embodiment of the present application;
[0046] Figure 10 A schematic structural diagram of a content detection device provided in an embodiment of the present application;
[0047] Figure 11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0049] In order to facilitate understanding and explanation of the technical solutions provided by the embodiments of the present application, the background technology of the present application will be described below.
[0050] After users upload multimedia materials, they can automatically generate and publish large amounts of multimedia data by combining these materials in various ways. However, not all published multimedia data is popular with users. Therefore, it is necessary to identify and analyze user-friendly content from this large amount of multimedia data so that high-quality multimedia data can be generated that is more appealing to users.
[0051] As an optional example, when the multimedia material is specifically an advertising multimedia material and the multimedia data is video data, the advertising multimedia material specifically refers to the advertising video material, and the multimedia data specifically refers to the advertising video data (hereinafter referred to as advertising video). Specifically, users will evaluate the advertising video through platform-defined incentive behaviors such as clicks, likes, and completion. When the click-through rate, likes, or completion rate of the advertising video are high, it can be determined that the advertising video is a high-quality video and the advertising video material in the advertising video is high-quality advertising video material. Otherwise, it is a low-quality advertising video and advertising video material. After determining the high-quality advertising video material, a higher-quality advertising video can be generated later. At present, a large amount of multimedia data can be first distributed and delivered, and then the user's behavior information on the delivered multimedia data can be obtained, such as clicks, likes, completion, etc. Then, the user's preference for the multimedia data is evaluated based on the user's behavior information. However, the delivery of a large amount of multimedia data will result in high delivery costs.
[0052] Based on this, the embodiment of the present application provides a training method for a content detection model, a content detection method, an apparatus and a device, which first extracts the content features of at least one category of the first multimedia data, clusters the content features of each category of the first multimedia data, and obtains multiple cluster centers of the content features of each category. Then, after extracting the content features of at least one category of the second multimedia data, the content features of each category of the second multimedia data are compared with the respective cluster centers of the content features of the corresponding category to obtain the cluster centers to which the content features of each category of the second multimedia data belong. According to the cluster centers to which the content features of each category of the second multimedia data belong, the content feature vector of the second multimedia data is obtained. The content detection model is trained using the obtained content feature vector of the second multimedia data, the user feature vector of the user account and the behavior category label of the user account for the second multimedia data. The content detection model after training is able to output the prediction result of the behavior category of the target user account for the target multimedia data. In this way, the content detection model can be used to predict the user's behavior category for the multimedia data without delivering the multimedia data, and then the user's liking for the multimedia data can be analyzed.
[0053] It should be noted that in the embodiments of the present application, the user feature vector of the user account and the behavior category label of the user account for the second multimedia data do not involve sensitive information of the user. The user feature vector of the user account and the behavior category label of the user account for the second multimedia data are obtained and used after user authorization. In one example, before obtaining the user feature vector of the user account and the behavior category label of the user account for the second multimedia data, a corresponding interface displays a prompt message related to obtaining data use authorization, and the user determines whether to agree to the authorization based on the prompt message.
[0054] In order to facilitate understanding of the training method of the content detection model provided in the embodiment of the present application, the following Figure 1 See the example scenario shown. Figure 1 As shown in the figure, this figure is a framework diagram of an exemplary application scenario provided in an embodiment of the present application.
[0055] In practical applications, at least one category of content features of the first multimedia data is first obtained. For example, the first multimedia data includes data of the title text category, data of the OCR (Optical Character Recognition) text category, data of the ASR (Automatic Speech Recognition) text category, or data of the video / image category. The content feature is a feature vector obtained based on the data. Different categories of data correspond to different categories of content features, that is, different categories of feature vectors. The first multimedia data is the collected multimedia data and can be used to determine the cluster centers of each category of content features. Then, after obtaining at least one category of content features of the first multimedia data, the content features of each category in the first multimedia data are clustered separately to obtain multiple cluster centers of the content features of each category. For example, there are five cluster centers corresponding to the content features of the title text category, namely cluster centers 01, 02, 03, 04, and 05.
[0056] After obtaining multiple cluster centers of content features of each category of each multimedia data, the content feature vector of the second multimedia data can be obtained based on the multiple cluster centers of content features of each category. In specific implementation, the content features of at least one category of the second multimedia data are first extracted, and then the content features of each category in the second multimedia data are compared with the respective cluster centers of the content features of the corresponding category that have been obtained to determine the cluster center to which the content features of each category of the second multimedia data belong. For example, the content features of the title text category data in the second multimedia data are compared with the 5 cluster centers that have been obtained to determine the cluster center to which the content features of the title text category data in the second multimedia data belong, such as cluster center A. Furthermore, based on the cluster center to which the content features of each category of the second multimedia data belong, the content feature vector of the second multimedia data is obtained. The content feature vector of the second multimedia data is used to train the content detection model.
[0057] In addition, a user feature vector of the user account is obtained and used to train the content detection model. Specifically, the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account with respect to the second multimedia data are used to train the content detection model. The content detection model, whether in the process of training or after training, is used to output a prediction result of the behavior category of the target user account with respect to the target multimedia data.
[0058] Those skilled in the art will understand that Figure 1 The framework diagram shown is only an example in which the embodiments of the present application can be implemented. The scope of application of the embodiments of the present application is not limited by any aspect of the framework.
[0059] To facilitate understanding of the present application, a training method for a content detection model provided in an embodiment of the present application is described below with reference to the accompanying drawings.
[0060] See also Figure 2 As shown in FIG, this figure is a flow chart of a method for training a content detection model provided by an embodiment of the present application. Figure 2 As shown, the method may include S201-S204:
[0061] S201: extracting content features of at least one category of second multimedia data, comparing the content features of each category of the second multimedia data with the cluster centers of the content features of the corresponding category, and obtaining the cluster center to which the content features of each category of the second multimedia data belong.
[0062] Before executing this step, it is necessary to first determine the cluster centers corresponding to the content features of at least one category of multimedia data. As an optional example, the multimedia data is advertising multimedia data. Figure 3a , Figure 3a This is a schematic diagram of a first multimedia data clustering provided in an embodiment of the present application. Figure 3a As shown, first multimedia data is collected. The first multimedia data is used to determine the cluster center of the multimedia data. For example, the first multimedia data is 50 million multimedia data. As an optional example, the first multimedia data is first advertisement multimedia data.
[0063] Then, the content features of at least one category of the first multimedia data are extracted. The categories of the first multimedia data include one or more of the following: title text category, OCR text category, ASR text category, and video / image category. In one or more embodiments, a pre-trained model can be directly used to extract the content features of at least one category of the first multimedia data, and then the extracted content features are transferred to the content detection model. For example, Figure 3aAs shown, the pre-trained model is a bidirectional pre-trained converter BERT model, and the corresponding extracted content features are Bert features. The BERT model can be used to extract title text category content features, OCR text category content features, and ASR text category content features. In addition, the pre-trained model can also be a picture-level deep learning model. The picture-level deep learning model can be used to extract video / image category content features. For example, based on the image dataset ImageNet, the corresponding extracted content features are ImageNet model features. Figure 3a As shown, the content features extracted from the data of the title text category are the title text Bert features, the content features extracted from the data of the OCR text category are the OCR text Bert features, the content features extracted from the data of the ASR text category are the ASR text Bert features, and the content features extracted from the data of the video / image category are the ImageNet model features.
[0064] Finally, the content features of each category of the first multimedia data are clustered separately to obtain multiple cluster centers of the content features of each category. In one or more embodiments, the cluster centers can be represented by ID serial numbers or other representation forms. For example, the multiple cluster centers of the content features of the title text category are 01, 02, 03, 04, and 05. The multiple cluster centers corresponding to the content features of the OCR text category are 06, 07, and 08. The multiple cluster centers corresponding to the content features of the ASR text category are 09, 10, 11, and 12. The multiple cluster centers corresponding to the content features of the video / image category are 13, 14, and 15.
[0065] After determining multiple cluster centers for the content features of each category of multimedia data, a content feature vector for the second multimedia data can be determined based on these cluster centers. The content feature vector for the second multimedia data is used to train the content detection model. It is understood that the second multimedia data is already delivered multimedia data, for example, 50,000 pieces of already delivered multimedia data. As an optional example, the second multimedia data is second advertising multimedia data.
[0066] In specific implementation, it is necessary to first extract at least one category of content features of the second multimedia data. Figure 3b , Figure 3bA schematic diagram of a clustering of second multimedia data provided in an embodiment of the present application. The categories of the second multimedia data also include one or more of a title text category, an OCR text category, an ASR text category, and a video / image category. In one or more embodiments, since the second multimedia data is used to train a content detection model, in order to improve the time of model training during the training of the content detection model, a pre-trained model can be directly used to extract content features of at least one category of the first multimedia data. For example, Figure 3b As shown, the pre-trained model is a bidirectional pre-trained transformer BERT model or a model based on the ImageNet image dataset, and the corresponding extracted content features are BERT features and ImageNet model features. As an optional example, based on the multimedia data being advertising multimedia data, the content detection model is an advertising content detection model.
[0067] Then, the content features of each category of the second multimedia data are compared with the respective cluster centers of the content features of the corresponding category to obtain the cluster center to which the content features of each category of the second multimedia data belong. For example, the content features of the obtained title text category data in the second multimedia data are compared with the multiple cluster centers of the title text category content features, and the obtained cluster center to which the content features of the title text category data belong is A. Finally, the content feature vector of the second multimedia data is obtained through the subsequent S202.
[0068] In one or more embodiments, the dimensions of the content features extracted using a pre-trained model are usually very high. In this case, the extracted content features may be first reduced in dimension, and then the reduced content features may be used for subsequent processing.
[0069] S202: Obtain a content feature vector of the second multimedia data according to the cluster center to which the content feature of each category of the second multimedia data belongs.
[0070] After obtaining the cluster centers to which the content features of each category of the second multimedia data belong, a content feature vector of the second multimedia data can be obtained based on the cluster centers to which the content features of each category of the second multimedia data belong. It will be appreciated that the content feature vector of the second multimedia data can be obtained based on the cluster centers to which the content features of each category of the second multimedia data belong through various achievable implementations.
[0071] In a possible implementation, an embodiment of the present application provides a specific implementation method for obtaining a content feature vector of the second multimedia data according to a cluster center to which the content feature of each category of the second multimedia data belongs.
[0072] First, based on the content features of each category, content feature vectors corresponding to multiple cluster centers of the content features of each category are calculated. Next, the content feature vector corresponding to the cluster center of the content features of each category of the second multimedia data is determined as the content feature vector of the second multimedia data. This method directly obtains the content feature vector of the second multimedia data without increasing model training time, thereby improving the training efficiency of the content detection model.
[0073] In a possible implementation, the embodiment of the present application provides a specific implementation method for obtaining the content feature vector of the second multimedia data according to the cluster center to which the content feature of each category of the second multimedia data belongs in S202. For details, please refer to B1-B5 below.
[0074] S203: Obtain a user feature vector of the user account.
[0075] In one or more embodiments, the obtained user feature vector of the user account is also used to train the content detection model.
[0076] In one possible implementation, the present embodiment provides a specific implementation method for obtaining a user feature vector of a user account, including:
[0077] A1: Collect user information of the user account, and generate a first user feature of the user account according to the user information of the user account.
[0078] Among them, the user information of the user account is used to represent the user's relevant information, including the user's identity information, the user's gender information, the user's age information, the user's province identification code (ie, province ID), the device identification code of the user account, i.e., the device ID, etc.
[0079] According to the user information of the user account, a first user feature of the user account may be generated. The first user feature of the user account is used to characterize the user account.
[0080] A2: Obtain the second user feature of the pre-trained user account.
[0081] In order to more accurately represent the user account information, in one or more embodiments, a second user feature of the user account is obtained. The second user feature is also used to represent the user account, which can make the representation of the user account more accurate.
[0082] As an optional example, the second user feature of the user account may be pre-trained, for example, the feature of the user account may be obtained from other services and used as the second user feature of the user account.
[0083] A3: Use the first user feature of the user account and the second user feature of the user account as a user feature vector of the user account.
[0084] In one or more embodiments, the user feature vector of the user account is composed of a feature vector corresponding to the first user feature of the user account and a feature vector corresponding to the second user feature of the user account. Thus, the obtained user feature vector of the user account can more accurately represent the user account.
[0085] It should be noted that in the embodiments of the present application, the user information of the user account, the first user feature of the user account, and the second user feature of the user account do not involve sensitive information of the user. The user information of the user account, the first user feature of the user account, and the second user feature of the user account are obtained and used after authorization by the user. In one example, before obtaining the user information of the user account, the first user feature of the user account, and the second user feature of the user account, the corresponding interface displays a prompt message related to obtaining data use authorization, and the user determines whether to agree to the authorization based on the prompt message.
[0086] S204: Using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data, a content detection model is trained. The content detection model is used to output a prediction result of the behavior category of the target user account for the target multimedia data.
[0087] After obtaining the content feature vector of the second multimedia data and the user feature vector of the user account, it is also necessary to obtain the behavior category label of the user account for the second multimedia data. It is understandable that the behavior category label of the user account for the second multimedia data can characterize the degree of liking of the user account for the second multimedia data. Among them, the behavior categories of the user account for the second multimedia data include clicks, likes or completions, etc. Taking likes as an example, the behavior category labels of the user account for the second multimedia data are likes and dislikes. If the label is likes, it means that the user account likes the multimedia data that is liked. Taking completion as an example, the behavior category label of the user account for the second multimedia data can be determined as a specific duration according to actual needs. For example, a label of less than or equal to 45 seconds and a label of more than 45 seconds.
[0088] Based on this, the content detection model is trained using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavioral category label of the user account for the second multimedia data. The trained content detection model is used to output a prediction result of the behavioral category of the target user account for the target multimedia data. As an optional example, based on the fact that the multimedia data is advertising multimedia data and the content detection model is an advertising content detection model, the target multimedia data is target advertising multimedia data. It can be understood that determining whether multimedia data is high-quality is not only related to the multimedia data itself, but also to the preferences of the user account. The multimedia data preferred by different user accounts may be different. Therefore, in the process of training the content detection model in the embodiment of the present application, not only the content feature vector of the second multimedia data is used, but also the user feature vector of the user account and the behavioral category label of the user account for the second multimedia data are used. That is, in the process of training the content detection model in the embodiment of the present application, both the factors of the multimedia data itself and the factors of the user are taken into account. The trained content detection model can reflect the preferences of different user accounts for multimedia data, making the content detection model more reasonable and accurate.
[0089] Because user account preferences may change over time, in one or more embodiments, the content detection model needs to be continuously retrained. That is, after a certain period of time, after recollecting the second multimedia data, the content detection model is retrained to improve the accuracy of the content detection model in predicting the current user account preferences.
[0090] In one or more embodiments, the content detection model can be trained with multiple labels, i.e., the user account's behavior category labels for the second multimedia data are multiple, such as click, like, and complete broadcast labels. In other embodiments, a click-based evaluation model can be trained based on the click behavior category label, a like-based evaluation model can be trained based on the like behavior category label, and a complete broadcast-based evaluation model can be trained based on the complete broadcast behavior category label. Finally, the content detection model is composed of a click-based evaluation model, a like-based evaluation model, and a complete broadcast-based evaluation model.
[0091] In some possible implementations, an embodiment of the present application provides a specific implementation method for training a content detection model using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data. Please see C1-C3 and D1-D4 below for details.
[0092] Based on the contents of S201-S204, an embodiment of the present application provides a method for training a content detection model, which first extracts content features of at least one category of the first multimedia data, clusters the content features of each category of the first multimedia data, and obtains multiple cluster centers of the content features of each category. Then, after extracting the content features of at least one category of the second multimedia data, the content features of each category of the second multimedia data are compared with the respective cluster centers of the content features of the corresponding category to obtain the cluster centers to which the content features of each category of the second multimedia data belong. Based on the cluster centers to which the content features of each category of the second multimedia data belong, a content feature vector of the second multimedia data is obtained. The content detection model is trained using the obtained content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data. The content detection model after training is able to output a prediction result of the behavior category of the target user account for the target multimedia data. In this way, the content detection model can be used to predict the user's behavior category for the multimedia data without delivering the multimedia data, and then the user's liking for the multimedia data can be analyzed.
[0093] It is understood that the above-mentioned S202 provides an implementation method for directly determining the content feature vector corresponding to the cluster center of the content feature of each category of the second multimedia data as the content feature vector of the second multimedia data. However, since the content feature vector obtained by this method is directly extracted from a pre-trained model, it is generally subject to overfitting, resulting in the content feature vector of the second multimedia data not accurately representing the second multimedia data.
[0094] Based on this, in one possible implementation, the embodiment of the present application provides another specific implementation method of obtaining the content feature vector of the second multimedia data according to the cluster center to which the content feature of each category of the second multimedia data belongs in S202, including:
[0095] B1: Obtain an initial content feature vector corresponding to the cluster center to which the content feature of each category of the second multimedia data belongs.
[0096] After determining the cluster center to which the content features of each category of the second multimedia data belong, set the initial content feature vector corresponding to the cluster center to which the content features of each category of the second multimedia data belong. The initial content feature vector is the initial value of the content feature vector corresponding to the cluster center and can be determined randomly. For example, the cluster center to which the content features of the title text category of the second multimedia data belong is 01, and the set initial content feature vector is represented by a1. The cluster center to which the content features of the OCR text category belong is 06, and the set initial content feature vector is represented by b1. The cluster center to which the content features of the ASR text category belong is 09, and the set initial content feature vector is represented by c1. The cluster center to which the content features of the video / image category belong is 13, and the set initial content feature vector is represented by d1.
[0097] B2: Determine the initial content feature vector corresponding to the cluster center to which the content feature of each category of the second multimedia data belongs as the content feature vector of the second multimedia data.
[0098] Furthermore, before training the content detection model, the initial content feature vector corresponding to the cluster center of the content features of each category of the second multimedia data is determined as the content feature vector of the second multimedia data for use in training the content detection model. It can be considered that the content feature vector of the second multimedia data is derived based on the content feature vector corresponding to the cluster center of the content features of each category of the second multimedia data. Furthermore, the content feature vector of the second multimedia data is adjusted as the content detection model is trained. For details, see B3-B4.
[0099] In one or more embodiments, the initial content feature vectors corresponding to the cluster centers of the content features of each category of the second multimedia data are concatenated, and the concatenated feature vector is the content feature vector of the second multimedia data. For example, the content feature vector of the second multimedia data is (a1, b1, c1, d1).
[0100] Based on the content of B1-B2, the training method of the content detection model provided in the embodiment of the present application further includes:
[0101] B3: During the process of training the content detection model, adjust the content feature vector of the second multimedia data.
[0102] After B2, in the process of training the content detection model, the content feature vector of the second multimedia data will be adjusted as the content detection model is iteratively trained. Since the content feature vector of the second multimedia data is obtained based on the content feature vector corresponding to the cluster center to which the content features of each category of the second multimedia data belong. The content feature vector corresponding to the cluster center to which the content features of each category currently belong can be re-determined based on the adjusted content feature vector of the second multimedia data. That is, the content feature vector corresponding to the cluster center to which the content features of each category belong is also adjusted accordingly. For example, the adjusted content feature vector of the second multimedia data is (a2, b2, c2, d2). The adjusted content feature vectors corresponding to the cluster centers to which the content features of each category belong (i.e., 01, 06, 09, and 13) are a2, b2, c2, and d2, respectively.
[0103] B4: Re-determine the content feature vector corresponding to the cluster center to which the content feature of each category belongs after adjustment as the initial content feature vector corresponding to the cluster center to which the content feature of the category belongs.
[0104] Before the next iteration of training the content detection model, the content feature vector corresponding to the cluster center of each category's content features after adjustment is re-determined as the initial content feature vector corresponding to the cluster center of the content features of the corresponding category. For example, the re-determined initial content feature vectors corresponding to the cluster center of each category's content features are a2, b2, c2, and d2, respectively. Furthermore, the adjusted content feature vector of the second multimedia data is (a2, b2, c2, d2), which continues to be used for training the content detection model this time.
[0105] B5: After the content detection model is trained, content feature vectors corresponding to multiple cluster centers of content features of each category are obtained.
[0106] After the content detection model is trained, the content feature vector of the second multimedia data is also adjusted. For example, the content feature vector of the second multimedia data after adjustment is (aa, bb, cc, dd). Based on the content feature vector of the second multimedia data finally obtained, the content feature vectors corresponding to the cluster centers of the content features of each category are obtained. For example, the cluster centers of the content features of each category are 01, 06, 09, and 13, and the corresponding content feature vectors are aa, bb, cc, and dd, respectively. It can be understood that the content feature vector corresponding to the cluster center of the content features of each category is the content feature vector after adjustment.
[0107] In one or more embodiments, the content detection model in this application is trained in real time. That is, after the second multimedia data is collected again, the content detection model is retrained. Therefore, before retraining the content detection model, the content feature vector of the second multimedia data is obtained. In this case, the content feature vector of the second multimedia data is still obtained from the initial content feature vector corresponding to the cluster center to which the content features of each category of the second multimedia data belong.
[0108] It is understandable that the cluster centers to which the content features of some categories in the re-collected second multimedia data belong may change. If the cluster centers to which the content features of some categories belong are used for the first time, the initial content feature vectors corresponding to these cluster centers are randomly initialized feature vectors. For example, the cluster center to which the content features of the OCR text category in the re-collected second multimedia data belong becomes 07, and its corresponding initial content feature vector is a randomly initialized feature vector, such as e1. The cluster center to which the content features of the video / image category belong becomes 14, and its corresponding initial content feature vector is also a randomly initialized feature vector, such as f1. If it is not used for the first time, the initial content feature vectors corresponding to these cluster centers are the content feature vectors corresponding to the cluster centers obtained after the last adjustment. For example, the cluster center to which the content features of the title text category of the second multimedia data belong is still 01, and its corresponding initial content feature vector is aa. The cluster center to which the content features of the ASR text category belong is still 09, and its corresponding initial content feature vector is cc.
[0109] Thus, after training the content detection model with multiple batches of large amounts of second multimedia data, since the cluster centers to which the content feature vectors of each category of each batch of second multimedia data belong may change, it is possible to finally obtain content feature vectors corresponding to multiple cluster centers of the content features of each category. For example, content feature vectors corresponding to multiple cluster centers (such as 01, 02, 03, 04, and 05) of the content features of the title text category are obtained. Content feature vectors corresponding to multiple cluster centers (such as 06, 07, and 08) corresponding to the content features of the OCR text category are obtained. Content feature vectors corresponding to multiple cluster centers (such as 09, 10, 11, and 12) corresponding to the content features of the ASR text category are obtained. Content feature vectors corresponding to multiple cluster centers (such as 13, 14, and 15) corresponding to the content features of the video / image category are obtained.
[0110] Based on the contents of B1-B5, the initial content feature vector corresponding to the cluster center to which the content features of each category of the second multimedia data belong is determined as the content feature vector of the second multimedia data. During the training process of the content detection model, the content feature vector of the second multimedia data is adjusted, that is, the initial content feature vector corresponding to the cluster center to which the content features of each category of the second multimedia data belong is adjusted. This makes the content feature vector of the second multimedia data more accurate in representing the second multimedia data and can more accurately characterize the second multimedia data. Furthermore, this can also make the trained content detection model more accurate.
[0111] See also Figure 4a , Figure 4a This is a schematic diagram of a content detection model provided in an embodiment of the present application. Figure 4a As shown, in one or more embodiments, the content detection model includes a first cross-feature extraction module and a connection module. Based on this, in one possible implementation, the present application embodiment provides a specific implementation method of training the content detection model in S204 using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data, including:
[0112] C1: Inputting the content feature vector of the second multimedia data and the user feature vector of the user account into the first cross-feature extraction module, so that the first cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a first feature vector.
[0113] The first cross-feature extraction module is used to perform cross-feature extraction on the input feature vector. Compared with the content feature vector of the independent second multimedia data and the user feature vector of the user account, the first feature vector can contain more feature vector information. Therefore, using the first feature vector to train the content detection model can achieve better training results. In addition, by using the first feature vector to train the content detection model, the model will learn the degree of influence of the user feature vectors of different user accounts on the content feature vectors of different second multimedia data, and explore the differences in preferences of different user accounts for different second multimedia data.
[0114] In order to facilitate the processing of feature vectors, the dimension of the feature vectors will be changed. Figure 4b As shown, Figure 4bA schematic diagram of another content detection model provided for an embodiment of the present application. In one or more embodiments, the content detection model further includes a plurality of fully connected layers, wherein the content feature vector of the second multimedia data is first input into the fully connected layer, the dimension of the content feature vector of the second multimedia data is changed, and then the content feature vector of the second multimedia data with the changed dimension is input into the first cross-feature extraction module. Similarly, the user feature vector of the user account is first input into the fully connected layer, the dimension of the user feature vector of the user account is changed, and then the user feature vector of the user account with the changed dimension is input into the first cross-feature extraction module, so that the first cross-feature extraction module performs cross-feature extraction on the input feature vector to obtain a first feature vector.
[0115] It is understandable that the embodiment of the present application does not limit the number and composition structure of the fully connected layers, which can be set according to actual conditions.
[0116] C2: Inputting the content feature vector of the second multimedia data and the user feature vector of the user account into the connection module, so that the connection module connects the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a second feature vector.
[0117] It can be understood that the connection module is used for cascade splicing, and splices the content feature vector of the second multimedia data and the user feature vector of the user account to obtain the second feature vector.
[0118] In one or more embodiments, the content detection model further includes a fully connected layer, and the obtained second feature vector is input into the fully connected layer to re-obtain the second feature vector.
[0119] C3: Train a content detection model using the first feature vector, the second feature vector, and the behavior category label of the user account for the second multimedia data.
[0120] After obtaining the first feature vector and the second feature vector, the content detection model is trained using the first feature vector, the second feature vector, and the behavior category label of the user account for the second multimedia data.
[0121] Based on the content of C1-C3, the content feature vector of the second multimedia data, and the user feature vector of the user account, a first feature vector is obtained using a first cross-feature extraction module, and a second feature vector is obtained using a connection module. Furthermore, a content detection model is trained using the first and second feature vectors, as well as the behavioral category label of the user account with respect to the second multimedia data. The user feature vector of the user account is utilized during model training, enabling the trained content detection model to output highly accurate prediction results for the behavioral category of the target user account with respect to the target multimedia data.
[0122] See also Figure 5a , Figure 5a This is a schematic diagram of another content detection model provided in an embodiment of the present application. Figure 5a As shown, in one or more embodiments, the content detection model includes a second cross-feature extraction module, a third cross-feature extraction module, and a connection module. Based on this, an embodiment of the present application provides a specific implementation method of training the content detection model in S204 using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data, including:
[0123] D1: Inputting the content feature vector of the second multimedia data and the first user feature into the second cross-feature extraction module, so that the second cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the first user feature to obtain a third feature vector.
[0124] Since the user feature vector of the user account can be composed of the first user feature of the user account and the second user feature of the user account, on the basis that the content detection model includes a second cross-feature extraction module, a third cross-feature extraction module and a connection module, the content feature vector of the second multimedia data and the first user feature can be input into the second cross-feature extraction module to obtain the third feature vector.
[0125] The third feature vector contains information about the content feature vector of the second multimedia data and information about the first user features. It is a combination of the content feature vector of the second multimedia data and the first user features. Compared to the independent content feature vector of the second multimedia data and the first user features, the third feature vector can contain more feature vector information. This allows for better training results when training a content detection model using the third feature vector. In addition, when training a content detection model using the third feature vector, the model will learn the degree of influence of users with different first user features on the content feature vectors of different second multimedia data, and uncover differences in preferences between different user accounts for different second multimedia data.
[0126] like Figure 5b As shown, Figure 5b A schematic diagram of another content detection model provided in an embodiment of the present application. In one or more embodiments, such as Figure 5bAs shown, the content detection model further includes multiple fully connected layers. The content feature vector of the second multimedia data is first input into the fully connected layer, the dimension of the content feature vector of the second multimedia data is changed, and the content feature vector of the second multimedia data with the changed dimension is then input into the second cross-feature extraction module. Similarly, the first user feature is first input into the fully connected layer, the dimension of the first user feature is changed, and the changed dimension of the first user feature is then input into the second cross-feature extraction module, so that the second cross-feature extraction module performs cross-feature extraction on the input feature vector to obtain a third feature vector.
[0127] D2: Inputting the content feature vector of the second multimedia data and the second user feature into the third cross-feature extraction module, so that the third cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the second user feature to obtain a fourth feature vector.
[0128] It is understandable that compared to the independent content feature vectors and second user features of the second multimedia data, the fourth feature vector can contain more feature vector information. This allows for better training results when using the fourth feature vector to train the content detection model. Furthermore, using the fourth feature vector to train the content detection model will allow the model to learn the degree to which users with different second user features influence the content feature vectors of different second multimedia data, thereby uncovering differences in preferences between different user accounts for different second multimedia data.
[0129] In one or more embodiments, the content detection model further includes a plurality of fully connected layers, wherein the content feature vector of the second multimedia data is first input into the fully connected layer, the dimension of the content feature vector of the second multimedia data is changed, and then the content feature vector of the second multimedia data with the changed dimension is input into the third cross-feature extraction module. Similarly, the second user feature is first input into the fully connected layer, the dimension of the second user feature is changed, and then the second user feature with the changed dimension is input into the third cross-feature extraction module, so that the third cross-feature extraction module performs cross-feature extraction on the input feature vector to obtain a fourth feature vector.
[0130] D3: Inputting the content feature vector of the second multimedia data, the first user feature and the second user feature into the connection module, so that the connection module connects the content feature vector of the second multimedia data, the first user feature and the second user feature to obtain a fifth feature vector.
[0131] In one or more embodiments, the content detection model further includes a fully connected layer, and the obtained fifth feature vector is input into the fully connected layer to re-obtain the fifth feature vector.
[0132] D4: Train a content detection model using the third eigenvector, the fourth eigenvector, the fifth eigenvector, and the behavior category label of the user account for the second multimedia data.
[0133] After obtaining the third eigenvector, the fourth eigenvector, and the fifth eigenvector, the content detection model is trained using the third eigenvector, the fourth eigenvector, the fifth eigenvector, and the behavior category label of the user account for the second multimedia data.
[0134] Based on the content of D1-D4, based on the content feature vector of the second multimedia data and the first user feature, the third feature vector is obtained using the second cross-feature extraction module. Based on the content feature vector of the second multimedia data and the second user feature, the third feature vector is obtained using the third cross-feature extraction module, and the fourth feature vector is obtained using the connection module. Based on the content feature vector of the second multimedia data, the first user feature and the second user feature, the fifth feature vector is obtained using the connection module. Furthermore, the content detection model is trained using the third feature vector, the fourth feature vector, the fifth feature vector and the behavior category label of the user account for the second multimedia data. The user feature vector of the user account is used in the process of training the model, so that the trained content detection model can output a highly accurate prediction result of the behavior category of the target user account for the target multimedia data.
[0135] After the content detection model is trained, the content detection model can be used to detect the content to obtain the prediction results of the target user account's behavior category for the target multimedia data. Figure 6 See the example scenario shown. Figure 6 As shown in the figure, this figure is a framework diagram of an exemplary application scenario provided in an embodiment of the present application.
[0136] like Figure 6 As shown, in practical applications, target multimedia data is first acquired, which is the multimedia data to be detected. Then, content features of at least one category of the target multimedia data are extracted. The content features of each category of the target multimedia data are compared with the cluster centers of the content features of the corresponding category to obtain the cluster centers to which the content features of each category of the target multimedia data belong. Furthermore, a content feature vector of the target multimedia data can be obtained based on the cluster centers to which the content features of each category of the target multimedia data belong. This content feature vector of the target multimedia data is input into the trained content detection model.
[0137] In addition, the user feature vector corresponding to the target user account must be obtained. The user feature vector corresponding to the target user account is used to input the trained content detection model.
[0138] Finally, the content feature vector of the target multimedia data and the user feature vector of the target user account are input into the content detection model to obtain a prediction result of the target user account's behavior category for the target multimedia data. The content detection model is trained according to the content detection model training method of any of the above embodiments.
[0139] It should be noted that in the embodiments of this application, the user feature vector of the target user account does not involve sensitive user information. The user feature vector of the target user account is obtained and used after user authorization. In one example, before obtaining the user feature vector of the target user account, a prompt related to obtaining data use authorization is displayed on the corresponding interface, and the user determines whether to agree to the authorization based on the prompt.
[0140] Those skilled in the art will understand that Figure 6 The framework diagram shown is only an example in which the embodiments of the present application can be implemented. The scope of application of the embodiments of the present application is not limited by any aspect of the framework.
[0141] To facilitate understanding of the content detection method provided in an embodiment of the present application, a content detection method provided in an embodiment of the present application is described below with reference to the accompanying drawings.
[0142] See also Figure 7 As shown in FIG, this figure is a flow chart of a content detection method provided by an embodiment of the present application. Figure 7 As shown, the method may include S701-S704:
[0143] S701: extracting content features of at least one category of target multimedia data, comparing the content features of each category of the target multimedia data with the cluster centers of the content features of the corresponding category, and obtaining the cluster center to which the content features of each category of the target multimedia data belong.
[0144] It can be understood that the target multimedia data is the multimedia data to be detected, and the target multimedia data can be one piece of multimedia data or multiple pieces of multimedia data.
[0145] In one or more embodiments, a pre-trained model may be used to extract content features of at least one category of target multimedia data.
[0146] After obtaining the content features of at least one category of the target multimedia data, the content features of each category of the target multimedia data are compared with the cluster centers of the content features of the corresponding category to obtain the cluster centers to which the content features of each category of the target multimedia data belong. The cluster centers of the content features of the corresponding category are obtained based on the first multimedia data in S201.
[0147] S702: Obtain a content feature vector of the target multimedia data according to the cluster center to which the content feature of each category of the target multimedia data belongs.
[0148] The obtained content feature vector of the target multimedia data is used to input into the trained content detection model for prediction.
[0149] It can be understood that S702 in the embodiment of the present application is similar to S202 in the above embodiment. For the sake of brevity, it will not be described in detail here. For detailed information, please refer to the description in the above embodiment.
[0150] S703: Obtain the user feature vector corresponding to the target user account.
[0151] The obtained user feature vector corresponding to the target user account is used to input into the trained content detection model for prediction. The content detection model detects that the user account of the target multimedia data is the target user account, that is, the user account that has clicked, liked, or completed the target multimedia data. This determines the target user account's preference for the target multimedia data, and based on the detection results, determines whether the target multimedia data is high-quality.
[0152] As an optional example, the target user account is a random user account.
[0153] Each piece of multimedia data has a biased audience, meaning only a subset of user accounts are interested. The target user accounts are those that are likely to be interested in the target multimedia data. Using user accounts that are likely to be interested in the target multimedia data to evaluate the target multimedia data makes the evaluation results more reasonable and accurate. As another alternative, target user accounts can be determined by training a user account recall model.
[0154] In a specific implementation, before obtaining the user feature vector corresponding to the target user account, the content detection method provided in the embodiment of the present application further includes:
[0155] The content feature vector of the target multimedia data is input into the user account recall model to obtain the target user account corresponding to the target multimedia data.
[0156] Among them, participants Figure 8 , Figure 8A training diagram of a user account recall model provided in an embodiment of the present application. The user account recall model is trained based on the content feature vector of the third multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the third multimedia data. During the training process of the user account recall model, the behavior category of the user account for the third multimedia data can be predicted by calculating the similarity between the content feature vector of the third multimedia data and the user feature vector of the user account. The predicted behavior category and the behavior category label are compared to train the user account recall model. As an optional example, the user account recall model is implemented by a deep neural network. As an optional example, the third multimedia data is third advertising multimedia data.
[0157] It should be noted that in the embodiments of the present application, the user feature vector of the user account and the user account's behavior category label for the third multimedia data do not involve the user's sensitive information. The user feature vector of the user account and the user account's behavior category label for the third multimedia data are obtained and used after user authorization. In one example, before obtaining the user feature vector of the user account and the user account's behavior category label for the third multimedia data, a corresponding interface displays a prompt message related to obtaining data use authorization, and the user determines whether to agree to the authorization based on the prompt message.
[0158] In one possible implementation, the present embodiment provides a specific implementation method for obtaining a user feature vector corresponding to a target user account, including:
[0159] E1: Collect user information of a target user account, and generate a first user feature of the target user account according to the user information of the target user account.
[0160] E2: Obtain the second user feature of the target user account obtained through pre-training.
[0161] E3: Use the first user feature of the target user account and the second user feature of the target user account as a user feature vector of the target user account.
[0162] E1-E3 in the embodiment of the present application are similar to A1-A3 in the above embodiment. For the sake of brevity, they will not be described in detail here. For detailed information, please refer to the description in the above embodiment.
[0163] It should be noted that in the embodiments of the present application, the user information of the target user account, the first user feature of the target user account, and the second user feature of the target user account do not involve sensitive information of the user. The user information of the target user account, the first user feature of the target user account, and the second user feature of the target user account are obtained and used after authorization by the user. In one example, before obtaining the user information of the target user account, the first user feature of the target user account, and the second user feature of the target user account, the corresponding interface displays a prompt message related to obtaining data use authorization, and the user determines whether to agree to the authorization based on the prompt message.
[0164] S704: Input the content feature vector of the target multimedia data and the user feature vector of the target user account into the content detection model to obtain the prediction result of the behavior category of the target user account for the target multimedia data. The content detection model is trained according to the training method of the content detection model of any of the above embodiments.
[0165] After obtaining the content feature vector of the target multimedia data and the user feature vector of the target user account, the content feature vector of the target multimedia data and the user feature vector of the target user account can be input into the content detection model to obtain a prediction result of the target user account's behavior category with respect to the target multimedia data. The obtained prediction result indicates the target user account's degree of liking for the target multimedia data.
[0166] In a possible implementation, the content detection method provided by the embodiment of the present application also includes: calculating the content detection evaluation result of the target multimedia data based on the prediction result of the target user account for the behavior category of the target multimedia data. As an optional example, when the prediction result of the target user account for the behavior category of the target multimedia data is an evaluation value, the content detection evaluation result of the target multimedia data is the average of the evaluation values of each behavior category. The evaluation value of each behavior category can be the average of the evaluation values of multiple target user accounts for the behavior category. For example, for a certain target multimedia data, the average value of the prediction results of the like behavior category of multiple target user accounts is 0.7. The average value of the prediction results of the click behavior category of multiple target user accounts is 0.4. If there are only the above two behavior categories, the content detection evaluation result of the target multimedia data obtained is 0.55.
[0167] It is understood that the obtained content detection evaluation result of the target multimedia data is a quantitative representation of the target user account's liking of the target multimedia data, and is also a quantitative representation of whether the target multimedia data is high-quality target multimedia data. For example, when the content detection evaluation result is greater than 0.5, it indicates that the target user account likes the target multimedia data and the target multimedia data is high-quality multimedia data.
[0168] In addition, when the first user feature of the target user account and the second user feature of the target user account are used as the user feature vector of the target user account, an embodiment of the present application provides a specific implementation method of inputting the content feature vector of the target multimedia data and the user feature vector of the target user account into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data, including:
[0169] The content feature vector of the target multimedia data, the first user feature of the target user account, and the second user feature of the target user account are input into the content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data.
[0170] Based on the contents of S701-S704, it can be seen that when checking the target content, the content features of at least one category of the target multimedia data are first extracted, and the content features of each category of the target multimedia data are compared with the respective cluster centers of the content features of the corresponding category to obtain the cluster centers to which the content features of each category of the target multimedia data belong. Based on the cluster centers to which the content features of each category of the target multimedia data belong, the content feature vector of the target multimedia data is obtained. In addition, the user feature vector corresponding to the target user account is obtained. Furthermore, the content feature vector of the target multimedia data and the user feature vector of the target user account are input into the content detection model to obtain the prediction result of the target user account's behavior category for the target multimedia data. In this way, the target multimedia data can be evaluated using the content detection model without having to deliver the target multimedia data, thereby reducing the delivery cost. In addition, since the influence of the user account on the multimedia data evaluation is taken into account during the training process of the content detection model, the prediction result of the target multimedia data using the content detection model is more reasonable and accurate.
[0171] Based on the content detection model training method provided in the above method embodiment, the embodiment of the present application also provides a content detection model training device, and the content detection model training device will be described below with reference to the accompanying drawings.
[0172] See also Figure 9 As shown in FIG, this figure is a structural diagram of a training device for a content detection model provided in an embodiment of the present application. Figure 9 As shown, the training device of the content detection model includes:
[0173] A first extraction unit 901 is configured to extract content features of at least one category of first multimedia data, and cluster the content features of each category of the first multimedia data to obtain multiple cluster centers of the content features of each category;
[0174] A second extraction unit 902 is configured to extract content features of at least one category of the second multimedia data, compare the content features of each category of the second multimedia data with the cluster centers of the content features of the corresponding category, and obtain the cluster center to which the content features of each category of the second multimedia data belong;
[0175] A first acquiring unit 903 is configured to obtain a content feature vector of the second multimedia data according to a cluster center to which the content feature of each category of the second multimedia data belongs;
[0176] A second acquiring unit 904 is configured to acquire a user feature vector of a user account;
[0177] The training unit 905 is used to train a content detection model using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data. The content detection model is used to output a prediction result of the behavior category of the target user account for the target multimedia data.
[0178] In a possible implementation, the first acquiring unit 903 includes:
[0179] A first acquiring subunit is configured to acquire an initial content feature vector corresponding to a cluster center to which content features of each category of the second multimedia data belong;
[0180] The first determining subunit is configured to determine an initial content feature vector corresponding to a cluster center to which the content features of each category of the second multimedia data belong as the content feature vector of the second multimedia data.
[0181] In a possible implementation, the apparatus further includes:
[0182] an adjusting unit, configured to adjust the content feature vector of the second multimedia data during training of the content detection model;
[0183] a determination unit, configured to re-determine the content feature vector corresponding to the cluster center to which the content feature of each category belongs after adjustment as the initial content feature vector corresponding to the cluster center to which the content feature of the category belongs;
[0184] The third acquisition unit is configured to obtain content feature vectors corresponding to multiple cluster centers of content features of each category after the training of the content detection model is completed.
[0185] In a possible implementation, the apparatus further includes:
[0186] a calculation unit, configured to calculate, based on the content features of each category, content feature vectors corresponding to a plurality of cluster centers of the content features of each category;
[0187] The first acquiring unit 903 includes:
[0188] The second determining subunit is configured to determine the content feature vector corresponding to the cluster center to which the content feature of each category of the second multimedia data belongs as the content feature vector of the second multimedia data.
[0189] In a possible implementation, the second acquiring unit 904 includes:
[0190] a collection subunit, configured to collect user information of a user account, and generate a first user feature of the user account based on the user information of the user account;
[0191] A second acquisition subunit is used to obtain a second user feature of the user account obtained through pre-training;
[0192] The third determining subunit is configured to use the first user feature of the user account and the second user feature of the user account as a user feature vector of the user account.
[0193] In a possible implementation, the content detection model includes a first cross-feature extraction module and a connection module, and the training unit 905 includes:
[0194] a first input subunit, configured to input the content feature vector of the second multimedia data and the user feature vector of the user account into the first cross-feature extraction module, so that the first cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a first feature vector;
[0195] a second input subunit, configured to input the content feature vector of the second multimedia data and the user feature vector of the user account into the connection module, so that the connection module connects the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a second feature vector;
[0196] The first training subunit is configured to train the content detection model by using the first feature vector, the second feature vector, and a behavior category label of the user account for the second multimedia data.
[0197] In a possible implementation, the content detection model includes a second cross-feature extraction module, a third cross-feature extraction module, and a connection module, and the training unit 905 includes:
[0198] a third input subunit, configured to input the content feature vector of the second multimedia data and the first user feature into the second cross-feature extraction module, so that the second cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the first user feature to obtain a third feature vector;
[0199] a fourth input subunit, configured to input the content feature vector of the second multimedia data and the second user feature into the third cross-feature extraction module, so that the third cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the second user feature to obtain a fourth feature vector;
[0200] a fifth input subunit, configured to input the content feature vector of the second multimedia data, the first user feature, and the second user feature into the connection module, so that the connection module connects the content feature vector of the second multimedia data, the first user feature, and the second user feature to obtain a fifth feature vector;
[0201] The second training subunit is configured to train the content detection model by using the third eigenvector, the fourth eigenvector, the fifth eigenvector, and the behavior category label of the user account for the second multimedia data.
[0202] Based on the content detection method provided in the above method embodiment, the embodiment of the present application further provides a content detection device, which will be described below with reference to the accompanying drawings.
[0203] See also Figure 10 As shown in FIG, this figure is a structural diagram of a content detection device provided by an embodiment of the present application. Figure 10 As shown, the content detection device includes:
[0204] An extraction unit 1001 is configured to extract content features of at least one category of the target multimedia data, compare the content features of each category of the target multimedia data with the respective cluster centers of the content features of the corresponding category, and obtain the cluster center to which the content features of each category of the target multimedia data belong;
[0205] A first obtaining unit 1002 is configured to obtain a content feature vector of the target multimedia data according to a cluster center to which the content feature of each category of the target multimedia data belongs;
[0206] The second acquisition unit 1003 is used to obtain a user feature vector corresponding to the target user account;
[0207] The first input unit 1004 is used to input the content feature vector of the target multimedia data and the user feature vector of the target user account into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data. The content detection model is trained according to the training method of the content detection model described in any of the above items.
[0208] In a possible implementation, the apparatus further includes:
[0209] A calculation unit is used to calculate a content detection evaluation result of the target multimedia data according to a prediction result of the behavior category of the target user account for the target multimedia data.
[0210] In a possible implementation, the apparatus further includes:
[0211] The second input unit is used to input the content feature vector of the target multimedia data into the user account recall model before obtaining the user feature vector corresponding to the target user account, so as to obtain the target user account corresponding to the target multimedia data; the user account recall model is trained based on the content feature vector of the third multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the third multimedia data.
[0212] In a possible implementation, the second acquiring unit 1003 includes:
[0213] a collection subunit, configured to collect user information of a target user account, and generate a first user feature of the target user account based on the user information of the target user account;
[0214] A first acquisition subunit is configured to acquire a second user feature of the target user account obtained through pre-training;
[0215] a determining subunit, configured to use the first user feature of the target user account and the second user feature of the target user account as a user feature vector of the target user account;
[0216] The first input unit 1004 is specifically configured to:
[0217] The content feature vector of the target multimedia data, the first user feature of the target user account, and the second user feature of the target user account are input into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data.
[0218] Based on the content detection model training method and content detection method provided in the above method embodiments, the present application also provides an electronic device, including: one or more processors; a storage device on which one or more programs are stored, when the one or more programs are executed by the one or more processors, the one or more processors implement the content detection model training method described in any of the above embodiments, or the content detection method described in any of the above embodiments.
[0219] Reference below Figure 11 , which shows a schematic structural diagram of an electronic device 1100 suitable for implementing an embodiment of the present application. The terminal device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (portable Android devices), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs (televisions) and desktop computers. Figure 11 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0220] like Figure 11 As shown, the electronic device 1100 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage device 1106 into a random access memory (RAM) 1103. Various programs and data required for the operation of the electronic device 1100 are also stored in the RAM 1103. The processing device 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0221] Typically, the following devices may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1106 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109. The communication device 1109 may allow the electronic device 1100 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 11The electronic device 1100 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0222] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1109, or installed from the storage device 1106, or installed from the ROM 1102. When the computer program is executed by the processing device 1101, the above-mentioned functions defined in the method of the embodiment of the present application are performed.
[0223] The electronic device provided in the embodiment of the present application and the training method and content detection method of the content detection model provided in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0224] Based on the content detection model training method and content detection method provided in the above method embodiments, an embodiment of the present application provides a computer-readable medium on which a computer program is stored, wherein when the program is executed by a processor, it implements the content detection model training method described in any of the above embodiments, or the content detection method described in any of the above embodiments.
[0225] It should be noted that the computer-readable medium mentioned above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0226] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0227] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0228] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the content detection model training method or content detection method.
[0229] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0230] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0231] The units involved in the embodiments described in this application may be implemented by software or hardware. In some cases, the name of a unit / module does not constitute a limitation of the unit itself. For example, a voice data acquisition module may also be described as a "data acquisition module."
[0232] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0233] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0234] According to one or more embodiments of the present application, [Example 1] provides a method for training a content detection model, which extracts content features of at least one category of first multimedia data, clusters the content features of each category of the first multimedia data, and obtains multiple cluster centers of the content features of each category. The method includes:
[0235] extracting content features of at least one category of the second multimedia data, comparing the content features of each category of the second multimedia data with respective cluster centers of the content features of the corresponding category, and obtaining the cluster center to which the content features of each category of the second multimedia data belong;
[0236] obtaining a content feature vector of the second multimedia data according to the cluster center to which the content feature of each category of the second multimedia data belongs;
[0237] Get the user feature vector of the user account;
[0238] A content detection model is trained using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data. The content detection model is used to output a prediction result of the behavior category of the target user account for the target multimedia data.
[0239] According to one or more embodiments of the present application, [Example 2] provides a method for training a content detection model, wherein obtaining a content feature vector of the second multimedia data based on the cluster center to which the content feature of each category of the second multimedia data belongs includes:
[0240] Obtaining an initial content feature vector corresponding to a cluster center to which content features of each category of the second multimedia data belong;
[0241] An initial content feature vector corresponding to the cluster center to which the content feature of each category of the second multimedia data belongs is determined as the content feature vector of the second multimedia data.
[0242] According to one or more embodiments of the present application, [Example 3] provides a method for training a content detection model, the method further comprising:
[0243] During training of the content detection model, adjusting the content feature vector of the second multimedia data;
[0244] Re-determine the content feature vector corresponding to the cluster center of the content feature of each category after adjustment as the initial content feature vector corresponding to the cluster center of the content feature of the category;
[0245] After the training of the content detection model is completed, content feature vectors corresponding to multiple cluster centers of the content features of each category are obtained.
[0246] According to one or more embodiments of the present application, [Example 4] provides a method for training a content detection model, the method further comprising:
[0247] Calculating content feature vectors corresponding to a plurality of cluster centers of the content features of each category according to the content features of each category;
[0248] The obtaining, according to the cluster center to which the content feature of each category of the second multimedia data belongs, a content feature vector of the second multimedia data includes:
[0249] The content feature vector corresponding to the cluster center to which the content feature of each category of the second multimedia data belongs is determined as the content feature vector of the second multimedia data.
[0250] According to one or more embodiments of the present application, [Example 5] provides a method for training a content detection model, wherein obtaining a user feature vector of a user account includes:
[0251] Collecting user information of a user account, and generating a first user feature of the user account based on the user information of the user account;
[0252] Obtaining a second user feature of the user account obtained through pre-training;
[0253] The first user feature of the user account and the second user feature of the user account are used as a user feature vector of the user account.
[0254] According to one or more embodiments of the present application, [Example 6] provides a method for training a content detection model, the content detection model including a first cross-feature extraction module and a connection module. The method for training the content detection model using a content feature vector of the second multimedia data, a user feature vector of the user account, and a behavior category label of the user account for the second multimedia data includes:
[0255] inputting the content feature vector of the second multimedia data and the user feature vector of the user account into the first cross-feature extraction module, so that the first cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a first feature vector;
[0256] inputting the content feature vector of the second multimedia data and the user feature vector of the user account into the connection module, so that the connection module connects the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a second feature vector;
[0257] The content detection model is trained using the first feature vector, the second feature vector, and a behavior category label of the user account for the second multimedia data.
[0258] According to one or more embodiments of the present application, [Example 7] provides a method for training a content detection model, the content detection model including a second cross-feature extraction module, a third cross-feature extraction module, and a connection module. Training the content detection model using a content feature vector of the second multimedia data, a user feature vector of a user account, and a behavior category label of the user account for the second multimedia data includes:
[0259] inputting the content feature vector of the second multimedia data and the first user feature into the second cross-feature extraction module, so that the second cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the first user feature to obtain a third feature vector;
[0260] inputting the content feature vector of the second multimedia data and the second user feature into the third cross-feature extraction module, so that the third cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the second user feature to obtain a fourth feature vector;
[0261] inputting the content feature vector of the second multimedia data, the first user feature, and the second user feature into the connection module, so that the connection module connects the content feature vector of the second multimedia data, the first user feature, and the second user feature to obtain a fifth feature vector;
[0262] The content detection model is trained using the third eigenvector, the fourth eigenvector, the fifth eigenvector, and the behavior category label of the user account for the second multimedia data.
[0263] According to one or more embodiments of the present application, [Example 8] provides a content detection method, the method comprising:
[0264] extracting content features of at least one category of the target multimedia data, comparing the content features of each category of the target multimedia data with respective cluster centers of content features of the corresponding category, and obtaining the cluster center to which the content features of each category of the target multimedia data belong;
[0265] Obtaining a content feature vector of the target multimedia data according to the cluster center to which the content feature of each category of the target multimedia data belongs;
[0266] Obtain the user feature vector corresponding to the target user account;
[0267] The content feature vector of the target multimedia data and the user feature vector of the target user account are input into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data. The content detection model is trained according to the training method of the content detection model described in any of the above items.
[0268] According to one or more embodiments of the present application, [Example 9] provides a content detection method, the method further comprising:
[0269] A content detection evaluation result of the target multimedia data is calculated according to a prediction result of the target user account for the behavior category of the target multimedia data.
[0270] According to one or more embodiments of the present application, [Example 10] provides a content detection method. Before obtaining a user feature vector corresponding to a target user account, the method further includes:
[0271] The content feature vector of the target multimedia data is input into the user account recall model to obtain the target user account corresponding to the target multimedia data; the user account recall model is trained based on the content feature vector of the third multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the third multimedia data.
[0272] According to one or more embodiments of the present application, [Example 11] provides a content detection method, wherein obtaining a user feature vector corresponding to a target user account includes:
[0273] Collecting user information of a target user account, and generating a first user feature of the target user account based on the user information of the target user account;
[0274] Obtaining a second user feature of the target user account obtained through pre-training;
[0275] Using the first user feature of the target user account and the second user feature of the target user account as a user feature vector of the target user account;
[0276] Inputting the content feature vector of the target multimedia data and the user feature vector of the target user account into a content detection model to obtain a prediction result of the behavior category of the target user account with respect to the target multimedia data includes:
[0277] The content feature vector of the target multimedia data, the first user feature of the target user account, and the second user feature of the target user account are input into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data.
[0278] According to one or more embodiments of the present application, [Example 12] provides a training device for a content detection model, the device comprising:
[0279] a first extraction unit, configured to extract content features of at least one category of the first multimedia data, and cluster the content features of each category of the first multimedia data to obtain a plurality of cluster centers of the content features of each category;
[0280] a second extraction unit, configured to extract content features of at least one category of the second multimedia data, compare the content features of each category of the second multimedia data with the respective cluster centers of the content features of the corresponding category, and obtain the cluster center to which the content features of each category of the second multimedia data belong;
[0281] a first acquiring unit, configured to obtain a content feature vector of the second multimedia data according to a cluster center to which the content feature of each category of the second multimedia data belongs;
[0282] A second acquisition unit is used to obtain a user feature vector of a user account;
[0283] A training unit is used to train a content detection model using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data. The content detection model is used to output a prediction result of the behavior category of the target user account for the target multimedia data.
[0284] According to one or more embodiments of the present application, [Example 13] provides a training device for a content detection model, wherein the first acquisition unit includes:
[0285] A first acquiring subunit is configured to acquire an initial content feature vector corresponding to a cluster center to which content features of each category of the second multimedia data belong;
[0286] The first determining subunit is configured to determine an initial content feature vector corresponding to a cluster center to which the content features of each category of the second multimedia data belong as the content feature vector of the second multimedia data.
[0287] According to one or more embodiments of the present application, [Example 14] provides a training device for a content detection model, the device further comprising:
[0288] an adjusting unit, configured to adjust the content feature vector of the second multimedia data during training of the content detection model;
[0289] a determination unit, configured to re-determine the content feature vector corresponding to the cluster center to which the content feature of each category belongs after adjustment as the initial content feature vector corresponding to the cluster center to which the content feature of the category belongs;
[0290] The third acquisition unit is configured to obtain content feature vectors corresponding to multiple cluster centers of content features of each category after the training of the content detection model is completed.
[0291] According to one or more embodiments of the present application, [Example 15] provides a training device for a content detection model, the device further comprising:
[0292] a calculation unit, configured to calculate, based on the content features of each category, content feature vectors corresponding to a plurality of cluster centers of the content features of each category;
[0293] The first acquiring unit includes:
[0294] The second determining subunit is configured to determine the content feature vector corresponding to the cluster center to which the content feature of each category of the second multimedia data belongs as the content feature vector of the second multimedia data.
[0295] According to one or more embodiments of the present application, [Example 16] provides a training device for a content detection model, wherein the second acquisition unit includes:
[0296] a collection subunit, configured to collect user information of a user account, and generate a first user feature of the user account based on the user information of the user account;
[0297] A second acquisition subunit is used to obtain a second user feature of the user account obtained through pre-training;
[0298] The third determining subunit is configured to use the first user feature of the user account and the second user feature of the user account as a user feature vector of the user account.
[0299] According to one or more embodiments of the present application, [Example 17] provides a training device for a content detection model, wherein the content detection model includes a first cross-feature extraction module and a connection module, and the training unit includes:
[0300] a first input subunit, configured to input the content feature vector of the second multimedia data and the user feature vector of the user account into the first cross-feature extraction module, so that the first cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a first feature vector;
[0301] a second input subunit, configured to input the content feature vector of the second multimedia data and the user feature vector of the user account into the connection module, so that the connection module connects the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a second feature vector;
[0302] The first training subunit is configured to train the content detection model by using the first feature vector, the second feature vector, and a behavior category label of the user account for the second multimedia data.
[0303] According to one or more embodiments of the present application, [Example 18] provides a training device for a content detection model, wherein the content detection model includes a second cross-feature extraction module, a third cross-feature extraction module, and a connection module, and the training unit includes:
[0304] a third input subunit, configured to input the content feature vector of the second multimedia data and the first user feature into the second cross-feature extraction module, so that the second cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the first user feature to obtain a third feature vector;
[0305] a fourth input subunit, configured to input the content feature vector of the second multimedia data and the second user feature into the third cross-feature extraction module, so that the third cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the second user feature to obtain a fourth feature vector;
[0306] a fifth input subunit, configured to input the content feature vector of the second multimedia data, the first user feature, and the second user feature into the connection module, so that the connection module connects the content feature vector of the second multimedia data, the first user feature, and the second user feature to obtain a fifth feature vector;
[0307] The second training subunit is configured to train the content detection model by using the third eigenvector, the fourth eigenvector, the fifth eigenvector, and the behavior category label of the user account for the second multimedia data.
[0308] According to one or more embodiments of the present application, [Example 19] provides a content detection device, the device including:
[0309] an extraction unit, configured to extract content features of at least one category of the target multimedia data, compare the content features of each category of the target multimedia data with respective cluster centers of content features of the corresponding category, and obtain the cluster center to which the content features of each category of the target multimedia data belong;
[0310] a first acquiring unit, configured to obtain a content feature vector of the target multimedia data according to a cluster center to which the content feature of each category of the target multimedia data belongs;
[0311] A second acquisition unit is used to obtain a user feature vector corresponding to the target user account;
[0312] A first input unit is used to input the content feature vector of the target multimedia data and the user feature vector of the target user account into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data. The content detection model is trained according to the training method of the content detection model described in any of the above items.
[0313] According to one or more embodiments of the present application, [Example 20] provides a content detection device, the device further comprising:
[0314] A calculation unit is used to calculate a content detection evaluation result of the target multimedia data according to a prediction result of the behavior category of the target user account for the target multimedia data.
[0315] According to one or more embodiments of the present application, [Example 21] provides a content detection device, the device further comprising:
[0316] The second input unit is used to input the content feature vector of the target multimedia data into the user account recall model before obtaining the user feature vector corresponding to the target user account, so as to obtain the target user account corresponding to the target multimedia data; the user account recall model is trained based on the content feature vector of the third multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the third multimedia data.
[0317] According to one or more embodiments of the present application, [Example 22] provides a content detection device, wherein the second acquisition unit includes:
[0318] a collection subunit, configured to collect user information of a target user account, and generate a first user feature of the target user account based on the user information of the target user account;
[0319] A first acquisition subunit is configured to acquire a second user feature of the target user account obtained through pre-training;
[0320] a determining subunit, configured to use the first user feature of the target user account and the second user feature of the target user account as a user feature vector of the target user account;
[0321] The first input unit is specifically configured to:
[0322] The content feature vector of the target multimedia data, the first user feature of the target user account, and the second user feature of the target user account are input into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data.
[0323] According to one or more embodiments of the present application, [Example 23] provides an electronic device, including:
[0324] one or more processors;
[0325] a storage device having one or more programs stored thereon,
[0326] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the above-described content detection model training methods or any of the above-described content detection methods.
[0327] According to one or more embodiments of the present application, [Example 24] provides a computer-readable medium on which a computer program is stored, wherein when the program is executed by a processor, it implements the training method of the content detection model as described in any of the above, or the content detection method as described in any of the above.
[0328] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0329] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0330] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0331] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0332] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for training a content detection model, characterized in that: Extracting content features of at least one category of first multimedia data, clustering the content features of each category of the first multimedia data, and obtaining multiple cluster centers of the content features of each category; the method includes: extracting content features of at least one category of the second multimedia data, comparing the content features of each category of the second multimedia data with respective cluster centers of the content features of the corresponding category, and obtaining the cluster center to which the content features of each category of the second multimedia data belong; obtaining a content feature vector of the second multimedia data according to the cluster center to which the content feature of each category of the second multimedia data belongs; Get the user feature vector of the user account; Training a content detection model using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account with respect to the second multimedia data, wherein the content detection model is configured to output a prediction result of the behavior category of the target user account with respect to the target multimedia data; The content detection model includes a first cross-feature extraction module and a connection module. The content detection model is trained by using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data, including: inputting the content feature vector of the second multimedia data and the user feature vector of the user account into the first cross-feature extraction module, so that the first cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a first feature vector; inputting the content feature vector of the second multimedia data and the user feature vector of the user account into the connection module, so that the connection module connects the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a second feature vector; The content detection model is trained using the first feature vector, the second feature vector, and a behavior category label of the user account for the second multimedia data.
2. The method according to claim 1, characterized in that The obtaining, according to the cluster center to which the content feature of each category of the second multimedia data belongs, a content feature vector of the second multimedia data includes: Obtaining an initial content feature vector corresponding to a cluster center to which content features of each category of the second multimedia data belong; An initial content feature vector corresponding to the cluster center to which the content feature of each category of the second multimedia data belongs is determined as the content feature vector of the second multimedia data.
3. The method according to claim 2, characterized in that The method further comprises: During training of the content detection model, adjusting the content feature vector of the second multimedia data; Re-determine the content feature vector corresponding to the cluster center of the content feature of each category after adjustment as the initial content feature vector corresponding to the cluster center of the content feature of the category; After the training of the content detection model is completed, content feature vectors corresponding to multiple cluster centers of the content features of each category are obtained.
4. The method according to claim 1, wherein The method further comprises: Calculating content feature vectors corresponding to a plurality of cluster centers of the content features of each category according to the content features of each category; The obtaining, according to the cluster center to which the content feature of each category of the second multimedia data belongs, a content feature vector of the second multimedia data includes: The content feature vector corresponding to the cluster center to which the content feature of each category of the second multimedia data belongs is determined as the content feature vector of the second multimedia data.
5. The method according to claim 1, wherein The obtaining of the user feature vector of the user account includes: Collecting user information of a user account, and generating a first user feature of the user account based on the user information of the user account; Obtaining a second user feature of the user account obtained through pre-training; The first user feature of the user account and the second user feature of the user account are used as a user feature vector of the user account.
6. The method according to claim 5, characterized in that The content detection model includes a second cross-feature extraction module, a third cross-feature extraction module, and a connection module. The content detection model is trained using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the second multimedia data, including: inputting the content feature vector of the second multimedia data and the first user feature into the second cross-feature extraction module, so that the second cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the first user feature to obtain a third feature vector; inputting the content feature vector of the second multimedia data and the second user feature into the third cross-feature extraction module, so that the third cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the second user feature to obtain a fourth feature vector; inputting the content feature vector of the second multimedia data, the first user feature, and the second user feature into the connection module, so that the connection module connects the content feature vector of the second multimedia data, the first user feature, and the second user feature to obtain a fifth feature vector; The content detection model is trained using the third eigenvector, the fourth eigenvector, the fifth eigenvector, and the behavior category label of the user account for the second multimedia data.
7. A content detection method, characterized in that: The method comprises: extracting content features of at least one category of the target multimedia data, comparing the content features of each category of the target multimedia data with respective cluster centers of content features of the corresponding category, and obtaining the cluster center to which the content features of each category of the target multimedia data belong; Obtaining a content feature vector of the target multimedia data according to the cluster center to which the content feature of each category of the target multimedia data belongs; Obtain the user feature vector corresponding to the target user account; The content feature vector of the target multimedia data and the user feature vector of the target user account are input into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data. The content detection model is trained according to the training method of the content detection model according to any one of claims 1-6.
8. The method according to claim 7, characterized in that The method further comprises: A content detection evaluation result of the target multimedia data is calculated according to a prediction result of the behavior category of the target user account for the target multimedia data.
9. The method according to claim 7, characterized in that Before obtaining the user feature vector corresponding to the target user account, the method further includes: The content feature vector of the target multimedia data is input into the user account recall model to obtain the target user account corresponding to the target multimedia data; the user account recall model is trained based on the content feature vector of the third multimedia data, the user feature vector of the user account, and the behavior category label of the user account for the third multimedia data.
10. The method according to claim 7, characterized in that The obtaining of the user feature vector corresponding to the target user account includes: Collecting user information of a target user account, and generating a first user feature of the target user account based on the user information of the target user account; Obtaining a second user feature of the target user account obtained through pre-training; Using the first user feature of the target user account and the second user feature of the target user account as a user feature vector of the target user account; Inputting the content feature vector of the target multimedia data and the user feature vector of the target user account into a content detection model to obtain a prediction result of the behavior category of the target user account with respect to the target multimedia data includes: The content feature vector of the target multimedia data, the first user feature of the target user account, and the second user feature of the target user account are input into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data.
11. A training device for a content detection model, characterized in that: The device comprises: a first extraction unit, configured to extract content features of at least one category of the first multimedia data, and cluster the content features of each category of the first multimedia data to obtain a plurality of cluster centers of the content features of each category; a second extraction unit, configured to extract content features of at least one category of the second multimedia data, compare the content features of each category of the second multimedia data with the respective cluster centers of the content features of the corresponding category, and obtain the cluster center to which the content features of each category of the second multimedia data belong; a first acquiring unit, configured to obtain a content feature vector of the second multimedia data according to a cluster center to which the content feature of each category of the second multimedia data belongs; A second acquiring unit, configured to acquire a user feature vector of a user account; a training unit, configured to train a content detection model using the content feature vector of the second multimedia data, the user feature vector of the user account, and the behavior category label of the user account with respect to the second multimedia data, the content detection model being configured to output a prediction result of the behavior category of the target user account with respect to the target multimedia data; The content detection model includes a first cross-feature extraction module and a connection module, and the training unit includes: a first input subunit, configured to input the content feature vector of the second multimedia data and the user feature vector of the user account into the first cross-feature extraction module, so that the first cross-feature extraction module performs cross-feature extraction on the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a first feature vector; a second input subunit, configured to input the content feature vector of the second multimedia data and the user feature vector of the user account into the connection module, so that the connection module connects the content feature vector of the second multimedia data and the user feature vector of the user account to obtain a second feature vector; The first training subunit is configured to train the content detection model by using the first feature vector, the second feature vector, and a behavior category label of the user account for the second multimedia data.
12. A content detection device, characterized in that: The device comprises: an extraction unit, configured to extract content features of at least one category of the target multimedia data, compare the content features of each category of the target multimedia data with respective cluster centers of content features of the corresponding category, and obtain the cluster center to which the content features of each category of the target multimedia data belong; a first acquiring unit, configured to obtain a content feature vector of the target multimedia data according to a cluster center to which the content feature of each category of the target multimedia data belongs; A second acquisition unit is used to obtain a user feature vector corresponding to the target user account; The first input unit is used to input the content feature vector of the target multimedia data and the user feature vector of the target user account into a content detection model to obtain a prediction result of the behavior category of the target user account for the target multimedia data, and the content detection model is trained according to the training method of the content detection model according to any one of claims 1-6.
13. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the content detection model training method as described in any one of claims 1-6, or the content detection method as described in any one of claims 7-10.
14. A computer-readable medium, characterized in that A computer program is stored thereon, wherein when the program is executed by a processor, the training method of the content detection model as described in any one of claims 1-6, or the content detection method as described in any one of claims 7-10 is implemented.
Citation Information
Patent Citations
Video retrieval method and device, electronic equipment and storage medium
CN111241345A
Multimedia data pushing method and device
CN113761364A