Methods, devices, electronic equipment, and storage media for generating account characteristics

By extracting and fusing features from the historical multimedia resources of the account to be identified, account features are generated, which solves the problem of low accuracy of multimedia resource features in the existing technology and improves the accuracy of resource features associated with multimedia resources.

CN115618026BActive Publication Date: 2026-05-26BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2022-09-15
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of multimedia resource features based on single-value behavior and dense features of user accounts is low, which leads to a low accuracy of resource features associated with multimedia resources.

Method used

By acquiring the historical multimedia resource set of the account to be identified, feature extraction and feature fusion are performed to generate account features. The account feature generation model is then used to train the feature extraction and fusion model, thereby improving the accuracy of the account features.

Benefits of technology

This improved the accuracy of account features, thereby improving the accuracy of resource features associated with multimedia resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115618026B_ABST
    Figure CN115618026B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method, apparatus, electronic device, and storage medium for generating account features. The method includes: acquiring a historical multimedia resource set of an account to be identified; performing feature extraction processing on each historical multimedia resource in the historical multimedia resource set to obtain multiple historical multimedia resource features; and performing feature fusion processing on the multiple historical multimedia resource features to obtain account features of the account to be identified. The account features are used in subsequent processing of associated multimedia resources of the account to be identified. This disclosure extracts multiple historical multimedia resource features from the historical multimedia resource set of the account to be identified and fuses these features to obtain fused features. These fused features can accurately characterize account features. Therefore, using these fused features as account features of the account to be identified improves the accuracy of account features, thereby improving the accuracy of determining resource features of associated multimedia resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and storage medium for generating account characteristics. Background Technology

[0002] In related technologies, in the scenario of predicting category information of multimedia resources, it is usually necessary to combine account characteristics to determine the resource characteristics of multimedia information, such as audio features and video features. In the process of determining account characteristics, the single-value behavior (sparse features) of the user account, such as the number of followers or the number of historical works of the user account, is generally used as account features, and the category information of multimedia resources is predicted in combination with the account features. On the basis of single-value behavior, dense features of the user account can also be added. For example, the average feature of multiple multimedia messages recently played by the user account can be used as the dense feature of the user, and then the dense feature of the user can be used as the account feature. However, the average feature of multimedia information cannot accurately represent the account features, and the accuracy of the account features obtained is low. Based on the account features, processing the associated multimedia resources of the account will lead to low accuracy of the resource features of the associated multimedia resources. Summary of the Invention

[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for generating account features, to at least solve the problem of low accuracy in determining resource features of associated multimedia resources in related technologies. The technical solution of this disclosure is as follows:

[0004] According to a first aspect of the present disclosure, a method for generating account features is provided, comprising:

[0005] Obtain the historical multimedia resource set of the account to be identified; the historical multimedia resource set includes multiple historical multimedia resources.

[0006] Feature extraction processing is performed on each historical multimedia resource in the historical multimedia resource set to obtain the historical multimedia resource features corresponding to each historical multimedia resource.

[0007] The features of multiple historical multimedia resources are fused to obtain the account features of the account to be identified. The account features are used in subsequent processing of the associated multimedia resources of the account to be identified.

[0008] In one exemplary embodiment, the feature fusion processing of multiple historical multimedia resource features to obtain the account features of the account to be identified includes:

[0009] The historical multimedia resource features are divided into a preset number of subsets of historical resource features, and each of the historical multimedia resource features is extracted by the feature extraction model in the account feature generation model.

[0010] Each set of historical resource feature subsets is input into the feature fusion model in the account feature generation model. The feature fusion model performs feature fusion processing on each set of historical resource feature subsets to obtain historical fusion features corresponding to each set of historical resource feature subsets.

[0011] The predetermined number of historical fusion features obtained are determined as the account features of the account to be identified.

[0012] In one exemplary embodiment, the training method for the account feature generation model includes:

[0013] Obtain the sample multimedia resource set of the sample account; the sample multimedia resource set includes multiple sample multimedia resources;

[0014] The sample multimedia resource set is input into the initial account feature generation model, which includes an initial feature extraction model and an initial feature fusion model.

[0015] The initial feature extraction model is used to perform feature extraction processing on each of the sample multimedia resources to obtain a sample multimedia resource feature set; the sample multimedia resource feature set includes the preset number of sample resource feature subsets.

[0016] The initial feature fusion model is used to perform feature fusion processing on each set of sample resource feature subsets to obtain sample fusion features corresponding to each set of sample resource feature subsets. A preset number of sample fusion features corresponding to the sample account are determined as the sample account features of the sample account.

[0017] Based on the characteristics of the sample accounts, determine the target loss value;

[0018] The model parameters of the initial account feature generation model are adjusted based on the target loss value to obtain the account feature generation model.

[0019] In one exemplary implementation, adjusting the model parameters of the initial account feature generation model based on the target loss value to obtain the account feature generation model includes:

[0020] Based on the target loss value, adjust the model parameters of the initial feature extraction model and the initial feature fusion model respectively until the preset training termination condition is met and the training ends.

[0021] The initial feature extraction model at the end of training is used as the feature extraction model, and the initial feature fusion model at the end of training is used as the feature fusion model.

[0022] Based on the feature extraction model and the feature fusion model, the account feature generation model is constructed.

[0023] In one exemplary implementation, the sample accounts are at least two, and the method further includes:

[0024] Obtain a preset number of sample resource feature subsets corresponding to each sample account; each sample multimedia resource subset is labeled with a sample category tag.

[0025] The sample multimedia resources in each of the sample resource feature subsets are subjected to category identification processing to obtain the sample category results of each sample multimedia resource;

[0026] The step of determining the target loss value based on the characteristics of the sample accounts includes:

[0027] The target loss value is determined based on the sample account characteristics, the sample category labels, and the sample category results.

[0028] In one exemplary implementation, determining the target loss value based on the sample account features, the sample category label, and the sample category result includes:

[0029] A first loss value is generated based on the sample category labels and sample category results corresponding to the sample multimedia resources.

[0030] Based on the characteristics of the sample accounts, a second loss value is generated;

[0031] The target loss value is determined based on the first loss value and the second loss value.

[0032] In one exemplary implementation, generating a second loss value based on the sample account features includes:

[0033] Obtain any two sample resource feature subsets corresponding to any sample account to obtain two first sample resource feature subsets;

[0034] Obtain a subset of sample resource features corresponding to each of any two sample accounts to obtain two second subsets of sample resource features;

[0035] Based on the sample fusion features corresponding to each of the two first sample resource feature subsets, the first similarity between the two first sample resource feature subsets is determined.

[0036] Based on the sample fusion features corresponding to each of the two second sample resource feature subsets, a second similarity between the two second sample resource feature subsets is determined.

[0037] The second loss value is determined based on the first similarity and the second similarity.

[0038] In one exemplary embodiment, the method further includes:

[0039] Obtain the multimedia resources to be identified for the account to be identified;

[0040] Based on the account characteristics and the multimedia resource to be identified, the category information of the multimedia resource to be identified is determined.

[0041] According to a second aspect of the present disclosure, an apparatus for generating account features is provided, comprising:

[0042] The historical resource acquisition module is configured to acquire the historical multimedia resource set of the account to be identified; the historical multimedia resource set includes multiple historical multimedia resources.

[0043] The historical feature extraction module is configured to perform feature extraction processing on each historical multimedia resource in the historical multimedia resource set to obtain historical multimedia resource features corresponding to each historical multimedia resource.

[0044] The account feature determination module is configured to perform feature fusion processing on multiple historical multimedia resource features to obtain the account features of the account to be identified, and the account features are used in subsequent processing of the associated multimedia resources of the account to be identified.

[0045] In one exemplary embodiment, the account characteristic determination module includes:

[0046] The resource feature segmentation submodule is configured to perform segmentation of multiple historical multimedia resource features to obtain a preset number of historical resource feature subsets, wherein each historical multimedia resource feature is extracted by the feature extraction model in the account feature generation model;

[0047] The historical fusion feature determination submodule is configured to execute a feature fusion model that inputs each set of historical resource feature subsets into the account feature generation model, and performs feature fusion processing on each set of historical resource feature subsets through the feature fusion model to obtain historical fusion features corresponding to each set of historical resource feature subsets;

[0048] The account feature determination submodule is configured to determine the predetermined number of historical fusion features as the account features of the account to be identified.

[0049] In one exemplary embodiment, the apparatus further includes: a training module for an account feature generation model, the training module comprising:

[0050] The sample resource set acquisition submodule is configured to acquire the sample multimedia resource set of the sample account; the sample multimedia resource set includes multiple sample multimedia resources.

[0051] The resource input submodule is configured to input the sample multimedia resource set into the initial account feature generation model, which includes an initial feature extraction model and an initial feature fusion model.

[0052] The sample feature set determination submodule is configured to perform feature extraction processing on each of the sample multimedia resources through the initial feature extraction model to obtain a sample multimedia resource feature set; the sample multimedia resource feature set includes the preset number of sample resource feature subsets.

[0053] The feature fusion processing submodule is configured to perform feature fusion processing on each set of sample resource feature subsets through the initial feature fusion model to obtain sample fusion features corresponding to each set of sample resource feature subsets, and determine a preset number of sample fusion features corresponding to the sample account as the sample account features of the sample account.

[0054] The target loss value determination submodule is configured to determine the target loss value based on the characteristics of the sample accounts;

[0055] The model generation submodule is configured to adjust the model parameters of the initial account feature generation model based on the target loss value to obtain the account feature generation model.

[0056] In one exemplary implementation, the model generation submodule includes:

[0057] The target loss value determination unit is configured to determine the target loss value based on the characteristics of the sample accounts;

[0058] The training unit is configured to train the initial feature extraction model and the initial feature fusion model based on the target loss value until the training ends when a preset training termination condition is met.

[0059] The feature fusion model determination unit is configured to perform the following: using the initial feature extraction model at the end of training as the feature extraction model, and using the initial feature fusion model at the end of training as the feature fusion model.

[0060] The model building unit is configured to perform the construction of the account feature generation model based on the feature extraction model and the feature fusion model.

[0061] In one exemplary embodiment, the sample accounts are at least two, and the device further includes:

[0062] The sample resource feature subset acquisition module is configured to acquire a preset number of sample resource feature subsets corresponding to each sample account; each sample multimedia resource in the sample resource feature subset is labeled with a sample category tag.

[0063] The sample category result determination module is configured to perform category identification processing on the sample multimedia resources in each subset of the sample resource features to obtain the sample category result for each sample multimedia resource.

[0064] In one exemplary embodiment, the target loss value determination unit includes:

[0065] The target loss value determination subunit is configured to determine the target loss value based on the sample account features, the sample category labels, and the sample category results.

[0066] In one exemplary implementation, the target loss value determination subunit includes:

[0067] The first loss value determination subunit is configured to generate a first loss value based on the sample category label and sample category result corresponding to the sample multimedia resource.

[0068] The second loss value determination subunit is configured to generate a second loss value based on the sample account features;

[0069] The loss value determination subunit is configured to perform the task of determining the target loss value based on the first loss value and the second loss value.

[0070] In one exemplary implementation, the second loss value determination subunit includes:

[0071] The first subset determination subunit is configured to perform the task of obtaining any two sample resource feature subsets corresponding to any sample account, thereby obtaining two first sample resource feature subsets;

[0072] The second subset determination subunit is configured to perform the task of obtaining a sample resource feature subset corresponding to each of any two sample accounts, resulting in two second sample resource feature subsets.

[0073] The first similarity determination subunit is configured to perform a first similarity determination between the two first sample resource feature subsets based on the sample fusion features corresponding to each of the two first sample resource feature subsets.

[0074] The second similarity determination subunit is configured to perform a second similarity determination between the two second sample resource feature subsets based on the sample fusion features corresponding to each of the two second sample resource feature subsets.

[0075] The second loss value determination subunit is configured to determine the second loss value based on the first similarity and the second similarity.

[0076] In one exemplary embodiment, the apparatus further includes:

[0077] The resource acquisition module is configured to acquire the multimedia resources to be identified for the account to be identified.

[0078] The category information recognition module is configured to determine the category information of the multimedia resource to be identified based on the account characteristics and the multimedia resource to be identified.

[0079] According to a third aspect of the present disclosure, an electronic device is provided, comprising:

[0080] processor;

[0081] Memory used to store the processor's executable instructions;

[0082] The processor is configured to execute the instructions to implement the account feature generation method described above.

[0083] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by an electronic device processor, the electronic device is enabled to perform the account feature generation method as described above.

[0084] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method for generating account features as described above.

[0085] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0086] This disclosure obtains a historical multimedia resource set of an account to be identified; the historical multimedia resource set includes multiple historical multimedia resources; feature extraction processing is performed on each historical multimedia resource in the historical multimedia resource set to obtain historical multimedia resource features corresponding to each historical multimedia resource; feature fusion processing is performed on multiple historical multimedia resource features to obtain account features of the account to be identified, and the account features are used in subsequent processing of the associated multimedia resources of the account to be identified. This disclosure extracts multiple historical multimedia resource features from the historical multimedia resource set of the account to be identified, and fuses these multiple historical multimedia resource features to obtain fused features. These fused features can accurately characterize account features; therefore, using these fused features as account features of the account to be identified improves the accuracy of account features; these account features are used in subsequent processing of the associated multimedia resources of the account to be identified, thereby improving the accuracy of determining the resource features of associated multimedia resources.

[0087] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0088] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0089] Figure 1 This is an application environment diagram illustrating a method for generating account features according to an exemplary embodiment.

[0090] Figure 2 This is a flowchart illustrating a method for generating account features according to an exemplary embodiment.

[0091] Figure 3 This is a flowchart illustrating a method for performing feature fusion processing on multiple historical multimedia resource features to obtain the account features of the account to be identified, according to an exemplary embodiment.

[0092] Figure 4 This is a flowchart illustrating a training method for an account feature generation model according to an exemplary embodiment.

[0093] Figure 5 This is a flowchart illustrating a method for determining a target loss value based on sample account characteristics, sample category labels, and sample category results, according to an exemplary embodiment.

[0094] Figure 6 This is a flowchart illustrating a method for generating a second loss value based on sample account features, according to an exemplary embodiment.

[0095] Figure 7 This is a flowchart illustrating a method for adjusting the model parameters of the initial account feature generation model based on a target loss value to obtain the aforementioned account feature generation model, according to an exemplary embodiment.

[0096] Figure 8 This is a flowchart illustrating a method for determining category information of a multimedia resource to be identified, according to an exemplary embodiment.

[0097] Figure 9 This is a schematic diagram illustrating the structure of a video category determination system according to an exemplary embodiment.

[0098] Figure 10 This is a block diagram illustrating an account feature generation apparatus according to an exemplary embodiment.

[0099] Figure 11 This is a block diagram illustrating an electronic device for determining account characteristics according to an exemplary embodiment. Detailed Implementation

[0100] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0101] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0102] It should be noted that the user information (including but not limited to multimedia resources published by users) and data (including but not limited to data used for display and analysis) involved in this disclosure are all information and data authorized by the users or fully authorized by all parties.

[0103] To improve the accuracy of resource characteristics associated with multimedia resources, this disclosure provides a method, apparatus, electronic device, and storage medium for generating account characteristics.

[0104] Please see Figure 1The diagram illustrates an application environment for a method of generating account features according to an exemplary embodiment. The application environment may include a server 01 and a client 02.

[0105] Specifically, in the embodiments of this specification, server 01 may include a standalone server, a distributed server, or a server cluster composed of multiple servers. It may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, model services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Server 01 may include a model communication unit, a processor, and a memory, etc. Specifically, server 01 can be used to obtain the historical multimedia resource set of the account to be identified; the historical multimedia resource set includes multiple historical multimedia resources; feature extraction processing is performed on each historical multimedia resource in the historical multimedia resource set to obtain historical multimedia resource features corresponding to each historical multimedia resource; and feature fusion processing is performed on multiple historical multimedia resource features to obtain the account features of the account to be identified, which are used in subsequent processing of the associated multimedia resources of the account to be identified; and the account features of the account to be identified are sent to client 02.

[0106] Specifically, in the embodiments of this specification, client 02 may include physical devices such as smartphones, desktop computers, tablets, laptops, digital assistants, smart wearable devices, and in-vehicle terminals, and may also include software running on the physical device, such as web pages provided to users by some service providers, or applications provided to users by such service providers. Specifically, client 02 can be used to query the account characteristics of the account to be identified.

[0107] Figure 2 This is a flowchart illustrating a method for generating account features according to an exemplary embodiment, such as... Figure 2 As shown, this method can be applied to Figure 1 The server 01 shown includes the following steps.

[0108] In step S201, the historical multimedia resource set of the account to be identified is obtained; the historical multimedia resource set includes multiple historical multimedia resources.

[0109] In this embodiment of the disclosure, the account to be identified can be the creator, publisher, or viewer of a historical multimedia resource set; the historical multimedia resource set can be a set of multimedia resources published by the account to be identified within a preset historical period; the preset historical period can be set according to actual needs, such as one week, one month, etc.; for example, the historical multimedia resource set can include N multimedia resources most recently published or viewed before the multimedia resource to be identified. Historical multimedia resources can be audio information or video information.

[0110] In step S203, feature extraction processing is performed on each historical multimedia resource in the above-mentioned historical multimedia resource set to obtain the historical multimedia resource features corresponding to each of the above-mentioned historical multimedia resources.

[0111] In this embodiment of the disclosure, the features in the historical multimedia resource feature set can be feature vectors of a preset dimension; the preset dimension can include, but is not limited to, 512-dimensional or 1024-dimensional features. When the historical multimedia resource is audio information, audio features of the audio information can be extracted, and the resulting historical multimedia resource features are audio features; when the historical multimedia resource is video information, video features of the video information can be extracted, and the resulting historical multimedia resource features are video features. The historical multimedia resource features can include, but are not limited to, text features, image features, audio features, etc., of each object in the historical multimedia resource; the objects can include, but are not limited to, physical objects such as people, animals, and plants in the historical multimedia resource, and can also include virtual objects such as cartoon characters, virtual characters, and fluids. During the feature extraction process, object features can be extracted from each historical multimedia resource, and the extracted object features are used as historical multimedia resource features.

[0112] In this embodiment of the disclosure, feature extraction processing is performed on each historical multimedia resource in the historical multimedia resource set to obtain the historical multimedia resource features corresponding to each historical multimedia resource. This may include the following steps S2031-S2035:

[0113] In step S2031, frame extraction is performed on each historical multimedia resource in the historical multimedia resource set to obtain the first historical multimedia resource set.

[0114] Specifically, in this embodiment of the disclosure, each historical multimedia resource can be processed by frame extraction to obtain a first historical multimedia resource corresponding to each historical multimedia resource; frame extraction can be performed uniformly based on the time corresponding to the multimedia resource frame or according to the objects in the multimedia resource frame.

[0115] In step S2033, image data augmentation processing is performed on the first historical multimedia resource set to obtain the second historical multimedia resource set;

[0116] Specifically, in this embodiment of the disclosure, image data augmentation processing is used to process each first historical multimedia resource in the first historical multimedia resource set. Image data augmentation processing may include operations such as image cropping, image color adjustment, and image contrast adjustment; for example, random cropping and color adjustment are performed on some or all multimedia resource frames in the first historical multimedia resources to enhance the diversity of multimedia resource frames.

[0117] In step S2035, feature extraction processing is performed on the second historical multimedia resource set to obtain the historical multimedia resource features corresponding to each historical multimedia resource.

[0118] Specifically, in this embodiment of the disclosure, feature extraction processing can be performed on each historical multimedia resource in the second historical multimedia resource set to obtain historical multimedia resource features corresponding to each historical multimedia resource, thereby improving the accuracy of historical multimedia resource features.

[0119] In this embodiment, features of each historical multimedia resource in the historical multimedia resource set can be extracted to obtain a historical multimedia resource feature set. The multimedia resource features may include, but are not limited to, text features, image features, and audio features of each object in the multimedia resource; objects may include, but are not limited to, physical objects such as people, animals, and plants in the multimedia resource, and may also include virtual objects such as cartoon characters, virtual characters, and fluids; for example, for a multimedia resource whose object is a cat, the features of the cat in the multimedia resource can be extracted. When each second historical multimedia resource includes objects, the object features of each second historical multimedia resource can be extracted to obtain an object feature set, and the obtained object feature set is used as the historical multimedia resource feature set; if any second historical multimedia resource includes multiple objects, the object features corresponding to each of the multiple objects in the second historical multimedia resource can be extracted; and the object features corresponding to the multiple objects are fused to obtain the historical multimedia resource features of any second historical multimedia resource.

[0120] In this embodiment of the disclosure, the initial feature extraction model can be trained based on the sample multimedia resource set of the sample account to obtain the feature extraction model; and the historical multimedia resource set can be input into the feature extraction model for feature extraction processing to obtain the historical multimedia resource features corresponding to each historical multimedia resource.

[0121] In this embodiment, the initial feature extraction model can be a deep neural network model such as I3d (inflated 3D network) or ECO (Efficient Convolutional Network). I3d is an inflated 3D model, while ECO is an efficient convolutional neural network model. ECO uses only a single frame image within a temporal neighborhood and acquires its appearance features using 2D convolution. Furthermore, to obtain the contextual relationships between long-term image frames, simple fractional fusion is insufficient. Therefore, ECO performs end-to-end fusion by performing 3D convolution on the feature maps of each sampled frame. The initial feature fusion model can be a multi-layer fully connected layer, or it can be a transformer model. From an overall framework perspective, a transformer is essentially an encoding / decoding framework.

[0122] In step S205, feature fusion processing is performed on multiple historical multimedia resource features to obtain the account features of the account to be identified. The account features are used in subsequent processing of the associated multimedia resources of the account to be identified.

[0123] In this embodiment, the features of various historical multimedia resources in the historical multimedia resource feature set can be fused to obtain fused features, which are then used as the account features of the account to be identified. The resulting fused features can accurately characterize the account features of the account to be identified, thereby improving the accuracy of account feature identification.

[0124] In this embodiment of the disclosure, the initial feature fusion model can be trained by feature extraction based on the sample multimedia resource set of the sample account to obtain the feature fusion model; and the historical multimedia resource feature set can be input into the feature fusion model for feature fusion processing to obtain the account features of the account to be identified.

[0125] In this embodiment of the disclosure, the historical multimedia resource feature set can be fused according to the trained feature fusion model to obtain the account features of the account to be identified, thereby improving the efficiency of determining the account features of the account to be identified.

[0126] In this embodiment of the disclosure, such as Figure 3 As shown, the above-mentioned feature fusion processing of multiple historical multimedia resource features to obtain the account features of the account to be identified includes the following steps S2051-S2055:

[0127] In step S2051, the above-mentioned historical multimedia resource features are divided to obtain a preset number of historical resource feature subsets. Each of the above-mentioned historical multimedia resource features is extracted by the feature extraction model in the account feature generation model.

[0128] Specifically, in this embodiment, the feature extraction model can be a sub-model of the account feature generation model; the initial feature extraction model can be trained based on the sample multimedia resource set of the sample account to obtain the feature extraction model; the historical multimedia resource features of each historical multimedia resource are extracted through the feature extraction model to obtain multiple historical multimedia resource features; then, the multiple historical multimedia resource features are divided into a preset number of groups according to a preset grouping strategy; the preset grouping strategy can be determined based on the browsing time, type, or other attribute information corresponding to each historical multimedia resource; for example, the preset grouping strategy can be to group historical multimedia resource features with a browsing time difference less than a preset threshold into one group, or to group historical multimedia resource features with a type similarity greater than a similarity threshold into one group; in addition, multiple historical multimedia resource features can be randomly combined according to a preset number to obtain a subset of historical resource features in a preset number of groups; the preset number can be set according to the actual situation, for example, the preset number can be set to 2.

[0129] In step S2053, each set of the above-mentioned historical resource feature subsets is input into the feature fusion model in the above-mentioned account feature generation model. The feature fusion model is used to perform feature fusion processing on each set of the above-mentioned historical resource feature subsets to obtain the historical fusion features corresponding to each set of the above-mentioned historical resource feature subsets.

[0130] Specifically, in this embodiment, the feature fusion model can be a sub-model of the account feature generation model. During training, the initial feature extraction model and the initial feature fusion model can be trained simultaneously to obtain the feature extraction model and the feature fusion model. The initial feature fusion model can be trained by feature extraction based on the sample multimedia resource set of the sample account to obtain the feature fusion model. Then, the features in each set of historical resource feature subsets are input into the feature fusion model for feature fusion processing to obtain the historical fusion features corresponding to each set of historical resource feature subsets. For any set of historical resource feature subsets, each feature in that set of historical resource feature subsets can be input into the feature fusion model for fusion processing to obtain the historical fusion features corresponding to that set of historical resource feature subsets. Each set of historical resource feature subsets corresponds to one historical fusion feature, thereby obtaining a preset number of historical fusion features.

[0131] In step S2055, the obtained preset number of historical fusion features are determined as the account features of the account to be identified.

[0132] In this embodiment of the disclosure, a set of a preset number of historical fusion features can be used as the account features of the account to be identified; alternatively, a preset number of historical fusion features can be concatenated to obtain the account features of the account to be identified; or a weight can be set for each historical fusion feature, and the account features of the account to be identified can be determined according to the weight corresponding to each historical fusion feature; each historical fusion feature corresponds to a subset of historical resource features, and the preset number of historical fusion features is essentially a determination based on the historical multimedia resource set. Obviously, using a preset number of historical fusion features can accurately characterize the account features of the account to be identified.

[0133] In this embodiment of the disclosure, a preset number of historical fusion features can be used as the account features of the account to be identified. Using the preset number of historical fusion features of the account to be identified to characterize the account features of the account to be identified can improve the accuracy of the account features.

[0134] In this embodiment of the disclosure, such as Figure 4 As shown, the training method for the above account feature generation model includes the following steps S401-S4011:

[0135] In step S401, the sample multimedia resource set of the sample account is obtained; the sample multimedia resource set includes multiple sample multimedia resources.

[0136] In this embodiment of the disclosure, when the sample multimedia resource is audio information, audio features of the audio information can be extracted, and the resulting sample multimedia resource features are audio features; when the sample multimedia resource is video information, video features of the video information can be extracted, and the resulting sample multimedia resource features are video features. The sample multimedia resource features may include, but are not limited to, text features, image features, and audio features of various objects in the sample multimedia resource; objects may include, but are not limited to, physical objects such as people, animals, and plants in the sample multimedia resource, and may also include virtual objects such as cartoon characters, virtual characters, and fluids. During the feature extraction process, object features can be extracted from each sample multimedia resource, and the extracted object features are used as the sample multimedia resource features. The sample multimedia resource and the historical multimedia resource have the same resource type; when the sample multimedia resource is video information, the historical multimedia resource is also video information; when the sample multimedia resource is audio information, the historical multimedia resource is also audio information.

[0137] In step S403, the above-mentioned sample multimedia resource set is input into the initial account feature generation model, which includes an initial feature extraction model and an initial feature fusion model;

[0138] In this embodiment of the disclosure, each sample multimedia resource in the above-mentioned sample multimedia resource set can be input into the initial account feature generation model. The initial account feature generation model can include two sub-models: an initial feature extraction model and an initial feature fusion model. During the training process, the two sub-models can be jointly trained.

[0139] In step S405, feature extraction processing is performed on each sample multimedia resource using the above-mentioned initial feature extraction model to obtain a sample multimedia resource feature set; the sample multimedia resource feature set includes the above-mentioned preset number of sample resource feature subsets;

[0140] In this embodiment of the disclosure, the features in the sample multimedia resource feature set can be feature vectors of a preset dimension; the preset dimension can include, but is not limited to, 512 dimensions or 1024 dimensions. The number of sample resource feature subsets during the training process is the same as the number of historical resource feature subsets during the application process, both being preset number groups. For any sample multimedia resource, one or more sample multimedia resource features can be extracted; the set of sample multimedia resource features corresponding to each sample multimedia resource is determined as the sample multimedia resource feature set.

[0141] In this embodiment of the disclosure, when the sample multimedia resource set corresponding to the sample account includes N sample multimedia resources, and the preset quantity corresponding to each sample multimedia resource set is 2, the feature vectors of the two sets of multimedia resource sets of the N sample multimedia resources with dimension D are used as the sample multimedia resource feature set.

[0142] In this embodiment of the disclosure, feature extraction processing is performed on each sample multimedia resource using the above-mentioned initial feature extraction model to obtain a sample multimedia resource feature set, including the following steps S4051-S4055:

[0143] In step S4051, frame extraction is performed on each sample multimedia resource in the above sample multimedia resource set to obtain the first sample multimedia resource set.

[0144] Specifically, in this embodiment of the disclosure, frame extraction can be performed on each sample multimedia resource to obtain a first sample multimedia resource corresponding to each sample multimedia resource; frame extraction can be performed uniformly based on the time corresponding to the multimedia resource frame or according to the objects in the multimedia resource frame. The frame extraction method for sample multimedia resources is the same as the frame extraction method for historical multimedia resources.

[0145] In step S4053, the first sample multimedia resource set is subjected to image data augmentation processing to obtain the second sample multimedia resource set;

[0146] Specifically, in this embodiment of the disclosure, image data augmentation processing is used to process each of the first sample multimedia resources in the first sample multimedia resource set. Image data augmentation processing may include operations such as image cropping, image color adjustment, and image contrast adjustment; for example, random cropping and color adjustment are performed on some or all multimedia resource frames in the first sample multimedia resources to enhance the diversity of multimedia resource frames.

[0147] In step S4055, the second sample multimedia resource feature set is processed by the above-mentioned initial feature extraction model to obtain the sample multimedia resource feature set.

[0148] Specifically, in this embodiment of the disclosure, feature extraction processing can be performed based on the second sample multimedia resource set to obtain the sample multimedia resource feature set, thereby improving the accuracy of the sample multimedia resource feature set.

[0149] In this embodiment of the disclosure, features of each sample multimedia resource in the sample multimedia resource set can be extracted to obtain a sample multimedia resource feature set. The multimedia resource features may include the features of each object within the multimedia resource.

[0150] In step S407, the initial feature fusion model is used to perform feature fusion processing on each of the above sample resource feature subsets to obtain the sample fusion features corresponding to each of the above sample resource feature subsets. A preset number of sample fusion features corresponding to the above sample account are determined as the sample account features of the above sample account.

[0151] In this embodiment of the disclosure, each sample account may have one or more sample account features.

[0152] In this embodiment of the disclosure, the sample multimedia resources corresponding to each sample account can be grouped to obtain a subset of sample resource features. Each subset of sample resource features is then input into an initial feature fusion model for feature fusion processing to obtain the sample fusion features corresponding to each subset of sample resource features. A preset number of sample fusion features corresponding to the same sample account are determined as the sample account features of the same sample account, thereby increasing the number of sample account features and improving the accuracy of the feature fusion model.

[0153] In step S409, the target loss value is determined based on the above-mentioned sample account characteristics;

[0154] In this embodiment of the disclosure, there are at least two sample accounts. Before determining the target loss value based on the characteristics of the sample accounts, the above method further includes the following steps S4081-S4083:

[0155] In step S4081, a preset number of sample resource feature subsets corresponding to each of the above sample accounts are obtained; the sample multimedia resources in each sample resource feature subset are labeled with sample category tags;

[0156] In this embodiment of the disclosure, the method for determining the subset of sample resource features of the preset quantity group is the same as the method for determining the subset of historical resource features of the preset quantity group; the preset quantity corresponding to each sample account is the same, for example, it can be set to 2.

[0157] In this embodiment of the disclosure, sample category labels can be annotated based on the content of sample multimedia resources. Sample multimedia resources in the same subset of sample resource features can correspond to the same or different sample category labels. Sample category labels can be determined based on object features in the sample multimedia resources. Text features, image features, audio features, etc., of objects can be extracted to determine the object's category, thereby obtaining sample category labels. For example, sample category labels can be "cat," "dog," "grass," etc. Sample category labels can be annotated manually or by machine. When using machine annotation, it is necessary to identify the objects in the sample multimedia resources before annotation.

[0158] In step S4083, the sample multimedia resources in each sample resource feature subset are subjected to category identification processing to obtain the sample category results of each sample multimedia resource.

[0159] In this embodiment of the disclosure, determining the target loss value based on sample account characteristics includes: determining the target loss value based on sample account characteristics, sample category labels, and sample category results.

[0160] In this embodiment of the disclosure, the sample category labels and sample category results of each sample multimedia resource in the sample multimedia resource set can also be obtained, and the target loss value can be determined by combining the sample account features, sample category labels and sample category results corresponding to each sample multimedia resource.

[0161] In this embodiment of the disclosure, the above-mentioned target loss value can be constructed based on sample account features, sample category labels, and sample category results, thereby improving the training speed and accuracy of the model.

[0162] In this embodiment of the disclosure, such as Figure 5 As shown, based on the above sample account characteristics, the above sample category labels, and the above sample category results, the above target loss value is determined, including the following steps S501-S505:

[0163] In step S501, a first loss value is generated based on the sample category labels and sample category results corresponding to the aforementioned sample multimedia resources.

[0164] In this embodiment of the disclosure, a first loss value can be constructed based on the differences between the sample category labels and sample category results corresponding to the same sample multimedia resource. Specifically, the sample category labels and sample category results can be of the same type; for example, both sample category labels and sample category results can be text or images. The similarity between the sample category labels and sample category results can be calculated, and the first loss value can be determined based on the similarity. When both the sample category labels and sample category results are text, the first loss value can be determined by calculating text similarity. The aforementioned first loss value is used to characterize the category loss of the sample multimedia resource.

[0165] In step S503, a second loss value is generated based on the above-mentioned sample account characteristics;

[0166] In this embodiment of the disclosure, the second loss value represents the comparative loss of the sample fusion features corresponding to each of the two sample resource feature subsets.

[0167] In this embodiment of the disclosure, such as Figure 6 As shown, based on the above sample account characteristics, a second loss value is generated, including the following steps S5031-S5039:

[0168] In step S5031, any two sample resource feature subsets corresponding to any sample account are obtained to obtain two first sample resource feature subsets;

[0169] In this embodiment of the disclosure, the two first sample resource feature subsets are multimedia resources corresponding to the same sample account. The similarity between the two first sample resource feature subsets is greater than or equal to a first value, that is, the two first sample resource feature subsets are highly similar.

[0170] In step S5033, a subset of sample resource features corresponding to each of any two sample accounts is obtained, resulting in two second subsets of sample resource features;

[0171] In this embodiment of the disclosure, the two second sample resource feature subsets come from two different sample accounts, and the similarity between the two second sample resource feature subsets is less than the first value, that is, the difference between the two first sample resource feature subsets is large.

[0172] In step S5035, based on the sample fusion features corresponding to the two first sample resource feature subsets, the first similarity between the two first sample resource feature subsets is determined.

[0173] In this embodiment of the disclosure, the first similarity characterizes the degree of similarity between two subsets of first sample resource features.

[0174] In step S5037, based on the sample fusion features corresponding to each of the two second sample resource feature subsets, the second similarity between the two second sample resource feature subsets is determined;

[0175] In this embodiment of the disclosure, the second similarity characterizes the degree of similarity between two subsets of second sample resource features.

[0176] In step S5039, the second loss value is determined based on the first similarity and the second similarity.

[0177] In this embodiment of the disclosure, the first difference between the first similarity and the second similarity can be used as the second loss value, and during the training process, the first difference needs to be made smaller and smaller; alternatively, the second difference between the second similarity and the first similarity can be used as the second loss value, and during the training process, the second difference needs to be made larger and larger.

[0178] In this embodiment of the disclosure, the feature vectors of two sets of multimedia resources with dimension D from N sample multimedia resources are used to calculate the distance between each pair of feature vectors, thus forming N*N matrices X. X(i,j) calculates the similarity of the feature vectors of the first set of multimedia resources of the i-th sample and the second set of multimedia resources of the j-th sample. The KL divergence is then calculated using a diagonal matrix as the contrastive loss, resulting in a second loss value. KL divergence (Kullback-Leibler Divergence) is generally used to measure the "distance" between two probability distribution functions.

[0179] In this embodiment of the disclosure, during model training, the first similarity between two first sample resource feature subsets is increased, while the second similarity between two second sample resource feature subsets is decreased. That is, two sample multimedia resources (two sets of multimedia resources obtained from sampling) of the same sample are as similar as possible, while different samples are as dissimilar as possible. Multimedia resource sets created by the same author are similar, while multimedia resource sets created by different authors are different. Similarly, multimedia resource sets consumed by the same consumer are similar, while multimedia resource sets consumed by different consumers are different.

[0180] In this embodiment of the disclosure, a second loss value can be constructed based on the similarity between two first sample resource feature subsets corresponding to the same sample account and the similarity between two second sample resource feature subsets corresponding to two sample accounts; thereby ensuring that the first similarity of the two first sample resource feature subsets determined by the trained model is greater than or equal to a preset threshold, and the second similarity of the two second sample resource feature subsets is less than the preset threshold; thereby improving the difference in account features corresponding to different sample accounts and improving the accuracy of account features.

[0181] In step S505, the target loss value is determined based on the first loss value and the second loss value.

[0182] In this embodiment of the disclosure, the sum of the first loss value and the second loss value can be determined as the target loss value.

[0183] In this embodiment of the disclosure, a first loss value can be constructed based on the sample category label, and a second loss value can be constructed based on the sample fusion features, thereby accurately constructing the target loss value corresponding to the model.

[0184] In step S4011, the model parameters of the initial account feature generation model are adjusted based on the target loss value to obtain the account feature generation model.

[0185] In this embodiment of the disclosure, the initial feature extraction model and the initial feature fusion model can be jointly trained using a sample multimedia resource set to obtain the feature extraction model and the feature fusion model, thereby improving the training accuracy and training efficiency of the two models.

[0186] In this embodiment of the disclosure, such as Figure 7 As shown, the model parameters of the initial account feature generation model are adjusted based on the target loss value to obtain the account feature generation model, including the following steps S40111-S40115:

[0187] In step S40111, the model parameters corresponding to the initial feature extraction model and the initial feature fusion model are adjusted based on the target loss value until the preset training termination condition is met and the training ends.

[0188] In this embodiment, the initial feature extraction model can be a deep neural model such as I3d or ECO. I3d is a dilated 3D model; ECO is an efficient convolutional neural model that uses only a single frame image within a temporal neighborhood and acquires its appearance features using 2D convolution. Furthermore, to obtain the contextual relationships between long-term image frames, simple fractional fusion is insufficient. Therefore, ECO performs end-to-end fusion by performing 3D convolution on the feature maps of each sampled frame. The initial feature fusion model can be a multi-layer fully connected layer, or it can be a transformer model. From an overall framework perspective, a transformer is essentially an encoding / decoding framework.

[0189] In this embodiment, the preset training termination condition may include the target loss value being less than a preset value or the number of model iterations being greater than a preset number; wherein, both the preset value and the preset number of iterations can be set according to actual conditions, for example, the preset number of iterations can be set to 200,000. During model training, the target loss value can be optimized based on an optimization algorithm, and the learning rate can be set to a fixed value, for example, 0.01; the optimization algorithm may include, but is not limited to, gradient descent; label smoothing can also be used to smooth the labels; label smoothing is a regularization method in the field of machine learning, usually used for classification problems, with the aim of preventing the model from predicting labels too confidently during training and improving the problem of poor generalization ability.

[0190] In step S40113, the initial feature extraction model at the end of training is used as the aforementioned feature extraction model, and the initial feature fusion model at the end of training is used as the aforementioned feature fusion model.

[0191] In this embodiment of the disclosure, the first model parameters corresponding to the initial feature extraction model at the end of training and the second model parameters corresponding to the initial feature fusion model at the end of training can be obtained. Then, the initial feature extraction model corresponding to the first model parameters is determined as the above-mentioned feature extraction model; and the initial feature fusion model corresponding to the second model parameters is determined as the above-mentioned feature fusion model.

[0192] In step S40115, the account feature generation model is constructed based on the above feature extraction model and the above feature fusion model.

[0193] In this embodiment of the disclosure, the feature extraction model and the feature fusion model can be combined to construct the above-mentioned account feature generation model. The resulting account feature generation model can not only perform feature extraction processing, but also feature fusion processing.

[0194] In this embodiment of the disclosure, a target loss value can be constructed based on a sample multimedia resource set; the initial feature extraction model and the initial feature fusion model are trained based on the target loss value, and the initial feature extraction model at the end of training is used as the feature extraction model, and the initial feature fusion model at the end of training is used as the feature fusion model, thereby quickly determining the feature extraction model and the feature fusion model.

[0195] In this embodiment of the disclosure, the method further includes: obtaining the associated business of the account to be identified; and determining the business characteristics of the associated business based on the account characteristics.

[0196] In this embodiment of the disclosure, the associated services may include, but are not limited to, resource feature determination services for associated multimedia resources and associated information search services. The business features of associated services can be determined through account features; for example, in the search service, both account features and search keywords can be used as business features, and search information can be obtained based on these features before being pushed to the target account, thereby improving the accuracy of the search information. In the associated services, the business features of the associated services can be determined by combining the account features of the target account, thereby improving the accuracy of the business features.

[0197] In this embodiment of the disclosure, such as Figure 8 As shown, obtaining the associated business of the aforementioned account to be identified includes the following steps S207-S209:

[0198] In step S207, the multimedia resources to be identified for the aforementioned account to be identified are obtained;

[0199] In this embodiment of the disclosure, the account to be identified can be the creator, publisher, or viewer of the multimedia resource to be identified; the multimedia resource to be identified has the same resource type as the historical multimedia resource, for example, both are video information or audio information.

[0200] Based on the above account characteristics, the business characteristics of the aforementioned related businesses are determined, including:

[0201] In step S209, based on the account characteristics and the multimedia resources to be identified, the category information of the multimedia resources to be identified is determined.

[0202] In this embodiment of the disclosure, the multimedia resource to be identified can also be subjected to frame extraction processing, image data augmentation processing, and other operations to obtain the processed multimedia resource to be identified. Based on the processed multimedia resource to be identified and the aforementioned characteristics of the account to be identified, the category information of the multimedia resource to be identified is determined.

[0203] In this embodiment of the disclosure, the category information of the multimedia resource to be identified is determined based on the processed multimedia resource to be identified and the aforementioned characteristics of the account to be identified, thereby improving the accuracy of the category information.

[0204] In this embodiment of the disclosure, based on the aforementioned account characteristics and the aforementioned multimedia resources to be identified, the category information of the multimedia resources to be identified is determined, including:

[0205] The aforementioned account characteristics and the aforementioned multimedia resources to be identified are input into the category information recognition model for category information recognition processing to obtain the category information of the aforementioned multimedia resources to be identified.

[0206] In this embodiment of the disclosure, the category information recognition model is trained on a preset model based on a training multimedia resource set of a training account. The category information recognition model can identify the target category label corresponding to the multimedia resource to be identified and the probability value of the target category label; thereby determining the category information of the multimedia resource to be identified based on the target category label.

[0207] In this embodiment of the disclosure, a category information recognition model can be obtained through training, and the category label recognition processing of the multimedia resources to be identified and the account features to be identified can be performed according to the model, thereby quickly and accurately determining the category information of the multimedia resources to be identified.

[0208] In this embodiment of the disclosure, the training method for the above-mentioned category information recognition model includes the following steps S2081-S2087:

[0209] In step S2081, the training multimedia resource set of the training account is input into the feature extraction model above for feature extraction processing to obtain the training multimedia resource feature set; the training multimedia resources in the training multimedia resource set are labeled with training category information tags.

[0210] In this embodiment, the training account and the sample account can be the same or different. The training account is the author, publisher, or viewer of the training multimedia resource set. There can be multiple training accounts, and each training account corresponds to a training multimedia resource set. The training multimedia resource set and the sample multimedia resource set can be the same or different.

[0211] In this embodiment of the disclosure, inputting the above-mentioned training multimedia resource set into the above-mentioned feature extraction model for feature extraction processing to obtain the training multimedia resource feature set may include the following steps S20811-S20815:

[0212] In step S20811, frame extraction is performed on each training multimedia resource in the above training multimedia resource set to obtain the first training multimedia resource set.

[0213] Specifically, in this embodiment of the present disclosure, frame extraction can be performed on each training multimedia resource to obtain a first training multimedia resource corresponding to each training multimedia resource; frame extraction can be performed uniformly based on the time corresponding to the multimedia resource frame or according to the objects in the multimedia resource frame.

[0214] In step S20813, the first training multimedia resource set is subjected to image data augmentation processing to obtain the second training multimedia resource set.

[0215] Specifically, in this embodiment of the present disclosure, image data augmentation processing is used to process each of the first training multimedia resources in the first training multimedia resource set. Image data augmentation processing may include operations such as image cropping, image color adjustment, and image contrast adjustment; for example, random cropping and color adjustment are performed on some or all multimedia resource frames in the first training multimedia resources to enhance the diversity of multimedia resource frames.

[0216] In step S20815, the second training multimedia resource set is input into the feature extraction model for feature extraction processing to obtain the training multimedia resource feature set.

[0217] Specifically, in this embodiment of the present disclosure, feature extraction processing can be performed based on the second training multimedia resource set to obtain a training multimedia resource feature set, thereby improving the accuracy of the training multimedia resource feature set.

[0218] In this embodiment of the disclosure, features of each training multimedia resource in the aforementioned training multimedia resource set can be extracted to obtain a training multimedia resource feature set. The training multimedia resource features may include the features of each object within the training multimedia resource.

[0219] In this embodiment of the disclosure, the above-mentioned inputting the training multimedia resource set into the feature extraction model for feature extraction processing to obtain the training multimedia resource feature set includes the following steps S208101-S208105:

[0220] In step S208101, the above-mentioned training multimedia resource set is divided into a target number of training multimedia resource subsets;

[0221] In this embodiment, the target number may be the same as or different from the preset number mentioned above. The method for partitioning the training multimedia resource subset is similar to the method for partitioning the historical multimedia resource subset and the sample multimedia resource subset.

[0222] In step S208103, the above-mentioned target number of training multimedia resource subsets are respectively input into the above-mentioned feature extraction model for feature extraction processing to obtain the training multimedia resource sub-feature set corresponding to each training multimedia resource subset;

[0223] In this embodiment of the disclosure, the method for extracting the training multimedia resource sub-feature set of each training multimedia resource subset through a feature extraction model is similar to the method for extracting the historical multimedia resource sub-feature set.

[0224] In step S208105, the training multimedia resource sub-feature sets corresponding to each of the above target quantity group training multimedia resource subsets are used as the above training multimedia resource feature sets.

[0225] In step S2083, the above-mentioned training multimedia resource feature set is input into the above-mentioned feature fusion model for feature fusion processing to obtain the training account features of the above-mentioned training account;

[0226] In this embodiment of the disclosure, the above-mentioned training multimedia resource feature set is input into the above-mentioned feature fusion model for feature fusion processing to obtain the training account features of the above-mentioned training account, including steps S20831-S20833:

[0227] In step S20831, each set of training multimedia resource sub-features is input into the above feature fusion model for feature fusion processing to obtain the training fusion features corresponding to each set of training multimedia resource sub-features.

[0228] In this embodiment of the disclosure, the multimedia resource sub-features in each training multimedia resource sub-feature set can be fused to obtain the corresponding training fused features.

[0229] In step S20833, the training fusion features of the target number mentioned above are determined as the training account features of the training account mentioned above.

[0230] In this embodiment of the disclosure, the training account features of each training account can be quickly and accurately determined through the above-described feature fusion model, thereby facilitating the subsequent training of the category information recognition model.

[0231] In step S2085, the above-mentioned training multimedia resource set and the above-mentioned training account features are input into a preset model for category information recognition processing, and the category information recognition result is output.

[0232] In this embodiment of the disclosure, the aforementioned preset model can be a machine learning model, used to train and obtain a model for determining category information.

[0233] In step S2087, the preset model is trained based on the difference between the training category information labels and the training category information results to obtain the category information recognition model.

[0234] In this embodiment of the disclosure, a loss value can be determined based on the difference between the training category information labels and the training category information results; and the model parameters of the preset model can be adjusted based on the loss value to obtain a category information recognition model.

[0235] In this embodiment of the disclosure, a preset model can be trained based on the extracted training account features and the aforementioned training multimedia resource set to obtain a category information recognition model with high accuracy and improve the recognition speed of category information.

[0236] In a specific embodiment, such as Figure 9 As shown, Figure 9This is a schematic diagram of a video category determination system. The system includes a trained video feature extraction model, a video feature fusion model, and a video category determination model. First, it acquires N recently posted videos labeled with category information tags from a user and divides these N videos into two video sets. Based on the initial feature extraction model, it extracts features from each video set to obtain a first video feature set and a second video feature set. The first and second video feature sets are then input into the initial feature fusion model for feature fusion to obtain account features. The first video feature set, the second video feature set, and the account features are then input into the initial category determination model for category information recognition processing to obtain the category information recognition result. Based on the aforementioned account features, the category information tags corresponding to each video, and the category... The initial feature extraction model, initial feature fusion model, and initial category determination model are trained based on the information recognition results. The initial feature extraction model at the end of training is used as the video feature extraction model, the initial feature fusion model at the end of training is used as the video feature fusion model, and the initial category determination model at the end of training is used as the video category determination model. Thus, an account feature generation model can be obtained based on the video feature extraction model and the video feature fusion model. Based on the account feature generation model, the historical video set of the account to be identified can be processed to obtain the account features of the account to be identified. This facilitates the video category identification processing based on the account features of the account to be identified and the video to be identified by the video category determination model, thereby quickly and accurately determining the category information of the video to be identified.

[0237] This disclosure obtains a historical multimedia resource set of an account to be identified; the historical multimedia resource set includes multiple historical multimedia resources; feature extraction processing is performed on each historical multimedia resource in the historical multimedia resource set to obtain historical multimedia resource features corresponding to each historical multimedia resource; feature fusion processing is performed on multiple historical multimedia resource features to obtain account features of the account to be identified, and the account features are used in subsequent processing of the associated multimedia resources of the account to be identified. This disclosure extracts multiple historical multimedia resource features from the historical multimedia resource set of the account to be identified, and fuses these multiple historical multimedia resource features to obtain fused features. These fused features can accurately characterize account features; therefore, using these fused features as account features of the account to be identified improves the accuracy of account features; these account features are used in subsequent processing of the associated multimedia resources of the account to be identified, thereby improving the accuracy of determining the resource features of associated multimedia resources.

[0238] Figure 10 This is a block diagram illustrating an account feature generation apparatus according to an exemplary embodiment. (Refer to...) Figure 10 The device includes:

[0239] The historical resource acquisition module 1010 is configured to acquire the historical multimedia resource set of the account to be identified; the historical multimedia resource set includes multiple historical multimedia resources.

[0240] The historical feature extraction module 1020 is configured to perform feature extraction processing on each historical multimedia resource in the historical multimedia resource set to obtain the historical multimedia resource features corresponding to each historical multimedia resource.

[0241] The account feature determination module 1030 is configured to perform feature fusion processing on multiple historical multimedia resource features to obtain the account features of the account to be identified. The account features are used in subsequent processing of the associated multimedia resources of the account to be identified.

[0242] In some embodiments, the account feature determination module 1030 includes: a resource feature division submodule, configured to divide multiple historical multimedia resource features to obtain a preset number of historical resource feature subsets, wherein each historical multimedia resource feature is extracted by a feature extraction model in the account feature generation model; a historical fusion feature determination submodule, configured to input each set of historical resource feature subsets into a feature fusion model in the account feature generation model, and perform feature fusion processing on each set of historical resource feature subsets through the feature fusion model to obtain historical fusion features corresponding to each set of historical resource feature subsets; and an account feature determination submodule, configured to determine the preset number of historical fusion features obtained as account features of the account to be identified.

[0243] In some embodiments, the apparatus further includes: a training module for an account feature generation model, the training module comprising: a sample resource set acquisition submodule, configured to acquire a sample multimedia resource set of a sample account; the sample multimedia resource set includes multiple sample multimedia resources; a resource input submodule, configured to input the sample multimedia resource set into an initial account feature generation model, the initial account feature generation model including an initial feature extraction model and an initial feature fusion model; a sample feature set determination submodule, configured to perform feature extraction processing on each sample multimedia resource through the initial feature extraction model to obtain a sample multimedia resource feature set; the sample multimedia resource feature set includes a preset number of sample resource feature subsets; a feature fusion processing submodule, configured to perform feature fusion processing on each set of sample resource feature subsets through the initial feature fusion model to obtain sample fusion features corresponding to each set of sample resource feature subsets, and determine the preset number of sample fusion features corresponding to the sample account as the sample account features of the sample account; a target loss value determination submodule, configured to determine a target loss value based on the sample account features; and a model generation submodule, configured to adjust the model parameters of the initial account feature generation model based on the target loss value to obtain an account feature generation model.

[0244] In some embodiments, there are at least two sample accounts, and the apparatus further includes: a sample resource feature subset acquisition module, configured to acquire a preset number of sample resource feature subsets corresponding to each sample account; each sample multimedia resource in the sample resource feature subset is labeled with a sample category label; a sample category result determination module, configured to perform category recognition processing on the sample multimedia resources in each sample resource feature subset to obtain the sample category result of each sample multimedia resource; in this embodiment, the target loss value determination unit includes: a target loss value determination subunit, configured to determine the target loss value based on the sample account features, sample category labels, and sample category results.

[0245] In some embodiments, the target loss value determination subunit includes: a first loss value determination subunit configured to generate a first loss value based on the sample category label corresponding to the sample multimedia resource and the sample category result; a second loss value determination subunit configured to generate a second loss value based on the sample account features; and a loss value determination subunit configured to determine a target loss value based on the first loss value and the second loss value.

[0246] In some embodiments, the second loss value determination subunit includes: a first subset determination subunit configured to perform operations to obtain any two sample resource feature subsets corresponding to any sample account, resulting in two first sample resource feature subsets; a second subset determination subunit configured to perform operations to obtain one sample resource feature subset corresponding to each of any two sample accounts, resulting in two second sample resource feature subsets; a first similarity determination subunit configured to perform operations to determine a first similarity between the two first sample resource feature subsets based on the sample fusion features corresponding to each of the two first sample resource feature subsets; a second similarity determination subunit configured to perform operations to determine a second similarity between the two second sample resource feature subsets based on the sample fusion features corresponding to each of the two second sample resource feature subsets; and a second loss value determination subunit configured to perform operations to determine a second loss value based on the first similarity and the second similarity.

[0247] In some embodiments, the apparatus further includes: a resource acquisition module configured to acquire multimedia resources to be identified for an account to be identified; and a category information identification module configured to determine category information of the multimedia resources to be identified based on account characteristics and the multimedia resources to be identified.

[0248] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0249] In one exemplary embodiment, an electronic device is also provided, including a processor; a memory for storing processor-executable instructions; wherein, when the processor is configured to execute the instructions stored in the memory, it implements the account feature generation method provided in any of the above embodiments.

[0250] The electronic device can be a terminal, a server, or a similar computing device. Taking a server as an example... Figure 11 This is a block diagram illustrating an electronic device according to an exemplary embodiment, such as... Figure 11 As shown, the server 1100 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1110 (CPUs 1110 may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), a memory 1130 for storing data, and one or more storage media 1120 (e.g., one or more mass storage devices) for storing application programs 1123 or data 1122. The memory 1130 and storage media 1120 may be temporary or persistent storage. The program stored in the storage media 1120 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the CPU 1110 may be configured to communicate with the storage media 1120 and execute the series of instruction operations stored in the storage media 1120 on the server 1100. Server 1100 may also include one or more power supplies 1160, one or more wired or wireless model interfaces 1150, one or more input / output interfaces 1140, and / or one or more operating systems 1121, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0251] The input / output interface 1140 can be used to receive or transmit data via a model. Specific examples of the model described above may include a wireless model provided by the communication vendor of server 1100. In one example, the input / output interface 1140 includes a model adapter (Network Interface Controller, NIC), which can connect to other model devices via a base station to communicate with the Internet. In one example, the input / output interface 1140 may be a radio frequency (RF) module for wireless communication with the Internet.

[0252] Those skilled in the art will understand that Figure 11The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 1100 may also include... Figure 11 The more or fewer components shown, or having the same Figure 11 The different configurations shown.

[0253] In one exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 1130 including instructions, which can be executed by a processor 1110 of device 1100 to perform the above-described method. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0254] In one exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the method for generating account features provided in any of the above embodiments.

[0255] This disclosure obtains a historical multimedia resource set of an account to be identified; the historical multimedia resource set includes multiple historical multimedia resources; feature extraction processing is performed on each historical multimedia resource in the historical multimedia resource set to obtain historical multimedia resource features corresponding to each historical multimedia resource; feature fusion processing is performed on multiple historical multimedia resource features to obtain account features of the account to be identified, and the account features are used in subsequent processing of the associated multimedia resources of the account to be identified. This disclosure extracts multiple historical multimedia resource features from the historical multimedia resource set of the account to be identified, and fuses these multiple historical multimedia resource features to obtain fused features, which can accurately characterize account features; therefore, using these fused features as account features of the account to be identified improves the accuracy of account features; these account features are used in subsequent processing of the associated multimedia resources of the account to be identified, thereby improving the accuracy of determining the resource features of associated multimedia resources.

[0256] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0257] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0258] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for generating account features, characterized in that, include: Obtain the historical multimedia resource set of the account to be identified; the historical multimedia resource set includes multiple historical multimedia resources. Feature extraction processing is performed on each historical multimedia resource in the historical multimedia resource set to obtain the historical multimedia resource features corresponding to each historical multimedia resource. Multiple historical multimedia resource features are input into an account feature generation model for feature fusion processing to obtain the account features of the account to be identified. The account features are used in subsequent processing of the associated multimedia resources of the account to be identified. The training method for the account feature generation model includes: The sample multimedia resource set of the sample account is input into the initial account feature generation model to obtain the sample account features and the sample category result of each sample multimedia resource; the sample category result is used to generate the first loss value. Obtain any two sample resource feature subsets corresponding to any sample account to obtain two first sample resource feature subsets; Obtain a subset of sample resource features corresponding to each of any two sample accounts to obtain two second subsets of sample resource features; Based on the sample account characteristics, a first similarity is determined between the two first sample resource feature subsets and a second similarity is determined between the two second sample resource feature subsets; the first similarity and the second similarity are used to determine the second loss value; The initial account feature generation model is trained based on the first loss value and the second loss value to obtain the account feature generation model.

2. The method according to claim 1, characterized in that, The step of inputting multiple historical multimedia resource features into an account feature generation model for feature fusion processing to obtain the account features of the account to be identified includes: The historical multimedia resource features are divided into a preset number of subsets of historical resource features, and each of the historical multimedia resource features is extracted by the feature extraction model in the account feature generation model. Each set of historical resource feature subsets is input into the feature fusion model in the account feature generation model. The feature fusion model performs feature fusion processing on each set of historical resource feature subsets to obtain historical fusion features corresponding to each set of historical resource feature subsets. The obtained historical fusion features are determined as the account features of the account to be identified.

3. The method according to claim 2, characterized in that, The method further includes: Obtain the sample multimedia resource set of the sample account; the sample multimedia resource set includes multiple sample multimedia resources; The initial account feature generation model includes an initial feature extraction model and an initial feature fusion model; the initial feature extraction model is used to perform feature extraction processing on each of the sample multimedia resources to obtain a sample multimedia resource feature set; the sample multimedia resource feature set includes a subset of sample resource features of the preset number of groups; The initial feature fusion model is used to perform feature fusion processing on each set of sample resource feature subsets to obtain sample fusion features corresponding to each set of sample resource feature subsets. A preset number of sample fusion features corresponding to the sample account are determined as the sample account features of the sample account. Based on the characteristics of the sample accounts, determine the target loss value; The model parameters of the initial account feature generation model are adjusted based on the target loss value to obtain the account feature generation model.

4. The method according to claim 3, characterized in that, The sample accounts are at least two, and the method further includes: Obtain a preset number of sample resource feature subsets corresponding to each sample account; each sample multimedia resource subset is labeled with a sample category tag. The sample multimedia resources in each of the sample resource feature subsets are subjected to category identification processing to obtain the sample category results of each sample multimedia resource; The step of determining the target loss value based on the characteristics of the sample accounts includes: The target loss value is determined based on the sample account characteristics, the sample category labels, and the sample category results.

5. The method according to claim 4, characterized in that, The step of determining the target loss value based on the sample account features, the sample category labels, and the sample category results includes: A first loss value is generated based on the sample category labels and sample category results corresponding to the sample multimedia resources. Based on the characteristics of the sample accounts, a second loss value is generated; The target loss value is determined based on the first loss value and the second loss value.

6. The method according to claim 5, characterized in that, The step of generating a second loss value based on the sample account features includes: Based on the sample fusion features corresponding to each of the two first sample resource feature subsets, the first similarity between the two first sample resource feature subsets is determined. Based on the sample fusion features corresponding to each of the two second sample resource feature subsets, a second similarity between the two second sample resource feature subsets is determined. The second loss value is determined based on the first similarity and the second similarity.

7. The method according to claim 1, characterized in that, The method further includes: Obtain the multimedia resources to be identified for the account to be identified; Based on the account characteristics and the multimedia resource to be identified, the category information of the multimedia resource to be identified is determined.

8. An apparatus for generating account features, characterized in that, include: The historical resource acquisition module is configured to acquire the historical multimedia resource set of the account to be identified; the historical multimedia resource set includes multiple historical multimedia resources. The historical feature extraction module is configured to perform feature extraction processing on each historical multimedia resource in the historical multimedia resource set to obtain historical multimedia resource features corresponding to each historical multimedia resource. The account feature determination module is configured to perform feature fusion processing by inputting multiple historical multimedia resource features into the account feature generation model to obtain the account features of the account to be identified. The account features are used in subsequent processing of the associated multimedia resources of the account to be identified. The training method for the account feature generation model includes: The sample multimedia resource set of the sample account is input into the initial account feature generation model to obtain the sample account features and the sample category result of each sample multimedia resource; the sample category result is used to generate the first loss value. Obtain any two sample resource feature subsets corresponding to any sample account to obtain two first sample resource feature subsets; Obtain a subset of sample resource features corresponding to each of any two sample accounts to obtain two second subsets of sample resource features; Based on the sample account characteristics, a first similarity is determined between the two first sample resource feature subsets and a second similarity is determined between the two second sample resource feature subsets; the first similarity and the second similarity are used to determine the second loss value; The initial account feature generation model is trained based on the first loss value and the second loss value to obtain the account feature generation model.

9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the method for generating account features as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of an electronic device, the electronic device is able to perform the method for generating account features as described in any one of claims 1-7.