Network training and feature representation methods, apparatuses, media, and devices

By combining feature extraction networks and feature interaction networks, the problem of insufficient feature fusion of multimedia elements is solved, more effective feature representation is achieved, and the feature representation capability of multimedia resources is improved.

CN116821674BActive Publication Date: 2026-07-21BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2023-05-22
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, the fusion of multimedia element features is insufficient, resulting in poor feature representation and failing to meet the needs of application tasks.

Method used

A feature extraction network is used to extract element features and perform feature interaction processing on each media element of the multimedia resource sample. By combining the element feature extraction network and the feature interaction network, the feature representation of the multimedia resource is achieved.

Benefits of technology

It enables in-depth mining and full integration of the features of different media elements in multimedia resources, improves the effectiveness of feature representation, and meets the business needs of downstream applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116821674B_ABST
    Figure CN116821674B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a network training and feature representation method, device, medium and equipment, and relates to the technical field of machine learning. The training method comprises: obtaining a plurality of media elements of each multimedia resource sample; inputting each media element of each multimedia resource sample into an element feature extraction network corresponding to the media element to obtain a plurality of element feature information of each multimedia resource sample; inputting the plurality of element feature information of each multimedia resource sample into a feature interaction network after splicing to obtain resource feature information of each multimedia resource sample; performing target behavior prediction processing corresponding to each media element according to the resource feature information of each multimedia resource sample to obtain a plurality of sample prediction information; and training the element feature extraction network based on the plurality of sample prediction information to obtain a trained element feature extraction network. The technical solution provided by the embodiment of the present disclosure can better represent the features of multimedia resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of machine learning technology, and in particular to methods, apparatus, media and devices for network training and feature representation. Background Technology

[0002] To improve the capabilities of machine learning algorithms and meet the needs of application tasks, it is necessary to find better representations of the original data, which is also the basic idea behind feature learning. For example, the quality of feature representation of video content directly affects the performance of recommendation and search services. Feature representation of multimedia resources such as videos involves the fusion of features from multiple media elements, including visual features and textual features.

[0003] In related technologies, feature fusion schemes for multiple media elements generally adopt a multi-stream mode, that is, feature extraction and mining are performed on the content of various media elements separately before feature splicing, but feature mining and fusion between different media elements are not sufficiently performed. Summary of the Invention

[0004] This disclosure provides methods, apparatus, media, and devices for network training and feature representation, to better extract features from different media elements in multimedia resources and to better mine features that are integrated among different media elements. The technical solution of this disclosure is as follows:

[0005] According to a first aspect of the present disclosure, a method for training a feature extraction network is provided, comprising:

[0006] Retrieve multiple media elements from each multimedia resource sample in a multi-multimedia resource sample;

[0007] Each media element of each multimedia resource sample is input into the element feature extraction network corresponding to each media element, and the element feature is extracted to obtain multiple element feature information of each multimedia resource sample.

[0008] The feature information of multiple elements of each multimedia resource sample is concatenated and input into the feature interaction network for feature interaction processing to obtain the resource feature information of each multimedia resource sample.

[0009] Based on the resource feature information of each multimedia resource sample, target behavior prediction processing corresponding to each media element is performed to obtain multiple sample prediction information;

[0010] Based on the prediction information of the multiple samples, the element feature extraction network corresponding to each media element is trained to obtain the trained element feature extraction network corresponding to each media element.

[0011] Optionally, obtaining multiple media elements of each multimedia resource sample from multiple multimedia resource samples includes:

[0012] Multiple frames of images are extracted from the target multimedia resource sample to obtain the target image of the target multimedia resource sample;

[0013] The target multimedia resource sample is subjected to speech recognition processing or text detection processing to obtain the target text of the target multimedia resource sample.

[0014] The target image and the target text are used as multiple media elements of the target multimedia resource sample; the target multimedia resource sample is any one of the multiple multimedia resource samples.

[0015] Optionally, the step of inputting each media element of each multimedia resource sample into an element feature network corresponding to each media element for element feature extraction processing to obtain multiple element feature information of each multimedia resource sample includes:

[0016] The first media element of each multimedia resource sample is input into the image feature extraction network corresponding to the first media element, and image feature extraction processing is performed to obtain the image feature information of each multimedia resource sample, wherein the first media element is an image;

[0017] The second media element of each multimedia resource sample is input into the text feature extraction network corresponding to the second media element, and the text feature extraction process is performed to obtain the text feature information of each multimedia resource sample, wherein the second media element is text;

[0018] The image feature information and text feature information of each multimedia resource sample constitute multiple element feature information of each multimedia resource sample.

[0019] Optionally, the step of concatenating multiple element feature information of each multimedia resource sample and inputting it into a feature interaction network for feature interaction processing to obtain resource feature information of each multimedia resource sample includes:

[0020] The image feature information and text feature information of each multimedia resource sample are concatenated to obtain the target feature information of each multimedia resource sample.

[0021] The target feature information of each multimedia resource sample is input into the feature interaction network for feature interaction processing to obtain the resource feature information of each multimedia resource sample; the feature interaction network adopts a multi-layer encoder and decoder network structure.

[0022] Optionally, the step of performing target behavior prediction processing corresponding to each media element based on the resource feature information of each multimedia resource sample to obtain multiple sample prediction information includes:

[0023] The resource feature information of each multimedia resource sample is input into the first target network corresponding to the first media element. The first target network performs topic classification prediction processing on each multimedia resource sample based on the resource feature information of each multimedia resource sample to obtain the first sample prediction information of each multimedia resource sample; the first media element is an image.

[0024] Optionally, the step of training the element feature extraction network corresponding to each media element based on the multiple sample prediction information to obtain the trained element feature extraction network corresponding to each media element includes:

[0025] Obtain first supervision information for each of the multimedia resource samples, wherein the first supervision information characterizes the topic analogy of each of the multimedia resource samples;

[0026] Based on the first sample prediction information and the first supervision information of each multimedia resource sample, the first loss value is calculated.

[0027] Based on the first loss value, the image feature extraction network corresponding to the first media element is trained to obtain the trained image feature extraction network.

[0028] Optionally, the step of performing target behavior prediction processing corresponding to each media element based on the resource feature information of each multimedia resource sample to obtain multiple sample prediction information includes:

[0029] Obtain the search text corresponding to each of the multimedia resource samples;

[0030] The search text corresponding to each of the multimedia resource samples is input into the second target network for text feature extraction processing to obtain the search text feature information of each of the multimedia resource samples.

[0031] The search text feature information and resource feature information of each multimedia resource sample are cross-calculated for similarity to obtain the feature similarity information of multiple target samples; the target sample corresponds to the combination of the search text feature information of the first multimedia resource sample and the resource feature information of the second multimedia resource sample; the first multimedia resource sample and the second multimedia resource sample are both any of the multimedia resource samples.

[0032] The feature similarity information of each target sample is input into a third target network corresponding to the second media element. The third target network performs click prediction processing on each target sample based on the target feature information of each target sample to obtain the second sample prediction information of each target sample. The second sample prediction information represents the probability that the second multimedia resource sample is clicked when the current search text is consistent with the search text corresponding to the first multimedia resource sample. It also represents the probability that the first multimedia resource sample and the second multimedia resource sample are the same multimedia resource sample. The second media element is text.

[0033] Optionally, the step of training the element feature extraction network corresponding to each media element based on the multiple sample prediction information to obtain the trained element feature extraction network corresponding to each media element includes:

[0034] Determine second supervision information for each of the target samples, wherein the second supervision information indicates whether the first multimedia resource sample associated with the target sample and the second multimedia resource sample are the same multimedia resource sample;

[0035] Based on the second sample prediction information and the second supervision information of each target sample, the second loss value is calculated.

[0036] Based on the second loss value, the text feature extraction network corresponding to the second media element is trained to obtain the trained text feature extraction network.

[0037] Optionally, the method further includes:

[0038] Based on the prediction information of multiple samples of each multimedia resource sample, the feature interaction network is trained to obtain the trained feature interaction network.

[0039] According to a second aspect of the present disclosure, a method for representing features of multimedia resources is provided, the method comprising:

[0040] To acquire multiple media elements of a multimedia resource;

[0041] Each of the multiple media elements is input into the element feature extraction network corresponding to each media element to perform element feature extraction processing, thereby obtaining multiple element feature information of the multimedia resource; the element feature network is obtained by the training method of the feature extraction network described in the first aspect;

[0042] The feature information of the multiple elements is input into the feature interaction network for feature interaction processing to obtain the resource feature information of the multimedia resource.

[0043] According to a third aspect of the present disclosure, a training apparatus for a feature extraction network is provided, the apparatus comprising:

[0044] The first acquisition module is configured to acquire multiple media elements of each multimedia resource sample from multiple multimedia resource samples;

[0045] The first element feature extraction module is configured to input each media element of each multimedia resource sample into the element feature extraction network corresponding to each media element, perform element feature extraction processing, and obtain multiple element feature information of each multimedia resource sample.

[0046] The first element feature interaction module is configured to concatenate multiple element feature information of each multimedia resource sample and input them into the feature interaction network for feature interaction processing to obtain the resource feature information of each multimedia resource sample.

[0047] The sample prediction module is configured to perform target behavior prediction processing corresponding to each media element based on the resource feature information of each multimedia resource sample, and obtain multiple sample prediction information.

[0048] The training module is configured to train the element feature extraction network corresponding to each media element based on prediction information from multiple samples, thereby obtaining the trained element feature extraction network corresponding to each media element.

[0049] According to a fourth aspect of the present disclosure, a feature representation apparatus for multimedia resources is provided, the apparatus comprising:

[0050] The second acquisition module is configured to acquire multiple media elements of multimedia resources;

[0051] The second element feature extraction module is configured to input each of the plurality of media elements into the element feature extraction network corresponding to each media element, perform element feature extraction processing, and obtain multiple element feature information of the multimedia resource; the element feature network is obtained by the training method of the feature extraction network described in any one of the first aspects;

[0052] The second element feature interaction module is configured to input the multiple element feature information into the feature interaction network, perform feature interaction processing, and obtain the resource feature information of the multimedia resource.

[0053] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the training method of the feature extraction network described in any one of the first aspects of the present disclosure or the feature representation method of multimedia resources described in any one of the second aspects.

[0054] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform a training method for a feature extraction network as described in any one of the first aspects of the present disclosure or a feature representation method for multimedia resources as described in any one of the second aspects.

[0055] According to a seventh aspect of the present disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, implement the training method for a feature extraction network as described in any one of the first aspects of the present disclosure or the feature representation method for multimedia resources as described in any one of the second aspects.

[0056] The technical solutions provided in this disclosure offer at least the following beneficial effects:

[0057] The technical solution provided in this disclosure, during the model training phase, inputs each media element of each multimedia resource sample into the element feature extraction network corresponding to the media element for element feature extraction processing, obtaining multiple element feature information of each multimedia resource sample, each element feature information corresponding to a media element; then, the multiple element feature information of each multimedia resource sample is concatenated and input into the same feature interaction network for feature interaction processing, obtaining resource feature information of each multimedia resource sample, which is a deep mining of the features of different media elements of a multimedia resource sample; using the resource feature information of each multimedia resource sample, target behavior prediction processing corresponding to each media element is performed to obtain multiple sample prediction information, and then the multiple sample prediction information of multiple multimedia resource samples can be used to train the element feature extraction network corresponding to each media element, thereby obtaining the trained multiple element feature extraction network. By utilizing the technical solutions provided in this disclosure, feature extraction networks are applied to various media elements in multimedia resources, and target processing tasks corresponding to each media element are used to train the feature extraction networks. This allows for more effective extraction of features from different media elements. Furthermore, by simultaneously inputting the element feature information corresponding to various media elements in multimedia resources into the same feature interaction network for feature interaction, the correlation between the element feature information of different media elements can be mined, thus more fully integrating the element feature information of multiple media elements. More effective feature extraction and more complete feature fusion of different media elements in multimedia resources result in more effective feature representations of multimedia resources, thereby meeting the ever-increasing business needs of downstream applications.

[0058] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0059] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0060] Figure 1 This is a flowchart illustrating a training method for a feature extraction network according to an exemplary embodiment;

[0061] Figure 2 This is a flowchart illustrating the acquisition of multiple media elements according to an exemplary embodiment;

[0062] Figure 3 This is a flowchart illustrating an element feature extraction method according to an exemplary embodiment;

[0063] Figure 4 This is a flowchart illustrating an element feature fusion interaction according to an exemplary embodiment;

[0064] Figure 5 This is a flowchart illustrating a training method for an image feature extraction network according to an exemplary embodiment;

[0065] Figure 6 This is a flowchart illustrating a training method for a text feature extraction network according to an exemplary embodiment;

[0066] Figure 7 This is a flowchart illustrating a method for representing the features of a multimedia resource according to an exemplary embodiment;

[0067] Figure 8 This is a schematic diagram of a network architecture according to an exemplary embodiment;

[0068] Figure 9 This is a block diagram illustrating a training apparatus for a feature extraction network according to an exemplary embodiment;

[0069] Figure 10 This is a block diagram illustrating a feature representation device for a multimedia resource according to an exemplary embodiment;

[0070] Figure 11 This is a block diagram of an electronic device for implementing a training method for a feature extraction network or a feature representation method for multimedia resources, according to an exemplary embodiment.

[0071] Figure 12 This is a block diagram of an electronic device for implementing a training method for a feature extraction network or a feature representation method for multimedia resources, according to an exemplary embodiment. Detailed Implementation

[0072] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0073] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0074] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0075] Figure 1 This is a flowchart illustrating a training method for a feature extraction network according to an exemplary embodiment. Figure 1 As shown, the method may include the following steps:

[0076] In step S110, multiple media elements of each multimedia resource sample in multiple multimedia resource samples are obtained.

[0077] In this embodiment of the disclosure, the multimedia resource sample can be an object that integrates multiple media elements, such as short videos and live data, used for training. The multiple media elements contained in each multimedia resource sample can include at least two of the following: images, text, audio and animation, data, links, etc.

[0078] In one embodiment of this disclosure, such as Figure 2 As shown, obtaining multiple media elements from each multimedia resource sample may include the following steps:

[0079] In step S111, multiple frames of images are extracted from the target multimedia resource sample to obtain the target image of the target multimedia resource sample.

[0080] Taking the target multimedia resource sample as any one of multiple multimedia resource samples as an example, when the target multimedia resource contains media elements such as images, multiple frames of images can be extracted from the target multimedia resource sample as target images for which image features are to be extracted.

[0081] Taking video as an example of a target multimedia resource, the cover image and N uniformly sampled images (N≥1, e.g., N=4) from the video can be extracted as target images of the target multimedia resource. Using multiple images as target images for extracting image features allows for a more comprehensive extraction of the video's visual features.

[0082] In step S112, speech recognition processing or text detection processing is performed on the target multimedia resource sample to obtain the target text of the target multimedia resource sample.

[0083] Taking video as an example of a target multimedia resource, speech recognition can be used to convert the speech in the video into text, or text detection can be used to identify subtitles and identifying text in video images.

[0084] In step S113, the target image and target text are used as multiple media elements of the target multimedia resource sample; the target multimedia resource sample is any one of the multiple multimedia resource samples.

[0085] In the above embodiments, taking images and text as examples, different acquisition methods are used for different types of media elements, which can effectively and accurately obtain the content corresponding to different media elements of multimedia resource samples, thereby helping to extract the features corresponding to different media elements.

[0086] In step S120, each media element of each multimedia resource sample is input into the element feature extraction network corresponding to each media element to perform element feature extraction processing, thereby obtaining multiple element feature information of each multimedia resource sample.

[0087] In this embodiment, different categories of media elements correspond to different element feature extraction networks, and these networks may differ in network type and structure. By utilizing their respective element feature extraction networks to perform feature extraction processing on the corresponding media elements, the features of each media element can be extracted more effectively and accurately.

[0088] In this embodiment of the disclosure, by performing element feature extraction processing on a multimedia resource sample, multiple element feature information of the multimedia resource sample can be obtained, and each element feature information corresponds to a media element.

[0089] In one embodiment of this disclosure, the element feature extraction network may include an image feature extraction network and a text feature extraction network. Specifically, as... Figure 3 As shown, step S120 may include the following:

[0090] In step S121, the first media element of each multimedia resource sample is input into the image feature extraction network corresponding to the first media element, and image feature extraction processing is performed to obtain the image feature information of each multimedia resource sample. The first media element is an image.

[0091] For example, ResNet-50 (a residual neural network) can be used as the image feature extraction network. The images of each multimedia resource sample are input into ResNet-50 for image feature extraction processing. The feature map of the last layer output by ResNet-50 is divided into block sequences as the image feature information of each multimedia resource sample.

[0092] In step S122, the second media element of each multimedia resource sample is input into the text feature extraction network corresponding to the second media element to perform text feature extraction processing, thereby obtaining the text feature information of each multimedia resource sample, where the second media element is text.

[0093] For example, BERT (a semantic representation model) can be used as a text feature extraction network. The text of each multimedia resource sample is input into BERT for text feature extraction processing, including text segmentation and embedding, to obtain the text feature information of each multimedia resource sample.

[0094] In step S123, the image feature information and text feature information of each multimedia resource sample constitute multiple element feature information of each multimedia resource sample.

[0095] In the above embodiments, taking images and text as examples, different types of media elements are input into their respective element feature extraction networks for corresponding element feature extraction, which can effectively and accurately extract the features of each media element in each multimedia resource sample.

[0096] In step S130, the feature information of multiple elements of each multimedia resource sample is concatenated and input into the feature interaction network for feature interaction processing to obtain the resource feature information of each multimedia resource sample.

[0097] In this embodiment, a single-stream feature fusion interaction method is adopted. This involves first concatenating the element feature information corresponding to each media element of each multimedia resource sample and then inputting it into the same feature interaction network for feature interaction processing. This yields the resource feature information of each multimedia resource sample. The feature interaction processing can perform inner product, Hadamard product, or bilinear cross processing on the element feature information corresponding to different media elements to examine the correlation between the element feature information corresponding to different media elements. This correlation is then used to update the multiple element feature information of the same multimedia resource sample, resulting in the resource feature information of that multimedia resource sample. Compared to a multi-stream fusion method that inputs the element feature information corresponding to each media element into its respective feature interaction network, mines the element feature information corresponding to each media element separately, and then fuses them, the single-stream feature fusion interaction method provided in this embodiment can perform feature interaction and mining on the element feature information corresponding to different media elements. This allows for a more thorough fusion of element feature information among multiple media elements, resulting in more accurate and effective resource feature information as the feature representation of the multimedia resource sample.

[0098] In one embodiment of this disclosure, such as Figure 4 As shown, step S130 may include the following steps:

[0099] In step S131, the image feature information and text feature information of each multimedia resource sample are concatenated to obtain the target feature information of each multimedia resource sample.

[0100] For example, the multiple element feature information of each multimedia resource sample includes image feature information and text feature information of each multimedia resource sample. The image feature information and text feature information of each multimedia resource sample are then processed through...<Sep token> Piece them together, and add them at the beginning and end.<Cls token> and<Eos token> As a marker, the data input to the feature interaction network can be obtained, which is the target feature information of each multimedia resource sample.

[0101] In step S132, the target feature information of each multimedia resource sample is input into the feature interaction network for feature interaction processing to obtain the resource feature information of each multimedia resource sample; the feature interaction network adopts a multi-layer encoder and decoder network structure.

[0102] For example, a Transformer model based on a multi-layered encoder-decoder and bidirectional attention mechanism is used as the feature interaction network. Target feature information, containing image and text feature information of each multimedia resource sample, is input into a multi-layered Transformer model for feature interaction processing. This allows for sufficient interaction between the image and text feature information of each multimedia resource sample, thereby obtaining the resource feature representation of each multimedia resource sample, which is the output of the Transformer model.<Cls token> The corresponding results.

[0103] In the above embodiments, using a multi-layer Transformer model as the feature interaction network can more fully mine and fuse the feature information between multiple media elements, thereby obtaining a more effective feature representation of multimedia resource samples. Furthermore, using a multi-layer Transformer model to fuse the feature information of multiple media elements, compared to using a multi-stream interaction fusion method, can also reduce the training load on the feature interaction network and improve training efficiency.

[0104] In step S140, based on the resource feature information of each multimedia resource sample, target behavior prediction processing corresponding to each media element is performed to obtain multiple sample prediction information.

[0105] In this embodiment of the disclosure, the resource feature information that characterizes the overall features of each multimedia resource sample is subjected to target behavior prediction processing corresponding to each media element to obtain multiple sample prediction information. Each sample prediction information also corresponds to a media element, and each sample prediction information can be used to train an element feature extraction network corresponding to a media element.

[0106] In this embodiment of the disclosure, the target behavior prediction processing can be behavior prediction for multimedia resource samples, such as topic segmentation, search, etc., so that supervision information can be constructed using historical behavior results to train the element feature extraction network.

[0107] In one embodiment of this disclosure, for media elements of the image category, such as Figure 5 As shown, step S140 may include:

[0108] In step S141, the resource feature information of each multimedia resource sample is input into the first target network corresponding to the first media element. The first target network performs topic classification prediction processing on each multimedia resource sample based on the resource feature information of each multimedia resource sample to obtain the first sample prediction information of each multimedia resource sample; the first media element is an image.

[0109] The first target network is a multi-classification task network model, where the classification target can be multiple selected topic categories. The first sample prediction information represents the prediction of the topic category for each multimedia resource sample, and may include the predicted probability that each multimedia resource sample belongs to each topic category.

[0110] In the above embodiments, considering that the topic category of multimedia resources such as videos is closely related to image features, a downstream topic classification model can be constructed to train the image feature extraction network corresponding to the image.

[0111] In another embodiment of this disclosure, for media elements of the text category, such as Figure 6 As shown, step S140 may further include:

[0112] In step S142, the search text corresponding to each multimedia resource sample is obtained.

[0113] It is feasible. The multiple multimedia resource samples selected in this embodiment are also the multimedia resources clicked by the user after entering the search text in the search service. Therefore, a correspondence between multimedia resource samples and search text can be constructed, and the search text can be used as a reference for the text features of multimedia resource samples.

[0114] In step S143, the search text corresponding to each multimedia resource sample is input into the second target network for text feature extraction processing to obtain the search text feature information of each multimedia resource sample.

[0115] The second target network can be a BERT model, including modules such as text encoding and embedding representation, which uses the output text feature vector as the search text feature information corresponding to each multimedia resource sample.

[0116] Furthermore, the text feature vectors output by the BERT model are regularized (e.g., L2 regularization) to obtain the search text feature information.

[0117] Furthermore, the resource feature information can be regularized (e.g., L2 regularization) to obtain regularized resource feature information.

[0118] In step S144, the search text feature information and resource feature information of each multimedia resource sample are cross-calculated for similarity, and used as feature similarity information of multiple target samples.

[0119] The target sample corresponds to a combination of the search text feature information of the first multimedia resource sample and the resource feature information of the second multimedia resource sample; both the first multimedia resource sample and the second multimedia resource sample are any multimedia resource samples.

[0120] Considering that the similarity between resource feature information and search text feature information is higher when they belong to the same multimedia resource sample than when they do not, a text feature extraction network corresponding to the text can be trained through comparative learning of resource feature information and search text feature information. This allows for better extraction of text features from multimedia resources. Comparative learning mainly includes a proxy task and an objective function. In the comparative learning of resource feature information and search text feature information, the proxy task involves cross-calculating the similarity between the search text feature information and the resource feature information of each multimedia resource sample to obtain feature similarity information for multiple target samples. The objective function is used to calculate the loss data between the predicted and true values ​​of the target samples, thereby guiding the model's learning direction.

[0121] For example, there are N multimedia resource samples, and the resource feature information is as follows: (1≤i≤N), the search text feature information are respectively (1≤j≤N), when i=j, and For the same multimedia resource sample, when i≠j, and These correspond to two different multimedia resource samples. For the target sample < , > The resource characteristic information of multimedia resource samples Search text feature information Performing a dot product operation yields the feature similarity information of the target samples, such as... Resource feature information for the i-th multimedia resource sample Search text feature information of the j-th multimedia resource sample The result of the dot product, Resource feature information for the j-th multimedia resource sample Search text feature information of the i-th multimedia resource sample The result of the dot product. It should be noted that when i ≠ j, and The meanings they represent are not the same, therefore, for N multimedia resource samples, a total of N will be generated. 2 One target sample.

[0122] In step S145, the feature similarity information of each target sample is input into the third target network corresponding to the second media element. The third target network performs click prediction processing on each target sample based on the target feature information of each target sample to obtain the second sample prediction information of each target sample.

[0123] The second sample prediction information represents the probability that the second multimedia resource sample will be clicked when the current search text is consistent with the search text corresponding to the first multimedia resource sample. It also represents the probability that the first multimedia resource sample and the second multimedia resource sample are the same multimedia resource sample. The second sample prediction information can also represent the prediction probability of whether the search text feature information and resource feature information associated with the target sample correspond to the same multimedia resource sample. The second media element is text.

[0124] The third target network can be a binary classification task network.

[0125] In the above embodiments, considering that the similarity between resource feature information and search text feature information is higher when they belong to the same multimedia resource sample than when they do not belong to the same multimedia resource sample, the text feature extraction network corresponding to the text can be trained by comparing and learning the resource feature information and search text feature information to better extract the text features of multimedia resources.

[0126] It should be noted that other text-related application tasks can also be constructed to train the text feature extraction network, but this disclosure does not limit this.

[0127] In step S150, based on the prediction information of multiple samples, the element feature extraction network corresponding to each media element is trained to obtain the trained element feature extraction network corresponding to each media element.

[0128] In this embodiment of the disclosure, each sample prediction information in the multiple sample prediction information of multiple multimedia resource samples is related to a media element, and is used to train the element feature extraction network corresponding to this media element.

[0129] In one embodiment of this disclosure, for media elements of the image category, based on step S141 provided in the foregoing embodiments, such as... Figure 5 As shown, step S250 may include:

[0130] In step S151, the first supervision information of each multimedia resource sample is obtained, and the first supervision information represents the topic category of each multimedia resource sample.

[0131] It is feasible to use the topic tags (such as hashtags) added by users when publishing multimedia resource samples as the topic category of the multimedia resource samples, and thus use them as the first supervision information for training the image feature extraction network.

[0132] In step S152, the first loss value is calculated based on the first sample prediction information and the first supervision information of each multimedia resource sample.

[0133] Yes, it is feasible. The loss function can be the cross-entropy loss function. Based on the operation rules of the cross-entropy loss function, the first sample prediction information of each multimedia resource sample, and the first supervision information of each multimedia resource sample, the first loss value can be calculated.

[0134] In step S153, based on the first loss value, the image feature extraction network corresponding to the first media element is trained to obtain the trained image feature extraction network.

[0135] In the above embodiments, the topic category of each multimedia resource sample is used as the first supervision information for training the image feature extraction network, and the image feature extraction network is trained based on the loss value of the first supervision information and the first sample prediction information. This can improve the training efficiency of the network and the ability of the image feature extraction network, thereby enabling better extraction of image features from multimedia resources.

[0136] In one embodiment of this disclosure, for media elements of the text category, based on steps S242 to S245 provided in the foregoing embodiments, such as... Figure 6 As shown, step S250 may further include:

[0137] In step S154, second supervision information for each target sample is determined. The second supervision information indicates whether the first multimedia resource sample and the second multimedia resource sample associated with the target sample are the same multimedia resource sample.

[0138] The second supervision information for each target sample is determined. The second supervision information indicates whether the search text feature information and resource feature information associated with each target sample correspond to the same multimedia resource sample, that is, whether the first multimedia resource sample and the second multimedia resource sample associated with the target sample are the same multimedia resource sample.

[0139] While obtaining multiple target samples through proxy tasks, corresponding labels can also be constructed based on the attributes of the target samples themselves, serving as second-level supervision information. Specifically, for the resource feature information of the i-th multimedia resource sample... It should be related to the search text feature information of the i-th multimedia resource sample. The closest, therefore the target sample < , The click probability represented by the second set of prediction information is higher than that of the target sample. , >and target samples< , The click probabilities represented by the second sample prediction information are all higher (i≠j). Therefore, the target sample < , The label of > is set to 1, and the target sample < , >and target samples< , The label for samples with i ≠ j is set to 0. The label of the target sample is also the second supervision information.

[0140] In step S155, the second loss value is calculated based on the second sample prediction information and the second supervision information of each target sample.

[0141] Yes, it is feasible. The loss function can be the cross-entropy loss function. Based on the operation rules of the cross-entropy loss function, the second sample prediction information of each target sample, and the second supervision information of each target sample, the second loss value can be calculated.

[0142] In step S156, based on the second loss value, the text feature extraction network corresponding to the second media element is trained to obtain the trained text feature extraction network.

[0143] In the above embodiments, a contrastive learning approach is used to construct target samples and second-supervisory information, eliminating the need for label annotation and effectively improving the training efficiency of the text feature extraction network. Simultaneously, training the text feature extraction network based on searched text feature information and resource feature information strengthens its capabilities, thereby enabling better extraction of text features from multimedia resources.

[0144] In another embodiment of this disclosure, while training the element feature extraction network corresponding to each element, the feature interaction network can also be trained based on the prediction information of multiple samples of each multimedia resource sample to obtain the trained feature interaction network.

[0145] Compared to the feature interaction fusion method of multi-stream mode, this embodiment only requires training one feature interaction network, which improves the overall training efficiency.

[0146] As can be seen from the technical solutions provided in the embodiments of this specification above, the technical solutions provided in the embodiments of this disclosure, in the model training stage, input each media element of each multimedia resource sample into the element feature extraction network corresponding to the media element, perform element feature extraction processing, and obtain multiple element feature information of each multimedia resource sample, each element feature information corresponding to a media element; then, the multiple element feature information of each multimedia resource sample is concatenated and input into the same feature interaction network, and feature interaction processing is performed to obtain resource feature information of each multimedia resource sample, the resource feature information being a deep mining of the features of different media elements of a multimedia resource sample; using the resource feature information of each multimedia resource sample, target behavior prediction processing corresponding to each media element is performed to obtain multiple sample prediction information, and then the multiple sample prediction information of multiple multimedia resource samples can be used to train the element feature extraction network corresponding to each media element, thereby obtaining the trained multiple element feature extraction networks. By utilizing the technical solutions provided in this disclosure, feature extraction networks are applied to various media elements in multimedia resources, and target processing tasks corresponding to each media element are used to train the feature extraction networks. This allows for more effective extraction of features from different media elements. Furthermore, by simultaneously inputting the element feature information corresponding to various media elements in multimedia resources into the same feature interaction network for feature interaction, the correlation between the element feature information of different media elements can be mined, thus more fully integrating the element feature information of multiple media elements. More effective feature extraction and more complete feature fusion of different media elements in multimedia resources result in more effective feature representations of multimedia resources, thereby meeting the ever-increasing business needs of downstream applications.

[0147] Figure 7 This is a flowchart illustrating a method for representing the features of a multimedia resource according to an exemplary embodiment. For example... Figure 7 As shown, the method may include the following steps:

[0148] In step 210, multiple media elements of the multimedia resource are acquired.

[0149] In this embodiment of the disclosure, the multimedia resource can be an object that integrates multiple media elements, such as short videos and live streaming data, used in a multimedia resource library, and can be applied to multimedia resource search, recommendation, and other services. The multiple media elements contained in the multimedia resource can include at least two of the following: images, text, audio / animation, data, links, etc.

[0150] In step S220, each of the multiple media elements is input into the element feature extraction network corresponding to each media element to perform element feature extraction processing, thereby obtaining multiple element feature information of the multimedia resource.

[0151] The element feature network is obtained through the training method of the feature extraction network provided in the above embodiments.

[0152] In this embodiment of the disclosure, the training of the element feature extraction network corresponding to each media element can adopt a feature extraction network training method provided in the foregoing embodiments, which will not be repeated here.

[0153] For example, taking images as an example, ResNet-50 (a residual neural network) can be used as the image feature extraction network. The image of the multimedia resource is input into ResNet-50 for image feature extraction processing. The feature map of the last layer output by ResNet-50 is divided into block sequences as the image feature information of the multimedia resource.

[0154] For example, taking text as an example, BERT (a semantic representation model) can be used as a text feature extraction network. The text of multimedia resources can be input into BERT for text feature extraction processing, including text segmentation and embedding, to obtain the text feature information of multimedia resources.

[0155] In step S230, multiple element feature information is input into the feature interaction network for feature interaction processing to obtain the resource feature information of the multimedia resources.

[0156] In this embodiment, a single-stream feature fusion interaction method is adopted. This involves first merging the element feature information corresponding to each of the various media elements of a multimedia resource into a single feature interaction network for feature fusion, thereby obtaining the resource feature information of the multimedia resource sample. Compared to a multi-stream fusion method that separately inputs the element feature information corresponding to each media element into its respective feature interaction network, mines the element feature information of each media element separately, and then fuses them, the single-stream feature fusion interaction method provided in this embodiment can perform feature interaction and mining on the element feature information corresponding to different media elements. This allows for a more thorough fusion of element feature information among multiple media elements, resulting in more accurate and effective resource feature information as a feature representation of the multimedia resource.

[0157] For example, a Transformer model based on a multi-layered encoder-decoder and bidirectional attention mechanism can be used as the feature interaction network. Target feature information, containing image and text features of multimedia resources, is input into a multi-layered Transformer model for feature interaction processing. This allows for sufficient interaction between the image and text features of the multimedia resources, thereby obtaining the resource feature representation of the multimedia resources, which is the output of the Transformer model.<Cls token> The corresponding results.

[0158] In another embodiment of this disclosure, the feature interaction network may also be obtained after training in a feature extraction network training method provided in the foregoing embodiments.

[0159] The specific implementation of each step in the methods described in the above embodiments has been described in detail in the embodiments of the method, and will not be elaborated further here.

[0160] Please see Figure 8The diagram illustrates a network architecture according to an exemplary embodiment. Taking video as an example, during the training phase, images and text for each video sample are obtained from multiple video samples used as training data. Images and text are two different types of media elements in a video. The images of each video sample are input into an image feature extraction network, and the text of each video sample is input into a text feature extraction network to obtain multiple element feature information for each video sample. The multiple element feature information for each video sample includes image feature information and text feature information for each video sample. The image feature information and text feature information of each video sample are then processed...<Sep token> Piece them together, and add them at the beginning and end.<Cls token> and<Eos token> As a marker, the data input to the feature interaction network can be obtained, which is the target feature information of each video sample. A multi-layer encoder-decoder model is used as the feature interaction network. The target feature information, containing both image and text features of each video sample, is input into a multi-layer encoder-decoder model for feature interaction processing. This allows for full interaction and fusion of the image and text features of each video sample, thereby obtaining the video feature representation of each video sample (i.e.,...). Figure 8 Cls token output (in the context of the Cls token).

[0161] In the embodiments of this disclosure, the image feature extraction network and the text feature extraction network are trained by performing different application processes on the video feature representation of each video sample. Specifically, to train the image feature extraction network, the video feature representation of each video sample is input into a first target network. The first target network performs topic classification prediction for each video sample based on the input video feature representation, obtaining the topic prediction information for each video sample. Then, it can calculate the loss value by combining the original topic category label of each video sample and adjust the model parameters of the image feature extraction network based on the loss value. Specifically, the text feature extraction network is trained by comparing and learning the search text feature information corresponding to each video sample with the video feature representation. First, the search text corresponding to each video sample is input into the second target network for feature extraction, obtaining the search text feature information for each video sample. Based on the comparison and learning between the video feature representation and the search text feature information of each video sample, target samples are constructed, and labels for these target samples are built as supervision information. The labels of the target samples indicate whether the video feature representation and search text feature information associated with the target sample correspond to the same video sample. The target samples are then input into the third target network for click prediction processing to obtain classification prediction information. The model parameters of the text feature extraction network can then be adjusted using the loss value between the classification prediction information and the supervision information. In the application phase, the network architecture does not include the modules shown in the dotted lines. The process from acquiring video images and text to obtaining video feature representations is consistent with the training phase and will not be elaborated here.

[0162] Once the video feature representation is obtained, it can be used for different downstream tasks, depending on business needs, and is not limited to topic classification and search businesses.

[0163] In addition, it should be noted that, Figure 8 The network architecture shown is merely one provided in this disclosure. In actual application scenarios, the target networks used in the training phase may have different network types or correspond to other types of services. The feature interaction networks used in the training and application phases are not limited to multi-layer encoder-decoder models.

[0164] Figure 9 This is a block diagram illustrating a training apparatus for a feature extraction network according to an exemplary embodiment. (Refer to...) Figure 9 The training device includes:

[0165] The first acquisition module 910 is configured to acquire multiple media elements of each multimedia resource sample in a plurality of multimedia resource samples, the plurality of media elements including images and text;

[0166] The first element feature extraction module 920 is configured to input each media element of each multimedia resource sample into the element feature extraction network corresponding to each media element, perform element feature extraction processing, and obtain multiple element feature information of each multimedia resource sample.

[0167] The first element feature interaction module 930 is configured to concatenate multiple element feature information of each multimedia resource sample and input them into the feature interaction network for feature interaction processing to obtain the resource feature information of each multimedia resource sample.

[0168] The sample prediction module 940 is configured to perform target behavior prediction processing corresponding to each media element based on the resource feature information of each multimedia resource sample, and obtain multiple sample prediction information.

[0169] The training module 950 is configured to train the element feature extraction network corresponding to each media element based on prediction information from multiple samples, thereby obtaining the trained element feature extraction network corresponding to each media element.

[0170] In one embodiment of this disclosure, the first acquisition module 910 may include:

[0171] The first media element acquisition unit is configured to extract multiple frames of images from the target multimedia resource sample to obtain the target image of the target multimedia resource sample;

[0172] The second media element acquisition unit is configured to perform speech recognition processing or text detection processing on the target multimedia resource sample to obtain the target text of the target multimedia resource sample.

[0173] The target image and the target text are used as multiple media elements of the target multimedia resource sample; the target multimedia resource sample is any one of the multiple multimedia resource samples.

[0174] In one embodiment of this disclosure, the first element feature extraction module 920 may include:

[0175] The first feature extraction unit is configured to input the first media element of each multimedia resource sample into the image feature extraction network corresponding to the first media element, perform image feature extraction processing, and obtain the image feature information of each multimedia resource sample, wherein the first media element is an image;

[0176] The second feature extraction unit is configured to input the second media element of each multimedia resource sample into the text feature extraction network corresponding to the second media element, perform text feature extraction processing, and obtain the text feature information of each multimedia resource sample, wherein the second media element is text;

[0177] The image feature information and text feature information of each multimedia resource sample constitute multiple element feature information of each multimedia resource sample.

[0178] In one embodiment of this disclosure, the first element feature interaction module 930 may include:

[0179] The feature fusion unit is configured to perform the concatenation of image feature information and text feature information of each multimedia resource sample to obtain target feature information of each multimedia resource sample.

[0180] The feature interaction unit is configured to input the target feature information of each multimedia resource sample into the feature interaction network, perform feature interaction processing, and obtain the resource feature information of each multimedia resource sample; the feature interaction network adopts a multi-layer encoder and decoder network structure.

[0181] In one embodiment of this disclosure, the sample prediction module 940 may include:

[0182] The first prediction unit is configured to input the resource feature information of each multimedia resource sample into a first target network corresponding to the first media element, and the first target network performs topic classification prediction processing on each multimedia resource sample based on the resource feature information of each multimedia resource sample to obtain the first sample prediction information of each multimedia resource sample; the first media element is an image.

[0183] In one embodiment of this disclosure, the training module 950 may include:

[0184] The first supervisory information acquisition unit is configured to acquire first supervisory information for each of the multimedia resource samples, wherein the first supervisory information characterizes the topic category of each of the multimedia resource samples.

[0185] The first loss calculation unit is configured to perform calculation of a first loss value based on the first sample prediction information and the first supervision information of each multimedia resource sample.

[0186] The first training unit is configured to train the image feature extraction network corresponding to the first media element based on the first loss value, so as to obtain the trained image feature extraction network.

[0187] In one embodiment of this disclosure, the sample prediction module 940 may further include:

[0188] The search text acquisition unit is configured to acquire the search text corresponding to each of the multimedia resource samples.

[0189] The search text feature extraction unit is configured to input the search text corresponding to each of the multimedia resource samples into the second target network, perform text feature extraction processing, and obtain the search text feature information of each of the multimedia resource samples.

[0190] The sample construction unit is configured to perform a cross-similarity calculation on the search text feature information and resource feature information of each multimedia resource sample, which is used as feature similarity information for multiple target samples; the target sample corresponds to the combination of the search text feature information of the first multimedia resource sample and the resource feature information of the second multimedia resource sample; the first multimedia resource sample and the second multimedia resource sample are both any of the multimedia resource samples.

[0191] The second prediction unit is configured to input the feature similarity information of each target sample into a third target network corresponding to the second media element. The third target network performs click prediction processing on each target sample based on the target feature information of each target sample to obtain second sample prediction information for each target sample. The second sample prediction information represents the probability that the second multimedia resource sample is clicked when the current search text is consistent with the search text corresponding to the first multimedia resource sample. It also represents the probability that the first multimedia resource sample and the second multimedia resource sample are the same multimedia resource sample. The second media element is text.

[0192] In one embodiment of this disclosure, the training module 950 may further include:

[0193] The second supervision information acquisition unit is configured to execute the second supervision information for determining each of the target samples, wherein the second supervision information indicates whether the first multimedia resource sample associated with the target sample and the second multimedia resource sample are the same multimedia resource sample.

[0194] The second loss calculation unit is configured to calculate a second loss value based on the second sample prediction information and the second supervision information of each target sample.

[0195] The second training unit is configured to train the text feature extraction network corresponding to the second media element based on the second loss value, thereby obtaining the trained text feature extraction network.

[0196] In one embodiment of this disclosure, the training device may further include:

[0197] The third training unit is configured to execute multiple sample prediction information based on each of the multimedia resource samples to train the feature interaction network, thereby obtaining the trained feature interaction network.

[0198] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0199] Figure 10 This is a block diagram illustrating a feature representation apparatus for a multimedia resource according to an exemplary embodiment. (Refer to...) Figure 10 The feature representation device includes:

[0200] The second acquisition module 1010 is configured to acquire multiple media elements of multimedia resources, including images and text.

[0201] The second element feature extraction module 1020 is configured to input each of the plurality of media elements into the element feature extraction network corresponding to each media element, perform element feature extraction processing, and obtain multiple element feature information of the multimedia resource; the element feature network is obtained by the training method of the feature extraction network provided in the above embodiment;

[0202] The second element feature interaction module 1030 is configured to input the multiple element feature information into the feature interaction network, perform feature interaction processing, and obtain the resource feature information of the multimedia resource.

[0203] In this embodiment of the disclosure, an electronic device is also provided, including: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement a training method for a feature extraction network or a feature representation method for multimedia resources as described in this embodiment of the disclosure.

[0204] Figure 11 This is a block diagram of an electronic device for implementing a training method for a feature extraction network or a feature representation method for multimedia resources, according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 11As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a training method for a feature extraction network or a feature representation method for multimedia resources. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.

[0205] Figure 12 This is a block diagram of an electronic device for implementing a training method for a feature extraction network or a feature representation method for multimedia resources, according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 12 As shown, this electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a training method for a feature extraction network or a feature representation method for multimedia resources.

[0206] Those skilled in the art will understand that Figure 11 and Figure 12 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0207] In this disclosure, a computer-readable storage medium including instructions is also provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform the training method of the feature extraction network or the feature representation method of multimedia resources in this disclosure.

[0208] In this embodiment of the disclosure, a computer program product is also provided, including computer instructions, which, when executed by a processor, implement the training method of the feature extraction network or the feature representation method of multimedia resources in this embodiment of the disclosure.

[0209] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0210] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0211] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for training a feature extraction network, characterized in that, The method includes: Retrieve multiple media elements from each multimedia resource sample in a multi-multimedia resource sample; Each media element of each multimedia resource sample is input into the element feature extraction network corresponding to each media element, and the element feature is extracted to obtain multiple element feature information of each multimedia resource sample. The feature information of multiple elements of each multimedia resource sample is concatenated and input into the feature interaction network for feature interaction processing to obtain the resource feature information of each multimedia resource sample. Based on the resource feature information of each multimedia resource sample, target behavior prediction processing is performed corresponding to each media element to obtain multiple sample prediction information corresponding to each multimedia resource sample; the multiple sample prediction information corresponds one-to-one with the multiple media elements in the corresponding multimedia resource sample; the business type corresponding to the target behavior prediction processing has a business association relationship with the corresponding media element; Based on the sample prediction information corresponding to each media element from the multiple sample prediction information corresponding to each multimedia resource sample, the element feature extraction network corresponding to each media element is trained to obtain the trained element feature extraction network corresponding to each media element. The plurality of sample prediction information includes second sample prediction information. The step of performing target behavior prediction processing corresponding to each media element based on the resource feature information of each multimedia resource sample to obtain the plurality of sample prediction information corresponding to each multimedia resource sample includes: Obtain the search text corresponding to each of the multimedia resource samples; The search text corresponding to each of the multimedia resource samples is input into the second target network for text feature extraction processing to obtain the search text feature information of each of the multimedia resource samples. The search text feature information and resource feature information of each multimedia resource sample are cross-calculated for similarity to obtain the feature similarity information of multiple target samples; each target sample corresponds to the combination of the search text feature information of the first multimedia resource sample and the resource feature information of the second multimedia resource sample; the first multimedia resource sample and the second multimedia resource sample are both any of the multimedia resource samples. The feature similarity information of each target sample is input into the third target network corresponding to the second media element. The third target network performs click prediction processing on each target sample based on the target feature information of each target sample to obtain the second sample prediction information of each target sample. The second sample prediction information represents the probability that the second multimedia resource sample will be clicked when the current search text is consistent with the search text corresponding to the first multimedia resource sample. The second media element is text. Each target sample is a training sample of the text feature extraction network corresponding to the second media element.

2. The training method for the feature extraction network according to claim 1, characterized in that, The step of obtaining multiple media elements from each multimedia resource sample in a plurality of multimedia resource samples includes: Multiple frames of images are extracted from the target multimedia resource sample to obtain the target image of the target multimedia resource sample; The target multimedia resource sample is subjected to speech recognition processing or text detection processing to obtain the target text of the target multimedia resource sample. The target image and the target text are used as multiple media elements of the target multimedia resource sample; the target multimedia resource sample is any one of the multiple multimedia resource samples.

3. The training method for the feature extraction network according to claim 1, characterized in that, The step involves inputting each media element of each multimedia resource sample into an element feature extraction network corresponding to each media element, performing element feature extraction processing, and obtaining multiple element feature information for each multimedia resource sample, including: The first media element of each multimedia resource sample is input into the image feature extraction network corresponding to the first media element, and image feature extraction processing is performed to obtain the image feature information of each multimedia resource sample, wherein the first media element is an image; The second media element of each multimedia resource sample is input into the text feature extraction network corresponding to the second media element, and the text feature extraction process is performed to obtain the text feature information of each multimedia resource sample, wherein the second media element is text; The image feature information and text feature information of each multimedia resource sample constitute multiple element feature information of each multimedia resource sample.

4. The training method for the feature extraction network according to claim 3, characterized in that, The step of concatenating multiple element feature information of each multimedia resource sample and inputting it into a feature interaction network for feature interaction processing to obtain resource feature information of each multimedia resource sample includes: The image feature information and text feature information of each multimedia resource sample are concatenated to obtain the target feature information of each multimedia resource sample. The target feature information of each multimedia resource sample is input into the feature interaction network for feature interaction processing to obtain the resource feature information of each multimedia resource sample; the feature interaction network adopts a multi-layer encoder and decoder network structure.

5. The training method for the feature extraction network according to claim 1, characterized in that, The multiple sample prediction information includes first sample prediction information; the step of performing target behavior prediction processing corresponding to each media element based on the resource feature information of each multimedia resource sample to obtain multiple sample prediction information corresponding to each multimedia resource sample includes: The resource feature information of each multimedia resource sample is input into the first target network corresponding to the first media element. The first target network performs topic classification prediction processing on each multimedia resource sample based on the resource feature information of each multimedia resource sample to obtain the first sample prediction information of each multimedia resource sample; the first media element is an image.

6. The training method for the feature extraction network according to claim 5, characterized in that, The step of training an element feature extraction network corresponding to each media element based on multiple sample prediction information corresponding to each of the multimedia resource samples, to obtain a trained element feature extraction network corresponding to each media element, includes: Obtain first supervision information for each of the multimedia resource samples, wherein the first supervision information characterizes the topic category of each of the multimedia resource samples; Based on the first sample prediction information and the first supervision information of each multimedia resource sample, the first loss value is calculated. Based on the first loss value, the image feature extraction network corresponding to the first media element is trained to obtain the trained image feature extraction network.

7. The training method for the feature extraction network according to claim 1, characterized in that, The step of training an element feature extraction network corresponding to each media element based on the sample prediction information corresponding to each of the multiple sample prediction information corresponding to each of the multimedia resource samples, to obtain the trained element feature extraction network corresponding to each media element, includes: Determine second supervision information for each of the target samples, wherein the second supervision information indicates whether the first multimedia resource sample associated with the target sample and the second multimedia resource sample are the same multimedia resource sample; Based on the second sample prediction information and the second supervision information of each target sample, the second loss value is calculated. Based on the second loss value, the text feature extraction network corresponding to the second media element is trained to obtain the trained text feature extraction network.

8. The training method for the feature extraction network according to claim 1, characterized in that, The method further includes: Based on the prediction information of multiple samples corresponding to each of the multimedia resource samples, the feature interaction network is trained to obtain the trained feature interaction network.

9. A method for representing the features of multimedia resources, characterized in that, The method includes: To acquire multiple media elements of a multimedia resource; Each of the plurality of media elements is input into the element feature extraction network corresponding to each media element, and element feature extraction processing is performed to obtain the element feature information of the multimedia resource; the element feature network is obtained by the training method of the feature extraction network according to any one of claims 1 to 8; The feature information of the multiple elements is input into the feature interaction network for feature interaction processing to obtain the resource feature information of the multimedia resource.

10. A training apparatus for a feature extraction network, characterized in that, The device includes: The first acquisition module is configured to acquire multiple media elements of each multimedia resource sample from multiple multimedia resource samples; The first element feature extraction module is configured to input each media element of each multimedia resource sample into the element feature extraction network corresponding to each media element, perform element feature extraction processing, and obtain multiple element feature information of each multimedia resource sample. The first element feature interaction module is configured to concatenate multiple element feature information of each multimedia resource sample and input them into the feature interaction network for feature interaction processing to obtain the resource feature information of each multimedia resource sample. The sample prediction module is configured to perform target behavior prediction processing based on the resource feature information of each multimedia resource sample, corresponding to each media element, to obtain multiple sample prediction information for each multimedia resource sample; the multiple sample prediction information corresponds one-to-one with the multiple media elements in the corresponding multimedia resource sample; the service type corresponding to the target behavior prediction processing has a service association relationship with the corresponding media element; The training module is configured to execute the sample prediction information corresponding to each media element from the multiple sample prediction information corresponding to each multimedia resource sample, and train the element feature extraction network corresponding to each media element to obtain the trained element feature extraction network corresponding to each media element. The plurality of sample prediction information includes second sample prediction information, and the sample prediction module further includes: The search text acquisition unit is configured to acquire the search text corresponding to each of the multimedia resource samples. The search text feature extraction unit is configured to input the search text corresponding to each of the multimedia resource samples into the second target network, perform text feature extraction processing, and obtain the search text feature information of each of the multimedia resource samples. The sample construction unit is configured to perform a cross-calculation of the similarity between the search text feature information and the resource feature information of each multimedia resource sample, and use them as feature similarity information for multiple target samples; each target sample corresponds to a combination of the search text feature information of the first multimedia resource sample and the resource feature information of the second multimedia resource sample; the first multimedia resource sample and the second multimedia resource sample are both any of the multimedia resource samples. The second prediction unit is configured to input the feature similarity information of each target sample into a third target network corresponding to the second media element. The third target network performs click prediction processing on each target sample based on the target feature information of each target sample to obtain the second sample prediction information of each target sample. The second sample prediction information represents the probability that the second multimedia resource sample is clicked when the current search text is consistent with the search text corresponding to the first multimedia resource sample. The second media element is text. Each target sample is a training sample of the text feature extraction network corresponding to the second media element.

11. The apparatus according to claim 10, characterized in that, The first acquisition module includes: The first media element acquisition unit is configured to extract multiple frames of images from the target multimedia resource sample to obtain the target image of the target multimedia resource sample; The second media element acquisition unit is configured to perform speech recognition processing or text detection processing on the target multimedia resource sample to obtain the target text of the target multimedia resource sample. The target image and the target text are used as multiple media elements of the target multimedia resource sample; the target multimedia resource sample is any one of the multiple multimedia resource samples.

12. The apparatus according to claim 10, characterized in that, The first element feature extraction module includes: The first feature extraction unit is configured to input the first media element of each multimedia resource sample into the image feature extraction network corresponding to the first media element, perform image feature extraction processing, and obtain the image feature information of each multimedia resource sample, wherein the first media element is an image; The second feature extraction unit is configured to input the second media element of each multimedia resource sample into the text feature extraction network corresponding to the second media element, perform text feature extraction processing, and obtain the text feature information of each multimedia resource sample, wherein the second media element is text; The image feature information and text feature information of each multimedia resource sample constitute multiple element feature information of each multimedia resource sample.

13. The apparatus according to claim 12, characterized in that, The first element feature interaction module includes: The feature fusion unit is configured to perform the concatenation of image feature information and text feature information of each multimedia resource sample to obtain target feature information of each multimedia resource sample. The feature interaction unit is configured to input the target feature information of each multimedia resource sample into the feature interaction network, perform feature interaction processing, and obtain the resource feature information of each multimedia resource sample; the feature interaction network adopts a multi-layer encoder and decoder network structure.

14. The apparatus according to claim 10, characterized in that, The plurality of sample prediction information includes first sample prediction information; the sample prediction module includes: The first prediction unit is configured to input the resource feature information of each multimedia resource sample into a first target network corresponding to the first media element, and the first target network performs topic classification prediction processing on each multimedia resource sample based on the resource feature information of each multimedia resource sample to obtain the first sample prediction information of each multimedia resource sample; the first media element is an image.

15. The apparatus according to claim 14, characterized in that, The training module includes: The first supervisory information acquisition unit is configured to acquire first supervisory information for each of the multimedia resource samples, wherein the first supervisory information characterizes the topic category of each of the multimedia resource samples. The first loss calculation unit is configured to perform calculation of a first loss value based on the first sample prediction information and the first supervision information of each multimedia resource sample. The first training unit is configured to train the image feature extraction network corresponding to the first media element based on the first loss value, so as to obtain the trained image feature extraction network.

16. The apparatus according to claim 10, characterized in that, The training module includes: The second supervision information acquisition unit is configured to execute the second supervision information for determining each of the target samples, wherein the second supervision information indicates whether the first multimedia resource sample associated with the target sample and the second multimedia resource sample are the same multimedia resource sample. The second loss calculation unit is configured to calculate a second loss value based on the second sample prediction information and the second supervision information of each target sample. The second training unit is configured to train the text feature extraction network corresponding to the second media element based on the second loss value, thereby obtaining the trained text feature extraction network.

17. The apparatus according to claim 10, characterized in that, The device further includes: The third training unit is configured to train the feature interaction network based on multiple sample prediction information corresponding to each of the multimedia resource samples, so as to obtain the trained feature interaction network.

18. A feature representation device for multimedia resources, characterized in that, The device includes: The second acquisition module is configured to acquire multiple media elements of multimedia resources; The second element feature extraction module is configured to input each of the plurality of media elements into the element feature extraction network corresponding to each media element, perform element feature extraction processing, and obtain multiple element feature information of the multimedia resource; the element feature network is obtained by the training method of the feature extraction network according to any one of claims 1 to 8. The second element feature interaction module is configured to input the multiple element feature information into the feature interaction network, perform feature interaction processing, and obtain the resource feature information of the multimedia resource.

19. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the training method of the feature extraction network as described in any one of claims 1 to 8 or the feature representation method of multimedia resources as described in claim 9.

20. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the training method of the feature extraction network as described in any one of claims 1 to 8 or the feature representation method of the multimedia resources as described in claim 9.