Model training method and device, data retrieval method and device and electronic equipment

By training the video and text feature extraction models and aligning them in the same feature space, the problem of requiring a bridge of the same type of features in the existing technology is solved, and the efficiency of video feature retrieval is improved.

CN120653802APending Publication Date: 2025-09-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410307810.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing video feature retrieval methods need to first retrieve data of the same type as the input data as a bridge, and cannot directly obtain the final retrieval results, resulting in low efficiency.

Method used

The video feature extraction model and the text feature extraction model are trained using the training dataset to generate a model that can align video and text features in the same feature space and perform retrieval directly in this feature space.

Benefits of technology

This achieves the goal of eliminating the need for similar feature bridges when retrieving video data or text data, thereby improving retrieval efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653802A_ABST
    Figure CN120653802A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the steps of obtaining a plurality of training sample groups; each training sample group comprises a plurality of sample data; the sample data comprises target sample data, positive sample data and negative sample data; the sample data comprises a sample video or a video tag; each video tag has a corresponding tag description text; the sample data of the at least one training sample group comprises a sample video and a video tag; and carrying out at least one training operation on the initial video feature extraction model and the initial text feature extraction model through the training data set until a training ending condition is met, and obtaining a trained video feature extraction model and a trained text feature extraction model. The video features extracted by the trained video feature extraction model and the text features extracted by the trained text feature extraction model are aligned to the same feature space, so that the retrieval efficiency is improved during retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of the Internet. Specifically, the present application relates to a model training method, a data retrieval method, a device and an electronic device. Background Art

[0002] With the development of internet technology, video-related feature retrieval, including video tag recognition or searching for videos by tag, has become an important part of video content characterization. Taking video tag recognition as an example, by automatically generating tags for massive amounts of user-generated content (UGC) videos, machines can provide video content features of varying granularity for downstream content distribution chains (such as recommendation systems and content operations), improving content distribution efficiency while significantly reducing the cost of manual content review.

[0003] Currently, feature retrieval for videos involves pre-storing multiple videos and their corresponding tag information in a database. When the target data is entered for a query, the database first searches for candidate data of the same type as the target data, and then determines the corresponding search results. For example, when the target video's search tag is entered, the database first searches for candidate videos similar to the target video, and then further filters the candidate videos' tags to obtain the search results. This current data retrieval method requires first searching for data of the same type as the input data as a bridge, and cannot directly obtain the final search results. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a model training method, a data retrieval method, a device, and an electronic device. The technical solutions provided by the present disclosure are as follows:

[0005] In one aspect, an embodiment of the present application provides a method for model training, the method comprising:

[0006] Obtain a training data set; the training data set includes multiple training sample groups; each training sample group includes multiple sample data; the sample data includes target sample data, positive sample data associated with the target sample data, and negative sample data not associated with the target sample data; the sample data includes sample videos or video labels; each video label has a corresponding label description text; the sample data of at least one training sample group includes a sample video and a video label;

[0007] Performing at least one training operation on the initial video feature extraction model and the initial text feature extraction model using the training data set until a training end condition is met, thereby obtaining a trained video feature extraction model and a trained text feature extraction model;

[0008] The training operations include:

[0009] The video features of each sample video included in the training data set are extracted by the initial video feature extraction model, and the text features of the label description text of each video label included in the training data set are extracted by the initial text feature extraction model;

[0010] For each training sample group, a loss value corresponding to the training sample group is determined based on the difference between sample features corresponding to each sample data in the training sample group; the sample features include video features or text features;

[0011] Based on the loss values ​​corresponding to each training sample group, the training loss is determined; based on the training loss, the parameters of the initial video feature extraction model and the initial text feature extraction model are adjusted, and the initial video feature extraction model and the initial text feature extraction model after the adjusted parameters are used as the initial video feature extraction model and the initial text feature extraction model corresponding to the next training operation.

[0012] In some possible implementations, obtaining a training dataset includes:

[0013] Obtain multiple sample videos; each sample video is annotated with at least one video tag;

[0014] Deduplication is performed on the video labels corresponding to each sample video to obtain multiple video labels;

[0015] Determine the similarity between each sample video and the similarity between the label description texts corresponding to each video label;

[0016] Based on the similarity between each sample video, the similarity between the label description texts corresponding to each video label, and the correspondence between each sample video and its respective video label, multiple sample videos and multiple sample labels are divided into multiple training sample groups to obtain a training data set.

[0017] In some possible implementations, the initial video feature extraction model includes a first video extraction module and a second video extraction module; the initial text feature extraction model includes a first text extraction module and a second text extraction module;

[0018] Determine the similarity between each sample video and the similarity between the label description texts corresponding to each video label, including:

[0019] Extracting initial video features of each sample video through a first video extraction module, and determining similarities between the initial video features;

[0020] Extracting initial text features of the tag description text corresponding to each video tag by a first text extraction module, and determining the similarity between the initial text features;

[0021] The parameters of the initial video feature extraction model and the initial text feature extraction model are adjusted based on the training loss, including:

[0022] Parameters of the second video extraction module and the second text extraction module are adjusted based on the training loss.

[0023] In some possible implementations, based on the similarity between the sample videos, the similarity between the label description texts corresponding to the video labels, and the correspondence between each sample video and its respective video label, the plurality of sample videos and the plurality of sample labels are divided into a plurality of training sample groups, including:

[0024] Each sample video is regarded as a video node, and each video tag is regarded as a tag node. Based on the similarity between each sample video, the similarity between the tag description texts corresponding to each video tag, and the correspondence between each sample video and its respective video tag, each node is connected to generate a relationship structure diagram between multiple sample videos and multiple sample tags;

[0025] At least some of the nodes in the relationship structure graph are used as anchor points; the nodes include video nodes and tag nodes;

[0026] For each anchor point, the positive sample nodes and negative sample nodes of the anchor point are determined based on the connection relationship between each node in the relationship structure graph;

[0027] The anchor point is used as the target sample data, the positive sample node corresponding to the anchor point is used as the positive sample data, and the negative sample node corresponding to the anchor point is used as the negative sample data to obtain the training sample group corresponding to the anchor point.

[0028] In some possible implementations, based on the similarity between sample videos, the similarity between the label description texts corresponding to the video labels, and the correspondence between each sample video and its respective video label, each node is connected to generate a relationship structure diagram between multiple sample videos and multiple sample labels, including:

[0029] Determine the similarity between each sample video, connect the video nodes whose similarity meets a first threshold, and form isomorphic edges between each video node;

[0030] Determine the similarity between the label description texts corresponding to the respective video labels, connect the label nodes whose similarity meets the second threshold, and form isomorphic edges between the label nodes;

[0031] Based on the corresponding relationship between each sample video and its respective video label, heterogeneous edges between video nodes and label nodes are formed to obtain a relationship structure graph.

[0032] In some possible implementations, for each anchor point, determining the positive sample node and the negative sample node of the anchor point based on the connection relationship between each node in the relationship structure graph includes:

[0033] In the relationship structure graph, the nodes whose number of connected edges with the anchor point is less than or equal to the preset number are regarded as positive sample nodes of the anchor point;

[0034] In the relationship structure graph, nodes whose number of connected edges with the anchor point exceeds a preset number are regarded as negative sample nodes of the anchor point.

[0035] In some possible implementations, the training data set further includes at least one training sample group in which all sample data are sample videos, and at least one training sample group in which all sample data are video labels;

[0036] The method also includes:

[0037] Among the multiple training sample groups, the training sample group whose sample data are all video labels is used as a label sample group, the training sample group whose sample data are all sample videos is used as a video sample group, and the sample group whose sample data includes sample videos and video labels is used as a mixed sample group;

[0038] When the number of training operations reaches a preset number, multiple target sample groups are selected from the multiple training sample groups based on the preset ratio of the number of label sample groups, video sample groups, and mixed sample groups;

[0039] A new training set is generated based on multiple target sample groups; and training operations after a preset number of times are performed based on the new training data set.

[0040] In some possible implementations, in the quantity ratio, the number of mixed sample groups is greater than the number of label sample groups, and the number of mixed sample groups is greater than the number of video sample groups.

[0041] In some possible implementations, based on a preset ratio of the number of label sample groups, video sample groups, and mixed sample groups, multiple target sample groups are selected from multiple training sample groups, including:

[0042] Determine the loss value corresponding to each training sample group when the number of training operations reaches a preset number;

[0043] The training sample group whose loss value is greater than the preset threshold is taken as the first sample group;

[0044] Based on the preset quantity ratios among the label sample group, the video sample group, and the mixed sample group, a plurality of target sample groups are selected from the plurality of first sample groups.

[0045] In some possible implementations, for each training sample group, determining a loss value corresponding to the training sample group based on differences between sample features corresponding to each sample data in the training sample group includes:

[0046] Determining first difference information between target sample data and positive sample data in the training sample group, and determining second difference information between the target sample data and negative sample data;

[0047] A loss value of the sample group is determined based on the first difference information and the second difference information.

[0048] In some possible implementations, extracting initial video features of each sample video includes:

[0049] Acquire features of at least two different modalities of the sample video; the features of the different modalities include at least two of the following: picture features of a video frame of the sample video, audio features of a sample audio of the sample video, content text features of a video content text of the sample video, and title text features of a video title of the sample video;

[0050] Fusing features of at least two different modalities to obtain initial video features of the sample video;

[0051] The tag description text includes the video tag, at least two different levels of categories to which the video tag belongs, and description information of the video tag; extracting initial text features of the tag description text corresponding to each video tag includes:

[0052] The video tag, at least two categories of different levels to which the video tag belongs, and description information of the video tag are converted into corresponding text features to obtain initial text features of the video tag.

[0053] On the other hand, an embodiment of the present application provides a data retrieval method, the method comprising:

[0054] Acquire target data; the target data includes target video data or target text data;

[0055] If the target data is a target video, the video features of the target video are extracted using the trained video feature extraction model; if the target data is a target text, the text features of the target text are extracted using the trained text feature extraction model; the video feature extraction model and the text feature extraction model are trained based on the above-mentioned model training method;

[0056] Based on the extracted video features or text features, retrieval results corresponding to the target data are retrieved in the database; the retrieval results include at least one of video data and label data; wherein the database stores candidate video features corresponding to a plurality of video data respectively; the database also stores candidate text features corresponding to a plurality of label data respectively; the label data includes video labels and label description texts; the candidate video features are extracted based on a video feature extraction model; the candidate text features are extracted based on a text feature extraction model.

[0057] On the other hand, an embodiment of the present application provides a model training device, comprising:

[0058] An acquisition module is configured to acquire a training data set; the training data set includes a plurality of training sample groups; each training sample group includes a plurality of sample data; the sample data includes target sample data, positive sample data associated with the target sample data, and negative sample data not associated with the target sample data; the sample data includes sample videos or video labels; each video label has a corresponding label description text; the sample data of at least one training sample group includes a sample video and a video label;

[0059] A training module is used to perform at least one training operation on the initial video feature extraction model and the initial text feature extraction model using a training data set until a training end condition is met, thereby obtaining a trained video feature extraction model and a trained text feature extraction model;

[0060] Among them, when performing training operations, the training module is specifically used to:

[0061] The video features of each sample video included in the training data set are extracted by the initial video feature extraction model, and the text features of the label description text of each video label included in the training data set are extracted by the initial text feature extraction model;

[0062] For each training sample group, a loss value corresponding to the training sample group is determined based on the difference between sample features corresponding to each sample data in the training sample group; the sample features include video features or text features;

[0063] Based on the loss values ​​corresponding to each training sample group, the training loss is determined; based on the training loss, the parameters of the initial video feature extraction model and the initial text feature extraction model are adjusted, and the initial video feature extraction model and the initial text feature extraction model after the adjusted parameters are used as the initial video feature extraction model and the initial text feature extraction model corresponding to the next training operation.

[0064] In some possible implementations, when acquiring the training data set, the first acquisition module is specifically configured to:

[0065] Obtain multiple sample videos; each sample video is annotated with at least one video tag;

[0066] Deduplication is performed on the video labels corresponding to each sample video to obtain multiple video labels;

[0067] Determine the similarity between each sample video and the similarity between the label description texts corresponding to each video label;

[0068] Based on the similarity between each sample video, the similarity between the label description texts corresponding to each video label, and the correspondence between each sample video and its respective video label, multiple sample videos and multiple sample labels are divided into multiple training sample groups to obtain a training data set.

[0069] In some possible implementations, the initial video feature extraction model includes a first video extraction module and a second video extraction module; the initial text feature extraction model includes a first text extraction module and a second text extraction module;

[0070] When determining the similarity between each sample video and the similarity between the label description texts corresponding to each video label, the first acquisition module is specifically used to:

[0071] Extracting initial video features of each sample video through a first video extraction module, and determining similarities between the initial video features;

[0072] Extracting initial text features of the tag description text corresponding to each video tag by a first text extraction module, and determining the similarity between the initial text features;

[0073] The parameters of the initial video feature extraction model and the initial text feature extraction model are adjusted based on the training loss, including:

[0074] Parameters of the second video extraction module and the second text extraction module are adjusted based on the training loss.

[0075] In some possible implementations, the first acquisition module, when dividing the plurality of sample videos and the plurality of sample labels into a plurality of training sample groups based on the similarity between the respective sample videos, the similarity between the label description texts corresponding to the respective video labels, and the correspondence between each sample video and its respective video label, is specifically configured to:

[0076] Each sample video is regarded as a video node, and each video tag is regarded as a tag node. Based on the similarity between each sample video, the similarity between the tag description texts corresponding to each video tag, and the correspondence between each sample video and its respective video tag, each node is connected to generate a relationship structure diagram between multiple sample videos and multiple sample tags;

[0077] At least some of the nodes in the relationship structure graph are used as anchor points; the nodes include video nodes and tag nodes;

[0078] For each anchor point, the positive sample nodes and negative sample nodes of the anchor point are determined based on the connection relationship between each node in the relationship structure graph;

[0079] The anchor point is used as the target sample data, the positive sample node corresponding to the anchor point is used as the positive sample data, and the negative sample node corresponding to the anchor point is used as the negative sample data to obtain the training sample group corresponding to the anchor point.

[0080] In some possible implementations, the first acquisition module, when connecting each node based on the similarity between each sample video, the similarity between the label description texts corresponding to each video label, and the correspondence between each sample video and its respective video label, generates a relationship structure diagram between multiple sample videos and multiple sample labels, is specifically configured to:

[0081] Determine the similarity between each sample video, connect the video nodes whose similarity meets a first threshold, and form isomorphic edges between each video node;

[0082] Determine the similarity between the label description texts corresponding to the respective video labels, connect the label nodes whose similarity meets the second threshold, and form isomorphic edges between the label nodes;

[0083] Based on the corresponding relationship between each sample video and its respective video label, heterogeneous edges between video nodes and label nodes are formed to obtain a relationship structure graph.

[0084] In some possible implementations, for each anchor point, when the first acquisition module determines the positive sample node and the negative sample node of the anchor point based on the connection relationship between each node in the relationship structure graph, it is specifically configured to:

[0085] In the relationship structure graph, the nodes whose number of connected edges with the anchor point is less than or equal to the preset number are regarded as positive sample nodes of the anchor point;

[0086] In the relationship structure graph, nodes whose number of connected edges with the anchor point exceeds a preset number are regarded as negative sample nodes of the anchor point.

[0087] In some possible implementations, the training data set further includes at least one training sample group in which all sample data are sample videos, and at least one training sample group in which all sample data are video labels;

[0088] The device also includes a selection module for:

[0089] Among the multiple training sample groups, the training sample group whose sample data are all video labels is used as a label sample group, the training sample group whose sample data are all sample videos is used as a video sample group, and the sample group whose sample data includes sample videos and video labels is used as a mixed sample group;

[0090] When the number of training operations reaches a preset number, multiple target sample groups are selected from the multiple training sample groups based on the preset ratio of the number of label sample groups, video sample groups, and mixed sample groups;

[0091] A new training set is generated based on multiple target sample groups; and training operations after a preset number of times are performed based on the new training data set.

[0092] In some possible implementations, in the quantity ratio, the number of mixed sample groups is greater than the number of label sample groups, and the number of mixed sample groups is greater than the number of video sample groups.

[0093] In some possible implementations, when the selection module selects multiple target sample groups from multiple training sample groups based on a preset ratio of the number of label sample groups, video sample groups, and mixed sample groups, it is specifically configured to:

[0094] Determine the loss value corresponding to each training sample group when the number of training operations reaches a preset number;

[0095] The training sample group whose loss value is greater than the preset threshold is taken as the first sample group;

[0096] Based on the preset quantity ratios among the label sample group, the video sample group, and the mixed sample group, a plurality of target sample groups are selected from the plurality of first sample groups.

[0097] In some possible implementations, for each training sample group, when the training module determines the loss value corresponding to the training sample group based on the difference between the sample features corresponding to each sample data in the training sample group, it is specifically configured to:

[0098] Determining first difference information between target sample data and positive sample data in the training sample group, and determining second difference information between the target sample data and negative sample data;

[0099] A loss value of the sample group is determined based on the first difference information and the second difference information.

[0100] In some possible implementations, when extracting the initial video features of each sample video, the acquisition module is specifically configured to:

[0101] Acquire features of at least two different modalities of the sample video; the features of the different modalities include at least two of the following: picture features of a video frame of the sample video, audio features of a sample audio of the sample video, content text features of a video content text of the sample video, and title text features of a video title of the sample video;

[0102] Fusing features of at least two different modalities to obtain initial video features of the sample video;

[0103] The tag description text includes the video tag, at least two different levels of categories to which the video tag belongs, and description information of the video tag; extracting initial text features of the tag description text corresponding to each video tag includes:

[0104] The video tag, at least two categories of different levels to which the video tag belongs, and description information of the video tag are converted into corresponding text features to obtain initial text features of the video tag.

[0105] On the other hand, an embodiment of the present application provides a data retrieval device, comprising:

[0106] The second acquisition module is used to acquire target data; the target data includes target video data or target text data;

[0107] An extraction module is configured to extract video features of the target video using a trained video feature extraction model if the target data is a target video; and to extract text features of the target text using a trained text feature extraction model if the target data is a target text. The video feature extraction model and the text feature extraction model are trained based on the above-mentioned model training method.

[0108] A retrieval module is used to retrieve retrieval results corresponding to target data in a database based on the extracted video features or text features; the retrieval results include at least one of video data and label data; wherein the database stores candidate video features corresponding to a plurality of video data respectively; the database also stores candidate text features corresponding to a plurality of label data respectively; the label data includes video labels and label description texts; the candidate video features are extracted based on a video feature extraction model; the candidate text features are extracted based on a text feature extraction model.

[0109] On the other hand, an embodiment of the present application further provides an electronic device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the computer program to implement the method provided in any optional embodiment of the present application.

[0110] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method provided in any optional embodiment of the present application.

[0111] On the other hand, an embodiment of the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the method provided in any optional embodiment of the present application.

[0112] The beneficial effects of the technical solution provided by the embodiments of the present application are as follows:

[0113] The training set includes multiple training sample groups, each training sample group includes target sample data, positive sample data associated with the target sample data, and negative sample data not associated with the target sample data, the sample data includes sample videos or video labels, and the sample data of at least one training sample group includes sample videos and video labels. The initial video feature extraction model and the initial text feature extraction model are trained by the training data set, so that in each group of training sample groups, the difference between the target sample data and the positive sample data is smaller, and the difference between the target sample data and the negative sample data is larger, so that the associated video features and text features are closer, so that the video features extracted by the trained video feature extraction model and the text features extracted by the trained text feature extraction model can be aligned to the same feature space, so that when retrieving video data or text data, the retrieval can be directly performed in the feature space where the video features and text features are aligned, and there is no need for features of the same type to be used as a bridge for indirect retrieval, which can effectively improve the retrieval efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.

[0115] Figure 1 A schematic diagram of the application environment of the model training method provided in an example;

[0116] Figure 2 A flowchart of a model training method provided in an embodiment of the present application;

[0117] Figure 3 A solution for obtaining a training sample set provided as an example of this application;

[0118] Figure 4 A schematic diagram of a solution for extracting features from sample videos and video tags provided in an embodiment of the present application;

[0119] Figure 5A schematic diagram of the relationship structure provided for an example of this application;

[0120] Figure 6 A schematic diagram of the relationship structure provided for an example of this application;

[0121] Figure 7 A schematic diagram of the relationship structure provided for an example of this application;

[0122] Figure 8 A schematic diagram of a solution for obtaining initial video features provided as an example of this application;

[0123] Figure 9 A flowchart of a data retrieval method provided in an embodiment of the present application;

[0124] Figure 10 A schematic diagram of a data retrieval solution provided for an example of this application;

[0125] Figure 11 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;

[0126] Figure 12 A schematic diagram of the structure of a data retrieval device provided in an embodiment of the present application;

[0127] Figure 13 A schematic structural diagram of an electronic device applicable to an embodiment of the present application. DETAILED DESCRIPTION

[0128] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.

[0129] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to the element and the other element establishing a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or as "B", or as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" can be implemented as parameter A including A1 or A2 or A3, and can also be implemented as parameter A including at least two of the three items A1, A2, and A3.

[0130] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0131] In order to better illustrate and understand the solutions provided by the embodiments of the present application, some relevant technical terms involved in the embodiments of the present application are first introduced:

[0132] UGC: User-generated Content, content produced by users

[0133] Transformer: A network structure model suitable for processing sequential data. It uses a full-attention structure instead of traditional recurrent neural networks (RNNs) or convolutional neural networks (CNNs). Its core idea is to model the relationship between each element in the input sequence and other elements as an attention weight matrix and calculate the representation of each element through a self-attention mechanism. This structure makes the Transformer very capable of capturing long-distance dependencies in sequences, thus achieving excellent performance in NLP tasks. The Transformer model is widely used in the field of NLP, including tasks such as machine translation, text summarization, and question-answering systems. It first introduced the self-attention mechanism and surpassed traditional sequence-to-sequence (Seq2Seq) models such as LSTM and GRU in many aspects. The success of the Transformer has also spawned many variant models based on the self-attention mechanism, such as BERT and GPT, which have made significant progress in NLP tasks.

[0134] Token: A token in Transformer refers to a word or subword in a text and is the basic unit used by the model when processing text.

[0135] OCR: Optical Character Recognition, which is the technology for detecting and recognizing text content from images;

[0136] ASR: Automatic Speech Recognition, a technology that converts human speech into text;

[0137] BFS, or Breadth-First-Search, is an algorithm used for graph traversal. The main purpose of this algorithm is to find the shortest path from a starting node to a target node in a graph, or to traverse all nodes in a graph.

[0138] Isomorphic graphs are graphs in which all nodes and edges are of the same type. In such graphs, nodes and edges can be viewed as entities with the same properties. For example, in a social network, there is only one node type, "people," and only one edge type, "acquaintances." Isomorphic graphs are the most traditional and simplest graph structure, as their simple structure makes them intuitive to work with.

[0139] Heterogeneous graphs are graphs in which nodes and edges are of at least two different types. In a heterogeneous graph, nodes can represent different types of entities, such as people, items, and locations, while edges represent different types of relationships between nodes, such as purchase relationships between people and items, or birthplace relationships between people and locations. Heterogeneous graphs are more complex and flexible than homogeneous graphs because they are more widely used in real-world applications, such as social network analysis, recommender systems, bioinformatics, and knowledge graphs.

[0140] With the development of internet technology, video-related feature retrieval, including video tag recognition or searching for videos by tag, has become an important part of video content characterization. Taking video tag recognition as an example, by automatically generating tags for massive amounts of user-generated content (UGC) videos, machines can provide video content features of varying granularity for downstream content distribution chains (such as recommendation systems and content operations), improving content distribution efficiency while significantly reducing the cost of manual content review.

[0141] Taking video tag recognition as an example, current tag generation methods can be roughly divided into the following two categories from the perspective of model implementation:

[0142] 1) Classification-based methods: The usual practice is to label a batch of videos with their corresponding content labels, construct a training set to train a classification model, and then use the model to label new videos. This type of method relies heavily on manual labeling. From the perspective of labels, it can be considered that each video labeled with a certain label is a positive sample of the label, and all other videos that are not labeled with the label are its negative samples. Negative samples are generally sufficient, but generally speaking, for each label, sufficient positive samples (such as more than 50) are required to train the model so that the model can accurately predict the label. Since it is generally difficult to obtain good model effects for labels with insufficient positive samples, it is generally necessary to filter out labels with insufficient positive samples and not train them, and only train the model to support labels with sufficient samples. On the other hand, when there are new labels that need to be supported, it is generally necessary to supplement the training data in a targeted manner and retrain the model to support them.

[0143] The label list supported by the classification method is a closed set, and it is difficult to quickly adjust the range of supported labels. When the target label list changes, the model needs to be retrained, and the iteration efficiency is low.

[0144] 2) Retrieval-based methods: This method is different from the closed set form supported by classification and can usually better support the open set form.

[0145] This type of method can be further divided into two types: the first is the v2v2t (video-to-video-to-tag) method. The usual practice is to extract features from the video based on a trained encoder, process the manually annotated data into the form of <video features, label list> and write it into the retrieval library. For videos that need to be labeled, first use the same encoder to extract features, search for videos with similar content in the retrieval library, and use the corresponding label lists of the top-K (k with the highest similarity) similar videos as candidate labels. The final result is obtained through a sorting module (such as voting). For new labels, it is only necessary to collect a sufficient number of video samples containing the target label and process them into the library to support the new label; the second is the v2t (video-to-tag) method. In this type of method, the retrieval does not need to use similar videos as a bridge, and the label text features can be directly retrieved based on the video features to obtain the final result. Unlike the v2v2t method, the feature spaces of the video and the labeled text must be consistent. The usual practice is to align the video and text features through bigrams (such as contrastive learning), triplet loss, or other distance learning methods. In the new feature space learned, the associated video and text (positive sample pairs) features are close, otherwise they are far apart.

[0146] Retrieval methods can flexibly support the addition and deletion of tags. However, the v2v2t method has two main shortcomings: a. This method assumes that videos with similar content have similar tags. Therefore, it uses similar videos as a bridge to obtain the tags of similar videos as a set of candidate tags. However, it also requires the design of a ranking module (such as voting based on tag frequency), which is not an end-to-end fast solution. b. New tags still require the labeling of a certain number of samples, which must be processed and then entered into the retrieval library before they can be supported. In contrast, the v2t method, when learning the shared feature space, focuses more on the relationship between video features and text features, often ignoring the isomorphic relationship between video and text features in the original feature space.

[0147] The model training method, data retrieval method, device, electronic device, computer-readable storage medium and computer program product provided in the present disclosure are intended to solve at least one of the above technical problems in the prior art.

[0148] The model training method of the present application can be implemented based on machine learning (ML) in artificial intelligence (AI).

[0149] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0150] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, large-scale model training, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, model training, natural language processing, and machine learning / deep learning.

[0151] Key technologies in speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction. Large model technology is revolutionizing the development of speech technology. Pre-trained models such as WavLM and UniSpeech, which leverage the Transformer architecture, possess strong generalization and versatility, making them highly suitable for a wide range of model training tasks.

[0152] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0153] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0154] The solutions provided in the embodiments of this application involve technologies such as artificial intelligence model training, which are specifically illustrated by the following embodiments.

[0155] The following describes several optional embodiments to illustrate the technical solutions provided by this application and the technical effects produced by the technical solutions of this application. It should be noted that the following embodiments can refer to, draw on, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.

[0156] In the specific implementation of this application, any data related to an object, such as video data, tag information, etc., when the embodiments of this application are applied to specific products or technologies, the permission or consent of the object must be obtained, and the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if any of the above-mentioned data related to an object is involved in the embodiments of this application, such data must be obtained with the authorization and consent of the object and in compliance with the relevant laws, regulations, and standards of the relevant countries and regions.

[0157] The model training method provided in the embodiment of the present application can be executed by any computer device, and optionally, can be executed by a server, wherein the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0158] Figure 1Schematic diagram of the application environment of the model training method provided in the embodiment of the present application. The application environment may include a terminal 101 and a server 102. The server 102 obtains a training data set, and performs at least one training operation on the initial video feature extraction model and the initial text feature extraction model through the training data set until the training end conditions are met, thereby obtaining a trained video feature extraction model and a text feature extraction model. The terminal 101 sends the target data to the server 102. If the target data is a target video, the server 102 extracts the video features of the target video through the trained video feature extraction model; if the target data is a target text, the server 102 extracts the text features of the target text through the trained text feature extraction model; based on the extracted video features or text features, the server 102 retrieves the retrieval results corresponding to the target data in the database, and returns the retrieval results to the terminal 101.

[0159] In the above application scenario, the server trains the text feature extraction model and the video feature extraction model. In other application scenarios, the terminal may train the text feature extraction model and the video feature extraction model, which is not limited to this.

[0160] Those skilled in the art will appreciate that a server may be an independent physical server, or a server cluster or distributed system consisting of multiple physical servers, or a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal may be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a laptop computer, a digital broadcast receiver, a MID (Mobile Internet Devices), a PDA (Personal Digital Assistant), a desktop computer, a smart home appliance, a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal, a vehicle-mounted computer, etc.), a smart speaker, a smart watch, etc. The terminal and the server may be directly or indirectly connected via wired or wireless communication, but are not limited thereto. The embodiments of the present invention may be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. The specific requirements may also be determined based on the actual application scenario and are not limited here.

[0161] The terminal (also referred to as a user terminal or user device) can be a smartphone, tablet computer, laptop computer, desktop computer, intelligent voice interaction device (such as a smart speaker), wearable electronic device (such as a smart watch), vehicle-mounted terminal, smart home appliance (such as a smart TV), AR / VR device, aircraft, etc., but is not limited to these. The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc.

[0162] The model training method provided in the embodiments of the present application can be executed by any electronic device.

[0163] Figure 2 The following is a flow chart of a model training method provided by an embodiment of the present application. Taking the execution subject as a server as an example, the model training method provided by the present application may include the following steps:

[0164] Step S201: Obtain a training data set.

[0165] The training data set includes multiple training sample groups; each training sample group includes multiple sample data; the sample data includes target sample data, positive sample data associated with the target sample data, and negative sample data not associated with the target sample data.

[0166] That is to say, the training sample group includes three sample data, one target sample data can be understood as an anchor point, one is a positive sample data set relative to the target sample data, and the last one is a negative sample data set relative to the target sample data.

[0167] In the specific implementation process, Figure 3 As shown in the figure, the square represents the sample video and the circle represents the video label. You can first obtain multiple sample videos, each sample video corresponds to at least one video label, and then use the multiple sample videos and multiple video labels as sample data, and then divide the multiple sample data into multiple training sample groups, and one training sample group includes three sample data.

[0168] The sample data includes sample videos or video labels; each video label has a corresponding label description text; and the sample data of at least one training sample group includes sample videos and video labels.

[0169] Specifically, the video tags may include tags associated with the video content, theme emotion, video type, objects included in the sample video, and the like.

[0170] Specifically, the label description text is text content generated for the label, and the corresponding content can be generated based on the label through a pre-trained text generation model.

[0171] For example, the video tags of a sample video include two video tags: dog and sofa. For "dog", the corresponding tag description may include: dog is a common canine mammal that is often kept at home; for "sofa", the corresponding tag description may include: sofa is a kind of soft furniture, which is a multi-seater chair with upholstered cushions.

[0172] That is to say, there may be such sample data: in a group of sample data, three sample data are sample videos; or three sample data are video labels; however, there is at least one group of sample data in which both sample videos and video labels exist.

[0173] Step S202: Perform at least one training operation on the initial video feature extraction model and the initial text feature extraction model using the training data set until a training end condition is met, thereby obtaining a trained video feature extraction model and a trained text feature extraction model.

[0174] The training operations include:

[0175] (1) The video features of each sample video included in the training data set are extracted by the initial video feature extraction model, and the text features of the label description text of each video label included in the training data set are extracted by the initial text feature extraction model.

[0176] Specifically, the initial video feature extraction model includes a first video extraction module and a second video extraction module; the initial text feature extraction model includes a first text extraction module and a second text extraction module. During the training process, only the parameters of the second video extraction module and the second text extraction module need to be adjusted. That is to say, the first video extraction module and the first text extraction module are preset models.

[0177] (2) For each training sample group, the loss value corresponding to the training sample group is determined based on the differences between the sample features corresponding to each sample data in the training sample group.

[0178] The sample features include video features or text features.

[0179] Specifically, the difference between the target sample data and the positive sample data is calculated, and the difference between the target sample data and the negative sample data is calculated, and then the loss value is determined.

[0180] (3) Determine the training loss based on the loss values ​​corresponding to each training sample group; adjust the parameters of the initial video feature extraction model and the initial text feature extraction model based on the training loss, and use the initial video feature extraction model and the initial text feature extraction model after adjusting the parameters as the initial video feature extraction model and the initial text feature extraction model corresponding to the next training operation.

[0181] Specifically, the parameters of the initial video feature extraction model and the initial text feature extraction model are adjusted based on the training loss, that is, the parameters of the second video extraction module and the second text extraction module are adjusted so that in each group of training samples, the difference between the target sample data and the positive sample data is smaller, and the difference between the target sample data and the negative sample data is larger, thereby making the associated video features and text features closer, so that the video features and text features can be aligned to the same feature space.

[0182] In the above embodiment, the training set includes multiple training sample groups, each training sample group includes target sample data, positive sample data associated with the target sample data, and negative sample data not associated with the target sample data, the sample data includes sample videos or video labels, and the sample data of at least one training sample group includes sample videos and video labels. The initial video feature extraction model and the initial text feature extraction model are trained by the training data set, so that in each group of training sample groups, the difference between the target sample data and the positive sample data is smaller, and the difference between the target sample data and the negative sample data is larger, so that the associated video features and text features are closer, so that the video features extracted by the trained video feature extraction model and the text features extracted by the trained text feature extraction model can be aligned to the same feature space. In this way, when retrieving video data or text data, the retrieval can be directly performed in the feature space where the video features and text features are aligned, and there is no need for features of the same type to be used as a bridge for indirect retrieval, which can effectively improve the retrieval efficiency.

[0183] In some possible implementations, step S201 of obtaining a training data set may include:

[0184] (1) Obtain multiple sample videos.

[0185] Each sample video is annotated with at least one video tag.

[0186] (2) De-duplication is performed on the video labels corresponding to each sample video to obtain multiple video labels.

[0187] Specifically, different sample videos may have the same video tag. The video tags corresponding to multiple different videos may be deduplicated to obtain multiple different video tags.

[0188] (3) Determine the similarity between each sample video and the similarity between the label description texts corresponding to each video label.

[0189] Specifically, determining the similarity between sample videos and determining the similarity between tag description texts corresponding to video tags may include:

[0190] Extracting initial video features of each sample video through a first video extraction module, and determining similarities between the initial video features;

[0191] The first text extraction module extracts initial text features of the tag description texts corresponding to the respective video tags, and determines the similarity between the initial text features.

[0192] like Figure 4 As shown, a first video extraction module and a first text extraction module, which do not require retraining, respectively extract initial text features and initial video features. Based on the initial video features, the similarity between each initial video feature is determined, and based on the initial text features, the similarity between each initial text feature is determined. A second video extraction module then performs feature extraction on the initial video features extracted by the first video extraction module to obtain video features; a second text extraction module then performs feature extraction on the initial text features extracted by the first text extraction module to obtain text features. A loss value is then determined based on the video and text feature pairs, and the parameters of the second video extraction module and the second text extraction module are adjusted.

[0193] (4) Based on the similarity between each sample video, the similarity between the label description texts corresponding to each video label, and the correspondence between each sample video and its respective video label, multiple sample videos and multiple sample labels are divided into multiple training sample groups to obtain a training data set.

[0194] Specifically, for each target sample data, the positive sample data and negative sample data of the target sample data can be determined by the similarity between each sample video, the similarity between the label description texts corresponding to each video label, and the correspondence between each sample video and its respective video label.

[0195] In some possible implementations, based on the similarity between the sample videos, the similarity between the label description texts corresponding to the video labels, and the correspondence between each sample video and its respective video label, the plurality of sample videos and the plurality of sample labels are divided into a plurality of training sample groups, which may include:

[0196] a. Treat each sample video as a video node and each video tag as a tag node. Based on the similarity between the sample videos, the similarity between the tag description texts corresponding to the video tags, and the correspondence between each sample video and its respective video tag, connect the nodes to generate a relationship structure diagram between multiple sample videos and multiple sample tags.

[0197] Specifically, in some possible implementations, based on the similarity between sample videos, the similarity between the label description texts corresponding to the video labels, and the corresponding relationship between each sample video and its respective video label, each node is connected to generate a relationship structure diagram between multiple sample videos and multiple sample labels, which may include:

[0198] a1. Determine the similarity between each sample video, connect the video nodes whose similarity meets a first threshold, and form isomorphic edges between each video node;

[0199] a2. Determine the similarity between the label description texts corresponding to the respective video labels, and connect the label nodes whose similarity meets the second threshold to form isomorphic edges between the label nodes;

[0200] a3. Based on the corresponding relationship between each sample video and its respective video label, heterogeneous edges are formed between the video node and the label node to obtain a relationship structure graph.

[0201] Among them, if the first threshold is met, that is, greater than or equal to the first threshold, it is considered that the two video nodes can be connected; when a video tag is a tag marked by a sample video, the video node and the video tag can be connected.

[0202] Specifically, such as Figure 5 As shown, diamond-shaped nodes represent video nodes, and circular nodes represent label nodes. Based on the initial text features and initial video features extracted by the first video feature extraction module and the first text feature extraction module, each node can find its k-nearest neighbors. Similar nodes are connected to form isomorphic edges, as shown by the solid lines in the figure. Based on the correspondence between each sample video and the annotated video label, the edges connecting the video nodes to the label nodes shown in the figure are formed, forming heterogeneous edges, as shown by the dotted lines in the figure.

[0203] b. Use at least some of the nodes in the relationship structure diagram as anchor points.

[0204] Among them, the nodes include video nodes and tag nodes.

[0205] That is to say, the video node can be used as an anchor point, or the label node can be used as an anchor point.

[0206] c. For each anchor point, determine the positive sample nodes and negative sample nodes of the anchor point based on the connection relationship between each node in the relationship structure graph.

[0207] Specifically, for each anchor point, based on the connection relationship between each node in the relationship structure graph, the positive sample node and the negative sample node of the anchor point are determined, including:

[0208] c1. Nodes in the relationship structure graph whose number of connected edges to the anchor point is less than or equal to a preset number are selected as positive sample nodes of the anchor point;

[0209] c2. Nodes in the relationship structure graph whose number of connected edges with the anchor point exceeds the preset number are regarded as negative sample nodes of the anchor point.

[0210] According to the neighbor relationship of the relationship structure graph, more similar and dissimilar sample relationships can be mined, which greatly expands the training samples. For example, Figure 6 As shown in the figure, based on the video labeling results, x and y are similar, and based on the isomorphism relationship, y and z are similar, so we can assume that x and z are also similar. However, t, which is farther away in the figure, is obviously not very similar to x. The definition of a triple (training sample group) can be expressed as follows:

[0211] Taking x as the anchor point, z is closer to x than t is to t.

[0212] This application can generate triples based on the BFS method, such as Figure 7 As shown:

[0213] Starting from node s, do a BFS step and get {x_(1,*)}.

[0214] Do another BFS on {x_(1,*)} and get {x_(2,*)}.

[0215] BFS stops until the number of steps reaches the preset number of edges (steps) t. At this time, for the starting point s, we get a set P = {{x_(1,*)},{x_(2,*)},…,{x_(t,*)}}, which is all the nodes that can be reached by BFS t times (t = 2 in the figure, all nodes in the P set are marked in red), which are also positive sample nodes; the set of other nodes on the heterogeneous graph is recorded as N, which is also the negative sample nodes.

[0216] At this time, it can be considered that for the anchor point s, any node i in P and any node j in N satisfy: the feature distance between any node in P and the anchor point is less than the feature distance between any node in N and the anchor point, that is, the distance between the anchor point and any positive sample node is less than the distance between the anchor point and any negative sample node.

[0217] d. Take the anchor point as the target sample data, the positive sample node corresponding to the anchor point as the positive sample data, and the negative sample node corresponding to the anchor point as the negative sample data to obtain the training sample group corresponding to the anchor point.

[0218] Specifically, the anchor point, positive sample node, and negative sample node are determined according to the relationship structure diagram, and every three nodes form a triplet, and the data corresponding to the triplet is the corresponding training sample group.

[0219] In the above embodiment, by taking each sample video as a video node and each video tag as a tag node, the nodes are connected based on the similarity between the sample videos, the similarity between the tag description texts corresponding to the video tags, and the correspondence between each sample video and its respective video tag, to generate a relationship structure diagram between multiple sample videos and multiple sample tags, and then at least some of the nodes in the relationship structure diagram are used as anchor points. For each anchor point, the positive sample node and the negative sample node of the anchor point are determined based on the connection relationship between the nodes in the relationship structure diagram. A group of training sample groups are generated based on the anchor points, positive sample nodes and negative sample nodes. This can help to discover the structural relationship between more sample videos and video tags, effectively reduce the amount of information that needs to be labeled in the training stage, and improve training efficiency.

[0220] In some possible implementations, the training data set further includes at least one training sample group in which all sample data are sample videos, and at least one training sample group in which all sample data are video labels.

[0221] The method also includes:

[0222] Among the multiple training sample groups, the training sample group whose sample data are all video labels is used as a label sample group, the training sample group whose sample data are all sample videos is used as a video sample group, and the sample group whose sample data includes sample videos and video labels is used as a mixed sample group;

[0223] When the number of training operations reaches a preset number, multiple target sample groups are selected from the multiple training sample groups based on the preset ratio of the number of label sample groups, video sample groups, and mixed sample groups;

[0224] A new training set is generated based on multiple target sample groups; and training operations after a preset number of times are performed based on the new training data set.

[0225] That is to say, during the training process, the training data set includes a video sample group, a label sample group, and a mixed sample group.

[0226] Specifically, when the number of training times reaches a preset number but the training end condition has not been reached, the number of different types of training sample groups in the training sample group may be adjusted.

[0227] In some possible implementations, in the quantity ratio, the number of mixed sample groups is greater than the number of label sample groups, and the number of mixed sample groups is greater than the number of video sample groups.

[0228] In the above embodiment, when the number of training operations reaches a preset number, multiple target sample groups are selected from multiple training sample groups based on the preset ratio between the number of label sample groups, video sample groups and mixed sample groups, and in the number ratio, the number of mixed sample groups is greater than the number of label sample groups, and the number of mixed sample groups is greater than the number of video sample groups. The training operations after the preset number of times are based on a new training data set, which can enable the model to use more mixed sample groups for training, thereby enhancing the ability of the trained model to align different types of sample data into a feature space.

[0229] In some possible implementations, based on a preset ratio of the number of label sample groups, video sample groups, and mixed sample groups, multiple target sample groups are selected from multiple training sample groups, including:

[0230] Determine the loss value corresponding to each training sample group when the number of training operations reaches a preset number;

[0231] The training sample group whose loss value is greater than the preset threshold is taken as the first sample group;

[0232] Based on the preset quantity ratios among the label sample group, the video sample group, and the mixed sample group, a plurality of target sample groups are selected from the plurality of first sample groups.

[0233] Specifically, the training sample group with a loss value greater than a preset threshold, that is, the first sample group, can be understood as a difficult sample group, which is equivalent to coarse-tuning the model in the previous preset number of training operations; when the number of training operations reaches the preset number, difficult samples are selected again, and based on the quantity ratio between the label sample group, the video sample group and the mixed sample group, the ability of the trained model to align different types of sample data into a feature space can be further enhanced.

[0234] In some possible implementations, for each training sample group, determining a loss value corresponding to the training sample group based on differences between sample features corresponding to each sample data in the training sample group includes:

[0235] Determining first difference information between target sample data and positive sample data in the training sample group, and determining second difference information between the target sample data and negative sample data;

[0236] A loss value of the sample group is determined based on the first difference information and the second difference information.

[0237] Specifically, the following formula can be used for calculation:

[0238] L triplet =max (0, m+d(A,P)-d(A,N)) (1)

[0239] Among them, d(A,P) represents the first difference information between the target sample data and the positive sample data; d(A,N) represents the second difference information between the target sample data and the negative sample data; m is a margin value, which is set to prevent the distance between the negative sample and the anchor sample from being too close.

[0240] In some possible implementations, extracting initial video features of each sample video includes:

[0241] (1) Obtain features of at least two different modalities of the sample video.

[0242] (2) Fuse the features of at least two different modalities to obtain the initial video features of the sample video.

[0243] Among them, the features of different modalities include at least two of the following: picture features of the video frame of the sample video, audio features of the sample audio of the sample video, content text features of the video content text of the sample video, and title text features of the video title of the sample video.

[0244] Specifically, the feature encoder of multimodal feature fusion can be pre-trained. Figure 8 As shown in the figure, the structure of the multimodal model can be that each modality has a separate encoder, and the modal features are extracted and fused separately (late-fusion form); or it can be that non-text modalities such as video frames are converted into tokens after passing through the feature encoder, and then fused and predicted together with the text tokens through the transformer (early-fusion form).

[0245] In some possible implementations, the tag description text includes the video tag, at least two different levels of categories to which the video tag belongs, and description information of the video tag.

[0246] As shown in Table 1 below:

[0247]

[0248]

[0249]

[0250] Extracting initial text features of the tag description text corresponding to each video tag may include:

[0251] The video tag, at least two categories of different levels to which the video tag belongs, and description information of the video tag are converted into corresponding text features to obtain initial text features of the video tag.

[0252] Therefore, the description text of the tag can be constructed using the following template:

[0253] [Label Chinese Name]: Belongs to the [Label Secondary Category] type under the [Label First Category] category, meaning: [Label Description Information]

[0254] For example, the description of "Interesting Experiment" is:

[0255] Interesting experiments: belong to the "Highlights - Fun" type under the "Life" category, which means: interesting experiments and challenging content without scientific experimental content, as long as they are interesting (this label does not apply to normal scientific experiments). If the video experiment project is interesting and has a big imagination and a strong curiosity, it can be included in the "Bizarre Experiments".

[0256] Each tag can generate a text description similar to this one. After the text is tokenized, it is input into a trained text model, such as BERT (Bidirectional Encoder Representations from Transformers), to obtain the corresponding initial text features.

[0257] In the above-mentioned model training method, the training set includes multiple training sample groups, each training sample group includes target sample data, positive sample data associated with the target sample data, and negative sample data not associated with the target sample data, the sample data includes sample video or video label, and the sample data of at least one training sample group includes sample video and video label. The initial video feature extraction model and the initial text feature extraction model are trained by the training data set, so that in each group of training sample groups, the difference between the target sample data and the positive sample data is smaller, and the difference between the target sample data and the negative sample data is larger, so that the associated video features and text features are closer, so that the video features extracted by the trained video feature extraction model and the text features extracted by the trained text feature extraction model can be aligned to the same feature space. In this way, when retrieving video data or text data, retrieval can be performed directly in the feature space where the video features and text features are aligned, and there is no need for features of the same type to be used as a bridge for indirect retrieval, which can effectively improve the retrieval efficiency.

[0258] Furthermore, by treating each sample video as a video node and each video tag as a tag node, the nodes are connected based on the similarity between the sample videos, the similarity between the tag description texts corresponding to the video tags, and the correspondence between each sample video and its respective video tag, to generate a relationship structure diagram between multiple sample videos and multiple sample tags, and then at least some of the nodes in the relationship structure diagram are used as anchor points. For each anchor point, the positive sample node and negative sample node of the anchor point are determined based on the connection relationship between the nodes in the relationship structure diagram, and a group of training sample groups are generated based on the anchor points, positive sample nodes and negative sample nodes. This can help to discover the structural relationship between more sample videos and video tags, effectively reduce the amount of information that needs to be labeled in the training stage, and improve training efficiency.

[0259] Furthermore, when the number of training operations reaches a preset number, multiple target sample groups are selected from multiple training sample groups based on the preset ratio between the number of label sample groups, video sample groups and mixed sample groups, and in the number ratio, the number of mixed sample groups is greater than the number of label sample groups, and the number of mixed sample groups is greater than the number of video sample groups. The training operations after the preset number of times are based on a new training data set, which can enable the model to use more mixed sample groups for training, thereby enhancing the ability of the trained model to align different types of sample data into a feature space.

[0260] Furthermore, the training sample group with a loss value greater than a preset threshold, that is, the first sample group, is understood as a difficult sample group, which is equivalent to coarse-tuning the model in the previous preset number of training operations; when the number of training operations reaches the preset number, difficult samples are selected again, and based on the quantity ratio between the label sample group, the video sample group and the mixed sample group, the ability of the trained model to align different types of sample data into a feature space can be further enhanced.

[0261] like Figure 9 As shown, in some possible implementations, a data retrieval method is provided, the method comprising:

[0262] Step S901: Acquire target data.

[0263] The target data includes target video data or target text data.

[0264] That is to say, you can enter a video to search for the corresponding video and tag, or you can enter text to search for the corresponding video and tag.

[0265] Step S902: If the target data is a target video, the video features of the target video are extracted using the trained video feature extraction model; if the target data is a target text, the text features of the target text are extracted using the trained text feature extraction model.

[0266] Among them, the video feature extraction model and the text feature extraction model are trained based on the model training method of the above embodiment.

[0267] That is to say, through the above-mentioned trained video feature extraction model and text feature extraction model, the video features and text features can be aligned to the same feature space, and then the corresponding text features or video features can be directly retrieved in the same feature space.

[0268] Step S903: searching the database for a search result corresponding to the target data based on the extracted video features or text features.

[0269] The search result includes at least one of video data and tag data.

[0270] In other words, you can search only for data of the same type, only for data of different types, or both.

[0271] For example, if the input target data is video data, the corresponding tag data can be directly retrieved from the database, or similar video data can be retrieved, or both can be retrieved.

[0272] Among them, the database stores candidate video features corresponding to multiple video data respectively; the database also stores candidate text features corresponding to multiple label data respectively; the label data includes video labels and label description texts; the candidate video features are extracted based on the video feature extraction model; the candidate text features are extracted based on the text feature extraction model.

[0273] In the above embodiment, the video features extracted by the trained video feature extraction model and the text features extracted by the trained text feature extraction model can be aligned to the same feature space. In this way, when searching video data or text data, the search can be performed directly in the feature space where the video features and text features are aligned. There is no need for features of the same type to serve as a bridge for indirect search, which can effectively improve the search efficiency.

[0274] In this application, after the model training is completed, for each tag in the tag library, it is only necessary to generate the corresponding tag description text, use the trained text feature extraction model to extract the text features, and import them into the database.

[0275] Similarly, for the target video to be identified, the video features are extracted by the above-mentioned video feature extraction module, and similar features are searched in the database in a feature retrieval manner to obtain the corresponding label features of the video, thereby obtaining the corresponding label description text.

[0276] Tags in the database are explicitly encoded as feature vectors and stored in the database. Therefore, for tags that need to be removed from the database, the corresponding features can be simply deleted. For tags that are added to the database, description text is generated, features are extracted, and the tags are added to the search database for validation. This eliminates the need to annotate data and update the model, allowing for flexible and efficient support of new tags.

[0277] Take the target data as video data as an example, Figure 10 As shown, the following will illustrate the above data retrieval method with examples:

[0278] Obtain multiple candidate tags and generate tag description text corresponding to each candidate tag;

[0279] Through the pre-trained text feature extraction model, feature extraction is performed on the label description texts corresponding to the multiple candidate labels to obtain the candidate text features corresponding to the multiple candidate labels;

[0280] Storing candidate text features corresponding to multiple candidate tags into a tag library;

[0281] Receive the target video to be retrieved;

[0282] Extract target video features of the target video through the trained video feature extraction model;

[0283] Retrieve target text features similar to target video features from multiple candidate text features;

[0284] The label description text corresponding to the target text feature is used as the label result of the target video.

[0285] In the above-mentioned data retrieval method, the video features extracted by the trained video feature extraction model and the text features extracted by the trained text feature extraction model can be aligned to the same feature space. In this way, when retrieving video data or text data, the retrieval can be performed directly in the feature space where the video features and text features are aligned. There is no need for features of the same type to serve as a bridge for indirect retrieval, which can effectively improve the retrieval efficiency.

[0286] like Figure 11 As shown, in some possible implementations, a model training device is provided, including:

[0287] The first acquisition module 1101 is configured to acquire a training data set; the training data set includes a plurality of training sample groups; each training sample group includes a plurality of sample data; the sample data includes target sample data, positive sample data associated with the target sample data, and negative sample data not associated with the target sample data; the sample data includes sample videos or video labels; each video label has a corresponding label description text; the sample data of at least one training sample group includes a sample video and a video label;

[0288] A training module 1102 is configured to perform at least one training operation on the initial video feature extraction model and the initial text feature extraction model using a training data set until a training end condition is met, thereby obtaining a trained video feature extraction model and a trained text feature extraction model;

[0289] The training module 1102, when performing the training operation, is specifically configured to:

[0290] The video features of each sample video included in the training data set are extracted by the initial video feature extraction model, and the text features of the label description text of each video label included in the training data set are extracted by the initial text feature extraction model;

[0291] For each training sample group, a loss value corresponding to the training sample group is determined based on the difference between sample features corresponding to each sample data in the training sample group; the sample features include video features or text features;

[0292] Based on the loss values ​​corresponding to each training sample group, the training loss is determined; based on the training loss, the parameters of the initial video feature extraction model and the initial text feature extraction model are adjusted, and the initial video feature extraction model and the initial text feature extraction model after the adjusted parameters are used as the initial video feature extraction model and the initial text feature extraction model corresponding to the next training operation.

[0293] In some possible implementations, when acquiring the training data set, the first acquisition module 1101 is specifically configured to:

[0294] Obtain multiple sample videos; each sample video is annotated with at least one video tag;

[0295] Deduplication is performed on the video labels corresponding to each sample video to obtain multiple video labels;

[0296] Determine the similarity between each sample video and the similarity between the label description texts corresponding to each video label;

[0297] Based on the similarity between each sample video, the similarity between the label description texts corresponding to each video label, and the correspondence between each sample video and its respective video label, multiple sample videos and multiple sample labels are divided into multiple training sample groups to obtain a training data set.

[0298] In some possible implementations, the initial video feature extraction model includes a first video extraction module and a second video extraction module; the initial text feature extraction model includes a first text extraction module and a second text extraction module;

[0299] When determining the similarity between sample videos and the similarity between tag description texts corresponding to video tags, the first acquisition module 1101 is specifically configured to:

[0300] Extracting initial video features of each sample video through a first video extraction module, and determining similarities between the initial video features;

[0301] Extracting initial text features of the tag description text corresponding to each video tag by a first text extraction module, and determining the similarity between the initial text features;

[0302] The training module 1102 adjusts the parameters of the initial video feature extraction model and the initial text feature extraction model based on the training loss, including:

[0303] Parameters of the second video extraction module and the second text extraction module are adjusted based on the training loss.

[0304] In some possible implementations, the first acquisition module 1101 is specifically configured to:

[0305] Each sample video is regarded as a video node, and each video tag is regarded as a tag node. Based on the similarity between each sample video, the similarity between the tag description texts corresponding to each video tag, and the correspondence between each sample video and its respective video tag, each node is connected to generate a relationship structure diagram between multiple sample videos and multiple sample tags;

[0306] At least some of the nodes in the relationship structure graph are used as anchor points; the nodes include video nodes and tag nodes;

[0307] For each anchor point, the positive sample nodes and negative sample nodes of the anchor point are determined based on the connection relationship between each node in the relationship structure graph;

[0308] The anchor point is used as the target sample data, the positive sample node corresponding to the anchor point is used as the positive sample data, and the negative sample node corresponding to the anchor point is used as the negative sample data to obtain the training sample group corresponding to the anchor point.

[0309] In some possible implementations, the first acquisition module 1101 connects the nodes based on the similarity between the sample videos, the similarity between the label description texts corresponding to the video labels, and the correspondence between each sample video and its respective video label to generate a relationship structure diagram between multiple sample videos and multiple sample labels, specifically for:

[0310] Determine the similarity between each sample video, connect the video nodes whose similarity meets a first threshold, and form isomorphic edges between each video node;

[0311] Determine the similarity between the label description texts corresponding to the respective video labels, connect the label nodes whose similarity meets the second threshold, and form isomorphic edges between the label nodes;

[0312] Based on the corresponding relationship between each sample video and its respective video label, heterogeneous edges between video nodes and label nodes are formed to obtain a relationship structure graph.

[0313] In some possible implementations, for each anchor point, when the first acquisition module 1101 determines the positive sample node and the negative sample node of the anchor point based on the connection relationship between each node in the relationship structure graph, it is specifically configured to:

[0314] In the relationship structure graph, the nodes whose number of connected edges with the anchor point is less than or equal to the preset number are regarded as positive sample nodes of the anchor point;

[0315] In the relationship structure graph, nodes whose number of connected edges with the anchor point exceeds a preset number are regarded as negative sample nodes of the anchor point.

[0316] In some possible implementations, the training data set further includes at least one training sample group in which all sample data are sample videos, and at least one training sample group in which all sample data are video labels;

[0317] The device also includes a selection module for:

[0318] Among the multiple training sample groups, the training sample group whose sample data are all video labels is used as a label sample group, the training sample group whose sample data are all sample videos is used as a video sample group, and the sample group whose sample data includes sample videos and video labels is used as a mixed sample group;

[0319] When the number of training operations reaches a preset number, multiple target sample groups are selected from the multiple training sample groups based on the preset ratio of the number of label sample groups, video sample groups, and mixed sample groups;

[0320] A new training set is generated based on multiple target sample groups; and training operations after a preset number of times are performed based on the new training data set.

[0321] In some possible implementations, when the selection module selects multiple target sample groups from multiple training sample groups based on a preset ratio of the number of label sample groups, video sample groups, and mixed sample groups, it is specifically configured to:

[0322] Determine the loss value corresponding to each training sample group when the number of training operations reaches a preset number;

[0323] The training sample group whose loss value is greater than the preset threshold is taken as the first sample group;

[0324] Based on the preset quantity ratios among the label sample group, the video sample group, and the mixed sample group, a plurality of target sample groups are selected from the plurality of first sample groups.

[0325] In some possible implementations, for each training sample group, when the training module 1102 determines the loss value corresponding to the training sample group based on the difference between the sample features corresponding to each sample data in the training sample group, it is specifically configured to:

[0326] Determining first difference information between target sample data and positive sample data in the training sample group, and determining second difference information between the target sample data and negative sample data;

[0327] A loss value of the sample group is determined based on the first difference information and the second difference information.

[0328] In some possible implementations, when extracting the initial video features of each sample video, the first acquisition module 1101 is specifically configured to:

[0329] Acquire features of at least two different modalities of the sample video; the features of the different modalities include at least two of the following: picture features of a video frame of the sample video, audio features of a sample audio of the sample video, content text features of a video content text of the sample video, and title text features of a video title of the sample video;

[0330] Fusing features of at least two different modalities to obtain initial video features of the sample video;

[0331] The tag description text includes the video tag, at least two different levels of categories to which the video tag belongs, and description information of the video tag; extracting initial text features of the tag description text corresponding to each video tag includes:

[0332] The video tag, at least two categories of different levels to which the video tag belongs, and description information of the video tag are converted into corresponding text features to obtain initial text features of the video tag.

[0333] The above-mentioned model training device, the training set includes multiple training sample groups, each training sample group includes target sample data, positive sample data associated with the target sample data and negative sample data not associated with the target sample data, the sample data includes sample video or video label, and the sample data of at least one training sample group includes sample video and video label. The initial video feature extraction model and the initial text feature extraction model are trained by the training data set, so that in each group of training sample groups, the difference between the target sample data and the positive sample data is smaller, and the difference between the target sample data and the negative sample data is larger, so that the associated video features and text features are closer, so that the video features extracted by the trained video feature extraction model and the text features extracted by the trained text feature extraction model can be aligned to the same feature space. In this way, when retrieving video data or text data, retrieval can be performed directly in the feature space where the video features and text features are aligned, and there is no need for features of the same type to be used as a bridge for indirect retrieval, which can effectively improve the retrieval efficiency.

[0334] Furthermore, by treating each sample video as a video node and each video tag as a tag node, the nodes are connected based on the similarity between the sample videos, the similarity between the tag description texts corresponding to the video tags, and the correspondence between each sample video and its respective video tag, to generate a relationship structure diagram between multiple sample videos and multiple sample tags, and then at least some of the nodes in the relationship structure diagram are used as anchor points. For each anchor point, the positive sample node and negative sample node of the anchor point are determined based on the connection relationship between the nodes in the relationship structure diagram, and a group of training sample groups are generated based on the anchor points, positive sample nodes and negative sample nodes. This can help to discover the structural relationship between more sample videos and video tags, effectively reduce the amount of information that needs to be labeled in the training stage, and improve training efficiency.

[0335] Furthermore, when the number of training operations reaches a preset number, multiple target sample groups are selected from multiple training sample groups based on the preset ratio between the number of label sample groups, video sample groups and mixed sample groups, and in the number ratio, the number of mixed sample groups is greater than the number of label sample groups, and the number of mixed sample groups is greater than the number of video sample groups. The training operations after the preset number of times are based on a new training data set, which can enable the model to use more mixed sample groups for training, thereby enhancing the ability of the trained model to align different types of sample data into a feature space.

[0336] Furthermore, the training sample group with a loss value greater than a preset threshold, that is, the first sample group, is understood as a difficult sample group, which is equivalent to coarse-tuning the model in the previous preset number of training operations; when the number of training operations reaches the preset number, difficult samples are selected again, and based on the quantity ratio between the label sample group, the video sample group and the mixed sample group, the ability of the trained model to align different types of sample data into a feature space can be further enhanced.

[0337] In some possible implementations, such as Figure 12 , an embodiment of the present application provides a data retrieval device, the device comprising:

[0338] The second acquisition module 1201 is used to acquire target data; the target data includes target video data or target text data;

[0339] Extraction module 1202, configured to extract video features of the target video using a trained video feature extraction model if the target data is a target video; and extract text features of the target text using a trained text feature extraction model if the target data is a target text; the video feature extraction model and the text feature extraction model are trained based on the above-mentioned model training method;

[0340] The retrieval module 1203 is used to retrieve retrieval results corresponding to the target data in the database based on the extracted video features or text features; the retrieval results include at least one of video data and label data; wherein the database stores candidate video features corresponding to a plurality of video data respectively; the database also stores candidate text features corresponding to a plurality of label data respectively; the label data includes video labels and label description texts; the candidate video features are extracted based on the video feature extraction model; the candidate text features are extracted based on the text feature extraction model.

[0341] In the above-mentioned data retrieval method, the video features extracted by the trained video feature extraction model and the text features extracted by the trained text feature extraction model can be aligned to the same feature space. In this way, when retrieving video data or text data, the retrieval can be performed directly in the feature space where the video features and text features are aligned. There is no need for features of the same type to serve as a bridge for indirect retrieval, which can effectively improve the retrieval efficiency.

[0342] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, and will not be repeated here.

[0343] In an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory. When the processor executes the computer program stored in the memory, the method of any optional embodiment of the present application can be implemented. Compared with the prior art, the method can directly search in the feature space where video features and text features are aligned, eliminating the need for indirect search using features of the same type, thereby effectively improving search efficiency.

[0344] In an alternative embodiment, an electronic device is provided, such as Figure 13 As shown, Figure 13 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure.

[0345] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0346] Bus 4002 may include a path for transmitting information between the above components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 13 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0347] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation here.

[0348] The memory 4003 is used to store the computer program for executing the embodiments of the present disclosure, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiments.

[0349] Based on the same principle as the method provided in the embodiment of the present application, the embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the corresponding content of the aforementioned method embodiment can be implemented.

[0350] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the corresponding content of the aforementioned method embodiment can be implemented.

[0351] It should be noted that the terms "first," "second," "third," "fourth," "1," "2," etc. (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that shown or described in the drawings.

[0352] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.

[0353] The above description is only an optional implementation method for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, the use of other similar implementation methods based on the technical ideas of this application also falls within the protection scope of the embodiments of this application.

Claims

1. A model training method, characterized in that: The method comprises: Acquire a training data set; the training data set includes multiple training sample groups; each training sample group includes multiple sample data; the sample data includes target sample data, positive sample data associated with the target sample data, and negative sample data not associated with the target sample data; the sample data includes sample videos or video labels; each video label has a corresponding label description text; the sample data of at least one training sample group includes a sample video and a video label; Performing at least one training operation on the initial video feature extraction model and the initial text feature extraction model using the training data set until a training end condition is met, thereby obtaining a trained video feature extraction model and a trained text feature extraction model; The training operation includes: Extracting video features of each sample video included in the training data set by the initial video feature extraction model, and extracting text features of the label description text of each video label included in the training data set by the initial text feature extraction model; For each training sample group, determining a loss value corresponding to the training sample group based on differences between sample features corresponding to each sample data in the training sample group; the sample features include the video features or the text features; Based on the loss values ​​corresponding to each training sample group, the training loss is determined; based on the training loss, the parameters of the initial video feature extraction model and the initial text feature extraction model are adjusted, and the initial video feature extraction model and the initial text feature extraction model after the adjusted parameters are used as the initial video feature extraction model and the initial text feature extraction model corresponding to the next training operation.

2. The method according to claim 1, characterized in that The obtaining of the training data set includes: Obtain multiple sample videos; each sample video is annotated with at least one video tag; Deduplication is performed on the video labels corresponding to each sample video to obtain multiple video labels; Determine the similarity between each sample video and the similarity between the label description texts corresponding to each video label; Based on the similarity between the sample videos, the similarity between the label description texts corresponding to the video labels, and the correspondence between each sample video and its respective video label, the multiple sample videos and multiple sample labels are divided into multiple training sample groups to obtain the training data set.

3. The method according to claim 2, characterized in that The initial video feature extraction model includes a first video extraction module and a second video extraction module; the initial text feature extraction model includes a first text extraction module and a second text extraction module; Determining the similarity between the sample videos and determining the similarity between the tag description texts corresponding to the video tags includes: Extracting initial video features of each sample video by the first video extraction module, and determining similarities between the initial video features; Extracting initial text features of the tag description texts corresponding to the respective video tags by the first text extraction module, and determining similarities between the initial text features; The adjusting the parameters of the initial video feature extraction model and the initial text feature extraction model based on the training loss includes: Parameters of the second video extraction module and the second text extraction module are adjusted based on the training loss.

4. The method according to claim 2, characterized in that The method of dividing the plurality of sample videos and the plurality of sample labels into a plurality of training sample groups based on the similarity between the sample videos, the similarity between the label description texts corresponding to the video labels, and the correspondence between each sample video and its respective video label comprises: Each sample video is regarded as a video node, and each video tag is regarded as a tag node. Based on the similarity between the sample videos, the similarity between the tag description texts corresponding to the video tags, and the corresponding relationship between each sample video and its respective video tag, the nodes are connected to generate a relationship structure diagram between the multiple sample videos and the multiple sample tags; At least some of the nodes in the relationship structure diagram are used as anchor points; the nodes include video nodes and tag nodes; For each anchor point, determining a positive sample node and a negative sample node of the anchor point based on the connection relationship between each node in the relationship structure graph; The anchor point is used as the target sample data, the positive sample node corresponding to the anchor point is used as the positive sample data, and the negative sample node corresponding to the anchor point is used as the negative sample data to obtain a training sample group corresponding to the anchor point.

5. The method according to claim 4, characterized in that Based on the similarity between the sample videos, the similarity between the label description texts corresponding to the video labels, and the corresponding relationship between each sample video and its respective video label, each node is connected to generate a relationship structure diagram between the multiple sample videos and the multiple sample labels, including: Determine the similarity between each sample video, connect the video nodes whose similarity meets a first threshold, and form isomorphic edges between each video node; Determine the similarity between the label description texts corresponding to the respective video labels, connect the label nodes whose similarity meets the second threshold, and form isomorphic edges between the label nodes; Based on the corresponding relationship between each sample video and its respective video label, heterogeneous edges are formed between the video nodes and the label nodes to obtain the relationship structure graph.

6. The method according to claim 4, characterized in that For each anchor point, determining the positive sample node and the negative sample node of the anchor point based on the connection relationship between the nodes in the relationship structure graph includes: Nodes in the relationship structure graph whose number of connected edges with the anchor point is less than or equal to a preset number are used as positive sample nodes of the anchor point; In the relationship structure graph, nodes whose number of connection edges with the anchor point exceeds the preset number are used as negative sample nodes of the anchor point.

7. The method according to claim 1, characterized in that The training data set further includes at least one training sample group in which all sample data are sample videos, and at least one training sample group in which all sample data are video labels; The method further comprises: Among the multiple training sample groups, the training sample group whose sample data are all video labels is used as a label sample group, the training sample group whose sample data are all sample videos is used as a video sample group, and the sample group whose sample data includes sample videos and video labels is used as a mixed sample group; When the number of training operations reaches a preset number, selecting a plurality of target sample groups from the plurality of training sample groups based on a preset ratio of the number of label sample groups, video sample groups, and mixed sample groups; A new training set is generated based on multiple target sample groups; and the training operations after the preset number of times are performed based on the new training data set.

8. The method according to claim 7, characterized in that In the quantity ratio, the quantity of the mixed sample group is greater than the quantity of the label sample group, and the quantity of the mixed sample group is greater than the quantity of the video sample group.

9. The method according to claim 7, characterized in that The selecting of a plurality of target sample groups from the plurality of training sample groups based on the preset quantity ratio among the label sample group, the video sample group, and the mixed sample group comprises: Determine the loss value corresponding to each training sample group when the number of training operations reaches the preset number; The training sample group whose loss value is greater than the preset threshold is taken as the first sample group; Based on the preset quantity ratios among the label sample group, the video sample group, and the mixed sample group, a plurality of target sample groups are selected from the plurality of the first sample groups.

10. The method according to claim 1, characterized in that For each training sample group, determining the loss value corresponding to the training sample group based on the difference between the sample features corresponding to each sample data in the training sample group includes: Determining first difference information between target sample data and positive sample data in the training sample group, and determining second difference information between the target sample data and the negative sample data; A loss value of the sample group is determined based on the first difference information and the second difference information.

11. The method according to claim 3, characterized in that The extracting of initial video features of each sample video includes: Acquire features of at least two different modalities of the sample video; the features of the different modalities include at least two of the following: picture features of a video frame of the sample video, audio features of a sample audio of the sample video, content text features of a video content text of the sample video, and title text features of a video title of the sample video; fusing the features of the at least two different modalities to obtain initial video features of the sample video; The tag description text includes the video tag, at least two categories of different levels to which the video tag belongs, and description information of the video tag; and extracting initial text features of the tag description text corresponding to each video tag includes: The video tag, at least two categories of different levels to which the video tag belongs, and description information of the video tag are converted into corresponding text features to obtain initial text features of the video tag.

12. A data retrieval method, characterized in that: The method comprises: Acquire target data; the target data includes target video data or target text data; If the target data is a target video, extracting video features of the target video using a trained video feature extraction model; if the target data is a target text, extracting text features of the target text using a trained text feature extraction model; the video feature extraction model and the text feature extraction model are trained based on the model training method according to any one of claims 1 to 11; Based on the extracted video features or text features, retrieval results corresponding to the target data are retrieved in the database; the retrieval results include at least one of video data and label data; wherein the database stores candidate video features corresponding to a plurality of video data respectively; the database also stores candidate text features corresponding to a plurality of label data respectively; the label data includes video labels and label description text; the candidate video features are extracted based on the video feature extraction model; the candidate text features are extracted based on the text feature extraction model.

13. A model training device, characterized in that: The device comprises: A first acquisition module is configured to acquire a training data set; the training data set includes a plurality of training sample groups; each training sample group includes a plurality of sample data; the sample data includes target sample data, positive sample data associated with the target sample data, and negative sample data not associated with the target sample data; the sample data includes sample videos or video labels; each video label has a corresponding label description text; the sample data of at least one training sample group includes a sample video and a video label; A training module, configured to perform at least one training operation on the initial video feature extraction model and the initial text feature extraction model using the training data set until a training end condition is met, thereby obtaining a trained video feature extraction model and a trained text feature extraction model; Wherein, when performing the training operation, the training module is specifically used to: Extracting video features of each sample video included in the training data set by the initial video feature extraction model, and extracting text features of the label description text of each video label included in the training data set by the initial text feature extraction model; For each training sample group, determining a loss value corresponding to the training sample group based on differences between sample features corresponding to each sample data in the training sample group; the sample features include the video features or the text features; Based on the loss values ​​corresponding to each training sample group, the training loss is determined; based on the training loss, the parameters of the initial video feature extraction model and the initial text feature extraction model are adjusted, and the initial video feature extraction model and the initial text feature extraction model after the adjusted parameters are used as the initial video feature extraction model and the initial text feature extraction model corresponding to the next training operation.

14. A data retrieval device, characterized in that: The device comprises: A second acquisition module is used to acquire target data; the target data includes target video data or target text data; an extraction module configured to extract video features of the target video using a trained video feature extraction model if the target data is a target video; and extract text features of the target text using a trained text feature extraction model if the target data is a target text; the video feature extraction model and the text feature extraction model being trained based on the model training method according to any one of claims 1 to 11; A retrieval module is used to retrieve retrieval results corresponding to the target data in a database based on the extracted video features or text features; the retrieval results include at least one of video data and label data; wherein the database stores candidate video features corresponding to a plurality of video data respectively; the database also stores candidate text features corresponding to a plurality of label data respectively; the label data includes video labels and label description text; the candidate video features are extracted based on the video feature extraction model; the candidate text features are extracted based on the text feature extraction model.

15. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the computer program to implement the method according to any one of claims 1 to 12.

16. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which implements the method according to any one of claims 1 to 12 when executed by a processor.

17. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

Citation Information

Cited By

  • Model optimization method, electronic equipment and computer readable storage medium

    CN121561424A