Method and device for training large model, equipment and medium
By obtaining sample data of comments and tag texts, using text encoder to extract features and dynamically fusion, calculate the two-way comparison loss value, and pre-training and fine-tuning of the big model, the technical problems of video aesthetic evaluation are solved, and effective evaluation of video aesthetics and user experience are achieved.
Patent Information
- Application Number
- CN202510757500.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-29
AI Technical Summary
The prior art lacks methods for evaluating video aesthetics, and it is impossible to effectively train large models that can evaluate video aesthetics.
By obtaining sample data containing comment text and label text, the first and second text encoders are used to extract features, dynamically fusion is combined with attention mechanism, bidirectional comparison loss values are calculated, and the big model is pre-trained and fine-tuned to obtain a model that can evaluate video aesthetics.
It realizes an effective evaluation of video aesthetics, improves user experience, and can display video content that meets users' aesthetic needs.
Smart Images

Figure CN120561594A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer data processing technology, and in particular to a method, apparatus, device and medium for training a large model. Background Art
[0002] As short video platforms continue to develop, attracting a large number of users, more and more users are posting videos on these platforms. Users' aesthetic requirements for video content are also becoming increasingly refined. However, existing technologies mostly detect whether a video is distorted, but do not evaluate the video's aesthetics.
[0003] Therefore, how to train a large model that can evaluate the aesthetics of videos is a technical problem that needs to be solved urgently. Summary of the Invention
[0004] The embodiments of this specification provide a method, apparatus, device, and medium for training a large model so as to perform aesthetic evaluation of a video based on the trained large model.
[0005] To solve the above technical problems, the embodiments of this specification provide a method for training a large model, including: Acquire a sample data set; the sample data set includes multiple groups of sample data; any one of the multiple groups of sample data includes comment text, label text, and video data; the comment text and label text in any one of the groups of sample data include aesthetic evaluation information of the video data in the any one of the groups of sample data; Using a first text encoder to perform feature extraction on the comment text to obtain comment text features; Using a second text encoder to perform feature extraction on the label text to obtain label text features; Based on the attention mechanism, the comment text features and the tag text features are dynamically fused to obtain fused text features; Calculating a bidirectional contrast loss value of the fused text features and the video features of the video data; Pre-training the large model based on the bidirectional contrast loss value to obtain a pre-trained large model; Fine-tune the pre-trained large model to obtain a fine-tuned large model.
[0006] The embodiments of this specification also provide a device for training a large model, including: A sample data acquisition module is configured to acquire a sample data set; the sample data set includes multiple groups of sample data; any one of the multiple groups of sample data includes comment text, label text, and video data; the comment text and label text in any one of the groups of sample data include aesthetic evaluation information for the video data in the any one of the groups of sample data; A comment text feature extraction module, configured to extract features from the comment text using a first text encoder to obtain comment text features; a label text feature extraction module, configured to extract features from the label text using a second text encoder to obtain label text features; A feature dynamic fusion module is used to dynamically fuse the comment text features and the tag text features based on an attention mechanism to obtain fused text features; A contrast loss value calculation module, used to calculate a bidirectional contrast loss value of the fused text features and the video features of the video data; A pre-training module, configured to pre-train the large model based on the bidirectional contrast loss value to obtain a pre-trained large model; The fine-tuning module is used to fine-tune the pre-trained large model to obtain a fine-tuned large model.
[0007] The embodiments of this specification also provide a device for training a large model, including: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Acquire a sample data set; the sample data set includes multiple groups of sample data; any one of the multiple groups of sample data includes comment text, label text, and video data; the comment text and label text in any one of the groups of sample data include aesthetic evaluation information of the video data in the any one of the groups of sample data; Using a first text encoder to perform feature extraction on the comment text to obtain comment text features; Using a second text encoder to perform feature extraction on the label text to obtain label text features; Based on the attention mechanism, the comment text features and the tag text features are dynamically fused to obtain fused text features; Calculating a bidirectional contrast loss value of the fused text features and the video features of the video data; Pre-training the large model based on the bidirectional contrast loss value to obtain a pre-trained large model; Fine-tune the pre-trained large model to obtain a fine-tuned large model.
[0008] An embodiment of this specification also provides a computer-readable storage medium having a computer program or instructions stored thereon, which can be executed by a processor to implement the steps of a method for training a large model.
[0009] At least one embodiment in the present specification can achieve the following beneficial effects: by obtaining a sample data set containing multiple groups of text data; wherein any group of sample data sets can contain comment text, label text and data text; the comment text and label text can contain aesthetic evaluation information for the video data; a first text encoder can be used to extract comment text features from the comment text, and a second text encoder can be used to extract label text features from the label text; the comment text features and the label text features are dynamically fused based on the attention mechanism to obtain fused text features, and then the bidirectional contrast loss value of the fused text features and the video features of the video data is calculated; the large model is pre-trained based on the bidirectional contrast loss value, so that the pre-trained large model can be fine-tuned to obtain the fine-tuned large model, so that each video data can be aesthetically evaluated based on the fine-tuned large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some of the embodiments described in this application. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.
[0011] Figure 1 This is a schematic diagram of the overall solution architecture of a method for training a large model in an actual application scenario provided by an embodiment of this specification; Figure 2 This is a flow chart of a method for training a large model provided in an embodiment of this specification; Figure 3 This is a flow chart of a method for training a large model provided in an embodiment of this specification; Figure 4 This is a schematic diagram of the structure of a device for training a large model provided in an embodiment of this specification; Figure 5 This is a structural diagram of a device for a method of training a large model provided in an embodiment of this specification. DETAILED DESCRIPTION
[0012] To make the purpose, technical solutions, and advantages of one or more embodiments of this specification more clear, the technical solutions of one or more embodiments of this specification will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of one or more embodiments of this specification.
[0013] In order to facilitate understanding of the various embodiments in this specification, some terms are explained below.
[0014] Generative AI is an artificial intelligence technology that can create entirely new content. It can generate similar, yet entirely new, text, images, audio, and video by learning patterns and regularities from large amounts of existing data. Specifically, video generation can involve generating video clips from text or images, video editing, or creating dynamic content.
[0015] Multimodal pre-training: This method allows a model to simultaneously learn and understand multiple different modalities and establish relationships between them. Multimodal pre-training enables the model to understand that expressions of the same concept in different modalities are equivalent, while also enabling information complementarity and joint reasoning between modalities.
[0016] Fine-tuning training refers to the process of further adjusting model parameters based on the pre-trained model using a small-scale dataset in a specific field to adapt the model to new tasks.
[0017] Video aesthetics assessment: It is a method of quantitatively evaluating the visual beauty of a video, covering analysis of aesthetic elements such as composition, color, dynamic rhythm, etc.
[0018] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.
[0019] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0020] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0021] The technical solutions provided by the embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0022] Figure 1 This is a schematic diagram of the overall solution architecture of a method for training a large model in an actual application scenario provided in an embodiment of this specification.
[0023] like Figure 1 As shown, the solution may include a sample dataset 1, a server 2, and a large model 3 for aesthetic evaluation. The sample dataset 1 may include multiple groups of sample data, and one group of sample data may include video data, label text, and comment text. The server 2 may obtain the sample dataset 1, and use the first text encoder and the second text encoder to extract features of the comment text and the label text respectively to obtain comment text features and label text features. The server 2 may also dynamically fuse the label text features and the comment text features to obtain fused text features; calculate the bidirectional contrast loss value of the fused text features and the video features of the video data, pre-train the large model based on the bidirectional contrast loss value, and then fine-tune the pre-trained large model to obtain the large model 3 for aesthetic evaluation, so that the large model 3 for aesthetic evaluation can be used to perform aesthetic evaluation on the video data. Figure 1 The server 2 may include but is not limited to any device, equipment, platform, equipment cluster, etc. with computing and processing capabilities.
[0024] Next, a method for training a large model provided in an embodiment of the specification will be described in detail with reference to the accompanying drawings.
[0025] Figure 2This is a flow chart of a method for training a large model provided in an embodiment of this specification. From a program perspective, the execution subject of the process can be a program installed on an application server or an application client.
[0026] like Figure 2 As shown, the method may include the following steps.
[0027] Step 202: Obtain a sample data set.
[0028] The sample data set includes multiple groups of sample data; any group of sample data in the multiple groups of sample data includes comment text, label text and video data; the comment text and label text in any group of sample data include aesthetic evaluation information of the video data in any group of sample data.
[0029] In an embodiment of the present specification, a set of sample data in a sample data set may be a video content sample comprising video data and text data, and the text data may include comment text and tag text. The video content samples in the sample data set may include a variety of different types of video data. Specifically, if divided according to the video description object, the sample data set may include video data of types such as people, scenery, architecture, and food; if divided according to the video subject matter, the sample data may include video data of themes such as documentaries, film and television drama clips, variety shows, news materials, casual videos, and artificial intelligence generated content (AIGC). AICG may be video content generated using generative AI, and generative AI may include at least one of the models for generating videos, such as Magi-1, SkyReels-V2, TTT-MLP, MorphStudio, and Stable Video.
[0030] In the embodiments of this specification, the aesthetic evaluation information may be obtained by evaluating the aesthetics of the video data based on expert experience; or, it may be obtained by users with professional knowledge performing an aesthetic evaluation of the video data. The comment text may be a text that uses natural language to describe the advantages or disadvantages of the video data. For example, the advantages may be "backlighting creates a sense of atmosphere", and the disadvantages may be "lens shaking causes a decrease in viewing experience", etc., which are texts used to comment on the aesthetics of the video data. The label text may be at least one label determined from a number of preset labels that matches the aesthetics of the video data, such as label texts such as push-in shot, rule of thirds composition, and low saturation. The preset labels may be a number of labels determined based on different video composition techniques, and videos with different aesthetics can be constructed through different video communication technologies.
[0031] Step 204: Use a first text encoder to perform feature extraction on the comment text to obtain comment text features.
[0032] In an embodiment of the present specification, the first text encoder may be a comment encoder for encoding comment text. The first text encoder may have a Transformer architecture with a first preset number of layers, for example, a 12-layer Transformer architecture. The first preset number of layers may represent the hierarchical depth, and the hierarchical depth may represent the deep feature extraction capability. Each layer of the Transformer architecture in the first text encoder may comment on the dependency between the words in the text and perform nonlinear feature changes. And the output of the previous layer of Transformer architecture may be used as the input of the next layer of Transformer architecture, until the last layer of Transformer architecture outputs the comment text features.
[0033] Step 206: Use a second text encoder to perform feature extraction on the label text to obtain label text features.
[0034] In the embodiments of this specification, the second text encoder and the first text encoder may be different encoders. The second text encoder may be a label encoder for encoding label text. The second text encoder may be a Transformer architecture with a second preset number of layers, and the second preset number of layers may be a different value from the first preset number of layers, for example, the first preset number of layers is 12, and the second preset number of layers may be 6; the second preset number of layers may also be the same as the first preset number of layers. The second text encoder may be used to encode technical terms, such as technical terms such as "symmetrical composition" and "shallow depth of field". The weight parameters in the second text encoder and the first text encoder are not shared, and both have independent weight parameters. This can avoid confusion between the semantics of certain words in the comment text and certain text in the label text, for example, "backlight" describes the atmosphere in the comment and serves as a technical indicator in the label.
[0035] Step 208: Based on the attention mechanism, the comment text features and the tag text features are dynamically fused to obtain fused text features.
[0036] In the embodiments of this specification, the server can dynamically adjust its attention to the comment text features and the tag text features during the feature fusion process through an attention mechanism to improve the accuracy of the fused text features. The feature dimensions of the comment text features and the tag text features can be the same, for example, 512 dimensions, 768 dimensions, 932 dimensions, etc.
[0037] Step 210: Calculate the bidirectional contrast loss value of the fused text features and the video features of the video data.
[0038] In the embodiments of this specification, the video features of the video data may be obtained by extracting features from the video using a video encoder. The bidirectional contrast loss value may represent the average of the contrast loss value from the fused text features to the video features and the contrast loss value from the video features to the fused text features.
[0039] Step 212: pre-training the large model based on the bidirectional contrast loss value to obtain a pre-trained large model.
[0040] In the embodiments of this specification, a larger bidirectional contrast loss value indicates a lower degree of match between the label text and the comment text and the video data; a smaller bidirectional contrast loss value indicates a higher degree of match between the label text and the comment text and the video data. If the bidirectional contrast loss value does not meet the preset requirements, for example, it is less than the preset contrast loss value, the server can adjust the parameters of the large model and continue to pre-train the large model after the parameter adjustment until the video features extracted by the pre-trained large model called by the server match the fused text features to a high degree, and the obtained contrast loss value meets the preset requirements, so that the pre-trained large model can perform cross-modal alignment between the video data and the label text and the comment text containing aesthetic evaluation information of the video data.
[0041] Step 214: fine-tune the pre-trained large model to obtain a fine-tuned large model.
[0042] In the embodiments of this specification, the server can fine-tune the pre-trained large model using some or all of the sample data in the sample data set. Alternatively, the server can re-acquire additional sample data to fine-tune the pre-trained large model. The fine-tuned large model can be used to perform aesthetic evaluation on video data. Based on the aesthetic evaluation results, video data that meets the user's aesthetic needs can be displayed to the user, improving the user experience.
[0043] It should be understood that the order of some steps in the methods described in one or more embodiments of this specification can be interchanged according to actual needs, or some steps can be omitted or deleted.
[0044] Figure 2The method is to obtain a sample data set containing multiple groups of text data; wherein any group of sample data sets may contain comment text, label text and data text; the comment text and label text may contain aesthetic evaluation information for the video data; a first text encoder may be used to extract comment text features from the comment text, and a second text encoder may be used to extract label text features from the label text; the comment text features and the label text features are dynamically fused based on the attention mechanism to obtain fused text features, and then the bidirectional contrast loss value of the fused text features and the video features of the video data is calculated; the large model is pre-trained based on the bidirectional contrast loss value, so that the pre-trained large model can be fine-tuned to obtain the fine-tuned large model, so that the aesthetic evaluation of each video data can be performed based on the fine-tuned large model.
[0045] based on Figure 2 The present specification also provides some specific implementation methods of the method, which are described below.
[0046] The server can process the video data before obtaining the sample data set, so as to improve the quality of each sample data in the sample data set, and further improve the performance of the pre-trained large model. Optionally, in the embodiment of this specification, a number of video data to be processed can be obtained; the technical parameters of each video data to be processed are adjusted to obtain each adjusted video data; the hash value of each adjusted video data is calculated; the hash value is used to perform deduplication processing on each adjusted video data to obtain deduplication video data; based on the structural similarity index of each deduplication video data, each deduplication video data is screened to obtain each screened video data; if there is a target type of video data in each screened video data, the target type of video data is subjected to spatiotemporal enhancement processing to obtain the target video data; the target video data may include the spatiotemporal enhanced video data and the non-target type of video data in the screened video data; based on the target video data, and the comment text and label text corresponding to the target video data, the sample data is obtained.
[0047] In the embodiments of this specification, technical parameters may include parameters such as video playback duration, video resolution, and video frame rate. The adjustment of technical parameters may specifically include: limiting the playback duration of the video content to a preset duration or a preset range of durations. The preset duration or the preset range of durations may be determined based on expert experience, or may be determined based on model performance and aesthetic expression, such as 5 seconds, 20 seconds, 15 seconds, etc., or within a preset range of durations such as 5-20 seconds, 7-15 seconds. The server may unify the resolution of the video content to a preset resolution so as to balance computing resources and visual quality. The preset resolution may be determined based on expert experience, or may be determined based on computing resources and visual quality requirements. The server may adjust the frame rate of the video content to a preset frame rate. The preset frame rate may be determined based on expert experience, or may be determined based on user needs, such as frame rates such as 25fps and 30fps.
[0048] In the embodiments of the present specification, the hash value may be obtained by hashing the adjusted video data using a hash function. The server may treat at least two adjusted video data with the same hash value as the same video data and perform deduplication processing. The Structural Similarity Index (SSIM) may represent the video quality, and the value range of SSIM may be 0 to 1. The higher the value, the higher the video quality. The server may filter out the video data whose SSIM of the deduplicated video data is greater than or equal to a preset SSIM value to obtain the filtered video data, thereby excluding video data with excessive distortion or low resolution in the video data and improving the quality of the video data. For example, video data whose SSIM is greater than or equal to 0.65 may be filtered out.
[0049] In the embodiments of this specification, the target type can be a video type predetermined based on expert experience, such as architectural video data. Spatiotemporal enhancement can specifically include randomly cropping each frame of the video data, for example, cropping 15% of the image area; performing frame rate dithering on the video data, for example, increasing or decreasing the frame rate by 3 fps; and performing color dithering on the video data, for example, increasing or decreasing the brightness or contrast by 5%.
[0050] Optionally, the embodiment of this specification describes using a first text encoder to perform feature extraction on the comment text to obtain comment text features, which may specifically include: processing the comment text based on the input length of the first text encoder to obtain target comment text; and using the first text encoder to encode the target comment text to obtain the comment text features.
[0051] In the embodiments of this specification, the input length may be a text length that the first text encoder can receive, determined based on the performance of the first text encoder, such as 32, 40, 36, etc. The target review text may be text that meets the input length of the first text encoder, thereby enabling the first text encoder to recognize the target review text and perform encoding processing.
[0052] As an implementation mode, optionally, the comment text is processed based on the input length of the first text encoder described in the embodiment of this specification to obtain the target comment text, which may specifically include: if the length of the comment text is greater than the input length, the comment text is truncated to obtain the target comment text; the length of the target comment text is the input length; if the length of the comment text is less than the input length, the comment text is padded to obtain the target comment text; the length of the target comment text is the input length; if the length of the comment text is equal to the input length, the comment text is used as the target comment text.
[0053] In the embodiments of this specification, the comment text is truncated, specifically, the comment text is summarized to obtain a target comment text that retains the core semantics of the comment text; or, the comment text is divided into multiple target comment texts that meet the input length and have complete semantics; the multiple target comment texts can be marked so that it can be determined that the multiple target comment texts are comments on a video data; thereby, the feature data of the multiple target comment texts can be fused to obtain the comment text features. The comment text is filled, which can be filled with preset words that have no practical meaning to obtain the target comment text. For example, the comment text is "backlight outlines the contours of the characters, but the background is overexposed". Assuming the input length is 32, the filling word is [PAD], and one [PAD] is one length. 18 [PAD] can be filled after the comment text so that the filled comment text can meet the input length requirement of the first text editor.
[0054] In the embodiments of this specification, the label text features may be obtained by encoding the label text using the second text encoder. Both the first text encoder and the second text encoder may be encoders included in the large model for extracting text features. The server may invoke the large model to encode the comment text and the label text using the first and second text encoders to obtain the corresponding comment text features and label text features.
[0055] As an implementation method, the server can use learnable parameters to fuse the label text features and the comment text features to improve the accuracy of the fusion result. Optionally, the embodiment of this specification describes a method of dynamically fusing the comment text features and the label text features based on the attention mechanism to obtain fused text features, which can specifically include: obtaining the learnable parameters learned using the large model; obtaining a first attention weight value corresponding to the comment text feature and a second attention weight value corresponding to the label text feature based on the learnable parameters, the comment text feature and the label text feature; performing weighted summation based on the first attention weight value, the comment text feature, the second attention weight value and the label text feature to obtain the fused text feature.
[0056] In the embodiments of this specification, a learnable parameter can be a parameter learned by the large model during training and can be used to dynamically adjust the weights of comment text features and tag text features, thereby dynamically adjusting the contributions of semantics and technology to the aesthetic evaluation of video data. A learnable parameter can be a parameter in the large model that can be adjusted and optimized through training of the large model.
[0057] In the embodiment of this specification, the first attention weight value can be calculated based on the formula: α=softmax(W·[h_text;h_tag]). Wherein, α represents the first attention weight; W represents the learnable parameter; h_text represents the comment text feature; h_tag represents the tag text feature; softmax represents the normalization function. The second attention weight value can be obtained by subtracting the first attention weight value from 1. Alternatively, the first attention weight value can also be calculated based on the formula: α=σ(W·[h_text;h_tag]+b). Wherein, α represents the first attention weight, α∈[0,1]; W represents the learnable parameter; h_text represents the comment text feature; h_tag represents the tag text feature; b represents the learnable bias value, which can also be learned during the training process of the large model; σ represents the Sigmoid function, which is used for normalization. In practical applications, the attention weight can be generated by the fully connected layer in the large model. The first and second attention weights can be visualized weights. For example, the first attention weight for person-type video data is 0.6, and the second attention weight is 0.4; the first attention weight for scenery-type video data is 0.2, and the second attention weight is 0.8, etc. The first and second attention weights can be used to reflect the differences in how aesthetic evaluation of different types of video data depends on different text modalities. Continuing with the above example, person-type video data relies more on comment text, while scenery-type video data relies more on label text.
[0058] In the embodiments of this specification, the server can calculate the fused text features using the formula: h_fused = α·h_text + (1-α)·h_tag. Here, α represents the first attention weight; h_text represents the comment text features; h_tag represents the tag text features; and h_fused represents the fused text features. This approach allows for dynamic adjustment of the first and second attention weights based on learnable parameters to achieve dynamic fusion of tag and comment text features, enabling the fused text features to better reflect the aesthetic evaluation characteristics of the video data.
[0059] As an implementation method, feature extraction can be performed on the video data to obtain video features before calculating the bidirectional contrast loss value. Optionally, before calculating the bidirectional contrast loss value of the fused text features and video features in the embodiments of this specification, the following steps can also be included: sampling the video data to obtain multiple video frames; extracting feature information of the multiple video frames based on the multiple video frames; and performing mean pooling on the feature information of the multiple video frames to obtain video features of the video data.
[0060] In the embodiments of this specification, the server may use a uniform sampling method to sample video data. Uniform sampling may mean extracting a preset number of images at equal intervals. Equal intervals may mean extracting a preset number of images at a preset time interval. For example, if 12 frames are uniformly sampled, 12 images may be extracted at 1-second intervals, and images played at the 0th, 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, 10th, 11th, and 12th seconds in the video may be extracted as sampling results. Equal intervals may also mean extracting a preset number of images at intervals of a preset number of frames. For example, if 8 frames are uniformly sampled, images may be extracted once at 10-frame intervals, and images at the 1st, 11th, 21st, 31st, 41st, 51st, 61st, 71st, and 81st frames in the video may be extracted as sampling results. Other existing sampling methods may also be used to sample and process video data, which are not specifically limited here.
[0061] In the embodiments of this specification, mean pooling can represent the determination of average feature information of feature information of multiple video frames in the time dimension, and compressing the average feature information to obtain a single vector representing the video features of the video data. This single vector can be used to represent the features of these multiple video frames. Compressing the average feature information can make the dimension of the video features obtained after compression the same as the dimension of the text features after fusion, and can also reduce the redundant information in the video features. For example, if the dimension of each video frame is 768 dimensions and the number of sampled video frames is 12, a 12*768-dimensional video frame feature can be obtained; assuming that the text features after fusion are 512 dimensions, the 12*768 dimensions can be transformed into a 512-dimensional video feature through mean pooling. In actual applications, the server can also compress the average features through a linear layer.
[0062] In practical applications, the video features can be obtained by encoding the video data by the server using a video encoder. The video encoder can also be an encoder included in the large model for extracting video features.
[0063] As an implementation mode, optionally, the extracting feature information of the multiple video frames based on the multiple video frames described in the embodiments of this specification may specifically include: segmenting the multiple video frames according to a preset segmentation size to obtain a number of segmented areas; performing feature extraction on each of the segmented areas based on pseudo 3D convolution to obtain a plurality of local spatiotemporal feature information; and processing the multiple local spatiotemporal feature information using a spatiotemporal attention mechanism to obtain feature information of the multiple video frames.
[0064] In the embodiments of this specification, the preset segmentation size can be determined based on expert experience or the performance of the video encoder. The preset segmentation size can be a size of 32×32, 48×48 pixels, etc. This allows for the generation of multiple video regions of the preset segmentation size, allowing the encoder to perform subsequent processing based on multiple video regions.
[0065] In an embodiment of the present specification, pseudo-3D convolution can extract features from a single video region using information at the same spatial position in the previous and next frames, and obtain local spatiotemporal features with two dimensions of spatial and temporal features. Pseudo-3D convolution can be performed by copying the 2D convolution kernel of the CLIP model into a 3D convolution kernel to obtain a pseudo-3D convolution kernel of a preset size, such as a size of 3×3×3. During the copying process, the 3D convolution kernel needs to be initialized. The initialization of the pseudo-3D convolution kernel can be to copy the weight value of the 2D convolution kernel to the middle time slice of the pseudo-3D convolution kernel, such as the time slice of T=2, and to initialize the dimensions of the pseudo-3D convolution kernel at other positions in the time dimension, such as the time slices of T=1 and T=3, to zero. This can retain the model's ability to process single-frame spatial information while also obtaining information at the same spatial position in adjacent frames, thereby capturing changes between adjacent frames of the video for feature extraction, improving the accuracy of the feature extraction results, and reducing resource consumption compared to real 3D convolution.
[0066] In an embodiment of the present specification, the spatiotemporal attention mechanism may be a self-attention that introduces the time dimension in the Transformer layer to calculate the correlation between frames. Specifically, multiple local spatiotemporal features obtained by pseudo-3D convolution processing may be spliced to obtain a feature sequence and input it into the Transformer layer of the video encoder, so that the Transformer layer can match the Query vector of each video area with the Key vector of other video areas in the sequence to calculate the attention score. The attention score can represent the correlation between each video area and the video area, and then the correlation of different spatiotemporal positions can be obtained, so that the feature data of the first preset dimension corresponding to each video frame can be obtained by weighting through the attention score. Among them, the Query vector and the Key vector can be obtained by linearly transforming each local spatiotemporal feature. Therefore, the video areas at different spatiotemporal positions can be associated and calculated through the spatiotemporal attention mechanism to obtain the feature information of each video frame.
[0067] As an implementation method, the bidirectional contrast loss value can be adjusted by the temperature coefficient so as to accelerate the convergence speed of the large model and improve the training efficiency of the large model. Optionally, the calculation of the bidirectional contrast loss value of the fused text features and video features described in the embodiments of this specification can specifically include: obtaining a temperature coefficient; the temperature coefficient is used as a scaling factor to adjust the similarity difference; based on the temperature coefficient, determining a first similarity matrix from the video feature to the fused text feature direction; based on the temperature coefficient, determining a second similarity matrix from the fused text feature to the video feature direction; using a preset loss function, based on the first similarity matrix and the second similarity matrix, determining the bidirectional contrast loss value.
[0068] In the embodiments of this specification, the smaller the temperature coefficient, the greater the degree to which the similarity gap is magnified; the larger the temperature coefficient, the greater the degree to which the similarity gap is narrowed; for example, when the temperature coefficient is 0.2, the difference between similarities 0.8 and 0.9 is greater than the difference between similarities 0.8 and 0.9 when the temperature coefficient is 0.3. The temperature coefficient can be a coefficient that is dynamically adjusted as the large model is trained.
[0069] In the embodiments of this specification, the first similarity matrix can be calculated based on the formula: S_ij=sim(v_i,t_j) / τ. Wherein, S_ij represents the first similarity matrix; vi_i represents the video feature; t_j represents the fused text feature; τ represents the temperature coefficient; sim represents the similarity calculation function. The second similarity matrix can be calculated based on the formula: S_ji=sim(t_j,v_i) / τ. Wherein, S_ji represents the second similarity matrix. The meaning of other characters can refer to the aforementioned explanation of the calculation formula of the first similarity matrix, and will not be elaborated here.
[0070] In the embodiments of this specification, the bidirectional contrast loss value can be calculated using the formula: L = (CE(S_ij, y) + CE(S_ji, y)) / 2. Here, y represents the true label matrix, which indicates the actual matching between the video data and the fused text for each sample in the sample data set; L represents the bidirectional contrast loss value; and CE represents the cross-entropy loss function. By calculating the bidirectional contrast loss value, the degree of matching between the large model and the video data, the commentary containing aesthetic evaluation information, and the label text can be determined, thereby accurately reflecting whether the large model meets the standards and whether further pre-training is required.
[0071] As an implementation method, the difficult samples in the sample data set can be weighted to improve the mining of difficult samples by the large model during the pre-training process, thereby further improving the performance of the large model after pre-training. Optionally, the method of determining the two-way contrast loss value based on the first similarity matrix and the second similarity matrix in the embodiment of this specification can specifically include: calculating the similarity of each group of sample data based on the first similarity matrix and the second similarity matrix; taking several groups of sample data whose similarity arrangement order is in a preset area as the first sub-sample data set; taking several groups of sample data whose similarity arrangement order is not in the preset area as the second sub-sample data set; using the preset loss function to perform weighted processing based on the first similarity matrix and the second similarity matrix corresponding to each group of sample data in the first sub-sample data set to obtain a first sub-bidirectional contrast loss value; using the preset loss function to determine the second sub-bidirectional contrast loss value based on the first similarity matrix and the second similarity matrix corresponding to each group of sample data in the second sub-sample data set; determining the two-way contrast loss value based on the first sub-bidirectional contrast loss value and the second sub-bidirectional contrast loss value.
[0072] In the embodiments of the present specification, the similarity can represent the degree of matching between the fused text features and the video features. The first similarity of each group of sample data is obtained from the first similarity matrix, and the second similarity of each group of sample data is obtained from the second similarity matrix; the first similarity and the second similarity are averaged to obtain the similarity of each group of samples. The server can arrange the similarities of each group of samples from large to small, and use the groups of sample data that are arranged at the back of a preset percentage as the first sub-sample data set; for example, the 10% of sample data that are arranged at the back as the first sub-sample data set. The server can also arrange the similarities of each group of samples from small to large, and use the groups of sample data that are arranged at the front of a preset percentage as the first sub-sample data set; for example, the 10% of sample data that are arranged at the front as the first sub-sample data set.
[0073] In an embodiment of the present specification, the second sub-sample data set may be a set consisting of the remaining sample data in the sample data set excluding the sample data in each group in the first sub-sample data set. The two-way contrast loss values of the first sub-sample data set and the second sub-sample data set can be calculated according to the above-mentioned two-way contrast loss value, which will not be described in detail here. The server can calculate the two-way contrast loss value based on the formula, L=γ*L1+L2. Wherein, L represents the two-way contrast loss value corresponding to the sample data set; L1 represents the first sub-two-way contrast loss value; L2 represents the second sub-two-way contrast loss value; γ represents the preset weight; the preset weight can be determined based on expert experience. This can achieve the mining of difficult samples and accelerate the convergence of the model.
[0074] As an implementation mode, optionally, the fine-tuning training of the pre-trained big model described in the embodiments of this specification to obtain the fine-tuned big model can specifically include: freezing the video encoder in the pre-trained big model to obtain the big model after the video encoder is frozen; obtaining a number of training data; the training data at least includes comment text, label text and video data; fine-tuning the big model after the frozen video encoder based on the several training data to obtain a first big model to be verified; judging whether the first model loss value of the first big model to be verified meets the preset conditions; if the first model loss value of the first big model to be verified meets the preset conditions, then using the first big model to be verified as the fine-tuned big model.
[0075] In the embodiments of this specification, the training data may also include an overall aesthetic score for the video data and attribute scores for various attributes of the video data. The training data may be obtained from a sample data set or from a data source, and the method for obtaining the training data is not specifically limited here. The overall aesthetic score may be an overall aesthetic score of the video content based on preset classification rules. For example, a 10-point or 100-point scale may be used. If the video content has professional-level composition but dull colors, the score may be determined to be 8 or 80 points. The attribute score may be an attribute score based on various attributes corresponding to the video content. Attributes may specifically include picture structure, shot size, lighting, tone, color, and depth of field; if the video content describes a person type, the attributes may also include expression, action, clothing, and makeup. The overall aesthetic score and attribute score may be determined by users with professional knowledge who score and annotate the video data.
[0076] In the embodiments of this specification, freezing the video encoder preserves the cross-modal alignment capabilities learned during pre-training, thereby avoiding overfitting on small data sets. The server can set requires_grad = False to train the pre-trained large model in a way that only the MLP layer can be trained; requires_grad can be a parameter in the video encoder; False can indicate that the requires_grad parameter can participate in the calculation during fine-tuning training but cannot be changed. The MLP layer can be a layer of a multi-layer perceptron structure that includes a fully connected layer and an activation function.
[0077] In an embodiment of the present specification, the first model loss value can be obtained by verifying the first large model to be verified based on the verification set. The server class determines whether the number of outliers in the verification result output by the first large model to be verified based on the verification set is greater than a preset number; if the number of outliers is less than or equal to the preset number, the first model loss value can be calculated based on the formula: Loss=1 / NΣ(y_pred-y_true)^2; wherein Loss represents the first model loss value; N represents the number of verification data in the verification set used; y_pred can represent the predicted value of the first large model to be verified; y_true can represent the true value corresponding to each verification data in the verification set; Σ represents the sum. If the number of outliers is greater than the preset number, the first model loss value can be calculated based on the formula: Loss={0.5*(y_pred - y_true)^2,if|y_pred-y_true|<1;|y_pred-y_true|- 0.5,otherwise}; where, if the absolute value of the difference between the model prediction value and the true value is less than 1, it can be calculated by 0.5*(y_pred - y_true)^2; if it is other cases, it can be calculated by |y_pred-y_true|-0.5. The physical meaning of each character in this formula is consistent with that in the above formula, and no further details are given here.
[0078] In the embodiments of this specification, the preset condition may mean that the first model loss value is less than or equal to a preset loss value; or that multiple first model loss values corresponding to the first large model to be verified meet the early stopping mechanism. The early stopping mechanism may mean that the first large model to be verified is verified using multiple validation sets to obtain multiple first model loss values; if a preset number of consecutive first model loss values among the multiple first model loss values do not decrease, it is determined that the first model loss value meets the preset condition, and fine-tuning training is stopped to obtain the large model after fine-tuning training.
[0079] In practical applications, the server can use the training data to perform multiple fine-tuning training on the pre-trained large model. The training data used in each fine-tuning training can contain some of the same data or different data. The amount of training data used in each fine-tuning training can be the same or different.
[0080] As an implementation mode, optionally, the method described in the embodiment of this specification may also include: if the loss value of the first model does not meet the preset conditions, obtaining a preset learning rate; using regularization technology to adjust the parameters of the first large model to be verified based on the preset learning rate to obtain a second large model to be verified; if the second model loss value of the second large model to be verified meets the preset conditions, then using the second large model to be verified as the fine-tuned large model.
[0081] In the embodiments of this specification, the learning rate can be determined based on the Cosine annealing mechanism. Specifically, the learning rate can be dynamically adjusted through the Cosine annealing mechanism so that the learning rate can smoothly decay from the initial value to 0 according to the cosine curve, thereby improving the convergence speed of the large model while also being able to fine-tune the parameters of the large model.
[0082] In the embodiments of this specification, regularization techniques may include weight decay and gradient clipping. Weight decay can reduce the weight by a preset value each time the model parameters are updated, thereby reducing model complexity. Gradient clipping adjusts the parameter gradient to avoid gradient explosion and improve the stability of fine-tuning training. The weight decay and gradient clipping parameters can be determined based on expert experience.
[0083] In practical applications, the label text, comment text, overall aesthetic rating, and attribute rating can be obtained by annotating the video data with multiple users with professional knowledge. To improve the accuracy of the aforementioned annotation content, the server can also process each annotation content. Specifically, for comment text, multiple comment text data for a video data set are obtained, and the presence of preset keywords in the comment data is identified. If the preset keywords are present, the preset keywords are removed to obtain multiple comment text data without the preset keywords. Similarity is calculated for the multiple comment text data without the preset keywords using a preset model or a preset similarity calculation method. It is determined whether the similarity is greater than the preset similarity. If the similarity is greater than the preset similarity, the multiple comment text data without the preset keywords are deduplicated to obtain a deduplicated comment text data set. The deduplicated comment text data set is merged to obtain target comment text data, which is used as the comment text data corresponding to the video data in the sample data. The preset keywords can represent words or characters without practical meaning, such as "very good" and "good". The remaining comment text data can be text that can specifically describe the content of the video, such as "slow motion amplifies emotional tension."
[0084] For label text, multiple groups of label text data for a video data are obtained, and one group of label text data includes at least one label text data; text label data whose label text data appears less than a preset number of times in the multiple groups of label text data can be determined, and the text label data is removed to obtain a number of label text data, and the multiple label text data do not contain repeated label text data, and semantic recognition is performed on the multiple label text data, and label text data with similar semantics are merged to obtain a plurality of target label text data, and the plurality of target label text data are used as the label text data corresponding to the video data in the sample data. For example, for label text data removal processing, if the preset number of times is 2, and "fisheye lens" only appears once, the label text data can be removed. For example, for label data merging, if the label text contains "top light" and "upper light source", it can be merged into "top light".
[0085] For the overall aesthetic score, multiple overall aesthetic scores for a video data set are obtained, the average of the multiple overall aesthetic scores is calculated, the deviation of each overall aesthetic score from the average is calculated, and overall aesthetic scores with a deviation greater than a preset deviation are removed as outliers; or the median of the multiple overall aesthetic scores is determined, the deviation of each overall aesthetic score from the median is calculated, and overall aesthetic scores with a deviation greater than a preset deviation are removed as outliers; the median or average of the remaining overall aesthetic scores after removing the outliers is used as the target overall aesthetic score for the video content. The preset deviation can be a value such as 2.5, 2.3, or 2.7. The overall aesthetic score can also be verified. Specifically, if the difference between the maximum and minimum values of the multiple overall aesthetic scores is greater than a preset difference, such as 5, 4.5, or 5.5, a manual review process can be triggered to re-evaluate the overall aesthetic score of the video content to ensure the accuracy of the overall aesthetic score and improve the scoring accuracy of the subsequently trained model. In practical applications, for the attribute scoring of video content, one attribute can also be scored multiple times to obtain multiple attribute scores corresponding to the attribute. Furthermore, attribute scores can be cleaned in the same manner as the overall aesthetic score to obtain target attribute scores, which will not be further elaborated here. The server can use the target overall aesthetic score and target attribute scores as scores for the video data in the training data. This method can be used to process multiple annotations from users with professional knowledge, improving the accuracy of the annotations, further enhancing the cross-modal alignment capabilities of the fine-tuned large model, and the accuracy of aesthetic scoring for video data.
[0086] In order to clearly explain the large model training process in the embodiments of this specification, Figure 3 The present invention provides a flow chart of a method for training a large model. Figure 3 The specific steps are as follows: Step 302: Obtain a sample data set.
[0087] The sample data set includes multiple groups of sample data, and each group of sample data may include an overall aesthetic score, an attribute score, a label text, a comment text, and video data.
[0088] Step 304: Use a video encoder to extract video features from the sample data set.
[0089] The video encoder in the embodiments of this specification may be obtained by extending a preset model, such as CLIPVit-B / 32, and may be a video encoder included in a large model.
[0090] Step 306: Use the first text encoder to extract comment text features from the sample data set.
[0091] Step 308: Use a second text encoder to extract label text features from the sample data set.
[0092] Step 310: Dynamically fuse the comment text features and the tag text features to obtain fused text features.
[0093] Step 312: Calculate the bidirectional contrast loss value based on the fused text features and video features.
[0094] Step 314: Update the model parameters of the large model based on the bidirectional contrast loss value to obtain the pre-trained large model.
[0095] In an embodiment of the present specification, the large model with updated parameters can be verified by a preset verification set. If the large model with updated parameters meets the training stop condition in terms of cross-modal alignment capability of comment text and label text with video data, the large model with updated parameters can be used as the pre-trained large model; if the training stop condition is not met, the large model with updated parameters can continue to be pre-trained until the training stop condition is met to obtain the pre-trained large model.
[0096] Step 316: Freeze the video encoder of the pre-trained large model.
[0097] Step 318: Fine-tune the large model after freezing the video encoder to obtain a fine-tuned large model.
[0098] Fine-tuning training can be used to train the model to regress to the ability to perform aesthetic scoring on video data. This allows the fine-tuned large model to be used to perform aesthetic scoring on video data, and videos with high aesthetic scores can be recommended or displayed to users.
[0099] The order of some steps in the methods described in one or more embodiments of this specification may be interchanged, or some steps may be omitted or deleted as needed. The aforementioned steps may contain the same or similar technical features as those in the preceding embodiments, and reference may be made to the description of the preceding embodiments for further understanding, and no further elaboration is required here.
[0100] This method, firstly, uses a dual text encoder to separate and process natural language comments and label text representing technical terminology, thus avoiding semantic interference. Furthermore, learnable attention weights are used to achieve multimodal adaptive fusion, resulting in fused text features with high feature representation accuracy.
[0101] Secondly, through pseudo-3D convolution expansion, lightweight video spatiotemporal modeling is implemented on the CLIP framework, which only adds a small amount of parameters. While ensuring the ability to extract spatiotemporal features, it can also reduce the amount of data calculation for spatiotemporal modeling.
[0102] Thirdly, by learning cross-modal alignment in the pre-training stage and focusing on aesthetic score regression in the fine-tuning stage, the accuracy of aesthetic scoring of video data by the fine-tuned large model is improved.
[0103] Fourthly, the large model is trained using video data containing annotated information such as overall aesthetic scores, attribute scores, comment text, and label text, which provides fine-grained supervision signals. It can also process each annotated information before training, thereby improving the quality of the annotated information and further improving the accuracy of the processed data of the fine-tuned large model.
[0104] Based on the same idea, the embodiments of this specification also provide a device corresponding to the above method. Figure 4 This is a schematic diagram of a device for training a large model provided in an embodiment of this specification. Figure 4 As shown, the device may include: The sample data acquisition module 402 is configured to acquire a sample data set; the sample data set includes multiple sets of sample data; any set of sample data in the multiple sets of sample data includes comment text, label text, and video data; the comment text and label text in any set of sample data include aesthetic evaluation information of the video data in the any set of sample data; A comment text feature extraction module 404 is configured to extract features from the comment text using a first text encoder to obtain comment text features; The label text feature extraction module 406 is used to extract features from the label text using a second text encoder to obtain label text features; A dynamic feature fusion module 408 is configured to dynamically fuse the comment text features and the tag text features based on an attention mechanism to obtain fused text features; A contrast loss value calculation module 410 is used to calculate a bidirectional contrast loss value of the fused text features and the video features of the video data; A pre-training module 412 is configured to pre-train the large model based on the bidirectional contrast loss value to obtain a pre-trained large model; The fine-tuning module 414 is used to fine-tune the pre-trained large model to obtain a fine-tuned large model.
[0105] based on Figure 5 The present specification also provides some specific implementation plans of the method, which are described below.
[0106] Optionally, the comment text feature extraction module can be used to: Processing the comment text based on the input length of the first text encoder to obtain a target comment text; The target review text is encoded using the first text encoder to obtain the review text features.
[0107] Optionally, the comment text feature extraction module can be used to: If the length of the comment text is greater than the input length, the comment text is truncated to obtain the target comment text; the length of the target comment text is the input length; If the length of the comment text is less than the input length, padding the comment text to obtain the target comment text; the length of the target comment text is the input length; If the length of the comment text is equal to the input length, the comment text is used as the target comment text.
[0108] Optionally, the feature dynamic fusion module can be specifically used to: Obtaining learnable parameters learned using the large model; Based on the learnable parameter, the comment text feature, and the label text feature, obtaining a first attention weight value corresponding to the comment text feature and a second attention weight value corresponding to the label text feature; The fused text feature is obtained by performing weighted summation based on the first attention weight value, the comment text feature, the second attention weight value and the label text feature.
[0109] Optionally, the device may further include a video feature extraction module, which may be specifically used to: Sampling the video data to obtain multiple video frames; Extracting feature information of the plurality of video frames based on the plurality of video frames; Mean pooling is performed on the feature information of the multiple video frames to obtain video features of the video data.
[0110] Optionally, the video feature extraction module may be specifically used to: Segmenting the plurality of video frames according to a preset segmentation size to obtain a plurality of segmentation regions; Performing feature extraction on each segmented area based on pseudo 3D convolution to obtain a plurality of local spatiotemporal feature information; The multiple local spatiotemporal feature information is processed using a spatiotemporal attention mechanism to obtain feature information of the multiple video frames.
[0111] Optionally, the contrast loss value calculation module may be specifically used to: Obtaining a temperature coefficient; the temperature coefficient is used as a scaling factor to adjust the similarity difference; Determining a first similarity matrix from the video feature to the fused text feature based on the temperature coefficient; Determining a second similarity matrix from the fused text feature to the video feature based on the temperature coefficient; The bidirectional comparison loss value is determined based on the first similarity matrix and the second similarity matrix using a preset loss function.
[0112] Optionally, the contrast loss value calculation module may be specifically used to: Calculating the similarity of each group of sample data based on the first similarity matrix and the second similarity matrix; Taking the groups of sample data whose similarity order is in a preset area as the first sub-sample data set; taking the groups of sample data whose similarity arrangement order is not in the preset area as the second sub-sample data set; Using the preset loss function, performing weighted processing based on the first similarity matrix and the second similarity matrix corresponding to each group of sample data in the first sub-sample data set to obtain a first sub-bidirectional contrast loss value; Determine a second sub-bidirectional contrast loss value by using the preset loss function and based on the first similarity matrix and the second similarity matrix corresponding to each group of sample data in the second sub-sample data set; The bidirectional contrast loss value is determined based on the first sub-bidirectional contrast loss value and the second sub-bidirectional contrast loss value.
[0113] Optionally, the fine-tuning module may be specifically used to: Freezing the video encoder in the pre-trained large model to obtain a large model after freezing the video encoder; Acquire a plurality of training data; the training data at least includes comment text, label text and video data; Fine-tune the large model after the frozen video encoder based on the plurality of training data to obtain a first large model to be verified; Determining whether the first model loss value of the first large model to be verified meets a preset condition; If the first model loss value of the first large model to be verified meets the preset condition, the first large model to be verified is used as the fine-tuned large model.
[0114] Optionally, the fine-tuning module may be specifically used to: If the loss value of the first model does not meet the preset condition, obtaining a preset learning rate; Using a regularization technique, adjusting parameters of the first large model to be verified based on the preset learning rate to obtain a second large model to be verified; If the second model loss value of the second large model to be verified meets the preset condition, the second large model to be verified is used as the fine-tuned large model.
[0115] Based on the same idea, the embodiments of this specification also provide devices corresponding to the above methods.
[0116] Figure 5 This is a schematic diagram of a device for training a large model provided in an embodiment of this specification. Figure 5 As shown, the device 500 may include: at least one processor 510; and, A memory 530 in communication with the at least one processor; wherein, The memory 530 stores instructions 520 executable by the at least one processor 510. The instructions are executed by the at least one processor 510 to enable the at least one processor 510 to: Acquire a sample data set; the sample data set includes multiple groups of sample data; any one of the multiple groups of sample data includes comment text, label text, and video data; the comment text and label text in any one of the groups of sample data include aesthetic evaluation information of the video data in the any one of the groups of sample data; Using a first text encoder to perform feature extraction on the comment text to obtain comment text features; Using a second text encoder to perform feature extraction on the label text to obtain label text features; Based on the attention mechanism, the comment text features and the tag text features are dynamically fused to obtain fused text features; Calculating a bidirectional contrast loss value of the fused text features and the video features of the video data; Pre-training the large model based on the bidirectional contrast loss value to obtain a pre-trained large model; Fine-tune the pre-trained large model to obtain a fine-tuned large model.
[0117] Based on the same idea, the embodiments of this specification also provide a computer-readable storage medium corresponding to the above method. The computer-readable storage medium stores a computer program or instructions, which can be executed by a processor to implement the above method for training a large model.
[0118] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. Figure 5 As for the device shown, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0119] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using physical hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using software called a "logic compiler." This is similar to the software compilers used during program development. Before compilation, the original code must be written in a specific programming language, called a Hardware Description Language (HDL). There are many types of HDL, including ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that simply by programming a method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0120] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the memory control logic. Those skilled in the art will also appreciate that, in addition to implementing the controller purely in computer-readable program code, the controller can also be implemented in the form of logic gates, switches, an application-specific integrated circuit, a programmable logic controller, an embedded microcontroller, etc. by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the means for implementing the various functions included therein can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can be considered both a software module implementing the method and a structure within the hardware component.
[0121] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0122] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0123] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0124] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0125] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0126] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0127] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0128] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0129] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0130] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0131] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0132] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0133] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for training a large model, comprising: Acquire a sample data set; the sample data set includes multiple groups of sample data; Any set of sample data in the plurality of sets of sample data includes comment text, label text, and video data; the comment text and label text in any set of sample data include aesthetic evaluation information of the video data in any set of sample data; Using a first text encoder to perform feature extraction on the comment text to obtain comment text features; Using a second text encoder to perform feature extraction on the label text to obtain label text features; Based on the attention mechanism, the comment text features and the tag text features are dynamically fused to obtain fused text features; Calculating a bidirectional contrast loss value of the fused text features and the video features of the video data; Pre-training the large model based on the bidirectional contrast loss value to obtain a pre-trained large model; Fine-tune the pre-trained large model to obtain a fine-tuned large model.
2. The method according to claim 1, wherein the first text encoder is used to extract features from the comment text to obtain comment text features, specifically comprising: Processing the comment text based on the input length of the first text encoder to obtain a target comment text; The target review text is encoded using the first text encoder to obtain the review text features.
3. The method according to claim 2, wherein processing the review text based on the input length of the first text encoder to obtain the target review text specifically comprises: If the length of the comment text is greater than the input length, truncating the comment text to obtain the target comment text; The length of the target review text is the input length; If the length of the comment text is less than the input length, padding the comment text to obtain the target comment text; The length of the target review text is the input length; If the length of the comment text is equal to the input length, the comment text is used as the target comment text.
4. The method according to claim 1, wherein the dynamic fusion of the comment text features and the tag text features based on the attention mechanism to obtain the fused text features specifically includes: Obtaining learnable parameters learned using the large model; Based on the learnable parameter, the comment text feature, and the label text feature, obtaining a first attention weight value corresponding to the comment text feature and a second attention weight value corresponding to the label text feature; The fused text feature is obtained by performing weighted summation based on the first attention weight value, the comment text feature, the second attention weight value and the label text feature.
5. The method according to claim 1, before calculating the bidirectional contrast loss value of the fused text features and the video features of the video data, further comprising: Sampling the video data to obtain multiple video frames; Extracting feature information of the plurality of video frames based on the plurality of video frames; Mean pooling is performed on the feature information of the multiple video frames to obtain video features of the video data.
6. The method according to claim 5, wherein extracting feature information of the plurality of video frames based on the plurality of video frames comprises: Segmenting the plurality of video frames according to a preset segmentation size to obtain a plurality of segmentation regions; Performing feature extraction on each segmented area based on pseudo 3D convolution to obtain a plurality of local spatiotemporal feature information; The multiple local spatiotemporal feature information is processed using a spatiotemporal attention mechanism to obtain feature information of the multiple video frames.
7. The method according to claim 1, wherein calculating the bidirectional contrast loss value of the fused text features and the video features of the video data comprises: Get the temperature coefficient; The temperature coefficient is used as a scaling factor to adjust the similarity difference; Determining a first similarity matrix from the video feature to the fused text feature based on the temperature coefficient; Determining a second similarity matrix from the fused text feature to the video feature based on the temperature coefficient; The bidirectional comparison loss value is determined based on the first similarity matrix and the second similarity matrix using a preset loss function.
8. The method according to claim 7, wherein determining the bidirectional contrast loss value based on the first similarity matrix and the second similarity matrix specifically comprises: Calculating the similarity of each group of sample data based on the first similarity matrix and the second similarity matrix; Taking the groups of sample data whose similarity order is in a preset area as the first sub-sample data set; taking the groups of sample data whose similarity arrangement order is not in the preset area as the second sub-sample data set; Using the preset loss function, performing weighted processing based on the first similarity matrix and the second similarity matrix corresponding to each group of sample data in the first sub-sample data set to obtain a first sub-bidirectional contrast loss value; Determine a second sub-bidirectional contrast loss value by using the preset loss function and based on the first similarity matrix and the second similarity matrix corresponding to each group of sample data in the second sub-sample data set; The bidirectional contrast loss value is determined based on the first sub-bidirectional contrast loss value and the second sub-bidirectional contrast loss value.
9. The method according to claim 1, wherein fine-tuning the pre-trained large model to obtain a fine-tuned large model specifically comprises: Freezing the video encoder in the pre-trained large model to obtain a large model after freezing the video encoder; Get some training data; The training data includes at least comment text, label text and video data; Fine-tune the large model after the frozen video encoder based on the plurality of training data to obtain a first large model to be verified; Determining whether the first model loss value of the first large model to be verified meets a preset condition; If the first model loss value of the first large model to be verified meets the preset condition, the first large model to be verified is used as the fine-tuned large model.
10. The method according to claim 9, further comprising: If the loss value of the first model does not meet the preset condition, obtaining a preset learning rate; Using a regularization technique, adjusting parameters of the first large model to be verified based on the preset learning rate to obtain a second large model to be verified; If the second model loss value of the second large model to be verified meets the preset condition, the second large model to be verified is used as the fine-tuned large model.
11. A device for training a large model, comprising: A sample data acquisition module is used to acquire a sample data set; the sample data set includes multiple groups of sample data; Any set of sample data in the plurality of sets of sample data includes comment text, label text, and video data; the comment text and label text in any set of sample data include aesthetic evaluation information of the video data in any set of sample data; A comment text feature extraction module, configured to extract features from the comment text using a first text encoder to obtain comment text features; a label text feature extraction module, configured to extract features from the label text using a second text encoder to obtain label text features; A feature dynamic fusion module is used to dynamically fuse the comment text features and the tag text features based on an attention mechanism to obtain fused text features; A contrast loss value calculation module, used to calculate a bidirectional contrast loss value of the fused text features and the video features of the video data; A pre-training module, configured to pre-train the large model based on the bidirectional contrast loss value to obtain a pre-trained large model; The fine-tuning module is used to fine-tune the pre-trained large model to obtain a fine-tuned large model.
12. A device comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Acquire a sample data set; the sample data set includes multiple groups of sample data; any one of the multiple groups of sample data includes comment text, label text, and video data; the comment text and label text in any one of the groups of sample data include aesthetic evaluation information of the video data in the any one of the groups of sample data; Using a first text encoder to perform feature extraction on the comment text to obtain comment text features; Using a second text encoder to perform feature extraction on the label text to obtain label text features; Based on the attention mechanism, the comment text features and the tag text features are dynamically fused to obtain fused text features; Calculating a bidirectional contrast loss value of the fused text features and the video features of the video data; Pre-training the large model based on the bidirectional contrast loss value to obtain a pre-trained large model; Fine-tune the pre-trained large model to obtain a fine-tuned large model.
13. A computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions are executed by a processor to implement the steps of the method for training a large model according to any one of claims 1 to 10.