Video resource characterization method, coding model training method and device
By inputting the content text, comment text and image feature vectors of video resources into the encoding model, and generating multimodal fusion vectors, the problem of insufficient accuracy in video resource representation by the existing video recommendation system is solved, and more accurate and personalized video recommendation is achieved.
Patent Information
- Application Number
- CN202510014461.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-16
AI Technical Summary
The existing video recommendation system has insufficient accuracy in the characterization of video resources, resulting in a lack of personalization of recommendation results and the inability to effectively integrate and utilize key information in user comments.
By obtaining the content text, comment text and multi-frame image features of video resources, text conversion and image feature extraction are performed, content vectors, comment vectors and image feature vectors are generated, and these vectors are input into the encoding model to generate multimodal fusion vectors.
The accuracy of video resource representation is improved, and by integrating the video's own information and user comment information, more accurate multimodal feature vectors are generated, thereby improving the personalization and accuracy of the recommendation system.
Smart Images

Figure CN120014648A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to technical fields such as large models, deep learning, data processing, and intelligent recommendation. Background Art
[0002] In recent years, the rapid growth of the number of videos has led to an increasingly personalized and diversified pursuit of video content by users. Among the numerous video resources, users hope to obtain content that matches their interests and needs. The realization of this goal depends on accurate representation of video resources. However, due to the lack of accuracy in the representation of video resources in current video recommendation systems, the recommendation results lack personalization. Therefore, how to improve the accuracy of video resource representation has become a problem that needs to be solved urgently. Summary of the invention
[0003] The present disclosure provides a method for characterizing video resources, and a method and device for training a coding model.
[0004] According to one aspect of the present disclosure, a method for characterizing a video resource is provided, comprising:
[0005] Get the video resource and the comment text for the video resource;
[0006] Get the content text contained in the video resource;
[0007] The content text is converted to obtain a content vector representation; and the comment text is converted to obtain a comment vector representation;
[0008] Determine an image feature vector of the video resource based on multiple frames of images in the video resource;
[0009] The content vector representation, comment vector representation and image feature vector are input into the encoding model to obtain a multimodal fusion vector of the video resource.
[0010] According to another aspect of the present disclosure, a method for training a coding model is provided, comprising:
[0011] Determine a plurality of sample pairs, the sample pairs comprising a positive sample pair and a negative sample pair; wherein the positive sample pair comprises information of two video samples of the same type, and the negative sample pair comprises information of two video samples of different types; wherein the information of the video sample comprises the video sample and a comment text for the video sample;
[0012] The contrastive learning training method is adopted to train the encoding model using multiple sample pairs.
[0013] According to one aspect of the present disclosure, a video resource characterization device is provided, comprising:
[0014] A first acquisition module is used to acquire a video resource and a comment text for the video resource;
[0015] The second acquisition module is used to acquire the content text contained in the video resource;
[0016] A vector conversion module is used to perform text conversion on the content text to obtain a content vector representation; and to perform text conversion on the comment text to obtain a comment vector representation; and to determine an image feature vector of the video resource based on multiple frames of images in the video resource;
[0017] The input module is used to input the content vector representation, the comment vector representation and the image feature vector into the encoding model to obtain the multimodal fusion vector of the video resource.
[0018] According to one aspect of the present disclosure, there is provided a training device for a coding model, comprising:
[0019] A sample pair determination module is used to determine a plurality of sample pairs, wherein the sample pairs include a positive sample pair and a negative sample pair; wherein the positive sample pair includes information of two video samples of the same type, and the negative sample pair includes information of two video samples of different types; wherein the information of the video sample includes the video sample and a comment text for the video sample;
[0020] The training module is used to train the encoding model using a plurality of sample pairs by adopting a contrastive learning training method.
[0021] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0022] at least one processor; and
[0023] a memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.
[0025] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.
[0026] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements any method according to the embodiments of the present disclosure when executed by a processor.
[0027] The present disclosure obtains video resources and their corresponding comment texts, extracts content texts from video resources, determines feature vectors of comment texts, content texts, and video frames, and finally inputs these feature vectors into a coding model to obtain a multimodal feature vector of the video resource. The multimodal feature vector generated by the present disclosure combines the video's own information with the user's comment information to characterize the video resource, which can improve the accuracy of the video resource characterization.
[0028] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0030] Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure;
[0031] Figure 2 is a flow chart of an implementation of a method for characterizing video resources according to an embodiment of the present disclosure;
[0032] Figure 3 is a flow chart of generating a video resource feature vector according to an embodiment of the present disclosure;
[0033] Figure 4 is a flowchart of a method for training a coding model according to an embodiment of the present disclosure;
[0034] Figure 5 is a schematic diagram of a process of coding model training according to an embodiment of the present disclosure;
[0035] Figure 6 is a structural diagram of a video resource representation device 600 according to an embodiment of the present disclosure;
[0036] Figure 7 is a structural diagram of a coding model training device 700 according to an embodiment of the present disclosure;
[0037] Figure 8 is a structural diagram of a coding model training device 800 according to an embodiment of the present disclosure;
[0038] Fig. 9 A schematic block diagram of an example electronic device 900 that may be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0039] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0040] The “and / or” in the embodiments of the present disclosure indicates that three relationships may exist. For example, A and / or B may indicate the following three situations: A exists alone, A and B exist at the same time, and B exists alone. The term “at least one” herein indicates any combination of at least two of any one or more of a plurality of types. For example, at least one of A, B, and C may indicate any one or more elements selected from the set consisting of A, B, and C. The terms “first” and “second” herein refer to and distinguish between multiple similar technical terms, and do not mean to limit the order or to limit the meaning to only two. For example, the first feature and the second feature refer to two types / two features. The first feature may be one or more, and the second feature may also be one or more.
[0041] In recent years, with the rapid increase in the amount of information, users' pursuit of information flow content is tending to be more personalized and diversified. For massive data or resources, users hope to find content that meets their interests, needs and even emotional resonance. However, in the face of this challenge, existing recommendation systems rely on more traditional representation methods, mainly based on the information of the resource itself, such as title, summary, tags, etc., as well as the most direct click interaction between users and resources.
[0042] Although this traditional recommendation strategy can reflect some user preferences to a certain extent, it ignores an important and rich interactive signal between users and resources - user comments. Comments, as a bridge between users and content, are more important than simply clicking a "like" or "dislike" button. User comments not only directly reflect the user's true feelings and feedback on the content, but also reveal the user's emotional tendencies, deep interest preferences, and their willingness and tendency to participate in social interactions.
[0043] However, existing recommendation systems often ignore these user comments and find it difficult to effectively mine and utilize the key information contained in the comments, resulting in the recommendation results often lacking sufficient interactivity and accuracy. Insufficient interactivity means that the recommended content is difficult to stimulate the user's enthusiasm for participation and cannot form an effective interactive cycle between users and content, and between users; while insufficient accuracy directly leads to a deviation between the recommended content and the user's real needs, reducing user satisfaction.
[0044] Therefore, how to effectively integrate and utilize comment signals to optimize the representation of resources and thus improve the performance of the recommendation system has become a key issue that needs to be addressed. In order to solve the above problem, the present disclosure proposes a method for representing video resources. Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure, such as Figure 1 As shown, the application scenario schematic diagram of the embodiment of the present disclosure may include but is not limited to a model training device 110, a feature vector generating device 120 and a resource recommendation device 130. The model training device 110 and the feature vector generating device 120, and the feature vector generating device 120 and the resource recommendation device 130 may communicate through any type of wired or wireless network. Specifically, the model training device 110 may be used to train a coding model, and send the trained coding model to the feature vector generating device 120, the feature vector generating device 120 receives the video resource sent by the resource recommendation device 130, and based on the video resource, the feature vector generating device 120 generates a feature vector using the trained coding model. The resource recommendation device 130 generates corresponding resource recommendations for the user based on the feature vector. Among them, the resource recommendation device 130 proposed in the embodiment of the present disclosure includes but is not limited to electronic devices such as mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, game consoles, e-book readers, multimedia playback devices, wearable devices, etc. In addition, the embodiments of the present disclosure do not impose any specific restrictions on the number of model training devices 110 . For example, the application scenario diagram of the embodiments of the present disclosure may include one or more model training devices 110 .
[0045] Figure 2 is a flowchart of a method for representing video resources according to an embodiment of the present disclosure, including:
[0046] S210, obtaining a video resource and a comment text for the video resource;
[0047] S220, obtaining the content text contained in the video resource;
[0048] S230, performing text conversion on the content text to obtain a content vector representation; and performing text conversion on the comment text to obtain a comment vector representation;
[0049] S240, determining an image feature vector of the video resource based on multiple frames of images in the video resource;
[0050] S250: Input the content vector representation, the comment vector representation and the image feature vector into the encoding model to obtain a multimodal fusion vector of the video resource.
[0051] From the above method, it can be seen that in the process of generating the feature vector corresponding to the video resource, the comment text element is incorporated to enrich the content dimension of the feature vector. The multimodal fusion vector finally generated not only covers various types of information of the video resource, such as content summary, image features, etc., but also incorporates the user's real feedback and comment information on the video resource, which improves the accuracy of the video resource representation vector; in the subsequent process, resource recommendation based on the more accurate video resource representation vector can further improve the accuracy of resource recommendation.
[0052] Figure 3 is a flow chart of generating a video resource feature vector according to an embodiment of the present disclosure.
[0053] like Figure 3 As shown, firstly, the video resource data 310 needs to be determined, wherein the video resource data 310 may include one or more comment texts, content texts and multiple frames of images of the video resource for the video resource.
[0054] In some implementations, the attention level of the comment text meets preset requirements; the attention level is represented by at least one of the number of likes, the number of reposts, and the number of replies.
[0055] For example, the preset requirements include: the number of likes, reposts or replies exceeds a preset threshold; or, the comments are sorted from high to low according to the number of likes, reposts or replies, and the comment texts in the first N (N is a positive integer) places are the comment texts that meet the preset requirements.
[0056] In the above method, when processing numerous comment texts of video resources, a strategic screening method is adopted to identify influential comments from a large amount of user feedback. Specifically, the attention of the comment text can be evaluated based on a series of objective and easily quantifiable indicators, such as the number of likes, forwarding times, and number of responses triggered by the comment text. These indicators not only intuitively reflect the interest of the comment text in the user group, but also provide an effective measurement standard to determine one or more representative and universal comment texts, which can represent more users' feedback and emotional tendencies, interest preferences, and social interaction willingness for video content. In the embodiment of the present disclosure, the top ten comment texts with the highest attention can be selected. In this way, the comment text representing the broad opinions and main views of the user group can be extracted; on this basis, the vector representation of the comment text is determined, and the vector representation is used as a factor for determining the multimodal fusion vector of the video resource, so that the determined multimodal fusion vector can include the user's direct feedback on the video content, and include the user's emotional tendencies, interest preferences, and social interaction willingness for the video content, thereby improving the accuracy and effectiveness of the representation of the video resource.
[0057] In some implementations, the content text included in the video resource includes at least one of a title text and a summary text.
[0058] In some embodiments, it further comprises:
[0059] Performing speech recognition on audio information in the video resource, and / or performing text recognition on text content in the video resource;
[0060] Key information is extracted from the result of speech recognition and / or the result of text recognition to obtain a summary text.
[0061] In the embodiments of the present disclosure, the audio information in the video resource can be recognized by automatic speech recognition (ASR) technology. As a key technology in the field of artificial intelligence, ASR technology uses complex algorithm models, a large amount of training data and powerful computing power to realize the recognition and conversion of voice content in the video.
[0062] In specific operations, ASR technology can capture and analyze diverse voice information such as dialogue, narration, and explanation in video resources, and convert these voice contents into clear and accurate text formats in real time or offline. In addition, ASR technology also has strong flexibility and adaptability, and can realize the recognition of professional terms in different languages, dialects, and specific fields.
[0063] In the embodiments of the present disclosure, text recognition can be performed on text content in video resources by using optical character recognition (OCR) technology. The core function of OCR technology is to convert text information in an image into editable and searchable text data.
[0064] In specific applications, when video resources contain text content, such as subtitles, labels, text descriptions, etc., OCR technology can be used to recognize these texts. OCR technology first locates the text area in the video through image processing algorithms, and then uses deep learning models to segment and recognize the text.
[0065] The text data generated by ASR and OCR technologies contains a large amount of information in the video resources. This information may cover the core content of the video resources, but it may also contain irrelevant content, repeated paragraphs or redundant information. To address this problem, the ERNIE Tiny large model can be used to process the text data generated by ASR or OCR technology to identify and extract key information in the text data, such as core ideas, important data, key events, etc.
[0066] When processing text data, the ERNIE Tiny model conducts in-depth analysis of the text to identify content that is closely related to the topic, while excluding irrelevant information, repeated parts, and information redundancy. On this basis, the ERNIE Tiny model further uses its summary generation capabilities to integrate and refine the identified key information to generate a text summary of the video resource. Extracting summary text through the ERNIE Tiny model not only solves the problem of redundant information in the original text, but also improves data input efficiency and semantic expression quality. In multimodal tasks, summary text provides more accurate feature representation vectors than original text, thereby improving the performance of the model and task results. Specifically, because the information in the summary text is more refined, the model requires less time and computing resources for processing, thereby speeding up the processing speed.
[0067] By performing speech recognition on the audio information in the video resources and text recognition on the text content in the video resources, the information content in the video can be captured. Furthermore, by extracting key information from these recognition results, a summary text corresponding to the video resource can be generated. This method can integrate and refine the video resource information, and retain the core points in the video while removing redundant information, thereby improving the efficiency of generating the subsequent video resource feature vector.
[0068] In the disclosed embodiment, the BAAI General Embedding (BGE) model 320 can be used to convert the above-mentioned comment text and content text into corresponding comment vector representation and content vector representation. Based on the principle of deep learning, the BGE model 320 uses a neural network architecture to extract information and encode features from the input text data. Specifically, when the comment text and content text are input into the BGE model 320, the model will perform pre-processing operations such as word segmentation and stop word removal on these texts to remove redundant information and retain key content.
[0069] Subsequently, the BGE model 320 uses pre-trained word embedding or sentence embedding technology to map each word or sentence into a high-dimensional vector space to form a "word vector" or "sentence vector". These vectors not only contain the semantic information of the vocabulary, but also capture the association and contextual relationship between the vocabulary. In this way, the comment text and content text are converted into their respective corresponding comment vector representation and content vector representation. In addition, the BGE model can focus on text representation, has semantic matching and context understanding capabilities, and provides accurate expression for text features.
[0070] In the embodiment of the present disclosure, since the content text includes the title text and the summary text, the content vector representation also includes the title vector representation and the summary vector representation.
[0071] In some implementations, determining an image feature vector of a video resource includes:
[0072] Extract multiple frames of images from video resources according to predetermined rules;
[0073] The image feature vectors of the multiple frames of images are determined, and the image feature vectors of the multiple frames of images are used as the image feature vectors of the video resource.
[0074] For the multiple frames of images in the video resource data 310, required images can be extracted from the video resources according to predetermined rules.
[0075] In the embodiment of the present disclosure, the preset rule may be: when extracting N frames of images, extract one frame of image per second in the first N / 2 seconds of the video resource to obtain N / 2 frames of image;
[0076] In the remaining time of the video resource, another N / 2 frames of images are extracted at an average time interval;
[0077] Wherein, N is an integer greater than 1.
[0078] In one example, if 10 frames of images need to be extracted from a video resource, images can be extracted from the first 5 seconds of the video resource at a frequency of one frame per second to obtain the first 5 frames of images. Then, the remaining 5 frames of images are extracted at evenly distributed time intervals during the remaining playback time of the video resource.
[0079] In the disclosed embodiment, the contrastive language-image pre-training (CLIP) model 330 may be used to convert the above-mentioned multiple frames of images into image feature vectors one by one.
[0080] Specifically, by inputting the above-mentioned multiple frames of images one by one into the CLIP model 330, the images are converted into corresponding feature vectors using the image parsing capability of the model. In this process, the CLIP model 330 can capture features such as color, texture, shape, and spatial relationship between objects in the image, and encode this information into a vector form. Finally, the image feature vectors corresponding to the multiple frames of images can be used as image feature vectors of video resources.
[0081] By extracting multiple frames of images from video resources through preset rules and obtaining the feature vectors of these images respectively, the image feature vectors of the entire video resource are determined, which improves the efficiency and accuracy of image feature extraction of video resources. At the same time, the CLIP model maps text information and image information to the same vector space, captures the semantic consistency between multimodal information, can handle diverse resource content, and improves the accuracy of generating matching feature vectors for multimodal information.
[0082] like Figure 3 As shown, by performing feature conversion processing on the video resource data, ten comment vector representations, one title vector representation, one summary vector representation, and ten image feature vectors can be obtained. In addition, in order to integrate these vector representations and feature vectors, a classification (CLS) vector can be introduced to form a complete feature sequence. Usually the CLS vector is added to a specific position of the feature sequence as a special mark. The feature sequence is input into the encoding model 340, and the encoding model 340 can further process the feature sequence.
[0083] In some embodiments, the encoding model 340 includes an encoder 341 of a conversion model, and the encoder 341 includes a four-layer encoding structure.
[0084] In the embodiment of the present disclosure, the encoding model 340 includes an encoder 341 of a Transformer model, such as Figure 3 As shown, in some examples, the encoder 341 can be composed of four identical encoding layers stacked together, each encoding layer including a multi-head attention unit, a feedforward unit, and a normalization unit.
[0085] In the process of generating multimodal fusion vectors for video resources, the feature sequence is first converted into a word embedding vector through the input embedding layer, and position encoding is added. Then, the elements in the feature sequence are sent to multiple encoding layers for processing. In each encoding layer, the multi-head self-attention unit calculates the attention weights between the input content and generates a weighted sum vector based on these weights. Then, the feedforward unit further processes and transforms the weighted sum vector. Finally, the output of the current layer is combined with the input through the normalization unit, and the data distribution is stabilized to generate the input of the next layer.
[0086] By performing the above structural design on the coding model 340, the extraction accuracy of information in the feature sequence and the generation efficiency of the video resource feature vector can be improved.
[0087] After the superposition processing of the four coding layers, the mapping module 342 in the coding model can be used to convert the processing result into a specific video resource multimodal fusion vector 350. The multimodal fusion vector 350 can improve the similarity between similar video resources and enhance the ability to distinguish between different resources. In practical applications, when video resources with high similarity to the multimodal fusion vector 350 are extracted, not only can video resources with similar title text or picture content be found, but also those video resources that have large differences in title information or picture display but mention the same event or contain similar text content in user comments can be identified.
[0088] Based on the above functions, the multimodal fusion vector 350 can be applied to the intelligent recommendation system. Specifically, based on the multimodal fusion vector 350, resources can be recalled and sorted for users.
[0089] In the resource recall stage, the system uses the user's recent interactive behaviors, such as clicking, browsing comments, liking comments, and replying to comments, as the basis for triggering resource recall. These behaviors not only reflect the user's immediate interests, but also contain clues to their long-term preferences. Then, the system determines the resources to be recalled for the user by calculating the similarity between the video resource feature vector and other candidate resource vectors. Among them, the recalled resources can be a single candidate resource that best matches the user's interests and has the highest similarity, or multiple candidate resources whose similarity exceeds a preset threshold.
[0090] In the resource sorting stage, if the recommendation system has screened out multiple candidate resources for the user in the previous resource recall stage, the system can sort these candidate resources based on the similarity between the feature vectors of these recalled resources and the feature vectors of resources involved in the user's historical behavior. Based on the sorting results, resources that the user is interested in or that meet their personalized needs can be presented first, thereby improving the accuracy of resource recommendations.
[0091] By using the above strategies, the accuracy of resource recommendations to users can be improved.
[0092] The disclosed embodiment also provides a method for training a coding model. Figure 4 is a flowchart of a method for training a coding model according to an embodiment of the present disclosure, including:
[0093] S410, determining a plurality of sample pairs, the sample pairs comprising a positive sample pair and a negative sample pair; wherein the positive sample pair comprises information of two video samples of the same type, and the negative sample pair comprises information of two video samples of different types; wherein the information of the video sample comprises the video sample and a comment text for the video sample;
[0094] S420: adopt a contrastive learning training method and use multiple sample pairs to train the encoding model.
[0095] By constructing multiple sample pairs including positive sample pairs and negative sample pairs, the coding model is trained by contrastive learning, so that the coding model can generate multimodal fusion vectors of video resources. Since the sample pairs contain video samples and comment texts for the video samples, the multimodal fusion vectors generated by the model can represent the video resources by combining the video's own information with the user's comment information, thereby improving the accuracy of the video resource representation.
[0096] In the disclosed embodiment, the model is trained using a contrastive learning method. The basic idea of contrastive learning is to make similar samples closer in the feature space, while dissimilar samples farther away. This is usually achieved by constructing positive and negative sample pairs: a pair of samples is selected, and whether the pair of samples is a positive sample pair or a negative sample pair is determined based on whether they belong to the same category or have some similarity. Through the learning process, the model is trained to make the feature representations of the positive sample pair closer in the feature space, while the feature representations of the negative sample pair are farther away.
[0097] Figure 5 FIG. 1 is a flow chart of coding model training according to an embodiment of the present disclosure. Figure 5 The corresponding processing process and data flow are schematically drawn for each video sample.
[0098] In some implementations, a contrastive learning training method is used to train the encoding model using multiple sample pairs, including:
[0099] The information of each video sample in the sample pair is processed in the following ways: obtaining the content text contained in the video sample, performing text conversion on the content text to obtain a content vector representation, and performing text conversion on the comment text for the video sample to obtain a comment vector representation; determining the image feature vector of the video sample based on multiple frames of images in the video resource; inputting the content vector representation, the comment vector representation and the image feature vector into the first encoding model to be trained to obtain a multimodal fusion vector of the video sample;
[0100] The similarity between the multimodal fusion vectors of the two video samples of the sample pair is calculated, and the first encoding model is adjusted according to the similarity and the type of the sample pair to obtain a trained second encoding model; wherein the type of the sample pair is a positive sample pair or a negative sample pair.
[0101] In the embodiment of the present disclosure, it is necessary to determine the information of each video sample in the sample pair, wherein the video sample can be Figure 5 The video sample item_a and the video sample item_b in . The video sample item_a and the video sample item_b can form a set of positive sample pairs, or form a set of negative sample pairs.
[0102] In some implementations, the method further includes obtaining a comment text for the video sample, including:
[0103] From a plurality of comment texts for the video sample, one or more comment texts whose attention levels meet preset requirements are obtained; the attention levels are represented by at least one of the number of likes, the number of reposts, and the number of replies.
[0104] For example, the preset requirements include: the number of likes, reposts or replies exceeds a preset threshold; or, the comments are sorted from high to low according to the number of likes, reposts or replies, and the comment texts in the first N (N is a positive integer) places are the comment texts that meet the preset requirements.
[0105] In the above method, when processing multiple comment texts of a video sample, a strategic screening method is adopted to identify influential comments from a large number of comment texts. This screening process is based on a series of objective and quantifiable indicators, such as the number of likes of the comment text, the frequency of forwarding, and the number of replies triggered, in order to evaluate the attention paid to each comment. These indicators not only intuitively reflect the interest and influence of the comment text among the user group, but also provide an effective measurement standard to determine one or more representative and universal comment texts that can represent more users' feedback and emotional tendencies, interest preferences, and willingness for social interaction towards the video content. For example Figure 5 As shown in the figure, in one example, the top ten comment texts ranked by attention can be selected. In this way, the comment texts reflecting the general views and main opinions of the user group can be extracted; based on this, the comment texts are converted into vector form, and this vector is used as a component of constructing the multimodal fusion vector of video resources, so that the generated multimodal fusion vector can reflect the user's direct evaluation of the video content, while incorporating the user's emotional tendencies, interest preferences, and willingness to participate in social interactions, thereby improving the accuracy and effectiveness of video resource representation.
[0106] In some implementations, the content text included in the video sample includes at least one of a title text and a summary text.
[0107] In some embodiments, it further comprises:
[0108] Performing speech recognition on audio information in the video sample, and / or performing text recognition on text content in the video sample;
[0109] Key information is extracted from the result of speech recognition and / or the result of text recognition to obtain a summary text.
[0110] In the embodiments of the present disclosure, the audio information in the video sample may be subjected to speech recognition by using the ASR technology, and the text content in the video sample may be subjected to text recognition by using the OCR technology.
[0111] The text data generated by ASR and OCR technologies contain rich information from video resources, which may touch upon the core content of the video, or may be mixed with irrelevant information, repetitive paragraphs or redundant data. To solve this problem, the ERNIE Tiny large model can be used to deeply process the text data generated by ASR or OCR, aiming to identify and extract the key elements in the text, including valuable information such as core ideas, important data, and key events.
[0112] Furthermore, the ERNIE Tiny model uses its summary generation capability to integrate and refine the identified key information to form an accurate text summary.
[0113] By processing the audio information in the video resources through speech recognition technology and parsing the text content in the video resources using text recognition technology, the information content in the video sample can be captured and extracted. Furthermore, by extracting key information from the information content, a summary text of the video sample can be obtained. The summary text can effectively integrate and refine the video resource information, while removing redundant content and retaining the core points of the video, thereby improving the efficiency of subsequent production of video resource feature vectors.
[0114] For each video sample, the content text contained therein is obtained. The content text and the comment text for the video sample can be converted into corresponding content vector representation and comment vector representation using the BGE model 510 .
[0115] In some implementations, determining an image feature vector of a video sample includes:
[0116] Extracting multiple frames of images from the video sample according to a predetermined rule;
[0117] The image feature vectors of the multiple frames of images are determined, and the image feature vectors of the multiple frames of images are used as the image feature vectors of the video samples.
[0118] In the embodiment of the present disclosure, the preset rule may be: when extracting N frames of images, extract one frame of image per second in the first N / 2 seconds of the video resource to obtain N / 2 frames of image;
[0119] In the remaining time of the video resource, another N / 2 frames of images are extracted at an average time interval;
[0120] Wherein, N is an integer greater than 1.
[0121] After determining multiple frames of images in the video sample, you can use Figure 5 The CLIP model 520 shown converts multiple frames of images into image feature vectors one by one.
[0122] According to preset rules, multiple frames of images are selected from the video sample, and the feature vectors of these images are obtained respectively to determine the image feature vector of the entire video sample. This process can improve the feature extraction efficiency and accuracy of the video sample.
[0123] For the acquired multiple-frame images of each video sample, the CLIP model 520 may be used to convert the multiple-frame images into corresponding image feature vectors.
[0124] After the information of each video sample is processed by feature conversion, a series of features corresponding to the information of these video samples can be obtained. Figure 5 As shown, a series of features corresponding to each video sample may include: vector representations of ten comment texts, a vector representation of a title, a vector representation of a summary, and feature vectors of ten images. In addition, in order to integrate these vector representations and feature vectors, a CLS vector may be introduced to form a complete feature sequence. The feature sequence is input into the first encoding model 530 to be trained to further process the feature sequence.
[0125] The content vector representation, comment vector representation and image feature vector corresponding to each video sample are input into the first encoding model 530 to be trained to generate a corresponding multimodal fusion vector.
[0126] In some embodiments, the encoding model includes an encoder of a conversion model, and the encoder includes a four-layer encoding structure.
[0127] In the embodiment of the present disclosure, the encoding model includes an encoder 531 of a Transformer model. Figure 5 As shown, the encoder 531 can be composed of four identical encoding layers stacked together, each encoding layer including a multi-head attention unit, a feedforward unit and a normalization unit.
[0128] In the process of generating multimodal fusion vectors for video samples, the feature sequence is first converted into a word embedding vector through the input embedding layer, and position encoding is added. Then, the elements in the feature sequence are sent to multiple encoding layers for processing. In each encoding layer, the multi-head self-attention unit calculates the attention weights between the input content and generates a weighted sum vector based on these weights. Then, the feedforward unit further processes and transforms the weighted sum vector. Finally, through the normalization unit, the output of the current layer is combined with the input and the data distribution is stabilized to generate the input of the next layer.
[0129] Using the above structure to design a coding model can improve the accuracy of information extraction in feature sequences and the efficiency of generating multimodal fusion vectors corresponding to video samples.
[0130] After the encoder 531 processes the content vector representation, comment vector representation, and image feature vector corresponding to each video sample, the mapping module 532 may generate a multimodal fusion vector corresponding to each video sample. Figure 5 It can be expressed as a multimodal fusion vector item_a and a multimodal fusion vector item_b.
[0131] In the disclosed embodiment, the information noise contrast estimation (InfoNCE) loss is used to adjust the first coding model 530 to be trained. InfoNCE loss learns meaningful representations of data by comparing the similarities between sample pairs. This does not rely on a large amount of labeled data, but uses the inherent structure and relationship of the data itself for training.
[0132] The core idea of InfoNCE loss is that in feature space, positive sample pairs should show higher similarity and be close to each other, while negative sample pairs should show lower similarity and be far away from each other. In the implementation process, InfoNCE loss calculates the similarity between each sample and its positive sample, and compares it with the similarity between its negative sample. Based on this loss, the model can maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs.
[0133] In one example, the InfoNCE loss is determined by calculating the similarity between the multimodal fusion vector item_a and the multimodal fusion vector item_b, combined with the type of sample pairs used to generate the two multimodal fusion vectors, where the type of the sample pair is a positive sample pair or a negative sample pair, and the similarity between the multimodal fusion vector item_a and the multimodal fusion vector item_b can be cosine similarity. The expression of InfoNCE loss is:
[0134]
[0135] In the above formula, x i The query sample is the sample to be classified or identified; is with x i Relevant positive samples; Represents x i Irrelevant negative samples; sim() represents the similarity between samples; τ represents the temperature parameter; N represents the number of sample pairs; K represents the number of negative samples.
[0136] Based on the above InfoNCE loss, the first encoding model 530 to be trained can be adjusted to obtain the trained second encoding model. The model training method of contrastive learning can improve the generalization ability of the second encoding model, so that the multimodal feature vector of the video sample generated by the model can be represented based on the video's own information and user comment information, thereby improving the accuracy of the video resource representation vector; in the subsequent process, using a more accurate video resource representation vector to perform resource recommendation can further improve the accuracy of resource recommendation.
[0137] In some implementations, determining a positive sample pair includes:
[0138] Acquire multiple video files watched by the user within a predetermined time period, and determine the multiple video files as a video file set of the user;
[0139] From the user's video file collection, two video files are randomly selected to determine a positive sample pair.
[0140] In the disclosed embodiment, it is necessary to obtain multiple video files watched by a user within a predetermined time period. This step can be achieved by analyzing the user's viewing history, browsing history, or downloading history. By collecting the video files watched by the user within the predetermined time period, a video file set of the user can be obtained, and this set includes multiple video files watched by the user within the specified time period.
[0141] From the user's video file collection, two video files are randomly selected to determine a positive sample pair. A positive sample pair refers to two video files that are considered to have similarity or correlation under certain specific conditions. Through this method, two video files with correlation or similarity can be extracted to determine a positive sample pair. In the model training stage, the use of positive sample pairs can enhance the model's ability to identify similar samples and improve the accuracy of the model's generation of video resource representation vectors.
[0142] In some implementations, determining a negative sample pair includes:
[0143] For multiple users, obtain multiple video files watched by each user within a predetermined time period to determine a video file set of each user;
[0144] From the video file sets of two users, one video file is randomly selected respectively to determine a negative sample pair.
[0145] In the disclosed embodiment, multiple video files watched by each user within a predetermined time period are obtained respectively. This step involves a detailed analysis of the user's viewing records, historical browsing data, or logs through related applications to capture the user's viewing behavior. Through this process, a video file collection can be constructed for each user, which accurately reflects the viewing preferences and interests of each user within a specified time period.
[0146] In order to construct a negative sample pair, one video file is randomly selected from two different user video file sets. If the two video files come from two users, and the two users have no obvious similarity in viewing behavior or the video files themselves have no direct correlation in terms of content, type, label, etc., then the two video files can be regarded as a negative sample pair. In this way, the two obtained video files can be made to have no correlation, and then the two video files are determined to be a negative sample pair. Introducing negative sample pairs during the model training process can enable the model to learn the differences in vector representations corresponding to different types of video files during the training process.
[0147] In addition, considering that popular resources appear more frequently in user viewing records, when constructing sample pairs, popular resources are prone to over-sampling, that is, the probability of popular resources being determined as positive sample pairs or negative sample pairs increases. Therefore, before constructing sample pairs, a preheating step can be taken, that is, setting a heat threshold and excluding those resources whose heat exceeds the threshold. This measure aims to rely more on resources with heat below the threshold to construct sample pairs, so as to reduce the problem of excessive penalties that popular resources may suffer due to oversampling.
[0148] In this way, not only can the accuracy of model training be improved, but also the generalization ability of the model can be enhanced, so that it can generate corresponding multimodal feature vectors when faced with resources of different popularity and types.
[0149] The disclosed embodiment also provides a video resource representation device, Figure 6 is a structural diagram of a video resource representation device 600 according to an embodiment of the present disclosure, including:
[0150] A first acquisition module 610 is used to acquire a video resource and a comment text for the video resource;
[0151] The second acquisition module 620 is used to acquire the content text contained in the video resource;
[0152] The vector conversion module 630 is used to perform text conversion on the content text to obtain a content vector representation; and to perform text conversion on the comment text to obtain a comment vector representation; and to determine an image feature vector of the video resource based on multiple frames of images in the video resource;
[0153] The input module 640 is used to input the content vector representation, the comment vector representation and the image feature vector into the encoding model to obtain a multimodal fusion vector of the video resource.
[0154] In some implementations, the attention level of the comment text meets preset requirements; the attention level is represented by at least one of the number of likes, the number of reposts, and the number of replies.
[0155] In some implementations, the content text included in the video resource includes at least one of a title text and a summary text.
[0156] In some implementations, the second acquisition module 620 is further configured to:
[0157] Performing speech recognition on audio information in the video resource, and / or performing text recognition on text content in the video resource;
[0158] Key information is extracted from the result of speech recognition and / or the result of text recognition to obtain a summary text.
[0159] In some implementations, the vector conversion module 630 is used to:
[0160] Extract multiple frames of images from video resources according to predetermined rules;
[0161] The image feature vectors of the multiple frames of images are determined, and the image feature vectors of the multiple frames of images are used as the image feature vectors of the video resource.
[0162] In some embodiments, the encoding model includes an encoder of a conversion model, and the encoder includes a four-layer encoding structure.
[0163] The disclosed embodiment also provides a training device for a coding model. Figure 7 is a schematic diagram of the structure of a coding model training device 700 according to an embodiment of the present disclosure, including:
[0164] The sample pair determination module 710 is used to determine a plurality of sample pairs, wherein the sample pairs include positive sample pairs and negative sample pairs; wherein the positive sample pairs include information of two video samples of the same type, and the negative sample pairs include information of two video samples of different types; wherein the information of the video samples includes the video samples and comment text for the video samples;
[0165] The training module 720 is used to train the coding model using a plurality of sample pairs by adopting a contrastive learning training method.
[0166] In some embodiments, the training module 720 is used to:
[0167] The information of each video sample in the sample pair is processed in the following ways: obtaining the content text contained in the video sample, performing text conversion on the content text to obtain a content vector representation, and performing text conversion on the comment text for the video sample to obtain a comment vector representation; determining the image feature vector of the video sample based on multiple frames of images in the video resource; inputting the content vector representation, the comment vector representation and the image feature vector into the first encoding model to be trained to obtain a multimodal fusion vector of the video sample;
[0168] The similarity between the multimodal fusion vectors of the two video samples of the sample pair is calculated, and the first encoding model is adjusted according to the similarity and the type of the sample pair to obtain a trained second encoding model; wherein the type of the sample pair is a positive sample pair or a negative sample pair.
[0169] Figure 8 8 is a schematic diagram of a structure of a training device 800 for a coding model according to an embodiment of the present disclosure. Figure 8 As shown, in some embodiments, a third acquisition module 830 is further included, which is used to:
[0170] From a plurality of comment texts for the video sample, one or more comment texts whose attention levels meet preset requirements are obtained; the attention levels are represented by at least one of the number of likes, the number of reposts, and the number of replies.
[0171] In some implementations, the content text included in the video sample includes at least one of a title text and a summary text.
[0172] In some implementations, the third acquisition module 830 is further configured to:
[0173] Performing speech recognition on audio information in the video sample, and / or performing text recognition on text content in the video sample;
[0174] Key information is extracted from the result of speech recognition and / or the result of text recognition to obtain a summary text.
[0175] In some embodiments, the training module 720 is used to:
[0176] Extracting multiple frames of images from the video sample according to a predetermined rule;
[0177] The image feature vectors of the multiple frames of images are determined, and the image feature vectors of the multiple frames of images are used as the image feature vectors of the video samples.
[0178] In some embodiments, the encoding model includes an encoder of a conversion model, and the encoder includes a four-layer encoding structure.
[0179] In some implementations, the sample pair determination module 710 is configured to:
[0180] Acquire multiple video files watched by the user within a predetermined time period, and determine the multiple video files as a video file set of the user;
[0181] From the user's video file collection, two video files are randomly selected to determine a positive sample pair.
[0182] In some implementations, the sample pair determination module 710 is configured to:
[0183] For multiple users, obtain multiple video files watched by each user within a predetermined time period to determine a video file set of each user;
[0184] From the video file sets of two users, one video file is randomly selected respectively to determine a negative sample pair.
[0185] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0186] In the technical solution disclosed in the present invention, the acquisition, storage and application of personal information of users involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0187] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0188] Fig. 9 A schematic block diagram of an example electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0189] like Fig. 9As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0190] A number of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0191] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as detection methods. For example, in some embodiments, the detection method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the detection method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the detection method in any other appropriate manner (e.g., by means of firmware).
[0192] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0193] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0194] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0195] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0196] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0197] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0198] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0199] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for representing video resources, comprising: Obtaining video resources and comment text for the video resources; Obtain the content text contained in the video resource; Performing text conversion on the content text to obtain a content vector representation; and performing text conversion on the comment text to obtain a comment vector representation; Determine an image feature vector of the video resource based on multiple frames of images in the video resource; The content vector representation, the comment vector representation and the image feature vector are input into a coding model to obtain a multimodal fusion vector of the video resource.
2. The method according to claim 1, wherein: The attention level of the comment text meets preset requirements; the attention level is represented by at least one of the number of likes, the number of reposts and the number of replies.
3. The method according to claim 1, wherein: The content text contained in the video resource includes at least one of a title text and a summary text.
4. The method according to claim 3, wherein: The method further comprises: Performing speech recognition on the audio information in the video resource, and / or performing text recognition on the text content in the video resource; Key information is extracted from the result of speech recognition and / or the result of text recognition to obtain the summary text.
5. The method according to claim 1, wherein: The step of determining the image feature vector of the video resource comprises: Extracting multiple frames of images from the video resource according to a predetermined rule; Determine the image feature vectors of the multiple frames of images, and use the image feature vectors of the multiple frames of images as the image feature vectors of the video resource.
6. The method according to claim 1, wherein: The encoding model includes an encoder of a conversion model, and the encoder includes a four-layer encoding structure.
7. A method for training an encoding model, comprising: Determine a plurality of sample pairs, wherein the sample pairs include a positive sample pair and a negative sample pair; wherein the positive sample pair includes information of two video samples of the same type, and the negative sample pair includes information of two video samples of different types; wherein the information of the video samples includes the video samples and comment text for the video samples; A contrastive learning training method is adopted to train the encoding model using the multiple sample pairs.
8. The method according to claim 7, wherein: The training method of using contrastive learning to train the coding model using the multiple sample pairs includes: The information of each video sample in the sample pair is processed in the following manners: obtaining content text contained in the video sample, performing text conversion on the content text to obtain a content vector representation, and performing text conversion on the comment text for the video sample to obtain a comment vector representation; determining an image feature vector of the video sample based on multiple frames of images in the video resource; inputting the content vector representation, the comment vector representation and the image feature vector into a first encoding model to be trained to obtain a multimodal fusion vector of the video sample; Calculate the similarity between the multimodal fusion vectors of the two video samples of the sample pair, and adjust the first encoding model according to the similarity and the type of the sample pair to obtain a trained second encoding model; wherein the type of the sample pair is a positive sample pair or a negative sample pair.
9. The method according to claim 7 or 8, further comprising, obtaining the comment text for the video sample, comprising: Obtaining one or more comment texts whose attention levels meet preset requirements from a plurality of comment texts for the video sample; The degree of attention is represented by at least one of the number of likes, the number of reposts, and the number of replies.
10. The method according to claim 8, wherein: The content text contained in the video sample includes at least one of a title text and a summary text.
11. The method according to claim 10, wherein: The method further comprises: Performing speech recognition on the audio information in the video sample, and / or performing text recognition on the text content in the video sample; Key information is extracted from the result of speech recognition and / or the result of text recognition to obtain the summary text.
12. The method according to claim 8, wherein: The determining of the image feature vector of the video sample comprises: Extracting multiple frames of images from the video sample according to a predetermined rule; Determine the image feature vectors of the multiple frames of images, and use the image feature vectors of the multiple frames of images as the image feature vectors of the video samples.
13. The method according to claim 7 or 8, wherein: The encoding model includes an encoder of a conversion model, and the encoder includes a four-layer encoding structure.
14. The method according to claim 7 or 8, wherein: Determining the positive sample pair includes: Acquire multiple video files watched by a user within a predetermined time period, and determine the multiple video files as a video file set of the user; Two video files are randomly selected from the video file collection of the user to determine the positive sample pair.
15. The method according to claim 7 or 8, wherein: Determining the negative sample pair includes: For multiple users, obtain multiple video files watched by each user within a predetermined time period to determine a video file set of each user; A video file is randomly selected from the video file sets of two users to determine the negative sample pair.
16. A video resource representation device, comprising: A first acquisition module, used to acquire video resources and comment texts for the video resources; A second acquisition module is used to acquire the content text contained in the video resource; A vector conversion module, used for performing text conversion on the content text to obtain a content vector representation; and performing text conversion on the comment text to obtain a comment vector representation; Determine an image feature vector of the video resource based on multiple frames of images in the video resource; An input module is used to input the content vector representation, the comment vector representation and the image feature vector into a coding model to obtain a multimodal fusion vector of the video resource.
17. A training device for a coding model, comprising: A sample pair determination module, used to determine a plurality of sample pairs, wherein the sample pairs include positive sample pairs and negative sample pairs; wherein the positive sample pairs include information of two video samples of the same type, and the negative sample pairs include information of two video samples of different types; wherein the video sample information includes the video sample and comment text for the video sample; The training module is used to train the encoding model using the multiple sample pairs by adopting a contrastive learning training method.
18. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 15.
19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-15.
20. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 15.
Citation Information
Cited By
Content interaction method for AI online education
CN120183260A