Video and copywriting correlation judgment method and device, equipment and medium

By acquiring target videos and text, determining feature sequences, and inputting them into a large model, the problem of inaccurate determination of the relevance between videos and text in existing technologies is solved, achieving higher determination accuracy and a lower false positive rate.

CN121935409APending Publication Date: 2026-04-28BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-10-28
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, methods for determining the relevance of videos and text are usually based solely on text or images, which makes it difficult to fully capture the various features of a video, resulting in inaccurate judgment results.

Method used

By acquiring the target video and target text, the target feature sequence in the text space is determined and input into the target large model, and the target judgment result is output, including one of multiple different candidate judgment results, thereby improving the judgment accuracy.

Benefits of technology

It enables the determination of the relevance between videos and text, improving the accuracy of the determination results and reducing the rate of false judgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935409A_ABST
    Figure CN121935409A_ABST
Patent Text Reader

Abstract

The invention provides a video and copywriting correlation judgment method and device, equipment and a medium. A specific implementation mode of the method comprises the steps of obtaining a target video and a target text; the target text comprises a target copywriting and guide information for the video; the guide information comprises a plurality of different judgment results to be selected; the to-be-selected judgment result indicates the correlation between the video and the copywriting; determining a target feature sequence in a text space based on the target video and the target text; inputting the target feature sequence into a target large model to obtain a target judgment result output by the target large model; the target judgment result is one of the plurality of different judgment results to be selected. According to the embodiment, the correlation between the video and the copywriting can be judged, the accuracy of the judgment result is improved, and the rate of misjudgment is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, device, and medium for determining the relevance of video and text. Background Technology

[0002] With the continuous development of the internet and media technology, many platforms have emerged to provide video services to users. The quantity and quality of videos disseminated online are also constantly improving, covering a wide range of themes and styles. To enhance user experience, video platforms need to ensure a high degree of relevance between the video's visual content and its corresponding text. Currently, methods for determining this relevance are typically based solely on text or images, making it difficult to comprehensively capture the diverse features of a video. Therefore, a solution for determining the relevance between video and text is needed. Summary of the Invention

[0003] This disclosure provides a method, apparatus, device, and medium for determining the relevance of video and text.

[0004] According to the first aspect, a method for determining the relevance of video and text is provided, the method comprising:

[0005] Acquire the target video and target text; the target text includes target text and guidance information for the video; the guidance information includes multiple different candidate judgment results; the candidate judgment results indicate the relevance between the video and the text.

[0006] Based on the target video and the target text, determine the target feature sequence in the text space;

[0007] The target feature sequence is input into the target large model to obtain the target determination result output by the target large model; the target determination result is one of the multiple different candidate determination results.

[0008] According to the second aspect, a device for determining the relevance of video and text is provided, the device comprising:

[0009] The acquisition module is used to acquire target video and target text; the target text includes target text and guidance information for the video; the guidance information includes multiple different candidate judgment results; the candidate judgment results indicate the relevance between the video and the text.

[0010] The determination module is used to determine a target feature sequence in the text space based on the target video and the target text;

[0011] The determination module is used to input the target feature sequence into the target large model and obtain the target determination result output by the target large model; the target determination result is one of the multiple different candidate determination results.

[0012] According to a third aspect, a computer-readable storage medium is provided, the storage medium storing a computer program that, when executed by a processor, implements the method described in any one of the first aspects.

[0013] According to a fourth aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any one of the first aspects.

[0014] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:

[0015] This disclosure provides a video-text relevance determination scheme. It acquires a target video and target text, where the target text includes target text for the video and guiding information. The guiding information includes multiple candidate determination results, each indicating the relevance between the video and the text. Based on the target video and target text, a target feature sequence in the text space is determined, and this sequence is input into a target large-scale model to obtain the target determination result output by the model. This target determination result is one of multiple candidate determination results. This approach enables the determination of the relevance between the video and the text, improves the accuracy of the determination result, and reduces the false positive rate.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This disclosure is a schematic diagram illustrating a scenario for determining the relevance of a video and text according to an exemplary embodiment;

[0019] Figure 2 This disclosure is a schematic diagram illustrating a large model training scenario according to an exemplary embodiment;

[0020] Figure 3 This is a flowchart illustrating a method for determining the relevance of video and text according to an exemplary embodiment of this disclosure;

[0021] Figure 4This disclosure is a flowchart illustrating a training method for a large model according to an exemplary embodiment;

[0022] Figure 5 This is a block diagram illustrating a video-text relevance determination device according to an exemplary embodiment of the present disclosure;

[0023] Figure 6 This is a schematic block diagram of another electronic device provided in some embodiments of this disclosure. Detailed Implementation

[0024] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0025] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0026] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0027] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0028] With the continuous development of the internet and media technology, many platforms have emerged to provide video services to users. The quantity and quality of videos disseminated online are also constantly improving, covering a wide range of themes and styles. To enhance user experience, video platforms need to ensure a high degree of relevance between the video's visual content and its corresponding text. However, current methods for determining this relevance are often based solely on text or images, making it difficult to comprehensively capture the diverse characteristics of a video.

[0029] This disclosure provides a method for determining the relevance between video and text. The method acquires a target video and target text, where the target text includes target text for the video and guiding information. The guiding information includes multiple candidate judgment results, each indicating the relevance between the video and the text. Based on the target video and target text, a target feature sequence in the text space is determined. This target feature sequence is then input into a target large-scale model to obtain the target judgment result output by the model. This target judgment result is one of multiple different candidate judgment results. This method enables the determination of the relevance between video and text, improves the accuracy of the judgment result, and reduces the false positive rate.

[0030] See Figure 1 This is a schematic diagram illustrating a scenario for determining the relevance of a video to text, according to an exemplary embodiment.

[0031] like Figure 1 As shown, firstly, the video A1 and text A2 to be judged, as well as guiding information A3, can be obtained. Specifically, video A1 can be a short video within a preset duration; text A2 can include text information describing the content of video A1, and text A2 can further include timestamp information; guiding information A3 can include multiple different candidate judgment results indicating the relevance between the video and the text. For example, guiding information A3 can include the following: [Judging whether the input video and text are related: 1. The text content is strongly related to the video, and the text timestamp is aligned with the video. 2. The text content is strongly related to the video, but the text timestamp is not aligned with the video. 3. The text content is strongly related to the video. 4. The text content is not related to the video.]

[0032] Next, video A1 can be input into a visual encoder, which extracts feature sequence B1 of video A1 in visual space. Feature sequence B1 is then input into an adapter, which maps it to text space, resulting in feature sequence B2 in text space. Simultaneously, text A2 and guiding information A3 are input into a text segmenter, which extracts feature sequences B3 of text A2 and guiding information A3 in text space. Feature sequences B2 and B3 are then merged and input into a large model M. Based on feature sequences B2 and B3, the large model M outputs the target judgment result selected from multiple different candidate judgment results (e.g., any number from 1 to 4).

[0033] For example, if the text A2 includes or excludes a timestamp, the large model M might output 3 or 4. Output 3 indicates that the text content is strongly related to the video footage, while output 4 indicates that the text content is not related to the video footage. If the text A2 includes a timestamp, the large model M might output 1, 2, or 4. Output 1 indicates that the text content is strongly related to the video footage, and the timestamp is aligned with the video footage; output 2 indicates that the text content is strongly related to the video footage, but the timestamp is not aligned with the video footage; and output 4 indicates that the text content is not related to the video footage.

[0034] It should be noted that the visual encoder and text segmenter can be models trained using any reasonable method, and the large model M and adapter can be obtained through pre-training and retraining. The training process for the large model M and adapter can be found in [reference needed]. Figure 2 The provided examples.

[0035] See Figure 2 This is a schematic diagram illustrating a large model training scenario according to an exemplary embodiment.

[0036] like Figure 2 As shown, firstly, we can obtain the large model M1 to be trained, the adapter, and the pre-trained visual encoder and text segmenter. We first pre-train the large model M1 to obtain the large model M2, enabling it to generate matching text for videos. For example, after pre-training, we input the guiding words into the text segmenter to obtain feature sequence C1, input the video into the visual encoder, and input the output of the visual encoder into the adapter to obtain feature sequence C2. We then merge feature sequences C1 and C2 and input them into the large model M2. The large model M2 can output relevant text generated for the aforementioned video, the content of which matches the video. Furthermore, the adapter can be pre-trained simultaneously during the pre-training of the large model M1.

[0037] After obtaining the large model M2, both the large model M2 and the pre-trained adapter can be retrained. First, the large model M2 can be used to obtain a training dataset. Specifically, multiple sets of sample data can be obtained, including sample videos and sample texts. The sample videos and sample texts included in these sample data sets may be related or unrelated, and whether the sample videos and sample texts in the sample data sets are related is unknown.

[0038] Next, using the large model M2, the relevance label for each sample data group is determined. For example, taking sample data group Z as an example, sample data group Z includes sample video Z1 and sample text Z2. First, feature sequences can be extracted from sample text Z2 and input into the large model M2 to obtain the perplexity output by the large model M2 as the relevance value X1. Then, feature sequences are extracted from sample text Z2 and sample video Z1 respectively and merged. The merged feature sequence is input into the large model M2 to obtain the perplexity output by the large model M2 as the relevance value X2. The difference between the relevance value X1 and the relevance value X2 can be calculated. If the difference is greater than or equal to a preset threshold Y1, the relevance label for sample video Z1 and sample text Z2 is determined to be relevant; if the difference is less than the preset threshold Y2, the relevance label for sample video Z1 and sample text Z2 is determined to be irrelevant.

[0039] After determining the relevance labels for each set of sample data, a training dataset can be constructed based on each set of sample data and its relevance labels. For example, sample data sets with irrelevant labels from multiple sample data sets can be labeled as Class 4 training data. Alternatively, sample videos and texts from a small number of different sample data sets with relevant labels from multiple sample data sets can be recombined to obtain sample data sets with irrelevant labels, which can then be labeled as Class 4 training data. Sample data sets with some relevant labels from multiple sample data sets can be labeled as Class 3 training data. Selecting some sample data sets with relevant labels from multiple sample data sets, the sample videos and texts in these sample data sets are parsed to obtain matching timestamps. Some timestamps are added to the matching sample texts and, together with the relevant sample videos, form a new sample data set, labeled as Class 1 training data. Another portion of timestamps, after adding time perturbations, is added to the matching sample texts and, together with the relevant sample videos, forms a new sample data set, labeled as Class 2 training data. The training data is categorized into four types: Type 1, where the text content is strongly correlated with the video footage and the timestamps are aligned with the video footage; Type 2, where the text content is strongly correlated with the video footage but the timestamps are not aligned with the video footage; Type 3, where the text content is strongly correlated with the video footage but has no timestamps; and Type 4, where the text content is unrelated to the video footage. These four types of training data form the training dataset.

[0040] Finally, the large model M2 and the adapter are trained using the training dataset, and the model parameters are updated to obtain the trained large model M. The large model M can then be applied to... Figure 1 In the scenario shown in the embodiment, the relevance between the video and the text is determined, and if the video and the text are related and the text includes a timestamp, it is determined whether the timestamps are aligned.

[0041] The present disclosure will now be described in detail with reference to specific embodiments.

[0042] Figure 3 This is a flowchart illustrating a method for determining the relevance of video and text according to an exemplary embodiment. This method can be applied to a terminal device. In this embodiment, for ease of understanding, an example is given using a terminal device capable of installing a large language model. Those skilled in the art will understand that the terminal device may include, but is not limited to, mobile terminal devices such as smartphones, smart wearable devices, tablets, etc. The method may include the following steps:

[0043] like Figure 3 As shown, in step 301, the target video and target text are acquired.

[0044] In this embodiment, the target text includes target text for the video and guidance information. The target text can be text information determined to be relevant to the target video; it may or may not be related to the target video. The guidance information may include multiple different candidate judgment results, which can indicate the relevance between the video and the text. Therefore, the candidate judgment results may include at least relevance information indicating the relevance between the video and the text. For example, the relevance information in the candidate judgment results may include, but is not limited to, indications that the video and text are strongly related, indications that the video and text are weakly related, and indications that the video and text are not related.

[0045] In some implementations, in addition to relevance information indicating the relevance between the video and the text, some of the candidate judgment results may also include alignment information for the text with timestamps. This alignment information indicates the alignment of the timestamps in the video and the text. For example, the alignment information in the candidate judgment results may include, but is not limited to, indications of timestamp alignment between the video and the text, and indications of timestamp misalignment between the video and the text.

[0046] In this embodiment, the target text may or may not have a timestamp. If the target text has a timestamp, and the relevance between the target video and the target text meets a preset condition, the target determination result may include alignment information between the target video and the timestamp. If the target text does not have a timestamp, the target determination result does not include alignment information between the target video and the timestamp. For example, multiple candidate determination results may include: video is related to text and the timestamp is aligned with the video; video is related to text but the timestamp is not aligned with the video; video is related to text but has no timestamp; video is unrelated to text. If the target text to be determined does not include a timestamp, the target determination result may be "video is related to text but has no timestamp" or "video is unrelated to text". If the target text to be determined includes a timestamp, the target determination result may be one of the three candidate determination results other than "video is related to text but has no timestamp".

[0047] It should be noted that the relevance between the target video and the target text satisfies the preset condition if the relevance is greater than a preset level. For example, if the relevance information in the candidate judgment results includes indications that the video and text are strongly related, and indications that the video and text are not related, then the relevance between the target video and the target text satisfies the preset condition if the video and text are strongly related. If the relevance information in the candidate judgment results includes indications that the video and text are strongly related, indications that the video and text are weakly related, and indications that the video and text are not related, then the relevance between the target video and the target text satisfies the preset condition if the video and text are either strongly related or weakly related.

[0048] In step 302, a target feature sequence in the text space is determined based on the target video and the target text.

[0049] In this embodiment, a target feature sequence in the text space can be determined based on the target video and the target text. Specifically, the target text can be segmented to obtain a first feature sequence in the text space; for example, a text segmenter can be used to process the target text. Simultaneously, the target video can be visually encoded to obtain a second feature sequence in the visual space, and this second feature sequence is mapped to the text space to obtain a third feature sequence in the text space. For example, an adapter can be used to map the second feature sequence to the text space. Finally, the first and third feature sequences are merged to obtain the target feature sequence.

[0050] In step 303, the target feature sequence is input into the target large model to obtain the target determination result output by the target large model.

[0051] In this embodiment, the target feature sequence can be input into the target large-scale model to obtain the target determination result output by the target large-scale model. This target determination result can be one of several different candidate determination results included in the guidance information. The target large-scale model can be a large-scale language model involving natural language processing. For the training process of the target large-scale model, please refer to [link to relevant documentation]. Figure 4 The provided examples.

[0052] This disclosure provides a method for determining the relevance between video and text. The method acquires a target video and target text, where the target text includes target text for the video and guiding information. The guiding information includes multiple candidate judgment results, each indicating the relevance between the video and the text. Based on the target video and target text, a target feature sequence in the text space is determined. This target feature sequence is then input into a target large-scale model to obtain the target judgment result output by the model. This target judgment result is one of multiple different candidate judgment results. This method enables the determination of the relevance between video and text, improves the accuracy of the judgment result, and reduces the false positive rate.

[0053] Figure 4 This is a flowchart illustrating a training method for a large model according to an exemplary embodiment. This embodiment describes the specific training process of the target large model, including the following steps:

[0054] like Figure 4 As shown, in step 401, the pre-trained initial large model is obtained.

[0055] In this embodiment, the initial large model can be a pre-trained large model capable of generating matching text for videos. Specifically, firstly, a large model to be trained can be obtained, and the large model to be trained can be pre-trained using sample videos, so that the large model can generate matching text related to the input video. The matching text can be a description or extension of the video content.

[0056] In step 402, multiple sets of sample data, including sample videos and sample texts, are obtained. In step 403, the relevance labels corresponding to the sample data sets are determined using the initial large model.

[0057] In this embodiment, multiple sets of sample data are acquired, each set including a sample video and a sample text. An initial large model can be used to process each set of sample data to obtain its corresponding relevance label. Specifically, for any set of sample data, the relevance label can be determined as follows: First, the initial large model can be used to process the sample text included in the set to obtain a first relevance value. For example, the sample text and preset guidance information included in the set can be input into a text segmenter to obtain the feature sequence corresponding to the sample text, and then the feature sequence corresponding to the sample text can be directly input into the initial large model to obtain the first relevance value. Alternatively, the sample text included in the set can be input into a text segmenter to obtain the feature sequence corresponding to the sample text, the video obtained based on random noise can be input into a visual encoder, and the output of the visual encoder can be input into an adapter to obtain the feature sequence corresponding to the video with random noise. The feature sequence corresponding to the sample text and the feature sequence corresponding to the video with random noise can be merged, and the result can be input into the initial large model to obtain the first relevance value.

[0058] It should be noted that the first relevant value can be the perplexity calculated by the initial large model based on the input feature sequence. This perplexity can be used to represent the sum of the probability losses for each character predicted by the large model in a given text. For example, given the text "QWERTYU" (where each letter represents a character in the text), the probability loss for predicting Q, the probability loss for predicting W, ..., the probability loss for predicting U can be determined by summing the probability losses corresponding to all letters in the given text as the aforementioned perplexity.

[0059] Next, the initial large model processes the sample videos and sample text included in the sample data set to obtain a second correlation value. For example, the sample text and preset guidance information included in the sample data set can be input into a text segmenter to obtain the feature sequence corresponding to the sample text. The sample videos included in the sample data set can be input into a visual encoder, and the output of the visual encoder can be input into an adapter to obtain the feature sequence corresponding to the sample videos. The feature sequences corresponding to the sample text and the sample videos are merged, and the result is input into the initial large model to obtain the second correlation value. This second correlation value is also the perplexity calculated by the initial large model based on the input feature sequences.

[0060] Finally, the difference between the first and second relevance values ​​is determined, and based on this difference, the relevance label corresponding to this set of sample data is determined. Since the initial large model has the ability to generate matching text for videos, and both the first and second relevance values ​​represent the perplexity calculated by the initial large model based on the input feature sequence, this perplexity reflects the sum of probability losses of the initial large model in predicting the input text feature sequence. Therefore, a larger first relevance value / second relevance value indicates a lower relevance between the processed text and the input video, and a smaller first relevance value / second relevance value indicates a higher relevance between the processed text and the input video. When only sample text is input, without sample video (or with noisy video), the obtained first relevance value can be used as a reference value for relevance.

[0061] In one implementation, after inputting sample text and sample video, if the difference between the obtained second correlation value and the first correlation value is greater than or equal to a preset first difference, it indicates that the correlation between the input sample text and sample video is low, and the correlation between the sample text and sample video can be determined as no correlation. If the difference between the obtained second correlation value and the first correlation value is less than a preset second difference, it indicates that the correlation between the input sample text and sample video is high, and the correlation between the sample text and sample video can be determined as correlation. Here, the first difference and the second difference can be the same or different. If the first difference and the second difference are different, the first difference should be greater than the second difference.

[0062] It should be noted that more intervals can be set for the difference between the second and first correlation values. The correlation label corresponding to the difference falling in different intervals is different. The correlation label corresponding to the sample data group can be determined based on this difference and the different intervals.

[0063] For example, we can set K1 and K2, where K1 is greater than K2. If the difference between the second correlation value and the first correlation value obtained from the sample data group Y is greater than or equal to K1, then the correlation label for the sample data group Y can be determined to be a strong correlation between video and text. If the difference between the second correlation value and the first correlation value is less than K1 but greater than or equal to K2, then the correlation label for the sample data group Y can be determined to be a weak correlation between video and text. If the difference between the second correlation value and the first correlation value is less than or equal to K2, then the correlation label for the sample data group Y can be determined to be no correlation between video and text.

[0064] In step 404, a training dataset is constructed based on the sample data group and the corresponding relevance labels of the sample data group. In step 405, the initial large model is updated using the training dataset to obtain the target large model.

[0065] In this embodiment, after determining the sample data groups and their corresponding relevance labels, a training dataset can be constructed using these sample data groups and their corresponding relevance labels. For example, sample data groups with irrelevant relevance labels can be extracted and labeled as the n1th class of training data. Alternatively, multiple sample data groups with relevant relevance labels can be extracted and recombined (i.e., combining sample videos and sample texts from different sample data groups) and labeled as the n1th class of training data. The remaining sample data groups with relevant relevance labels are divided into three parts: Z1, Z2, and Z3. Sample data group Z1 is labeled as the n2th class of training data. Sample data groups Z2 and Z3 are analyzed for their sample videos and sample texts to obtain timestamps. Sample data group Z2 is taken, its timestamps are added to the sample texts, and it is labeled as the n3th class of training data. The timestamps corresponding to sample data group Z3 are perturbed, and the perturbed timestamps are added to its sample texts and labeled as the n4th class of training data. The training dataset is constructed from the training data of class n1, class n2, class n3, and class n4.

[0066] After obtaining the training dataset, the initial large model can be trained using the training dataset to update the parameters of the initial large model and obtain the target large model.

[0067] Because this embodiment uses a pre-trained initial large model to determine the relevance labels corresponding to the sample data groups, and constructs a training dataset based on the relevance labels, the trained target large model can accurately determine the relevance between the video and the text, reducing the false judgment rate.

[0068] It should be noted that although the operations of the methods of this disclosure embodiment are described in a specific order in the above embodiments, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowcharts may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0069] Corresponding to the aforementioned embodiments of the video-text relevance determination method, this disclosure also provides embodiments of the video-text relevance determination device.

[0070] like Figure 5 As shown, Figure 5 This disclosure is a block diagram of a video-text relevance determination device according to an exemplary embodiment. The device may include: an acquisition module 501, a determination module 502, and a determination module 503.

[0071] The acquisition module 501 is used to acquire the target video and target text. The target text includes the target text and guidance information for the video. The guidance information includes multiple candidate judgment results, which indicate the relevance between the video and the text.

[0072] The determination module 502 is used to determine the target feature sequence in the text space based on the target video and the target text.

[0073] The determination module 503 is used to input the target feature sequence into the target large model and obtain the target determination result output by the target large model. The target determination result is one of several different candidate determination results.

[0074] In some implementations, the candidate determination results include relevance information used to indicate the relevance between the video and the text. Among the multiple different candidate determination results, some candidate determination results also include alignment information provided for the text with timestamps, which is used to indicate the alignment between the timestamps in the video and the text.

[0075] In other implementations, the target text may or may not have a timestamp. If the target text has a timestamp and the relevance between the target video and the target text meets a preset condition, the target determination result includes alignment information between the target video and the timestamp.

[0076] In other embodiments, the determining module 502 is configured to: perform word segmentation on the target text to obtain a first feature sequence of the target text in the text space; perform visual encoding on the target video to obtain a second feature sequence of the target video in the visual space; map the second feature sequence to the text space to obtain a third feature sequence in the text space; and merge the first feature sequence and the third feature sequence to obtain the target feature sequence.

[0077] In other implementations, the target large model can be trained as follows: A pre-trained initial large model is obtained, which has the ability to generate matching text for videos. Multiple sets of sample data, including sample videos and sample text, are obtained. Using the initial large model, relevance labels corresponding to the sample data sets are determined. Based on the sample data sets and their corresponding relevance labels, a training dataset is constructed. Using the training dataset, the initial large model is updated to obtain the target large model.

[0078] In other implementations, for any given set of sample data, the relevance label corresponding to that set of sample data can be determined using an initial large model as follows: The initial large model is used to process the sample text included in the sample data set to obtain a first relevance value; the initial large model is then used to process the sample videos and sample text included in the sample data set to obtain a second relevance value. The difference between the first and second relevance values ​​is determined, and based on this difference, the relevance label corresponding to that set of sample data is determined.

[0079] In other implementations, a training dataset can be constructed based on a set of sample data and the corresponding correlation labels as follows: a set of reference sample data that meets the preset correlation conditions is extracted from multiple sets of sample data, the timestamps corresponding to the reference sample data sets are obtained, and a training dataset is constructed based on the timestamps and the reference sample data sets.

[0080] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the embodiments of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0081] This disclosure provides an electronic device with several embodiments. The electronic device includes a processor and a memory, and can be used to implement a client or server. The memory stores computer-executable instructions (e.g., one or more computer program modules) non-transitoryly. The processor executes the computer-executable instructions, which, when executed by the processor, can perform one or more steps in the video-text relevance determination method described above, thereby implementing the video-text relevance determination method described above. The memory and processor can be interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0082] For example, a processor can be a central processing unit (CPU), a graphics processing unit (GPU), or other form of processing unit with data processing and / or program execution capabilities. For instance, a CPU can be based on x86 or ARM architectures. A processor can be a general-purpose processor or a special-purpose processor, capable of controlling other components in an electronic device to perform desired functions.

[0083] For example, the memory may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB storage, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and the processor may run one or more computer program modules to implement various functions of the electronic device. Various application programs and various data, as well as various data used and / or generated by the application programs, may also be stored in the computer-readable storage medium.

[0084] It should be noted that, in the embodiments of this disclosure, the specific functions and technical effects of the electronic devices can be referred to the description of the correlation determination method between video and text above, and will not be repeated here.

[0085] Figure 6 This is a schematic block diagram of another electronic device provided in some embodiments of this disclosure. The electronic device 920 is, for example, suitable for implementing the video and text relevance determination method provided in embodiments of this disclosure. The electronic device 920 can be a terminal device, etc., and can be used to implement a client or server. The electronic device 920 can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. It should be noted that... Figure 6 The illustrated electronic device 920 is merely an example and does not impose any limitation on the functionality and scope of use of the embodiments of this disclosure.

[0086] like Figure 6 As shown, the electronic device 920 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 921, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 922 or a program loaded from a storage device 928 into a random access memory (RAM) 923. The RAM 923 also stores various programs and data required for the operation of the electronic device 920. The processing unit 921, ROM 922, and RAM 923 are interconnected via a bus 924. An input / output (I / O) interface 925 is also connected to the bus 924.

[0087] Typically, the following devices can be connected to I / O interface 925: input devices 926 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 927 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 928 including, for example, magnetic tapes, hard disks, etc.; and communication devices 929. Communication device 929 allows electronic device 920 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 6 An electronic device 920 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 920 may alternatively implement or have more or fewer devices.

[0088] For example, according to embodiments of this disclosure, the above-described method for determining the relevance of video and text can be implemented as a computer software program. For instance, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for executing the above-described method for determining the relevance of video and text. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 929, or installed from a storage device 928, or installed from a ROM 922. When the computer program is executed by a processing device 921, it can implement the functions defined in the method for determining the relevance of video and text provided in embodiments of this disclosure.

[0089] This disclosure provides a storage medium in some embodiments. For example, the storage medium may be a non-transitory computer-readable storage medium for storing non-transitory computer-executable instructions. When the non-transitory computer-executable instructions are executed by a processor, the video-text relevance determination method described in the embodiments of this disclosure can be implemented. For example, when the non-transitory computer-executable instructions are executed by a processor, one or more steps in the video-text relevance determination method described above can be performed.

[0090] For example, the storage medium can be used in the aforementioned electronic device; for instance, the storage medium may include the memory in the electronic device.

[0091] For example, the storage medium may include a memory card for a smartphone, a storage component for a tablet computer, a hard disk for a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, or other suitable storage media.

[0092] For example, the description of the storage medium can be found in the description of the memory in the embodiments of the electronic device, and will not be repeated here. The specific functions and technical effects of the storage medium can be found in the description of the relevance determination method between video and text above, and will not be repeated here.

[0093] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0094] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0095] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for determining the relevance of a video to its text, the method comprising: Obtain the target video and target text; The target text includes targeted copy and guidance information for the video; The guidance information includes multiple different candidate judgment results; The candidate determination results indicate the relevance between the video and the text; Based on the target video and the target text, determine the target feature sequence in the text space; The target feature sequence is input into the target large model to obtain the target determination result output by the target large model; The target determination result is one of the multiple different candidate determination results.

2. The method according to claim 1, wherein, The candidate judgment results include relevance information used to indicate the relevance between the video and the text; some of the multiple different candidate judgment results also include alignment information provided for text with timestamps, the alignment information being used to indicate the alignment between the timestamps in the video and the text.

3. The method according to claim 1, wherein, The target text may or may not have a timestamp; if the target text has a timestamp and the relevance between the target video and the target text meets a preset condition, then the target determination result includes the alignment information between the target video and the timestamp.

4. The method according to claim 1, wherein, The step of determining the target feature sequence in the text space based on the target video and the target text includes: The target text is segmented to obtain the first feature sequence of the target text in the text space; The target video is subjected to visual encoding processing to obtain a second feature sequence of the target video in visual space; The second feature sequence is mapped to the text space to obtain the third feature sequence in the text space; The first feature sequence and the third feature sequence are merged to obtain the target feature sequence.

5. The method according to claim 1, wherein, The target large model is trained in the following manner: Obtain a pre-trained initial large model, which has the ability to generate matching text for videos; Obtain multiple sets of sample data, including sample videos and sample text. Using the initial large model, determine the correlation labels corresponding to the sample data groups; A training dataset is constructed based on the sample data set and the corresponding correlation labels of the sample data set. Using the training dataset, the initial large model is updated to obtain the target large model.

6. The method according to claim 5, wherein, For any given set of sample data, using the initial large model, determine the relevance label corresponding to that set of sample data, including: The sample texts included in the sample data set are processed using the initial large model to obtain a first relevance value; The sample videos and sample texts included in the sample data set are processed using the initial large model to obtain a second relevance value; Determine the difference between the first correlation value and the second correlation value; Based on the difference, the correlation label corresponding to the sample data group is determined.

7. The method according to claim 5, wherein, The step of constructing a training dataset based on the sample data set and the corresponding relevance labels of the sample data set includes: Extract a portion of the reference sample data sets whose correlation meets preset conditions from the multiple sets of sample data sets; Obtain the timestamp corresponding to the reference sample data group; A training dataset is constructed based on the timestamps and the reference sample data set.

8. A device for determining the relevance of a video to text, the device comprising: The acquisition module is used to acquire the target video and target text; The target text includes targeted copy and guidance information for the video; The guidance information includes multiple different candidate judgment results; the candidate judgment results indicate the relevance between the video and the text. The determination module is used to determine a target feature sequence in the text space based on the target video and the target text; The determination module is used to input the target feature sequence into the target large model and obtain the target determination result output by the target large model; The target determination result is one of the multiple different candidate determination results.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-7.

10. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-7.