Video clip screening and screening model training method and device, medium

By combining multi-dimensional visual elements and dialogue semantic information into the video clip selection model for scoring and adjusting the model parameters, the problem of low efficiency in video clip selection in existing technologies is solved, and the effect of efficient and automated selection of core video clips is achieved.

CN122200504APending Publication Date: 2026-06-12BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610309917.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-13
Publication Date
2026-06-12

Smart Images

  • Figure CN122200504A_ABST
    Figure CN122200504A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a method and device for screening a video clip and training a screening model, and a medium, wherein the method comprises: inputting first sample data and a first clip screening prompt word into a first video clip screening model, obtaining first video association information of a first core video clip output by the first video clip screening model, performing format matching on the information format of the first video association information and the information format of a plurality of standard video association information, determining a first loss value according to the format matching degree, performing content matching on the information content of the first video association information and the information content of the plurality of standard video association information pre-labeled, determining a second loss value according to the content matching degree, determining a target loss value according to the first loss value and the second loss value, and training the first video clip screening model according to the target loss value to obtain a target video clip screening model. In the technical solution, the efficiency and reliability of the video clip screening are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a method, apparatus, and medium for screening video clips and training a screening model. Background Technology

[0002] With the rise of short videos, selecting video clips has become a mainstream demand. For example, selecting core video clips of the same storyline and using them as lead-generating video clips for corresponding videos.

[0003] In related technologies, core video segments are selected through manual screening, with individuals repeatedly watching and comparing the footage to choose the most suitable segments. This method of selecting core video segments is relatively inefficient. Summary of the Invention

[0004] To solve the above-mentioned technical problems, or at least partially solve them, this disclosure provides a method, apparatus, and medium for selecting video segments and training a selection model.

[0005] This disclosure provides a training method for a video segment selection model. The method includes: inputting first sample data and first segment selection prompts into a first video segment selection model, wherein the first sample data includes multiple video segments belonging to the same storyline and multiple segments of dialogue semantic information corresponding one-to-one with the multiple video segments; the first segment selection prompts contain instructions for the first video segment selection model to: score each video segment based on multi-dimensional visual elements and dialogue semantic information to obtain a score value; select the core video segment with the highest score among the multiple video segments; and output text of multiple types of standard video association information of the core video segment, wherein the multiple types of standard video association information include at least one of the following of the core video segment: selection reason, score value, plot type, and video segment identification information; wherein the multiple types of standard video association information are output by the first video segment selection model instructed by the first segment selection prompts. The process involves: obtaining video association information; acquiring the first video association information of the first core video segment output by the first video segment selection model, wherein the first video association information is the video association information actually output by the first video segment selection model; matching the information format of the first video association information with the information format of the multi-class standard video association information, and determining a first loss value based on the format matching degree; matching the information content of the first video association information with the information content of the pre-annotated multi-class standard video association information, and determining a second loss value based on the content matching degree; determining a target loss value based on the first loss value and the second loss value, and determining whether the target loss value is greater than a first preset loss value threshold; and adjusting the matrix parameters of the visual layer in the first video segment selection model until the target loss value is not greater than the first preset loss value threshold, thereby obtaining the trained target video segment selection model.

[0006] This disclosure provides a method for filtering video clips, comprising: inputting video clip data and pre-set clip filtering prompts into a pre-trained target video clip filtering model, wherein the video clip data includes multiple video clips belonging to the same storyline and multiple segments of dialogue semantic information corresponding one-to-one with the multiple video clips; obtaining third video association information of a third core video clip output by the target video clip filtering model, wherein the third core video clip is the core video clip with the highest score among the multiple video clips after the target video clip filtering model scores each video clip based on multi-dimensional visual elements and dialogue semantic information to obtain a score value, and the third video association information of the third core video clip includes: reason for selection, score value, plot type, and video clip identification information, wherein the target video clip filtering model is trained using the above-described training method for video clip filtering models.

[0007] This disclosure also provides a training device for a video clip selection model. The device includes: a first input module for inputting first sample data and first clip selection prompts into a first video clip selection model, wherein the first sample data includes multiple video clips belonging to the same storyline and multiple segments of dialogue semantic information corresponding one-to-one with the multiple video clips; the first clip selection prompts include text instructing the first video clip selection model to: score each video clip based on multi-dimensional visual elements and dialogue semantic information to obtain a score value, select the core video clip with the highest score among the multiple video clips, and output text of multiple types of standard video association information of the core video clip, wherein the multiple types of standard video association information include at least one of the following of the core video clip: selection reason, score value, plot type, and video clip identification information, wherein the multiple types of standard video association information are video association information output by the first video clip selection model instructed by the first clip selection prompts; and an acquisition module. The system comprises the following modules: a first determination module, used to obtain first video association information of the first core video segment output by the first video segment selection model, wherein the first video association information is the video association information actually output by the first video segment selection model; a first determination module, used to perform format matching between the information format of the first video association information and the information format of the multi-class standard video association information, and determine a first loss value based on the format matching degree; a second determination module, used to perform content matching between the information content of the first video association information and the information content of the pre-labeled multi-class standard video association information, and determine a second loss value based on the content matching degree; and a training processing module, used to determine a target loss value based on the first loss value and the second loss value, and determine whether the target loss value is greater than a first preset loss value threshold. If it is greater than the first preset loss value threshold, the matrix parameters of the visual layer in the first video segment selection model are adjusted until the target loss value is not greater than the first preset loss value threshold, thereby obtaining the trained target video segment selection model.

[0008] This disclosure also provides a video clip filtering device, comprising: a second input module for inputting video clip data and pre-set clip filtering prompts into a pre-trained target video clip filtering model, wherein the video clip data includes multiple video clips belonging to the same storyline and multiple segments of dialogue semantic information corresponding one-to-one with the multiple video clips; and an acquisition module for acquiring third video association information of a third core video clip output by the target video clip filtering model, wherein the third core video clip is the core video clip with the highest score among the multiple video clips after the target video clip filtering model scores each video clip based on multi-dimensional visual elements and dialogue semantic information to obtain a score value, and the third video association information of the third core video clip includes: reason for selection, score value, plot type, and video clip identification information, wherein the target video clip filtering model is trained using the training method for video clip filtering models described above.

[0009] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement a training method for a video segment selection model as provided in this disclosure, or a video segment selection method.

[0010] This disclosure also provides a computer-readable storage medium storing a computer program for executing a training method for a video segment selection model as provided in this disclosure, or a video segment selection method.

[0011] The technical solution provided in this disclosure has the following advantages compared with the prior art: The video segment selection and selection model training scheme provided in this embodiment inputs first sample data and first segment selection prompts into a first video segment selection model. The first sample data includes multiple video segments belonging to the same storyline and multiple segments of dialogue semantic information corresponding one-to-one with each video segment. The first segment selection prompts instruct the first video segment selection model to: score each video segment based on multi-dimensional visual elements and dialogue semantic information to obtain a score value; select the core video segment with the highest score among the multiple video segments; and output text containing multiple types of standard video association information for the core video segment. The multiple types of standard video association information include at least one of the following for the core video segment: selection reason, score value, plot type, and video segment identification information. The multiple types of standard video association information are the video segments output by the first video segment selection model as instructed by the first segment selection prompts. The first video segment selection model obtains the first video association information of the first core video segment output by the first video segment selection model. This first video association information is the actual video association information output by the first video segment selection model. Then, the format of the first video association information is matched with the format of multiple types of standard video association information. A first loss value is determined based on the format matching degree. The content of the first video association information is matched with the content of pre-labeled multiple types of standard video association information. A second loss value is determined based on the content matching degree. Finally, a target loss value is determined based on the first and second loss values. It is then determined whether the target loss value is greater than a first preset loss threshold. If it is greater than the first preset loss threshold, the matrix parameters of the visual layer in the first video segment selection model are adjusted until the target loss value is no greater than the first preset loss threshold, thus obtaining the trained target video segment selection model. This technical solution improves the efficiency and reliability of video segment selection. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0013] Figure 1 A flowchart illustrating a training method for a video segment selection model provided in this embodiment of the disclosure; Figure 2 A flowchart illustrating a training method for another video segment selection model provided in this embodiment of the disclosure; Figure 3 A flowchart illustrating a training method for another video segment selection model provided in this disclosure embodiment; Figure 4This is a flowchart illustrating a video clip selection method according to an embodiment of the present disclosure; Figure 5 A schematic diagram of the structure of a training device for a video clip selection model provided in an embodiment of this disclosure; Figure 6 A schematic diagram of the structure of a video clip filtering device provided in an embodiment of this disclosure; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0016] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0017] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0018] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0019] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0020] To address the aforementioned issues, this disclosure provides a method for training a video segment selection model, which will be described below with reference to specific embodiments.

[0021] Figure 1 This is a flowchart illustrating a training method for a video segment selection model provided in an embodiment of the present disclosure. The method can be executed by a training device for the video segment selection model, which can be implemented in software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, the method includes: Step 101: Input the first sample data and the first segment selection prompts into the first video segment selection model. The first sample data includes multiple video segments belonging to the same storyline and multiple segments of dialogue semantic information corresponding to each video segment. The first segment selection prompts instruct the first video segment selection model to: score each video segment based on multi-dimensional visual elements and dialogue semantic information to obtain a score value, select the core video segment with the highest score among multiple video segments, and output the text of multiple types of standard video association information of the core video segment. The multiple types of standard video association information include at least one of the following of the core video segment: reason for selection, score value, plot type, and video segment identification information. The multiple types of standard video association information is the video association information output by the first video segment selection model as instructed by the first segment selection prompts.

[0022] The first sample data includes multiple video clips belonging to the same storyline. The same storyline refers to different plots, characters, events, and content segments belonging to the same coherent narrative thread. They revolve around the same core background, main plot, or core setting, and are logically, temporally, or causally related to each other. They are not independent, fragmented content, but ultimately collectively support a complete story framework. In some possible embodiments, the first sample data may also include video clip links belonging to the same storyline. The first video clip filtering model has the function of retrieving video clips from the video library based on these links.

[0023] The first sample data also includes: multiple segments of dialogue semantic information that correspond one-to-one with multiple video clips. The dialogue semantic information is obtained by semantic recognition of the dialogue information. Alternatively, in addition to semantic recognition of the dialogue information, pre-set plot description information can also be obtained as a type of dialogue semantic information. The dialogue semantic information is used to help the model understand the plot of the video clips.

[0024] In the embodiments of this disclosure, the first segment selection prompt refers to keywords or phrases that can accurately filter out video segments that meet the requirements during video segment selection. The first segment selection prompt includes instructions for the first video segment selection model to: score each video segment based on multi-dimensional visual elements and dialogue semantic information to obtain a score value; select the core video segment with the highest score among multiple video segments; and output text containing multiple types of standard video association information for the core video segment. The multi-dimensional visual elements may include elements such as character actions, facial expressions, environmental changes, camera movements, and conflict intensity in the scene. These visual elements help the model understand the plot type and story information, and find the video segments with the greatest dissemination value. For example, core video clips are typically video segments involving action conflict, chases, tension, comedic moments, plot twists, or emotional outbursts. The basic model architecture of the first video clip selection model can be understood as an artificial intelligence model. This AI model can score each video clip based on its understanding and the instructions from the first clip selection prompts. In some possible embodiments, the first clip selection prompts can also indicate the priority levels of multiple plot types, define the definition information for each plot type, and define the priority level of each plot type within the multiple plot types. Scoring can be strictly restricted to the priority level, ensuring that plot types with higher priority levels receive higher scores than those with lower priority levels. In short, scoring rules can be arbitrarily defined based on the needs of the scenario, using the first clip selection prompts.

[0025] The reason for selection indicates why the core video segment was selected. It can be generated by the first video segment selection model itself. For example, the reason for selection could be because there is action scene, the protagonist's secret was discovered, or there is emotional scene. The score indicates the reliability of the core story segment. The plot type is related to predefined categories. For example, the plot type can include action type, ethical type, etc. The video segment identification information can be pre-set video sequence number, video link, or other information that can uniquely identify the video. Users can obtain the corresponding core video segment based on the video segment identification information.

[0026] The first video segment selection model can be an open-source artificial intelligence model, etc. In this embodiment, the first sample data and the first segment selection prompt words are input into the first video segment selection model. It is important to emphasize that the multi-class standard video association information is the video association information output by the first video segment selection model as instructed by the first segment selection prompt words; that is, the multi-class standard video association information indicates an ideal video association information that the first video segment selection model is expected to output. The multi-class standard video association information is the pre-standardized standard video association information of the first video segment.

[0027] Step 102: Obtain the first video association information of the first core video segment output by the first video segment filtering model, wherein the first video association information is the video association information actually output by the first video segment filtering model.

[0028] In the embodiments of this disclosure, first video association information of the first core video segment selected by the first video segment filtering model is obtained. The first core video segment and the first video association information are in one-to-one correspondence, and the first core video segment can be one or more. The first video association information is the video association information actually output by the first video segment filtering model.

[0029] Step 103: Match the information format of the first video association information with the information formats of multiple types of standard video association information, and determine the first loss value based on the format matching degree.

[0030] In this embodiment, the information format of the first video-related information is matched with the information formats of multiple types of standard video-related information. A first loss value is determined based on the format matching degree. The first loss value reflects the matching degree between the first video-related information and multiple types of standard video-related information in terms of information completeness and other formats.

[0031] In one embodiment of this disclosure, the information format of the first video association information is matched with the information format of the multiple types of standard video association information, and a first loss value is determined based on the format matching degree, including: Determine whether the text format of the first video-related information is consistent with the predefined text format, and determine the first reference value based on the determination result; In this embodiment, it is determined whether the text format of the first video association information is consistent with the predefined text format. For example, the predefined text format is the data exchange format JSON. The multiple types of standard video association information are output in JSON format, which includes fields and corresponding field values. The number of fields corresponds to the number of multiple types of standard video association information, and the field value of each field is the video association information of the corresponding class.

[0032] In this embodiment, the text format of the first video association information can be extracted as the main evaluation object. For example, the text format of the first video association information can be represented as: ,in, The text format is the text format of the first video association information, and Prase is the JASON format of the first video association information C output by the first video segment filtering model.

[0033] In this embodiment, it is determined whether the text format of the first video-related information is consistent with a predefined text format. For example, when it is determined that the text format of the first video-related information is consistent with the predefined text format, a pre-set first preset value is used as a first reference value. The first preset value is a predefined value, such as S1. When it is determined that the text format of the first video-related information is inconsistent with the predefined text format, a second preset value is used as the first reference value. The second preset value is less than the first preset value. The second preset value can be S2, which is less than S1. S2 can be greater than 0, thereby avoiding zero reward in extreme cases and ensuring the stability of the training process.

[0034] And / or, Determine whether the first video association information contains multiple types of standard video association information, and determine the second reference value based on the determination result; In this embodiment, it is determined whether the first video association information contains multiple types of standard video association information, that is, the second reference value is determined based on the completeness of the categories of video association information included in the first video association information.

[0035] Continuing with the predefined text format JSON as the data exchange format, in this embodiment, the extracted JSON content undergoes a field integrity check. The focus is on whether the fields included in the first video association information contain fields corresponding to multiple types of standard video association information. For example, multiple fields may include the reason for selection, score, plot type, and video segment identification information. Each time the first video association information is identified as containing one of these fields, a complete score can be obtained, thereby encouraging the model to output results with complete structure and information.

[0036] In this embodiment, the output of formula (1) can be used as the second reference value, wherein, in formula (1), The second reference value is k, which is a field included in the first video association information. This is an indicator function; it returns 1 if the field exists, and 0 otherwise.

[0037] Formula (1) In this embodiment, in order to improve the calculation efficiency of the second reference value, the number of information classes of multiple standard video association information included in the first video association information can be directly obtained according to the judgment result. The product of the number of information classes and the preset coefficient is used as the second reference value. The preset coefficient can be a positive number less than 1, such as 0.15.

[0038] And / or, Determine whether the plot type in the first video's associated information belongs to a preset plot type set, and determine the third reference value based on the determination result; In this embodiment, considering that a set of plot types will be preset based on the application scenario, the plot types output by the model will be constrained to be within the range of the preset set of plot types. The preset set of plot types includes multiple plot types, which can be labeled according to the actual needs of the scenario. For example, it may include action plots, love plots, family ethics plots, etc.

[0039] In this embodiment, when it is determined that the plot type in the first video-related information belongs to a preset plot type set, a pre-set third preset value is used as a third reference value. When it is determined that the plot type in the first video-related information does not belong to the preset plot type set, a pre-set fourth preset value is used as a third reference value. The fourth preset value is less than the third preset value. Both the third and fourth preset values ​​can be set according to the needs of the scenario. For example, the third preset value is 0 and the fourth preset value is -1.

[0040] And / or, Compare the byte length of the filtered reason in the first video association information with the size of the preset byte length range, and determine the fourth reference value based on the comparison result; In this embodiment, to avoid the model generating text that is too short or too long, a soft constraint is applied to the reasons for filtering the output. If the content is too short, it may mean that there is insufficient information and will be deducted appropriately. If the content is too long, it may contain redundant information and will also be slightly penalized. This mechanism encourages the model to achieve a balance between sufficient information and concise expression.

[0041] In this embodiment, the byte length of the filtered reason in the first video association information is compared with the size relationship of a preset byte length range. A fourth reference value is determined based on the comparison result. In some possible embodiments, when it is found that the byte length of the filtered reason is within the preset byte length range, a pre-set fifth preset value is used as the fourth reference value. When it is found that the byte length of the filtered reason is not within the preset byte length range, a pre-set sixth preset value is used as the fourth reference value. The sixth preset value is less than the fifth preset value. Both the fifth and sixth preset values ​​can be set according to the needs of the scenario. For example, the fifth preset value is 0 and the sixth preset value is -0.05.

[0042] After obtaining the aforementioned reference values, in this embodiment, a first loss value is determined based on at least one of the first, second, third, and fourth reference values. For example, if the reference values ​​include multiple of the first, second, third, and fourth reference values, the average of the multiple reference values ​​can be taken, and the average can be normalized to a normalized value between 0 and 1. The obtained normalized value is then used as the first loss value.

[0043] For example, when the reference values ​​include the first reference value, the second reference value, the third reference value, and the fourth reference value, the first loss value can be calculated using the following formula (2), where S is the first loss value and clip indicates that the result is restricted to the interval [0,1]. As the first reference value, This is the second reference value. This is the third reference value. This is the fourth reference value.

[0044] Formula (2) Step 104: Match the information content of the first video association information with the information content of the pre-labeled multi-class standard video association information, and determine the second loss value based on the content matching degree.

[0045] In this embodiment, the information content of the first video association information is matched with the information content of multiple pre-labeled standard video association information, and a second loss value is determined based on the content matching degree. That is, the second loss value is used to reflect the accuracy of the information content of the first video association information.

[0046] It should be noted that the method for determining the second loss value based on the information content matching degree between the first video association information and the pre-annotated standard video association information of the first video association information differs in different application scenarios, as shown in the following example: In some possible examples, the semantic matching degree between the information content of the first video association information and the pre-annotated standard video association information of the first video association information is directly calculated, and the semantic matching degree is determined as the second loss value.

[0047] In some possible examples, when the first core video segment is at least one, such as Figure 2 As shown, the information content of the first video association information is matched with the standard information content of pre-labeled multi-class standard video association information, and a second loss value is determined based on the content matching degree, including: Step 201: Match the plot type in each first video association information with the plot type in the corresponding standard video association information to determine the number of first core video segments corresponding to the successfully matched plot types.

[0048] In this embodiment, the number of first core video segments corresponding to the successfully matched video segment identifier information is first counted, that is, the number of correctly selected first core video segments is determined.

[0049] Step 202: Calculate the ratio of the number of successfully matched first core video segments to the total number of first core video segments.

[0050] In this embodiment, the ratio of the number of successfully matched first core video segments to the total number of first core video segments is calculated.

[0051] In this embodiment, the output of the first video segment selection model is assumed to be... Where N is the total number of the first core video segments output. The reason for selecting the i-th first core video segment is... For the i-th plot type, the pre-labeled standard video association information belonging to the same storyline is: , where M is the number of core video segments corresponding to the pre-annotated standard video association information.

[0052] Therefore, in this embodiment, determining whether each first core video segment belongs to a correct prediction involves determining the standard video association information corresponding to each first story core segment based on the video segment identification information. The correctness of the prediction result can be expressed as follows: , among which, = At that time, determine =1 indicates that the i-th first core video segment belongs to a correct prediction, and the i-th first core video segment is a successfully matched first core video segment; otherwise... =0 indicates that the i-th first core video segment belongs to an incorrect prediction, and the i-th first core video segment is a first core video segment that failed to match. Therefore, As a quantity ratio.

[0053] Step 203: Obtain the semantic similarity between the filtered reasons in each first video association information and the filtered reasons in the corresponding standard video association information.

[0054] Since the reasons for filtering are usually natural language text, which cannot be directly matched using strict string matching, this embodiment performs matching based on the semantic dimension. In this embodiment, the semantic similarity between the reasons for filtering in each first video association information and the reasons for filtering in the corresponding standard video association information is obtained.

[0055] Step 204: Determine whether the semantic similarity is greater than the preset semantic similarity threshold, and determine the fifth reference value based on the determination result.

[0056] In this embodiment, it is determined whether the semantic similarity is greater than a preset semantic similarity threshold, and a fifth reference value is determined based on the determination result. In this embodiment, when it is determined that the semantic similarity is greater than the preset semantic similarity threshold, a pre-set seventh preset value is determined as the fifth reference value. When it is determined that the semantic similarity is not greater than the preset semantic similarity threshold, a pre-set eighth preset value is determined as the fifth reference value. The seventh preset value is greater than the eighth preset value. The seventh and eighth preset values ​​can be set according to the needs of the scenario. For example, the seventh preset value is 1 and the eighth preset value is 0.

[0057] Step 205: Calculate the mean of the reference values ​​of all fifth reference values ​​corresponding to all first core video segments.

[0058] In this embodiment, the mean of the reference values ​​of all fifth reference values ​​corresponding to all first core video segments is also calculated.

[0059] In this embodiment, let the vector encoding of the reason for the i-th first core video segment being selected be as follows: The corresponding vector encoding of the filtering reason in the standard video association information (that is, determining the standard filtering reason corresponding to the first core video segment by matching the video segment identifier information in the first video association information with the video segment identifier information in the labeled standard video association information) is as follows: ,in, Identify any text encoding model. In this embodiment, calculate... and The semantic similarity Ri is used to determine the fifth reference value when Ri is greater than t, given a preset semantic similarity threshold t. If the value is 1, then determine the fifth reference value. The value is 0. Calculate the mean of the reference values ​​for all fifth reference values ​​corresponding to all first core video segments. .

[0060] Step 206: Determine the second loss value based on the quantity ratio and the average of the reference values.

[0061] In this embodiment, the second loss value can be determined based on the quantity ratio or the average of reference values. For example, the second loss value can be determined. ,in, These are the preset weighting coefficients.

[0062] Step 105: Determine the target loss value based on the first loss value and the second loss value, and determine whether the target loss value is greater than the first preset loss value threshold.

[0063] Step 106: When the target loss value is greater than the first preset loss threshold, adjust the matrix parameters of the visual layer in the first video segment selection model until the target loss value is not greater than the first preset loss threshold, thus obtaining the trained target video segment selection model. In the embodiments of this disclosure, the target loss value is determined based on the first loss value and the second loss value. For example, the average of the first loss value and the second loss value is determined as the target loss value; or, the first loss value is multiplied by a preset first weight value to obtain a first product value, the second loss value is multiplied by a preset second weight value to obtain a second product value, and the first product value and the second product value are summed to obtain the target loss value.

[0064] In this embodiment, the trained target video segment selection model is obtained based on the target loss value.

[0065] In some possible embodiments, it is determined whether the target loss value is greater than a first preset loss threshold. The first preset loss threshold can be customized. If it is greater than the first preset loss threshold, the matrix parameters of a preset matrix in the visual layer of the first video segment selection model are adjusted until the target loss value is no greater than the first preset loss threshold, thus obtaining the trained target video segment selection model.

[0066] In this embodiment, when the target loss value is greater than a preset loss threshold, the matrix parameters of the preset matrix in the visual layer of the first video segment selection model are adjusted until the target loss value is no greater than the preset loss threshold, thus obtaining the trained target video segment selection model. In other words, in this embodiment, to improve the model training effect, the full model parameters in the first video segment selection model are not adjusted; instead, the matrix parameters of the preset matrix in the visual layer are adjusted. The visual layer is a partial model layer in the first video segment selection model, and the preset matrix can be a partial matrix such as the self-attention matrix in the visual layer. The preset matrix can be selected using algorithms such as Low-Rank Adaptation (LoRA). Therefore, in this embodiment, by inputting video segments, dialogue semantics, and other content, multimodal information is fully utilized to achieve the ability to intelligently select core story video segments in the storyline. Furthermore, during the model training phase, reinforcement learning is used to achieve customized optimization of the storyline's multimodal large model, enabling the model to accurately locate the core video segments in the storyline, and the output results are complete, logically coherent, and meet expectations. In some possible embodiments, the language layer in the first video segment selection model can also be trained based on the target loss value to obtain the target video segment selection model.

[0067] In one embodiment of this disclosure, to further ensure the effectiveness of the target video segment selection model, the first segment selection prompts can be fine-tuned synchronously during the training phase. Specifically, when the target loss value exceeds a preset loss threshold, the user can be prompted to input guiding words that highlight multi-dimensional information about the video segment. For example, the user can be prompted to input guiding words that include multi-dimensional signals such as character actions, facial expressions, environmental changes, camera movement, and conflict intensity. Additionally, the user can be prompted to input guiding words that instruct the model to differentiate scores based on multi-dimensional signals such as character actions, facial expressions, environmental changes, camera movement, and conflict intensity. This ensures that different video segments have interpretable and non-overlapping scoring reasons, thereby achieving true discriminative power. Thus, the trained target video segment selection model can automatically select the most exciting segments within the same storyline, significantly improving selection efficiency and objectivity.

[0068] In one embodiment of this disclosure, in order to further improve the stability of the trained target video segment selection model, the basic model structure of the first video segment selection model can be trained before inputting the first sample data and the first segment selection prompt words into the first video segment selection model.

[0069] In this embodiment, as Figure 3 As shown, before inputting the first sample data and the first segment selection prompts into the first video segment selection model, the method further includes: Step 301: Input the second sample data and the first segment selection prompts into the first video segment selection model. The second sample data includes multiple video segments belonging to the same storyline and multiple dialogue semantic information corresponding to the multiple video segments.

[0070] In this embodiment, the second sample data and the first segment filtering prompt words are input into the first video segment filtering model. The second sample data includes multiple video segments belonging to the same storyline and multiple lines of semantic information corresponding to the multiple video segments. The second sample data can be the same as or different from the first sample data.

[0071] Step 302: Obtain the second video association information of the second core video segment selected by the first video segment filtering model.

[0072] In this embodiment, the second video association information of the second core video segment selected by the first video segment filtering model is obtained.

[0073] Step 303: Calculate the information matching degree between the second video association information and the multi-class standard video association information, and determine the third loss value based on the information matching degree.

[0074] In this embodiment, a third loss value is determined based on the matching degree of the second video association information and the multi-class standard video association information of the second core video segment. That is, in this embodiment, the matching degree of the second video association information and the pre-annotated multi-class standard video association information of the second core video segment is directly determined. For example, the semantic similarity of the second video association information and the pre-annotated multi-class standard video association information of the second core video segment is calculated, and the semantic similarity is used as the matching degree.

[0075] In this embodiment, the matching degree can be directly used as the third loss value.

[0076] Step 304: When the third loss value is greater than the second preset loss value threshold, adjust the model parameters of other model layers in the first video segment selection model other than the visual layer until the third loss value is not greater than the second preset loss value threshold.

[0077] In this implementation, the first preset model layer is a separate model layer from the visual layer of the initial video clip selection model. The training workload for these other model layers is relatively small. These other model layers can be the language layer (Large Language Model, LLM) of the first video clip selection model, and the preset matrix can be an LLM layer, thus improving training efficiency. This implementation freezes the training of the visual layer, significantly reducing the training workload. The initial video clip selection model can also be an open-source artificial intelligence model.

[0078] In summary, the training method for the video segment selection model in this embodiment of the present disclosure inputs first sample data and first segment selection prompts into the first video segment selection model. The first sample data includes multiple video segments belonging to the same storyline and multiple segments of dialogue semantic information corresponding one-to-one with the multiple video segments. The first segment selection prompts include instructions for the first video segment selection model to: score each video segment based on multi-dimensional visual elements and dialogue semantic information to obtain a score value; select the core video segment with the highest score among the multiple video segments; and output text containing multiple types of standard video association information for the core video segment. The multiple types of standard video association information include at least one of the following for the core video segment: selection reason, score value, plot type, and video segment identification information. The multiple types of standard video association information are the video links output by the first video segment selection model as instructed by the first segment selection prompts. The first video segment selection model obtains the first video association information of the first core video segment output by the first video segment selection model. This first video association information is the actual video association information output by the first video segment selection model. Then, the format of the first video association information is matched with the format of multiple types of standard video association information. A first loss value is determined based on the format matching degree. The content of the first video association information is matched with the content of pre-labeled multiple types of standard video association information. A second loss value is determined based on the content matching degree. Finally, a target loss value is determined based on the first and second loss values. It is then determined whether the target loss value is greater than a first preset loss threshold. If it is greater than the first preset loss threshold, the matrix parameters of the visual layer in the first video segment selection model are adjusted until the target loss value is no greater than the first preset loss threshold, thus obtaining the trained target video segment selection model. This technical solution improves the efficiency and reliability of video segment selection.

[0079] To implement the above embodiments, this disclosure also proposes a method for selecting video segments. Figure 4 This is a flowchart illustrating a video segment selection method according to an embodiment of the present disclosure, as shown below. Figure 4 As shown, the method includes: Step 401: Input the video segment data and the pre-set segment selection prompts into the pre-trained target video segment selection model. The video segment data includes multiple video segments belonging to the same storyline and multiple dialogue semantic information corresponding to the multiple video segments.

[0080] The video clip data includes multiple video clips belonging to the same storyline, and multiple segments of dialogue semantic information corresponding to each of the multiple video clips. The clip selection prompts are the first clip selection prompts in the above embodiment, or the updated first clip selection prompts.

[0081] In one embodiment of this disclosure, video segment data and pre-set segment selection prompts are input into a pre-trained target video segment selection model.

[0082] Step 402: Obtain the third video association information of the third core video segment output by the target video segment selection model. The third core video segment is the core video segment with the highest score among multiple video segments selected by the target video segment selection model after scoring each video segment based on multi-dimensional visual elements and dialogue semantic information. The third video association information of the third core video segment includes: the reason for selection, the score, the plot type, and the video segment identification information. The target video segment selection model is trained using the training method of the video segment selection model.

[0083] In this embodiment, the third video association information of the third core video segment output by the target video segment filtering model is obtained. The third core video segment is a high-energy video segment among multiple video segments. The third video association information of the third core video segment includes: the reason for being filtered, the score, the plot type, and the video segment identification information. Thus, the filtering model based on the trained video segments directly filters out the video segments used to describe the core story plot of the same storyline, improving the efficiency of video segment filtering.

[0084] In one embodiment of this disclosure, when there are multiple third core video segments, the multiple third core video segments are spliced ​​and edited to obtain a promotional video for the storyline. Thus, a promotional video for the storyline is generated in one step. This is a video segment with good promotional effect selected based on multimodal information, which helps to improve the video promotion effect.

[0085] In summary, the video segment selection method of this disclosure involves inputting video segment data and pre-set segment selection prompts into a pre-trained target video segment selection model. The video segment data includes multiple video segments belonging to the same storyline and multiple segments of dialogue semantic information corresponding one-to-one with each video segment. The method obtains the third video association information of the third core video segment output by the target video segment selection model. This third core video segment is the highest-scoring core video segment among the multiple video segments selected after the target video segment selection model scores each video segment based on multi-dimensional visual elements and dialogue semantic information. The third video association information of the third core video segment includes: the reason for selection, the score, the plot type, and video segment identification information. This improves the efficiency and reliability of video segment selection.

[0086] To implement the above embodiments, this disclosure also proposes a training apparatus for a video segment selection model.

[0087] Figure 5 This is a schematic diagram of a training device for a video segment selection model provided in an embodiment of this disclosure. The device can be implemented by software and / or hardware, and is generally integrated into an electronic device. Figure 5 As shown, the device includes: a first input module 510, an acquisition module 520, a first determination module 530, a second determination module 540, and a training processing module 550, wherein, The first input module 510 is used to input first sample data and first segment filtering prompts into the first video segment filtering model. The first sample data includes multiple video segments belonging to the same storyline and multiple segments of dialogue semantic information corresponding to the multiple video segments. The first segment filtering prompts include text instructing the first video segment filtering model to: score each video segment based on multi-dimensional visual elements and dialogue semantic information to obtain a score value, filter out the core video segment with the highest score among the multiple video segments, and output text of multiple types of standard video association information of the core video segment. The multiple types of standard video association information include at least one of the following of the core video segment: reason for being filtered, score value, plot type, and video segment identification information. The multiple types of standard video association information is the video association information output by the first video segment filtering model instructed by the first segment filtering prompts. The acquisition module 520 is used to acquire the first video association information of the first core video segment output by the first video segment selection model, wherein the first video association information is the video association information actually output by the first video segment selection model. The first determining module 530 is used to perform format matching between the information format of the first video association information and the information format of the multiple types of standard video association information, and determine a first loss value based on the format matching degree. The second determining module 540 is used to perform content matching between the information content of the first video association information and the information content of the pre-annotated multi-class standard video association information, and determine a second loss value based on the content matching degree. The training processing module 550 is used to determine a target loss value based on the first loss value and the second loss value, and to determine whether the target loss value is greater than a first preset loss value threshold. If it is greater than the first preset loss value threshold, the matrix parameters of the visual layer in the first video segment selection model are adjusted until the target loss value is not greater than the first preset loss value threshold, thereby obtaining the trained target video segment selection model.

[0088] The training apparatus for the video segment selection model provided in this disclosure can execute the training method for the video segment selection model provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0089] To implement the above embodiments, this disclosure also proposes a video clip filtering device.

[0090] Figure 6 This is a schematic diagram of a video clip filtering device provided in an embodiment of this disclosure. Figure 6 As shown, the device includes: a second input module 610 and an acquisition module 620, wherein, The second input module 610 is used to input video segment data and pre-set segment selection prompts into a pre-trained target video segment selection model. The video segment data includes multiple video segments belonging to the same storyline and multiple lines of semantic information corresponding to the multiple video segments. The acquisition module 620 is used to acquire the third video association information of the third core video segment output by the target video segment selection model. The third core video segment is the core video segment with the highest score among multiple video segments selected by the target video segment selection model after scoring each video segment based on multi-dimensional visual elements and dialogue semantic information. The third video association information of the third core video segment includes: the reason for selection, the score, the plot type, and the video segment identification information. The target video segment selection model is trained by the above-mentioned video segment selection model training method. The video segment selection device provided in this embodiment can execute the video segment selection method provided in any embodiment of this disclosure and has the corresponding functional modules and beneficial effects of the execution method.

[0091] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program / instruction, which, when executed by a processor, implements the training method of the video segment selection model or the video segment selection method in the above embodiments.

[0092] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.

[0093] The following is a detailed reference. Figure 7The diagram illustrates a structural schematic suitable for implementing the electronic device 700 in the embodiments of this disclosure. The electronic device 700 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0094] like Figure 7 As shown, the electronic device 700 may include a processor (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a memory 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processor 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0095] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0096] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from memory 708, or installed from ROM 702. When the computer program is executed by processor 701, it performs the functions defined in the training method of the video segment selection model or the video segment selection method of embodiments of this disclosure.

[0097] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0098] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0099] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0100] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the training method for the video segment selection model or the video segment selection method.

[0101] Electronic devices can be programmed with computer program code in one or more programming languages ​​or combinations thereof to perform the operations of this disclosure. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0103] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0104] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0105] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0106] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0107] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0108] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A training method for a video clip selection model, characterized in that, include: The first sample data and the first segment selection prompt are input into the first video segment selection model. The first sample data includes multiple video segments belonging to the same storyline and multiple segments of dialogue semantic information corresponding to each of the multiple video segments. The first segment selection prompt contains text instructing the first video segment selection model to: score each video segment based on multi-dimensional visual elements and dialogue semantic information to obtain a score value, select the core video segment with the highest score among the multiple video segments, and output the text of multiple types of standard video association information of the core video segment. The multiple types of standard video association information include at least one of the following of the core video segment: reason for selection, score value, plot type, and video segment identification information. The multiple types of standard video association information are the video association information output by the first video segment selection model instructed by the first segment selection prompt. Obtain the first video association information of the first core video segment output by the first video segment filtering model, wherein the first video association information is the video association information actually output by the first video segment filtering model; The information format of the first video association information is matched with the information format of the multiple types of standard video association information, and a first loss value is determined based on the format matching degree. The information content of the first video association information is matched with the information content of the pre-labeled multi-class standard video association information, and a second loss value is determined based on the content matching degree. A target loss value is determined based on the first loss value and the second loss value, and it is determined whether the target loss value is greater than a first preset loss value threshold. When the target loss value is greater than the first preset loss threshold, the matrix parameters of the visual layer in the first video segment selection model are adjusted until the target loss value is no greater than the first preset loss threshold, thus obtaining the trained target video segment selection model.

2. The method as described in claim 1, characterized in that, Before inputting the first sample data and the first segment filtering prompt text into the first video segment filtering model, the method further includes: The second sample data and the first segment filtering prompt text are input into the first video segment filtering model. The second sample data includes multiple video segments belonging to the same storyline and multiple segments of dialogue semantic information corresponding to the multiple video segments. Obtain the second video association information of the second core video segment output by the first video segment filtering model; Calculate the information matching degree between the second video association information and the multi-class standard video association information, and determine the third loss value based on the information matching degree; When the third loss value is greater than the second preset loss value threshold, the model parameters of other model layers in the first video segment selection model other than the visual layer are adjusted until the third loss value is not greater than the second preset loss value threshold.

3. The method as described in claim 1, characterized in that, The step of matching the information format of the first video association information with the information format of the multiple types of standard video association information, and determining a first loss value based on the format matching degree, includes: Determine whether the text format of the first video-related information is consistent with a predefined text format, and determine a first reference value based on the determination result; and / or, Determine whether the first video association information contains the multi-type standard video association information, and determine the second reference value based on the determination result; and / or, Determine whether the plot type in the first video association information belongs to a preset plot type set, and determine a third reference value based on the determination result; and / or, Compare the byte length of the filtered reason in the first video association information with the size relationship of the preset byte length range, and determine the fourth reference value based on the comparison result; The first loss value is determined based on at least one of the first reference value, the second reference value, the third reference value, and the fourth reference value.

4. The method as described in claim 3, characterized in that, The step of determining the first reference value based on the judgment result includes: When it is determined that the text format of the first video association information is consistent with the predefined text format, the first preset value is used as the first reference value. When it is determined that the text format of the first video-related information is inconsistent with the pre-defined text format, a pre-set second preset value is used as the first reference value, wherein the second preset value is less than the first preset value.

5. The method as described in claim 3, characterized in that, The step of determining the second reference value based on the judgment result includes: The number of information classes of the multiple standard video association information included in the first video association information is determined based on the judgment result; The product of the number of information types and the preset coefficient is used as the second reference value.

6. The method as described in claim 3, characterized in that, The determination of the third reference value based on the judgment result includes: When it is determined that the plot type in the first video association information belongs to a preset plot type set, a pre-set third preset value is used as the third reference value; When it is determined that the plot type in the first video association information does not belong to the preset plot type set, a pre-set fourth preset value is used as the third reference value, wherein the fourth preset value is less than the third preset value.

7. The method as described in claim 3, characterized in that, The step of determining the fourth reference value based on the comparison results includes: When the byte length of the reason for being filtered is found to be within the preset byte length range, the fifth preset value is used as the fourth reference value. When it is determined that the byte length of the filtered reason does not fall within the preset byte length range, a pre-set sixth preset value is used as the fourth reference value, wherein the sixth preset value is less than the fifth preset value.

8. The method as described in claim 1, characterized in that, The step of matching the information content of the first video association information with the pre-labeled standard information content of the multiple types of standard video association information, and determining the second loss value based on the content matching degree, includes: Match the plot type in each of the first video association information with the plot type in the corresponding standard video association information to determine the number of first core video segments corresponding to the successfully matched plot types; Calculate the ratio of the number of successfully matched first core video segments to the total number of first core video segments; Obtain the semantic similarity between the filtered reason in each of the first video association information and the filtered reason in the corresponding standard video association information; Determine whether the semantic similarity is greater than a preset semantic similarity threshold, and determine a fifth reference value based on the determination result; Calculate the mean of the reference values ​​of all the fifth reference values ​​corresponding to all the first core video segments; The second loss value is determined based on the ratio of the quantities and the average of the reference values.

9. The method as described in claim 8, characterized in that, The determination of the fifth reference value based on the judgment result includes: When it is determined that the semantic similarity is greater than the preset semantic similarity threshold, the pre-set seventh preset value is determined to be the fifth reference value; When it is determined that the semantic similarity is not greater than the preset semantic similarity threshold, the pre-set eighth preset value is determined to be the fifth reference value, wherein the seventh preset value is greater than the eighth preset value.

10. A method for selecting video clips, characterized in that, include: The video clip data and pre-set clip selection prompts are input into a pre-trained target video clip selection model. The video clip data includes multiple video clips belonging to the same storyline and multiple segments of dialogue semantic information that correspond one-to-one with the multiple video clips. Obtain the third video association information of the third core video segment output by the target video segment selection model, wherein the third core video segment is the core video segment with the highest score among the multiple video segments selected by the target video segment selection model after scoring each video segment based on multi-dimensional visual elements and dialogue semantic information to obtain a score value. The third video association information of the third core video segment includes: the reason for selection, the score value, the plot type, and the video segment identification information. The target video segment selection model is trained by the training method of the video segment selection model as described in any one of claims 1-9.

11. The method as described in claim 10, characterized in that, The method further includes: When there are multiple third core video segments, the multiple third core video segments are spliced ​​and edited to obtain the promotional video of the storyline.

12. A training device for a video clip selection model, characterized in that, include: The first input module is used to input first sample data and first segment filtering prompts into the first video segment filtering model. The first sample data includes multiple video segments belonging to the same storyline and multiple segments of dialogue semantic information corresponding one-to-one with the multiple video segments. The first segment filtering prompts include text instructing the first video segment filtering model to: score each video segment based on multi-dimensional visual elements and dialogue semantic information to obtain a score value, filter out the core video segment with the highest score value among the multiple video segments, and output text of multiple types of standard video association information of the core video segment. The multiple types of standard video association information include at least one of the following of the core video segment: reason for being filtered, score value, plot type, and video segment identification information. The multiple types of standard video association information are the video association information output by the first video segment filtering model instructed by the first segment filtering prompts. The acquisition module is used to acquire the first video association information of the first core video segment output by the first video segment selection model, wherein the first video association information is the video association information actually output by the first video segment selection model. The first determining module is used to perform format matching between the information format of the first video association information and the information format of the multiple types of standard video association information, and determine a first loss value based on the format matching degree. The second determining module is used to perform content matching between the information content of the first video association information and the information content of the pre-labeled multi-class standard video association information, and determine a second loss value based on the content matching degree. The training processing module is used to determine a target loss value based on the first loss value and the second loss value, and to determine whether the target loss value is greater than a first preset loss value threshold. If it is greater than the first preset loss value threshold, the matrix parameters of the visual layer in the first video segment selection model are adjusted until the target loss value is not greater than the first preset loss value threshold, thereby obtaining the trained target video segment selection model.

13. A video clip filtering device, characterized in that, include: The second input module is used to input video segment data and pre-set segment selection prompts into a pre-trained target video segment selection model. The video segment data includes multiple video segments belonging to the same storyline and multiple segments of dialogue semantic information corresponding to the multiple video segments. The acquisition module is used to acquire the third video association information of the third core video segment output by the target video segment selection model. The third core video segment is the core video segment with the highest score among the multiple video segments selected by the target video segment selection model after scoring each video segment based on multi-dimensional visual elements and dialogue semantic information. The third video association information of the third core video segment includes: the reason for selection, the score, the plot type, and the video segment identification information. The target video segment selection model is trained by the training method of the video segment selection model as described in any one of claims 1-9.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program for executing a training method for a video segment selection model as described in any one of claims 1-9, or a video segment selection method as described in claim 10 or 11.