Film series intelligent dynamic collection method and device based on multi-modal large model

By integrating film and television content features and series structure rule bases into a multimodal large model, the automated, intelligent, and dynamic collection of series has been achieved. This solves the problems of low efficiency and insufficient accuracy of manual operation, improves collection efficiency and accuracy, reduces costs, and enhances user experience.

CN121388198APending Publication Date: 2026-01-23SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511487367.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

In existing technologies, the collection of series relies on manual operation, which results in high labor costs, low efficiency, and insufficient accuracy. It is difficult to respond quickly to newly added content and cannot accurately identify potentially related film and television content.

Method used

It adopts a multimodal large model to integrate the text, audio and visual features of film and television content, combined with a pre-set series structure rule base, to achieve automated, intelligent and dynamic collection. Through multimodal feature extraction, reasoning and dynamic mapping comparison, adjustments are made in combination with user behavior feedback.

Benefits of technology

It improves the accuracy and efficiency of data collection, reduces labor costs, enhances user experience, and adapts to different types and changing user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121388198A_ABST
    Figure CN121388198A_ABST
Patent Text Reader

Abstract

The invention discloses a film series intelligent dynamic collection method and device based on a multi-modal large model, and belongs to the technical field of film and television content processing, and the method comprises the steps: extracting text information, audio information and visual information of film and television content; the text information is analyzed to extract key semantic features, the audio information is analyzed to extract emotion tone features, the visual information is analyzed to extract image hue, shot language and scene composition visual style features, and multi-modal feature vectors are obtained; inputting the multi-modal feature vectors into a pre-trained artificial intelligence large model, and inferring an association relationship between the film and television contents and a series episode category to which the film and television contents belong in combination with a preset series episode structure rule base; and carrying out dynamic mapping comparison on the inferred incidence relation between the film and television contents and the category of the series with the stock series, and storing a comparison result. The collection efficiency and accuracy are improved, the labor cost is reduced, and the experience of watching the series of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of film and television content processing, and in particular to a film series intelligent dynamic collection method and device based on a multi-modal large model, a server and a storage medium. BACKGROUND

[0002] In the field of film and television content processing, the collection and management of series is crucial to improving user viewing experience and enhancing platform content appeal. Currently, the traditional series collection method relies on manual operation, and the operation personnel need to manually select and organize related films according to limited information such as text introduction, actor lineup, and broadcast time of film and television content to build series.

[0003] This method of the prior art has many drawbacks: first, the labor cost is high, and manual selection and collection of massive film and television content requires a lot of time and manpower; second, the efficiency is low, and it is difficult to quickly respond to newly added film and television content, resulting in lagging series updates; third, the accuracy is insufficient, and manual judgment is easily affected by subjective factors, which cannot accurately identify series content with potential correlations, such as derivative series based on the same IP but with different narrative perspectives or forms, which are easily missed or incorrectly collected, seriously affecting users' complete acquisition and viewing experience of series content.

[0004] Therefore, the prior art still needs to be improved and developed. SUMMARY

[0005] To solve the above technical problems, the present application provides a film series intelligent dynamic collection method and device based on a multi-modal large model, a server and a storage medium, which fuses the text, audio, visual and other multi-modal features of film and television content, combines a preset series structure rule library, and uses the powerful reasoning and analysis capabilities of a large model to realize automatic, intelligent and dynamic collection of film series, improve collection efficiency and accuracy, reduce labor costs, and improve users' viewing experience of series.

[0006] The technical solution of the present application is as follows: A film series intelligent dynamic collection method based on a multi-modal large model, comprising: Obtaining film and television content, performing multi-modal feature extraction on the film and television content, extracting text information, audio information and visual information of the film and television content; Respectively analyzing and extracting key semantic features from the text information, analyzing and extracting emotional tone features from the audio information, and analyzing and extracting picture tone, shot language, scene composition visual style features from the visual information to obtain a multi-modal feature vector; The extracted multi-modal feature vector is input into a pre-trained artificial intelligence large model, and a preset series structure rule library is combined to infer the correlation between the video content and the series category to which the video content belongs. According to the established inventory series database, the inferred correlation between the video content and the series category to which the video content belongs is dynamically mapped and compared with the inventory series, and the updated series information is stored in the inventory series database.

[0007] The film series intelligent dynamic collection method based on the multi-modal large model, wherein the step of obtaining the video content and performing multi-modal feature extraction on the video content includes: A series structure rule library containing a plurality of series structure rules is constructed in advance, the plurality of series structure rules including IP derivative rules, regional content rules, and content type rules; the series structure rule library can be flexibly configured and updated according to actual needs.

[0008] The film series intelligent dynamic collection method based on the multi-modal large model, wherein the step of obtaining the video content and performing multi-modal feature extraction on the video content includes: Obtaining video content, wherein the video content includes newly published and stored video content and all video content in the inventory series database obtained periodically; Performing multi-modal feature extraction on the obtained video content to extract text information, audio information, and visual information of the video content.

[0009] The film series intelligent dynamic collection method based on the multi-modal large model, wherein the step of respectively analyzing and extracting key semantic features from the text information, analyzing and extracting emotional tone features from the audio information, and analyzing and extracting picture tone, shot language, scene composition visual style features from the visual information to obtain a multi-modal feature vector includes: Using natural language processing technology to perform word segmentation, part-of-speech tagging, and semantic analysis on the text information to extract key semantic features including character names, story locations, and core plots; An audio emotion analysis algorithm is used to process the audio information to analyze the emotional tone features of the dialogue tone and background music; Computer vision technology is used to analyze the visual information to extract picture tone, shot language, and scene composition visual style features to form a multi-modal feature vector.

[0010] The method further includes the following steps: inputting the extracted multi-modal feature vector into the pre-trained artificial intelligence large model; The pre-trained artificial intelligence large model filters out relevant series rules according to the preliminary matching of the multi-modal feature vector in the preset series structure rule library; and analyzes the fit degree of the characteristics of the video content and the rules to determine the series belonging of the video content, and infers the correlation between the video contents and the series category to which they belong.

[0011] The method further includes the following steps: According to the established inventory series database, the inferred correlation between the video contents and the series category to which they belong are dynamically mapped and compared with the inventory series in the inventory series database; The characteristics of the video content itself are associated, and user behavior feedback data is introduced as a dynamic adjustment factor; when a user frequently associates a new series with an existing series, the weight of the association is automatically increased; when a new series is detected to have a conflict or unreasonable collection with an existing series, the rules are adjusted accordingly; after comparison and adjustment, the updated series information is stored in the inventory series database.

[0012] The method further includes the following steps after the step of dynamically mapping and comparing the inferred correlation between the video contents and the series category to which they belong with the inventory series according to the established inventory series database, and storing the updated series information in the inventory series database: A visual human intervention interface is set up to receive user operation instructions through the human intervention interface to review the collection results; Through the system interface, user operation instructions are received to label and correct error collection cases, and automatically record error cases and modification records; The recorded error cases and modification records are fed back to the training data of the artificial intelligence large model to optimize the training process of the artificial intelligence large model in reverse.

[0013] An intelligent dynamic collection device for film series based on a multi-modal large model, wherein the device comprises: A series structure rule library module for pre-constructing a series structure rule library containing a plurality of series structure rules, the plurality of series structure rules including IP derivation rules, regional content rules, and content type rules; the series structure rule library can be flexibly configured and updated according to actual needs; An acquisition and feature extraction module for acquiring video content, performing multi-modal feature extraction on the video content, and extracting text information, audio information, and visual information of the video content; A multi-modal feature extraction module for respectively analyzing and extracting key semantic features from the text information, analyzing and extracting emotional tone features from the audio information, and analyzing and extracting picture tone, shot language, and scene composition visual style features from the visual information to obtain multi-modal feature vectors; A large model reasoning and matching module for inputting the extracted multi-modal feature vectors into a pre-trained artificial intelligence large model, combining a pre-set series structure rule library, and reasoning out the correlation between video contents and the series category to which the video contents belong; A dynamic mapping, comparison, and updating module for dynamically mapping and comparing the correlation between video contents and the series category to which the video contents belong, which are reasoned out, with inventory series, and storing the updated series information after comparison into the inventory series database.

[0014] A server, comprising a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, and the one or more programs comprise a method for executing any one of the methods.

[0015] A computer-readable storage medium, wherein when the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute any one of the methods.

[0016] As can be seen from the above, the present application provides an intelligent dynamic collection method, device, server, and storage medium for film series based on a multi-modal large model, which adopts multi-modal feature extraction, large model reasoning combined with series structure rule library matching, dynamic mapping, comparison, and updating. 1) Improved collection accuracy: By fusing multi-modal feature analysis and series structure rule library, the correlation between video contents can be more comprehensively and accurately identified, avoiding incorrect collection caused by single information analysis, and compared with traditional methods, the collection accuracy is improved by about 50%.

[0017] 2) Improved collection efficiency: The automatic series collection process greatly reduces manual operation and improves the processing speed of massive video content by about 5 times, enabling quick response to new content and timely updating of the series library.

[0018] 3) Reduced cost: Reduces a large amount of manpower investment, reduces the labor cost of video content management, and improves the operation efficiency of enterprises.

[0019] 4) Optimize user experience: Based on user behavior feedback, dynamically adjust the collection results, and provide more accurate and complete series content recommendations for users, enhance user satisfaction and stickiness to the platform.

[0020] 5) Strong adaptability: The series structure rule library can be flexibly configured and updated to adapt to different types, different styles of video content, and changing user needs and market trends. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0022] Figure 1 is the process schematic diagram of the film series intelligent dynamic collection method based on the multi-modal large model of the embodiment 1 of the present application.

[0023] Figure 2 is the process schematic diagram of the film series intelligent dynamic collection method based on the multi-modal large model of the embodiment 2 of the present application.

[0024] Figure 3 The principle block diagram of the film series intelligent dynamic collection device embodiment based on the multi-modal large model provided by the present application.

[0025] Figure 4 is the internal structure principle block diagram of the server provided by the embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical scheme and advantages of the present application more clear and definite, the present application will be further described in detail below with reference to the drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0027] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.

[0028] Current technology and traditional methods of categorizing TV series rely on manual operation. Operations personnel must manually select and organize relevant films and television programs based on limited information such as text descriptions, cast lists, and broadcast times. This method has several drawbacks: First, it is labor-intensive; manual selection and categorization of a massive amount of content requires significant time and manpower. Second, it is inefficient, struggling to quickly respond to newly added content, leading to delays in series updates. Third, it lacks accuracy; human judgment is easily influenced by subjective factors, making it difficult to accurately identify potentially related series. For example, derivative series based on the same IP but using different narrative perspectives or presentation styles are easily missed or incorrectly categorized, severely impacting users' complete access to and viewing experience of series content.

[0029] Although some existing technologies have developed methods for classifying and associating film and television content using algorithms, most of them are based on analysis of single textual or visual information, lacking a comprehensive consideration of the multi-dimensional characteristics of film and television content. They are difficult to effectively identify the complex internal connections between series content and cannot meet the needs of the film and television content management field for accurate and efficient collection of series.

[0030] To address the aforementioned technical problems, this invention provides a method for intelligent dynamic aggregation of film series based on a multimodal large model, as described in the following embodiments.

[0031] Example 1 like Figure 1 As shown in the figure, an intelligent dynamic aggregation method for film series based on a multimodal large model according to an embodiment of the present invention includes the following steps: Step S100: Obtain film and television content, and perform multimodal feature extraction on the film and television content to extract text information, audio information and visual information of the film and television content; In a specific embodiment of the present invention, film and television content can be obtained from video platforms (such as film and television databases and streaming media platforms), wherein the film and television content includes: newly released film and television content and all film and television content in the existing series database obtained periodically; The acquired film and television content undergoes multimodal feature extraction, extracting textual, audio, and visual information. Specifically, in this embodiment, multimodal refers to simultaneously extracting information from three different dimensions: text, audio, and visual, covering the textual, audio, and visual information of the film and television content.

[0032] In particular embodiments, the text information of the video content includes plot summaries, scripts, etc., the audio information includes dialogues, background music, etc., and the visual information includes screen images, shot transitions, etc.

[0033] Specifically, the present application relates to text information extraction, which is to extract key information related to text in video content, such as subtitle content (dialogue, narration) of a film, text in the opening and closing credits (film title, credits), and text elements in the picture (such as the title on a poster, road sign text in a scene), which will ultimately be converted into analyzable text data.

[0034] Specifically, the present application relates to audio information extraction, which is to extract sound-related features in videos, including the frequency of character dialogue, the rhythm and melody and volume of background music, environmental sounds (such as rain, footsteps, explosions), etc. In short, it is to convert what is heard into audio parameters (such as sound wave frequency, decibel value). Specifically, the present application relates to visual information extraction, which is to extract core features related to video pictures, such as facial expressions of characters, body movements, objects in the scene (such as tables and chairs, buildings), color style of the picture (such as cool color tone, warm color tone), shot movement (such as push shot, pull shot), etc. It is equivalent to breaking down what is seen into identifiable visual element data.

[0035] Step S200, respectively analyzing and extracting key semantic features from the text information, analyzing and extracting emotional tone features from the audio information, and analyzing and extracting picture tone, shot language, and scene composition visual style features from the visual information to obtain a multi-modal feature vector; In this embodiment, the natural language processing (NLP) technique can be used to analyze the semantic features of the text information; an audio emotion analysis algorithm can be used to process the audio information and analyze the features such as the tone of the dialogue and the emotional tone of the background music; and a computer vision (CV) technique can be used to analyze the visual information and extract visual style features such as picture tone, shot language, and scene composition, to form a multi-modal feature vector.

[0036] In further embodiments of the present application, the step S200 specifically includes: S201, using natural language processing technology to perform word segmentation, part-of-speech tagging, and semantic analysis on the text information, and extracting key semantic features including character names, story locations, and core plots; In this embodiment, the language analysis tool natural language processing (NLP) technique is used to mine key plot information from text data, and the operation and results include the following two parts: First, the basic processing link can first segment the text information, such as "the main character meets the villain in New York" into "main character / in / New York / meet / villain"; and perform part-of-speech tagging to label the attributes of each word, such as "main character" is a noun and "meet" is a verb, so that the scattered text becomes structured. Then, the core extraction link can understand the meaning behind the text through semantic analysis, and finally extract 3 types of key information, character names (such as "main character: Xiaoming, villain: Lao Wang"), story locations (such as "New York, suburban laboratory"), and core plots (such as "Xiaoming conflicts with Lao Wang in New York"), forming "semantic features" that can summarize the core of the plot. This step is to convert chaotic video and television text such as dialogues and subtitles into clear key information of the plot, facilitating quick understanding of the context of the video content.

[0037] In the embodiment of the present step, the audio emotion analysis algorithm of the sound analysis tool is used to judge the emotion conveyed by the audio, including analysis of two types of sound: first, the tone of the dialogue, specifically analyzing the tone, speed, and volume of the character's speech, such as "slow speed, low tone" analysis may be "sad", "fast speed, high volume" analysis may be "angry"; second, the emotional tone of the background music: judging the style and emotion of the music, such as "soft piano music" may be "warm", "rapid drum music" may be "tense". In the embodiment of the present step, these analysis results will be integrated to form features that can reflect the emotion of the audio, such as "the emotional tone of this audio: tense". Through analysis, the emotion of the sound can be understood, and the emotional atmosphere of the video segment such as warmth, tension, and sadness can be judged through audio.

[0038] In the embodiment of the present step, the computer vision technology is used to analyze the visual information and extract visual style features such as picture color, shot language, and scene composition to form a multi-modal feature vector.

[0039] In the embodiment of the present invention, the picture analysis tool such as computer vision technology can be used to disassemble visual elements and extract features that can reflect the style of the picture, including three types: first, picture color, specifically judging the color tendency of the overall picture, such as "predominantly blue" is a cool color, and "predominantly red" is a warm color; second, shot language, specifically identifying the shooting method of the shot, such as "the camera slowly approaches the character" is a push shot (highlighting the character's emotion), and "the camera shoots from high to low" is a low-angle shot (showing the overall scene); third, scene composition, specifically analyzing the arrangement of elements in the picture, such as "the character is in the center of the picture" (highlighting the main character) and "the picture is symmetrical on both sides" (creating a sense of stability). ​In the embodiment of the present application, the above extracted key semantic features, emotional tone features, picture color tone, shot language, scene composition visual style features are converted into multi-modal feature vectors that can be recognized by a computer.

[0040] In step S300, the extracted multi-modal feature vectors are input into a pre-trained artificial intelligence large model, and a preset series structure rule library is combined to infer the correlation between the video content and the series category to which the video content belongs. Before the embodiment of the present step is implemented, a series structure rule library containing a plurality of series structure rules needs to be pre-constructed, and the plurality of series structure rules include IP derivation rules, regional content rules, and content type rules.

[0041] In the embodiment of the present application, a database containing a plurality of series structure rules is constructed, including IP derivation rules (such as judgment standards and correlation relationships of different types such as explicit main transmission, prequel, sequel, and homin creation), regional content rules (such as rules for collecting video content in different language versions and cultural backgrounds), and content type rules (such as classification standards of film special series, television series, and animation series). The rule library can be flexibly configured and updated according to actual needs.

[0042] Then, the present application inputs the extracted multi-modal feature vectors into the large model, and the large model infers the correlation between the video content and the series category to which the video content belongs based on the trained algorithm and model and in combination with the series structure rule library. The specific process is as follows: the large model first performs preliminary matching in the rule library according to the multi-modal feature vectors, and screens out possible related series rules; then, the large model further analyzes the degree of fit between the features of the video content and the rules to determine the accurate series attribution of the video content.

[0043] Further, the step S300 specifically includes: S301, inputting the extracted multi-modal feature vectors into a pre-trained artificial intelligence large model; In the embodiment of the present step, the multi-modal feature vectors are a plurality of data sets covering multiple dimensions extracted from the video content in the previous steps, such as visual dimension picture color tone, scene layout, and character costume style, audio dimension background music melody and dialogue style, and text dimension plot outline keywords and character relationship description. These information will be converted into a numerical vector form that can be recognized by a computer, such as [0.23, 0.89, 0.11,...]. The pre-trained artificial intelligence large model of the embodiment is a model trained by a large amount of video data (including series and non-series works), and has the ability to process multi-modal data, for example, a multi-modal large model based on a Transformer architecture (such as an improved version of CLIP), which has learned the association rules between feature vectors and video attributes (such as whether it is a series or a type).

[0044] S302, the pre-trained artificial intelligence large model filters out related series rules according to the preliminary matching of the multi-modal feature vector in the preset series structure rule library. In the embodiment of the application, the preset series structure rule library is a database constructed in advance to store series core feature rules, including IP derivative rules (clear judgment standards and association relationships of different types such as main story, prequel, spin-off, and fan fiction), regional content rules (rules for collecting video content in different language versions and cultural backgrounds), content type rules (classification standards such as movie special series, TV series, and animation series), etc. The rule library can be flexibly configured and updated according to actual needs.

[0045] The preset series structure rule library of the embodiment can also include common structures of series, such as continuous character lineup (such as the same character running through multiple works), coherent plot timeline (such as the ending of a prequel being the beginning of a sequel), unified worldview (such as the same fictional background, such as a science fiction universe or an ancient dynasty), consistent theme style (such as all being suspense and reasoning or family comedy), etc. Each rule corresponds to a specific feature vector judgment standard. In this step, the preliminary matching process is that the model compares the input multi-modal feature vector with each rule in the rule library: for example, if the feature vector of video A contains an 80% character name repetition rate and a 60% scene repetition rate, the model will match the rule of continuous character lineup + unified scene in the rule library, while excluding the rule of complete character replacement and large worldview difference. The purpose of this step is to filter out rules unrelated to the current video content, reduce the computational load of subsequent analysis, ensure that the subsequent fit degree analysis is only around possible related series rules, and improve the determination efficiency.

[0046] S303, and analyze the fit degree of the features of the video content and the rules to determine the series attribution of the video content and infer the association relationship between the video contents and the series category to which they belong.

[0047] In the present application, the fit degree of the analysis features and the rules is a quantitative evaluation of the relevant rules screened out in step S302. Specifically, the well-trained artificial intelligence model sets a fit degree scoring standard (such as a full score of 100) for each relevant rule. For example, for the plot timeline coherence rule, if the plot timeline connection error between movie B and movie C is less than 10%, the fit degree is 90 points; if there is a part of the timeline fault, the score is 60 points. By synthesizing the fit degree scores of all relevant rules (such as weighted average), the overall fit degree of the movie content and the rule set of a series (such as 85 points, with 80 points as the threshold for meeting the series attribution) is obtained. In the present application, the determination of the series attribution is based on the fit degree result. Specifically, when the overall fit degree of movie D and the rule set of the XX series reaches the threshold, it is determined that movie D belongs to the XX series; if it does not reach the threshold, it is determined to be a non-series or belongs to other series with a higher matching degree. In the present embodiment, the inference of the association relationship and the series category is further outputted with detailed results. Specifically, the association relationship includes that movie E is a prequel of movie F, movie G and movie H are parallel works of the same series, etc., which is based on the analysis of the timeline and plot causality features; and the series category includes continuous plot category (such as “Investigation Case Files” and “Legend of the Condor Heroes”), single-unit plot category (such as “Detective Conan” where each case is independent but the characters are continuous), derivative work category (such as “Iron Man” and “The Avengers” which belong to the Marvel Cinematic Universe derivative series), etc., which is based on the judgment of the feature differences of different categories of series in the rule library. The purpose of this step is to complete the closed loop from data analysis to result output, not only to determine whether the movie content belongs to a series, but also to clearly present its association logic with other works and specific series type, providing decision basis for movie classification, recommendation, copyright management, etc.

[0048] Step S400: According to the established inventory series database, the inferred association relationship between the movie content, the belonging series category, and the inventory series are dynamically mapped and compared, and the updated series information is stored in the inventory series database.

[0049] In the present embodiment, according to the established inventory series database, the newly inferred series are dynamically mapped and compared with the inventory series. In the comparison process, in addition to considering the feature association of the movie content itself, user behavior feedback data (such as user viewing time, collection record, comment content) are introduced as dynamic adjustment factors. If a user frequently associates a new movie with an existing series, the weight of the association is automatically increased; if it is found that the new movie and the existing series have conflicts or unreasonable attribution, the rules are adjusted accordingly. After the comparison and adjustment are completed, the updated series information is stored in the inventory series database.

[0050] Further, the step S400 specifically comprises: S401, according to the established inventory series database, the inferred association between the video content and the series category to which it belongs is dynamically mapped and compared with the inventory series in the inventory series database; In the present application, the inventory series database is a core data pool that has stored a large amount of historical series information, which contains the basic information of each inventory series (such as series name, work list, premiere time), core attributes (such as belonging category: continuous drama category / unit drama category), internal association (such as work A is the 3rd of series B, work C is the prequel of series B), etc., and the data structure is standardized, supporting quick retrieval according to series category, association type, etc. The dynamic mapping and comparison in the present application is a precise matching process divided by dimension and level, specifically divided into two steps: First, series category mapping, specifically, the category of the new video content inferred in S303 (such as derivative works) is matched with the category label of the inventory series in the database, and all inventory series in the same category are selected, for example, if the new content is a Marvel-style derivative drama, then the Marvel Cinematic Universe series, DC Extended Universe series, etc. in the derivative works category in the database are located first; Second, association mapping, specifically, based on the association clues of the new content, such as new drama D and already inferred drama E are the same world view sequel, further comparison is made in the above same category inventory series, if there is already F series in the database to which drama E belongs, then the consistency of the characteristics (such as world view, character) of new drama D and F series is verified to confirm whether it can be included in F series; if there is no record of drama E in the database, then it is judged whether new drama D and drama E constitute a new series prototype, and comparison is made with the inventory single work in the database that has no association to exclude potential omission of attribution, such as whether there is an inventory work that has not been found to have association with new drama D. The role of this step is to let the newly inferred series information find its attribution, either to be included in the existing inventory series to enrich the content system of the series, or to be identified as a new series to lay the foundation for subsequent database update, while avoiding the confusion of splitting the same series into multiple entries.

[0051] S402, the characteristics of the video content itself are associated, and user behavior feedback data is introduced as a dynamic adjustment factor; when a user frequently associates a new drama with an existing series, the weight of the association is automatically increased; when it is detected that there is a conflict or unreasonable collection between the new drama and the existing series, the rules are adjusted accordingly; after the comparison and adjustment are completed, the updated series information is stored in the inventory series database.

[0052] On the basis of the mapping comparison in step S401, on the one hand, the feature correlation verification of the video content itself is strengthened, and the consistency of the core features such as the picture style, the dialogue system and the core role of the new drama and the target inventory series is checked again to ensure that there is no obvious contradiction in the objective feature level; On the other hand, user behavior feedback data is introduced, which includes active operations of the user on the video platform, such as marking the new drama G as a sequel of the series H, mentioning the association between the new drama G and the series H in the comment, and passive behaviors such as 70% of the users watching the new drama G and then watching the works of the series H; these data can reflect the subjective cognition of the user on the association relationship of the new drama, and make up for the possible deviation of the pure model reasoning, such as the small plot association that is not identified by the model. Based on the consistency of the feature association, the application also takes the user behavior feedback as a dynamic adjustment factor and optimizes in two scenarios: one is positive weight promotion, that is, when the user frequently operates the association between the new drama and a certain inventory series, such as 1000+ users marking the new drama I as belonging to the series J within a week, and the association has no conflict with the objective features, the confidence weight of the association is automatically promoted, such as from 60% to 85%, so that the association is more preferentially reflected in subsequent recommendation and classification; The second is conflict adjustment. When unreasonable collection is detected, it may be model reasoning deviation, such as misclassifying the costume drama K into the modern urban series L, or user misoperation, a small number of users marking the non-associated drama M as belonging to the series N, then triggering the regularized adjustment, first detecting the feature conflict through the feature conflict detection rule, such as the time background feature of the costume drama K conflicts with the modern background feature of the series L, then determining that the collection is wrong, and then combining the user feedback weight rule, such as the proportion of mislabeling users is less than 5%, automatically canceling the wrong association, or initiating manual review prompt, such as when the conflict feature is not obvious, prompting the administrator to confirm. In the embodiment of the application, after the adjustment is completed, the finally confirmed information is updated in the database specification format. If the new drama is classified into the inventory series, the new drama information is added to the work list of the series, and the association relationship is supplemented, such as the new drama I is the 5th part of the series J; if the new drama and other works constitute a new series, a new series entry is created in the database, and the series name, work list, category, association relationship and other information are recorded; at the same time, the source of this update is recorded, such as model reasoning + user feedback adjustment, and the time, which is convenient for subsequent tracing and data auditing.

[0053] In a further embodiment of the application, the film series intelligent dynamic collection method based on a multi-modal large model further comprises the following steps after step S400: S501, setting a visual manual intervention interface, receiving user operation instructions through the manual intervention interface, and reviewing the collection result; In the embodiment of the present application, the visual artificial intervention interface is a clear and easy-to-understand operation interface, such as a table showing new drama names, machine-determined series, associated relationships and confidence weights, or a graph directly showing the associated link between the new drama and the existing series. Then, the artificial, such as a platform administrator or a professional reviewer, receives operation instructions through the interface, such as checking the list of cases to be reviewed, marking suspicious aggregation results, and checking the machine output aggregation results, such as the aggregation of new drama A into series B, to determine whether it conforms to the actual situation and avoid machine misjudgment that is not automatically detected by the system.

[0054] S502, receiving user operation instructions through the system interface to mark and correct error aggregation cases, and automatically recording error aggregation cases and modification records; In the embodiment of the present application, after the artificial discovers an error aggregation in the system interface, such as the misaggregation of new drama C into series D, which should actually belong to series E, the artificial sends operation instructions, such as marking the error type as category confusion and correcting the series as E, and the system will complete the marking and correction according to the instructions; at the same time, the system automatically records error aggregation details (error aggregation content, error type) and modification records (modifier, modification time, corrected content), forming a traceable error file.

[0055] S503, feeding the recorded error aggregation cases and modification records into the training data of the artificial intelligence large model to reversely optimize the training process of the artificial intelligence large model.

[0056] In the embodiment of this step, the error aggregation cases recorded in S502, such as the misaggregation of new drama C, and the modification records (correct aggregation and reasons), are arranged into data in a format conforming to the model training format and supplemented into the training data set of the artificial intelligence large model; the model learns the characteristics of these error aggregation cases (such as the difference characteristics of new drama C and series D and the matching characteristics of new drama C and series E) in subsequent iterative training, adjusts the internal judgment logic, reduces the occurrence of similar errors (such as category confusion), and realizes reverse optimization.

[0057] As can be seen from the above, the present application provides a film series drama intelligent dynamic aggregation method based on a multi-modal large model, which realizes the automatic, intelligent and dynamic aggregation of film series dramas by fusing the multi-modal features of film and television content such as text, audio and vision, combining a preset series drama structure rule library, using the powerful reasoning and analysis capability of the large model, improving the aggregation efficiency and accuracy, reducing the labor cost, and improving the user experience of watching series dramas.

[0058] The present application will be further described in detail through specific application embodiments: For example, Figure 2As shown, the structural block diagram of the multi-modal large model-based film series intelligent dynamic collection system of the embodiment shows the connection relationship and data flow between multi-modal feature extraction, series structure rule library, large model reasoning and matching, dynamic mapping comparison and updating, and artificial review and optimization. The multi-modal large model-based film series intelligent dynamic collection method provided in the embodiment 2 is described by taking the collection of newly stored video content as an example, which includes the following steps: Step S11, when new video content is released and stored, the system automatically triggers the collection step and enters step S12.

[0059] Step S12, acquiring video content, acquiring text information, audio information and visual information of the video content through a multi-modal feature extraction module. For example, for a newly released TV series, its plot synopsis and dialogue script are extracted as text information, the dialogue voice and background music in the series are obtained as audio information, and multiple key frames are intercepted as visual information.

[0060] Step S13, using natural language processing (NLP) technology to analyze the semantics of the text information and extract key semantic features; using an audio sentiment analysis algorithm to process the audio information and analyze features such as dialogue tone and background music emotional tone; using computer vision (CV) technology to analyze visual information and extract visual style features such as picture color tone, shot language, scene composition, etc., to form a multi-modal feature vector; Specifically, the natural language processing (NLP) technology can be used to process text information such as word segmentation, part-of-speech tagging, and semantic analysis, and extract key semantic features such as character names, story locations, and core plot; the audio sentiment analysis algorithm is used to process the audio information and analyze the tension of the dialogue, the happy or sad tone of the background music, and other emotional features; the CV technology is used to process the visual information and extract visual style features such as color tone style (such as warm color tone, cold color tone), shot type (such as close-up, panoramic), etc., to form a multi-modal feature vector.

[0061] Step S14, inputting the multi-modal feature vector into an artificial intelligence large model for reasoning and matching. The pre-trained artificial intelligence large model combines the series structure rule library for reasoning. Assuming that the TV series is a derivative work of a well-known IP, the large model analyzes its association with other works under the IP in terms of characters, plot, etc. according to the IP derivative rules, and determines that it belongs to the spin-off category of the IP series.

[0062] Step S15, dynamic mapping comparison and updating are performed, and specifically, the newly determined series are compared and mapped with the inventory series database. At the same time, the behavior data of the user for the television series, such as the viewing time length and collection, are collected, and if it is found that the user has a long viewing time length for the television series and has a collection behavior, and the television series has a high correlation with several works in the existing series, the weight of the correlation is automatically increased.

[0063] Step S16, after the comparison and adjustment are completed, the updated series information is stored in the inventory series database.

[0064] Step S17, then artificial review and optimization are performed, and the operation personnel are prompted by the system to review the collection result, and the operation personnel view the collection situation through a visual interface, and if it is found that the collection is wrong, the error case can be directly marked and corrected, and the system feeds back the error case to the training data of the large model for optimizing the large model.

[0065] Specific application embodiment 3, the series intelligent dynamic collection method based on the multi-modal large model provided by the specific application embodiment 3 takes the re-collection of the inventory film and television content as an example, and includes the following steps: Step S21, re-collection operations are periodically performed on all film and television contents in the inventory series database.

[0066] Step S22, according to the same steps in the specific application embodiment 2, multi-modal feature extraction, large model reasoning and matching, dynamic mapping comparison and updating are performed on each inventory film and television content.

[0067] Step S23, in the comparison process, the original collection result is adjusted according to the latest user behavior data and the updated content of the series structure rule library. For example, if a new type of derivative drama judgment rule is added to the series structure rule library, the inventory film and television content that meets the rule is re-determined to belong to a series.

[0068] Step S24, similarly, the result of the re-collection is checked and corrected by the operation personnel through the artificial review and optimization module, so that the collection accuracy and integrity of the inventory series are ensured.

[0069] As can be seen from the above, the present application adopts an innovative analysis mechanism of multi-modal feature fusion, specifically including: 1), cross-modal information deep mining is adopted, the present application breaks through the limitation of traditional single mode analysis, and constructs a multi-modal feature fusion system of text semantics, audio emotion and visual style. Through natural language processing technology, the deep logic of the movie plot is analyzed, the audio emotion analysis algorithm is used to capture the dialogue emotion and background music atmosphere, and the computer vision technology is used to extract the visual elements such as picture composition and color change, so as to form a multi-modal feature vector with strong correlation, and realize the all-round and deep feature extraction of the movie content.

[0070] 2) a dynamic weight distribution strategy is adopted, the present application designs a dynamic weight distribution algorithm for different types of movie content. For example, for the plot-oriented movie, the weight of the text semantic feature is increased; for the film with visual special effects as the highlight, the weight of the visual style feature is increased. The strategy makes the large model adjust the importance of each modal feature flexibly during the reasoning process, and improves the accuracy of the series drama association judgment.

[0071] The present application also carries out intelligent construction of series drama structure rule base, specifically including: 1) the rule dynamic updating and expansion are adopted, the series drama structure rule base of the present application supports real-time updating based on the dynamic of the film and television industry and the change of user demand. The operation personnel can add new series drama type rules (such as the emerging interactive drama series rules) or modify the existing rule parameters through the simple operation interface, so that the system can always adapt to the changing film and television market and maintain the accurate identification ability of various series dramas.

[0072] 2) the rule conflict resolution mechanism is adopted, when there are multiple rules in the series drama structure rule base that are applicable to the same movie content, the priority judgment and conflict resolution algorithm is introduced. By setting the rule priority and combining the multi-modal feature analysis result, the most reasonable collection rule is automatically selected, the error collection caused by rule conflict is avoided, and the uniqueness and accuracy of series drama collection are ensured.

[0073] And the present application also adopts optimization and reinforcement of large model reasoning, including: 1) transfer learning and domain adaptation are adopted, the present application is based on a general large model, which is adapted to the field of movie series drama collection through transfer learning technology. A large amount of labeled movie series drama data is used to fine-tune the model, so that the large model can accurately understand the series drama association logic between movie contents, and has higher precision and efficiency in series drama identification task compared with directly using general large model.

[0074] 2) Incremental learning and optimization are adopted, the artificial intelligence large model of the embodiment of the present application has incremental learning ability, can incorporate the error cases corrected by artificial review, new series drama type data in time into training, and continuously optimize model parameters. With the continuous accumulation of data and model training, the accuracy and generalization ability of the large model to series drama collection will be continuously improved, realizing self-evolution and optimization.

[0075] And the present application also adopts the intelligent decision system of dynamic mapping comparison, specifically including: 1) User behavior driven correlation optimization is adopted, the present application deeply integrates user viewing time, collection, comment and other behavior data into the dynamic mapping comparison process. A user behavior analysis model is constructed, and the series drama correlation preference behind the user behavior is mined through machine learning algorithm. For example, if a large number of users frequently and continuously watch two seemingly unrelated films, the system will increase the correlation weight of the two films, automatically adjust the series drama collection result, so that it is more in line with the actual needs of users.

[0076] 2) Real-time dynamic updating strategy is adopted, the dynamic mapping comparison module of the present application adopts the combination of real-time monitoring and periodic batch updating. On the one hand, the newly entered video and television contents are mapped and compared in real time to ensure that they are quickly and accurately classified into corresponding series dramas; on the other hand, the inventory series drama library is updated in batches in real time, reflecting the influence of user behavior changes and rule library updates on the classification results, ensuring that the series drama library is always in the latest and most accurate state.

[0077] The present application also adopts the efficient review mechanism of man-machine cooperation, specifically including: 1) Visual interactive interface design is adopted, the present application provides intuitive and convenient visual interactive interface through the artificial review and optimization module. The operation personnel can quickly check the series drama classification result through the interface, directly mark the error classification cases on the interface by using the graphical marking tool, which is simple and efficient. At the same time, the interface supports visual editing of series drama rules, which is convenient for operation personnel to adjust the rule library according to actual situation.

[0078] 2) Review data feedback model training is adopted, the present application receives the marking data, correction opinions and other information generated by artificial review, and feeds back the information to the large model training process after system arrangement. Through the construction of data feedback link, the artificial experience is converted into effective data for model training, realizing the closed loop optimization of man-machine cooperation, and continuously improving the accuracy and reliability of the large model in series drama classification under the intervention of artificial.

[0079] As can be seen from the above, through the embodiment of the present application, the following effects can be realized: 1) Improved collection accuracy: By fusing multi-modal feature analysis and series structure rule library, it can more comprehensively and accurately identify the association between video content, avoiding errors caused by single information analysis. Compared with traditional methods, the collection accuracy is improved by about 50%.

[0080] 2) Improved collection efficiency: The automatic series collection process greatly reduces manual operation, and the processing speed of mass video content collection is improved by about 5 times, which can quickly respond to new content and update the series library in time.

[0081] 3) Reduced cost: Reducing a large amount of manpower investment, reducing the labor cost of video content management, and improving the operating efficiency of enterprises.

[0082] 4) Optimized user experience: Based on user behavior feedback, dynamically adjust the collection results, which can better meet the viewing needs and preferences of users, provide more accurate and complete series content recommendations for users, and enhance user satisfaction and stickiness to the platform.

[0083] 5) Enhanced adaptability: The series structure rule library can be flexibly configured and updated, which can adapt to different types, different styles of video content, and changing user needs and market trends.

[0084] Exemplary device As shown in Figure 3 The embodiment of the present application provides an intelligent dynamic collection device for film series based on a multi-modal large model, which comprises: A series structure rule library module 310 is used to pre-construct a series structure rule library containing a plurality of series structure rules, including IP derivative rules, regional content rules, and content type rules. The series structure rule library can be flexibly configured and updated according to actual needs. An acquisition and feature extraction module 320 is used to acquire video content, and perform multi-modal feature extraction on the video content to extract text information, audio information, and visual information of the video content. A multi-modal feature extraction module 330 is used to analyze and extract key semantic features from the text information, analyze and extract emotional tone features from the audio information, and analyze and extract picture tone, shot language, scene composition visual style features from the visual information to obtain a multi-modal feature vector. A large model inference and matching module 340 is used to input the extracted multi-modal feature vector into a pre-trained artificial intelligence large model, combine the preset series structure rule library, and infer the association between video content and the series category to which it belongs. The dynamic mapping comparison and updating module 350 is configured to dynamically map and compare the inferred correlation between the video contents and the series drama categories to the inventory series dramas according to the established inventory series drama database, and store the updated series drama information in the inventory series drama database.

[0085] Based on the above embodiments, the application further provides a server, the principle block diagram of which can be shown as follows. Figure 4 The server comprises a processor, a memory, a network interface, a display screen, and a database connected through a system bus.

[0086] The memory stores one or more programs configured to be executed by the processor, such as the multi-modal feature extraction step, the large model inference and matching step, the dynamic mapping comparison and updating step, and the artificial review and optimization step, to realize the multi-modal large model-based movie series intelligent dynamic collection method of the above embodiments.

[0087] The server refers to an intelligent television, an intelligent tablet, etc. with data processing capability. The memory can be an internal memory, a flash memory, a hard disk, or a cloud storage space, used for storing program codes and storing various data such as pre-set multiple intelligent body images applicable to different groups of people and pre-set conversational customization processes. The processor can be a central processing unit, used for executing algorithm logic in the program. The program contains the multi-modal large model-based movie series intelligent dynamic collection method.

[0088] In further embodiments, the server of the present embodiment comprises a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors. The one or more programs contain instructions for performing the following operations: Obtaining video contents, performing multi-modal feature extraction on the video contents, extracting text information, audio information, and visual information of the video contents; Respectively analyzing and extracting key semantic features from the text information, analyzing and extracting emotional tone features from the audio information, and analyzing and extracting picture tone, shot language, scene composition visual style features from the visual information to obtain multi-modal feature vectors; Inputting the extracted multi-modal feature vectors into a pre-trained artificial intelligence large model, combining a pre-set series drama structure rule library, and inferring the correlation between the video contents and the series drama categories to which the video contents belong; According to the established inventory series drama database, the inferred correlation between the video contents and the series drama categories to which the video contents belong is dynamically mapped and compared with the inventory series dramas, and the updated series drama information is stored in the inventory series drama database, as described above.

[0089] Before the steps of obtaining the video content and performing multi-modal feature extraction on the video content to extract text information, audio information and visual information of the video content, the method further comprises: pre-constructing a series structure rule library containing a plurality of series structure rules, the plurality of series structure rules including IP derivative rules, regional content rules, and content type rules; the series structure rule library can be flexibly configured and updated according to actual needs.

[0090] Before the steps of obtaining the video content and performing multi-modal feature extraction on the video content to extract text information, audio information and visual information of the video content, the method further comprises: obtaining video content, wherein the video content includes newly released and stored video content and all video content in a stock series database obtained periodically; performing multi-modal feature extraction on the obtained video content to extract text information, audio information and visual information of the video content.

[0091] The step of respectively analyzing the text information to extract key semantic features, analyzing the audio information to extract emotional tone features, and analyzing the visual information to extract picture tone, shot language, scene composition and visual style features to obtain a multi-modal feature vector comprises: performing word segmentation, part-of-speech tagging and semantic analysis on the text information using natural language processing technology to extract key semantic features including character names, story locations and core plots; processing the audio information using an audio emotion analysis algorithm to analyze features such as dialogue tone and emotional tone of background music; analyzing the visual information using computer vision technology to extract visual style features such as picture tone, shot language and scene composition, and forming a multi-modal feature vector.

[0092] The step of inputting the extracted multi-modal feature vector into a pre-trained artificial intelligence large model and combining a pre-set series structure rule library to infer the correlation between video content and the series category to which the video content belongs comprises: inputting the extracted multi-modal feature vector into a pre-trained artificial intelligence large model; The pre-trained artificial intelligence large model preliminarily matches the multi-modal feature vector in the pre-set series structure rule library to filter out related series rules; and analyze the fit degree of the features of the video content and the rules to determine the series attribution of the video content, and infer the correlation between the video content and the series category to which the video content belongs.

[0093] The step of dynamically mapping and comparing the inferred association relationship between the video content and the series category to which the video content belongs with the inventory series according to the established inventory series database includes: The step of dynamically mapping and comparing the inferred association relationship between the video content and the series category to which the video content belongs with the inventory series according to the established inventory series database includes: The characteristics of the video content itself are associated, and user behavior feedback data is introduced as a dynamic adjustment factor; when a user frequently associates a new series with an existing series, the weight of the association is automatically increased; when a new series is detected to have a conflict or unreasonable collection with an existing series, the rules are adjusted accordingly; after the comparison and adjustment are completed, the updated series information is stored in the inventory series database.

[0094] The step of dynamically mapping and comparing the inferred association relationship between the video content and the series category to which the video content belongs with the inventory series according to the established inventory series database includes: A visual manual intervention interface is set up, and user operation instructions are received through the manual intervention interface to review the collection results; Through the system interface, user operation instructions are received to mark and correct the error collection cases, and automatically record the error cases and modification records; The recorded error cases and modification records are fed back to the training data of the artificial intelligence large model, and the training process of the artificial intelligence large model is optimized in reverse, as described above.

[0095] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0096] The above only describes the embodiments of the present application and is not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and variations, for example, a distributed large model collaborative aggregation scheme based on federated learning, that is, on the basis of the original centralized large model, the federated learning technology is introduced, allowing multiple platforms (such as different regional video streaming platforms) of scattered video content resources to collaboratively train the large model without sharing the original data. Each platform trains the model using local data and only uploads model parameter update information, and the central server aggregates and updates the global model. This variant scheme can protect the data privacy of each platform and integrate multi-party data resources to improve the accuracy and generalization ability of the large model for collecting diversified video content series, and is particularly suitable for scenarios with widely distributed video resources and sensitive data. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-modal large model-based film series intelligent dynamic collection method, characterized in that, The method comprises the following steps: Obtaining video and television content, performing multi-modal feature extraction on the video and television content, and extracting text information, audio information and visual information of the video and television content; Respectively analyzing and extracting key semantic features from the text information, analyzing and extracting emotional tone features from the audio information, and analyzing and extracting picture tone, shot language, scene composition visual style features from the visual information to obtain a multi-modal feature vector; Inputting the extracted multi-modal feature vector into a pre-trained artificial intelligence large model, combining a pre-set series structure rule library, and inferring the correlation between the video and television content and the series category to which the video and television content belong; According to the established inventory series database, the inferred correlation between the video and television content and the series category to which the video and television content belong are dynamically mapped and compared with the inventory series, and the updated series information is stored in the inventory series database.

2. The multi-modal large model-based film series intelligent dynamic collection method according to claim 1, characterized in that, The method further comprises the following steps before the step of obtaining video and television content, performing multi-modal feature extraction on the video and television content, and extracting text information, audio information and visual information of the video and television content: Pre-constructing a series structure rule library containing a plurality of series structure rules, wherein the plurality of series structure rules include IP derivation rules, regional content rules and content type rules; and the series structure rule library can be flexibly configured and updated according to actual needs.

3. The multi-modal large model-based film series intelligent dynamic collection method according to claim 1, characterized in that, The method further comprises the following steps of obtaining video and television content, performing multi-modal feature extraction on the video and television content, and extracting text information, audio information and visual information of the video and television content: Obtaining video and television content, wherein the video and television content includes newly released and stored video and television content and all video and television content in the inventory series database obtained periodically; Performing multi-modal feature extraction on the obtained video and television content, and extracting text information, audio information and visual information of the video and television content.

4. The multi-modal large model-based film series intelligent dynamic collection method according to claim 1, characterized in that, The method further comprises the following steps of respectively analyzing and extracting key semantic features from the text information, analyzing and extracting emotional tone features from the audio information, and analyzing and extracting picture tone, shot language, scene composition visual style features from the visual information to obtain a multi-modal feature vector: Using natural language processing technology to perform word segmentation, part-of-speech tagging and semantic analysis on the text information, and extracting key semantic features including character names, story locations and core plots; Using an audio emotion analysis algorithm to process the audio information, and analyzing emotional tone features of dialogue tone and background music; Using computer vision technology to analyze the visual information, and extracting picture tone, shot language, scene composition visual style features to form a multi-modal feature vector.

5. The multi-modal large model-based movie series intelligent dynamic collection method according to claim 1, characterized in that, The method further comprises the following steps of inputting the extracted multi-modal feature vector into a pre-trained artificial intelligence large model, and combining a pre-set series structure rule library to infer the correlation between the video and television content and the series category to which the video and television content belong: Inputting the extracted multi-modal feature vector into a pre-trained artificial intelligence large model; The pre-trained artificial intelligence large model performs preliminary matching of the multi-modal feature vector in the pre-set series structure rule library, and screens out related series rules; And analyze the characteristics of the film and television content and the degree of conformity of the rules, determine the series of film and television content, infer the relationship between the film and television content and the category of the series.

6. The multi-modal large model-based movie series intelligent dynamic collection method according to claim 1, characterized in that, The step of dynamically mapping and comparing the inferred relationship between the film and television content and the category of the series with the inventory series according to the established inventory series database includes: According to the established inventory series database, the inferred relationship between the film and television content and the category of the series are dynamically mapped and compared with the inventory series in the inventory series database; The characteristics of the film and television content are associated, and user behavior feedback data is introduced as a dynamic adjustment factor; when a user frequently associates a new series with an existing series, the weight of the association is automatically increased; when a new series is detected to have conflicts or unreasonable collection with an existing series, the rules are adjusted accordingly; after comparison and adjustment, the updated series information is stored in the inventory series database.

7. The multi-modal large model-based movie series intelligent dynamic collection method according to claim 1, characterized in that, The step of dynamically mapping and comparing the inferred relationship between the film and television content and the category of the series with the inventory series according to the established inventory series database includes: Set up a visual manual intervention interface to receive user operation instructions through the manual intervention interface and review the collection results; Through the system interface, receive user operation instructions to mark and correct error collection cases, and automatically record error cases and modification records; The recorded error cases and modification records are fed back to the training data of the artificial intelligence large model to optimize the training process of the artificial intelligence large model.

8. A multi-modal large model-based film series intelligent dynamic collection device, characterized in that, The device comprises: A series structure rule library module for pre-constructing a series structure rule library containing a plurality of series structure rules, including IP derivative rules, regional content rules, and content type rules; the series structure rule library can be flexibly configured and updated according to actual needs; An acquisition and feature extraction module for acquiring film and television content and extracting multi-modal features of the film and television content, including text information, audio information, and visual information; A multi-modal feature extraction module for analyzing and extracting key semantic features from the text information, analyzing and extracting emotional tone features from the audio information, and analyzing and extracting picture tone, shot language, scene composition visual style features from the visual information to obtain multi-modal feature vectors; A large model inference and matching module for inputting the extracted multi-modal feature vectors into a pre-trained artificial intelligence large model, combining the pre-set series structure rule library, and inferring the relationship between the film and television content and the category of the series. The dynamic mapping comparison and updating module is configured to compare the inferred correlation between the video and television contents and the series category to which the video and television contents belong with the inventory series according to the established inventory series database, and store the updated series information into the inventory series database.

9. A server, characterized by A non-transitory computer-readable medium storing code for execution by one or more processors comprises one or more programs for implementing the method of any of claims 1-7.

10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is enabled to perform the method of any of claims 1-7.

Citation Information

Patent Citations

  • A method, a device, a device and a storage medium for establishing a drama relationship

    CN109543071A

  • Short video classification method and system, equipment and storage medium

    CN113743277A

  • Video set determination method and device, electronic equipment and storage medium

    CN114363660A

  • Vision and language-based annotation association type short video emotion recognition method and system

    CN114882412A

  • Video classification method and device, computer equipment and computer readable storage medium

    CN117475351A