An adaptive hierarchical multi-category character portrait generation method

Through the adaptive differential hierarchical multi-category character portrait generation method, combined with visual information and large language model, the problem that traditional methods are difficult to capture the emotional changes and behavioral evolution of characters is solved, and efficient analysis is achieved in complex plots and multi-character scenes is improved, and the depth and accuracy of video analysis are improved.

CN119763024BActive Publication Date: 2025-05-06XIANGJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510262399.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-05-06
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

Traditional character portrait generation methods are difficult to effectively capture the emotional changes and behavioral evolution of characters in different plot development stages, especially when dealing with complex plots and multi-character scenes.

Method used

Adaptive differential hierarchical multi-category character portrait generation method is adopted. Through the fusion of visual information and large language models, multi-dimensional keyframes in the video are extracted, semantic difference degree, object change and time difference factors are calculated, keyframe extraction thresholds are dynamically adjusted, and multi-grained differential descriptions are generated using the GPT-4V model to generate a role portrait that dynamically reflects the emotional changes and behavioral evolution of the character.

Benefits of technology

It realizes accurate capture of the emotional changes and behavioral evolution of characters in complex plots and multi-character scenes, improves the depth and accuracy of video analysis, and provides more flexible and efficient tools, suitable for film and television production, virtual character creation, and animation character analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119763024B_ABST
    Figure CN119763024B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of information processing technology, and specifically discloses an adaptive difference-based hierarchical multi-category character portrait generation method. Compared with the traditional video analysis method, it introduces multi-dimensional difference calculation, such as semantic difference, object change and time difference factor, and combines with the adaptive threshold adjustment mechanism, so that the key frame extraction of the fixed time window can dynamically adapt to the changes of different video contents. Especially in complex scenes and multi-character videos, this adaptability can effectively capture important plots and character changes in the plot, thereby improving the accuracy and flexibility of key frame extraction, and solving the problem that the analysis method used in the traditional character portrait production method is difficult to effectively capture the emotional changes and behavioral evolution of the characters at different plot development stages when dealing with complex plots and multi-character scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information processing technology, and specifically discloses a method for generating multi-category role portraits with adaptive differentiation and hierarchy. Background Art

[0002] With the rapid growth of digital media content, especially in the rapid development of animation, film and television, and the virtual world, video content analysis has become an important research direction in the field of artificial intelligence.

[0003] Traditional video analysis methods mostly focus on simple scene cutting, object detection and action recognition tasks, usually relying on specific visual features or manually set rules, and have limited understanding of video content. However, with the development of deep learning technology, especially the application of convolutional neural networks (CNN) and natural language processing (NLP) in video analysis, the understanding of video content is not limited to object detection at the image level, but further extended to high-level content analysis such as video semantics, character behavior, and emotional changes.

[0004] Character portrait generation is an important task in video analysis, especially in animation and film and television works, which often requires detailed modeling of the personality traits, emotional evolution, and behavioral patterns of multiple characters.

[0005] Traditional character portrait generation methods mostly rely on manually designed rules or the application of fixed templates, which lack flexibility and dynamic adjustment capabilities. In addition, in traditional character portrait production methods, the video frame-based analysis methods usually ignore important information such as timing differences, scene transitions, and character behavior dynamics. This makes it difficult to effectively capture the emotional changes and behavioral evolution of characters at different stages of plot development when dealing with complex plots and multi-character scenes.

[0006] The present invention provides an adaptive differential hierarchical multi-category character portrait generation method to solve the above problems. Summary of the invention

[0007] The purpose of the present invention is to solve the problem that the analysis method used in the traditional character portrait production method is difficult to effectively capture the emotional changes and behavioral evolution of the characters at different stages of plot development when dealing with complex plots and multi-character scenes.

[0008] In order to achieve the above object, the basic scheme of the present invention provides an adaptive difference hierarchical multi-category character portrait generation method, comprising the following steps:

[0009] Step S1: obtaining a video to be analyzed and preprocessing the video, wherein the preprocessing step includes sequentially performing preliminary editing of the video and segmenting the video to obtain a number of fixed-length windows;

[0010] Step S2: extract dynamic multi-dimensional key frames in sequence according to the obtained fixed time window to obtain corresponding key frame sets, and combine all key frame sets to form a global key frame set sequence;

[0011] Step S3: sequentially inputting the key frame sets corresponding to the fixed-length windows in the key frame set sequence into the large language model: GPT-4V model, providing prompt words to guide the GPT-4V model to generate text descriptions for the input fixed-length windows, and completing differential descriptions;

[0012] Step S4: Input all differential descriptions and their corresponding timestamps into the GPT-4V model in sequence, and add character portrait prompts to guide the GPT-4V model to extract the differences in character information over time, extract differential descriptions for multiple main characters respectively and aggregate them to obtain differential portraits of multiple characters, and use the GPT-4V model to perform multi-perspective comparison of different perspectives of the same character until several character differential portraits are generated, and then input the character differential portraits and the corresponding plot change descriptions into the GPT-4V model for global integration to obtain multi-category character portraits.

[0013] Furthermore, in step S1, the preliminary video editing steps include: video detection, identifying and deleting highly repetitive clips in the video.

[0014] Further, when implementing step S2, the following steps are included:

[0015] Step S21: sequentially establish a key frame set for the fixed time window, and initialize the first frame as the current key frame;

[0016] Step S22: obtaining multi-dimensional frame-level semantic information of the current key frame and the next frame in the fixed time window of this iteration;

[0017] Step S23: Calculate the multi-dimensional difference and obtain a comprehensive difference score based on the time difference factor according to the multi-dimensional frame-level semantic information obtained in step S22;

[0018] Step S24: Calculate the adaptive threshold and complete the key frame determination. When the next frame meets the determination condition, the next frame replaces the current key frame.

[0019] Step S25: Repeat steps S21 to S24 until all fixed-length windows are processed to obtain a global key frame set sequence.

[0020] Further, in step S22, the multi-dimensional frame-level semantic information includes: k and the next frame f i The semantic vector and , main object, number of main objects, number of main objects that appear in the next frame and number of main objects that disappear, and timestamp difference between the current key frame and the next frame to obtain.

[0021] Furthermore, when implementing step S23, the calculation of the multi-dimensional difference includes:

[0022] b1) through the current key frame f k The semantic vector and the next frame f i The semantic vector Calculating semantic difference :

[0023] ;

[0024] b2) Calculate the difference between objects :

[0025] ;

[0026] in, and Respectively represent f i and f k The number of main objects in and They are f i The number of newly appeared objects and the number of disappeared objects in , α and β are adjustable weight parameters;

[0027] b3) Calculate the time difference factor d time :

[0028] ;

[0029] in, Indicates the timestamp difference between the current key frame and the next frame, T max represents the interval threshold, and otherwise represents other situations.

[0030] Furthermore, when implementing step S23, the comprehensive difference score S(f i ,f k ) is calculated through semantic difference, object difference and time difference factor, and the calculation formula is as follows:

[0031] ;

[0032] Among them, ω1, ω2 and ω3 are all adjustable weight parameters.

[0033] Further, when implementing step S24, the following steps are included:

[0034] c1) Set the initial static threshold ;

[0035] c2) Utilize d semantic (f i ,f k ), d object (f i ,f k ) and d time Calculating dynamic thresholds , the calculation formula is as follows:

[0036] ;

[0037] Among them, α1, α2 and α3 represent weights;

[0038] c3) Compare S(f i ,f k )and ,like , then determine f i is the current key frame, otherwise, maintain f k is the current keyframe.

[0039] Further, when implementing step S3, the following steps are included:

[0040] Step S31: Generate the first content description;

[0041] The method includes inputting a key frame set of a first fixed-length window into a GPT-4V model, and providing a prompt word to guide the GPT-4V model to generate a text description k1 for the content of the fixed-length window;

[0042] Step S32: Based on the GPT-4V model, generate multi-granularity differential information about the current fixed-length window and the next fixed-length window, and output differential descriptions for the obtained multi-granularity differential information respectively, and finally merge them into the differential description k2;

[0043] Step S33: Repeat step S32 to obtain a differential description sequence (k1, ..., k m ), m represents the number of fixed-length windows, m=1, 2, 3,…

[0044] The principle and effect of this scheme are:

[0045] 1. Compared with the prior art, the present invention can extract richer semantic information from videos by integrating visual information and large language models, especially using a powerful image encoder and an advanced language generation model: the GPT-4V model, thereby generating more accurate and detailed multi-category character portraits.

[0046] 2. Compared with the prior art, the present invention introduces a multi-dimensional key frame extraction mechanism, performs intelligent analysis based on factors such as semantic difference, object change and time difference factor, and dynamically adjusts the key frame extraction threshold to ensure that important changes in the video are captured.

[0047] 3. Compared with the prior art, the present invention uses the GPT-4V model to generate multi-granular differential descriptions, so that the character portrait can dynamically reflect the emotional changes and behavioral evolution of the character during the development of the plot. This method can generate multi-perspective character portraits from multiple perspectives, greatly improving the depth and accuracy of video analysis. It can be widely used in multiple fields such as film and television production, virtual character creation, and animation character analysis, providing content creators and analysts with more intelligent, flexible and efficient tools, and solving the problem that the analysis method used in the traditional character portrait production method is difficult to effectively capture the emotional changes and behavioral evolution of the character at different stages of plot development when dealing with complex plots and multi-character scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0049] Figure 1 A flowchart of a method for generating multi-category character portraits by adaptively differentiating hierarchical layers is shown in an embodiment of the present application. DETAILED DESCRIPTION

[0050] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.

[0051] An adaptive differential hierarchical multi-category character portrait generation method, implementing Figure 1 As shown:

[0052] The steps include:

[0053] Step S1: Obtain the video to be analyzed and preprocess the video. The preprocessing step includes sequentially performing preliminary editing of the video and segmenting the video to obtain a number of fixed time windows. In this embodiment, the proposed video can represent not only animation videos, but also other film and television videos.

[0054] Specifically, the steps of preliminary video editing include: video detection, identification and deletion of highly repetitive clips in the video, such as the opening song, ending song, and mid-film breaks, so as to obtain a more compact and continuous core plot video, ensuring high quality of subsequent key frame extraction and analysis.

[0055] During the video segmentation process, the window length is adaptively set according to parameters such as the video frame rate and scene transition rate. For example, a shorter window of 2 seconds can be used for action movies, and a longer window of 3 seconds can be used for feature films as a fixed-length window. The fixed-length window is used for subsequent differential analysis of local content.

[0056] Step S2: Extract dynamic multi-dimensional key frames in sequence according to the obtained fixed-length windows to obtain corresponding key frame sets, and combine all key frame sets to form a global key frame set sequence.

[0057] Specifically, the steps include:

[0058] Step S21: sequentially establish a set of key frames for a fixed time window, and initialize the first frame as the current key frame.

[0059] The specific operations are as follows:

[0060] At the beginning of each iteration, the mth fixed-length window is selected in turn, where m=1, 2, 3, ..., and all frames of the selected fixed-length window are constructed into a set V, and the first frame of V is used as the current key frame;

[0061] Step S22: Obtain multi-dimensional frame-level semantic information of the current key frame and the next frame in the fixed time window of this iteration, including:

[0062] a1) Set the current key frame f k and the next frame f i , input into the image encoder of CLIP-Large, and the semantic vectors are obtained through the image encoder. and ;

[0063] a2) Detect f using the pre-trained target detection model k and f i Detect the main objects in the game, including but not limited to the detection of people, vehicles, and props, and count the number of main objects;

[0064] Among them, the pre-trained target detection model used is YOLO v11.

[0065] a3) Compare f k and f i The main object in the statistics f i The number of new main objects that appear and the number of main objects that disappear;

[0066] a4) Record f k The timestamp and f i The timestamp difference ;

[0067] Step S23: Calculate the multi-dimensional difference and obtain a comprehensive difference score based on the multi-dimensional frame-level semantic information obtained in step S22, which specifically includes the following steps:

[0068] b1) Through semantic vector and Calculating semantic difference , the calculation formula is as follows:

[0069] ;

[0070] b2) Calculate the difference between objects , the calculation formula is as follows:

[0071] ;

[0072] in, and Respectively represent f i and f k The number of main objects in and They are f i The number of newly appeared objects and the number of disappeared objects in , α and β are adjustable weight parameters;

[0073] b3) Calculate the time difference factor d time , the calculation formula is as follows:

[0074] ;

[0075] Among them, T max represents the interval threshold;

[0076] b4) through d semantic (f i ,f k ) 、d object (f i ,f k ) and d timeCalculate the comprehensive difference score S(f i ,f k ), the calculation formula is as follows:

[0077] ;

[0078] Among them, ω1, ω2 and ω3 are all adjustable weight parameters.

[0079] Step S24: Calculate the adaptive threshold and complete the key frame determination. When the next frame meets the determination conditions, the next frame replaces the current key frame. Specifically, the steps include:

[0080] c1) Set the initial static threshold ;

[0081] In this embodiment, the initial static threshold Determined according to the complexity of the video content and usually set to 0.2.

[0082] c2) Utilize d semantic (f i ,f k ) 、d object (f i ,f k ) and d time Calculating dynamic thresholds , the calculation formula is as follows:

[0083] ;

[0084] Among them, α1, α2 and α3 represent weights;

[0085] c3) Compare S(f i ,f k )and ,like , then determine f i is the current key frame, otherwise, maintain f k is the current keyframe.

[0086] Step S25: Repeat steps S21 to S24 until all fixed-length windows are processed to obtain a global key frame set sequence V key , each fixed-length window corresponds to a set of key frames, which facilitates subsequent description and differential analysis.

[0087] Step S3: sequentially input the key frame sets corresponding to the fixed-length windows in the key frame set sequence into the large language model: GPT-4V model, provide prompt words to guide the GPT-4V model to generate a text description for the input fixed-length window, and complete the differential description.

[0088] Specifically, the steps include:

[0089] Step S31: Generate the first content description;

[0090] The method includes inputting a set of key frames of a first fixed-length window into a GPT-4V model, and providing prompt words to guide the GPT-4V model to generate a text description k1 for the content of the fixed-length window. The generated text description content may include a summary of scenes, characters, actions, etc.;

[0091] Step S32: Based on the GPT-4V model, generate multi-granularity differential information about the current fixed-length window and the next fixed-length window, and output differential descriptions for the obtained multi-granularity differential information, specifically including the following steps:

[0092] d1) When transitioning from the current fixed-length window to the next fixed-length window, the following information is used as the input of the GPT-4V model: the key frame set of the current fixed-length window, the key frame set of the next fixed-length window; the text description k1 of the current fixed-length window; additional prompts, after input, guide the GPT-4V model to output multi-granularity differential information, including scene change information, role change information, object or prop change information, camera perspective or lens motion information;

[0093] d2) The GPT-4V model is required to output differential descriptions for the above multi-granular differential information points, such as "scene differential summary", "character emotion differential", "action differential", etc., and merge them into the final differential description k2;

[0094] Step S33: Repeat step S32 to obtain a differential description sequence (k1, ..., k m );

[0095] Step S4: Input all differential descriptions and their corresponding timestamps into the GPT-4V model in sequence, and add character portrait prompts to guide the GPT-4V model to extract the differences in character information over time, extract differential descriptions for multiple main characters respectively and aggregate them to obtain differential portraits of multiple characters, and use the GPT-4V model to perform multi-perspective comparison of different perspectives of the same character until several character differential portraits are generated, and then input the character differential portraits and the corresponding plot change descriptions into the GPT-4V model for global integration to obtain multi-category character portraits.

[0096] Specifically, the steps include:

[0097] Summarize the differential descriptions of all fixed-length windows, including collecting k1 to k mAll differential descriptions and corresponding timestamps are input into the GPT-4V model, and "character portrait prompts" are added to guide the GPT-4V model to extract character information, such as differences in character appearance, behavior, emotions, motivations, etc. over time;

[0098] When there are multiple main characters in a video, differential descriptions can be extracted and aggregated to form differential portraits of multiple characters. , n represents the number of characters, which is input into the GPT-4V model. The GPT-4V model compares the different perspectives of the same character (narrator perspective, opponent perspective, etc.) from multiple perspectives.

[0099] Continue the above process until a number of character difference portraits are generated, and the dynamic changes of the characters in the plot are recorded through the character difference portraits;

[0100] All the obtained character differential portraits and the corresponding self-written plot change descriptions are input into the GPT-4V model for global integration to complete the generation of multi-category character portraits, which can highlight the character's overall personality traits, background information and key time.

[0101] Compared with traditional video analysis methods, the present invention introduces multi-dimensional difference calculation, such as semantic difference, object change and time difference factor, and combines the adaptive threshold adjustment mechanism, so that the key frame extraction of fixed time window can dynamically adapt to the changes of different video contents. Especially in complex scenes and multi-role videos, this adaptability can effectively capture important plots and character changes in the plot, thereby improving the accuracy and flexibility of key frame extraction.

[0102] At the same time, the present invention uses the GPT-4V model to generate multi-granular differential descriptions to refine the changes in video content. By performing differential processing on video frames, differential descriptions such as scene changes, character actions, and emotional transitions are gradually generated, and these descriptions are used to dynamically update character portraits. This innovation can capture the details of the character's emotions, behaviors, and psychological changes in complex plots, breaking through the limitations of traditional static character descriptions.

[0103] At the same time, the present invention introduces the concept of multi-perspective comparative analysis in character portrait generation. By analyzing the character from different perspectives (such as the character's own perspective, the narration perspective, etc.), the three-dimensional sense and multi-dimensional expression of the character portrait are further deepened. At the same time, the character's personality and behavior evolution in different works or scenes can also be dynamically updated through differential portraits. This method is particularly suitable for multi-role and cross-work analysis, which promotes the innovation and development of character portrait generation technology.

[0104] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. An adaptive differential hierarchical multi-category character portrait generation method, characterized in that: The steps include: Step S1: obtaining a video to be analyzed and preprocessing the video, wherein the preprocessing step includes sequentially performing preliminary editing of the video and segmenting the video to obtain a number of fixed-length windows; Step S2: extract dynamic multi-dimensional key frames in sequence according to the obtained fixed time window to obtain corresponding key frame sets, and combine all key frame sets to form a global key frame set sequence; Step S3: sequentially inputting the key frame sets corresponding to the fixed-length windows in the key frame set sequence into the large language model: GPT-4V model, providing prompt words to guide the GPT-4V model to generate text descriptions for the input fixed-length windows, and completing differential descriptions; Step S4: Input all differential descriptions and their corresponding timestamps into the GPT-4V model in sequence, and add character portrait prompts to guide the GPT-4V model to extract the differences in character information over time, extract differential descriptions for multiple main characters respectively and aggregate them to obtain differential portraits of multiple characters, and use the GPT-4V model to perform multi-perspective comparison of different perspectives of the same character until several character differential portraits are generated, and then input the character differential portraits and corresponding plot change descriptions into the GPT-4V model for global integration to obtain multi-category character portraits; When implementing step S2, the following steps are included: Step S21: sequentially establish a key frame set for the fixed time window, and initialize the first frame as the current key frame; Step S22: obtaining multi-dimensional frame-level semantic information of the current key frame and the next frame in the fixed time window of this iteration; Step S23: Calculate the multi-dimensional difference and obtain a comprehensive difference score based on the time difference factor according to the multi-dimensional frame-level semantic information obtained in step S22; Step S24: Calculate the adaptive threshold and complete the key frame determination. When the next frame meets the determination condition, the next frame replaces the current key frame. Step S25: Repeat steps S21 to S24 until all fixed-length windows are processed to obtain a global key frame set sequence.

2. The method for generating multi-category character portraits by adaptive differentiation and hierarchicalization according to claim 1, characterized in that: In step S1, the preliminary video editing steps include: video detection, identifying and deleting highly repetitive clips in the video.

3. The method for generating multi-category character portraits by adaptive differentiation and hierarchicalization according to claim 1, characterized in that: In step S22, the multi-dimensional frame-level semantic information includes: k and the next frame f i The semantic vector and , main object, number of main objects, number of main objects that appear in the next frame and number of main objects that disappear, and timestamp difference between the current key frame and the next frame to obtain.

4. The method for generating multi-category character portraits by adaptive differentiation and hierarchicalization according to claim 1, characterized in that: When implementing step S23, the calculation of the multi-dimensional difference includes: b1) through the current key frame f k The semantic vector and the next frame f i The semantic vector Calculating semantic difference : ; b2) Calculate the difference between objects : ; in, and Respectively represent f i and f k The number of main objects in and They are f i The number of newly appeared objects and the number of disappeared objects in , α and β are adjustable weight parameters; b3) Calculate the time difference factor d time : ; in, Indicates the timestamp difference between the current key frame and the next frame, T max represents the interval threshold, and otherwise represents other situations.

5. The method for generating multi-category character portraits by adaptive differentiation and hierarchicalization according to claim 4, characterized in that: When implementing step S23, the comprehensive difference score S(f i ,f k ) is calculated through semantic difference, object difference and time difference factor, and the calculation formula is as follows: ; Among them, ω1, ω2 and ω3 are all adjustable weight parameters.

6. The method for generating multi-category character portraits by adaptive differentiation and hierarchicalization according to claim 5, characterized in that: When implementing step S24, the following steps are included: c1) Set the initial static threshold ; c2) Utilize d semantic (f i ,f k ) 、d object (f i ,f k ) and d time Calculating dynamic thresholds , the calculation formula is as follows: ; Among them, α1, α2 and α3 represent weights; c3) Compare S(f i ,f k )and ,like , then determine f i is the current key frame, otherwise, maintain f k is the current keyframe.

7. The method for generating multi-category character portraits by adaptive differentiation and hierarchicalization according to claim 6, characterized in that: When implementing step S3, the following steps are included: Step S31: Generate the first content description; The method includes inputting a key frame set of a first fixed-length window into a GPT-4V model, and providing a prompt word to guide the GPT-4V model to generate a text description k1 for the content of the fixed-length window; Step S32: Based on the GPT-4V model, generate multi-granularity differential information about the current fixed-length window and the next fixed-length window, and output differential descriptions for the obtained multi-granularity differential information respectively, and finally merge them into the differential description k2; Step S33: Repeat step S32 to obtain a differential description sequence (k1, ..., k m ), m represents the number of fixed-length windows, m=1, 2, 3,…

Citation Information

Patent Citations

  • Method and device for capturing character social relation evolution and related products

    CN115966002A

  • Human movement behavior generation method and system based on large language model

    CN118673102A