Method and system for role alignment in cross-scene long video based on large model
By extracting facial and clothing feature vectors from long videos and combining them with multimodal information and temporal modeling, the instability of character recognition in complex scenarios under existing technologies is solved, achieving efficient and accurate character alignment across scenarios. This is applicable to film and television production, video retrieval, and intelligent security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 北京中科闻歌科技股份有限公司
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-10
AI Technical Summary
Existing video character recognition technologies show significant performance degradation in complex scenarios such as facial occlusion, large-angle pose changes, and low light conditions, and are difficult to adapt to the characteristics of diverse appearances and large scene spans of the same character in long videos such as movies, TV series, and variety shows.
By acquiring the keyframe set of the target video, feature vectors of the target face and clothing are extracted. Cross-modal role representation is constructed by combining multimodal information. Temporal modeling and causal reasoning modules are introduced to associate role identities. A large model is used for high-order semantic feature extraction and a small model is used for fast detection. The frame sparsity is dynamically adjusted to adapt to long video processing.
It significantly enhances the stability and adaptability of character identities across long-form videos, improves accuracy in complex narratives and multi-character interaction scenarios, reduces computational resource consumption, and is suitable for various business scenarios such as film and television production, video retrieval, and intelligent security.
Smart Images

Figure CN121838232A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of character recognition technology, and in particular to a method and system for character alignment in long videos across different scenes based on a large model. Background Technology
[0002] With the continuous advancement of artificial intelligence technology and the widespread application of large-scale models, video content understanding and character analysis technologies have also ushered in new development opportunities. The technological development and maturity of multimodal large-scale models provide a new technical foundation for cross-modal semantic association and fine-grained visual understanding, and it has broad application prospects in many fields such as film and television production, content review, video retrieval, and intelligent security. For example, in the film and television industry, automatic character clustering can assist in plot analysis, character trajectory tracing, and video structuring; in video platforms, users can quickly locate relevant segments through character tags; and in security scenarios, it can realize cross-scene association and analysis of specific targets under multiple cameras and multiple time periods.
[0003] Existing video character clustering methods mostly rely on face recognition, re-identification (Re-ID), or visual feature extraction techniques. For example, the method and apparatus for identifying video objects in a video (authorization number: CN111242189B) obtains at least one trace of a video object from the video, clusters it to obtain target object clusters corresponding to each video object, determines the target object matching the target object cluster from the target object information database, and identifies the corresponding target object information as the object information of the video object; while the method, apparatus, computer equipment and storage medium for classifying film and television characters (authorization number: CN114282057B) crops character images from film and television video images through subject detection, extracts features, divides feature subsets according to video source, performs multiple rounds of clustering, and filters out features with large deviations.
[0004] While existing technologies have laid a foundation for video character recognition and clustering, they still have many limitations in practical applications. For example, face recognition-based methods show significant performance degradation in complex scenarios such as facial occlusion, large-angle pose changes, and low light conditions; while re-recognition methods that rely on visual features are sensitive to changes in clothing and styling, making it difficult to adapt to the characteristics of diverse appearances and large scene spans of the same character in long videos such as movies, TV series, and variety shows. Summary of the Invention
[0005] According to a first aspect of the present invention, a method for character alignment in long cross-scene videos based on a large model is provided, the method comprising the following steps: S1, acquire the target video, wherein the target video is a long video across scenes provided to the target user; S2, Obtain a keyframe set from the target video, wherein the keyframe set includes several keyframes; S3, based on the set of keyframes, obtain the target face feature vector; S4, Based on the set of keyframes, obtain the feature vector of the target clothing; S5. Identify the target person based on the target clothing feature vector and the target face feature vector.
[0006] According to a second aspect of the present invention, a character alignment system for cross-scene long videos based on a large model is provided, the system comprising: The first execution module is used to acquire the target video, wherein the target video is a long video across scenes provided by the target user; The second execution module is used to obtain a key frame set from the target video, wherein the key frame set includes several key frames; The third execution module is used to obtain the target face feature vector based on the key frame set; The fourth execution module is used to obtain the target clothing feature vector based on the keyframe set; The fifth execution module is used to identify the target person based on the target clothing feature vector and the target face feature vector.
[0007] The present invention has at least the following beneficial effects: This invention provides a method for character alignment in long, cross-scene videos based on a large model, the method comprising the following steps: The process involves: acquiring a target video, wherein the target video is a long video provided to the target user across multiple scenes; obtaining a keyframe set from the target video, wherein the keyframe set includes several keyframes; obtaining a target facial feature vector based on the keyframe set; obtaining a target clothing feature vector based on the keyframe set; and identifying the target person based on the target clothing feature vector and the target facial feature vector. This method can construct a cross-modal character representation by fusing multimodal information such as face, clothing, visual appearance, and color tone, effectively overcoming the vulnerability of single features (such as face or appearance) in complex scenarios such as occlusion, pose changes, lighting changes, and clothing changes, significantly enhancing the stability and adaptability of character identity association in long videos across multiple scenes. Furthermore, it can introduce temporal modeling and causal reasoning modules, using face clustering anchor points as the center for pre-processing. Post-temporal deduction, combined with time decay weights and weighted similarity calculations, enables long-term consistent tracking of character identities, avoiding clustering errors caused by transient feature changes and improving accuracy in complex narratives and multi-character interaction scenarios. Furthermore, it employs frame rate adaptive sampling and keyframe detection strategies, dynamically adjusting frame sparsity based on video length and scene changes. This significantly reduces computational resource consumption while maintaining clustering accuracy, making it particularly suitable for efficient processing of long videos (such as movies, TV dramas, and variety shows). Finally, it can extract high-order semantic features using large models (such as InsightFace and BGE-VL) and combine them with small models (such as the YOLO series) for rapid detection and segmentation, achieving a balance between accuracy and efficiency. This allows it to adapt to various business needs, such as film production, video retrieval, and intelligent security. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 The flowchart illustrates a method for aligning characters in a long, cross-scene video based on a large model, as provided in this embodiment of the invention. Detailed Implementation
[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0011] This invention provides a method for character alignment in long videos across different scenes based on a large model, such as... Figure 1 As shown, the method includes the following steps: S1, acquire the target video, wherein the target video is a long video across scenes provided by the target user.
[0012] S2, Obtain a keyframe set from the target video, wherein the keyframe set includes several keyframes.
[0013] Specifically, step S2 also includes the following steps: S21, obtain the target video duration t and target video frame rate f; where the target video duration is in minutes.
[0014] S22, based on t and f, determine the number of frames per second (k) of the target video, where k satisfies the following condition: .
[0015] S23, based on k and the target video, obtain the keyframe set, which can be further understood as: obtaining the target video A = {A1, ..., A...} i , ..., A m}, A i It is the set of video frames within the i-th second of the target video, where i ranges from 1 to m and m = t × 60; from each A i k video frames are extracted from each frame, and each extracted video frame is used as a keyframe.
[0016] In another specific embodiment, the method further includes the following steps: S100, obtain A i = (A i1 , ..., A ij , ..., A in A ij It is A i The j-th video frame within the range, where j ranges from 1 to n, and n is A. i The number of video frames within.
[0017] Preferably, n=24.
[0018] S101, according to A ij and A i(j-1) , obtain A ij The corresponding key change △S ij , △S ij The following conditions must be met: , where H 1 ij It is A ij With A i(j-1) Histogram differences between them, H 2ij It is A ij With A i(j-1) The change in optical flow between them, α is the change in optical flow between H 1 ij The corresponding first adjustment parameter, β, is H 2 ij The corresponding second adjustment parameter; those skilled in the art set the first and second adjustment parameters according to actual needs, which will not be elaborated here.
[0019] Furthermore, α + β = 1.
[0020] S102, according to △S ij Adjust k, where k satisfies the following condition: ,in, , .
[0021] S3, based on the set of keyframes, obtain the target face feature vector.
[0022] Specifically, step S3 also includes the following steps: S31, based on the keyframe set B={B1, ..., B...} g , ..., B z}, from B g B was identified in the middle g Initial face bounding box information V g =(V gw V gh ), B g This is the g-th keyframe, where g ranges from 1 to z, and z is the number of keyframes. V gw It is B g The initial width of the face bounding box, V gh It is B g The initial height of the face bounding box.
[0023] S32, according to B g , obtain B g The corresponding initial face recognition confidence level C g , where C g Meets the following conditions: W 0 It is the first confidence parameter, U 0 It is the second confidence parameter, φ(B) g ) is B g Its characteristics.
[0024] S33, when C g When >△C, B g As the first intermediate frame D g Where △C is the preset confidence threshold.
[0025] S34, when C g When ≤△C, B g Filter it out.
[0026] S35, according to D g , obtain D g The corresponding face bounding box area S g , among which, S g The following conditions must be met: S g =V gw ×V gh ; S36 when S g ≥S 0 g , for D g Processing is performed to obtain the target face feature vector; this can be further understood as: converting D... g Perform face alignment processing to obtain D g The corresponding second intermediate frame will be D g The corresponding second intermediate frame is input into the InsightFace model to extract D. g The initial face feature vector of the corresponding second intermediate frame and all D g The initial face feature vector of the corresponding second intermediate frame is processed to obtain the target face feature vector; the target face feature vector has a dimension of 512.
[0027] S36 when S g <S 0 g D g Filter it out.
[0028] S4. Obtain the target clothing feature vector based on the set of keyframes.
[0029] Specifically, step S4 also includes the following steps: S41, B g Inputting into the YOLO-Seg model to obtain B g The corresponding binary mask M for the initial clothing area g , of which M g M meets the following conditions: g =YOLO-Seg(B g B 0 g ), where B 0 g It is the keyframe B after filtering. g The initial face bounding box.
[0030] S42, when M gWhen =1, for B g Feature extraction is performed to obtain B g The corresponding feature vector of the target clothing.
[0031] S43, when M g When =0, B was not applied. g Feature extraction is performed.
[0032] Furthermore, step S42 also includes the following steps: S421, B g The corresponding initial clothing region is input into the BGE-VL visual semantic feature encoding model to obtain U g The corresponding first clothing feature vector L 1 g , where L 1 g The following conditions must be met: L 1 g =BGE-VL(B g ×M g ).
[0033] S422, according to U g The corresponding initial clothing area, obtain U g The corresponding first clothing feature vector L 2 g , where L 2 g =(L 2 gh L 2 gs L 2 gv ), L 2 gh =(L 21 gh L 22 gh ), L 2 gs =(L 21 gs L 22 gs ), L 2 gv =(L 21 gv L 22 gv ), L 21 gh It is the first characteristic of hue in HSV color characteristics, L 22 gh It is the second characteristic of hue in the HSV color characteristics, L 21 gsIt is the first characteristic of saturation in HSV color features, L 22 gs It is the second characteristic of saturation in HSV color characteristics, L 21 gv Lightness is the first characteristic of HSV color features. 22 gv It is the second characteristic of brightness in the HSV color characteristics.
[0034] S423, L 1 g and L 2 g Feature fusion processing is performed to obtain B g The corresponding target clothing feature vector; further understood as: U g The corresponding target clothing feature vector L g It is L 1 g With L 2 g It is obtained by concatenating vectors.
[0035] S5. Identify the target person based on the target clothing feature vector and the target face feature vector.
[0036] Specifically, step S5 also includes the following steps: S51, which identifies all target face bounding boxes in the target video.
[0037] S52, calculate the overlap area Q between any initial face recognition bounding box in the target video and the preset face recognition bounding box. r , where Q r The following conditions must be met: Among them, E r E0 is the r-th initial face recognition bounding box of the remaining keyframes after filtering and processing in the keyframe set, and E0 is the preset face recognition bounding box.
[0038] S53, when Q r When Q0 > 0, a mapping relationship is established between the initial face and the preset face, and the initial face is used as the target face; where Q0 is the preset threshold.
[0039] Furthermore, Q0 = 0.7.
[0040] S54, perform DBSCAN clustering algorithm on all the target face feature vectors corresponding to the target faces to obtain key face feature clusters U={U1, ..., U2}. a , ..., U b}; where U aIt is the a-th key face feature cluster, where the value of a ranges from 1 to b, and b is the number of key face clusters; any U x It is obtained through the DBSCAN(P, ϵ, min-samples) function relationship, where P is the set of all target face feature vectors, ϵ is the neighborhood radius, and min-samples is the minimum number of samples required for the core point.
[0041] S55, based on the key facial feature clusters and the target clothing feature vector, the target face is identified. Further, the best keyframe from any one of the key facial feature clusters is selected as the designated keyframe, where the timestamp preceding and following the timestamp of the designated keyframe is used as a preset time period. The target similarity is determined to identify the target face when the target similarity is greater than a preset similarity threshold, wherein the target similarity sim meets the following conditions: , Where t0 is the timestamp of the specified keyframe, and V x V is the target face feature vector of the x-th keyframe within a preset time period. y It is the target face feature vector of the y-th keyframe within a preset time period; V 0 x V is the feature vector of the target clothing in the x-th keyframe within a preset time period. 0 y η is the target clothing feature vector of the y-th keyframe within the preset time period; η is the weight of the first face feature, δ is the weight of the second clothing feature; λ is the attenuation coefficient; t is a certain time period within the preset time period.
[0042] This embodiment provides a method for character alignment in cross-scene long videos based on a large model. The method includes the following steps: acquiring a target video, wherein the target video is a cross-scene long video provided by a target user; acquiring a keyframe set from the target video, wherein the keyframe set includes several keyframes; acquiring a target face feature vector based on the keyframe set; acquiring a target clothing feature vector based on the keyframe set; and identifying the target character based on the target clothing feature vector and the target face feature vector. This method can construct a cross-modal character representation by fusing multimodal information such as face, clothing, visual appearance, and color tone, effectively overcoming the vulnerability of single features (such as face or appearance) in complex scenarios such as occlusion, pose changes, lighting changes, and clothing changes. It significantly enhances the stability and adaptability of character identity association in long-video cross-scene scenarios. Furthermore, it can introduce temporal modeling and... The causal reasoning module uses face clustering anchors as the center to perform sequential inference, combining time decay weights and weighted similarity calculations to achieve long-term consistent tracking of character identities. This avoids clustering errors caused by transient feature changes and improves accuracy in complex narratives and multi-character interaction scenarios. Furthermore, it employs frame rate adaptive sampling and keyframe detection strategies to dynamically adjust frame sparsity based on video length and scene changes. This significantly reduces computational resource consumption while maintaining clustering accuracy, making it particularly suitable for efficient processing of long videos (such as movies and TV series). Finally, it can extract high-order semantic features using large models (such as InsightFace and BGE-VL) and combine them with small models (such as the YOLO series) for rapid detection and segmentation, achieving a balance between accuracy and efficiency. This allows it to adapt to various business needs, such as film production, video retrieval, and intelligent security.
[0043] Embodiments of the present invention also provide a character alignment system for cross-scene long videos based on a large model, the system comprising: The first execution module is used to acquire the target video, wherein the target video is a long video across scenes provided by the target user; The second execution module is used to obtain a key frame set from the target video, wherein the key frame set includes several key frames; The third execution module is used to obtain the target face feature vector based on the key frame set; The fourth execution module is used to obtain the target clothing feature vector based on the keyframe set; The fifth execution module is used to identify the target person based on the target clothing feature vector and the target face feature vector.
[0044] While specific embodiments of the invention have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of the invention. It should also be understood that various modifications can be made to the embodiments without departing from the scope and spirit of the invention. The scope of the invention is defined by the appended claims.
Claims
1. A method for character alignment in long, cross-scene videos based on a large model, characterized in that, The method includes the following steps: S1, acquire the target video, wherein the target video is a long video across scenes provided to the target user; S2, Obtain a keyframe set from the target video, wherein the keyframe set includes several keyframes; S3, Based on the set of keyframes, obtain the target face feature vector; S4. Obtain the target clothing feature vector based on the keyframe set; S5. Identify the target person based on the target clothing feature vector and the target face feature vector.
2. The method for character alignment in cross-scene long videos based on a large model according to claim 1, characterized in that, Step S2 also includes the following steps: S21, obtain the target video duration t and the target video frame rate f; the target video duration is in minutes; S22, based on t and f, determine the number of frames per second (k) of the target video, where k satisfies the following condition: ; S23, based on k and the target video, obtain the keyframe set.
3. The method for character alignment in cross-scene long videos based on a large model according to claim 2, characterized in that, In step S23, the target video A = {A1, ..., A2} is obtained. i , ..., A m }, A i It is the set of video frames within the i-th second of the target video, where i ranges from 1 to m and m = t × 60; from each A i k video frames are extracted from each frame, and each extracted video frame is used as a keyframe.
4. The method for character alignment in cross-scene long videos based on a large model according to claim 2, characterized in that, The method further includes the following steps: S100, obtain A i = (A i1 , ..., A ij , ..., A in A ij It is A i The j-th video frame within the range, where j ranges from 1 to n, and n is A. i The number of video frames within; S101, according to A ij and A i(j-1) , obtain A ij The corresponding key change △S ij , △S ij Meets the following conditions: , where H 1 ij It is A ij With A i(j-1) Histogram differences between them, H 2 ij It is A ij With A i(j-1) The change in optical flow between H, α is the change in optical flow between H 1 ij The corresponding first adjustment parameter, β, is H 2 ij The corresponding second adjustment parameter; S102, according to △S ij Adjust k.
5. The method for character alignment in cross-scene long videos based on a large model according to claim 4, characterized in that, n=24。 6. The method for character alignment in cross-scene long videos based on a large model according to claim 1, characterized in that, Step S3 also includes the following steps: S31, based on the keyframe set B={B1, ..., B...} g , ..., B z }, from B g B was identified in the middle g Initial face bounding box information V g =(V gw V gh ), B g This is the g-th keyframe, where g ranges from 1 to z, and z is the number of keyframes. V gw It is B g The initial width of the face bounding box, V gh It is B g The initial height of the face bounding box; S32, according to B g , obtain B g The corresponding initial face recognition confidence level C g , where C g Meets the following conditions: W 0 It is the first confidence parameter, U 0 It is the second confidence parameter, φ(B) g ) is B g Features; S33, when C g When >△C, B g As the first intermediate frame D g Where △C is the preset confidence threshold; S34, when C g When ≤△C, B g Filter it out; S35, according to D g , obtain D g The corresponding face bounding box area S g , of which S g The following conditions must be met: S g =V gw ×V gh ; S36, when S g ≥S 0 g , for D g The process is performed to obtain the target face feature vector; S37, when S g <S 0 g D g Filter it out.
7. The method for character alignment in cross-scene long videos based on a large model according to claim 1, characterized in that, In step S36, D g Perform face alignment processing to obtain D g The corresponding second intermediate frame will be D g The corresponding second intermediate frame is input into the InsightFace model to extract D. g The initial face feature vector of the corresponding second intermediate frame and all D g The initial face feature vector of the corresponding second intermediate frame is processed to obtain the target face feature vector; the target face feature vector has a dimension of 512.
8. The method for character alignment in cross-scene long videos based on a large model according to claim 1, characterized in that, Step S4 also includes the following steps: S41, U g Inputting it into the YOLO-Seg model yields U g The corresponding binary mask M for the initial clothing area g , of which M g M meets the following conditions: g =YOLO-Seg(B g B 0 g ), where B 0 g It is the keyframe U after filtering. g The initial face bounding box; S42, when M g When =1, for B g Feature extraction is performed to obtain B g The corresponding target clothing feature vector; S43, when M g When =0, B was not applied. g Feature extraction is performed.
9. The method for character alignment in cross-scene long videos based on a large model according to claim 1, characterized in that, Step S42 also includes the following steps: S421, U g The corresponding initial clothing region is input into the BGE-VL visual semantic feature encoding model to obtain U g The corresponding first clothing feature vector L 1 g , where L 1 g The following conditions must be met: L 1 g =BGE-VL(U g ×M g ); S422, according to U g The corresponding initial clothing area, obtain U g The corresponding first clothing feature vector L 2 g , where L 2 g =(L 2 gh L 2 gs L 2 gv ), L 2 gh =(L 21 gh L 22 gh ), L 2 gs =(L 21 gs L 22 gs ), L 2 gv =(L 21 gv L 22 gv ), L 21 gh It is the first characteristic of hue in HSV color characteristics, L 22 gh It is the second characteristic of hue in the HSV color characteristics, L 21 gs It is the first characteristic of saturation in HSV color features, L 22 gs It is the second characteristic of saturation in HSV color characteristics, L 21 gv Lightness is the first characteristic of HSV color features. 22 gv It is the second characteristic of lightness in the HSV color characteristics; S423, L 1 g and L 2 g Perform feature fusion processing to obtain U g The corresponding feature vector of the target clothing.
10. A character alignment system for cross-scene long videos based on a large model, characterized in that, The system includes: The first execution module is used to acquire the target video, wherein the target video is a long video across scenes provided by the target user; The second execution module is used to obtain a key frame set from the target video, wherein the key frame set includes several key frames; The third execution module is used to obtain the target face feature vector based on the key frame set; The fourth execution module is used to obtain the target clothing feature vector based on the keyframe set; The fifth execution module is used to identify the target person based on the target clothing feature vector and the target face feature vector.
Citation Information
Patent Citations
Feature extraction method, device and terminal equipment
CN111242189B
Film and television character classification method, device, computer equipment and storage medium
CN114282057B