A video processing method, a visual analysis model training method, and an electronic device

The visual analysis model, which utilizes multi-granularity recognition and cross-attention mechanisms, addresses the need for detailed interaction in video content in existing technologies. It enables multi-dimensional analysis of video content and provides detailed information, thereby enhancing the user experience.

CN120431503BActive Publication Date: 2026-03-31HONOR DEVICE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing algorithms for video-based intelligent interaction scenarios are insufficient to meet users' diverse and detailed interactive needs for video content, and cannot effectively perform multi-granularity recognition and analysis.

Method used

A multi-granularity recognition method is used to finely segment videos into scenes, shots, and behaviors. A visual analysis model is used to learn through a cross-attention mechanism to obtain recognition results at the scene granularity, shot granularity, and behavior granularity of the video. Combined with voice and text descriptions, more accurate scene segmentation is achieved.

Benefits of technology

It enables multi-dimensional and detailed understanding and analysis of video content, providing more detailed video information, supporting video understanding and analysis in various application scenarios, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431503B_ABST
    Figure CN120431503B_ABST
Patent Text Reader

Abstract

The application discloses a video processing method, a visual analysis model training method and an electronic device, and relates to the technical field of video processing, and comprises the following steps: an electronic device performs visual analysis on a to-be-processed video, and obtains a behavior division result and a shot division result of the to-be-processed video. The behavior division result represents the division result of video clips corresponding to different behaviors of a subject in the to-be-processed video, and the shot division result represents the division result of video clips corresponding to different shots in the to-be-processed video. The electronic device obtains a scene division result corresponding to the to-be-processed video based on the content description of the video clips corresponding to different shots contained in the shot division result. The scene division result represents the division result of video clips corresponding to different scenes in the to-be-processed video. In the scheme, the understanding and division of the video content are more dimensional and detailed, and more detailed video content information can be provided to support video understanding and analysis in various application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to a video processing method, a training method for a visual analysis model, and an electronic device. Background Technology

[0002] Video is gradually becoming the primary medium for people to obtain information, and video-based intelligent interaction scenarios are becoming increasingly common. These scenarios can include video content location, generation of highlight moments, and intelligent video creation, among others.

[0003] These video-based intelligent interactive scenarios require the identification and analysis of video content to obtain corresponding output results. For example, outputting video content tags for video content localization; outputting video segment tags to extract highlights; or outputting several target segments of the video to achieve intelligent video production.

[0004] As videos contain increasingly rich and diverse content, users have more detailed interactive needs. However, existing algorithms for intelligent video interaction scenarios have limited functionality and are insufficient to meet users' interactive requirements. Summary of the Invention

[0005] This application provides a video processing method, a visual analysis model training method, and an electronic device. The electronic device can perform multi-granularity recognition on the video to be processed, and obtain recognition results at the scene granularity, shot granularity, and behavior granularity of the video to be processed. The understanding and division of video content is more multi-dimensional and detailed, and can provide more detailed video content information to support video understanding and analysis in various application scenarios in the later stage.

[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions.

[0007] Firstly, a video processing method is provided, the method comprising:

[0008] Electronic devices acquire videos to be processed.

[0009] Electronic devices perform visual analysis on the video to be processed to obtain the behavior segmentation results and shot segmentation results of the video to be processed.

[0010] Among them, the behavior segmentation result represents the segmentation result of video segments corresponding to different behaviors of the subject in the video to be processed, and the shot segmentation result represents the segmentation result of video segments corresponding to different shots in the video to be processed.

[0011] The electronic device obtains the scene segmentation result corresponding to the video to be processed based on the content description of the video segments corresponding to different shots included in the shot segmentation result.

[0012] The scene segmentation result represents the segmentation result of video segments corresponding to different scenes in the video to be processed; a scene contains one or more shots.

[0013] For a video, the granularity, from coarse to fine, can be scene-level, shot-level, and action-level. That is, scene changes are always accompanied by shot changes, and a single shooting scene may include one or more shots. A single shot may contain one or more different behaviors of the subject in the video, or multiple shots may contain the same behavior of the subject. The subject can be a person or an animal. When the subject is a person, the behavior can include different actions in different scenarios such as walking, eating, dancing, running, shooting, and jumping. When the subject is an animal, the behavior can include different actions in different scenarios such as walking, eating, hunting, drinking, and running. A video can be segmented at scene-level, shot-level, and action-level granularity to support data use in subsequent video understanding and analysis. Understandably, the granularity of video segmentation can also include other granularities. For example, other granularities can include the subject's emotions, such as anger, sadness, happiness, excitement, depression, etc., and video segments can be identified and segmented according to different emotions. Other granularities can include the subject's facial expressions, such as smiling, laughing, crying, etc., and video segments can be identified and segmented according to different facial expressions. For example, other granularities can include the number of subjects, such as segmenting video segments based on whether the subjects include a single person or multiple people. For example, other granularities can include the video's shooting color tone, such as segmenting video segments based on different video shooting color tones.

[0014] In this application, the electronic device can perform multi-granularity recognition on the video to be processed, obtaining recognition results at the scene granularity, shot granularity, and behavior granularity. This provides a more multi-dimensional and detailed understanding and segmentation of video content, offering more detailed video content information to support video understanding and analysis in various application scenarios. For example, it can enable the search and location of more detailed parts of the video content, and can process (edit, composite, etc.) content featuring the same behavior or shot within the video content. Simultaneously, it improves the user experience of using the video processing functions provided by the electronic device.

[0015] In one possible implementation of the first aspect, the electronic device performs visual analysis on the video to be processed to obtain behavior segmentation results and shot segmentation results of the video to be processed, including:

[0016] The electronic device performs visual analysis on the video to be processed, and obtains at least one behavior label and at least one shot label from the video to be processed.

[0017] The electronic device uses at least one behavior label and the timestamp of the video segment corresponding to the behavior label as the behavior segmentation result.

[0018] The electronic device uses at least one lens tag and the timestamp of the video segment corresponding to the lens tag as the lens segmentation result.

[0019] Among them, the types of behavior tags include multiple preset behavior tags, non-preset behavior tags, or no behavior tags, and the types of shot tags include no shot switching tags and shot tags corresponding to multiple preset shot switching types.

[0020] In this application, when the subject is a human, the various preset behavior tags can include behavior tags for different scenarios such as walking, eating, dancing, running, shooting, and jumping. When the subject is an animal, the various preset behavior tags can include behavior tags for different scenarios such as walking, eating, hunting, drinking, and running. Non-preset behavior tags are tags corresponding to behaviors other than the various preset behaviors. No behavior tag may mean that there is no tag corresponding to the subject in the video frame. The shot tags corresponding to the various preset shot transition types can include tags for different types of shot changes such as shot transition (cut), shot zoom-in, shot zoom-out, and gradual transition (fade).

[0021] The shot segmentation result of the video to be processed can be expressed as follows:

[0022] [00:00:00.000-00:01:00.960s, camera movement (1),

[0023] 00:01:01.000-00:01:48.150s, camera change (2),

[0024] 00:01:58.000-00:02:05.400s, camera movement (3)].

[0025] The behavior segmentation results can be represented as:

[0026] [00:01:58.000-00:02:33.140s, Walking (action1),

[0027] 00:02:35.750-00:03:00.440s, eating (action2)).

[0028] After obtaining at least one behavior label and at least one shot label from the video to be processed through visual analysis, the electronic device obtains the behavior segmentation result and shot segmentation result of the video to be processed based on the start and end timestamps of the video segments corresponding to each label. This provides a more multi-dimensional and detailed understanding and segmentation of video content, and can provide more detailed video content information to support video understanding and analysis in various application scenarios in the future.

[0029] In another possible implementation of the first aspect, the electronic device performs visual analysis on the video to be processed to obtain at least one behavior tag and at least one shot tag for the video to be processed, including:

[0030] The electronic device acquires multiple video segments of the video to be processed in a sequential manner. The duration of each video segment is the same as the preset sliding window length.

[0031] The electronic device performs visual analysis on each video segment to be processed, and obtains at least one behavior tag and at least one shot tag for each video segment to be processed.

[0032] If adjacent video segments to be processed overlap, the electronic device performs deduplication and merging of the behavior tags and shot tags corresponding to the overlapping segments to obtain the behavior tags and shot tags of all video segments to be processed.

[0033] In this application, the electronic device can segment the video to be processed based on the sliding window length and perform deduplication on the results, which enables the video processing method to be applied to the processing of videos of different durations, making the application scenarios of the video processing method more extensive.

[0034] In another possible implementation of the first aspect, the electronic device obtains the scene segmentation result corresponding to the video to be processed based on the content description of the video segments corresponding to different shots included in the shot segmentation result, including:

[0035] The electronic device performs video segmentation on the video to be processed based on the timestamps of the video segments corresponding to each lens label in the lens segmentation results, and obtains multiple video segments after segmentation.

[0036] Among these, the shots of video clips that are temporally adjacent are different.

[0037] The electronic device performs video content recognition on each video segment and obtains a content description for each segment. The content description includes the scene content corresponding to the video segment.

[0038] The electronic device obtains the scene segmentation results corresponding to the video to be processed based on the content description of each video segment.

[0039] In this application, the scene segmentation result can be represented as:

[0040] [00:00:00.000-00:01:00.960s, Scene 1]

[0041] 00:01:01.000-00:01:48.150s, Scene 2

[0042] 00:01:58.000-00:02:05.400s, Scene 3.

[0043] The electronic device obtains the scene segmentation results corresponding to the video to be processed based on the content description of the video segments corresponding to different shots included in the shot segmentation results. By obtaining the description of the video content, the understanding and segmentation of the video content becomes more detailed, and the scene segmentation becomes more accurate.

[0044] In another possible implementation of the first aspect, the electronic device obtains the scene segmentation result corresponding to the video to be processed based on the content description of each video segment, including:

[0045] The electronic device determines the scene description for each video segment based on the content description of each video segment.

[0046] The scene description includes a description of the environment scene corresponding to the video clip.

[0047] If the scene descriptions of adjacent video segments have a similarity greater than or equal to a similarity threshold, the adjacent video segments are determined to belong to the same scene label, and the scene labels of the video segments corresponding to different scenes are obtained.

[0048] The scene segmentation results are obtained by using different scene labels and the timestamps of the videos to be processed corresponding to the scene labels.

[0049] In this application, if the similarity of scene descriptions of adjacent video segments is greater than or equal to a similarity threshold, meaning that the scene descriptions of adjacent video segments are highly similar, then the adjacent video segments are considered to describe the same scene and are labeled with the same scene tag. The electronic device performs scene recognition and segmentation based on the scene description corresponding to each video content description, and obtains the start and end timestamps of scene changes to obtain the scene segmentation result of the video to be processed, resulting in a more accurate scene segmentation result.

[0050] In another possible implementation of the first aspect, the electronic device obtains the scene segmentation result corresponding to the video to be processed based on the content description of the video segments corresponding to different shots included in the shot segmentation result, including:

[0051] Electronic devices perform voice analysis on the video to be processed to obtain a voice-text description of the video.

[0052] The voice text description includes one or more of the following: the text description corresponding to human dialogue in the video to be processed, the text description corresponding to the narrator's voice in the video to be processed, and the text description corresponding to the sound effects in the video to be processed.

[0053] The electronic device performs video segmentation on the video to be processed based on the timestamps of the video segments corresponding to each lens label in the lens segmentation results, and obtains multiple video segments after segmentation.

[0054] Among these, the shots of video clips that are temporally adjacent are different.

[0055] The electronic device performs video content recognition on each video segment and obtains a content description for each video segment.

[0056] The electronic device obtains the scene segmentation results corresponding to the video to be processed based on the voice and text description of the video and the content description of each video segment.

[0057] For example, the video content description can be represented as:

[0058] Shot 1, 00:00:00.000-00:01:00.960s: The video begins with an aerial view overlooking a cracked wasteland, showing a vast expanse of dry land. A tornado begins to form in this wasteland, stirring up dust and debris. The tornado gradually grows larger and stronger, eventually dominating the frame. The background remains consistent, with the desolate terrain contrasting sharply with the dynamics of the tornado.

[0059] Scene 2, 00:01:01.000-00:01:48.150s: A tornado forms and moves across a barren, open area. It rapidly grows from a small, concentrated column of dust into a spiral structure that increases significantly in both height and width. The background is a vast, flat landscape with sparse vegetation and a cloudy sky.

[0060] Scene 3, 00:01:58.000-00:02:05.400s: The video begins with a wide-angle shot of a wasteland, where a tornado is forming and rising. The sky is overcast, and as the tornado rises, text describing the scene appears on the screen. The camera then focuses on a skeleton on the ground, describing the location as being in the shadow of Mount Kilimanjaro.

[0061] For example, a speech-text segment can be represented as:

[0062] In the audio text segment 1, 00:00:00.000-00:00:30.040s, on the Amboseli Plain, in the shadow of Mount Kilimanjaro, the seasonal rains that should have come for the past two years have not arrived.

[0063] Audio text segment 2, 00:31:00.000-00:01:00.960s: This year, they should have appeared long ago. This is the worst drought in half a century. Amboseli is usually a haven for elephants.

[0064] In audio text segment 3, 00:01:01.000-00:01:48.150s, it says, "These plains should have been green and covered with grass. Now, there's nothing but dust. This family is forced to constantly migrate in search of anything they can eat."

[0065] In audio text segment 4, 00:01:58.000-00:02:05.400s, it says, "The calf must keep up with the pace of migration, sometimes without even time to suckle. As the grass disappears, the elephant can only grab withered branches from the dust. Adult elephants may be able to survive on these, but they cannot sustain the calf for long."

[0066] In this application, the electronic device obtains the scene segmentation result corresponding to the video to be processed based on the content description of the video segments corresponding to different shots included in the shot segmentation result, and the speech text segment obtained by speech analysis of the video to be processed. Combining the text descriptions of the two dimensions for scene recognition makes the understanding and segmentation of video content more detailed and the scene segmentation more accurate.

[0067] In another possible implementation of the first aspect, the audio-text description of the video to be processed includes multiple audio-text segments divided by time / time sequence.

[0068] The electronic device obtains the scene segmentation results corresponding to the video to be processed based on the audio-text description of the video and the content description of each video segment, including:

[0069] The electronic device uses each audio text segment as a reference description and determines the scene label for each video segment based on the content description of each video segment.

[0070] If a description of a voice text segment covers the content descriptions of multiple adjacent video segments, the adjacent video segments are determined to belong to the same scene label.

[0071] The scene segmentation results are obtained by using different scene labels and the timestamps of the videos to be processed corresponding to the scene labels.

[0072] In this application, the electronic device can use each voice text segment as a reference description, determine the scene label of each video segment based on the content description of each video segment, and perform scene recognition by combining the two-dimensional text description. Through multi-dimensional data processing, the scene segmentation results are more accurate.

[0073] In another possible implementation of the first aspect, the electronic device uses each speech text segment as a reference description and determines the scene label for each video segment based on the content description of each video segment, including:

[0074] If the content description of at least one video segment matches the audio text segment, the audio text segment is used as a reference description for at least one video segment to determine the scene label of at least one video segment.

[0075] In this application, if at least one video segment's content description matches an audio text segment, meaning a scene is described in two dimensions—audio text and video content description—the electronic device can perform scene recognition on the two-dimensional text descriptions to obtain the scene description corresponding to each audio text segment / video content description. Based on the scene descriptions, scene recognition and segmentation are performed, and the start and end timestamps of scene transitions are obtained to obtain the scene segmentation result for the video to be processed. Combining scene recognition and segmentation from both audio text and video content description dimensions, and through multi-dimensional data processing, the resulting scene segmentation result is more accurate.

[0076] In another possible implementation of the first aspect, the plurality of speech-text segments include a first speech-text segment, and the method further includes:

[0077] If the content descriptions of all video clips do not match the first audio text segment, discard the first audio text segment.

[0078] In this application, if a speech text segment has a very low similarity to any text describing the video content—that is, if the speech text segment does not match any video content description and there is no similar scene / subject description—then the speech text segment is determined to be a text segment with low relevance to the video content. To avoid interference from this type of speech text segment in scene recognition, it can be discarded and not used as data support for scene recognition. Discarding speech text segments without valid information can greatly reduce their impact on video scene recognition, thereby making video scene recognition more accurate.

[0079] In another possible implementation of the first aspect, the method further includes:

[0080] The electronic device stores the behavior segmentation results, shot segmentation results, and scene segmentation results of the video to be processed into a video tag database. The video tag database stores the behavior segmentation results, shot segmentation results, and scene segmentation results for multiple videos.

[0081] In this application, the electronic device can store scene segmentation results, shot segmentation results, and behavior segmentation results into a video tag database. The electronic device can provide various user-oriented video processing application scenarios, using the video tags of each video in the video tag database for corresponding data support.

[0082] In another possible implementation of the first aspect, the video tag database includes behavior tags for behavior segmentation, shot tags for shot segmentation, and scene tags for scene segmentation. The method further includes:

[0083] The electronic device receives a video content query request.

[0084] The video content query request may include one or more of the following: scene description, shot description, or behavior description.

[0085] Electronic devices retrieve target tags that match the video content query request from scene segmentation, shot segmentation, and behavior segmentation results of multiple videos in the video tag database.

[0086] The target tags include one or more of the following: behavior tags, shot tags, and scene tags.

[0087] The electronic device obtains the target video segment corresponding to the timestamp of the target tag.

[0088] The electronic device outputs the target video clip.

[0089] In this application, the electronic device can perform multi-granularity recognition on the video to be processed, obtaining recognition results at the scene granularity, shot granularity, and behavior granularity. This provides a more detailed understanding and segmentation of the video content, offering more detailed video content information to support video understanding and analysis in various application scenarios. For example, in video content query scenarios, it can enable the search and location of more detailed parts of the video content, while also improving the user experience of using the video processing functions provided by the electronic device.

[0090] In another possible implementation of the first aspect, the video tag database includes behavior tags for behavior segmentation, shot tags for shot segmentation, and scene tags for scene segmentation. The method further includes:

[0091] The electronic device receives a video editing request.

[0092] The video editing request includes one or more descriptions of the scene, shot, or behavior of the video to be edited.

[0093] Electronic devices retrieve target tags that match the video editing request from scene segmentation, shot segmentation, and behavior segmentation results of multiple videos in the video tag database.

[0094] The target tags include one or more of the following: behavior tags, shot tags, and scene tags.

[0095] The electronic device acquires at least one target video segment corresponding to the timestamp of the target tag.

[0096] The electronic device outputs a video synthesized from at least one target video segment.

[0097] In this application, the electronic device can perform multi-granularity recognition on the video to be processed, obtaining recognition results at the scene granularity, shot granularity, and behavior granularity. This provides a more detailed understanding and segmentation of the video content, offering more detailed video content information to support video understanding and analysis in various subsequent application scenarios. For example, in a video editing request scenario, it can process (edit, composite, etc.) content that exhibits the same behavior or shot within the video content, while also improving the user experience of using the video processing functions provided by the electronic device.

[0098] In another possible implementation of the first aspect, the electronic device performs visual analysis on the video to be processed to obtain behavior segmentation results and shot segmentation results of the video to be processed, including:

[0099] The electronic device inputs the video to be processed into the visual analysis model for visual analysis, and obtains the behavior segmentation results and shot segmentation results of the video to be processed.

[0100] The visual analysis model is a model obtained by learning and training using multiple cross-attention mechanisms. The input of the visual analysis model is the video to be processed, and the output of the visual analysis model includes frame-level behavior prediction labels and frame-level shot prediction labels. Behavior segmentation results are obtained by merging adjacent image frames with the same behavior prediction labels, and shot segmentation results are obtained by merging adjacent image frames with the same shot prediction labels.

[0101] The training process of the visual analysis model includes taking turns using frame-level feature vectors, shot-level feature vectors, and behavior-level feature vectors as query vectors, key vectors, and numerical vectors in the cross-attention mechanism for cross-learning to obtain behavior prediction labels, shot prediction labels, frame-level behavior prediction labels, and frame-level shot prediction labels. Based on these labels, the model loss is calculated, and the visual analysis model is trained until the model loss is less than a loss threshold.

[0102] In this application, the electronic device inputs the video to be processed into a visual analysis model for visual analysis, obtaining behavior segmentation results and shot segmentation results. The behavior segmentation results represent the segmentation of different behaviors of the subject in the video, while the shot segmentation results represent the segmentation of different shots in the video. Further, based on the content descriptions of the video segments corresponding to different shots included in the shot segmentation results, the electronic device obtains the scene segmentation results corresponding to the video. The scene segmentation results represent the segmentation of different scenes in the video. The electronic device can perform multi-granularity recognition on the video to be processed, obtaining scene-granularity recognition results, shot-granularity recognition results, and behavior-granularity recognition results. It can simultaneously perform multi-granularity temporal segmentation of the video, including coarse-to-fine segmentation of scenes, shots, and behaviors, providing a more detailed understanding and segmentation of video content and offering more detailed video content information to support video understanding and analysis in various application scenarios later on. The visual analysis model utilizes multiple cross-attention mechanisms to allow frame-level high-dimensional semantic information to interact with shot-level spatial detail information and behavior-level semantic information, thereby improving the prediction accuracy of frame-level shot prediction labels and frame-level behavior prediction labels.

[0103] Secondly, a training method for a visual analysis model is provided. The visual analysis module includes a feature extraction module and a learning interaction module. This method includes:

[0104] The electronic device inputs the sample video into the feature extraction module of the visual analysis model to obtain the first shot feature vector, the first line feature vector, and the first frame feature vector, respectively.

[0105] The electronic device inputs the first lens feature vector, the first line feature vector, and the first frame feature vector into the learning interaction module, and uses a cross-attention mechanism to learn and output a predicted label.

[0106] The prediction labels include frame-level behavior prediction labels and frame-level shot prediction labels.

[0107] Electronic devices train a visual analysis model to obtain a trained visual analysis model that meets the training conditions.

[0108] Among them, meeting the training conditions includes the visual analysis model's loss being less than the loss threshold, or the visual analysis model reaching convergence, or the visual analysis model reaching the required number of training iterations threshold.

[0109] The feature extraction module can be a backbone network obtained after video classification pre-training, such as VideoMaskedAutoencoders (VideoMAE), Multimodal Video Understanding Model (InternVideo2), Video Vision Model (ViViT), or 3D Convolutional Model.

[0110] In this application, the visual analysis model utilizes multiple cross-attention mechanisms during training to allow frame-level high-dimensional semantic information to interact with shot-level spatial detail information and behavior-level semantic information, thereby improving the prediction accuracy of frame-level shot prediction labels and frame-level behavior prediction labels. The trained visual analysis model can output relatively accurate frame-level shot prediction labels and frame-level behavior prediction labels, enabling electronic devices to obtain relatively accurate behavior segmentation and shot segmentation results. The visual analysis model can simultaneously perform multi-granular temporal segmentation of videos, including coarse-to-fine segmentation of shots and behaviors, resulting in a more detailed understanding and segmentation of video content, supporting video understanding and analysis in various application scenarios.

[0111] In one possible implementation of the second aspect, the electronic device inputs the sample video into the feature extraction module of the visual analysis model to obtain the first shot feature vector, the first line feature vector, and the first frame feature vector, respectively, including:

[0112] The electronic device inputs the sample video into the feature extraction module of the visual analysis model. The shallow network of the feature extraction module outputs the first shot feature vector, and the deep network of the feature extraction module outputs the first frame feature vector and the first line feature vector.

[0113] In this application, camera angle changes often cause changes in the content of the image, and these changes inevitably result in differences in pixel distribution and pixel values. Shallow features typically contain more spatial detail information, and the pixel differences represented by shallow features can effectively identify camera angle changes. Outputting the first camera action feature through a shallow network improves the efficiency of acquiring camera action features. Deep features typically contain high-dimensional abstract speech information. Outputting action feature vectors (first action feature vectors) and frame feature vectors (first frame feature vectors) through a deep network extracts feature vectors containing more semantic information and are more accurate.

[0114] In another possible implementation of the second aspect, the learning interaction module includes an initial layer and an interaction layer. The initial layer includes a first behavior learning module, a first frame learning module, and a first shot learning module. The interaction layer includes a second behavior learning module, a second shot learning module, a second frame learning module, and a third frame learning module.

[0115] The electronic device inputs the first shot feature vector, the first line feature vector, and the first frame feature vector into the learning interaction module, performs learning using a cross-attention mechanism, and outputs predicted labels, including:

[0116] The electronic device inputs the first lens feature vector into the first lens learning module of the initial layer to obtain the second lens feature vector; the receptive field of the second lens feature vector is larger than that of the first lens feature vector.

[0117] The electronic device inputs the first frame feature vector into the first frame learning module of the initial layer to obtain the second frame feature vector; the receptive field of the second frame feature vector is larger than that of the first frame feature vector.

[0118] The electronic device uses the feature vector of the second frame as the query vector in the cross-attention mechanism, and the feature vector of the second shot as the key vector and numerical vector in the cross-attention mechanism, and inputs them into the second shot learning module of the interaction layer to obtain the feature vector of the third shot.

[0119] Electronic devices obtain lens prediction labels for sample videos based on the feature vector of the third lens.

[0120] The learning module can be a neural network model such as a deep learning model (transformer) or a dilated convolution model (dilatedconv).

[0121] In this application, multiple cross-attention mechanisms are used to allow frame-level high-dimensional semantic information (second frame feature vector) to interact with shot-level spatial detail information (second shot feature vector), so that the second shot learning module can be biased towards learning behavioral change information between each frame, thereby improving the prediction accuracy of shot prediction labels.

[0122] In another possible implementation of the second aspect, the electronic device inputs the first lens feature vector, the first action feature vector, and the first frame feature vector into the learning interaction module, performs learning using a cross-attention mechanism, and outputs a predicted label, further including:

[0123] The electronic device inputs the first row of feature vectors into the first row learning module of the initial layer to obtain the second row of feature vectors; the receptive field of the second row of feature vectors is larger than that of the first row of feature vectors.

[0124] The electronic device uses the second frame feature vector as the query vector in the cross-attention mechanism, and the second row feature vector as the key vector and numerical vector in the cross-attention mechanism, and inputs them into the second row learning module of the interaction layer to obtain the third row feature vector.

[0125] Electronic devices obtain behavioral prediction labels for sample videos based on third-behavior feature vectors.

[0126] In this application, multiple cross-attention mechanisms are used to allow frame-level high-dimensional semantic information (second frame feature vector) to interact with behavior-level semantic information (second behavior feature vector), so that the second behavior learning module can be biased towards learning behavior change information between each frame, thereby improving the prediction accuracy of behavior prediction labels.

[0127] In another possible implementation of the second aspect, the electronic device inputs the first lens feature vector, the first action feature vector, and the first frame feature vector into the learning interaction module, performs learning using a cross-attention mechanism, and outputs a predicted label, further including:

[0128] The electronic device uses the feature vector of the third lens as the query vector in the cross-attention mechanism, and the feature vector of the second frame as the key vector and numerical vector in the cross-attention mechanism, and inputs them into the second frame learning module of the interaction layer to obtain the feature vector of the third frame.

[0129] The electronic device uses the third frame feature vector as the query vector in the cross-attention mechanism, and the second frame feature vector as the key vector and numerical vector in the cross-attention mechanism, and inputs them into the third frame learning module of the interaction layer to obtain the fourth frame feature vector.

[0130] The electronic device obtains the frame-level shot prediction label corresponding to the sample video based on the feature vector of the third frame.

[0131] The electronic device obtains the frame-level behavior prediction label corresponding to the sample video based on the feature vector of the fourth frame.

[0132] In this application, multiple cross-attention mechanisms are used to allow frame-level high-dimensional semantic information (second frame feature vector) to interact with shot-level spatial detail information (third shot feature vector) and behavior-level semantic information (third behavior feature vector), respectively. This allows the second frame learning module to focus on shot change information and the third frame learning module to focus on behavior change information, thereby improving the prediction accuracy of frame-level shot prediction labels and frame-level behavior prediction labels.

[0133] In another possible implementation of the second aspect, the visual analysis module also includes a multilayer perceptron;

[0134] The electronic device obtains frame-level shot prediction labels corresponding to the sample video based on the feature vector of the third frame, including:

[0135] The electronic device inputs the feature vector of the third frame into a multilayer perceptron for dimensionality reduction processing to obtain the frame-level shot prediction label corresponding to the sample video.

[0136] Furthermore, the electronic device obtains frame-level behavior prediction labels corresponding to the sample video based on the feature vector of the fourth frame, including:

[0137] The electronic device inputs the feature vector of the fourth frame into a multilayer perceptron for dimensionality reduction processing to obtain the frame-level behavior prediction label corresponding to the sample video.

[0138] In this application, after the feature vector is reduced in dimensionality by the output of the multilayer perceptron, frame-level behavior prediction labels and frame-level shot prediction labels are obtained. The model loss is calculated based on the frame-level behavior prediction labels and frame-level shot prediction labels for model training, which can improve model efficiency and model training efficiency, while also reducing the risk of model overfitting.

[0139] In another possible implementation of the second aspect, the electronic device trains a visual analysis model to obtain the trained visual analysis model, including:

[0140] The electronic device calculates the lens prediction loss based on the lens prediction label and the lens standard label corresponding to the sample video.

[0141] Electronic devices calculate behavior prediction loss based on behavior prediction labels and the behavior standard labels corresponding to sample videos.

[0142] The electronic device calculates the frame-level shot prediction loss based on the frame-level shot prediction label and the standard shot label corresponding to the sample video.

[0143] Electronic devices calculate frame-level behavior prediction loss based on frame-level behavior prediction labels and behavior standard labels corresponding to sample videos.

[0144] Electronic devices train a visual analysis model based on the sum of one or more of the following losses: lens prediction loss, behavior prediction loss, frame-level lens prediction loss, and frame-level behavior prediction loss, until the loss is less than a loss threshold, and then obtain the trained visual analysis model.

[0145] The loss can include lens prediction loss calculated based on lens prediction labels and standard lens labels, and behavior prediction loss calculated based on behavior prediction labels and standard behavior labels. The loss functions for calculating lens prediction loss and behavior prediction loss can be multi-class loss functions such as cross-entropy loss, or direct negative logarithmic loss. The loss can also include the sum of frame-level behavior prediction loss calculated based on frame-level behavior prediction labels and standard behavior labels, and frame-level lens prediction loss calculated based on frame-level lens prediction labels and standard lens labels. The loss functions for calculating frame-level behavior prediction loss and frame-level lens prediction loss can be multi-class loss functions such as cross-entropy loss (one-hot or many-hot).

[0146] In this application, the electronic device can train a visual analysis model based on the sum of one or more losses until the loss of the visual analysis model is less than a loss threshold, or until the visual analysis model converges, or until the training iterations are reached, ultimately obtaining a well-trained visual analysis model. In some embodiments, the electronic device can also acquire commonly used loss functions such as cross-attention loss to train the visual analysis model and obtain a well-trained visual analysis model. The visual analysis model trained based on the sum of one or more losses has better performance, and its output predicted labels (behavior segmentation results / scene segmentation results) have higher accuracy.

[0147] In another possible implementation of the second aspect, the visual analysis module further includes a timing processing module, and the method further includes:

[0148] After the electronic device acquires the frame-level behavior prediction labels of the sample video, adjacent image frames with the same frame-level behavior prediction labels are merged in time sequence to obtain the merged behavior prediction labels, which are used as the behavior segmentation results.

[0149] Frame-level behavior prediction loss is calculated based on the merged behavior prediction labels and the behavior standard labels corresponding to the sample videos.

[0150] After the electronic device acquires the frame-level shot prediction labels of the sample video, adjacent image frames with the same frame-level shot prediction labels are merged in chronological order to obtain the merged shot prediction labels, which serve as the shot segmentation results.

[0151] Frame-level shot prediction loss is calculated based on the merged shot prediction labels and the standard shot labels corresponding to the sample videos.

[0152] In this application, adjacent image frames with the same frame-level action prediction label can be merged, and the frame-level action prediction loss can be calculated based on the merged action prediction label and the action standard label corresponding to the sample video. Similarly, adjacent image frames with the same frame-level shot prediction label can be merged, and the frame-level shot prediction loss can be calculated based on the merged shot prediction label and the shot standard label corresponding to the sample video. Post-processing the frame-level action prediction label and frame-level shot prediction label data limits the amount of data required to calculate the loss, reduces the computational effort required to calculate the loss, and improves the training efficiency of the model.

[0153] In another possible implementation of the second aspect, the visual analysis module further includes a sliding window processing module, and the method further includes:

[0154] The electronic device inputs the sample video into the sliding window processing module to obtain at least one sample video segment. The duration of the at least one sample video segment is the same as the preset sliding window length.

[0155] The electronic device inputs the sample video into the feature extraction module of the visual analysis model to obtain the first shot feature vector, the first line feature vector, and the first frame feature vector, including:

[0156] For each sample video segment, the electronic device inputs the sample video segment into the feature extraction module of the visual analysis model to obtain the first shot feature vector, the first line feature vector, and the first frame feature vector, respectively.

[0157] In this application, the electronic device can input sample videos into the sliding window processing module, and divide the video to be processed into segments based on the sliding window length. This allows the visual analysis model to be used for processing sample videos of different durations and for model training, making the application scenarios of the visual analysis model more extensive.

[0158] In another possible implementation of the second aspect, when there are repeated segments in multiple sample video segments, the predicted label is output, including:

[0159] After the electronic device acquires the frame-level action prediction labels and frame-level shot prediction labels of all sample video segments, the frame-level action prediction labels and frame-level shot prediction labels of the duplicate image frames are deduplicated to obtain the deduplicated frame-level action prediction labels and frame-level shot prediction labels of all sample video segments.

[0160] In this application, when there are duplicate segments in multiple sample video segments, the electronic device can perform deduplication processing on the frame-level behavior prediction labels and frame-level shot prediction labels of the sample video segments. The behavior segmentation results (prediction results) and shot segmentation results (shot prediction results) obtained based on the frame-level behavior prediction labels and frame-level shot prediction labels of the sample video segments are more accurate, and the model training effect based on the prediction results is better.

[0161] Thirdly, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in any of the first aspects or the steps of the method described in any of the second aspects.

[0162] Fourthly, a computer-readable storage medium is provided that stores instructions which, when executed by a processor, implement the steps of the method described in any one of the first aspects or the method described in any one of the second aspects.

[0163] Fifthly, a computer program product comprising instructions is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the method described in any one of the first aspects or the steps of the method described in any one of the second aspects.

[0164] In a sixth aspect, embodiments of this application provide a chip including a processor, the processor being configured to invoke a computer program in memory to perform steps of the method as described in any of the first aspects or the method as described in any of the second aspects.

[0165] It is understood that the beneficial effects achieved by the electronic device described in the third aspect, the computer-readable storage medium described in the fourth aspect, the computer program product described in the fifth aspect, and the chip described in the sixth aspect can be referred to the beneficial effects of the first aspect and any possible design or the second aspect and any possible design, which will not be repeated here. Attached Figure Description

[0166] Figure 1 A schematic diagram illustrating lens segmentation based on pixels, provided as an embodiment of this application;

[0167] Figure 2 A schematic diagram illustrating scene segmentation based on speech semantics, provided as an embodiment of this application;

[0168] Figure 3 A schematic diagram illustrating a multi-granularity video partitioning method provided in this application embodiment;

[0169] Figure 4A schematic diagram illustrating a multi-granularity partitioning method provided in an embodiment of this application;

[0170] Figure 5 A schematic diagram of the interface for displaying detailed video information in a gallery application, provided as an embodiment of this application;

[0171] Figure 6 A schematic diagram of a video content positioning scenario provided for an embodiment of this application;

[0172] Figure 7 A schematic diagram of a smart image processing scene provided for an embodiment of this application;

[0173] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0174] Figure 9 A software structure block diagram of an electronic device provided in an embodiment of this application;

[0175] Figure 10 A flowchart illustrating a video processing method provided in an embodiment of this application;

[0176] Figure 11 A schematic diagram of the network structure of a visual analysis model provided in an embodiment of this application;

[0177] Figure 12 This is a schematic diagram illustrating how to obtain recognition results based on time-series processing, as provided in an embodiment of this application.

[0178] Figure 13 This is a schematic diagram illustrating how to crop an input video according to an embodiment of this application.

[0179] Figure 14 A flowchart illustrating another video processing method provided in an embodiment of this application;

[0180] Figure 15 A schematic diagram illustrating the input and output of a large language model provided in an embodiment of this application;

[0181] Figure 16 A possible structural schematic diagram of the electronic device provided in the embodiments of this application;

[0182] Figure 17 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation

[0183] In the description of the embodiments of this application, the terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to be a limitation of this application. As used in the specification and appended claims of this application, the singular expressions "a," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" refer to one or more (including two). The term "and / or" is used to describe the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can indicate: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0184] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes direct connections and indirect connections, unless otherwise stated. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated.

[0185] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0186] Video is gradually becoming the primary medium for people to obtain information, and video-based intelligent interaction scenarios are becoming increasingly common. These scenarios can include video content location, generation of highlight moments, and intelligent video creation, among others.

[0187] These video-based intelligent interaction scenarios require the identification and analysis of video content to obtain corresponding output results. For example, in the application scenario of video content localization, the electronic device specifically identifies the content of candidate videos and obtains the corresponding content tags; when the electronic device receives a video content search operation, it searches based on the content tags of each candidate video to locate the corresponding video segment that matches the video content search operation. For example, in the application scenario of generating highlights, the electronic device specifically receives an operation to generate highlights, identifies highlights (image frames or video segments) in candidate videos, and generates a highlight segment (video segment) synthesized from the highlights. For example, in the application scenario of intelligent video production, when the electronic device receives an operation to generate intelligent video production, it identifies the content of candidate videos, obtains video segments that meet the requirements from the candidate videos, and generates a video segment synthesized from these video segments.

[0188] As video content becomes increasingly diverse, users have more refined interactive needs. Currently, some classic algorithms for intelligent video interaction scenarios offer relatively limited granularity in video processing, failing to meet user interaction requirements. To illustrate with specific application scenarios, in one particular scenario—video content location—electronic devices can only search at the granular level of video theme or scene. For example, searching for videos about basketball games or videos filmed at the beach. They cannot provide the corresponding search capabilities for finer-grained or other dimensions. For instance, searching for a specific action of a subject in the video, such as searching for a basketball shot, or searching for a specific shot within a scene, such as searching for a panoramic shot. In another scenario—specifically, in intelligent video creation—electronic devices can combine similar scene clips into a single clip. For example, combining videos filmed at the beach into a single video. However, they cannot provide the corresponding capabilities for more fine-grained or other dimensions of intelligent video creation. For example, video clips featuring the same behavior can be intelligently processed into a single video, such as video clips of someone running on the beach or video clips of the main character eating.

[0189] For example, Figure 1 A schematic diagram of lens segmentation based on pixels is provided. For example... Figure 1 As shown, traditional methods of segmenting lenses using a single granularity often rely solely on pixel differences for lens segmentation. Figure 1 The image uses rectangular boxes to represent the positions of lens changes based on pixel division. This division method is rather coarse, and the results obtained are not very accurate. Figure 2A schematic diagram of scene segmentation based on speech semantics is given. For example... Figure 2 As shown, traditional single-granularity scene segmentation methods are often based on semantic scene recognition using Automatic Speech Recognition (ASR) technology, identifying multiple different scenes (including scene A, scene B, scene C, scene D, and scene E). This scene recognition method relies on the accuracy of speech semantic detection, and the resulting scene segmentation results are not ideal.

[0190] As users become increasingly sensitive to the granularity of video content, it is essential to perform more comprehensive and detailed analysis of the content and information contained in videos in order to provide users with a better intelligent video interaction experience, such as recommendation, editing, and search. This makes multi-granular segmentation of videos crucial.

[0191] Considering users' perception needs for content granularity in real-world usage scenarios and the separable granularity of video content, this application summarizes several video segmentation granularities in its embodiments. For example, video segmentation granularity can include scene granularity, shot granularity, and behavior granularity. Further, video segmentation granularity can include video shooting scene (shooting space and / or shooting time), video shooting tone, video shooting shot, the subject in the video, the subject's behavior, the subject's emotion, the subject's expression, and the number of subjects in the video. Among these, the segmentation granularity of the subject's behavior, emotion, expression, and number is relatively fine. In scenarios where granularity requirements are not high, these granularities can all be considered as behavior granularity to reduce the computational power required for video processing. This application mainly describes video processing methods for scene granularity, shot granularity, and behavior granularity. It is understood that the video processing algorithm can be adjusted according to the actual scenario to achieve video segment recognition at other granularities.

[0192] The video processing method provided in this application embodiment involves an electronic device inputting the video to be processed into a visual analysis model for visual analysis, obtaining behavior segmentation results and shot segmentation results for the video. The behavior segmentation results represent the segmentation of different behaviors of the subject in the video, while the shot segmentation results represent the segmentation of different shots in the video. Further, based on the content descriptions of the video segments corresponding to different shots included in the shot segmentation results, the electronic device obtains scene segmentation results for the video. The scene segmentation results represent the segmentation of different scenes in the video. The electronic device can perform multi-granularity recognition on the video to be processed, obtaining scene-granularity recognition results, shot-granularity recognition results, and behavior-granularity recognition results. This provides a more detailed understanding and segmentation of video content, offering more detailed video content information to support video understanding and analysis in various application scenarios. For example, it can enable the search and location of more detailed parts of the video content, and can process (editing, compositing, etc.) content with the same behavior or shot appearing in the video content. It also improves the user experience of using the video processing functions provided by the electronic device.

[0193] Among them, for multi-granularity recognition processing of videos, based on the characteristics of the video, the multi-granularity can include the shooting scene, shooting lens, and the behavior of the main body in the video. Figure 3 A hierarchical diagram illustrating multi-granularity video partitioning is provided. For example... Figure 3As shown, for a video, the granularity, from coarse to fine, can be scene granularity, shot granularity, and action granularity. That is, scene changes are always accompanied by shot changes, and a shooting scene may include one or more shots. A single shot may contain one or more different behaviors of the subject in the video, or multiple shots may contain the same behavior of the subject in the video. The subject can be a person or an animal. When the subject is a person, the behavior can include different behaviors in different scenarios such as walking, eating, dancing, running, shooting, and jumping. When the subject is an animal, the behavior can include different behaviors in different scenarios such as walking, eating, hunting, drinking, and running. For a video, segments can be divided at scene granularity, shot granularity, and action granularity to support data use in subsequent video understanding and analysis scenarios. It is understandable that the granularity of video segmentation can also include other granularities. For example, other granularities can include the subject's emotions, such as anger, sadness, happiness, excitement, depression, etc., and video segments can be identified and segmented according to different emotions. Other granularities can also include the subject's facial expressions, such as smiling, laughing, crying, etc., and video segments can be identified and segmented according to different facial expressions. For example, other granularities can also include the number of subjects, such as segmenting video segments according to whether the subjects include a single person or multiple people. For example, other granularities can also include the video's shooting color tone, such as segmenting video segments according to different video shooting color tones. This embodiment uses relatively classic scenes, shots, and behaviors as examples for illustration. Corresponding learning and recognition can also be performed for other granularities to obtain corresponding recognition results.

[0194] Specifically, electronic devices can obtain behavior segmentation results, shot segmentation results, and scene segmentation results of the video to be processed using the video processing method provided in this application embodiment. For example, Figure 4 A schematic diagram of multi-granularity video partitioning is provided. For example... Figure 4 Video 1 shown in (a) is a 4-minute and 5-second elephant documentary. At approximately 58 seconds, Video 1 records an aerial view of a tornado in the elephants' environment; at approximately 1 minute and 15 seconds, Video 1 records a frontal view of a tornado in the elephants' environment; at approximately 2 minutes and 2 seconds, Video 1 records footage of elephants walking in a group; at approximately 3 minutes and 22 seconds, Video 1 records footage of cracked soil, and so on.

[0195] The electronic device uses the video processing method provided in the embodiments of this application to perform multi-granularity recognition on video 1, and obtains the scene segmentation result, shot segmentation result and behavior segmentation result of video 1.

[0196] The scene segmentation result represents the segmentation of different scenes in the video to be processed. The scene segmentation result can include scene labels and timestamps of the video segments corresponding to each scene label. For example, based on scene-level granularity recognition, the scene segmentation result of video 1 can be represented as:

[0197] [00:00:00.000-00:01:48.960s, Scene 1]

[0198] 00:01:58.000-00:03:00.440s, Scene 2

[0199] 00:03:30.120-00:04:04.960s, Scene 3.

[0200] According to the playback sequence of Video 1, there are two scene changes in Video 1: from Scene 1 to Scene 2, and from Scene 2 to Scene 3.

[0201] The video segment from the beginning of Video 1 to approximately 1 minute and 48 seconds corresponds to Scene 1. Scene 1 depicts a tornado occurring in the elephant's habitat. Scene 1 may include footage of the tornado. For example... Figure 4 As shown in (b), a schematic diagram of two image frames in the video clip corresponding to Scene 1 is given. These include image frame 1, which is viewed from above the tornado, and image frame 2, which is viewed from the front of the tornado.

[0202] The video clip from approximately 1 minute 58 seconds to 3 minutes in Video 1 corresponds to Scene 2. Scene 2 is a scene featuring elephants as the main subject. Scene 2 can include footage with elephants as the main subject; for example, scenes showing elephants walking, eating, or drinking water all fall under Scene 2. Figure 4 As shown in (b), a schematic diagram of three image frames in the video clip corresponding to Scene 2 is given. These include image frame 3, which shows elephants walking in a herd, image frame 4, which shows elephants walking alone, and image frame 5, which shows elephants eating.

[0203] The video clip from approximately 3 minutes and 20 seconds to approximately 4 minutes and 4 seconds (the end) of Video 1 corresponds to Scene 3. Scene 3 is a shot of the earth. Scene 3 may include shots of the earth, such as... Figure 4 As shown in (b), a schematic diagram of an image frame from the video clip corresponding to Scene 3 is given. This includes image frame 6, which shows the cracked soil.

[0204] The shot segmentation result represents the segmentation of video segments generated based on different shots during shot transitions in the video to be processed. The shot segmentation result can include shot labels and the timestamps of the corresponding video segments for each shot label. For example, based on shot-level granularity recognition, the shot segmentation result for video 1 can be represented as:

[0205] [00:00:00.000-00:01:00.960s, Shot 1]

[0206] 00:01:01.000-00:01:48.960s, Shot 2.

[0207] 00:01:58.000-00:02:05.400s, Shot 3.

[0208] 00:02:06.000-00:02:33.140s, Shot 4.

[0209] 00:02:35.750-00:03:00.440s, Shot 5.

[0210] 00:03:30.120-00:04:04.960s, Shot 6.

[0211] According to the playback sequence of Video 1, there were five camera cuts in Video 1.

[0212] Specifically, the video clip from the beginning of video 1 to about one minute in is a top-down shot of scene one (shot one). For example... Figure 4 As shown in (c), a schematic diagram of an image frame from a video clip corresponding to shot one is given. This includes image frame 1, which is a top-down view of the tornado.

[0213] Around 1 minute into the video, the camera angle changes from a top-down view of scene one (Shot 1) to a frontal view of scene one (Shot 2). In other words, the segment from approximately 1 minute to approximately 1 minute and 48 seconds into the video corresponds to Shot 2. Figure 4 As shown in (c), a schematic diagram of an image frame from the video clip corresponding to shot two is given. This includes image frame 2, which is viewed from the front of the tornado.

[0214] Around 1 minute and 58 seconds into Video 1, the camera angle changes from a direct shot of Scene 1 (Shot 2) to a panoramic shot of Scene 2 (Shot 3). In other words, the video segment from approximately 1 minute and 58 seconds to approximately 2 minutes into Video 1 corresponds to the panoramic shot of Scene 2. Figure 4As shown in (c), a schematic diagram of an image frame from the video clip corresponding to shot three is given. This includes image frame 3, which shows a herd of elephants walking.

[0215] Around 2 minutes and 6 seconds into Video 1, the camera switches from a panoramic shot of Scene 2 (Shot 3) to a close-up shot of Scene 2 (Shot 4). Within the same scene, Shot 3 and Shot 4 capture the same content: an elephant walking. This is a zoom-in shot. In other words, the video segment from approximately 2 minutes and 6 seconds to approximately 2 minutes and 33 seconds into Video 1 corresponds to a close-up shot of Scene 2. Figure 4 As shown in (c), a schematic diagram of an image frame from the video clip corresponding to shot four is given. This includes image frame 4, which shows an elephant walking alone.

[0216] Around 2 minutes and 35 seconds into Video 1, a camera angle changes from a close-up shot of Scene 2 (Shot 4) to another close-up shot of Scene 2 (Shot 5). Within the same scene, Shot 4 and Shot 5 capture different content, resulting in a camera angle change. In other words, the video segment from approximately 2 minutes and 35 seconds to approximately 3 minutes into Video 1 corresponds to another close-up shot of Scene 2. Figure 4 As shown in (c), a schematic diagram of an image frame from the video clip corresponding to shot five is given. This includes image frame 5, which depicts an elephant eating.

[0217] Around 3 minutes and 30 seconds into Video 1, the camera shifts from shot five of Scene 2 to shot six of Scene 3. In other words, the video segment from approximately 3 minutes and 30 seconds to approximately 4 minutes and 4 seconds (the end) corresponds to shot six of Scene 3. Figure 4 As shown in (c), a schematic diagram of an image frame from the video clip corresponding to shot six is ​​given. This includes image frame 6, which shows the cracked soil.

[0218] The behavior segmentation result represents the segmentation of the video to be processed based on the behaviors of different subjects. The behavior segmentation result can include behavior labels and timestamps of the video segments corresponding to each behavior label.

[0219] For example, based on behavior-level recognition, the behavior segmentation result of video 1 can be represented as:

[0220] [00:01:58.000-00:02:33.140s, walking (action1), 00:02:35.750-00:03:00.440s, eating (action2)].

[0221] According to the playback sequence of Video 1, two actions involving an elephant (the main subject) occur in Video 1. These include the video segment corresponding to the elephant walking (action 1) and the video segment corresponding to the elephant eating (action 2).

[0222] Specifically, the behavior corresponding to the video segment from approximately 1 minute 58 seconds to approximately 2 minutes 33 seconds in Video 1 is elephant walking (behavior one). For example... Figure 4 As shown in (d), a schematic diagram of two image frames in a video clip corresponding to behavior one is given. These include image frame 3, showing a herd of elephants walking, and image frame 4, showing a single elephant walking alone.

[0223] The video clip from approximately 2 minutes 35 seconds to 3 minutes in video 1 depicts an elephant feeding. For example... Figure 4 As shown in (d), a schematic diagram of an image frame from the video clip corresponding to behavior two is given. This includes image frame 5, which shows an elephant eating.

[0224] In some scenarios, when the electronic device is equipped with a camera, it can process the video after recording it, obtaining scene segmentation, shot segmentation, and behavior segmentation results. Alternatively, after receiving and storing a video, the electronic device can process the stored video to obtain scene segmentation, shot segmentation, and behavior segmentation results.

[0225] In some embodiments, after acquiring the scene segmentation results, shot segmentation results, and behavior segmentation results of a video, the electronic device can store these results in a designated storage space. For example, the scene segmentation results, shot segmentation results, and behavior segmentation results can be stored in a video tag database. The electronic device can provide various user-oriented video processing application scenarios, using the video tags of each video in the video tag database for corresponding data support.

[0226] In one embodiment, a mobile phone is used as an example. After the mobile phone captures a video and stores it in a gallery application, or after storing a video in a gallery application, the mobile phone can process the video to obtain multi-granularity recognition results. The mobile phone can store the multi-granularity recognition results of the video in a video tag database. The video tag database can be a local database on the mobile phone or a cloud database. When the mobile phone needs to obtain the multi-granularity recognition results of the video, it can communicate with the cloud video tag database to obtain the corresponding tag data. In one example, the mobile phone can also obtain one or more tags from the video's scene tags, shot tags, and behavior tags based on the multi-granularity recognition results, and add them to the video details in the gallery application.

[0227] For example, Figure 5 This document provides a schematic diagram illustrating a scenario for viewing videos corresponding to tags in a gallery application. (Reference) Figure 5 As shown in (a) of the diagram, a schematic diagram of a display interface 10 for video details of video 1 is provided. The display interface 10 includes a video preview area and a video details area located below the video preview area. The video details include a display area 11 for information such as the video's storage date, video name, and storage path. Display area 11 also includes an edit button, allowing the electronic device to edit the video name or storage path in response to user interaction with the edit button. The video details also include a display area 12 for information such as video size, resolution, duration, video compression encoding, audio encoding, and video frame rate. Figure 5 As shown in (a), Video 1 has a file size of 1.69GB, a resolution of 1920*1080dpi, a duration of 4 minutes and 5 seconds, uses H.264 compression encoding, employs Advanced Audio Coding (ACC) audio encoding, and has a frame rate of 48fps. That is, Video 1 displays at a frame rate of 48 frames per second. Video details include display area 13 corresponding to the video tags. Display area 13 can include one or more of the video's scene tags, shot tags, and action tags. For example, such as... Figure 5 As shown in (a), display area 13 includes scene tags for video 1 such as "#ElephantLivingEnvironment", "#ElephantMainScene", and "#LandScene"; and behavior tags for video 1 such as "#ElephantWalking" and "#ElephantFeeding". In addition, display area 13 also includes a button for manually adding video tags, allowing the electronic device to receive user input to add video tag information.

[0228] In some embodiments, the electronic device displays one or more video tags in display area 13. The electronic device can receive a user's selection operation on a tag, and in response to the selection operation, display and play the video segment corresponding to that tag. Figure 5 As shown in (b), the electronic device can receive a user's selection operation on "Elephant Feeding" and, in response to the selection operation, play the target segment corresponding to "Elephant Feeding" in the playback interface 14.

[0229] In some embodiments, reference Figure 5As shown in (c), a schematic diagram of a playback interface 15 for video 1 is provided. The display area 16 of the playback interface 15 may include a playback progress bar for video 1. The playback progress bar for video 1 may display identifiers corresponding to one or more tags included in video 1, such as... Figure 5 (c) illustrates the black dot in the playback progress bar. The electronic device can respond to the user's sliding of the play button within the progress bar. If the play button is detected sliding to a specific label, the corresponding label is displayed. For example... Figure 5 As shown in (c), when the play button is detected to have been slid to a certain position, the electronic device displays the "Elephant Feeding" label corresponding to that position on the interface. Furthermore, if the electronic device detects that the user has remained on that label for a certain duration, such as 1 second, it can determine that the user needs to view the video clip corresponding to that label. The electronic device can then display the video (or video clip) corresponding to the "Elephant Feeding" label on the playback interface 17 and play the video clip corresponding to the "Elephant Feeding" label, such as... Figure 5 As shown in (d).

[0230] In some embodiments, reference Figure 5 As shown in (e), a schematic diagram of a camera interface 18 for a photo gallery application is provided. The electronic device can also categorize videos (or video clips) corresponding to tags that appear most frequently into an album based on the tags of each video. For example... Figure 5 As shown in (e), the photo album includes an album labeled "Elephant Feeding" and an album labeled "Basketball Shooting." In response to a user selecting the "Elephant Feeding" album, the electronic device can select a video or video clip from that album for playback. Figure 5 As shown in (f), the electronic device displays a playback interface 17 for the video corresponding to "Elephant Feeding" and plays the video clip corresponding to the "Elephant Feeding" tag. Alternatively, the electronic device may respond to the user's selection of the "Elephant Feeding" album by displaying a details interface for the "Elephant Feeding" album, wherein the details interface includes thumbnails of the videos (or video clips) included in the "Elephant Feeding" album. The electronic device may respond to the user's selection of a particular thumbnail by playing the corresponding video clip.

[0231] In some embodiments, a video may include a large number of video tags, and it may be impossible to display all of them in the video details; or, it may be impossible to mark all video tags with a playback progress bar; or, it may be impossible to categorize all video tags into albums. In one implementation, considering that frequently occurring scenes / shots / actions in the video are important, the electronic device can acquire a limited number of high-frequency video tags and display them in the video details / mark them in the playback progress bar / build a categorized album. For example, if "elephant eating" appears frequently, then "elephant eating" will be added to the video details for display / marked in the playback progress bar / built into a categorized album. Alternatively, in another implementation, considering that highlights in a video are generally rare segments with low frequency, the electronic device can also acquire a limited number of low-frequency video tags and display them in the video details / mark them in the playback progress bar / build a categorized album. In yet another implementation, if the frequency of the video tags in the video is relatively even, the electronic device can randomly select a limited number of video tags to display in the video details / mark them in the playback progress bar / build a categorized album. Alternatively, electronic devices can display video tags corresponding to the highlights of a video in the video details / mark them in the playback progress bar / create categorized albums. The filtering and display of video tags in the video details can be determined according to actual product needs; this embodiment does not limit the filtering rules for video tags in the video details.

[0232] The following examples illustrate the data support role of the scene segmentation results, shot segmentation results, and behavior segmentation results obtained by the video processing method provided in this application in subsequent video processing applications.

[0233] In a typical use case, for example, an electronic device can provide video content location functionality. When the electronic device is a mobile phone, it can receive video content location requests initiated by the user through the graphical interface and perform a video content search. Alternatively, the electronic device can receive video content location requests initiated by the user from the negative first page of the user's desktop and perform a video content search. This embodiment does not limit how the user initiates the video content location request.

[0234] Let's take the example of an electronic device receiving a video content location operation initiated by a user on a graphical interface and then searching for and locating the video content. Figure 6 A scenario diagram illustrating video content localization is provided. Among them, Figure 6(a) shows a schematic diagram of the main interface 20 of a gallery application. The main interface 20 of the gallery application includes a search bar 21. The electronic device receives a video content location operation initiated by the user in the search bar 21. For example, the user enters "find videos of elephants eating" in the search bar. The electronic device performs content recognition on the input content, determines that the subject of the video is "elephant", then the corresponding video scene is a scene with elephants as the main subject (i.e., scene two), and the corresponding behavior tag is "eating". The electronic device queries the video tag database for video tag data that matches the scene tag "scene two" and the behavior tag "eating". If there is video tag data that matches it, then based on the timestamp of the video segment corresponding to the video tag data, the corresponding video segment is obtained as the target video segment. For example, such as Figure 6 As shown in (b), the scene tag matches the segment of video 1 [00:01:58.000-00:03:00.440s, Scene 2], and further matches the behavior tag matches the segment of video 1 [00:02:35.750-00:03:00.440s, Eating].

[0235] In one implementation, after locating the target video segment, the electronic device can first locate the position of the complete video corresponding to that target video segment in the image library, such as... Figure 6 As shown in (c). After the electronic device locates the complete video, it locates the start position of the timestamp of the target video segment (e.g., ...). Figure 6 As shown in (d) at 02:35), the target segment is played in the playback interface 22 of video 1, such as Figure 6 As shown in (d).

[0236] Alternatively, in one implementation, if multiple matching video tag data exist, that is, multiple matching target video segments, the electronic device can copy the original videos corresponding to the multiple target video segments into a photo album, forming a collection of videos of elephants eating. For one of the videos—for example, this video could be the earliest video stored in the gallery application in chronological order, or the latest video stored in the image application in chronological order—the electronic device can locate and play the video segment of the elephant eating within that video (the target video segment).

[0237] The video processing method provided in this embodiment can refine the video content positioning to the level of the behavior of the main body in the video, thereby improving the user experience of using the video content positioning function.

[0238] In another typical use case, for example, an electronic device can provide intelligent video creation / one-click video creation functions. When the electronic device is a mobile phone, it can receive intelligent video creation operations initiated by the user on the image interface and perform intelligent video creation processing. Alternatively, the electronic device can receive intelligent video creation operations initiated by the user on the negative one page of the desktop and perform intelligent video creation processing. This embodiment does not limit how the user initiates the video content location operation.

[0239] Let's take the example of an electronic device receiving a user's intelligent image processing operation from the image interface and performing intelligent image processing. Figure 7 A schematic diagram of a smart integrated scene is presented. Among them, Figure 7 (a) shows a schematic diagram of the main interface 20 of a gallery application. The main interface 20 of the gallery application includes a search bar 21. The electronic device receives a smart video creation operation initiated by the user in the search bar 21. For example, the user enters "create a collection of videos of blowing out candles" in the search bar. The electronic device performs content recognition on the input content, determines that the main behavior corresponding to the video is "blowing out candles", and the behavior tag corresponding to the video is "blowing out candles". The electronic device queries the video tag database for video tag data that matches the behavior tag "blowing out candles". If there is at least one video tag data that matches it, then based on the timestamp of the video segment corresponding to the video tag data, it obtains at least one target video segment, and performs a synthesis process on the at least one target video segment to generate a video segment containing multiple candle-blowing segments. For example, as Figure 7 As shown in (b), based on the behavior tag "blowing out candles", video segment 1 [00:01:10.000-00:02:10.000s, blowing out candles], video segment 2 [00:00:25.000-00:01:30.000s, blowing out candles], and video segment 3 [00:00:15.000-00:01:45.000s, blowing out candles] are matched. Further, based on the video segments matched with the behavior tag, a video collection of candle-blowing videos is generated. The electronic device generates a video collection with a duration equal to the sum of the durations of segment 1, segment 2, and segment 3. The total duration of this video collection is 210 seconds (3 minutes and 30 seconds).

[0240] In one implementation, while the electronic device is generating a 210-second compilation of videos of people blowing out candles, it can display a prompt message on the main interface 20 of the gallery application. This prompt message instructs the user to wait for the video compilation to be generated. Figure 7 As shown in (c), the prompt message could be "Generating a candle-blowing video for you, please wait...". In one implementation, after generating a 210-second collection of candle-blowing videos, the electronic device can store the video collection (video 9) in a gallery application, such as... Figure 7As shown in (d). In some embodiments, the electronic device can also play the video collection in the video collection playback interface 23 after storing the video collection, such as... Figure 7 As shown in (e). The video processing method provided in this embodiment can refine the video content down to the behavioral granularity of the main subject in the video, thereby improving the user experience of using the intelligent video editing function.

[0241] In one implementation, the electronic device can also add the original videos containing candle-blowing clips to a new album, forming a candle-blowing video collection. For example, videos 2, 5, and 6 can be added to a new album, which can be named "Blowing Out Candles." Alternatively, the electronic device can add the corresponding video clips for blowing out candles—that is, the candle-blowing clips in video 2, video 5, and video 6—to a new album, forming a candle-blowing video collection. This achieves behavior-level video classification processing.

[0242] The video processing method provided in this application can also be applied to other intelligent video interaction scenarios, such as scenarios related to obtaining video highlight moments; for example, intelligent segmentation / intelligent addition of transition effects based on scenes, shots, and subject behavior in the video; for example, video recommendation scenarios based on subject behavior in the video, and so on. The video processing method provided in this embodiment can obtain recognition results at the scene granularity, shot granularity, and behavior granularity of the video, and does not limit the specific application scenarios based on the recognition results.

[0243] The video processing method provided in this application can be applied to electronic devices with image processing capabilities. Electronic devices can also be referred to as terminals, user equipment (UE), mobile stations (MS), mobile terminals (MT), etc. Electronic devices can be mobile phones, smart TVs, wearable devices, tablets, computers with wireless transceiver capabilities, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, or wireless terminals in smart homes, etc. The embodiments of this application do not limit the specific technologies or device forms used in the electronic devices.

[0244] Figure 8 A schematic diagram of the structure of the electronic device 100 is shown.

[0245] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, sensor module 180, camera 193, display screen 194, etc.

[0246] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0247] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0248] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0249] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0250] In this embodiment, an algorithm module is deployed in the processor 110. The processor 110 can act as the execution body of the video processing method, used to obtain multi-granularity recognition results of the video to be processed. Furthermore, the processor 110 can also store the multi-granularity recognition results of the video to be processed in a video tag database. The processor 110 can also respond to user-triggered operations in other intelligent video interaction scenarios, performing corresponding data processing based on the multi-granularity recognition results of the video.

[0251] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0252] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0253] In this embodiment, the electronic device 100 can display the main interface of the gallery application on the display screen 194, such as... Figure 5 The display interface 10 shown in (a) is shown. After video processing, the electronic device 100 can also display the application effects of the video processing, such as in a video positioning scenario, like... Figure 5 (c) and Figure 5 As shown in (d), the electronic device 100 can display the video positioning results on the display screen 194. For example, in a smart video production / one-click video production scenario, such as... Figure 6 (c)- Figure 6 As shown in (d), the electronic device 100 can also display a video collection obtained through smart video creation / one-click video creation on the display screen 194. Or, as... Figure 5 As shown, the electronic device 100 can also display a video details display interface on the display screen 194.

[0254] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0255] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0256] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0257] In this embodiment, the electronic device 100 can capture video using the camera 193. After capturing the video, the electronic device can store it in a gallery application. After capturing the video, the electronic device can perform video processing to obtain multi-granularity recognition results. Furthermore, the electronic device can store the multi-granularity recognition results of the video in a designated storage space. For example, it can store the multi-granularity recognition results of the video in a local or cloud-based video tag database.

[0258] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0259] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0260] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0261] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0262] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0263] Electronic device 100 can implement audio functions through audio module 170, speaker, receiver, microphone, headphone jack, and application processor, such as video playback, music playback, and recording.

[0264] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0265] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses a layered architecture. Taking the system as an example, the software structure of electronic device 100 is illustrated.

[0266] Figure 9 This is a software structure block diagram of the electronic device 100 according to an embodiment of the present invention.

[0267] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, [the following is omitted as the text is incomplete and likely refers to a specific implementation or feature]. The system is divided into four layers, from top to bottom: Application layer, Framework (FWK), Hardware Abstraction Layer (HAL), Kernel layer, and Hardware layer.

[0268] The application layer can include a series of application packages. For example... Figure 7 As shown, the application layer may include applications with video processing capabilities, such as camera applications and gallery applications. In some embodiments, the application layer may also include other functional applications. For example, the application layer may also include applications such as calendar, call, map, Bluetooth, music, video, and SMS.

[0269] like Figure 9 As shown, the framework layer FWK provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer includes some predefined functions. For example, the application programming interface of the framework layer FWK may include interfaces corresponding to camera applications, or interfaces for video processing.

[0270] In some embodiments, the framework layer (FWK) may also include a content provider and a view system. The content provider stores and retrieves data, making this data accessible to applications. The data may include video, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc. The view system includes visual controls, such as controls for displaying text and controls for displaying images. The view system can be used to build applications. The display interface may consist of one or more views. For example, a display interface including a text notification icon may include a view for displaying text and a view for displaying images.

[0271] The Hardware Abstraction Layer (HAL) deploys HAL interfaces corresponding to upper-layer applications. These HAL interfaces enable connection and communication between upper-layer applications and kernel-level modules. For example... Figure 7 As shown, the Hardware Abstraction Layer (HAL) includes a camera HAL interface and a video processing HAL interface. Gallery applications can communicate with the video processing module deployed in the kernel layer through the corresponding application programming interface for video processing and the video processing HAL interface.

[0272] The kernel layer is the layer between hardware and software, and it houses video processing modules and drivers. In one implementation, when the video tag database is deployed locally on the electronic device, it can be deployed within the kernel layer. The kernel layer drivers can include drivers corresponding to the hardware. For example, if the hardware layer includes a display screen, the kernel layer drivers include display drivers; if the hardware layer includes a camera, the kernel layer drivers include camera drivers.

[0273] The kernel-level video processing module can serve as the execution entity for the video processing method provided in this application embodiment. (Combined with...) Figure 7 The provided electronic device architecture diagram will briefly illustrate the interaction flow of the video processing method in this application embodiment.

[0274] Taking video recording and processing as an example, when the camera application receives a user's command to record a video, it sends a recording command to the corresponding driver (e.g., the camera driver) through the application programming interface (API) of the framework layer and the camera interface of the hardware abstraction layer. The camera driver then controls the camera to record. When the camera application receives a user's command to stop recording, it still sends a stop recording command to the camera driver through the layered interfaces, and the camera driver controls the camera to stop recording. After recording stops, the camera driver stores the video captured by the camera to the gallery application through the layered interfaces.

[0275] When the gallery application detects a newly stored video, it sends video processing instructions to the video processing module through the application programming interface (API) at the framework layer and the video processing HAL interface at the hardware abstraction layer. The video processing module, deployed at the kernel layer, acquires the video to be processed, performs multi-granularity recognition on the video, and obtains the scene segmentation, shot segmentation, and behavior segmentation results. The video processing module then stores these results in the video tag database.

[0276] For example, in a subsequent application scenario of video processing, such as when a gallery application receives a user's video location request, it sends a video location instruction to the video processing module through interfaces provided by the framework layer and the hardware abstraction layer. This video location instruction carries the query content for video location. For example, the query content could be "search for videos of elephants eating". After receiving the video location instruction, the video processing module uses the query content carried in the video location instruction to query the video tag database for behavior tags matching "elephant eating" to determine the corresponding target video segment. The video processing module can then combine at least one target video segment into a single video segment and store it in the gallery application. Alternatively, the video processing module can return the timestamp of the original video corresponding to the target video segment to the gallery application and display the revelation location of the target video segment corresponding to the original video on the gallery application's display interface.

[0277] The above embodiments have described the video processing method from the perspectives of its applicable scenarios and the interaction of various modules in the execution device. The following describes the algorithm implementation of the video processing method provided in this application. Specifically, the following details how the video processing module / electronic device obtains the multi-granularity recognition results of the video to be processed.

[0278] In some embodiments, Figure 10 A flowchart illustrating a video processing method is provided. Taking an electronic device as the executing entity as an example, the video processing method includes:

[0279] S201, Electronic device acquires video to be processed.

[0280] In this embodiment, the video to be processed can be a video captured by an electronic device. For example, the video to be processed is a video captured and stored by the electronic device. Alternatively, the video to be processed can also be a video acquired by the electronic device. For example, a video acquired by the electronic device from a local database or other devices.

[0281] In this embodiment, the video type to be processed is not limited. The video to be processed can be a first-person perspective video or a third-person perspective video. The subject of the video to be processed is also not limited; it can be a video featuring humans or animals. The duration of the video to be processed is also not limited; it can be a short video or a full-length film. The size or resolution of the video to be processed is also not limited; it can be a low-resolution video, such as a smooth video (vertical resolution of 360P) or a standard-definition video (vertical resolution of 480P); it can also be a high-resolution video, such as a high-definition video (vertical resolution of 720P), a high-definition video (vertical resolution of 1080P), a 2K resolution video, a 4K resolution video, or an 8K resolution video. The video compression encoding and audio encoding of the video to be processed are also not limited. The video to be processed can be a video using the H.264 video compression coding standard, or a video using H.261, H.263, AV1, Advanced Video Coding (AVC), or other video compression coding standards. The video to be processed can also be a video using the AAC audio coding standard, or a video using the Pulse Code Modulation (PCM), Waveform Audio File Format (WAV), MPEG-1 Audio Layer 3 (MP3), or other audio coding standards.

[0282] S202, The electronic device performs visual analysis on the video to be processed to obtain the shot segmentation results and behavior segmentation results of the video to be processed.

[0283] In some embodiments, the electronic device may employ a preset visual analysis algorithm to obtain shot segmentation results and behavior segmentation results of the video to be processed.

[0284] In one implementation, the electronic device can employ an image content recognition algorithm to divide the video to be processed into multiple image frames, determining whether a camera movement has occurred based on the image content of adjacent image frames. For example, image frame 2 is the next image frame in time sequence adjacent to image frame 1. Image frame 1 contains a full-body shot of a person, while image frame 2 contains a shot of the upper body of a person. It can be determined that a camera movement has occurred at image frame 2, and the type of camera movement may be a zoom-in. The content of the five image frames following image frame 2 is highly similar, and no camera movement has occurred. Image frame 8 is the next image frame in time sequence adjacent to image frame 7. Image frame 8 contains a view of a room, but does not include the person appearing in image frame 1 or image frames 2-7. It can be determined that a camera movement has occurred at image frame 8, and the type of camera movement is a cut. The content of the five image frames following image frame 8 is highly similar, and no camera movement has occurred. Image frame 14 is the next image frame in time sequence adjacent to image frame 13, and image frame 15 is the next image frame in time sequence adjacent to image frame 14. Image frame 15 contains a view of a different room than that in image frame 13. The shooting scene in image frame 15 is completely different from that in image frame 13, and the camera angle has changed. Image frame 14 contains a view of one room from image frame 13 with a certain degree of transparency, and also contains a view of another room from image frame 15 with a certain degree of transparency. Therefore, the type of camera angle change that occurs between image frames 13 and 15 can be determined to be a gradual fade.

[0285] Electronic devices can label the corresponding lens type of an image frame based on the lens change type. For example, if image frame 2 undergoes a lens change relative to image frame 1, and the lens change type is zoom-in, then the lens label for image frame 2 would be 2. If image frame 1 and image frame 2 are different shots, the segment corresponding to image frame 1 is a video segment shot using lens 1. If the content of the five image frames following image frame 2 is highly similar and no lens change occurs, then the lens labels for image frames 3-7 can be labeled as those corresponding to no lens change; for example, the lens label for no lens change would be 0. Based on the lens labels of each image frame, it can also be determined that the video segment corresponding to image frames 2-7 is the video segment shot using lens 2 after the lens change. The start timestamp of the video segment using lens 2 corresponds to the timestamp of image frame 2, and the end timestamp corresponds to the timestamp of image frame 7.

[0286] Image frame 8 undergoes a shot change relative to image frame 7, and the type of shot change is a cut. For example, if the shot label corresponding to a cut is 1, then the shot label for image frame 8 is 1. Image frame 8 and image frame 7 represent different shots, and the segment corresponding to image frame 8 is a video segment shot based on shot 3. The content of the five image frames following image frame 8 is highly similar, and no shot change occurs. Therefore, the shot labels for image frames 9-13 can be marked as the labels corresponding to no shot change; for example, the shot label corresponding to no shot change is 0. Based on the shot labels of each image frame, it can also be determined that the video segment corresponding to image frames 8-13 is the video segment shot based on shot 3 after the shot change. The start timestamp of the video segment based on shot 3 corresponds to the timestamp of image frame 8, and the end timestamp corresponds to the timestamp of image frame 13.

[0287] The shooting scenes of image frame 15 and image frame 13 are completely different; the camera angle has changed, and the type of camera angle change is a fade. For example, the lens label corresponding to a fade is 3, and the intermediate frames of all image frames included in the fade can be labeled with lens label 3. For example, the lens label of image frame 14 is labeled as 3. Image frame 15 is the completed fade, and the lens label of image frame 14 has been marked as a fade. Therefore, the lens label of image frame 15 can be marked with the lens label corresponding to the lens without change. The video clip starting with the timestamp of image frame 14 is the video clip shot based on lens 4.

[0288] In the above embodiments, the specific values ​​set for the lens label are merely examples, and this embodiment does not limit them.

[0289] In one implementation, the electronic device can also employ an image content recognition algorithm to identify the subject in image frames. When an image frame is found to contain a subject, the algorithm analyzes the content of multiple consecutive image frames to determine the subject's behavior and labels the identified behavior. For example, if image frame 1 and images 2 through 7 are found to contain a person (subject), the algorithm identifies the subject's behavior in images 1 through 7. If it determines that the person in images 1 through 7 is eating, and "eating" is a behavior that the image content recognition algorithm can recognize, then the behavior label for images 1 through 7 is determined to be "eating." Conversely, if no subject is identified in images 8 through 15, then the behavior label for images 8 through 15 is determined to be "no behavior," and the behavior label for "no behavior" can be "NULL." If image frames 16-20 are identified as containing a subject, but the subject behavior appearing in image frames 16-20 is not identified, that is, the subject behavior appearing in image frames 16-20 is a behavior that the image content recognition algorithm cannot identify, in this case, the behavior label corresponding to image frames 16-20 can be set to a fixed value, indicating that there is a subject behavior but the specific behavior cannot be identified.

[0290] Due to the rich variety of camera angle changes and subject behaviors, the image content recognition algorithms described above may have limitations in their recognition capabilities. In other feasible approaches, electronic devices can employ visual analysis models with learning capabilities to recognize camera angle changes and subject behaviors in the video being processed. For example, the visual analysis model in this embodiment can be a neural network model.

[0291] For example, in one implementation, Figure 11 A schematic diagram of the network structure of a visual analysis model is given. For example... Figure 11 As shown, the input of this visual analysis model is the video to be processed, and the output can include behavior labels, shot labels, frame-level behavior labels, and frame-level shot labels of the video to be processed.

[0292] like Figure 11 As shown, the visual analysis model includes a feature extraction module and a learning interaction module. The learning interaction module further comprises an initial layer and an interaction layer. The initial layer includes a first-shot learning module, a first-frame learning module, and a first-action learning module. The interaction layer includes a second-shot learning module, a second-action learning module, a second-frame learning module, and a third-frame learning module. The visual analysis model also includes a multilayer perceptron (MLP) for outputting low-dimensional feature vectors.

[0293] Electronic devices input video (such as the video to be processed) into the feature extraction module of a visual analysis model to obtain feature vectors of different granularities from the video. This feature extraction module can be based on convolutional neural networks, recurrent neural networks, or transformer encoders. For example, the feature extraction module can be a backbone network obtained after pre-training for video classification, such as Video Masked Autoencoders (VideoMAE), multimodal video understanding models (InternVideo2), video vision models (ViViT), or 3D convolutional models. The shallow network based on the feature extraction module can output shallow features. These shallow features typically contain more spatial detail information and can represent pixel differences. For example, camera changes often cause changes in the content of the scene, and these changes will inevitably result in differences in pixel distribution and pixel values. The pixel differences represented by shallow features can effectively identify camera changes. In other words, such as... Figure 11 As shown, the shallow network of the feature extraction model outputs a shot feature vector (first shot feature vector) for shot recognition. The deep network of the feature extraction module outputs deep features. Deep features typically contain high-dimensional abstract speech information, which can assist in learning behavior labels and shot labels. Figure 11 As shown, the deep network through the feature extraction module can output behavioral feature vectors (first behavioral feature vectors) and frame feature vectors (first frame feature vectors).

[0294] In some other embodiments, for example, when the granularity of video segmentation also includes the video shooting tone, feature vectors of tone granularity can be learned based on shallow features; when the granularity of video segmentation also includes the subject's expression, emotion, or number of subjects, the process can be similar to the feature vector processing of behavior granularity, learning feature vectors of expression granularity based on deep features, learning feature vectors of emotion granularity based on deep features, learning feature vectors of subject number granularity based on deep features, and so on.

[0295] In the multi-granularity recognition process of the video to be processed, after obtaining the first shot feature vector, the first behavior feature vector, and the first frame feature vector, the corresponding third and fourth frame feature vectors can be obtained through multiple cross-attention mechanisms in the initial layer and the interaction layer. The third frame feature vector refers to the frame-level shot feature vector, and the fourth frame feature vector refers to the frame-level behavior feature vector. The third and fourth frame feature vectors are then subjected to dimensionality reduction processing by a multilayer perceptron (MLP) to obtain frame-level shot prediction labels and frame-level behavior prediction labels.

[0296] Frame-level shot prediction labels refer to the shot label prediction results for each frame, and frame-level action prediction labels refer to the action label prediction results for each frame. Further, based on the frame-level shot prediction labels, temporally adjacent identical shot prediction labels are merged and organized to obtain the shot segmentation results of the video to be processed; based on the frame-level action prediction labels, temporally adjacent identical action prediction labels are merged and organized to obtain the action segmentation results of the video to be processed.

[0297] In this embodiment, the learning module is used to output feature vectors with a wider field of view and richer features. The frame learning module, behavior learning module, and shot learning module can all be deep learning models (transformers), dilated convolutional models, etc. In one example, the frame learning module and shot learning module can use dilated convolutions, and the behavior learning module can use a transformer.

[0298] After the second behavior feature vector, second shot feature vector, and second frame feature vector output by the learning module are processed in the input value interaction layer, multiple cross-attention mechanisms are used to allow frame-level high-dimensional semantic information to interact with shot-level spatial detail information and behavior-level semantic information, respectively. This ensures that the two sub-branches of the frame branch focus on shot change information and behavior change information, respectively, thereby outputting frame-level behavior prediction labels and frame-level shot prediction labels. The frame-level behavior prediction labels and frame-level shot prediction labels learned through multiple cross-attention mechanisms exhibit high accuracy.

[0299] Simultaneously, the second shot learning module can output shot-level prediction labels (shot prediction labels), and the second behavior learning module can output behavior-level prediction labels (behavior prediction labels). Both shot prediction labels and behavior prediction labels can be used as auxiliary data to assist in determining the final shot segmentation results and behavior segmentation results using frame-level behavior prediction labels and frame-level shot prediction labels.

[0300] In this embodiment, as Figure 11 As shown, the visual analysis model also includes a temporal processing module. After obtaining frame-level behavior prediction labels and frame-level shot prediction labels, these are input into the temporal processing module. The temporal processing module can obtain the start and end times of frames with the same labels according to the temporal sequence, and generate timestamps for the video segments corresponding to the labels. This yields the behavior segmentation results and shot segmentation results. For example, Figure 12 A schematic diagram is provided illustrating a method for obtaining recognition results based on time-series processing. For example, Figure 12As shown in (a), the behavior labels for image frames 1-3 are all behavior label 1, and the behavior labels for image frames 4-9 are all behavior label 2. Therefore, we take the timestamp of image frame 1 as the start time of video segment 1 corresponding to behavior label 1, and the timestamp of image frame 3 as the end time of video segment 1 corresponding to behavior label 1; we take the timestamp of image frame 4 as the start time of video segment 2 corresponding to behavior label 2, and the timestamp of image frame 9 as the end time of video segment 2 corresponding to behavior label 2, thus obtaining the behavior segmentation result. Figure 12 As shown in (b), the shot labels for image frames 1-5 are "no shot change"; the shot label for image frame 6 is "zoom-in"; and the shot labels for image frames 7-9 are "no shot change". Therefore, the timestamp of image frame 1 is taken as the start time of video segment 3 corresponding to shot 1, and the timestamp of image frame 5 is taken as the end time of video segment 3 corresponding to shot 1; the timestamp of image frame 6, where a shot change occurs, is taken as the start time of video segment 4 corresponding to action shot 2, and the timestamp of image frame 9 is taken as the end time of video segment 4 corresponding to action shot 2, thus obtaining the shot segmentation result.

[0301] Understandably, the duration of the video input to the visual analysis model can be arbitrary. For example... Figure 11The visual analysis model shown may also include a sliding window processing module. This module is used to trim the input video to be processed into multiple video segments with a duration equal to the sliding window length. The visual analysis model processes each video segment. For example, a video analysis length (sliding window length t) is set for the visual analysis model. If the duration of the video to be processed exceeds the video analysis length, the video can be trimmed according to the sliding window length with a certain overlap rate. Each trimmed video segment with a duration equal to the video analysis length is then fed into the visual analysis model for processing to obtain the corresponding behavior segmentation results and shot segmentation results. The video to be processed can be represented as Bx(T / t)xHxW, where B (Batch) is the number of video segments calculated based on the actual duration T (Time) of the video to be processed and the sliding window length t; H (Height) is the height of the video to be processed, which can also represent the vertical resolution of the video frame; and W (Width) is the width of the video to be processed, which can also represent the horizontal resolution of the video frame. If the actual duration T of the video to be processed is not an integer multiple of the sliding window duration t, that is, t cannot be divided by T, then a duration supplementation operation can be performed on the video / sample video to be processed. The actual duration of the input video is preset to ensure that the duration of the input video is divisible by the sliding window length. When the duration of the input video is divisible by the sliding window length, the input video can be directly cropped into segments of the sliding window length with an overlap rate of 0. In another case, if the actual duration T of the video to be processed is not an integer multiple of the sliding window duration t, the input video can also be cropped into segments of the sliding window length according to a certain overlap rate. The overlap rate can be a percentage greater than 10% and less than 50%, for example, an overlap rate of 25%. For example, Figure 13 A schematic diagram illustrating the cropping process of the input video is provided. Figure 13 (a) provides an example of cropping five video segments when the duration of the input video is divisible by the sliding window length. The duration of the input video can be the original duration or the duration after padding. Figure 13 (b) provides an example of cropping the input video into 5 video segments when the input video length is not divisible by the sliding window length. There are duplicate segments among the 5 video segments.

[0302] Furthermore, after obtaining frame-level action prediction labels and frame-level shot prediction labels for all video segments of the video to be processed, duplicate image frames in the frame-level action prediction labels and frame-level shot prediction labels can be deduplicated. This allows for the acquisition of recognition results from the deduplicated frame-level action prediction labels and frame-level shot prediction labels. Alternatively, after obtaining action segmentation results and shot segmentation results based on the frame-level action prediction labels and frame-level shot prediction labels, duplicate segments in the action segmentation results and shot segmentation results are deduplicated according to the overlap rate to obtain the final action segmentation results and shot segmentation results corresponding to the input video. This embodiment does not limit the specific value of the overlap rate.

[0303] In this embodiment, the electronic device can segment the video to be processed based on the sliding window length and perform deduplication on the results, which allows the video processing method to be applied to the processing of videos of different durations, making the application scenarios of the video processing method more extensive.

[0304] The following section uses the training process of the visual analysis model to illustrate the interactions between the various modules within the model. It particularly focuses on how the interaction layer employs multiple cross-attention mechanisms to improve the accuracy of prediction results.

[0305] The accuracy of a video analysis model determines the accuracy of multi-granularity recognition results for videos. Therefore, this embodiment provides a training method for a visual analysis model to obtain a more accurate visual analysis model for multi-granularity recognition.

[0306] Combination Figure 11 The network structure of the given visual analysis model is intended to illustrate the model training process.

[0307] In some embodiments, the electronic device inputs an input video (such as a sample video) into a visual analysis model. The electronic device obtains the duration of the sample video. If the duration of the sample video exceeds the video analysis length T, it trims the sample video according to a preset overlap rate to obtain multiple sample video segments. These multiple sample video segments are sequentially fed into the visual analysis model for processing to obtain the recognition results for each segment. Simultaneously, the multiple sample video segments are aligned and reassembled according to the overlap rate to obtain the complete sample video.

[0308] Specifically, the sample video clips are input into the feature extraction module of the visual analysis model to obtain the first shot feature vector, the first line feature vector, and the first frame feature vector of the sample video clips.

[0309] The first shot feature vector is processed by the first shot learning module of the initial layer to obtain the second shot feature vector. The second shot feature vector has a larger field of view and more features compared to the first shot feature vector. The second row feature vector is processed by the first row learning module of the initial layer to obtain the second row feature vector. The second row feature vector has a higher field of view and more features compared to the first row feature vector. The first frame feature vector is processed by the first frame learning module of the initial layer to obtain the second frame feature vector. The second frame feature vector has a higher field of view and more features compared to the first frame feature vector.

[0310] After obtaining the feature vectors of the second shot, the second action, and the second frame, the interaction layer learns based on multiple cross-attention mechanisms to obtain more accurate feature vectors.

[0311] This process includes: using the feature vector of the second frame as the query vector, and the feature vector of the second shot as the key and value vectors, inputting them into the second shot learning module using an attention mechanism for learning, to obtain the feature vector of the third shot. The third shot feature vector is a feature vector obtained by the second shot learning module based on frame-level high-dimensional semantic information and shot-level spatial detail information. Shot-level shot prediction labels obtained based on the third shot feature vector are more accurate.

[0312] The second frame feature vector is used as the query vector, and the second row feature vector is used as the key and value vectors. An attention mechanism is employed to input these vectors into the second row learning module for learning, resulting in the third row feature vector. The third row feature vector is obtained by the second row learning module based on frame-level high-dimensional semantic information and behavior-level semantic information. Behavior-level behavior prediction labels obtained based on the third row feature vector are more accurate.

[0313] The feature vector of the third shot is used as the query vector, and the feature vector of the second frame is used as the key and value vectors. An attention mechanism is employed to input these features into the second frame learning module for learning, resulting in the third frame feature vector (also known as the frame-level shot feature vector). The third frame feature vector is obtained by the second frame learning module based on frame-level high-dimensional semantic information and shot-level spatial detail information. Frame-level shot prediction labels obtained based on the third frame feature vector are more accurate.

[0314] The third frame feature vector is used as the query vector, and the second frame feature vector is used as the key and value vectors. An attention mechanism is employed to input these vectors into the third frame learning module for learning, resulting in the fourth frame feature vector (also a frame-level behavioral feature vector). The fourth frame feature vector is obtained by the third frame learning module based on frame-level high-dimensional semantic information and behavioral-level semantic information. Frame-level behavioral prediction labels obtained based on the fourth frame feature vector are more accurate.

[0315] Multiple learning modules in the interaction layer utilize multiple cross-attention mechanisms to allow frame-level high-dimensional semantic information to interact with shot-level spatial detail information and behavior-level semantic information, respectively. This allows the second-frame learning module to focus on shot change information and the third-frame learning module to focus on behavior change information, thereby improving the prediction accuracy of frame-level shot prediction labels and frame-level behavior prediction labels.

[0316] After an electronic device obtains a lens prediction label based on a third lens feature vector, a behavior prediction label based on a third behavior feature vector, and outputs frame-level lens prediction labels and frame-level behavior prediction labels by processing the third frame feature vector through an MLP, a visual analysis model can be trained based on the losses calculated for each prediction label. These losses can include lens prediction losses calculated based on the lens prediction label and the standard lens label, and behavior prediction losses calculated based on the behavior prediction label and the standard behavior label. The loss functions for calculating the lens prediction loss and behavior prediction loss can be multi-class losses such as cross-entropy loss, or direct negative logarithmic loss. The losses can also include the sum of frame-level behavior prediction losses calculated from frame-level behavior prediction labels and standard behavior labels, and frame-level lens prediction losses calculated from frame-level lens prediction labels and standard lens labels. The loss functions for calculating the frame-level behavior prediction loss and frame-level lens prediction loss can be multi-class loss functions such as cross-entropy loss (one-hot or many-hot).

[0317] Electronic devices can train a visual analysis model based on the sum of one or more losses until the loss of the visual analysis model is less than a loss threshold, or until the visual analysis model converges, or until the required number of training iterations is reached, ultimately obtaining a well-trained visual analysis model. In some embodiments, electronic devices can also use commonly used loss functions such as cross-attention loss to train the visual analysis model and obtain a well-trained visual analysis model.

[0318] It's understandable that training a visual analysis model primarily involves training the outputs of each learning module in the interaction layer. Therefore, the feature extraction module can be frozen and not trained.

[0319] In some embodiments, some distinctions are made between shot prediction labels (shot labels), behavior prediction labels (behavior labels), frame-level shot prediction labels (frame-level shot labels), and frame-level behavior prediction labels (frame-level behavior labels).

[0320] Frame-level shot prediction labels refer to each image frame corresponding to a shot prediction label; while shot prediction labels are prediction labels at the shot granularity, and multiple image frames may correspond to one shot prediction label. Frame-level action prediction labels refer to each image frame corresponding to a action prediction label; action prediction labels are prediction labels at the action granularity, and multiple image frames may correspond to one action prediction label.

[0321] The behavior labels (behavior prediction labels) can include preset behavior labels, no behavior labels, and non-preset behavior labels. For example, preset behaviors can include different behaviors set for different domains, scenarios, and subjects. For example, behaviors for animals could include "eating," "walking," "jumping," "hunting," "sleeping," etc., while behaviors for humans in a basketball scenario could include "shooting," "passing," "dunking," "dribbling," etc. Different behaviors can be assigned different behavior labels, which can be represented by numeric or alphanumeric characters. For example, "eating" corresponds to behavior label 1, "behavior" to behavior label 2, "shooting" to behavior label 8, and so on. If no subject exists or no behavior exists, there is no behavior label. The no behavior label can be NULL or other specified characters. If a subject is identified and exhibits behavior, but this behavior is not a preset behavior, a non-preset behavior label is used. For example, a non-preset behavior label can be INF or other specified characters. This embodiment does not limit the specific identifier settings for behavior labels.

[0322] Shot tags (shot prediction tags) can include shot tags corresponding to various shot change types and tags for no shot change. The tag for no shot change can be 0 or other specified characters. Various shot change types can include cuts, fades, zoom-ins, zoom-outs, etc. Different shot change types can be assigned different shot tags, which can be represented by numeric or alphanumeric characters. For example, the shot tag for a cut frame is 1, the shot tag for a zoom-in frame is 2, the shot tag for all frames (or intermediate frames) involved in a fade is 3, and the shot tag for a zoom-out frame is 4. This embodiment does not limit the specific identifier settings for the shot tags.

[0323] S203, the electronic device obtains the scene segmentation result of the video to be processed based on the shot segmentation result of the video to be processed.

[0324] In this embodiment, the electronic device obtains the shot segmentation result of the video to be processed. The shot segmentation result includes the shot tags of the video to be processed and the timestamps of the video segments corresponding to the shot tags. For example, the shot segmentation result of a video segment (such as the video to be processed) can be represented as:

[0325] [00:00:00.000-00:01:00.960s, camera movement (1),

[0326] 00:01:01.000-00:01:48.150s, camera change (2),

[0327] 00:01:58.000-00:02:05.400s, camera movement (3)].

[0328] There were three camera changes throughout the video, involving several different types of camera transitions. These included a cut (camera cut, camera zoom-in, camera zoom-in, camera zoom-in, camera fade, camera fade, etc.).

[0329] Electronic devices can perform video segmentation based on the shot division results of the video to be processed. The video to be processed is divided into multiple video segments based on shot transition points (shot split points). The electronic device can use tools such as ffmpeg or open-source computer vision libraries (OpenCV) to perform video segmentation, obtaining multiple segmented video segments. This embodiment does not limit the specific algorithm or tool used for video segmentation.

[0330] After the electronic device acquires multiple video segments after the video to be processed has been cut into segments, it can obtain a video content description for each video segment.

[0331] In some embodiments, the electronic device can obtain the video content description (text) corresponding to each video segment based on some classic large language models or description algorithms. For example, the large language model can be an open-source video content description generation model such as the InternLM model, the InternVL model, the VideoLLaMA 2 model, and the InternVideo2 model. The description algorithm can be the description algorithm provided by the open-source real-time digital human dialogue system (VideoChat). This embodiment does not limit which specific algorithm is used to obtain the video content description.

[0332] For example, after segmenting the video based on the shot division results, a video content description is generated for each video segment. The resulting video content description can be represented as:

[0333] Shot 1, 00:00:00.000-00:01:00.960s: The video begins with an aerial view overlooking a cracked wasteland, showing a vast expanse of dry land. A tornado begins to form in this wasteland, stirring up dust and debris. The tornado gradually grows larger and stronger, eventually dominating the frame. The background remains consistent, with the desolate terrain contrasting sharply with the dynamics of the tornado.

[0334] Scene 2, 00:01:01.000-00:01:48.150s: A tornado forms and moves across a barren, open area. It rapidly grows from a small, concentrated column of dust into a spiral structure that increases significantly in both height and width. The background is a vast, flat landscape with sparse vegetation and a cloudy sky.

[0335] Scene 3, 00:01:58.000-00:02:05.400s: The video begins with a wide-angle shot of a wasteland, where a tornado is forming and rising. The sky is overcast, and as the tornado rises, text describing the scene appears on the screen. The camera then focuses on a skeleton on the ground, describing the location as being in the shadow of Mount Kilimanjaro.

[0336] The electronic device obtains the scene segmentation result for the video to be processed based on the video content description of each video segment (shot segment) of the video to be processed. In one feasible approach, the electronic device can perform semantic recognition of the scene based on the video content description to obtain the scene description corresponding to each video content description. Based on the scene description, scene recognition and segmentation are performed, and the start and end timestamps of scene changes are obtained to obtain the scene segmentation result of the video to be processed.

[0337] In some specific implementations, electronic devices can use a large language model to perform scene recognition on the video to be processed, and obtain the scene segmentation results of the video. The input to the large language model includes multiple video content descriptions of the video to be processed generated by the electronic device.

[0338] The large language model performs semantic recognition based on multiple video content descriptions, obtains the scenes contained in the video to be processed, merges temporally adjacent clips of the same scene, and outputs scene labels corresponding to each scene in the video to be processed, as well as the timestamps of the video clips corresponding to the scene labels. The large language model can be a common large language model, such as the Llama series, the Qwen2 series, etc. This embodiment does not limit which large language model is used for scene recognition.

[0339] It should be noted that, regardless of the large language model used, in order to better suit the processing requirements of this embodiment, a model prompt can be set. The prompt refers to the input text or instruction provided by the model to guide it in generating a specific type of response. In this embodiment, the large language model's prompt can be:

[0340] "You are an expert in analyzing video scene changes. Based on the video content description, please merge the clips of the same scene in time sequence and finally output the scene labels and time ranges of each scene."

[0341] The time range and description of the video content for each video segment are as follows:

[0342] In the audio text segment 1, 00:00:00.000-00:00:30.040s, on the Amboseli Plain, in the shadow of Mount Kilimanjaro, the seasonal rains that should have come for the past two years have not arrived.

[0343] Audio text segment 2, 00:31:00.000-00:01:00.960s: This year, they should have appeared long ago. This is the worst drought in half a century. Amboseli is usually a haven for elephants.

[0344] In audio text segment 3, 00:01:01.000-00:01:48.150s, it says, "These plains should have been green and covered with grass. Now, there's nothing but dust. This family is forced to constantly migrate in search of anything they can eat."

[0345] In audio text segment 4, 00:01:58.000-00:02:05.400s, it says, "The calf must keep up with the pace of migration, sometimes without even time to suckle. As the grass disappears, the elephant can only grab withered branches from the dust. Adult elephants may be able to survive on these, but they cannot sustain the calf for long."

[0346] The prompt for the large language model above is just an example. Based on actual data processing needs, the prompt for the large language model can be updated so that the large language model can output the required output data based on the input data.

[0347] In this embodiment, the scene segmentation result of the video to be processed is obtained through either a large language model or other scene recognition algorithms.

[0348] The result of this scene segmentation can be represented as:

[0349] [00:00:00.000-00:01:00.960s, Scene 1]

[0350] 00:01:01.000-00:01:48.150s, Scene 2

[0351] 00:01:58.000-00:02:05.400s, Scene 3.

[0352] Therefore, through steps S202 and S203, the behavior segmentation results, shot segmentation results, and behavior segmentation results of the video to be processed can be obtained, realizing the acquisition of multi-granularity recognition results. This provides effective data support for subsequent video processing.

[0353] In this embodiment, the electronic device inputs the video to be processed into a visual analysis model for visual analysis, obtaining behavior segmentation results and shot segmentation results. The behavior segmentation results represent the segmentation of different behaviors of the subject in the video, while the shot segmentation results represent the segmentation of different shots in the video. Further, based on the content descriptions of the video segments corresponding to different shots included in the shot segmentation results, the electronic device obtains the scene segmentation results corresponding to the video. The scene segmentation results represent the segmentation of different scenes in the video. The electronic device can perform multi-granularity recognition on the video to be processed, obtaining scene-granularity recognition results, shot-granularity recognition results, and behavior-granularity recognition results. It can simultaneously perform multi-granularity temporal segmentation of the video, including coarse-to-fine segmentation of scenes, shots, and behaviors, and can output type labels for shot changes (e.g., transition, fade, camera movement, NULL) and behavior labels (preset behavior labels, non-preset behavior labels NULL). This provides a more detailed understanding and segmentation of video content, offering more detailed video content information to support video understanding and analysis in various application scenarios later on. For example, it can enable the search and location of more detailed parts of video content, and can process (edit, composite, etc.) content that appears in the same action or in the same shot in the video content. At the same time, it can also improve the user experience of using the video processing functions provided by electronic devices.

[0354] In some embodiments, a video often includes both visuals and audio. For example, when the video is a documentary, the audio often includes introductory narration about the documentary and sound effects captured during filming. The narration introduces the visual content of the documentary and is closely related to it. Sound effects are often specific sounds produced in certain scenes within the documentary; therefore, sound effects are also coupled to some extent with the visual content. For example, if the visual content is a waterfall, the corresponding sound effects might include the sound of water flowing down from top to bottom and the sound of water hitting rocks. Similarly, if the visual content features elephants, the corresponding sound effects might include the sound of elephants roaring.

[0355] Based on this, in some embodiments, the electronic device can also analyze the audio of the video and use the effective information obtained from the audio to assist in scene recognition (or other granular recognition) of the video.

[0356] In one feasible way Figure 14 A flowchart illustrating another video processing method is provided. Again using an electronic device as the execution subject, the video processing method includes:

[0357] S301, Electronic device acquires video to be processed.

[0358] The embodiments provided in S201 above can be referred to, and will not be repeated in this embodiment.

[0359] S302, The electronic device performs visual analysis on the video to be processed to obtain the shot segmentation results and behavior segmentation results of the video to be processed.

[0360] The embodiments provided in S202 above can be referred to, and will not be repeated in this embodiment.

[0361] S303, the electronic device performs voice analysis on the video to be processed and obtains the voice text description of the video to be processed.

[0362] In this embodiment, the electronic device can employ a speech analysis algorithm to perform speech analysis on the video to be processed, thereby obtaining a speech-text description of the video. For example, the electronic device can use the AI ​​speech recognition model Whisper, the open-source speech recognition toolkit Kaldi, or the speech emotion recognition model emotion2vec to perform speech analysis on the video to be processed, thereby obtaining a speech-text description of the video. This embodiment does not limit the specific speech analysis algorithm used to perform speech analysis on the video to be processed.

[0363] In some embodiments, the audio-text description of the video to be processed may include multiple audio-text segments. These segments can be divided based on speech intervals, pauses, etc. For example, if blank audio in the video exceeds a certain duration, audio description segmentation can be performed. The certain duration can be a time interval of 1 second, 2 seconds, etc.

[0364] For example, a video clip from a documentary about elephants includes narration. After speech analysis, the resulting audio-text description of the narration may include multiple audio-text segments.

[0365] Audio text segment 1, 00:00:00.000-00:00:30.040s, “On the Amboseli plain, in the shadow of Mount Kilimanjaro, the seasonal rains that should have come for the past two years have not arrived.”

[0366] Audio text segment 2, 00:31:00.000-00:01:00.960s, “This year, they should have appeared long ago. This is the worst drought in half a century. Amboseli is usually a haven for elephants.”

[0367] Audio text segment 3, 00:01:01.000-00:01:48.150s, “These plains should have been green and covered with grass. Now, there is nothing but dust. This family is forced to migrate constantly in search of anything to eat.”

[0368] Audio text segment 4, 00:01:58.000-00:02:05.400s, “Young elephants must keep up with the pace of migration, sometimes without even time to suckle. As the grass disappears, elephants can only grab withered branches from the dust. Adult elephants may be able to survive on these, but they cannot support calves for long.”

[0369] In some embodiments, different national / regional language models can be pre-defined in the speech analysis model / speech analysis algorithm to generate speech text descriptions corresponding to different speech. These different national / regional languages ​​can include various major and minor languages ​​such as English, French, German, Japanese, Korean, Thai, and Arabic. For example, if the language of the human voice in the video to be processed is English, the generated initial speech text description will be in English. If the language of the human voice in the video to be processed is French, the generated initial speech text description will be in French. If the language of the human voice in the video to be processed is Korean, the generated initial speech text description will be in Korean. If the language of the human voice in the video to be processed is Japanese, the generated initial speech text description will be in Japanese.

[0370] In this embodiment, the electronic device combines the speech-text description processed by speech analysis with the video content description corresponding to the shot segmentation results of the video to be processed to obtain the scene segmentation results of the video to be processed. The language of the video content description is Chinese. Therefore, if the language of the initial speech-text description of the video to be processed obtained by the electronic device after speech analysis is not Chinese, the electronic device can also use a language translation model to translate the initial speech-text description of the video to be processed into a Chinese speech-text description in order to obtain the scene segmentation results of the video to be processed.

[0371] S304, The electronic device obtains the scene segmentation results of the video to be processed based on the shot segmentation results and the voice and text description.

[0372] In this embodiment, the electronic device obtains the shot segmentation result of the video to be processed. The shot segmentation result includes the shot tags of the video to be processed and the timestamps of the video segments corresponding to the shot tags. For example, the shot segmentation result of a video segment (such as the video to be processed) can be represented as:

[0373] [00:00:00.000-00:01:00.960s, camera movement (1),

[0374] 00:01:01.000-00:01:48.150s, camera change (2),

[0375] 00:01:58.000-00:02:05.400s, camera movement (3)].

[0376] There were three camera changes throughout the video, involving several different types of camera transitions. These included a cut (camera cut, camera zoom-in, camera zoom-in, camera zoom-in, camera fade, camera fade, etc.).

[0377] Electronic devices can perform video segmentation based on the shot division results of the video to be processed. The video to be processed is divided into multiple video segments based on shot transition points (shot split points). The electronic device can use tools such as ffmpeg or open-source computer vision libraries (OpenCV) to perform video segmentation, obtaining multiple segmented video segments. This embodiment does not limit the specific algorithm or tool used for video segmentation.

[0378] After the electronic device acquires multiple video segments after the video to be processed has been cut into segments, it can obtain a video content description for each video segment.

[0379] In some embodiments, the electronic device can obtain the video content description (text) corresponding to each video segment based on some classic large language models or description algorithms. For example, the large language model can be an open-source video content description generation model such as the InternLM model, the InternVL model, the VideoLLaMA 2 model, and the InternVideo2 model. The description algorithm can be the description algorithm provided by the open-source real-time digital human dialogue system (VideoChat). This embodiment does not limit which specific algorithm is used to obtain the video content description.

[0380] For example, after segmenting the video based on the shot division results, a video content description is generated for each video segment. The resulting video content description can be represented as:

[0381] Shot 1, 00:00:00.000-00:01:00.960s: The video begins with an aerial view overlooking a cracked wasteland, showing a vast expanse of dry land. A tornado begins to form in this wasteland, stirring up dust and debris. The tornado gradually grows larger and stronger, eventually dominating the frame. The background remains consistent, with the desolate terrain contrasting sharply with the dynamics of the tornado.

[0382] Scene 2, 00:01:01.000-00:01:48.150s: A tornado forms and moves across a barren, open area. It rapidly grows from a small, concentrated column of dust into a spiral structure that increases significantly in both height and width. The background is a vast, flat landscape with sparse vegetation and a cloudy sky.

[0383] Scene 3, 00:01:58.000-00:02:05.400s: The video begins with a wide-angle shot of a wasteland, where a tornado is forming and rising. The sky is overcast, and as the tornado rises, text describing the scene appears on the screen. The camera then focuses on a skeleton on the ground, describing the location as being in the shadow of Mount Kilimanjaro.

[0384] The electronic device obtains the scene segmentation result for the video to be processed based on the video content description of each video segment (shot segment) of the video to be processed, and the voice text description of the video to be processed obtained through voice analysis processing.

[0385] In one feasible approach, the electronic device can perform scene recognition on two-dimensional text descriptions to obtain the scene description corresponding to each audio text segment / each video content description. Based on the scene description, scene recognition and segmentation are performed, and the start and end timestamps of scene transitions are obtained to obtain the scene segmentation result of the video to be processed.

[0386] In one scenario, if a speech text segment has very low similarity to any description of the video content—meaning it doesn't match any video content description and there are no similar scenes / subjects—then this speech text segment is considered to have low relevance to the video content. To avoid interference from this type of speech text segment in scene recognition, it can be discarded and not used as data support for scene recognition.

[0387] For example, the audio-text segment might be a dialogue between unrelated characters in the video, which is largely irrelevant to the overall video content. For instance, in an elephant documentary, there might be a background audio dialogue between two people: Person A: "Have you eaten?" Person B: "No." Speech analysis would identify this dialogue and generate the corresponding audio-text segment "Have you eaten? No." However, because this dialogue has low relevance to the video content (elephant documentary), this audio-text segment can be discarded. Discarding audio-text segments without relevant information significantly reduces their impact on video scene recognition, thus making video scene recognition more accurate.

[0388] In some specific implementations, electronic devices can use a large language model to perform scene recognition on the video to be processed, obtaining scene segmentation results. The input to the large language model includes multiple speech-text segments of the video to be processed acquired by the electronic device, as well as multiple video content descriptions generated by the electronic device. (Reference) Figure 15 A schematic diagram illustrating the input and output of a large language model is provided. The large language model performs semantic recognition based on these two inputs (speech text segments and video content descriptions), obtains the scenes contained in the video to be processed, merges temporally adjacent clips of the same scene, and outputs scene labels corresponding to each scene in the video to be processed, as well as the timestamps of the video clips corresponding to the scene labels. The large language model can be any common large language model, such as the Llama series, the Qwen2 series, etc. This embodiment does not limit the choice of which large language model to use.

[0389] It should be noted that, regardless of the large language model used, in order to better suit the processing requirements of this embodiment, a model prompt can be set. The prompt refers to the input text or instruction provided by the model to guide it in generating a specific type of response. In this embodiment, the large language model's prompt can be:

[0390] "You are an expert in analyzing video scene changes. Based on the audio text segment and video content description, please merge the shot clips of the same scene in time sequence and finally output the scene label and time range of each scene."

[0391] The time range and description of the video content for each video segment are as follows:

[0392] In the audio text segment 1, 00:00:00.000-00:00:30.040s, on the Amboseli Plain, in the shadow of Mount Kilimanjaro, the seasonal rains that should have come for the past two years have not arrived.

[0393] Audio text segment 2, 00:31:00.000-00:01:00.960s: This year, they should have appeared long ago. This is the worst drought in half a century. Amboseli is usually a haven for elephants.

[0394] In audio text segment 3, 00:01:01.000-00:01:48.150s, it says, "These plains should have been green and covered with grass. Now, there's nothing but dust. This family is forced to constantly migrate in search of anything they can eat."

[0395] In audio text segment 4, 00:01:58.000-00:02:05.400s, the calf must keep up with the migration, sometimes without even time to suckle. As the grass disappears, the elephant can only grab withered branches from the dust. Adult elephants may be able to survive on these, but they cannot sustain the calf for long.

[0396] The time range and description of each audio text segment are as follows:

[0397] Shot 1, 00:00:00.000-00:01:00.960s: The video begins with an aerial view overlooking a cracked wasteland, showing a vast expanse of dry land. A tornado begins to form in this wasteland, stirring up dust and debris. The tornado gradually grows larger and stronger, eventually dominating the frame. The background remains consistent, with the desolate terrain contrasting sharply with the dynamics of the tornado.

[0398] Scene 2, 00:01:01.000-00:01:48.150s: A tornado forms and moves across a barren, open area. It rapidly grows from a small, concentrated column of dust into a spiral structure that increases significantly in both height and width. The background is a vast, flat landscape with sparse vegetation and a cloudy sky.

[0399] Scene 3, 00:01:58.000-00:02:05.400s: The video begins with a wide-angle shot of a wasteland, where a tornado is forming and rising. The sky is overcast, and as the tornado rises, text describing the scene appears on the screen. The camera then focuses on a skeleton on the ground, describing the location as being in the shadow of Mount Kilimanjaro.

[0400] The prompt for the large language model above is just an example. Based on actual data processing needs, the prompt for the large language model can be updated so that the large language model can output the required output data based on the input data.

[0401] In this embodiment, regardless of whether a large language model or other scene recognition algorithms are used, the scene segmentation result of the video to be processed is obtained. This scene segmentation result can be expressed as:

[0402] [00:00:00.000-00:01:00.960s, Scene 1]

[0403] 00:01:01.000-00:01:48.150s, Scene 2

[0404] 00:01:58.000-00:02:05.400s, Scene 3.

[0405] Therefore, through S302 and S304, the behavior segmentation results, shot segmentation results, and behavior segmentation results of the video to be processed can be obtained, realizing the acquisition of multi-granularity recognition results. This provides effective data support for subsequent video processing.

[0406] In this embodiment, the electronic device inputs the video to be processed into a visual analysis model for visual analysis, obtaining behavior segmentation results and shot segmentation results. The behavior segmentation results represent the segmentation of different behaviors of the subject in the video, while the shot segmentation results represent the segmentation of different shots in the video. Further, based on the content descriptions of the video segments corresponding to different shots included in the shot segmentation results, the electronic device obtains the scene segmentation results corresponding to the video. The scene segmentation results represent the segmentation of different scenes in the video. The electronic device can perform multi-granularity recognition on the video to be processed, obtaining scene-granularity recognition results, shot-granularity recognition results, and behavior-granularity recognition results. It can simultaneously perform multi-granularity temporal segmentation of the video, including coarse-to-fine segmentation of scenes, shots, and behaviors, and can output type labels for shot changes (e.g., transition, fade, camera movement, NULL) and behavior labels (preset behavior labels, non-preset behavior labels NULL). This provides a more detailed understanding and segmentation of video content, offering more detailed video content information to support video understanding and analysis in various application scenarios later on. For example, it can enable the search and location of more detailed parts of video content, and can process (edit, composite, etc.) content that appears in the same action or in the same shot in the video content. At the same time, it can also improve the user experience of using the video processing functions provided by electronic devices.

[0407] It should be noted that the personal information used in the technical solution of this application is limited to information for which individual consent has been obtained, including but not limited to notifying and reminding users to read the relevant user agreement (notification) and sign the agreement (authorization) which includes authorization of relevant user information before users use the function.

[0408] The technical solutions disclosed in this application involve the collection, storage, use, processing, transmission, provision, and disclosure of users' personal information, all of which comply with relevant laws and regulations and do not violate public order and good morals.

[0409] Figure 16 A schematic diagram of a possible structure of the electronic device involved in the above embodiments is shown. Figure 16 The electronic device 1600 shown includes a processing module 1601, a display module 1602, and a storage module 1603.

[0410] The processing module 1601 may be a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The processor may include an application processor and a baseband processor. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0411] For example, the processing module 1601 can be as follows: Figure 8 The processor 110 shown; the display module 1602 can be as follows: Figure 8 The display screen 194 shown; the storage module 1603 can be as follows: Figure 8 The internal memory 121 shown. The electronic device provided in this application embodiment can be Figure 8 The electronic device 100 shown.

[0412] This application also provides a chip system (e.g., a system-on-a-chip (SoC)). Figure 17 As shown, the chip system includes at least one processor 1701 and at least one interface circuit 1702. The processor 1701 and the interface circuit 1702 are interconnected via lines. For example, the interface circuit 1702 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 1702 can be used to send signals to other devices (e.g., the processor 1501 or the camera of an electronic device). Exemplarily, the interface circuit 1702 can read instructions stored in the memory and send those instructions to the processor 1701. When the instructions are executed by the processor 1701, the electronic device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete components, which are not specifically limited in this application embodiment.

[0413] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device, cause the electronic device to perform various functions or steps performed by the electronic device 100 in the above method embodiment.

[0414] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the electronic device 100 in the above method embodiments. For example, the computer may be the aforementioned electronic device 100.

[0415] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0416] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0417] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0418] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0419] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0420] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method of video processing, the method comprising: The method comprises: An electronic device acquires a to-be-processed video; The electronic device performs visual analysis on the to-be-processed video to acquire a behavior division result and a shot division result of the to-be-processed video; wherein the behavior division result represents the division result of video segments corresponding to different behaviors of a subject in the to-be-processed video, and the shot division result represents the division result of video segments corresponding to different shots in the to-be-processed video; The electronic device acquires a scene division result corresponding to the to-be-processed video based on the content description of the video segments corresponding to different shots contained in the shot division result; wherein the electronic device acquires a speech text description of the to-be-processed video; the speech text description of the to-be-processed video comprises a plurality of speech text segments divided in time / chronologically; the electronic device determines a scene label of each video segment based on the content description of each video segment by taking each speech text segment as a reference description; if the description of one speech text segment covers the content description of a plurality of adjacent video segments, it is determined that the adjacent video segments belong to one scene label, and different scene labels and the time stamp of the to-be-processed video corresponding to the scene label are taken as the scene division result; the scene division result represents the division result of video segments corresponding to different scenes in the to-be-processed video; one scene comprises one or more shots.

2. The method of claim 1, wherein, The electronic device performs visual analysis on the to-be-processed video to acquire a behavior division result and a shot division result of the to-be-processed video, comprising: The electronic device performs visual analysis on the to-be-processed video to acquire at least one behavior label and at least one shot label of the to-be-processed video; The electronic device takes the at least one behavior label and the time stamp of the video segment corresponding to the behavior label as the behavior division result; The electronic device takes the at least one shot label and the time stamp of the video segment corresponding to the shot label as the shot division result; Wherein, the type of the behavior label comprises a plurality of preset behavior labels, a non-preset behavior label or no behavior label, and the type of the shot label comprises a no-shot switching label and a shot label corresponding to a plurality of preset shot switching types.

3. The method of claim 2, wherein, The electronic device performs visual analysis on the to-be-processed video to acquire at least one behavior label and at least one shot label of the to-be-processed video, comprising: The electronic device acquires a plurality of to-be-processed video segments of the to-be-processed video in chronological order; the time length of the plurality of to-be-processed video segments is the same as a preset sliding window length; The electronic device performs visual analysis on each to-be-processed video segment to acquire at least one behavior label and at least one shot label of each to-be-processed video segment; If adjacent to-be-processed video segments have overlapping segments, the electronic device performs de-duplication and merging processing on the behavior labels and the shot labels corresponding to the overlapping segments to obtain the behavior labels and the shot labels of all to-be-processed video segments.

4. The method of claim 2, wherein, The electronic device obtains a scene division result corresponding to the to-be-processed video based on content descriptions of video clips corresponding to different shots contained in the shot division result, and the obtaining includes: The electronic device performs a video clip division operation on the to-be-processed video based on time stamps of video clips corresponding to each shot label in the shot division result, and obtains a plurality of divided video clips; the shots of time-sequentially adjacent video clips in the plurality of video clips are different; The electronic device performs video content recognition on each video clip to obtain a content description of each video clip; the content description includes picture content corresponding to the video clip; The electronic device obtains a scene division result corresponding to the to-be-processed video based on the content description of each video clip.

5. The method of claim 3, wherein, The electronic device obtains a scene division result corresponding to the to-be-processed video based on the content description of each video clip, and the obtaining includes: The electronic device determines a scene description of each video clip based on the content description of each video clip; the scene description includes a description of an environmental scene corresponding to the video clip; If the similarity of the scene descriptions of adjacent video clips is greater than or equal to a similarity threshold, it is determined that the adjacent video clips belong to one scene label, and a scene label of a video clip corresponding to a different scene is obtained; The scene division result is obtained by using different scene labels and time stamps of the to-be-processed video corresponding to the scene labels.

6. The method of claim 1 or 2, wherein, The electronic device obtains a speech text description of the to-be-processed video, and the obtaining includes: The electronic device performs speech analysis on the to-be-processed video to obtain a speech text description of the to-be-processed video; the speech text description includes one or more of a text description corresponding to a human voice dialogue in the to-be-processed video, a text description corresponding to a voiceover voice of the to-be-processed video, and a text description corresponding to an audio effect in the to-be-processed video; The electronic device obtains a scene division result corresponding to the to-be-processed video based on content descriptions of video clips corresponding to different shots contained in the shot division result, and the obtaining further includes: The electronic device performs a video clip division operation on the to-be-processed video based on time stamps of video clips corresponding to each shot label in the shot division result, and obtains a plurality of divided video clips; the shots of time-sequentially adjacent video clips in the plurality of video clips are different; The electronic device performs video content recognition on each video clip to obtain a content description of each video clip.

7. The method of claim 1, wherein, The electronic device determines a scene label of each video clip based on the content description of each video clip by using each speech text segment as a reference description, and the determining includes: If the content description of at least one video clip matches a speech text segment, the speech text segment is used as a reference description of the at least one video clip, and a scene label of the at least one video clip is determined.

8. The method of claim 7, wherein, The plurality of speech text segments include a first speech text segment, and the method further includes: if the content description of all video clips does not match the first voice text segment, discarding the first voice text segment.

9. The method according to any one of claims 1-6, characterized in that, The method further comprises: The electronic device stores the behavior division result, the shot division result and the scene division result of the video to be processed into a video label database; the video label database stores the behavior division result, the shot division result and the scene division result corresponding to a plurality of videos.

10. The method of claim 9, wherein, In the video label database, the behavior division result includes a behavior label, the shot division result includes a shot label, and the scene division result includes a scene label, and the method further comprises: The electronic device receives a video content query request; the video content query request includes one or more of a scene description, a shot description or a behavior description; The electronic device queries a target label matching the video content query request from the scene division result, the shot division result and the behavior division result of a plurality of videos in the video label database; the target label includes one or more of the behavior label, the shot label or the scene label; The electronic device acquires a target video clip corresponding to the timestamp based on the timestamp corresponding to the target label; The electronic device outputs the target video clip.

11. The method of claim 9, wherein, In the video label database, the behavior division result includes a behavior label, the shot division result includes a shot label, and the scene division result includes a scene label, and the method further comprises: The electronic device receives a video clip request; the video clip request includes one or more of a scene description, a shot description or a behavior description of a video to be clipped; The electronic device queries a target label matching the video clip request from the scene division result, the shot division result and the behavior division result of a plurality of videos in the video label database; the target label includes one or more of the behavior label, the shot label or the scene label; The electronic device acquires at least one target video clip corresponding to the timestamp based on the timestamp corresponding to the target label; The electronic device outputs a video synthesized by the at least one target video clip.

12. The method of any one of claims 1-6, wherein, The electronic device performs visual analysis on the video to be processed to acquire the behavior division result and the shot division result of the video to be processed, comprising: The electronic device inputs the video to be processed into a visual analysis model to perform visual analysis and acquire the behavior division result and the shot division result of the video to be processed; The visual analysis model is a model obtained by learning and training using a plurality of cross-attention mechanisms, the input of the visual analysis model is the video to be processed, and the output of the visual analysis model includes frame-level behavior prediction labels and frame-level shot prediction labels; the behavior division result is obtained by merging adjacent image frames with the same behavior prediction label, and the shot division result is obtained by merging adjacent image frames with the same shot prediction label; The training process of the visual analysis model comprises: The frame-granularity feature vector, the shot-granularity feature vector, and the behavior-granularity feature vector are alternately taken as a query vector, a key vector, and a value vector in a cross-attention mechanism for cross learning to obtain a behavior prediction label, a shot prediction label, a frame-level behavior prediction label, and a frame-level shot prediction label. A model loss is calculated based on the behavior prediction label, the shot prediction label, the frame-level behavior prediction label, and the frame-level shot prediction label. The visual analysis model is trained based on the model loss until the model loss of the visual analysis model is less than a loss threshold.

13. A method of training a visual analytics model, the method comprising: The visual analysis model includes a feature extraction module and a learning interaction module, and the method includes: The electronic device inputs a sample video into the feature extraction module of the visual analysis model to obtain a first shot feature vector, a first behavior feature vector, and a first frame feature vector; The electronic device inputs the first shot feature vector, the first behavior feature vector, and the first frame feature vector into the learning interaction module, learns by using a cross-attention mechanism, and outputs a prediction label; the prediction label includes a frame-level behavior prediction label and a frame-level shot prediction label; The electronic device trains the visual analysis model to obtain a trained visual analysis model that meets a training condition; The training condition includes that the model loss of the visual analysis model is less than a loss threshold, or the visual analysis model reaches convergence, or the number of times of training of the visual analysis model reaches a number threshold.

14. The method of claim 13, wherein, The electronic device inputs a sample video into the feature extraction module of the visual analysis model to obtain a first shot feature vector, a first behavior feature vector, and a first frame feature vector, including: The electronic device inputs the sample video into the feature extraction module of the visual analysis model, and the feature extraction module outputs the first shot feature vector by using a shallow network, and outputs the first frame feature vector and the first behavior feature vector by using a deep network.

15. The method of claim 14, wherein, The learning interaction module includes an initial layer and an interaction layer; the initial layer includes a first behavior learning module, a first frame learning module, and a first shot learning module; and the interaction layer includes a second behavior learning module, a second shot learning module, a second frame learning module, and a third frame learning module. The electronic device inputs the first shot feature vector, the first behavior feature vector, and the first frame feature vector into the learning interaction module, learns by using a cross-attention mechanism, and outputs a prediction label, including: The electronic device inputs the first shot feature vector into the first shot learning module of the initial layer to obtain a second shot feature vector; the receptive field of the second shot feature vector is larger than that of the first shot feature vector; The electronic device inputs the first frame feature vector into the first frame learning module of the initial layer to obtain a second frame feature vector; the receptive field of the second frame feature vector is larger than that of the first frame feature vector; and The electronic device inputs the first frame feature vector into the first frame learning module of the initial layer to obtain a second frame feature vector; the receptive field of the second frame feature vector is larger than that of the first frame feature vector. The electronic device inputs the second shot feature vector, the second behavior feature vector and the second frame feature vector into the interaction layer, and learns by using a cross attention mechanism to output a prediction label. The electronic device obtains a shot prediction label of the sample video based on the third shot feature vector.

16. The method of claim 15, wherein, The electronic device inputs the first shot feature vector, the first behavior feature vector and the first frame feature vector into the learning interaction module, learns by using a cross attention mechanism, and outputs a prediction label, and further includes: The electronic device inputs the first behavior feature vector into the first behavior learning module of the initial layer to obtain a second behavior feature vector; a receptive field of the second behavior feature vector is larger than a receptive field of the first behavior feature vector; The electronic device inputs the second frame feature vector as a query vector in the cross attention mechanism, and inputs the second behavior feature vector as a key vector and a value vector in the cross attention mechanism into the second behavior learning module of the interaction layer to obtain a third behavior feature vector; The electronic device obtains a behavior prediction label of the sample video based on the third behavior feature vector.

17. The method of claim 16, wherein, The electronic device inputs the first shot feature vector, the first behavior feature vector and the first frame feature vector into the learning interaction module, learns by using a cross attention mechanism, and outputs a prediction label, and further includes: The electronic device inputs the third shot feature vector as a query vector in the cross attention mechanism, and inputs the second frame feature vector as a key vector and a value vector in the cross attention mechanism into the second frame learning module of the interaction layer to obtain a third frame feature vector; The electronic device inputs the third behavior feature vector as a query vector in the cross attention mechanism, and inputs the second frame feature vector as a key vector and a value vector in the cross attention mechanism into the third frame learning module of the interaction layer to obtain a fourth frame feature vector; The electronic device obtains a frame-level shot prediction label corresponding to the sample video based on the third frame feature vector; The electronic device obtains a frame-level behavior prediction label corresponding to the sample video based on the fourth frame feature vector.

18. The method of claim 17, wherein, The visual analysis model further includes a multi-layer perception; The electronic device obtains a frame-level shot prediction label corresponding to the sample video based on the third frame feature vector, including: The electronic device inputs the third frame feature vector into the multi-layer perception for dimension reduction processing to obtain the frame-level shot prediction label corresponding to the sample video; The electronic device obtains a frame-level behavior prediction label corresponding to the sample video based on the fourth frame feature vector, including: The electronic device inputs the fourth frame feature vector into the multi-layer perception for dimension reduction processing to obtain the frame-level behavior prediction label corresponding to the sample video.

19. The method of claim 18, wherein, The electronic device trains the visual analysis model to obtain a trained visual analysis model, including: The electronic device calculates a shot prediction loss based on the shot prediction label and a standard shot label corresponding to the sample video; The electronic device calculates a behavior prediction loss based on the behavior prediction label and a standard behavior label corresponding to the sample video; The electronic device calculates a frame-level shot prediction loss based on the frame-level shot prediction label and a standard shot label corresponding to the sample video; The electronic device calculates a frame-level behavior prediction loss based on the frame-level behavior prediction label and a standard behavior label corresponding to the sample video; The electronic device trains the visual analysis model based on a sum of one or more losses of the shot prediction loss, the behavior prediction loss, the frame-level shot prediction loss, and the frame-level behavior prediction loss until the loss is less than the loss threshold, and obtains the trained visual analysis model.

20. The method of claim 19, wherein, The visual analysis model further includes a time sequence processing module, and the method further includes: After the electronic device obtains the frame-level behavior prediction label of the sample video, adjacent image frames with the same frame-level behavior prediction label are merged in time sequence to obtain a merged behavior prediction label as a behavior division result; A frame-level behavior prediction loss is calculated based on the merged behavior prediction label and a standard behavior label corresponding to the sample video; After the electronic device obtains the frame-level shot prediction label of the sample video, adjacent image frames with the same frame-level shot prediction label are merged in time sequence to obtain a merged shot prediction label as a shot division result; A frame-level shot prediction loss is calculated based on the merged shot prediction label and a standard shot label corresponding to the sample video.

21. The method according to any one of claims 13-20, characterized by, The visual analysis model further includes a sliding window processing module, and the method further includes: The electronic device inputs the sample video into the sliding window processing module to obtain at least one sample video segment of the sample video; the duration of the at least one sample video segment is the same as a preset sliding window length; The electronic device inputs the sample video into the feature extraction module of the visual analysis model to obtain a first shot feature vector, a first behavior feature vector, and a first frame feature vector, including: For each of the sample video segments, the electronic device inputs the sample video segment into the feature extraction module of the visual analysis model to obtain the first shot feature vector, the first behavior feature vector, and the first frame feature vector.

22. The method of claim 21, wherein, In the case that the plurality of sample video segments exist repeated segments, the output prediction label includes: After the electronic device obtains the frame-level behavior prediction label and the frame-level shot prediction label of all the sample video segments, the frame-level behavior prediction label and the frame-level shot prediction label of the repeated image frames are de-duplicated to obtain the frame-level behavior prediction label and the frame-level shot prediction label of all the sample video segments after de-duplication.

23. An electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program, when executed by the processor, causes the electronic device to perform the method of any one of claims 1-22. The processor executes the computer program to implement the steps of the method of any one of claims 1-12 or the steps of the method of any one of claims 13-22.

24. A computer readable storage medium having stored thereon computer programs / instructions, characterized in that, The computer program / instructions, which when executed by the processor, implement the steps of the method of any one of claims 1-12 or the steps of the method of any one of claims 13-22.

25. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions, which when executed by the processor, implement the steps of the method of any one of claims 1-12 or the steps of the method of any one of claims 13-22.

Citation Information

Patent Citations

  • Action segmentation model processing method and device, computer equipment and storage medium

    CN113591529A

  • Video scene segmentation method and device, equipment and storage medium

    CN116453006A

  • Video processing method and device, computer readable medium and electronic equipment

    CN117061815A

  • Video processing method, electronic equipment, chip system and storage medium

    CN118474448A