Video processing method and device and electronic equipment

By determining the video clip sequence, its text content and action tag information, and using a multimodal big model to extract and understand key video sub-segments, the problem of low processing efficiency of e-commerce live videos in the prior art is solved, and efficient understanding and summary generation of long videos and emergencies are achieved.

CN120602748APending Publication Date: 2025-09-05BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510813014.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing video processing methods are difficult to effectively process e-commerce live videos that last for a long time, especially the accuracy of understanding of emergencies, resulting in inefficient video processing.

Method used

By determining the video clip sequence, its text content and action tag information, a multimodal model is used to extract key video sub-segments and perform video understanding processing to generate the summary content of the key video sub-segments.

Benefits of technology

Improve the extraction accuracy and processing efficiency of key video sub-snippets in videos that last for a long time, understand sudden events in the video, and generate high-quality video summary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602748A_ABST
    Figure CN120602748A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method and device and electronic equipment, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, and can be applied to scenes of digital human, content generation based on artificial intelligence and the like. The specific implementation scheme is as follows: determining a video clip sequence of a to-be-processed video, and text contents and at least one piece of action label information corresponding to video clips in the video clip sequence; according to the text content and the at least one piece of action label information, key video sub-clips are extracted from the video clip; performing video understanding processing on the key video sub-clips to obtain sub-clip abstract contents of the key video sub-clips; wherein the video clip sequence is determined, so that the video processing method can be suitable for videos with relatively long duration time and the like; by considering at least one piece of action label information, the video processing method can understand sudden events and the like in the video, so that the video processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as computer vision, deep learning, and large models, and can be applied to scenarios such as digital humans and artificial intelligence-based content generation, and in particular to a video processing method, device, and electronic device. Background Art

[0002] With the continuous development of technologies such as artificial intelligence, computer vision, and large models, video processing technology has been widely used in many fields such as digital humans, artificial intelligence-based content generation technology (AIGC), video generation, and content search. Summary of the Invention

[0003] The present disclosure provides a video processing method, device and electronic device.

[0004] According to one aspect of the present disclosure, a video processing method is provided, comprising: determining a video clip sequence of a video to be processed, as well as text content and at least one action label information corresponding to the video clips in the video clip sequence; extracting key video sub-segments from the video clips based on the text content and the at least one action label information; and performing video understanding processing on the key video sub-segments to obtain sub-segment summary content of the key video sub-segments.

[0005] According to another aspect of the present disclosure, a video processing device is provided, comprising: a determination module for determining a video segment sequence of a video to be processed, as well as text content and at least one action label information corresponding to the video segments in the video segment sequence; an extraction module for extracting key video sub-segments from the video segments based on the text content and the at least one action label information; and an understanding processing module for performing video understanding processing on the key video sub-segments to obtain sub-segment summary content of the key video sub-segments.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the video processing method proposed above in the present disclosure.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the video processing method proposed above in the present disclosure.

[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the steps of the video processing method proposed above in the present disclosure when executed by a processor.

[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0011] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;

[0012] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;

[0013] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;

[0014] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure;

[0015] Figure 5 is a schematic diagram of extracting key video sub-segments;

[0016] Figure 6 It is a schematic diagram of understanding of key video sub-segments;

[0017] Figure 7 is a schematic diagram according to a fifth embodiment of the present disclosure;

[0018] Figure 8 is a block diagram of an electronic device for implementing the video processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0020] With the continuous development of technologies such as artificial intelligence, computer vision, and large models, video processing technology has been widely used in many fields such as digital humans, artificial intelligence-based content generation technology (AIGC), video generation, and content search.

[0021] The current video processing method mainly uses a dual-tower architecture or a two-stage architecture model to understand and process the video in order to extract the content summary of the video.

[0022] The model in the above solution has difficulty in understanding and processing videos of long-lasting e-commerce live broadcasts, and its accuracy in understanding sudden events in e-commerce live broadcasts is low, which reduces the efficiency of video processing.

[0023] To address the above problems, the present disclosure provides a video processing method, device, and electronic device.

[0024] Figure 1 This is a schematic diagram according to the first embodiment of the present disclosure. It should be noted that the video processing method of the embodiment of the present disclosure can be applied to a video processing device, which can be configured in an electronic device so that the electronic device can perform video processing functions.

[0025] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, and other hardware devices with various operating systems, touch screens and / or display screens.

[0026] The video processing device may also be software in an electronic device, such as video processing software, etc. In the following embodiments, the execution subject is an electronic device as an example for description.

[0027] like Figure 1 As shown, the video processing method may include the following steps:

[0028] Step 101: Determine a video segment sequence of a video to be processed, as well as text content and at least one action tag information corresponding to the video segments in the video segment sequence.

[0029] In the embodiments of the present disclosure, the video to be processed can be any resource that can be represented in video form, including multiple video frames, such as video, short video, live stream, digital human video, digital human conversation video, video formed by video conversation, etc.

[0030] In the disclosed embodiments, the sequence of video segments of the video to be processed can be determined by segmenting the video to be processed to obtain multiple video segments; and then arranging the multiple video segments according to their timestamps to obtain a sequence of video segments. The segmentation method can be, for example, segmentation into equal time periods or segmentation into unequal time periods. It should be noted that the segmentation method for the video to be processed is not specifically limited and can be set according to actual needs.

[0031] In the embodiment of the present disclosure, the text content corresponding to the video clip may be text subtitles in the video clip, and / or a speech recognition result of the speech content in the video clip.

[0032] In the embodiment of the present disclosure, at least one action tag information is used to describe the relevant action of at least one character in the video to be processed. The action tags in the action tag information may include, for example, waving, pointing at the camera, shouting, displaying a product, etc.

[0033] Step 102: extract key video sub-segments from the video segment based on the text content and at least one action tag information.

[0034] In an embodiment of the present disclosure, in one example, for any video segment in a video segment sequence, the process of the electronic device executing step 102 may, for example, be to input the text content corresponding to the video segment, at least one action label information, and each video frame in the video segment into a first multimodal large model, and obtain at least one set of key start and end timestamps output by the first multimodal large model; and determine at least one key video sub-segment in the video segment based on the video segment and the at least one set of key start and end timestamps.

[0035] A set of key start and end timestamps may include a start timestamp and an end timestamp. The start timestamp and the end timestamp may be timestamps in a video segment, such as the Nth second in the video segment. The start timestamp may be used to define the start time point of a key video sub-segment in the video segment, and the end timestamp may be used to define the end time point of the key video sub-segment in the video segment.

[0036] In another example of the embodiment of the present disclosure, for any video segment in a video segment sequence, the electronic device may perform step 102 by, for example, sampling the video segment to obtain at least one video frame in the video segment; inputting text content, at least one video frame, and at least one action label information corresponding to the video segment into a first multimodal large model to obtain at least one set of key start and end timestamps output by the first multimodal large model; and determining at least one key video sub-segment in the video segment based on the video segment and the at least one set of key start and end timestamps.

[0037] Taking into account the number of parameters and computational accuracy of the first multimodal large model, key video sub-segments are determined by combining the first multimodal large model, the text content corresponding to the video segment, at least one video frame, and at least one action tag information, thereby improving the accuracy of extracting key video sub-segments. The input of at least one action tag information enables the first multimodal large model to focus on the action tag information in the video segment, thereby extracting key video sub-segments involving the action tag information, further improving the accuracy of extracting key video sub-segments.

[0038] In the embodiment of the present disclosure, in order to ensure that the first multimodal large model can accurately understand the current task, thereby extracting key video sub-segments involving action label information, and further improving the extraction accuracy of key video sub-segments, the prompt text of the first multimodal large model can be set so that the prompt text can indicate that the key video sub-segments need to contain action-related semantic information.

[0039] Among them, action-related semantic information may include at least one of the following: object interaction information, character interaction information, action reinforcement information, and emotion expression information.

[0040] Examples of item interaction include product usage and display, blackboard explanations, and scene interpretation using props. Examples of character interaction include gestures like "like / follow" facing the camera, coordinating with the host, and proactively responding to audience comments. Examples of action reinforcement include countdowns to link, waving to attract attention, tapping the table, and leaning forward. Examples of emotional expression include expressions of nervousness, humor, and surprise.

[0041] Step 103: Perform video understanding processing on the key video sub-segment to obtain the sub-segment summary content of the key video sub-segment.

[0042] In an embodiment of the present disclosure, the electronic device may input key video sub-segments into a video understanding model to obtain sub-segment summary content output by the video understanding model.

[0043] In the disclosed embodiment, after determining the sub-segment summary content of the key video sub-segments, the video to be processed can be used as a base, combined with the sub-segment summary content of each key video sub-segment to generate a digital human video. In other words, the video is generated by replacing the real person in the video to be processed with a digital human.

[0044] In the embodiment of the present disclosure, after determining the sub-segment summary contents of the key video sub-segments, content generation processing may be performed on the video to be processed in combination with the sub-segment summary contents of each key video sub-segment.

[0045] In the embodiment of the present disclosure, after determining the sub-segment summary content of the key video sub-segment, the video to be processed can be annotated based on the summary content of each sub-segment, so that the object can select the key video sub-segment corresponding to the sub-segment summary content of interest for viewing according to actual needs.

[0046] In the disclosed embodiment, after determining the sub-segment summary content of the key video sub-segment, the sub-segment summary content can be arranged sequentially. For example, the sub-segment summary content related to product A can be placed before the sub-segment summary content related to product B. The video generation process is then re-performed based on the arrangement order, the key video sub-segment, and the video to be processed.

[0047] The video processing method of the embodiment of the present disclosure determines a video clip sequence of a video to be processed, as well as text content and at least one action label information corresponding to the video clips in the video clip sequence; extracts key video sub-clips from the video clips based on the text content and the at least one action label information; performs video understanding processing on the key video sub-clips to obtain sub-clips summary content of the key video sub-clips; wherein, the determination of the video clip sequence enables the video processing method to be applicable to videos of longer duration, etc.; the consideration of at least one action label information enables the video processing method to understand sudden events in the video, etc., thereby improving video processing efficiency.

[0048] In order to further improve the extraction accuracy of key video sub-segments, the electronic device can recognize the voice content of the video segment to determine the text content, and perform action tag information extraction processing on the video segment to determine the action tag information. Figure 2 As shown, Figure 2 is a schematic diagram according to a second embodiment of the present disclosure, Figure 2 The illustrated embodiment may include the following steps:

[0049] Step 201: Determine a video segment sequence of a video to be processed.

[0050] Step 202: perform recognition processing on the speech content of the video segments in the video segment sequence to obtain text content corresponding to the video segments.

[0051] In an embodiment of the present disclosure, the electronic device may obtain the voice content of each video segment in a video segment sequence; input the voice content into a speech recognition model, and obtain text content output by the speech recognition model.

[0052] Step 203: extract action label information from the video clip to obtain at least one action label information corresponding to the video clip.

[0053] In an embodiment of the present disclosure, in one example, the electronic device may perform step 203 by inputting the video clip into a label information extraction model and obtaining at least one action label information output by the label information extraction model.

[0054] In another example of the embodiment of the present disclosure, the process of the electronic device executing step 203 may, for example, be to perform human key point detection processing on at least one video frame in the video clip to obtain human key point information in the at least one video frame; and determine at least one action label information corresponding to the video clip based on the human key point information in the at least one video frame.

[0055] In the disclosed embodiments, human body key point information may include at least one of the following: body key point information and hand key point information. The body key point information may include position information of each body key point in the video frame. Examples of body key points include head key points, upper arm key points, forearm key points, upper limb joint key points, thigh key points, calf key points, foot key points, and lower limb joint key points. These are not specifically limited here and can be set based on actual needs.

[0056] The hand key point information may include the position information of each hand key point in the video frame. For example, the hand key points include the finger base key points, the proximal knuckle key points, the distal knuckle key points, and the fingertip key points. There is no specific limitation here and it can be set according to actual needs.

[0057] Among them, the action tag information is mainly reflected in the key points of the human body. The setting of multiple key points of the human body can enable the electronic device to determine the action tag information based on the multiple key points of the human body, thereby further improving the accuracy of the determination of the action tag information.

[0058] In the disclosed embodiment, to reduce the amount of data processing during action tag information extraction, reduce the number of video frames to be processed, and improve video processing efficiency, at least one video frame in a video clip is at least one video frame obtained by sampling the video clip. The sampling frame rate may be, for example, 2 FPS (frame rate).

[0059] In an embodiment of the present disclosure, the electronic device performs human key point detection processing on at least one video frame in a video clip. The process of obtaining human key point information in at least one video frame can, for example, be to input the video frame into a human posture detection model for each video frame in the at least one video frame, and obtain the human key point information output by the human posture detection model.

[0060] Among them, considering the parameter quantity and calculation accuracy of the human posture detection model, combining the human posture detection model and the video frame to determine the human body key point information in the video frame can further improve the accuracy of the determined human body key point information.

[0061] Among them, human key point detection processing is performed on at least one video frame in the video clip, so that the electronic device can accurately determine the action situation in the video clip based on the position and movement of the human key points in the video clip, thereby further improving the accuracy of the determined action label information.

[0062] In an embodiment of the present disclosure, the process of an electronic device determining at least one action label information corresponding to a video clip based on the human key point information in at least one video frame can, for example, be to perform time window sliding processing on the human key point information in at least one video frame to obtain at least one time window and a sequence of human key point information within the time window; for each time window in the at least one time window, determine the action label information based on the sequence of human key point information within the time window.

[0063] Among them, for each time window in at least one time window, the electronic device determines the action label information based on the human body key point information sequence in the time window. For example, the process can be: performing feature extraction processing on the human body key point information sequence in the time window to obtain the action trajectory feature; querying the trajectory feature library based on the action trajectory feature to obtain the first reference trajectory feature in the trajectory feature library that matches the action trajectory feature; and determining the action label information based on the action label corresponding to the first reference trajectory feature and the start and end timestamps of the human body key point information sequence in the time window.

[0064] The reference trajectory features and the action trajectory features in the trajectory feature library can be extracted using the same trajectory feature extraction model.

[0065] Among them, time window sliding processing is performed on the human body key point information in at least one video frame, and the motion trajectory features extracted according to the human body key point information sequence in at least one time window are matched with the reference trajectory features in the trajectory feature library to determine the action label information. The action label information can be determined by combining more action-related information, thereby further improving the accuracy of the determined action label information.

[0066] Step 204 : extract key video sub-segments from the video segment based on the text content and at least one action tag information.

[0067] Step 205 : Perform video understanding processing on the key video sub-segment to obtain the sub-segment summary content of the key video sub-segment.

[0068] It should be noted that the details of steps 204 to 205 can be found in Figure 1 Steps 102 to 103 in the illustrated embodiment will not be described in detail here.

[0069] The video processing method of the embodiment of the present disclosure determines a video clip sequence of a video to be processed; performs recognition processing on the voice content of the video clips in the video clip sequence to obtain text content corresponding to the video clips; performs action label information extraction processing on the video clips to obtain at least one action label information corresponding to the video clips; extracts key video sub-segments from the video clips based on the text content and at least one action label information; performs video understanding processing on the key video sub-segments to obtain sub-segment summary content of the key video sub-segments; wherein, the action label extraction processing on the video clips enables the key video sub-segments in the video clips to be extracted in combination with the extracted at least one action label, thereby improving the extraction accuracy of the key video sub-segments, thereby further improving the video processing efficiency.

[0070] In the case where adjacent video segments in a video segment sequence overlap, the same key video sub-segment may be extracted from adjacent video segments, that is, the same key video sub-segment is split into two. To avoid the above phenomenon, the two key video sub-segments can be fused to further improve the accuracy of the key video sub-segments obtained. Figure 3 As shown, Figure 3 is a schematic diagram according to a third embodiment of the present disclosure, Figure 3 The illustrated embodiment may include the following steps:

[0071] Step 301: Determine a video segment sequence of a video to be processed, as well as text content and at least one action tag information corresponding to the video segments in the video segment sequence.

[0072] Step 302: When there is overlap between adjacent video segments in the video segment sequence and when there is overlap between key video sub-segments in the adjacent video segments, obtain at least two overlapping key video sub-segments.

[0073] In the embodiments of the present disclosure, overlapping between adjacent video segments refers to the presence of at least two identical video frames in the adjacent video segments. Overlapping between key video sub-segments in adjacent video segments refers to the presence of overlapping between a key video sub-segment in a first video segment and another key video sub-segment in a second video segment in the adjacent video segments.

[0074] Step 303: Perform fusion processing on at least two key video sub-segments.

[0075] In an embodiment of the present disclosure, in one example, the electronic device may perform step 303 by, for example, determining a candidate video segment including at least two key video sub-segments in the video to be processed; and determining the candidate video segment as the key video sub-segment after fusion processing.

[0076] Among them, the electronic device can determine the start timestamp and end timestamp of at least two key video sub-segments in the video to be processed; determine the minimum start timestamp among multiple start timestamps; determine the maximum end timestamp among multiple end timestamps; obtain a candidate video segment in the video to be processed based on the minimum start timestamp and the maximum end timestamp; the start timestamp of the candidate video segment is the minimum start timestamp; the end timestamp of the candidate video segment is the maximum end timestamp.

[0077] Among them, using the candidate video segments including at least two key video sub-segments in the video to be processed as the key video sub-segments after fusion processing can reduce the number of key video sub-segments and improve the extraction accuracy of key video sub-segments.

[0078] In another example of the embodiment of the present disclosure, the process of the electronic device executing step 303 may, for example, be to determine a candidate video segment including at least two key video sub-segments in the video to be processed; extract candidate key video sub-segments from the candidate video segment based on the text content and at least one action label information corresponding to the candidate video segment; and determine the candidate key video sub-segments as the key video sub-segments after fusion processing.

[0079] The electronic device may obtain the voice content of the candidate video clip; perform recognition processing on the voice content to obtain text content; and perform action label information extraction processing on the candidate video clip to obtain at least one action label information.

[0080] Among them, according to the text content and at least one action label information corresponding to the candidate video segment, the candidate key video sub-segments are extracted from the candidate video segment and used as the key video sub-segments after fusion processing, which can further improve the extraction accuracy of the key video sub-segments.

[0081] Step 304 : extract key video sub-segments from the video segment based on the text content and at least one action tag information.

[0082] Step 305 : Perform video understanding processing on the key video sub-segment to obtain the sub-segment summary content of the key video sub-segment.

[0083] It should be noted that the details of step 301, step 304 to step 305 can be found in Figure 1 Steps 101 to 103 in the illustrated embodiment will not be described in detail here.

[0084] The video processing method of the embodiment of the present disclosure determines a video clip sequence of a video to be processed, as well as text content and at least one action label information corresponding to the video clips in the video clip sequence; when there is overlap between adjacent video clips in the video clip sequence and when there is overlap between key video sub-clips in adjacent video clips, obtains at least two overlapping key video sub-clips; performs fusion processing on the at least two key video sub-clips; extracts key video sub-clips from the video clips based on the text content and at least one action label information; performs video understanding processing on the key video sub-clips to obtain sub-clips summary content of the key video sub-clips; wherein, when there is overlap between adjacent video clips in the video clip sequence, the same key video sub-clips may be extracted from the adjacent video clips, and fusion processing on the same key video sub-clips can further improve the accuracy of the determined key video sub-clips.

[0085] In order to further improve the accuracy of the sub-segment summary content of the key video sub-segment, the second multimodal large model can be combined to determine at least one action-related element content of the key video sub-segment, and then the sub-segment summary content can be determined in combination with the action-related element content. Figure 4 As shown, Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure, Figure 4 The illustrated embodiment may include the following steps:

[0086] Step 401: Determine a video segment sequence of a video to be processed, as well as text content and at least one action tag information corresponding to the video segments in the video segment sequence.

[0087] Step 402: extract key video sub-segments from the video segment based on the text content and at least one action tag information.

[0088] Step 403: Determine the text content corresponding to the key video sub-segment.

[0089] In an embodiment of the present disclosure, in one example, the electronic device may perform recognition processing on the voice content of the key video sub-segment to obtain text content corresponding to the key video sub-segment.

[0090] In another example, each character in the text content corresponding to the video clip to which the key video sub-clip belongs is assigned a timestamp to indicate the correspondence between the characters in the text content and the video frames in the video clip. Accordingly, based on the start and end timestamps of the key video sub-clip, the characters corresponding to each timestamp from the start to the end timestamp can be determined. The characters are then concatenated in timestamp order to obtain the text content corresponding to the key video sub-clip.

[0091] Step 404 : Perform video frame sampling processing on the key video sub-segment to obtain at least one video frame in the key video sub-segment.

[0092] Step 405: Input the text content and at least one video frame into the second multimodal large model, and obtain at least one action-related element content output by the second multimodal large model.

[0093] In the disclosed embodiments, it should be noted that, to further improve the accuracy of determining action-related element content, a prompt text can be provided for the second multimodal large model. This prompt text can indicate the extraction of at least one action-related element content. In other words, the prompt text can indicate which action-related element content needs to be extracted and the specific content under each action-related element content.

[0094] In an embodiment of the present disclosure, to further improve the accuracy of determining action-related element content, the electronic device may further perform the following process: performing optical character recognition on at least one video frame to obtain recognized content in the at least one video frame. Accordingly, the electronic device may perform step 405 by, for example, inputting text content, at least one video frame, and recognized content into a second multimodal large model, and obtaining at least one action-related element content output by the second multimodal large model.

[0095] The action-related elements may include at least one of the following: the role of the person performing the action, the action type, the associated object, and the event content. For example, in an e-commerce live broadcast, the role of the person performing the action may be the host, assistant host, or vice host.

[0096] Examples of action types include displaying a product, shouting loudly, or imitating a demonstration. Examples of associated objects include products or props that the performer touches, displays, or points to. Examples of event content include promotional guidance and emphasizing product quality.

[0097] Among them, the consideration of multiple related elements can extract more features in at least one video frame of the key video sub-segment, thereby further improving the accuracy of the determined sub-segment summary content.

[0098] Among them, the consideration of the identification content in at least one video frame allows the electronic device to consider more features, thereby further improving the accuracy of the determined action-related element content.

[0099] Step 406 : Determine the sub-segment summary content of the key video sub-segment based on the at least one action-related element content.

[0100] In the embodiment of the present disclosure, the process of the electronic device executing step 406 may be, for example, inputting at least one action-related element content into the large language model and obtaining the sub-segment summary content output by the large language model.

[0101] Among them, considering the parameter quantity and calculation accuracy of the large language model, determining the sub-segment summary content in combination with the large language model and at least one action-related element content can further improve the accuracy of the determined sub-segment summary content.

[0102] The video processing method of the embodiment of the present disclosure determines a video clip sequence of a video to be processed, as well as text content and at least one action label information corresponding to the video clips in the video clip sequence; extracts key video sub-segments from the video clips based on the text content and the at least one action label information; determines the text content corresponding to the key video sub-segments; performs video frame sampling processing on the key video sub-segments to obtain at least one video frame in the key video sub-segments; inputs the text content and the at least one video frame into a second multimodal large model to obtain at least one action-related element content output by the second multimodal large model; wherein, at least one action-related element content of the key video sub-segment is determined in combination with the second multimodal large model, and then the sub-segment summary content is determined in combination with the action-related element content, which can further improve the accuracy of the sub-segment summary content of the key video sub-segment.

[0103] The following examples are given to illustrate this. Figure 5 The figure shows the extraction diagram of key video sub-segments. Figure 5 The following steps may be included. Step 501: Obtain a video stream (i.e., the video to be processed). Step 502: Coarsely slice the video stream to obtain video segments (i.e., a sequence of video segments). Step 503: Perform automatic speech recognition (ASR) on the audio in the video segment to obtain speech subtitles (text content). Step 504: Sample the video segment to obtain at least one video frame. Step 505: Perform key point detection on the at least one sampled video frame to obtain body key points and hand key points. Step 506: Perform sliding window detection on the body key points and hand key points to perform fine-grained posture / gesture recognition, and perform structured processing to obtain action label information. Step 507: Input the action label information, speech subtitles, and at least one video frame into the MLLM (multimodal large model) to obtain the start and end timestamps of the key video sub-segments; after post-processing in combination with the start and end timestamps, obtain the highlight segment (i.e., the key video sub-segment).

[0104] like Figure 6 The following is a schematic diagram of understanding key video sub-segments. Figure 6The following steps may be included. Step 601, performing ASR processing on the audio in the highlight video clip (i.e., the key video sub-clip) to obtain voice subtitles. Step 602, performing sampling processing on the highlight video clip to obtain at least one video frame. Step 603, performing optical character recognition (OCR) on at least one video frame to obtain recognition content. Step 604, inputting the voice subtitles, at least one video frame and the recognition content into the MLLM (multimodal large model) to obtain action-related elements such as the actor, action, object, action intention, etc. Step 605, inputting the action-related element content into the large language model (Large Language Model, LLM) to obtain the segment caption (i.e., the sub-segment summary content of the key video sub-clip).

[0105] In order to implement the above embodiment, the present disclosure also provides a video processing device. Figure 7 As shown, Figure 7 7 is a schematic diagram of a fifth embodiment of the present disclosure. The video processing device 70 may include: a determination module 701 , an extraction module 702 , and a comprehension processing module 703 .

[0106] Among them, the determination module 701 is used to determine the video segment sequence of the video to be processed, as well as the text content and at least one action label information corresponding to the video segment in the video segment sequence; the extraction module 702 is used to extract key video sub-segments from the video segment based on the text content and the at least one action label information; the understanding processing module 703 is used to perform video understanding processing on the key video sub-segments to obtain the sub-segment summary content of the key video sub-segments.

[0107] As a possible implementation method of an embodiment of the present disclosure, the determination module 701 includes: a first determination unit, a recognition processing unit and an extraction processing unit; the first determination unit is used to determine the video clip sequence of the video to be processed; the recognition processing unit is used to recognize the voice content of the video clips in the video clip sequence to obtain the text content corresponding to the video clip; the extraction processing unit is used to extract action label information from the video clip to obtain at least one action label information corresponding to the video clip.

[0108] As a possible implementation method of an embodiment of the present disclosure, the extraction and processing unit is specifically used to perform human key point detection processing on at least one video frame in the video clip to obtain human key point information in the at least one video frame; and determine at least one action label information corresponding to the video clip based on the human key point information in the at least one video frame.

[0109] As a possible implementation method of an embodiment of the present disclosure, the extraction and processing unit is specifically used to input each video frame in the at least one video frame into a human posture detection model to obtain the human key point information output by the human posture detection model.

[0110] As a possible implementation of the embodiment of the present disclosure, the human body key point information includes at least one of the following: body key point information, hand key point information.

[0111] As a possible implementation method of an embodiment of the present disclosure, the first determination unit is specifically used to perform time window sliding processing on the human key point information in the at least one video frame to obtain at least one time window and a sequence of human key point information within the time window; for each time window in the at least one time window, determine the action label information according to the sequence of human key point information within the time window.

[0112] As a possible implementation manner of an embodiment of the present disclosure, the first determination unit is further specifically configured to perform feature extraction processing on the human body key point information sequence within the time window to obtain a motion trajectory feature; query a trajectory feature library based on the motion trajectory feature to obtain a first reference trajectory feature in the trajectory feature library that matches the motion trajectory feature; and determine the action label information based on the action label corresponding to the first reference trajectory feature and the start and end timestamps of the human body key point information sequence within the time window.

[0113] As a possible implementation manner of the embodiment of the present disclosure, the at least one video frame in the video clip is at least one video frame obtained by sampling the video clip.

[0114] As a possible implementation method of an embodiment of the present disclosure, the extraction module 702 includes: a first sampling processing unit, a first acquisition unit and a second determination unit; the first sampling processing unit is used to perform sampling processing on the video clip to obtain at least one video frame in the video clip; the first acquisition unit is used to input the text content corresponding to the video clip, the at least one video frame and the at least one action label information into the first multimodal large model to obtain at least one set of key start and end timestamps output by the first multimodal large model; the second determination unit is used to determine at least one key video sub-segment in the video clip based on the video clip and the at least one set of key start and end timestamps.

[0115] As a possible implementation method of an embodiment of the present disclosure, the prompt text for the key video sub-segment extraction task in the first multimodal large model indicates that the key video sub-segment needs to contain action-related semantic information; the action-related semantic information includes at least one of the following: object interaction information, character interaction information, action reinforcement information, and emotional expression information.

[0116] As a possible implementation method of an embodiment of the present disclosure, there is overlap between adjacent video segments in the video segment sequence; the device also includes: an acquisition module and a fusion processing module; the acquisition module is used to obtain at least two overlapping key video sub-segments when there is overlap between key video sub-segments in adjacent video segments; the fusion processing module is used to perform fusion processing on the at least two key video sub-segments.

[0117] As a possible implementation of the embodiment of the present disclosure, the fusion processing module is specifically used to determine a candidate video segment in the video to be processed that includes the at least two key video sub-segments; and determine the candidate video segment as the key video sub-segment after fusion processing.

[0118] As a possible implementation method of an embodiment of the present disclosure, the fusion processing module is specifically used to determine the candidate video segments in the video to be processed that include the at least two key video sub-segments; extract candidate key video sub-segments from the candidate video segments based on the text content and at least one action label information corresponding to the candidate video segments; and determine the candidate key video sub-segments as the key video sub-segments after fusion processing.

[0119] As a possible implementation method of an embodiment of the present disclosure, the understanding processing module 703 includes: a third determination unit, a second sampling processing unit, a second acquisition unit and a fourth determination unit; the third determination unit is used to determine the text content corresponding to the key video sub-segment; the second sampling processing unit is used to perform video frame sampling processing on the key video sub-segment to obtain at least one video frame in the key video sub-segment; the second acquisition unit is used to input the text content and the at least one video frame into a second multimodal large model to obtain at least one action-related element content output by the second multimodal large model; the fourth determination unit is used to determine the sub-segment summary content of the key video sub-segment based on the at least one action-related element content.

[0120] As a possible implementation method of an embodiment of the present disclosure, the understanding processing module 703 also includes: a third acquisition unit, used to perform optical symbol recognition processing on the at least one video frame to obtain the recognition content in the at least one video frame; the second acquisition unit, specifically used to input the text content, the at least one video frame and the recognition content into a second multimodal large model to obtain at least one action-related element content output by the second multimodal large model.

[0121] As a possible implementation of the embodiment of the present disclosure, the action-related element content includes at least one of the following: the role of the action performer, the action type, the associated object, and the event content.

[0122] As a possible implementation manner of the embodiment of the present disclosure, the fourth determining unit is specifically configured to input the at least one action-related element content into a large language model, and obtain the sub-segment summary content output by the large language model.

[0123] The video processing device of the embodiment of the present disclosure determines a video clip sequence of a video to be processed, as well as text content and at least one action label information corresponding to the video clips in the video clip sequence; extracts key video sub-segments from the video clips based on the text content and the at least one action label information; performs video understanding processing on the key video sub-segments to obtain sub-segment summary content of the key video sub-segments; wherein, the determination of the video clip sequence enables the video processing method to be applicable to videos of longer duration, etc.; the consideration of at least one action label information enables the video processing method to understand sudden events in the video, etc., thereby improving video processing efficiency.

[0124] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, comply with relevant laws and regulations, and do not violate public order and good morals.

[0125] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0126] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0127] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 can also be stored in the RAM 803. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0128] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0129] The computing unit 801 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the video processing method. For example, in some embodiments, the video processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the video processing method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the video processing method by any other appropriate means (e.g., by means of firmware).

[0130] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0131] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0132] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0134] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0135] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0136] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0137] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A video processing method, comprising: Determining a video segment sequence of a video to be processed, as well as text content and at least one action label information corresponding to the video segments in the video segment sequence; extracting a key video sub-segment from the video segment according to the text content and the at least one action tag information; Video understanding processing is performed on the key video sub-segment to obtain sub-segment summary content of the key video sub-segment.

2. The method according to claim 1, wherein The step of determining a video clip sequence of a video to be processed, and text content and at least one action tag information corresponding to the video clips in the video clip sequence, includes: Determining a video segment sequence of the video to be processed; Recognizing the speech content of the video segments in the video segment sequence to obtain text content corresponding to the video segments; Perform action label information extraction processing on the video clip to obtain at least one action label information corresponding to the video clip.

3. The method according to claim 2, wherein: The performing action label extraction processing on the video clip to obtain at least one action label information corresponding to the video clip includes: Performing human key point detection processing on at least one video frame in the video clip to obtain human key point information in the at least one video frame; At least one action label information corresponding to the video segment is determined according to the human body key point information in the at least one video frame.

4. The method according to claim 3, wherein: The performing human key point detection processing on at least one video frame in the video clip to obtain human key point information in the at least one video frame includes: For each video frame in the at least one video frame, the video frame is input into a human posture detection model to obtain the human key point information output by the human posture detection model.

5. The method according to claim 3 or 4, wherein: The human body key point information includes at least one of the following: body key point information, hand key point information.

6. The method according to claim 3, wherein: The determining, based on the human body key point information in the at least one video frame, at least one action label information corresponding to the video clip includes: Performing a time window sliding process on the human body key point information in the at least one video frame to obtain at least one time window and a sequence of human body key point information within the time window; For each time window in at least one time window, action label information is determined according to a human body key point information sequence within the time window.

7. The method according to claim 6, wherein: Determining the action label information according to the human body key point information sequence within the time window includes: Performing feature extraction processing on the human body key point information sequence within the time window to obtain motion trajectory features; Querying a trajectory feature library according to the motion trajectory feature to obtain a first reference trajectory feature in the trajectory feature library that matches the motion trajectory feature; The action label information is determined according to the action label corresponding to the first reference trajectory feature and the start and end timestamps of the human body key point information sequence in the time window.

8. The method according to claim 3, wherein: The at least one video frame in the video segment is at least one video frame obtained by sampling the video segment.

9. The method according to claim 1, wherein: The extracting a key video sub-segment from the video segment according to the text content and the at least one action tag information includes: performing sampling processing on the video clip to obtain at least one video frame in the video clip; Inputting the text content corresponding to the video clip, the at least one video frame, and the at least one action label information into a first multimodal large model, and obtaining at least one set of key start and end timestamps output by the first multimodal large model; At least one key video sub-segment in the video segment is determined according to the video segment and the at least one set of key start and end timestamps.

10. The method according to claim 9, wherein: Prompt text for a key video sub-segment extraction task in the first multimodal large model, indicating that the key video sub-segment needs to contain action-related semantic information; The action-related semantic information includes at least one of the following: object interaction information, character interaction information, action reinforcement information, and emotion expression information.

11. The method according to claim 1, wherein There are overlaps between adjacent video segments in the video segment sequence; the method further includes: When key video sub-segments in adjacent video segments overlap, obtaining at least two overlapping key video sub-segments; A fusion process is performed on the at least two key video sub-segments.

12. The method according to claim 11, wherein The fusing the at least two key video sub-segments includes: Determining candidate video segments in the video to be processed that include the at least two key video sub-segments; The candidate video segment is determined as the key video sub-segment after fusion processing.

13. The method according to claim 11, wherein The fusing the at least two key video sub-segments includes: Determining candidate video segments in the video to be processed that include the at least two key video sub-segments; Extracting candidate key video sub-segments from the candidate video segment according to text content and at least one action label information corresponding to the candidate video segment; The candidate key video sub-segment is determined as the fused key video sub-segment.

14. The method according to claim 1, wherein The performing video understanding processing on the key video sub-segment to obtain the sub-segment summary content of the key video sub-segment includes: Determining text content corresponding to the key video sub-segment; Performing video frame sampling processing on the key video sub-segment to obtain at least one video frame in the key video sub-segment; Inputting the text content and the at least one video frame into a second multimodal large model, and obtaining at least one action-related element content output by the second multimodal large model; The sub-segment summary content of the key video sub-segment is determined according to the at least one action-related element content.

15. The method according to claim 14, wherein The performing video understanding processing on the key video sub-segment to obtain the sub-segment summary content of the key video sub-segment further includes: performing optical symbol recognition processing on the at least one video frame to obtain recognition content in the at least one video frame; Correspondingly, inputting the text content and the at least one video frame into the second multimodal large model and obtaining at least one action-related element content output by the second multimodal large model includes: The text content, the at least one video frame, and the recognition content are input into a second multimodal large model to obtain at least one action-related element content output by the second multimodal large model.

16. The method according to claim 14 or 15, wherein: The action-related element content includes at least one of the following: the role of the action performer, the action type, the associated object, and the event content.

17. The method according to claim 14, wherein: The determining, based on the at least one action-related element content, the sub-segment summary content of the key video sub-segment includes: The at least one action-related element content is input into a large language model, and the sub-segment summary content output by the large language model is obtained.

18. A video processing device, comprising: A determination module, configured to determine a video segment sequence of a video to be processed, as well as text content and at least one action label information corresponding to the video segments in the video segment sequence; an extraction module, configured to extract a key video sub-segment from the video segment based on the text content and the at least one action tag information; The understanding processing module is configured to perform video understanding processing on the key video sub-segment to obtain a sub-segment summary content of the key video sub-segment.

19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 17.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 17.

21. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 17.