Video processing method and device, electronic equipment and computer readable storage medium
By calculating the segmentation and category probability of video frames and selecting video frames above the segmentation threshold as endpoints, the problem of inaccurate video segmentation is solved, and accurate segmentation and content recognition in video processing are achieved.
Patent Information
- Application Number
- CN202511421709.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies struggle to effectively separate different types of video segments from the video to be processed, affecting the accuracy and efficiency of subsequent video processing.
By acquiring video features, calculating the segmentation probability and category probability of video frames, selecting video frames above the segmentation threshold as endpoints, extracting video segments, and determining the category of the video segment based on the category probability of the frame.
It achieves precise video segmentation, improves the accuracy and efficiency of video processing, and can better identify and process different types of video content.
Smart Images

Figure CN120980295A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to a video processing method and apparatus, electronic device and computer-readable storage medium. Background Technology
[0002] By segmenting the video to be processed, video segments belonging to different categories can be separated. These segments can then be used for subsequent processing, such as video splicing, video editing, and video compositing. Therefore, the method of segmenting the video to be processed is of great importance. Summary of the Invention
[0003] This application provides a video processing method and apparatus, an electronic device, and a computer-readable storage medium.
[0004] Firstly, a video processing method is provided, the method comprising:
[0005] Get the video to be processed;
[0006] Based on the video features of the video to be processed, the segmentation probability of each video frame of the video to be processed is obtained. The segmentation probability represents the probability that a video frame is the endpoint of a video segment. The endpoint includes the start frame and the end frame.
[0007] Select two or more video frames from the video to be processed whose segmentation probability is greater than or equal to the segmentation threshold. The two or more video frames include a first video frame and a second video frame with adjacent timestamps.
[0008] The target video segment is obtained by extracting the video segment between the first video frame and the second video frame from the video to be processed.
[0009] In conjunction with any embodiment of this application, before obtaining the target video segment, the method further includes:
[0010] Based on the video features of the video to be processed, the category probability of each video frame from the first video frame to the second video frame is obtained, and the category probability represents the probability of the video frame belonging to a category.
[0011] After obtaining the target video segment, the method further includes:
[0012] The category of the target video segment is obtained based on the category probability of each video frame from the first video frame to the second video frame.
[0013] In conjunction with any embodiment of this application, obtaining the category of the target video segment based on the category probabilities of each video frame from the first video frame to the second video frame includes:
[0014] Based on the category probabilities of each video frame from the first video frame to the second video frame, the category probabilities of each category in the target video segment are summed to obtain the category probability sum of each category;
[0015] The category corresponding to the maximum value of the sum of the category probabilities is determined as the category of the target video segment.
[0016] In conjunction with any embodiment of this application, obtaining the segmentation probability of each video frame of the video to be processed based on the video features of the video to be processed includes:
[0017] Based on the video features, determine the image differences between two adjacent video frames in the video to be processed;
[0018] Based on the image differences, the segmentation probability of two adjacent video frames is determined, and the segmentation probability of two adjacent video frames is positively correlated with the image differences.
[0019] In any embodiment of this application, the video features include color features;
[0020] The step of determining the image differences between two adjacent video frames in the video to be processed based on the video features includes:
[0021] Based on the color features, determine the color difference between two adjacent video frames in the video to be processed;
[0022] The image difference between two adjacent video frames in the video to be processed is obtained based on the color difference between the two adjacent video frames in the video to be processed, and the image difference is positively correlated with the color difference.
[0023] In conjunction with any embodiment of this application, obtaining the segmentation probability of each video frame of the video to be processed based on the video features of the video to be processed includes:
[0024] Based on the video features, determine the correlation between two non-adjacent video frames in the video to be processed;
[0025] Based on the correlation, the segmentation probability of each video frame between the two non-adjacent video frames is obtained, and the segmentation probability of the video frames between the two non-adjacent video frames is negatively correlated with the correlation.
[0026] In conjunction with any embodiment of this application, before determining the correlation between two non-adjacent video frames in the video to be processed based on the video features, the method further includes:
[0027] Determine the resolution of each video frame in the video to be processed;
[0028] The step of determining the correlation between two non-adjacent video frames in the video to be processed based on the video features includes:
[0029] If there are video frames with a resolution less than or equal to a resolution threshold in two adjacent frames of the video to be processed, the correlation between two non-adjacent video frames in the video to be processed is determined.
[0030] In conjunction with any embodiment of this application, before obtaining the segmentation probability of each video frame of the video to be processed based on the video features of the video to be processed, the method further includes:
[0031] The video to be processed is processed using a neural network to obtain the video features of the video to be processed.
[0032] Secondly, a video processing apparatus is provided, the apparatus comprising:
[0033] The acquisition unit is used to acquire the video to be processed.
[0034] The first processing unit is configured to obtain the segmentation probability of each video frame of the video to be processed based on the video features of the video to be processed. The segmentation probability represents the probability that a video frame is the endpoint of a video segment, and the endpoint includes a start frame and an end frame.
[0035] The second processing unit is used to select two or more video frames from the video to be processed, the video frames having a segmentation probability greater than or equal to the segmentation threshold, the two or more video frames including a first video frame and a second video frame with adjacent timestamps.
[0036] The cropping unit is used to crop a video segment from the video to be processed between the first video frame and the second video frame to obtain a target video segment.
[0037] In conjunction with any embodiment of this application, the first processing unit is further configured to obtain the category probability of each video frame from the first video frame to the second video frame based on the video features of the video to be processed, wherein the category probability represents the probability of the video frame belonging to a category.
[0038] The second processing unit is further configured to obtain the category of the target video segment based on the category probability of each video frame from the first video frame to the second video frame.
[0039] In conjunction with any embodiment of this application, the second processing unit is configured to:
[0040] Based on the category probabilities of each video frame from the first video frame to the second video frame, the category probabilities of each category in the target video segment are summed to obtain the category probability sum of each category;
[0041] The category corresponding to the maximum value of the sum of the category probabilities is determined as the category of the target video segment.
[0042] In conjunction with any embodiment of this application, the first processing unit is configured to:
[0043] Based on the video features, determine the image differences between two adjacent video frames in the video to be processed;
[0044] Based on the image differences, the segmentation probability of two adjacent video frames is determined, and the segmentation probability of two adjacent video frames is positively correlated with the image differences.
[0045] In any embodiment of this application, the video features include color features;
[0046] The first processing unit is configured to:
[0047] Based on the color features, determine the color difference between two adjacent video frames in the video to be processed;
[0048] The image difference between two adjacent video frames in the video to be processed is obtained based on the color difference between the two adjacent video frames in the video to be processed, and the image difference is positively correlated with the color difference.
[0049] In conjunction with any embodiment of this application, the first processing unit is configured to:
[0050] Based on the video features, determine the correlation between two non-adjacent video frames in the video to be processed;
[0051] Based on the correlation, the segmentation probability of each video frame between the two non-adjacent video frames is obtained, and the segmentation probability of the video frames between the two non-adjacent video frames is negatively correlated with the correlation.
[0052] In conjunction with any embodiment of this application, the first processing unit is further configured to:
[0053] Determine the resolution of each video frame in the video to be processed;
[0054] If there are video frames with a resolution less than or equal to a resolution threshold in two adjacent frames of the video to be processed, the correlation between two non-adjacent video frames in the video to be processed is determined.
[0055] In any embodiment of this application, the first processing unit is further configured to process the video to be processed using a neural network to obtain the video features of the video to be processed.
[0056] Thirdly, an electronic device is provided, characterized in that it includes: a processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein, when the processor executes the computer instructions, the electronic device performs a method as described in the first aspect above and any possible implementation thereof.
[0057] Fourthly, another electronic device is provided, comprising: a processor, a transmitting device, an input device, an output device, and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein, when the processor executes the computer instructions, the electronic device performs the method as described in the first aspect above and any possible implementation thereof.
[0058] Fifthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, the computer program including program instructions that, when executed by a processor, cause the processor to perform a method as described in the first aspect above and any possible implementation thereof.
[0059] In a sixth aspect, a computer program product is provided, the computer program product comprising a computer program or instructions, wherein, when the computer program or instructions are executed on a computer, the computer performs the method described in the first aspect and any possible implementation thereof.
[0060] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application.
[0061] In this embodiment, when the video processing device acquires a video to be processed, it obtains the segmentation probability of each video frame of the video to be processed based on the video characteristics of the video to be processed. The segmentation probability represents the probability that a video frame is an endpoint of a video segment, and the endpoints include a start frame and an end frame. Then, it selects two or more video frames from the video to be processed whose segmentation probability is greater than or equal to a segmentation threshold, thus selecting two or more video frames that can serve as endpoints. Finally, it uses two frames with adjacent timestamps from the two or more video frames as the start frame and the end frame, respectively, to extract the target video segment from the video to be processed, thereby achieving the segmentation of the video to be processed. Attached Figure Description
[0062] To more clearly illustrate the technical solutions in the embodiments of this application or the background art, the accompanying drawings used in the embodiments of this application or the background art will be described below.
[0063] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0064] Figure 1 A flowchart illustrating a video processing method provided in an embodiment of this application;
[0065] Figure 2 A schematic diagram of a neural network structure provided in an embodiment of this application;
[0066] Figure 3 A schematic diagram of the structure of a holed deep neural network module provided in an embodiment of this application;
[0067] Figure 4 This is a schematic diagram of the structure of a color similarity calculation module provided in an embodiment of this application;
[0068] Figure 5 This is a schematic diagram of the structure of a video processing apparatus provided in an embodiment of this application;
[0069] Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0070] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0071] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0072] In this document, the term "embodiment" means that a particular video feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0073] The execution subject of this application embodiment is a video processing device, which can be any electronic device capable of executing the technical solutions disclosed in the method embodiments of this application. Optionally, the video processing device can be one of the following: a mobile phone, a computer, a tablet computer, or a wearable smart device.
[0074] It should be understood that the method embodiments of this application can also be implemented by a processor executing computer program code. The embodiments of this application are described below with reference to the accompanying drawings. Please refer to... Figure 1 , Figure 1 This is a flowchart illustrating a video processing method provided in an embodiment of this application.
[0075] 101. Obtain the video to be processed.
[0076] In this embodiment, the video to be processed can be any video. The video to be processed can be an offline video or an online video. Offline videos can be videos captured by a camera or a mobile smart device. Online videos can be videos captured in real-time by a camera. The video to be processed can be a video containing any content; for example, it could be a basketball game video, a dance video, or a video that includes both dance and basketball gameplay.
[0077] In one implementation of acquiring a video to be processed, the video processing device receives the video to be processed from user input via an input component. The input component includes at least one of the following: a keyboard, a mouse, a touchscreen, a touchpad, or an audio input device.
[0078] In another implementation of acquiring the video to be processed, the video processing device receives the video to be processed sent by the terminal. The terminal can be any of the following: mobile phone, computer, tablet computer, or server.
[0079] In another implementation of acquiring the video to be processed, the video processing device obtains the video to be processed by downloading the video from the Internet.
[0080] In another implementation of acquiring the video to be processed, the video processing device has a communication connection with the camera, and the camera acquires the video captured by the camera as the video to be processed through the communication connection.
[0081] 102. Based on the video features of the video to be processed, the segmentation probability of each video frame of the video to be processed is obtained.
[0082] In this embodiment of the application, the segmentation probability characterizes the probability that a video frame is the endpoint of a video segment, wherein the endpoint includes the start frame and the end frame.
[0083] Optionally, endpoints are used to distinguish video segments with different categories. Specifically, in the video to be processed, endpoints are the boundaries between video segments with different categories. For example, if the video to be processed includes frames 1, 2, 3, 4, 5, and 6, where frames 1, 3, 4, and 6 are all endpoints, then based on these endpoints, the video to be processed can be divided into a first video segment and a second video segment. The first video segment includes frames 1, 2, and 3, and the second video segment includes frames 4, 5, and 6. In this case, the first video segment and the second video segment have different categories.
[0084] In this embodiment of the application, the category of a video segment represents the category of the content of the video segment. For example, if the category of the first video is dance and the category of the second video is food, then the content of the first video is related to dance and the content of the second video is related to food.
[0085] In this embodiment, the video features of the video to be processed carry feature information of the video frames in the video to be processed, as well as timestamp information of the video frames in the video to be processed. The feature information of the video frames includes at least one of the following: texture information of the video frame, color information of the video frame, shape information of objects within the video frame, brightness information of the video frame, and spatial relationship information of different objects within the video frame. The timestamp of the video frame can be determined based on the timestamp information of the video frame, wherein the timestamp of the video frame represents the playback time of the video frame in the video to be processed.
[0086] For example, the video to be processed includes a first frame, a second frame, and a third frame, where the timestamp of the first frame is t1, the timestamp of the second frame is t2, and the timestamp of the third frame is t3. In this case, the video features of the video to be processed carry the feature information of the first frame, the feature information of the second frame, the feature information of the third frame, the timestamp of the information of the first frame is t1, the timestamp of the feature information of the second frame is t2, and the timestamp of the feature information of the third frame is t3.
[0087] In one possible implementation, the video processing device obtains the video features of the video to be processed by performing feature extraction processing on the video to be processed.
[0088] When the video processing device obtains the video features of the video to be processed, it can determine the feature information of each video frame based on the video features of the video to be processed, and then use the feature information of each video frame to determine the segmentation probability of each video frame in the video to be processed.
[0089] 103. Select two or more video frames from the above-mentioned videos to be processed, whose segmentation probability is greater than or equal to the segmentation threshold.
[0090] In this embodiment, the higher the segmentation probability of a video frame, the higher the probability that the video frame is an endpoint. Therefore, to improve the accuracy of identifying endpoints from the video to be processed, video frames with high segmentation probabilities can be selected as endpoints. The video processing device determines whether the segmentation probability of a video frame is high or low based on a segmentation threshold. Specifically, if the segmentation probability of a video frame is greater than or equal to the segmentation threshold, it indicates that the segmentation probability of the video frame is high, and thus the video frame is determined to be an endpoint; if the segmentation probability of a video frame is less than the segmentation threshold, it indicates that the segmentation probability of the video frame is low, and thus the video frame is determined not to be an endpoint.
[0091] Based on the segmentation probability of each video frame in the video to be processed, the video processing device can select two or more video frames whose segmentation probability is greater than or equal to the segmentation threshold from the video to be processed, that is, select two or more video frames that can be used as endpoints from the video to be processed.
[0092] The video processing device selects two or more video frames from the video to be processed that can serve as endpoints, including a first video frame and a second video frame with adjacent timestamps. That is, the first video frame and the second video frame are any two frames with adjacent timestamps among two or more video frames. It should be understood that the fact that the first video frame and the second video frame have adjacent timestamps among two or more video frames does not mean that the first video frame and the second video frame have adjacent timestamps in the video to be processed.
[0093] For example, the video to be processed includes a first frame, a second frame, a third frame, a fourth frame, a fifth frame, and a sixth frame. The video processing device selects two or more video frames from the video to be processed that can serve as endpoints, including the first frame, the third frame, the fourth frame, and the sixth frame. In this case, the first frame and the third frame are two frames with adjacent timestamps among two or more video frames, the third frame and the fourth frame are two frames with adjacent timestamps among two or more video frames, and the third frame and the sixth frame are two frames with adjacent timestamps among two or more video frames.
[0094] 104. Extract the video segment between the first video frame and the second video frame from the video to be processed to obtain the target video segment.
[0095] Since both the first and second video frames are endpoints, the video processing device can obtain the target video segment by extracting the video segment between the first and second video frames from the video to be processed.
[0096] For example, the video to be processed includes a first frame, a second frame, a third frame, a fourth frame, a fifth frame, and a sixth frame. The video processing device selects two or more video frames from the video to be processed that can serve as endpoints, including the first frame, the third frame, the fourth frame, and the sixth frame. When the first video frame is the first frame and the second video frame is the third frame, the target video segment extracted by the video processing device from the video to be processed includes the first frame, the second frame, and the third frame. When the first video frame is the third frame and the second video frame is the fourth frame, the target video segment extracted by the video processing device from the video to be processed includes the third frame and the fourth frame. When the first video frame is the fourth frame and the second video frame is the sixth frame, the target video segment extracted by the video processing device from the video to be processed includes the fourth frame, the fifth frame, and the sixth frame.
[0097] In this embodiment, when the video processing device acquires a video to be processed, it obtains the segmentation probability of each video frame of the video to be processed based on the video characteristics of the video to be processed. The segmentation probability represents the probability that a video frame is an endpoint of a video segment, and the endpoints include a start frame and an end frame. Then, it selects two or more video frames from the video to be processed whose segmentation probability is greater than or equal to a segmentation threshold, thus selecting two or more video frames that can serve as endpoints. Finally, it uses two frames with adjacent timestamps from the two or more video frames as the start frame and the end frame, respectively, to extract the target video segment from the video to be processed, thereby achieving the segmentation of the video to be processed.
[0098] As an optional implementation, the video processing device further performs the following steps before obtaining the target video segment:
[0099] 201. Based on the video features of the video to be processed, obtain the category probability of each video frame from the first video frame to the second video frame.
[0100] In this embodiment, the category probability represents the probability of a video frame belonging to a category, where the category of the video frame represents the category of the video frame's content. For example, if the category of a video frame is dance, then the content of that video frame includes dance.
[0101] In this embodiment, each video frame from the first video frame to the second video frame includes the first video frame, the second video frame, and the video frame in the video to be processed located between the first video frame and the second video frame. That is, each video frame from the first video frame to the second video frame is a video frame in the target video segment. For example, if the video to be processed includes a first frame, a second frame, a third frame, a fourth frame, a fifth frame, and a sixth frame, and the first video frame is the first frame and the second video frame is the third frame, then each video frame from the first video frame to the second video frame includes: the first frame, the second frame, and the third frame.
[0102] Since the video features of the video to be processed include the feature information of each video frame, the video processing device determines the feature information of each video frame from the first video frame to the second video frame based on the video features of the video to be processed, and then can use the feature information of each video frame from the first video frame to the second video frame to determine the category probability of each video frame from the first video frame to the second video frame.
[0103] For example, the video to be processed includes the first frame, the second frame, the third frame, the fourth frame, the fifth frame, and the sixth frame. Each video frame from the first video frame to the second video frame includes the first frame, the second frame, and the third frame. Then, the class probability of each video frame from the first video frame to the second video frame includes the class probability of the first frame, the class probability of the second frame, and the class probability of the third frame.
[0104] In this embodiment, after obtaining the target video segment, the video processing device further performs the following steps:
[0105] 202. Based on the above-mentioned category probabilities of each video frame in the target video segment, the category of the target video segment is obtained.
[0106] The video processing device determines the category of each video frame based on its category probability, and then obtains the category of the target video segment based on the category of each video frame. In one possible implementation, the video processing device determines the category of each video frame based on its category probability. Based on the categories of each video frame, the category with the most corresponding video frames is determined as the category of the target video segment.
[0107] For example, the target video segment includes video frame a, video frame b, and video frame c. Based on the category probability of each video frame, the video processing device determines that video frame a is categorized as dance, video frame b as dance, and video frame c as food. At this point, there are 2 video frames categorized as dance and 1 video frame categorized as food. Therefore, the video processing device determines the target video to be categorized as dance.
[0108] In this embodiment, the video processing device obtains the category probability of each video frame from the first video frame to the second video frame based on the video characteristics of the video to be processed. Then, when the target video is extracted from the video to be processed, the category of the target video can be obtained based on the category probability of each video frame from the first video frame to the second video frame.
[0109] As an optional implementation, the video processing device performs the following steps during step 202:
[0110] 301. Based on the category probabilities of each video frame from the first video frame to the second video frame, sum the category probabilities of each category in the target video segment to obtain the category probability sum of each category.
[0111] 302. Determine the category corresponding to the maximum sum of the above category probabilities as the category of the above target video segment.
[0112] For example, the target video segment includes video frames a, b, and c. The category probabilities for video frame a include a probability of 0.8 for dance and a probability of 0.2 for food. The category probabilities for video frame b include a probability of 0.6 for dance and a probability of 0.4 for food. The category probabilities for video frame c include a probability of 0.5 for dance and a probability of 0.5 for food. In this case, the target video segment includes categories for dance and food. The sum of the category probabilities for dance is 0.8 + 0.6 + 0.5 = 1.9, and the sum of the category probabilities for food is 0.2 + 0.4 + 0.5 = 1.1. Since the sum of the category probabilities for dance is greater than that for food, the sum of the category probabilities for dance is the maximum sum of category probabilities. Therefore, the category corresponding to the maximum sum of category probabilities is dance, meaning the category of the target video segment is dance.
[0113] In this embodiment, the video processing device sums the category probabilities of each category in the target video segment based on the category probabilities of each video frame from the first video frame to the second video frame, obtaining a sum of category probabilities for each category. Then, the category corresponding to the maximum value of the sum of category probabilities is determined as the category of the target video segment, thereby achieving the determination of the category of the target video segment based on the category probabilities of each video frame in the target video segment.
[0114] As an optional implementation, the video processing device performs the following steps during step 102:
[0115] 401. Based on the above video features, determine the image differences between two adjacent video frames in the video to be processed.
[0116] In this embodiment of the application, image differences represent differences in image content. For example, the image differences between video frame a and video frame b represent differences in the semantics of video frame a and video frame b.
[0117] Optionally, the image difference is numerical, meaning the video processing device quantifies the differences in image content between two adjacent video frames in the video to be processed based on video features, thus obtaining the image difference. In one possible implementation, the larger the numerical value of the image difference, the greater the difference in the image content.
[0118] In this embodiment of the application, two adjacent video frames in the video to be processed are two frames with adjacent timestamps in the video to be processed. For example, the video to be processed includes a first frame, a second frame, and a third frame. In this case, two adjacent video frames in the video to be processed can be the first frame and the second frame, or they can be the second frame and the third frame.
[0119] Since video features carry feature information of each video frame in the video to be processed, the video processing device can determine the feature information of two adjacent video frames in the video to be processed based on the video features, and then use the feature information of two adjacent video frames in the video to be processed to determine the image differences between two adjacent video frames in the video to be processed.
[0120] 402. Based on the above image differences, determine the above segmentation probability of the above two adjacent video frames.
[0121] The greater the difference between two adjacent video frames, the greater the probability that these two video frames belong to different categories of video segments. Therefore, the video processing device can determine the segmentation probability of two adjacent video frames based on the difference between their images.
[0122] In this embodiment, the segmentation probability of two adjacent video frames is positively correlated with the image difference between the two adjacent video frames. That is, the greater the image difference between two adjacent video frames, the greater the segmentation probability of the two adjacent video frames.
[0123] In this embodiment, the video processing device determines the image differences between two adjacent video frames in the video to be processed based on video features. Then, when the segmentation probability of two adjacent video frames is positively correlated with the image differences between them, the segmentation probability of two adjacent video frames is determined based on the image differences. This allows the segmentation probability of each video frame in the video to be processed to be determined separately, and also improves the accuracy of the segmentation probability of each video frame in the video to be processed.
[0124] As an optional implementation, the video features include color features. During step 401, the video processing apparatus performs the following steps:
[0125] 501. Based on the above color characteristics, determine the color difference between two adjacent video frames in the above video to be processed.
[0126] In this embodiment, color difference represents the difference in image colors. Optionally, the color difference is a numerical value, meaning the video processing device quantifies the difference in image colors between two adjacent video frames in the video to be processed based on color features to obtain the color difference. In one possible implementation, the larger the numerical value of the color difference, the greater the difference in image colors it represents.
[0127] 502. Based on the color difference between two adjacent video frames in the video to be processed, the image difference between two adjacent video frames in the video to be processed is obtained.
[0128] The greater the color difference between two adjacent video frames, the greater the difference in the image content of the two video frames. Therefore, the video processing device can determine the image difference between two adjacent video frames based on the color difference between them.
[0129] In this embodiment, the image difference between two adjacent video frames is positively correlated with the color difference between two adjacent video frames; that is, the greater the color difference between two adjacent video frames, the greater the image difference between two adjacent video frames.
[0130] In this embodiment, the video processing device determines the color difference between two adjacent video frames in the video to be processed based on video features. Then, if the image difference between two adjacent video frames and the color difference between two adjacent video frames are positively correlated, the image difference between two adjacent video frames can be determined based on the color difference, thereby improving the accuracy of the image difference between two adjacent video frames.
[0131] As an optional implementation, the video processing device performs the following steps during step 102:
[0132] 601. Based on the above video features, determine the correlation between two non-adjacent video frames in the above video to be processed.
[0133] In this embodiment, two non-adjacent video frames in the video to be processed are two video frames with non-adjacent timestamps. For example, if the video to be processed includes a first frame, a second frame, and a third frame, then the two non-adjacent video frames in the video to be processed are the first frame and the third frame.
[0134] In this embodiment, the correlation between two non-adjacent video frames represents the correlation between the image content of the two non-adjacent video frames. Optionally, the correlation between two non-adjacent video frames is the similarity between the two non-adjacent video frames. For example, let's say two non-adjacent video frames are video frame a and video frame b. If the image content of both video frame a and video frame b is dance, then video frame a and video frame b have a high correlation. If the image content of video frame a is dance and the image content of video frame b is food, then video frame a and video frame b have a low correlation.
[0135] Since video features carry feature information of each video frame in the video to be processed, the video processing device can determine the feature information of two non-adjacent video frames in the video to be processed based on the video features, and then use the feature information of two non-adjacent video frames in the video to be processed to determine the correlation between two non-adjacent video frames in the video to be processed.
[0136] 602. Based on the above correlation, the segmentation probability of each video frame between the two non-adjacent video frames is obtained.
[0137] The greater the correlation between two non-adjacent video frames, the lower the probability that these two frames belong to different video segments, meaning a higher probability that they belong to the same video segment. In other words, the higher the probability that video frames located between these two non-adjacent frames belong to the same video segment, and the lower the probability of segmenting video frames located between these two non-adjacent frames. Therefore, the video processing device can determine the segmentation probability of each video frame between two non-adjacent video frames based on their correlation. Specifically, when the segmentation probability of video frames between two non-adjacent video frames is negatively correlated with the aforementioned correlation, the video processing device obtains the segmentation probability of each video frame between two non-adjacent video frames based on their correlation.
[0138] In this embodiment, the video processing apparatus determines the correlation between two non-adjacent video frames in the video to be processed based on the aforementioned video features. Then, when the segmentation probability of video frames between two non-adjacent video frames is negatively correlated with the aforementioned correlation, the segmentation probability of each video frame between two non-adjacent video frames can be obtained based on the correlation between the two non-adjacent video frames, thereby improving the accuracy of the segmentation probability of each video frame between two non-adjacent video frames.
[0139] As an optional implementation, the video processing device further performs the following steps before performing step 601:
[0140] 701. Determine the resolution of each video frame in the above-mentioned video to be processed.
[0141] After step 701 is completed, the video processing device performs the following steps during step 601:
[0142] 702. If there are video frames with a resolution less than or equal to the resolution threshold in two adjacent frames of the above-mentioned video to be processed, determine the correlation between two non-adjacent video frames in the above-mentioned video to be processed.
[0143] The sharpness of a video frame affects the accuracy of its feature information. Specifically, there is a positive correlation between low video frame sharpness and the accuracy of its feature information; that is, the lower the video frame sharpness, the lower the accuracy of its feature information. Therefore, when two adjacent video frames contain a low-sharpness frame, determining the segmentation probability of those two frames based on their feature information will result in low accuracy in the segmentation probability.
[0144] In this embodiment of the application, in order to improve the accuracy of the segmentation probability, the video processing device determines the correlation between two non-adjacent video frames in the video to be processed when there are video frames with low resolution in two adjacent frames. Thus, the segmentation probability of each video frame between two non-adjacent video frames can be obtained based on the correlation between the two non-adjacent video frames.
[0145] In this embodiment, the video processing device determines whether the clarity of a video frame is low or high based on a clarity threshold. Specifically, if the clarity of a video frame is less than or equal to the clarity threshold, it indicates that the video frame has low clarity; if the clarity of a video frame is greater than the clarity threshold, it indicates that the video frame has high clarity. Therefore, when there are video frames with clarity less than or equal to the clarity threshold in two adjacent frames of the video to be processed, the video processing device determines the correlation between two non-adjacent video frames in the video to be processed. Based on the correlation between the two non-adjacent video frames, the device can obtain the segmentation probability of each video frame between the two non-adjacent video frames, thereby improving the accuracy of the segmentation probability of each video frame between the two non-adjacent video frames.
[0146] Steps 401 and 402, and steps 601 and 602, respectively provide two different implementation methods for determining the segmentation probability of each video frame in the video to be processed. As an optional implementation method, the video processing apparatus combines the two implementation methods for determining the segmentation probability provided in steps 401 and 402, and steps 601 and 602, to determine the segmentation probability of each video frame in the video to be processed.
[0147] In one possible implementation, the video processing device, after determining the image difference between two adjacent video frames in the video to be processed and the correlation between two non-adjacent video frames in the video to be processed, determines the segmentation probability of video frames in the video to be processed based on the image difference and the correlation, wherein the segmentation probability of video frames in the video to be processed is positively correlated with the image difference and negatively correlated with the correlation.
[0148] For example, the video to be processed includes a first frame, a second frame, and a third frame. The video processing device determines the image difference between the first frame and the second frame, as well as the correlation between the first frame and the third frame. Then, based on the image difference and the correlation, it determines the segmentation probability of the second frame, wherein the segmentation probability of the second frame is positively correlated with the image difference and negatively correlated with the correlation.
[0149] In this embodiment, the video processing device combines the two implementation methods for determining the segmentation probability provided in steps 401 and 402, and steps 601 and 602, to determine the segmentation probability of each video frame in the video to be processed, thereby improving the accuracy of the segmentation probability of each video frame.
[0150] Optionally, the video processing device combines the two implementation methods for determining the segmentation probability provided in steps 401 and 402, and steps 601 and 602. When determining the segmentation probability of each video frame in the video to be processed, steps 501 and 502 are executed during the execution of step 401 to further improve the accuracy of the segmentation probability of each video frame.
[0151] Optionally, the video processing device combines the two implementation methods for determining the segmentation probability provided in steps 401 and 402, and steps 601 and 602. When determining the segmentation probability of each video frame in the video to be processed, step 701 is executed before step 601, and step 702 is executed during the execution of step 601 after step 701 is completed, so as to further improve the accuracy of the segmentation probability of each video frame.
[0152] Optionally, the video processing device combines the two implementation methods for determining the segmentation probability provided in steps 401 and 402, and steps 601 and 602. When determining the segmentation probability of each video frame in the video to be processed, steps 501 and 502 are executed during the execution of step 401, and step 701 is executed before the execution of step 601. After the execution of step 701, step 702 is executed during the execution of step 601, thereby further improving the accuracy of the segmentation probability of each video frame.
[0153] As an optional implementation, before performing step 102, the video processing device obtains the video features of the video to be processed by performing the following steps: 801. Processing the video to be processed using a neural network to obtain the video features of the video to be processed.
[0154] In this embodiment, the neural network has the ability to extract features from the video. Optionally, Figure 2 A schematic diagram of the neural network structure, such as Figure 2 As shown, the neural network includes: a dilated deepconvolution neural network (DDCNN) module, an average pooling module, a color similarities calculation module, a concat layer, a temporal self-attention module, a fully connected layer (Dense), and a sigmoid function.
[0155] Figure 3 This is a schematic diagram of the structure of a holed deep neural network module, such as... Figure 3 As shown, DDCNN includes four 1×3×3 convolutional kernels (Conv) and four 3×1×1 dilated convolutional kernels (Dilation). It should be understood that both the 1×3×3 convolutional kernel and the 3×1×1 dilated convolutional kernel are three-dimensional convolutional kernels. The size of a three-dimensional convolutional kernel can be represented as T×H×W, where T is the temporal dimension, H is the height, and W is the width.
[0156] Four 1×3×3 convolutional kernels process the input data of DDCNN. Then, four 3×1×1 dilated convolutional kernels process the output data of the four 1×3×3 kernels. Specifically, one 3×1×1 dilated convolutional kernel processes the output data of one 1×3×3 kernel, and any two different 3×1×1 dilated convolutional kernels process different data. Finally, the output data of DDCNN is obtained by concatenating the output data of the four 3×1×1 dilated convolutional kernels. Figure 3 Output in the middle.
[0157] Figure 4 This is a schematic diagram of the structure of a color similarity calculation module, such as... Figure 4As shown, four global average pooling modules process the input data of the color similarity calculation module, and then the output data of the four global average pooling modules are concatenated. A fully connected layer (Dense) then processes the concatenated data. The pre-similarity calculation module (Cosine Similarities) calculates the color similarity from the output data of the fully connected layer, and finally, the fully connected layer (Dense) processes the color similarity to obtain the output data of the color similarity calculation module (i.e.,...). Figure 4 (Output in the text). It should be understood that the output data of the color similarity calculation module is color similarity, specifically, the color similarity between two video frames. This color similarity is negatively correlated with the aforementioned color difference, meaning that the color difference can be determined based on this color similarity.
[0158] The temporal self-attention module can determine the correlation between two non-adjacent video frames by executing step 601 above, thereby achieving cross-frame segmentation probability determination in the temporal dimension. Optionally, the temporal self-attention module can be expressed by the following formula:
[0159]
[0160] In this context, Q, K, and V represent features of video frames in the video to be processed. It should be understood that Q, K, and V can be the same, meaning they can be features of the same video frame. Any two of Q, K, and V can also be the same; that is, Q and K can be features of the same video frame, but Q and K are both different from V; or Q and V can be features of the same video frame, but Q and V are both different from K; or K and V can be features of the same video frame, but K and V are both different from Q. Q, K, and V can also be features of three different video frames in the video to be processed. k Let Q be the number of channels, and softmax(·) denotes the logistic regression (softmax) function.
[0161] Input the video to be processed to Figure 2 After the neural network shown, Figure 2The neural network shown processes the video to be processed. The output of the temporal self-attention module is the video feature of the video to be processed. Then, the video features are processed through two different recognition branches (each including a fully connected layer and a sigmoid function), outputting the segmentation probability and class probability of each video frame in the video to be processed, respectively. In other words, this neural network can sequentially perform feature extraction processing to obtain the video features of the video to be processed, and extract target video segments from the video to be processed based on the video features, and determine the class of the target video segments, thereby reducing the cost of extracting target video segments from the video to be processed and the cost of determining the class of the target video segments.
[0162] Optionally, the neural network is trained before feature extraction from the video to be processed. The training process includes: processing the training video using the neural network to obtain video features; processing the video features of the training video through two different recognition branches to obtain the predicted segmentation probability and class probability of each video frame in the training video; segmenting the training video according to the predicted segmentation probability of each video frame to obtain at least one segmented video segment; determining the class of each segmented video segment according to the class probability of each video frame in the training video, and then using the class of the segmented video segment as the predicted class of each video frame in that segmented video segment, thereby obtaining the predicted class of each video frame in the training video.
[0163] The first difference between the predicted segmentation probability and the true segmentation probability of each video frame in the training video is determined. Optionally, this first difference can be determined using the cross-entropy loss function. The second difference between the predicted class and the true class of each video frame in the training video is also determined. Optionally, this second difference can be determined using the cross-entropy loss function. Based on the first and second differences, a training loss is obtained, where both the first and second differences are positively correlated with the training loss. Based on the training loss, the parameters of the neural network are adjusted until the training loss converges, completing the training of the neural network.
[0164] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0165] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, while using clear signs / information to inform users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, personal information processing may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0166] The methods of the embodiments of this application have been described in detail above, and the apparatus of the embodiments of this application is provided below.
[0167] Please see Figure 5 , Figure 5 This is a schematic diagram of a video processing device provided in an embodiment of this application. The video processing device 1 includes: an acquisition unit 11, a first processing unit 12, a second processing unit 13, and a cropping unit 14. Specifically:
[0168] Acquisition unit 11 is used to acquire the video to be processed;
[0169] The first processing unit 12 is configured to obtain the segmentation probability of each video frame of the video to be processed based on the video features of the video to be processed. The segmentation probability represents the probability that a video frame is the endpoint of a video segment. The endpoint includes a start frame and an end frame.
[0170] The second processing unit 13 is used to select two or more video frames from the video to be processed, the video frames having a segmentation probability greater than or equal to the segmentation threshold, the two or more video frames including a first video frame and a second video frame with adjacent timestamps.
[0171] The cropping unit 14 is used to crop a video segment from the video to be processed between the first video frame and the second video frame to obtain a target video segment.
[0172] In any embodiment of this application, the first processing unit 12 is further configured to obtain the category probability of each video frame from the first video frame to the second video frame based on the video features of the video to be processed, wherein the category probability represents the probability of the video frame belonging to a category.
[0173] The second processing unit 13 is further configured to obtain the category of the target video segment based on the category probability of each video frame from the first video frame to the second video frame.
[0174] In any embodiment of this application, the second processing unit 13 is configured to:
[0175] Based on the category probabilities of each video frame from the first video frame to the second video frame, the category probabilities of each category in the target video segment are summed to obtain the category probability sum of each category;
[0176] The category corresponding to the maximum value of the sum of the category probabilities is determined as the category of the target video segment.
[0177] In any embodiment of this application, the first processing unit 12 is configured to:
[0178] Based on the video features, determine the image differences between two adjacent video frames in the video to be processed;
[0179] Based on the image differences, the segmentation probability of two adjacent video frames is determined, and the segmentation probability of two adjacent video frames is positively correlated with the image differences.
[0180] In any embodiment of this application, the video features include color features;
[0181] The first processing unit 12 is configured to:
[0182] Based on the color features, determine the color difference between two adjacent video frames in the video to be processed;
[0183] The image difference between two adjacent video frames in the video to be processed is obtained based on the color difference between the two adjacent video frames in the video to be processed, and the image difference is positively correlated with the color difference.
[0184] In any embodiment of this application, the first processing unit 12 is configured to:
[0185] Based on the video features, determine the correlation between two non-adjacent video frames in the video to be processed;
[0186] Based on the correlation, the segmentation probability of each video frame between the two non-adjacent video frames is obtained, and the segmentation probability of the video frames between the two non-adjacent video frames is negatively correlated with the correlation.
[0187] In conjunction with any embodiment of this application, the first processing unit 12 is further configured to:
[0188] Determine the resolution of each video frame in the video to be processed;
[0189] If there are video frames with a resolution less than or equal to a resolution threshold in two adjacent frames of the video to be processed, the correlation between two non-adjacent video frames in the video to be processed is determined.
[0190] In any embodiment of this application, the first processing unit 12 is further configured to process the video to be processed using a neural network to obtain the video features of the video to be processed.
[0191] In this embodiment, when the video processing device acquires a video to be processed, it obtains the segmentation probability of each video frame of the video to be processed based on the video characteristics of the video to be processed. The segmentation probability represents the probability that a video frame is an endpoint of a video segment, and the endpoints include a start frame and an end frame. Then, it selects two or more video frames from the video to be processed whose segmentation probability is greater than or equal to a segmentation threshold, thus selecting two or more video frames that can serve as endpoints. Finally, it uses two frames with adjacent timestamps from the two or more video frames as the start frame and the end frame, respectively, to extract the target video segment from the video to be processed, thereby achieving the segmentation of the video to be processed.
[0192] In some embodiments, the functions or modules of the apparatus provided in this application can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0193] Figure 6 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device 2 includes a processor 21 and a memory 22. Optionally, the electronic device 2 also includes an input device 23 and an output device 24. The processor 21, memory 22, input device 23, and output device 24 are coupled together via connectors, which include various interfaces, transmission lines, or buses, etc., and are not limited in this embodiment. It should be understood that in the various embodiments of this application, coupling refers to mutual connection in a specific way, including direct connection or indirect connection through other devices, such as through various interfaces, transmission lines, buses, etc.
[0194] Processor 21 may include one or more processors, such as one or more central processing units (CPUs). If the processor is a CPU, it may be a single-core CPU or a multi-core CPU. Optionally, processor 21 may be a processor group consisting of multiple CPUs, with the multiple processors coupled to each other via one or more buses. Optionally, the processor may also be other types of processors, etc., which are not limited in this embodiment.
[0195] The memory 22 can be used to store computer program instructions, as well as various types of computer program code, including program code for executing the scheme of this application. Optionally, the memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), which is used for related instructions and data.
[0196] Input device 23 is used to input data and / or signals, and output device 24 is used to output data and / or signals. Input device 23 and output device 24 can be independent devices or an integrated device.
[0197] It is understood that in this embodiment of the application, the memory 22 can be used not only to store related instructions, but also to store related data, etc. This embodiment of the application does not limit the specific data stored in the memory.
[0198] Understandable, Figure 6 This is merely a simplified design of an electronic device. In practical applications, the electronic device may also include other necessary components, including, but not limited to, any number of input / output devices, processors, memories, etc., and all electronic devices that can implement the embodiments of this application are within the protection scope of this application.
[0199] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0200] Those skilled in the art will readily understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. Those skilled in the art will also readily understand that the various embodiments of this application have different focuses, and for the sake of convenience and brevity, the same or similar parts may not be repeated in different embodiments. Therefore, parts not described or not described in detail in one embodiment can be referred to the descriptions in other embodiments.
[0201] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some video features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0202] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0203] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0204] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital versatile discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
[0205] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A video processing method, characterized in that, The method includes: Get the video to be processed; The video to be processed is input into a neural network so that the neural network processes the video to obtain the video features of the video to be processed; The video features of the video to be processed are input into the first recognition branch of the neural network, so that the first recognition branch obtains the segmentation probability of each video frame in the video to be processed based on the video features of the video to be processed. The segmentation probability represents the probability that a video frame is the endpoint of a video segment. The endpoint includes the start frame and the end frame. Select two or more video frames from the video to be processed, the segmentation probability being greater than or equal to the segmentation threshold. The two or more video frames include the first video frame and the second video frame with adjacent timestamps among the two or more video frames. The target video segment is obtained by extracting the video segment between the first video frame and the second video frame from the video to be processed. Alternatively, the video features of the video to be processed can be input into the second recognition branch of the neural network, so that the second recognition branch can obtain the category probability of each video frame in the video to be processed based on the video features of the video to be processed, wherein the category probability represents the probability of the video frame belonging to a category. Based on the video features of the video to be processed, the segmentation probability of each video frame of the video to be processed is obtained. The segmentation probability represents the probability that a video frame is the endpoint of a video segment. The endpoint includes the start frame and the end frame. Select two or more video frames from the video to be processed, the segmentation probability being greater than or equal to the segmentation threshold. The two or more video frames include the first video frame and the second video frame with adjacent timestamps among the two or more video frames. The target video segment is obtained by extracting the video segment between the first video frame and the second video frame from the video to be processed. The category of the target video segment is obtained based on the category probability of each video frame from the first video frame to the second video frame.
2. The method according to claim 1, characterized in that, The step of obtaining the category of the target video segment based on the category probabilities of each video frame from the first video frame to the second video frame includes: Based on the category probabilities of each video frame from the first video frame to the second video frame, the category probabilities of each category in the target video segment are summed to obtain the category probability sum of each category; The category corresponding to the maximum value of the sum of the category probabilities is determined as the category of the target video segment.
3. The method according to claim 1 or 2, characterized in that, The step of obtaining the segmentation probability of each video frame of the video to be processed based on the video features of the video to be processed includes: Based on the video features, determine the image differences between two adjacent video frames in the video to be processed; Based on the image differences, the segmentation probability of two adjacent video frames is determined, and the segmentation probability of two adjacent video frames is positively correlated with the image differences.
4. The method according to claim 3, characterized in that, The video features include color features; The step of determining the image differences between two adjacent video frames in the video to be processed based on the video features includes: Based on the color features, determine the color difference between two adjacent video frames in the video to be processed; The image difference between two adjacent video frames in the video to be processed is obtained based on the color difference between the two adjacent video frames in the video to be processed, and the image difference is positively correlated with the color difference.
5. The method according to claim 1 or 2, characterized in that, The step of obtaining the segmentation probability of each video frame of the video to be processed based on the video features of the video to be processed includes: Based on the video features, determine the correlation between two non-adjacent video frames in the video to be processed; Based on the correlation, the segmentation probability of each video frame between the two non-adjacent video frames is obtained, and the segmentation probability of the video frames between the two non-adjacent video frames is negatively correlated with the correlation.
6. The method according to claim 5, characterized in that, Before determining the correlation between two non-adjacent video frames in the video to be processed based on the video features, the method further includes: Determine the resolution of each video frame in the video to be processed; The step of determining the correlation between two non-adjacent video frames in the video to be processed based on the video features includes: If there are video frames with a resolution less than or equal to a resolution threshold in two adjacent frames of the video to be processed, the correlation between two non-adjacent video frames in the video to be processed is determined.
7. A video processing apparatus, characterized in that, The device includes: The acquisition unit is used to acquire the video to be processed. The first processing unit is configured to input the video to be processed into a neural network, so that the neural network processes the video to be processed to obtain the video features of the video to be processed; The first processing unit is further configured to input the video features of the video to be processed into the first recognition branch in the neural network, so that the first recognition branch obtains the segmentation probability of each video frame of the video to be processed based on the video features of the video to be processed. The segmentation probability characterizes the probability that a video frame is the endpoint of a video segment, and the endpoint includes a start frame and an end frame. The second processing unit is used to select two or more video frames from the video to be processed, the video frames having a segmentation probability greater than or equal to the segmentation threshold, the two or more video frames including a first video frame and a second video frame with adjacent timestamps. The cropping unit is used to crop a video segment from the video to be processed between the first video frame and the second video frame to obtain a target video segment; Alternatively, the first processing unit is configured to input the video features of the video to be processed into the second recognition branch in the neural network, so that the second recognition branch obtains the category probability of each video frame in the video to be processed based on the video features of the video to be processed, wherein the category probability represents the probability of the video frame belonging to a category. The first processing unit is further configured to obtain the segmentation probability of each video frame of the video to be processed based on the video features of the video to be processed. The segmentation probability represents the probability that a video frame is the endpoint of a video segment, and the endpoint includes a start frame and an end frame. The second processing unit is used to select two or more video frames from the video to be processed, the video frames having a segmentation probability greater than or equal to a segmentation threshold. The two or more video frames include a first video frame and a second video frame with adjacent timestamps among the two or more video frames. The extraction unit is used to extract a video segment from the video to be processed between the first video frame and the second video frame to obtain a target video segment. The second processing unit is further configured to obtain the category of the target video segment based on the category probability of each video frame from the first video frame to the second video frame.
8. An electronic device, characterized in that, include: A processor and a memory, the memory being used to store computer program code, the computer program code including computer instructions, wherein, when the processor executes the computer instructions, the electronic device performs the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 6.