Video processing methods, apparatus, electronic devices and storage media

By splitting video segments and setting the video features with the highest proportion of feature information, combined with a preset transformation model and convolutional network, the problem of one-sided video segment analysis results is solved, and accurate prediction of video excitement is achieved.

CN116668765BActive Publication Date: 2026-04-03BEIJING IQIYI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, fixed-duration analysis of video clips leads to biased analysis results, which cannot accurately reflect the quality of the video. Furthermore, manual adjustment is required to determine the appropriate video duration, which is inconvenient and prone to errors.

Method used

By splitting the video segment to be analyzed into multiple sub-video segments, using preset split lengths and feature processing, video features are generated. A preset transformation model and convolutional network are used to predict the video's excitement score, setting the target segment's feature information to have the highest proportion, and combining it with contextual feature information.

Benefits of technology

It enables accurate analysis of the quality of video clips, avoiding the problems of missing context and large errors caused by fixed-duration analysis, and improves the accuracy of video quality scores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116668765B_ABST
    Figure CN116668765B_ABST
Patent Text Reader

Abstract

This invention provides a video processing method, apparatus, electronic device, and storage medium, comprising: splitting a video segment to be analyzed according to a preset splitting length to obtain multiple first sub-video segments; performing feature processing on each first sub-video segment to obtain feature information corresponding to the first sub-video segment; generating video features of the video segment to be analyzed based on the feature information corresponding to the first sub-video segments, wherein the feature information of the target segment has the highest proportion in the video features; and obtaining a video quality score corresponding to the target segment in the video segment to be analyzed based on the video features of the video segment to be analyzed. This invention improves the accuracy of the video quality score of the target segment by splitting the video segment to be analyzed into multiple short video segments and obtaining video features that both connect the feature information of the preceding and following segments and highlight the features of the target segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, and in particular to a video processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] Currently, video applications often display exciting clips of videos on their interface to attract users to watch the videos. These exciting clips are extracted from the videos and are of high quality. This requires analyzing the video clips to determine their quality and thus identifying the most exciting segments.

[0003] To be applicable to videos from various sources, existing solutions use a fixed-length video segment method to analyze the video quality of video clips. That is, by slicing video segments into fixed-length segments, performing quality prediction processing on these segments, and obtaining the quality level of the fixed-length video segments.

[0004] However, since the duration of the video segment used for predicting video quality is fixed, it is necessary to determine the appropriate duration of the video segment to be analyzed based on the video content and analyze the video segment independently. This will result in the loss of contextual video content information, and the analysis results of the video segment will be highly biased, making it impossible for the obtained video quality to accurately reflect the quality of the video. Summary of the Invention

[0005] The purpose of this invention is to provide a video processing method, apparatus, electronic device, and storage medium to solve the problem that existing methods of independently analyzing video segments result in highly biased analysis results that cannot accurately reflect the quality of the video. The specific technical solution is as follows:

[0006] In a first aspect of this invention, a video processing method is provided, which may include:

[0007] According to the preset splitting length, the video segment to be analyzed is split into multiple first sub-video segments, wherein the first sub-video segments include the target segments in the video segment to be analyzed;

[0008] Perform feature processing on each of the first sub-video segments to obtain the feature information corresponding to the first sub-video segment;

[0009] Based on the feature information corresponding to the first sub-video segment, video features of the video segment to be analyzed are generated; wherein, the feature information of the target segment accounts for the highest proportion among the video features;

[0010] Based on the video features of the video segment to be analyzed, obtain the video quality score corresponding to the target segment in the video segment to be analyzed.

[0011] Optionally, the step of performing feature processing on each of the first sub-video segments to obtain feature information corresponding to the first sub-video segment includes:

[0012] Each of the first sub-video segments is captured using a sliding window of a preset size, and a preset frame image is extracted from the first sub-video segment;

[0013] Generate the image features of the preset frame image, and store the correspondence between the image features of the preset frame image and each of the first sub-video segments;

[0014] Based on the image features of the preset frame images and the corresponding relationship, the feature information corresponding to the first sub-video segment is obtained.

[0015] Optionally, generating video features of the video segment to be analyzed based on the feature information corresponding to the first sub-video segment includes:

[0016] The feature information corresponding to multiple first sub-video segments is respectively input into a preset conversion model so that the preset conversion model can convert the feature information.

[0017] Obtain the feature vector of the preset output quantity ratio of multiple first sub-video segments output by the preset conversion model;

[0018] The feature vectors of multiple first sub-video segments are merged according to a preset output quantity ratio to generate the video features of the video segment to be analyzed.

[0019] Optionally, before generating the video features of the video segment to be analyzed based on the feature information corresponding to the first sub-video segment, the method further includes:

[0020] Invoke multiple preset conversion models corresponding to the first sub-video segment;

[0021] Based on the position of the target segment in the video segment to be analyzed, a preset output ratio of multiple first sub-video segments corresponding to a preset conversion model is determined. The preset output ratio of the feature information of the first sub-video segments is inversely proportional to the distance between the first sub-video segments and the target segment, and the preset output ratio of the feature information of the target segment is the highest.

[0022] Optionally, obtaining the video excitement score corresponding to the target segment in the video segment to be analyzed based on the video features of the video segment to be analyzed includes:

[0023] The video features of the video segment to be analyzed are input into a pre-trained convolutional network to predict the video excitement score;

[0024] Obtain the video brilliance score corresponding to the target segment in the video segment to be analyzed output by the convolutional network.

[0025] Optionally, before inputting the video features of the video segment to be analyzed into a pre-trained convolutional network for video excitement score prediction, the method further includes:

[0026] Obtain training data for training the convolutional network, wherein the training data includes video features of historical video segments to be analyzed and the actual video quality corresponding to the historical video segments to be analyzed;

[0027] The video features of the historical video segment to be analyzed are input into the convolutional network to predict the video excitement score, and the predicted video excitement score output by the convolutional network is obtained.

[0028] Call the loss function to determine the score loss between the predicted video quality score and the actual video quality score;

[0029] Based on the fractional loss, the gradient of the network parameters of the convolutional network is calculated, and the network parameters are updated based on the gradient.

[0030] After determining the score loss between the predicted video excitement score and the actual video excitement score after updating the network parameters, and if the score loss is less than a preset loss threshold, the trained convolutional network is obtained.

[0031] Optionally, after obtaining the video quality score corresponding to the target segment in the video segment to be analyzed based on the video features of the video segment to be analyzed, the method further includes:

[0032] According to the second preset splitting length, each of the first sub-video segments is split to obtain multiple second sub-video segments, wherein the second preset splitting length is less than the preset splitting length;

[0033] Obtain the video quality scores of multiple second sub-video segments;

[0034] The video quality scores of multiple second sub-video segments are weighted and averaged to generate the video quality score of the target segment in the video segment to be analyzed.

[0035] In a second aspect of the present invention, a video processing apparatus is provided, the apparatus comprising:

[0036] The first processing module is used to split the video segment to be analyzed according to a preset splitting length to obtain multiple first sub-video segments, wherein the first sub-video segments include the target segments in the video segment to be analyzed.

[0037] The second processing module is used to perform feature processing on each of the first sub-video segments to obtain the feature information corresponding to the first sub-video segments.

[0038] The first generation module is used to generate video features of the video segment to be analyzed based on the feature information corresponding to the first sub-video segment; wherein, the feature information of the target segment accounts for the highest proportion among the video features;

[0039] The first acquisition module is used to acquire the video quality score corresponding to the target segment in the video segment to be analyzed based on the video features of the video segment to be analyzed.

[0040] A third aspect of the present invention also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus.

[0041] Memory, used to store computer programs;

[0042] The processor, when executing a program stored in memory, performs any of the video processing methods described above.

[0043] In a fourth aspect of the invention, a computer-readable storage medium is also provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform any of the video processing methods described above.

[0044] This invention, by splitting the video segment to be analyzed into multiple short video segments, can flexibly obtain the target segment of the required duration, avoiding the need for manual selection of appropriate video duration based on video context due to the fixed processing duration of the processing device. By setting the target segment to have the highest proportion in the video features of the video segment to be analyzed, video features that both connect the feature information of the preceding and following segments and highlight the features of the target segment are obtained. Based on these video features, the video quality score of the target segment is obtained, improving the accuracy of the video quality score of the target segment. This solves the problems of missing context due to short video length and excessive error in quality score due to long video length, thus achieving accurate analysis of the quality of video segments. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0046] Figure 1 This is one of the flowcharts of a video processing method provided in an embodiment of the present invention;

[0047] Figure 2 yes Figure 1 The flowchart of step 102 of the video processing method provided in this embodiment of the invention;

[0048] Figure 3 yes Figure 1 The flowchart of step 103 of the video processing method provided in this embodiment of the invention;

[0049] Figure 4 yes Figure 1 The flowchart of step 104 of the video processing method provided in this embodiment of the invention;

[0050] Figure 5 This is a flowchart illustrating a video processing method provided in an embodiment of the present invention;

[0051] Figure 6 This is the second step flowchart of a video processing method provided in an embodiment of the present invention;

[0052] Figure 7 This is a structural block diagram of a video processing device provided in an embodiment of the present invention;

[0053] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0054] The technical solutions of the embodiments of the present invention will now be described with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application can be thoroughly understood and its scope can be fully conveyed to those skilled in the art.

[0055] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0056] In various embodiments of the present invention, it should be understood that the sequence number of each process described below does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0057] The video processing method, apparatus, electronic device, and storage medium provided by the present invention will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0058] Reference Figure 1 The diagram illustrates one of the steps of a video processing method provided by an embodiment of the present invention, which may include:

[0059] Step 101: According to the preset splitting length, the video segment to be analyzed is split into multiple first sub-video segments, wherein the first sub-video segments include the target segments in the video segment to be analyzed.

[0060] In this embodiment of the invention, to address the problem of inaccurate analysis results caused by independently analyzing video segments within a fixed short time period in current video processing methods, it is necessary to consider two issues: First, the appropriate duration for reflecting video content requires manual adjustment, which is inconvenient. Second, using a single score to represent the quality of all videos within a longer time period leads to significant errors in the analysis results. Furthermore, the current method of independently analyzing video segments with fixed input durations is not convenient for analyzing video quality and cannot objectively evaluate the quality of videos within that time period. Therefore, this embodiment of the invention proposes to input longer video segments, break them down, and combine the video information from the longer segments to analyze and predict the video quality score of the target segment.

[0061] Specifically, in this embodiment of the invention, the video segment to be analyzed is split according to a preset splitting length to obtain multiple first sub-video segments. Each first sub-video segment includes a target segment within the video segment to be analyzed; the target segment is the video segment from which a video quality score is to be obtained. It should be noted that the preset splitting length is determined based on the duration of the target segment to be analyzed, and the duration of the target segment is determined based on the user's video analysis requirements, and is not specifically limited here.

[0062] For example, if the target segment duration is determined to be 4 seconds, then based on the location of the target segment, a 20-second video segment including the target segment is obtained. This 20-second video segment is then split according to a preset splitting length (4 seconds), that is, the video segment to be analyzed is split into five first sub-video segments of length 4 seconds each. Each of these five 4-second first sub-video segments includes the target segment within the video segment to be analyzed. Of course, the above is only a specific example. In actual use, the duration of the video segment to be analyzed is determined based on the processing time threshold of the video processing device, and the duration of the target segment is determined based on actual analysis requirements. These details will not be elaborated upon here.

[0063] In this embodiment of the invention, by splitting the video segment to be analyzed into multiple short video segments, the target segment of the required duration can be flexibly obtained, avoiding the problem that the video processing device has a fixed processing duration and requires manual selection of the appropriate video duration based on the video context, which is inconvenient.

[0064] Step 102: Perform feature processing on each first sub-video segment to obtain the feature information corresponding to the first sub-video segment.

[0065] In this embodiment of the invention, in order to achieve accurate video quality prediction of the target segment, feature processing is performed on multiple sub-video segments after splitting. This yields video features that connect the preceding and following segments while highlighting the target segment. Furthermore, based on these video features, a video quality score for the target segment is obtained, thereby improving the accuracy of the video quality score for the target segment.

[0066] Specifically, this embodiment performs feature processing on each first sub-video segment to obtain the feature information corresponding to the first sub-video segment. For example, for each 4-second first sub-video segment, preset frame images are uniformly extracted as input data using an equally spaced frame extraction method. The open-source image feature extraction model Clip model is used to extract features from the input preset frame images to obtain image features, and then the feature information corresponding to each first sub-video segment is output based on the image features. It should be noted that the above example is only a specific illustration. In actual use, step 102 can also use any other image feature extraction model capable of extracting image features to perform feature processing on each first sub-video segment to obtain the feature information of the first sub-video segment; these will not be elaborated upon here.

[0067] Step 103: Generate video features of the video segment to be analyzed based on the feature information corresponding to the first sub-video segment; among them, the feature information of the target segment accounts for the highest proportion in the video features.

[0068] In this embodiment of the invention, in order to obtain video features that both connect the preceding and following segments and highlight the target segment, the target segment is set to have the highest proportion of video features in the video segment to be analyzed, thereby obtaining the video quality score of the target segment based on these video features.

[0069] Specifically, the feature information corresponding to multiple first sub-video segments is input into a preset conversion model, so that the preset conversion model can convert the feature information and output feature vectors of multiple first sub-video segments with a preset output quantity ratio; among them, the feature information of the target segment has the highest preset output quantity ratio, so that the feature information of the target segment has the highest proportion in the video features; the feature vectors of multiple first sub-video segments with a preset output quantity ratio are merged to generate the video features of the video segment to be analyzed.

[0070] It should be noted that in this embodiment, the preset transformation model is used to process the features corresponding to the video segments in order to output a video feature that globally expresses the content of the video segment to be analyzed. Therefore, this embodiment can use the transformer model. The transformer model is a transformation model that relies entirely on self-attention to calculate the input and output, without using a recurrent neural network for sequence alignment. This invention uses the transformer model to perform multi-head self-attention on the feature information corresponding to each first sub-video segment in the time and space dimensions, thereby outputting a video feature that globally expresses the content of the video segment to be analyzed.

[0071] For example, multiple transformer models are called to transform the feature information of each first sub-video segment. The input feature information is operated on with the transformation matrix to obtain multiple feature matrices. The multiple feature matrices are then operated on with the pre-generated attention matrix to output multiple feature vectors of a preset output ratio for the first sub-video segments. The above feature vectors are then merged to generate the video features of the video segment to be analyzed.

[0072] Specifically, based on the position of the first sub-video segment within the entire video segment to be analyzed, the output proportions of multiple transformer models are pre-set. The proportion of feature information output from the first sub-video segment is inversely proportional to the distance between the first sub-video segment and the target segment. That is, the greater the distance between the first sub-video segment and the target segment, the lower the proportion of feature information output; conversely, the smaller the distance, the higher the proportion of feature information output. This setting ensures that the proportion of feature information output from the target segment is the highest, highlighting the features of the target segment. Furthermore, it considers the features of the video segments in the context of the target segment, thereby obtaining a video feature that globally expresses the content of the video segment to be analyzed, and thus predicting the video quality of the target segment.

[0073] Specifically, in this embodiment, the video quality analysis result of the 4-second target segment is obtained by processing the 20-second video segment to be analyzed. The number of output features of the transformer model corresponding to each first sub-video segment (0-4s, 4-8s, 8-12s, 12-16s, 16-20s) is preset. Specifically, the output ratio of feature information of the first sub-video segment is set according to the position of the target segment. If the target segment is in the middle of the video segment to be analyzed, the feature output is in the ratio of 1:3:8:3:1. The feature information of each first sub-video segment is input into the corresponding transformer model. The transformer model will convert the feature information into a preset number of feature vectors. The feature output ratio of 0-4s and 16-20s is the same, so the same transformer1 is used to save data processing resources. The number of output features is 1. Similarly, the same transformer 2 is used for the feature output ratios of 4-8s and 12-16s, with a feature output of 16*24. The feature output ratio of 8-12s is the highest, using a separate transformer 3, with a feature output of 16*64. The five sets of feature vectors output by the three transformer models are then merged to finally generate 16*8+16*24+16*64+16*24+16*8, a total of 16*128-dimensional features, which are used to express the video features of the entire 20s video segment to be analyzed.

[0074] In this embodiment of the invention, by setting the proportion of feature information corresponding to each split sub-video segment, the proportion of the target segment in the video features of the generated video segment to be analyzed is maximized. This yields video features that connect the preceding and following segments while highlighting the target segment. Based on these video features, a video quality score for the target segment is obtained, improving the accuracy of the video quality score. This solves the problems of missing contextual content due to short video duration and excessive error in quality score due to long video duration, thus achieving accurate analysis of the quality of video segments.

[0075] Step 104: Based on the video features of the video segment to be analyzed, obtain the video excitement score corresponding to the target segment in the video segment to be analyzed.

[0076] In this embodiment of the invention, the video features of the video segment to be analyzed are input into a pre-trained convolutional network to predict the video brilliance, so that the convolutional network outputs a video brilliance score. The obtained video brilliance score is the video brilliance score corresponding to the target segment in the video segment to be analyzed.

[0077] It should be noted that this embodiment uses a pre-trained convolutional network to predict the video features of the video segment to be analyzed, thereby achieving accurate scoring of the video's entertainment value and predicting the video entertainment value score corresponding to the target segment within the video segment to be analyzed. The convolutional network is pre-trained based on the predicted and actual video entertainment value scores of historical video segments to be analyzed.

[0078] Specifically, in this embodiment, the video features of the video segment to be analyzed are input into a pre-trained convolutional network. The video highlighting module within the convolutional network outputs a video highlighting score, which is the video highlighting score corresponding to the target segment in the video segment to be analyzed. In this embodiment, the convolutional network can be a regression neural network. The fully connected layer of the convolutional network merges the five sets of feature vectors output by the three transformer models to generate a total of 16*128 dimensions of video features (16*8+16*24+16*64+16*24+16*8). This clustering analysis result is then used by the pre-trained highlighting module to predict and score the clustering analysis result, thus obtaining the video highlighting score corresponding to the target segment in the video segment to be analyzed. It should be noted that the convolutional network is obtained through iterative training using training data including a certain number of historical video segments to be analyzed. The specific training process is described in the following embodiments and will not be repeated here.

[0079] In this embodiment of the invention, a trained convolutional network is used to obtain the video quality score of the target segment in the video segment to be analyzed based on the video features of the video segment to be analyzed, thereby achieving accurate analysis of the quality of the video segment and improving the accuracy of the video quality score of the target segment. Furthermore, the convolutional network can achieve quality scoring by processing any video feature, making the technical solution provided by this embodiment of the invention more applicable.

[0080] The video processing method provided by this invention splits a video segment to be analyzed into multiple first sub-video segments according to a preset splitting length. Feature processing is then performed on each first sub-video segment to obtain corresponding feature information. Based on this feature information, video features of the video segment to be analyzed are generated. Finally, based on these video features, a video quality score corresponding to the target segment within the video segment to be analyzed is obtained. In this embodiment, by splitting the video segment to be analyzed into multiple short video segments, the required target segment length can be flexibly obtained, avoiding the need for manual selection of a suitable video length based on the video context due to the fixed processing time of the processing device. By setting the target segment to have the highest proportion in the video features of the video segment to be analyzed, video features that both connect the feature information of the preceding and following segments and highlight the features of the target segment are obtained. Based on these video features, a video quality score for the target segment is obtained, improving the accuracy of the target segment's video quality score. This solves the problems of missing contextual content due to short video lengths and excessively large quality score errors due to long video lengths, achieving accurate analysis of the quality of video segments.

[0081] Reference Figure 2 ,yes Figure 1 The flowchart of step 102 of the video processing method provided in this embodiment of the invention is shown. The video processing method disclosed in this embodiment is... Figure 1 One feasible implementation of step 102 in the illustrated embodiment specifically includes:

[0082] Step 201: Use a sliding window of a preset size to capture each first sub-video segment, and extract preset frame images from the first sub-video segments.

[0083] In this embodiment of the invention, a preset frame image is uniformly extracted from each 4-second first sub-video segment using an equally spaced frame extraction method, based on a sliding window of a preset size. Specifically, the sliding window of the preset size slides across multiple sub-video segments at a uniform speed. Since the sliding window is a pre-set fixed-size window, when the sliding window slides over a sub-video segment, the preset frame image is extracted according to the equally spaced frame extraction method. The frame extraction method is set according to the sliding speed of the sliding window. For example, in each 4-second sub-video segment, 16 frames are uniformly extracted using a frame extraction method of 4 frames per second. Of course, the above is only a specific example. In actual use, any other preset frame image that can be uniformly extracted at equal intervals can be used to perform feature processing on each first sub-video segment. No specific limitation is made here.

[0084] This embodiment uses a sliding processing method to uniformly extract preset frame images at equal intervals, ensuring that the feature information corresponding to the first sub-video segment can accurately reflect the video content information, thereby improving the accuracy of the video excitement score.

[0085] Step 202: Generate image features of preset frame images and store the correspondence between the image features of preset frame images and each first sub-video segment.

[0086] In this embodiment of the invention, image features of preset frame images are generated based on preset frame images extracted from each first sub-video segment and a preset image feature extraction model. Since the target segment can be determined as needed, the preset frame images of the first sub-video segments can be reused during video processing. Therefore, this embodiment stores the correspondence between the image features of the preset frame images and each first sub-video segment, so that when using a first sub-video segment, image features can be directly obtained based on the correspondence, without needing to repeatedly extract features from the first sub-video segment.

[0087] For example, the preset image feature extraction model in this embodiment can use the Clip model to extract features for each first sub-video segment. The model used for each frame is the same. In order to obtain video features that can connect the feature information of the preceding and following segments and highlight the features of the target segment, this is achieved by setting the target segment to have the highest proportion in the video features of the video segment to be analyzed. Therefore, during the use of the model, if the target segment of the video segment to be analyzed is the 4-second segment from 12 to 16 seconds, then the video features of the video segment to be analyzed from 4 to 24 seconds need to be obtained. Since the preset frame images have already been extracted and preset frames generated for each first sub-video segment when analyzing the video quality of the 8-12 second target segment, the preset image feature extraction model has already been used. In this video quality analysis of the 12-16s target segment, the image features can be utilized from the first sub-video segments (4-8s, 8-12s, 12-16s, and 16-20s) of the previously analyzed video segment. Then, a preset image feature extraction model is used to extract image features from the newly divided 20-24s video segment. This means that when the first video segment required for the target segment overlaps with a historical first sub-video segment, preset frame images from the historical first sub-video segment can be reused during video processing. This reuse of image feature data extracted by the Clip model saves video processing resources and further improves data processing speed.

[0088] It should be noted that the above examples are only specific illustrations. In actual use, any other image feature extraction model capable of extracting image features can be used to process the features of each first sub-video segment to obtain the feature information of the first sub-video segment. No specific limitations are made here.

[0089] Step 203: Based on the image features and correspondence of the preset frame images, obtain the feature information corresponding to the first sub-video segment.

[0090] In this embodiment, the open-source preset image feature extraction model Clip model is used. After extracting the feature information from each frame of the input image, the feature information of each frame of the output image is processed according to the image features and correspondence of the preset frame images to obtain the feature information corresponding to the first sub-video segment.

[0091] Compared with the prior art, the embodiments of the present invention, while achieving the beneficial effects of the first embodiment, adopt a sliding processing method. During the sliding window process, the image features of repeatedly used video segments are stored, and the feature extraction results can be used repeatedly, effectively reducing the problem of repeated data calculation, saving processing resource costs, and ensuring a significant improvement in algorithm performance.

[0092] Reference Figure 3 ,yes Figure 1 The flowchart of step 103 of the video processing method provided in this embodiment of the invention is shown. The video processing method disclosed in this embodiment is... Figure 1 One feasible implementation of step 103 in the illustrated embodiment specifically includes:

[0093] Step 301: Input the feature information corresponding to multiple first sub-video segments into the preset conversion model so that the preset conversion model can convert the feature information.

[0094] In this embodiment of the invention, in order to obtain feature information that connects the preceding and following segments while highlighting the video features of the target segment, the proportion of the target segment in the video features of the video segment to be analyzed is set to be the highest, thereby obtaining the video quality score of the target segment based on these video features. Specifically, the feature information corresponding to multiple first sub-video segments is input into a preset conversion model, so that the preset conversion model performs conversion processing on the feature information and outputs feature vectors of multiple first sub-video segments in a preset output ratio.

[0095] For example, in this embodiment, the number of output features of the transformer model corresponding to each first sub-video segment, i.e., 0-4s, 4-8s, 8-12s, 12-16s, and 16-20s, is preset. For instance, the output number ratio is set according to the position of the target segment. If the target segment is located in the middle of the video segment to be analyzed, the features can be output in a ratio of 1:3:8:3:1 to maximize the preset output number ratio of the target segment's feature information, thus obtaining the feature vector of the first sub-video segment with the preset output number ratio.

[0096] Step 302: Obtain the feature vector of the preset output quantity ratio of multiple first sub-video segments output by the preset conversion model.

[0097] Step 303: Merge the feature vectors of multiple first sub-video segments according to a preset output quantity ratio to generate video features of the video segment to be analyzed.

[0098] Specifically, the feature vectors of the first sub-video segment are merged to generate video features for the video segment to be analyzed. For example, for any 20-second video segment to be analyzed, to ensure more reasonable use of the 20-second video content, when outputting the transformer model, the features are output in a ratio of 1:3:8:3:1 based on the time distance of each 4-second video segment relative to the target segment. This uses the features of the 20-second video and increases the dependency weight of the corresponding features of the target segment. Thus, the five sets of feature vectors output by the three transformer models are merged to obtain the video features that represent the entire 20-second video, so as to predict the video excitement score of the target segment based on these video features.

[0099] Specifically, before step 103 generates video features of the video segment to be analyzed based on the feature information corresponding to the first sub-video segment, it also includes:

[0100] Call multiple preset conversion models corresponding to the first sub-video segment;

[0101] Based on the position of the target segment within the video segment to be analyzed, the preset output ratios of multiple first sub-video segments corresponding to preset conversion models are determined. Among them, the preset output ratio of feature information of the first sub-video segments is inversely proportional to the distance between the first sub-video segments and the target segment, with the preset output ratio of feature information of the target segment being the highest.

[0102] Compared with the prior art, the embodiments of the present invention, based on achieving the beneficial effects of the first and second embodiments, set the proportion of feature information corresponding to each split video segment, so that the proportion of the target segment in the video features of the generated video segment to be analyzed is the highest. This results in video features that both connect the preceding and following segments and highlight the target segment. Based on these video features, the video quality score of the target segment is obtained, improving the accuracy of the video quality score of the target segment. This solves the problems of missing contextual content due to short video length and excessive error in quality score due to long video length, thus achieving accurate analysis of the quality of video segments.

[0103] Reference Figure 4 ,yes Figure 1 The flowchart of step 104 of the video processing method provided in this embodiment of the invention is shown. The video processing method disclosed in this embodiment is... Figure 1 One feasible implementation of step 104 in the illustrated embodiment specifically includes:

[0104] Step 401: Input the video features of the video segment to be analyzed into a pre-trained convolutional network to predict the video excitement score.

[0105] Specifically, by inputting the video features of the video segment to be analyzed into a pre-trained convolutional network to score the video's quality, the analysis result of the video quality corresponding to the target segment within the video segment to be analyzed can be predicted.

[0106] Step 402: Obtain the video excitement score corresponding to the target segment in the video segment to be analyzed from the output of the convolutional network.

[0107] It should be noted that the convolutional network is pre-trained based on the predicted video quality scores and the actual video quality scores from historical video processing. The convolutional network is used to obtain the video quality score corresponding to the target segment in the video segment to be analyzed and then process it.

[0108] In this embodiment of the invention, a trained convolutional network is used to obtain the video quality score of the target segment in the video segment to be analyzed based on the video features of the video segment to be analyzed, thereby achieving accurate analysis of the quality of the video segment and improving the accuracy of the video quality score of the target segment. Furthermore, the convolutional network can achieve quality scoring by processing any video feature, making the technical solution provided by this embodiment of the invention more applicable.

[0109] Specifically, the steps for pre-training a convolutional network include:

[0110] Obtain training data for training the convolutional network, including video features of historical video segments to be analyzed and the actual video quality corresponding to the historical video segments to be analyzed;

[0111] The video features of the historical video segments to be analyzed are input into a convolutional network to predict the video excitement score, and the predicted video excitement score output by the convolutional network is obtained.

[0112] Call the loss function to determine the score loss between the predicted video highlight score and the actual video highlight score;

[0113] Based on the fractional loss, calculate the gradient of the network parameters of the convolutional network, and update the network parameters based on the gradient;

[0114] After determining the updated network parameters, the score loss between the predicted video excitement score and the actual video excitement score is calculated. If the score loss is less than a preset loss threshold, the trained convolutional network is obtained.

[0115] It should be noted that in this embodiment, the convolutional network can be a regressive neural network. A training algorithm is used to train the video features of the historical video segments to be analyzed until the loss between the video excitement score output by the convolutional network and the actual video excitement score meets the requirements of the preset loss function, thus obtaining the trained convolutional network.

[0116] To score video quality using a convolutional network (CNN), a predefined loss function, MSEloss, is called to determine the score loss between the predicted video quality score and the actual video quality score. Backward gradient propagation is then performed on the quality scoring module and the transformer model within the CNN, calculating the gradient of the loss function with respect to each parameter and updating the network parameters based on the gradient until the predicted score loss of the adjusted CNN is less than a predefined loss threshold. This completes the training of the CNN, and the trained CNN is obtained. Iterative loss training further improves the prediction accuracy of the CNN.

[0117] To help those skilled in the art to better understand the video processing method described in the above steps, please refer to... Figure 5 The following is an example of a flowchart illustrating a video processing method, specifically including:

[0118] Based on the preset segmentation length, the video segment to be analyzed is segmented, referring to... Figure 5 Taking a 20-second video segment as an example, it is divided into five 4-second video segments. Each video segment includes the target segment (8-12 seconds) within the video segment to be analyzed. For each 4-second video segment, a preset frame image is extracted evenly using an equally spaced frame extraction method. An image feature extraction model is used to extract features from the preset frame images to obtain corresponding image features. The feature information of the corresponding image features of multiple video segments is input into a preset conversion model to transform the feature information and output feature vectors of a preset output ratio for multiple video segments. The feature vectors of the preset output ratio for multiple video segments are merged to generate video features of the video segment to be analyzed. The video features of the video segment to be analyzed are input into a pre-trained video excitement scoring model to predict the video excitement level, so that the video excitement scoring model outputs a video excitement score. This video excitement score is the video excitement score corresponding to the target segment in the video segment to be analyzed.

[0119] Compared with existing technologies, this invention is applicable to the needs of different video processing scenarios. It uses a pre-trained convolutional network to accurately score the video features of the video segment to be analyzed, thereby obtaining the video performance score of the target segment in the video segment to be analyzed. This achieves accurate analysis of the video segment's performance, improves the accuracy of the target segment's video performance score, and the convolutional network can achieve performance scoring by processing any video feature, making the technical solution provided by this invention more applicable.

[0120] Reference Figure 6 The second flowchart of a video processing method provided by an embodiment of the present invention is shown, specifically including:

[0121] Step 101: According to the preset splitting length, the video segment to be analyzed is split into multiple first sub-video segments, wherein the first sub-video segments include the target segments in the video segment to be analyzed.

[0122] Step 102: Perform feature processing on each first sub-video segment to obtain the feature information corresponding to the first sub-video segment.

[0123] Step 103: Generate video features of the video segment to be analyzed based on the feature information corresponding to the first sub-video segment; among them, the feature information of the target segment accounts for the highest proportion in the video features.

[0124] Step 104: Based on the video features of the video segment to be analyzed, obtain the video excitement score corresponding to the target segment in the video segment to be analyzed.

[0125] Steps 101-104 above will not be repeated here, referring to the previous discussion.

[0126] Step 105: According to the second preset splitting length, split each first sub-video segment to obtain multiple second sub-video segments, wherein the second preset splitting length is less than the preset splitting length.

[0127] In this embodiment of the invention, in order to further improve the accuracy of the video performance score of the target segment, each first sub-video segment is further split into smaller segments, so that the granularity of the video performance calculation result is smaller, thereby reducing the error of the predicted video performance score of the target segment.

[0128] Specifically, in this embodiment, each first sub-video segment is split according to a second preset splitting length to obtain multiple second sub-video segments, wherein the second preset splitting length is less than a preset splitting length. It should be noted that the second preset splitting length is determined based on the duration of the target segment to be analyzed. The duration of the target segment depends on the user's video analysis requirements. The fact that the second preset splitting length is less than the preset splitting length allows for further splitting of the target segment; no specific limitations are imposed here.

[0129] Step 106: Obtain the video quality scores of multiple second sub-video segments.

[0130] In this embodiment of the invention, obtaining video performance scores for multiple second sub-video segments may specifically include: performing feature processing on each second sub-video segment to obtain feature information corresponding to the second sub-video segment, and inputting the feature information corresponding to the second sub-video segment into a pre-trained convolutional network to obtain video performance scores for multiple second sub-video segments.

[0131] Step 107: Perform a weighted average of the video performance scores of multiple second sub-video segments to generate the video performance score of the target segment in the video segments to be analyzed.

[0132] Specifically, the video quality scores of multiple second sub-video segments are weighted and averaged to generate the video quality score of the target segment in the video segment to be analyzed. For example, the 4-second first sub-video segment is further divided into 2-second second sub-video segments to obtain the quality analysis results of each 2-second second sub-video segment. The weighted average of the quality scores of the two second sub-video segments of the 4-second target segment is used to represent the quality analysis score of the target segment.

[0133] Preferably, for the 30-32s brilliance score of the target segment in the video segment to be analyzed, the 28-32s brilliance analysis score can be obtained by using the 20-40s video segment to be analyzed, and the 30-34s brilliance analysis score can be calculated by using the 22-42s video segment to be analyzed. Then, the average of the 28-32s and 30-34s brilliance scores is taken as the final brilliance analysis score of the target segment within the 30-32s time period.

[0134] Compared with the prior art, the embodiments of the present invention, while achieving the beneficial effects of the above embodiments, further reduce the analysis granularity of the short video by further splitting the video segment to be analyzed, and perform weighted averaging on the video performance scores of the second sub-video segments of the target segment, thereby generating a more accurate video performance score of the target segment in the video segment to be analyzed, improving the accuracy of the video performance score of the target segment, solving the problems of missing contextual content due to short video duration and excessive error in performance score due to long video duration, and achieving accurate analysis of the performance of video segments.

[0135] Reference Figure 7 , Figure 7 This is a structural block diagram of a video processing apparatus 500 provided in an embodiment of this application, such as... Figure 7 As shown, the device may include:

[0136] The first processing module 501 is used to split the video segment to be analyzed according to a preset splitting length to obtain multiple first sub-video segments, wherein the first sub-video segments include the target segments in the video segment to be analyzed.

[0137] The second processing module 502 is used to perform feature processing on each of the first sub-video segments to obtain feature information corresponding to the first sub-video segments.

[0138] The first generation module 503 is used to generate video features of the video segment to be analyzed based on the feature information corresponding to the first sub-video segment; wherein, the feature information of the target segment accounts for the highest proportion among the video features;

[0139] The first acquisition module 504 is used to acquire the video excitement score corresponding to the target segment in the video segment to be analyzed based on the video features of the video segment to be analyzed.

[0140] Optionally, the second processing module 502 includes:

[0141] The extraction submodule is used to extract a preset frame image from each of the first sub-video segments according to a sliding window of a preset size;

[0142] The storage submodule is used to generate the image features of the preset frame image and store the correspondence between the image features of the preset frame image and each of the first sub-video segments;

[0143] The feature acquisition submodule is used to acquire feature information corresponding to the first sub-video segment based on the image features of the preset frame image and the correspondence.

[0144] Optionally, the first generation module 503 includes:

[0145] The processing submodule is used to input the feature information corresponding to multiple first sub-video segments into a preset conversion model, so that the preset conversion model performs conversion processing on the feature information.

[0146] The feature vector acquisition submodule is used to acquire feature vectors representing a preset output ratio of multiple first sub-video segments from the preset conversion model.

[0147] The generation submodule is used to merge the feature vectors of multiple first sub-video segments with a preset output quantity ratio to generate the video features of the video segment to be analyzed.

[0148] Optionally, the first generation module 503 further includes:

[0149] The submodule is invoked to call multiple preset conversion models corresponding to the first sub-video segment;

[0150] The setting submodule is used to determine the output ratio of multiple first sub-video segments corresponding to a preset conversion model based on the position of the target segment in the video segment to be analyzed. The output ratio of the feature information of the first sub-video segments is inversely proportional to the distance between the first sub-video segments and the target segment, and the preset output ratio of the feature information of the target segment is the highest.

[0151] Optionally, the first acquisition module 504 includes:

[0152] The prediction submodule is used to input the video features of the video segment to be analyzed into a pre-trained convolutional network to predict the video excitement score;

[0153] The score acquisition submodule is used to acquire the video brilliance score corresponding to the target segment in the video segment to be analyzed output by the convolutional network.

[0154] Optionally, the first acquisition module 504 further includes:

[0155] The first acquisition submodule is used to acquire training data for training the convolutional network, wherein the training data includes video features of historical video segments to be analyzed;

[0156] The second acquisition submodule is used to input the training data into the convolutional network to predict the video excitement score and obtain the predicted video excitement score output by the convolutional network.

[0157] The loss determination submodule is used to call the loss function to determine the score loss between the predicted video quality score and the actual video quality score;

[0158] The parameter adjustment submodule is used to perform backpropagation on the convolutional network based on the score loss, and adjust the network parameters of the convolutional network.

[0159] The third acquisition submodule is used to acquire the trained convolutional network when the score loss is less than a preset threshold.

[0160] Optionally, the device further includes:

[0161] The third processing module is used to split each of the first sub-video segments according to the second preset splitting length to obtain multiple second sub-video segments, wherein the second preset splitting length is less than the preset splitting length;

[0162] The second acquisition module is used to acquire video quality scores for multiple second sub-video segments;

[0163] The second generation module performs a weighted average of the video quality scores of multiple second sub-video segments to generate the video quality score of the target segment in the video segment to be analyzed.

[0164] This invention provides a video processing apparatus that, according to a preset splitting length, splits a video segment to be analyzed into multiple first sub-video segments. Each first sub-video segment undergoes feature processing to obtain corresponding feature information. Based on this feature information, video features of the video segment to be analyzed are generated. Based on these video features, a video quality score corresponding to a target segment within the video segment to be analyzed is obtained. This invention, by splitting the video segment to be analyzed into multiple short video segments, flexibly obtains target segments of the required duration, avoiding the need for manual selection of appropriate video durations based on video context due to fixed processing times in processing devices. Since the target segment has the highest proportion of video features in the video segment to be analyzed, video features that connect the preceding and following segments while highlighting the target segment's features are obtained. Based on these video features, a video quality score for the target segment is obtained, improving the accuracy of the target segment's video quality score. This solves the problems of missing context due to short video durations and excessively large quality score errors due to long video durations, achieving accurate analysis of video segment quality.

[0165] This invention also provides an electronic device. Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 8 As shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.

[0166] Memory 603 is used to store computer programs;

[0167] When processor 601 executes a program stored in memory 603, it performs the following steps:

[0168] Based on the preset splitting length, the video segment to be analyzed is split into multiple first sub-video segments, where the first sub-video segments include the target segments in the video segment to be analyzed.

[0169] Perform feature processing on each first sub-video segment to obtain the feature information corresponding to the first sub-video segment.

[0170] Based on the feature information corresponding to the first sub-video segment, video features of the video segment to be analyzed are generated; among these, the feature information of the target segment accounts for the highest proportion of the video features.

[0171] Based on the video features of the video segment to be analyzed, obtain the video excitement score corresponding to the target segment in the video segment to be analyzed.

[0172] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0173] The communication interface is used for communication between the aforementioned terminal and other devices.

[0174] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0175] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0176] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the video processing methods described in the above embodiments.

[0177] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or third-party database to another website, computer, server, or third-party database via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or third-party database that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid state disks (SSDs)).

[0178] It should be noted that, in this document, relational terms such as "first" and "first" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0179] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0180] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A video processing method, characterized in that, include: According to the preset splitting length, the video segment to be analyzed is split into multiple first sub-video segments, wherein the first sub-video segments include the target segments in the video segment to be analyzed; Perform feature processing on each of the first sub-video segments to obtain the feature information corresponding to the first sub-video segment; Based on the feature information corresponding to the first sub-video segment, video features of the video segment to be analyzed are generated; wherein, the feature information of the target segment accounts for the highest proportion among the video features; Based on the video features of the video segment to be analyzed, obtain the video quality score corresponding to the target segment in the video segment to be analyzed.

2. The method according to claim 1, characterized in that, The step of performing feature processing on each of the first sub-video segments to obtain the feature information corresponding to the first sub-video segment includes: Each of the first sub-video segments is captured using a sliding window of a preset size, and a preset frame image is extracted from the first sub-video segment; Generate the image features of the preset frame image, and store the correspondence between the image features of the preset frame image and each of the first sub-video segments; Based on the image features of the preset frame images and the corresponding relationship, the feature information corresponding to the first sub-video segment is obtained.

3. The method according to claim 1, characterized in that, The step of generating video features for the video segment to be analyzed based on the feature information corresponding to the first sub-video segment includes: The feature information corresponding to multiple first sub-video segments is respectively input into a preset conversion model so that the preset conversion model can convert the feature information. Obtain the feature vector of the preset output quantity ratio of multiple first sub-video segments output by the preset conversion model; The feature vectors of multiple first sub-video segments are merged according to a preset output quantity ratio to generate the video features of the video segment to be analyzed.

4. The method according to claim 3, characterized in that, Before generating the video features of the video segment to be analyzed based on the feature information corresponding to the first sub-video segment, the method further includes: Invoke multiple preset conversion models corresponding to the first sub-video segment; Based on the position of the target segment in the video segment to be analyzed, a preset output ratio of multiple first sub-video segments corresponding to a preset conversion model is determined. The preset output ratio of the feature information of the first sub-video segments is inversely proportional to the distance between the first sub-video segments and the target segment, and the preset output ratio of the feature information of the target segment is the highest.

5. The method according to claim 1, characterized in that, The step of obtaining the video excitement score corresponding to the target segment in the video segment to be analyzed based on the video features of the video segment to be analyzed includes: The video features of the video segment to be analyzed are input into a pre-trained convolutional network to predict the video excitement score; The video quality score of the video segment to be analyzed is obtained from the output of the convolutional network, and the video quality score is used as the video quality score of the target segment in the video segment to be analyzed.

6. The method according to claim 5, characterized in that, Before inputting the video features of the video segment to be analyzed into a pre-trained convolutional network for video performance score prediction, the method further includes: Obtain training data for training the convolutional network, wherein the training data includes video features of historical video segments to be analyzed and the actual video quality corresponding to the historical video segments to be analyzed; The video features of the historical video segment to be analyzed are input into the convolutional network to predict the video excitement score, and the predicted video excitement score output by the convolutional network is obtained. Call the loss function to determine the score loss between the predicted video quality score and the actual video quality score; Based on the fractional loss, the gradient of the network parameters of the convolutional network is calculated, and the network parameters are updated based on the gradient. Determine the score loss between the predicted video excitement score and the actual video excitement score after updating the network parameters. If the score loss is less than a preset loss threshold, obtain the trained convolutional network.

7. The method according to claim 1, characterized in that, After obtaining the video quality score corresponding to the target segment in the video segment to be analyzed based on the video features of the video segment to be analyzed, the method further includes: According to the second preset splitting length, each of the first sub-video segments is split to obtain multiple second sub-video segments, wherein the second preset splitting length is less than the preset splitting length; Obtain the video quality scores of multiple second sub-video segments; The video quality scores of the second sub-video segments of the multiple target segments are weighted and averaged to generate the video quality score of the target segments in the video segments to be analyzed.

8. A video processing apparatus, characterized in that, include: The first processing module is used to split the video segment to be analyzed according to a preset splitting length to obtain multiple first sub-video segments, wherein the first sub-video segments include the target segments in the video segment to be analyzed. The second processing module is used to perform feature processing on each of the first sub-video segments to obtain the feature information corresponding to the first sub-video segments. The first generation module is used to generate video features of the video segment to be analyzed based on the feature information corresponding to the first sub-video segment; wherein, the feature information of the target segment accounts for the highest proportion among the video features; The first acquisition module is used to acquire the video excitement score corresponding to the target segment in the video segment to be analyzed based on the video features of the video segment to be analyzed.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the video processing method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the video processing method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Video wonderful degree evaluation method and related equipment

    CN110267119A

  • Video highlight detection method and device, computer equipment and storage medium

    CN115205723A