Video processing method and device, electronic equipment and medium
By grouping the video set and calibrating multiple models, the uncertainty rate of video pairs is determined, and the target video set is selected, which solves the problem of poor training video effectiveness and improves training results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2023-07-21
- Publication Date
- 2026-05-12
AI Technical Summary
The effectiveness of training videos in existing technologies is poor, mainly because the training videos contain a large number of easily calibrated video pairs, resulting in poor training performance.
By grouping the video set to be processed, multiple preset calibration models are used to calibrate the videos in the video pairs respectively, obtaining multiple calibration information, and determining the uncertainty rate of the video pairs based on this calibration information, and selecting video pairs whose uncertainty rates meet the conditions as the target video set.
This improves the effectiveness of training videos, avoids a large number of easily calibrated videos, and enhances the training effect of the model.
Smart Images

Figure CN117115511B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a video processing method, apparatus, electronic device, and medium. Background Technology
[0002] With the development of computer technology, more and more data of various types are being generated on the Internet, including video data. For some video platforms, it is usually necessary to calibrate the videos on the platform in order to further process the videos based on the calibration scores. Accordingly, the calibration model is usually trained using training videos, and then the trained calibration model is directly used to calibrate the videos to be calibrated.
[0003] Training videos are usually obtained by labeling video pairs. However, in existing technologies, the video pairs in training videos are often randomly constructed, and there are a large number of video pairs that are relatively easy to label. This results in poor training performance when using such training videos. Therefore, how to improve the effectiveness of training videos has become an urgent problem to be solved. Summary of the Invention
[0004] This disclosure provides a video processing method, apparatus, electronic device, and medium to at least address the problem of how to improve the effectiveness of training videos. The technical solution of this disclosure is as follows:
[0005] According to a first aspect of the present disclosure, a video processing method is provided, comprising:
[0006] The videos in the video set to be processed are grouped to obtain multiple video pairs;
[0007] For any of the video pairs, the first video and the second video contained in the video pair are calibrated using multiple preset calibration models to obtain multiple calibration information; the calibration information includes first calibration information for the first video and second calibration information for the second video;
[0008] Based on the first calibration information and the second calibration information in each of the calibration information, the uncertainty rate of the video pair is determined; the uncertainty rate is used to characterize the difficulty of calibrating the video pair;
[0009] Video pairs whose uncertainty rates meet the calibration conditions are identified as target video pairs, and a target video set is obtained based on the target video pairs.
[0010] Optionally, determining the uncertainty rate of the video pair based on the first calibration information and the second calibration information in each of the calibration information includes:
[0011] When the absolute value of the information value of each of the calibration information is greater than a specified threshold, the calibration information that includes the first calibration information being greater than the second calibration information is taken as the first calibration result, and the calibration information that includes the first calibration information not being greater than the second calibration information is taken as the second calibration result; the information value of the calibration information is the difference between the first calibration information and the second calibration information in the calibration information;
[0012] The number of the first calibration results is determined as the first quantity, and the number of the second calibration results is determined as the second quantity;
[0013] The uncertainty rate of the video pair is determined based on the first quantity, the second quantity, and the total quantity of the first quantity and the second quantity.
[0014] Optionally, the calibration condition is that the uncertainty rate is equal to 1 / 2.
[0015] Optionally, the videos in the set of videos to be processed are grouped to obtain multiple video pairs, including:
[0016] Identify the video attributes of each video contained in the video set to be processed;
[0017] Based on the video attributes of each video, two videos with the same video attributes are grouped into the same group to obtain the multiple video pairs.
[0018] Optionally, the first video and the second video contained in the video pair are calibrated using multiple preset calibration models to obtain multiple calibration information, including:
[0019] The first video and the second video are respectively input into the plurality of preset calibration models, so that the plurality of preset calibration models calibrate the first video and the second video respectively;
[0020] The output data of each of the multiple preset calibration models are obtained and used as the multiple calibration information.
[0021] Optionally, the plurality of preset calibration models are generated based on multiple subsets of data; the method further includes:
[0022] Sampling is performed on the specified video set according to different sampling strategies;
[0023] For any of the sampling strategies, a subset of videos sampled according to the sampling strategy is determined as a data subset, thus obtaining the plurality of data subsets.
[0024] Optionally, the sampling locations and / or sampling numbers of the different sampling strategies may differ.
[0025] According to a second aspect of the present disclosure, a video processing apparatus is provided, comprising:
[0026] The grouping module is configured to group the videos contained in the video set to be processed, resulting in multiple video pairs.
[0027] The calibration module is configured to perform calibration on any given video pair using multiple preset calibration models, respectively, for the first video and the second video contained in the video pair, to obtain multiple calibration information; the calibration information includes first calibration information for the first video and second calibration information for the second video;
[0028] An uncertainty rate determination module is configured to determine the uncertainty rate of the video pair based on the first calibration information and the second calibration information in each of the calibration information; the uncertainty rate is used to characterize the difficulty of calibrating the video pair;
[0029] The target determination module is configured to determine video pairs whose uncertainty rates meet calibration conditions as target video pairs, so as to obtain a target video set based on the target video pairs.
[0030] Optionally, the uncertainty determination module is specifically configured to execute:
[0031] When the absolute value of the information value of each of the calibration information is greater than a specified threshold, the calibration information that includes the first calibration information being greater than the second calibration information is taken as the first calibration result, and the calibration information that includes the first calibration information not being greater than the second calibration information is taken as the second calibration result; the information value of the calibration information is the difference between the first calibration information and the second calibration information in the calibration information;
[0032] The number of the first calibration results is determined as the first quantity, and the number of the second calibration results is determined as the second quantity;
[0033] The uncertainty rate of the video pair is determined based on the first quantity, the second quantity, and the total quantity of the first quantity and the second quantity.
[0034] Optionally, the calibration condition is that the uncertainty rate is equal to 1 / 2.
[0035] Optionally, the grouping module is specifically configured to perform:
[0036] Identify the video attributes of each video contained in the video set to be processed;
[0037] Based on the video attributes of each video, two videos with the same video attributes are grouped into the same group to obtain the multiple video pairs.
[0038] Optionally, the calibration module is specifically configured to execute:
[0039] The first video and the second video are respectively input into the plurality of preset calibration models, so that the plurality of preset calibration models calibrate the first video and the second video respectively;
[0040] The output data of each of the multiple preset calibration models are obtained and used as the multiple calibration information.
[0041] Optionally, the plurality of preset calibration models are generated based on multiple subsets of data; the device further includes:
[0042] The sampling module is configured to sample a specified video set according to different sampling strategies.
[0043] The subset determination module is configured to perform, for any of the sampling strategies, determine a subset of video samples obtained by sampling according to the sampling strategy as a data subset, thereby obtaining the plurality of data subsets.
[0044] Optionally, the sampling locations and / or sampling numbers of the different sampling strategies may differ.
[0045] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0046] processor;
[0047] Memory used to store the processor's executable instructions;
[0048] The processor is configured to execute the instructions to implement the method as described in any one of the first aspects.
[0049] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided, wherein instructions in the storage medium, when executed by a processor of an electronic device, cause the electronic device to perform the method as described in any one of the first aspects.
[0050] According to a sixth aspect of the present disclosure, a computer program product is provided, the computer program product including readable program instructions that, when executed by a processor of an electronic device, cause the electronic device to perform the method as described in any one of the first aspects.
[0051] The technical solutions provided by the embodiments of this disclosure bring at least the following beneficial effects: In the embodiments of this disclosure, multiple video pairs are obtained by grouping the videos contained in the video set to be processed; for any video pair, the first video and the second video contained in the video pair are calibrated by multiple preset calibration models to obtain multiple calibration information; the calibration information includes first calibration information for the first video and second calibration information for the second video; based on the first calibration information and the second calibration information in each of the calibration information, the uncertainty rate of the video pair is determined; the uncertainty rate is used to characterize the difficulty of calibrating the video pair; the video pairs whose uncertainty rates meet the calibration conditions are determined as target video pairs, so as to obtain a target video set based on the target video pairs. In this way, by using multiple preset calibration models to calibrate video pairs in the video set to be processed, the uncertainty rate of any video pair can be obtained through the calibration information of multiple models. Thus, the embodiments of this disclosure can filter video pairs based on the uncertainty rate to obtain target video pairs that meet the calibration conditions. Since the uncertainty rate can characterize the difficulty of calibrating video pairs, the embodiments of this disclosure can filter out target video pairs that meet the requirements from the video set to be processed, avoiding the existence of a large number of easily calibrated videos in the target video set and improving the effectiveness of the target video set.
[0052] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0053] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0054] Figure 1 This is a flowchart illustrating a video processing method according to an exemplary embodiment;
[0055] Figure 2 This is a schematic diagram illustrating another video processing method according to an exemplary embodiment;
[0056] Figure 3 This is a block diagram illustrating a video processing apparatus according to an exemplary embodiment;
[0057] Figure 4 This is a block diagram illustrating an apparatus for video processing according to an exemplary embodiment;
[0058] Figure 5 This is a block diagram illustrating another apparatus for video processing according to an exemplary embodiment. Detailed Implementation
[0059] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0060] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0061] Figure 1 This is a flowchart illustrating a video processing method according to an exemplary embodiment, such as... Figure 1 As shown, the method may include the following steps:
[0062] Step 101: Group the videos in the video set to be processed to obtain multiple video pairs.
[0063] The video data contained in the aforementioned video set to be processed may be collected from any authorized video platform or may be a video set to be processed provided by an authorized user. This embodiment of the disclosure does not impose any restrictions on this.
[0064] The grouping operation described above can be performed by randomly grouping all videos in the video set to be processed, with each video group containing two videos to form a video pair. Alternatively, the groups can be grouped according to a preset grouping rule. This embodiment of the disclosure does not limit the specific implementation of the grouping.
[0065] Step 102: For any of the video pairs, the first video and the second video contained in the video pair are calibrated using multiple preset calibration models to obtain multiple calibration information; the calibration information includes first calibration information for the first video and second calibration information for the second video.
[0066] The aforementioned preset calibration model can be a model used to calibrate the quality of a video. It can be pre-trained using randomly constructed training samples. The number of the aforementioned preset calibration models can be 3, 4, 5, etc., and can be set according to actual needs. This embodiment of the disclosure does not limit this. Specifically, the aforementioned multiple calibration models can be trained based on different training samples, thus the aforementioned multiple preset calibration models can be different models.
[0067] Specifically, the aforementioned pre-defined calibration models can be trained using either single-stimulus labeled training samples or pair-ranked training samples. Single-stimulus labeling refers to labeling only one sample at a time, while pair-ranking labeling refers to randomly selecting two samples and performing binary classification on them (high or low). Understandably, a calibration model trained on single-stimulus labeled samples can calibrate each of the two videos in a video pair separately, outputting corresponding first and second calibration information. Similarly, a calibration model trained on pair-ranked samples can perform binary classification on the two videos in a video pair, directly outputting first and second calibration information. Specifically, the number of calibration information obtained for any video pair is consistent with the number of pre-defined calibration models, and correspondingly, the number of first and second calibration information obtained for any video pair is also consistent with the number of pre-defined calibration models.
[0068] In one optional embodiment, when multiple preset calibration models are all trained based on single-stimulus labeled samples, the operation of calibrating the first and second videos contained in the video pair using multiple preset calibration models to obtain multiple calibration information may specifically include the following steps:
[0069] S11. Input the first video and the second video into the plurality of preset calibration models respectively, so that the plurality of preset calibration models calibrate the first video and the second video respectively.
[0070] S12. Obtain the output data of each of the multiple preset calibration models as the multiple calibration information.
[0071] Specifically, the output data of any preset calibration model can represent the calibration information of the input video. Accordingly, when the input video is a first video, the output data is the first calibration information; when the input video is a second video, the output data is the second calibration information. Furthermore, by obtaining the output data of multiple preset calibration models, multiple calibration information generated for the video pair can be obtained. The first calibration information represents the quality of the first video, and the second calibration information represents the quality of the second video. It should be noted that the training samples used to train the preset calibration models can use the same scoring system, thereby ensuring that the scoring system of the calibration information output by each preset calibration model for any video pair is the same. This can be a 5-point system, a 10-point system, or even a percentage system, etc. This embodiment does not limit this.
[0072] Step 103: Based on the first calibration information and the second calibration information in each of the calibration information, determine the uncertainty rate of the video pair; the uncertainty rate is used to characterize the difficulty of calibrating the video pair.
[0073] Step 104: Determine the video pairs whose uncertainty rates meet the calibration conditions as target video pairs, and obtain the target video set based on the target video pairs.
[0074] The aforementioned uncertainty rate refers to the uncertainty in calibrating video pairs. It can be understood as the inconsistency in the calibration of the same video pair by different calibration models. Specifically, when the calibration information of the first and second videos in a video pair is relatively consistent among different calibration models, it indicates that the uncertainty of the video pair is low. Conversely, when the calibration information of the first and second videos is inconsistent among different calibration models, it indicates that the uncertainty of the video pair is high, and correspondingly, the difficulty of calibrating it is also high.
[0075] One approach is to determine whether the calibration information output by each preset calibration model is consistent by maintaining the order of the calibration information. This means that the order of the first and second calibration information output by each preset calibration model for the same video pair is consistent. For example, if the first calibration information output by all preset calibration models for the same video pair is higher or lower than the second calibration information, it indicates that the order of the video pair is consistent, the uncertainty of the video pair is 0, and the calibration difficulty is the lowest. Conversely, if the order of the first and second calibration information output by several preset calibration models is different from that of other preset calibration models, it indicates that the order of the calibration information is inconsistent, the uncertainty of the video pair is greater than 0, and the calibration difficulty is relatively high.
[0076] Understandably, a higher uncertainty rate for a video pair indicates greater uncertainty in the calibration of that video pair by each preset calibration model, thus indicating greater difficulty in calibrating the video pair. However, when the calibration of video pairs is difficult, using the video pair for training the model to be trained yields greater gains. Therefore, embodiments of this disclosure can filter all video pairs based on their uncertainty rates to select those that meet the calibration conditions. These calibration conditions can be an uncertainty rate greater than a preset threshold or a maximum uncertainty rate; the specific conditions can be set according to actual needs, and embodiments of this disclosure do not impose any limitations on this.
[0077] Furthermore, in this embodiment of the present disclosure, video pairs that meet the calibration conditions can be selected based on the uncertainty rate of all video pairs and the above calibration conditions. The target video pairs can be added to the target video set. Alternatively, other video pairs in the above-mentioned video set to be processed, except for the target video pairs, can be deleted, leaving only the target video pairs, thereby obtaining the above-mentioned target video set.
[0078] Furthermore, the aforementioned target video set can be used as a training video set. After manually annotating the target video pairs in the target video set, the annotated target video set can be used to train any calibration model to be trained. Since the uncertainty rate of each target video pair in the target video set meets the calibration conditions, the calibration of them is more difficult. Therefore, using the target video set to train the calibration model to be trained is more effective and can improve the training gain of the model.
[0079] In summary, the video processing method provided in this disclosure involves grouping videos in a set of videos to be processed to obtain multiple video pairs; for any given video pair, a first video and a second video contained in the video pair are calibrated using multiple preset calibration models to obtain multiple calibration information; the calibration information includes first calibration information for the first video and second calibration information for the second video; based on the first and second calibration information in each of the calibration information, the uncertainty rate of the video pair is determined; the uncertainty rate is used to characterize the difficulty of calibrating the video pair; and video pairs whose uncertainty rates meet the calibration conditions are determined as target video pairs, so as to obtain a target video set based on the target video pairs. In this way, by using multiple preset calibration models to calibrate video pairs in the video set to be processed, the uncertainty rate of any video pair can be obtained through the calibration information of multiple models. Thus, the embodiments of this disclosure can filter video pairs based on the uncertainty rate to obtain target video pairs that meet the calibration conditions. Since the uncertainty rate can characterize the difficulty of calibrating video pairs, the embodiments of this disclosure can filter out target video pairs that meet the requirements from the video set to be processed, avoiding the existence of a large amount of easily calibrated data in the target video set and improving the effectiveness of the target video set.
[0080] In one optional embodiment, the operation of determining the uncertainty rate of the video pair based on the first calibration information and the second calibration information in each of the calibration information may specifically include the following steps:
[0081] Step 201: When the absolute value of the information value of each of the calibration information is greater than the specified threshold, the calibration information that includes the first calibration information being greater than the second calibration information is taken as the first calibration result, and the calibration information that includes the first calibration information not being greater than the second calibration information is taken as the second calibration result; the information value of the calibration information is the difference between the first calibration information and the second calibration information in the calibration information.
[0082] The specified threshold mentioned above refers to the just noticeable difference (JND), which is the minimum difference that the human eye can perceive. Specifically, if the absolute value of the information value of the labeling information of a video pair is less than the JND, it indicates that the video pair cannot be classified by human, thus making it impossible to use the video pair to train the model. In other words, the video pair cannot be used as a training sample. Therefore, in this embodiment of the disclosure, the video pairs can be pre-screened using JND, and subsequent operations can be performed only if the absolute value of the information value of the labeling information is greater than the specified threshold.
[0083] Specifically, the information value of the above calibration information refers to the difference between the first calibration information and the second calibration information. Correspondingly, the absolute value of the information value of the above calibration information refers to the absolute value of the difference. The specified threshold can be set according to the actual situation, and can be set to 0.1, 0.2, 0.4, etc. This disclosure embodiment does not limit this.
[0084] Furthermore, for any video to PC i,j =(x i x j When the absolute value of each information value is greater than a specified threshold, the calibration results of the above-mentioned multiple preset calibration models can be either a first calibration result or a second calibration result. The first calibration result... Accordingly, the second calibration result in, Indicates the first video x i The calibration information, which is the first calibration information mentioned above, accordingly, Indicates the second video x j The calibration information is the second calibration information mentioned above.
[0085] Step 202: Determine the number of the first calibration results as the first quantity, and determine the number of the second calibration results as the second quantity.
[0086] Step 203: Determine the uncertainty rate of the video pair based on the first quantity, the second quantity, and the total quantity of the first quantity and the second quantity.
[0087] Here, the first quantity refers to the number of first calibration results among multiple calibration information, and the second quantity refers to the number of second calibration results among multiple calibration information. Specifically, since the absolute value of each information value is greater than a specified threshold, the number of first and second calibration results can characterize whether the calibration results of each preset calibration model for the same video pair are consistent. For example, when the number of first or second calibration results is 0, it indicates that all preset calibration models have consistent calibration results for the same video pair, and the uncertainty rate of the video pair is 0.
[0088] Since, when the absolute value of each calibration value is greater than the specified threshold, the calibration results of the multiple preset calibration models for any video pair only have two possibilities: a first calibration result and a second calibration result. Therefore, the total number is usually consistent with the number of preset calibration models. Specifically, the ratio of the first number to the total number can be used as the uncertainty rate, or the ratio of the second number to the total number can be used as the uncertainty rate. This disclosure does not limit this approach.
[0089] For example, taking the ratio of the first quantity to the total quantity as the uncertainty rate, the uncertainty rate can be calculated using the following formula:
[0090]
[0091] Wherein, m represents the number of calibration result types. Since, when the absolute value of each calibration value is greater than the specified threshold, the calibration results of the multiple preset calibration models for any video pair only have two types: the first calibration result and the second calibration result, the value of m in this embodiment can be 2. Of course, this embodiment only lists the case where the absolute value of each calibration value is greater than the specified threshold. In other cases, m can be taken according to actual needs, and this embodiment does not impose any restrictions on this. u represents the uncertainty rate of any video pair, p1 represents the first calibration result, and correspondingly, num... p1 The number representing the first calibration result, i.e., the first quantity, is num above. pk Characterizing the number of k-th calibration results, thus the above Represents the total quantity of the first quantity and the second quantity.
[0092] In this embodiment, when the absolute value of the information value of each calibration information is greater than a specified threshold, calibration information including first calibration information greater than second calibration information is taken as the first calibration result, and calibration information including first calibration information not greater than second calibration information is taken as the second calibration result; the information value of the calibration information is the difference between the first calibration information and the second calibration information in the calibration information; the number of the first calibration results is determined as the first quantity, and the number of the second calibration results is determined as the second quantity; based on the first quantity, the second quantity, and the total number of the first quantity and the second quantity, the uncertainty rate of the video pair is determined. Thus, since the differences between the videos in the video pair are small when the absolute value of the information value of each calibration information is not greater than the specified threshold, the effectiveness of the video pair is correspondingly low. Therefore, this embodiment can further improve the effectiveness of the video by determining the first calibration result and the second calibration result only when the absolute value of the information value is greater than the specified threshold. At the same time, determining the uncertainty rate of the video pair by the number of the first calibration result and the number of the second calibration result can ensure the accuracy of the uncertainty rate calculation.
[0093] In one alternative embodiment, the calibration condition described above is that the uncertainty rate is equal to 1 / 2.
[0094] Specifically, when the uncertainty rate is equal to 1 / 2, it indicates that among the calibration results of each preset calibration model for the same video pair, the number of first calibration results and second calibration results are the same. That is, there are 1 / 2 preset calibration models whose calibration information for the first video is higher than that for the second video. Correspondingly, the other 1 / 2 preset calibration models have higher calibration information for the second video than for the first video. At this time, the uncertainty of the calibration model in calibrating the video pair is the greatest, that is, the difficulty in calibrating the video pair is the greatest, and the gain when using the video pair to train the model to be trained is the greatest. Therefore, in this embodiment of the disclosure, the above calibration condition can be directly set to an uncertainty rate of 1 / 2.
[0095] In this embodiment of the disclosure, the calibration condition is that the uncertainty rate is equal to 1 / 2. By setting the calibration condition to an uncertainty rate of 1 / 2, the uncertainty of the determined target video pair is maximized, making its calibration most difficult. This maximizes the gain when using the video pair to train the model, thereby maximizing the effectiveness of the selected target video pair and ultimately maximizing the training effect of the model using the target video pair.
[0096] In one optional embodiment, the operation of grouping the videos in the video set to be processed to obtain multiple video pairs may specifically include the following steps:
[0097] Step 301: Identify the video attributes of each video contained in the video set to be processed.
[0098] Step 302: Based on the video attributes of each video, two videos with the same video attributes are grouped into the same group to obtain the multiple video pairs.
[0099] The aforementioned video attributes can be attribute information of the video in different dimensions or from different angles, such as video type, video size, etc. The number of video attributes can be one or more, and can be set according to actual needs; this disclosure does not limit this. Specifically, the operation of identifying video attributes can select different identification methods based on different video attributes. For example, the video type can be identified using a video type identification model, and the video size can be identified by reading the video length. It is understood that by grouping two videos with the same video attributes into the same group, the video pairs in the same group can be made more similar to a certain extent, thereby further increasing the difficulty of calibrating the constructed video pairs and improving the efficiency of the constructed video pairs.
[0100] In one optional embodiment, the aforementioned plurality of preset calibration models are generated based on a plurality of data subsets; this embodiment may further include the following steps:
[0101] Step 401: Sample the specified video set according to different sampling strategies.
[0102] Step 402: For any of the sampling strategies, determine the video subset obtained by sampling according to the sampling strategy as a data subset, and obtain the multiple data subsets.
[0103] The aforementioned specified video set can be pre-constructed and contains a large number of training samples. These samples may include single-stimulus labeled video samples or double-stimulus labeled video pairs; this embodiment does not impose any limitations on this. The double-stimulus labeled video pairs can be randomly constructed. The aforementioned sampling strategy refers to the sampling method for the specified video set and may include parameters such as sampling location and sampling quantity. Different sampling strategies result in different data subsets.
[0104] Furthermore, after obtaining multiple data subsets, the initial candidate models can be trained using each subset to obtain the aforementioned multiple preset calibration models. Specifically, one sampling strategy corresponds to one data subset, and one data subset corresponds to one preset calibration model; therefore, the number of preset calibration models is the same as the number of sampling strategies and data subsets. The initial candidate models can be pre-built and can be models used for calibrating video data.
[0105] Specifically, since the sampling strategies corresponding to each data subset are different, the training video samples contained in each data subset are also different. Thus, by training the candidate model using multiple different data subsets, the model parameters of the multiple preset calibration models obtained will also have different degrees of difference. Therefore, the above-mentioned multiple preset calibration models are multiple different preset calibration models. This allows different preset calibration models to be used to calibrate the same video pair, thereby improving the accuracy of calibrating the same video pair.
[0106] In this embodiment, a specified video set is sampled according to different sampling strategies. For any sampling strategy, a subset of videos obtained by sampling according to the sampling strategy is determined as a data subset, resulting in multiple data subsets. Thus, multiple different data subsets can be obtained from the same specified video set using different sampling strategies. Multiple different preset calibration models can be obtained from these multiple different data subsets, allowing different preset calibration models to be used to calibrate the same video pair, thereby improving the accuracy of calibration for the same video pair.
[0107] In one alternative embodiment, each training video in the specified video set carries reference calibration information. Specifically, this embodiment can generate the plurality of preset calibration models through the following steps:
[0108] S21. Each of the data subsets is taken as a target subset, and the initial candidate model is trained based on the target subset.
[0109] The aforementioned reference calibration information can be obtained in advance through manual annotation.
[0110] In this embodiment of the disclosure, each data subset can be used as a target subset in turn, and the target subset can be used to train the above-mentioned candidate model.
[0111] The above-described operation of training the initial candidate model based on the target subset may specifically include the following steps in this embodiment:
[0112] S22. Input the training videos in the target subset into the initial candidate model in sequence.
[0113] S23. Obtain the prediction calibration information generated by any of the training videos in the target subset for the candidate model.
[0114] S24. Adjust the model parameters of the candidate model based on the difference between each of the predicted calibration information and the corresponding reference calibration information until the training stopping condition is met, and then use the current candidate model as a preset calibration model.
[0115] For steps S22 to S23 above, the candidate model can be used to calibrate each training video in the target subset in sequence to obtain the prediction calibration information of each training video. Specifically, any training video can be used as the input value of the candidate model, and the output value of the candidate model can be used as the prediction calibration information of the training video.
[0116] The training termination condition mentioned above can be that the number of training iterations reaches a preset threshold, or the difference between the predicted calibration information and the corresponding reference calibration information reaches a preset difference threshold, or the preset loss function value of the model reaches a minimum, etc. It can be set according to actual needs, and this disclosure embodiment does not limit it.
[0117] Specifically, if the above-mentioned training stop conditions are not met, the adjusted candidate model can be used to continue the training operation until the training stop conditions are met and the training is determined to be over. The current candidate model can be used as a preset calibration model.
[0118] In one alternative embodiment, the different sampling strategies have different sampling locations and / or sampling quantities.
[0119] Here, the sampling position refers to the location of the sampled training video, and the sampling quantity refers to the number of sampled training videos. Specifically, the training videos included in the specified video set can be arranged sequentially, so the sampling position can include the start position and / or end position of the sampling, or the sampling position can be multiple scattered positions. The sampling position and sampling quantity can be set according to actual needs.
[0120] In this embodiment of the disclosure, the different sampling strategies involve different sampling positions and / or sampling numbers. Therefore, different data subsets can be obtained through different sampling positions and / or sampling numbers, ensuring that the multiple pre-calibrated models trained are different.
[0121] In one alternative embodiment, the present disclosure may further include the following steps:
[0122] S31. Obtain the target labeling information of each target video pair contained in the target video set.
[0123] S32. Based on the target video pairs and the target calibration information of each target video pair, train the calibration model to be trained to obtain the target calibration model; the target calibration model is used to calibrate the videos.
[0124] The target calibration information mentioned above can be obtained through manual annotation, or by sequentially outputting each target video pair contained in the target video set and determining the target calibration information of each target video pair by receiving input information from relevant personnel. Alternatively, the target calibration information can also be obtained through other calibration algorithms or calibration models. This disclosure embodiment does not limit this.
[0125] The aforementioned calibration model to be trained can be any model used to calibrate the video, it can be pre-built, it can be any one of the above-mentioned multiple preset calibration models, or it can be other calibration models.
[0126] Specifically, since the target video pairs included in the above-mentioned target video set are selected based on uncertainty rate and are video pairs with high labeling difficulty, this embodiment of the disclosure uses the target video set as the training sample of the labeling model to be trained. Compared with randomly mined training video pairs, this can improve the effectiveness of model training.
[0127] Specifically, the above training operation may involve sequentially inputting each target video pair into the calibration model to be trained, adjusting the model parameters of the calibration model to be trained based on the difference between the predicted calibration information output by the calibration model to be trained and the corresponding target calibration information, until a preset training stopping condition is reached, at which point the current calibration model to be trained is used as the target calibration model. The training stopping condition may be that the number of training iterations reaches a preset threshold, the difference between the predicted calibration information and the corresponding target calibration information reaches a preset difference threshold, or the preset loss function value of the model reaches its minimum, etc., and can be set according to actual needs; this embodiment does not impose any limitations on this.
[0128] In one scenario, the aforementioned video labeling refers to Video Quality Assessment (VQA). Specifically, with the rise of video-based social media, the user's video consumption experience (Quality of Experience, QoE) has gradually become an indispensable and important component. The viewing quality of a video directly or indirectly affects the user's video consumption experience; therefore, how to measure the physical quality of a video has become one of the core issues in the audio-visual field. VQA, as an objective quality labeling method, can effectively label the physical quality of videos. It has undergone technological evolution from manually designed features to deep network-based techniques and has made significant progress on mainstream industry datasets (KonIQ10k / YouTube UGC, etc.). However, unlike typical visual tasks such as image classification, which generally have millions of training data sets, the VQA field is limited by the difficulty and high cost of labeling. The amount of training data is generally in the thousands, and a small number are in the tens of thousands. The lack of datasets limits the large-scale application of VQA algorithms in industry.
[0129] Rank Image Quality Assessment (RankIQA) addresses the limitation of small datasets in the VQA domain by automatically constructing pairings based on different levels of distortion to aid training. However, real-world distortion is often a mixture of different types, and pre-training with a single pair has poor transferability in real-world scenarios. Better generalization performance is often achieved by adding relative-ranking information within the batch to assist model training. However, the source of relative-ranking information still comes from the Mean Opinion Score (MOS) labels obtained from single stimulus annotations, which remains limited by the small dataset size.
[0130] Existing technologies typically use a pair-rank method for labeling, which is simpler than single-stimulus labeling, requiring only a binary classification of good or bad, significantly reducing the time and number of participants required for single-stimulus labeling. However, in traditional pair-rank labeling, randomly constructing pairs is highly inefficient. Since even the best VQA algorithms can maintain order for about 70%, theoretically, 30% of the 100 pairs labeled are hard cases, meaning 30% of the effective data contributes to the model's learning, resulting in low labeling efficiency.
[0131] Figure 2 This is a schematic diagram illustrating another video processing method according to an exemplary embodiment, such as... Figure 2As shown, this embodiment of the present disclosure changes the sampling strategy for a specified video set. Based on different sampling strategies (data sampling strategies 1-4), the initial candidate model (AI model) can be trained, resulting in multiple different VQA models (VQA sub-models 1-4). VQA has relatively strong randomness in order preservation (Spearman rank-order correlation coefficient, SROCC) within the subdivided MOS small score range (4-5 points). This embodiment of the present disclosure uses multiple models to label the videos in the video set to be processed, and filters them by specifying thresholds and labeling conditions. It can mine hard samples based on the principle of maximizing entropy, improve the effectiveness of pair labeling, and reduce labeling costs. Meanwhile, after conducting calibration experiments on the videos of the video set to be processed and the target video set using the same calibration model, this embodiment of the present disclosure shows that the order preservation property of the calibration model for the video set to be processed is greater than 75%, meaning that there are less than 25% of valid video pairs in the video set to be processed. The order preservation property of the calibration model for the target video set is 45%, meaning that there are 55% of valid video pairs in the target video set. It can be seen that this embodiment of the present disclosure improves the effectiveness of training videos. At the same time, this embodiment of the present disclosure shows through experiments that the training efficiency of any calibration model using the target video set can be improved by 40% compared to the video set to be processed.
[0132] Figure 3 This is a block diagram illustrating a video processing apparatus according to an exemplary embodiment, such as... Figure 3 As shown, the device 50 may include:
[0133] Grouping module 501 is configured to group the videos contained in the video set to be processed, resulting in multiple video pairs;
[0134] The calibration module 502 is configured to perform calibration on any video pair using multiple preset calibration models, respectively, for the first video and the second video contained in the video pair, to obtain multiple calibration information; the calibration information includes first calibration information for the first video and second calibration information for the second video;
[0135] Uncertainty rate determination module 503 is configured to determine the uncertainty rate of the video pair based on the first calibration information and the second calibration information in each of the calibration information; the uncertainty rate is used to characterize the difficulty of calibrating the video pair;
[0136] The target determination module 504 is configured to determine video pairs whose uncertainty rates meet calibration conditions as target video pairs, so as to obtain a target video set based on the target video pairs.
[0137] In one alternative embodiment, the uncertainty determination module 503 is specifically configured to perform:
[0138] When the absolute value of the information value of each of the calibration information is greater than a specified threshold, the calibration information that includes the first calibration information being greater than the second calibration information is taken as the first calibration result, and the calibration information that includes the first calibration information not being greater than the second calibration information is taken as the second calibration result; the information value of the calibration information is the difference between the first calibration information and the second calibration information in the calibration information;
[0139] The number of the first calibration results is determined as the first quantity, and the number of the second calibration results is determined as the second quantity;
[0140] The uncertainty rate of the video pair is determined based on the first quantity, the second quantity, and the total quantity of the first quantity and the second quantity.
[0141] In one alternative embodiment, the calibration condition is that the uncertainty rate is equal to 1 / 2.
[0142] In one alternative embodiment, the grouping module 501 is specifically configured to perform:
[0143] Identify the video attributes of each video contained in the video set to be processed;
[0144] Based on the video attributes of each video, two videos with the same video attributes are grouped into the same group to obtain the multiple video pairs.
[0145] In one alternative embodiment, the calibration module 502 is specifically configured to perform:
[0146] The first video and the second video are respectively input into the plurality of preset calibration models, so that the plurality of preset calibration models calibrate the first video and the second video respectively;
[0147] The output data of each of the multiple preset calibration models are obtained and used as the multiple calibration information.
[0148] In one alternative embodiment, the plurality of preset calibration models are generated based on a plurality of data subsets; the device 50 further includes:
[0149] The sampling module is configured to sample a specified video set according to different sampling strategies.
[0150] The subset determination module is configured to perform, for any of the sampling strategies, determine a subset of video samples obtained by sampling according to the sampling strategy as a data subset, thereby obtaining the plurality of data subsets.
[0151] In one alternative embodiment, the different sampling strategies have different sampling locations and / or sampling quantities.
[0152] In summary, the video processing apparatus provided in this embodiment of the present disclosure obtains multiple video pairs by grouping the videos contained in the video set to be processed; for any video pair, a first video and a second video contained in the video pair are calibrated using multiple preset calibration models to obtain multiple calibration information; the calibration information includes first calibration information for the first video and second calibration information for the second video; based on the first calibration information and the second calibration information in each of the calibration information, the uncertainty rate of the video pair is determined; the uncertainty rate is used to characterize the difficulty of calibrating the video pair; video pairs whose uncertainty rates meet the calibration conditions are determined as target video pairs, so as to obtain a target video set based on the target video pairs. In this way, by using multiple preset calibration models to calibrate video pairs in the video set to be processed, the uncertainty rate of any video pair can be obtained through the calibration information of multiple models. Thus, the embodiments of this disclosure can filter video pairs based on the uncertainty rate to obtain target video pairs that meet the calibration conditions. Since the uncertainty rate can characterize the difficulty of calibrating video pairs, the embodiments of this disclosure can filter out target video pairs that meet the requirements from the video set to be processed, avoiding the existence of a large number of easily calibrated videos in the target video set and improving the effectiveness of the target video set.
[0153] According to one embodiment of this disclosure, an electronic device is provided, including: a processor and a memory for storing processor-executable instructions, wherein the processor is configured to perform, when executing, steps in the video processing method as described in any of the above embodiments.
[0154] According to one embodiment of this disclosure, a computer-readable storage medium is also provided, which, when executed by a processor of an electronic device, enables the electronic device to perform steps in the video processing method as described in any of the above embodiments.
[0155] According to one embodiment of this disclosure, a computer program product is also provided, which includes readable program instructions that, when executed by a processor of an electronic device, enable the electronic device to perform steps in the video processing method as described in any of the above embodiments.
[0156] Figure 4This is a block diagram illustrating an apparatus for video processing according to an exemplary embodiment. The apparatus 800 may include a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output interface 812, a sensor component 814, a communication component 816, and a processor 820. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the video processing method described above. In an exemplary embodiment, a storage medium including instructions is also provided, such as the memory 804 including instructions, which can be executed by the processor 820 of the apparatus 800 to complete the method described above. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.
[0157] Figure 5 This is a block diagram illustrating another apparatus for video processing according to an exemplary embodiment.
[0158] The device 900 may include a processing component 922, a memory 932, an input / output interface 958, a network interface 950, and a power supply component 926. The device 900 may be provided as a server. The application program stored in the memory 932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 922 is configured to execute instructions to perform the aforementioned video processing method.
[0159] All user information (including but not limited to user device information, user personal information, etc.) and related data involved in this disclosure are information authorized by the user or by the parties involved.
[0160] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0161] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A video processing method, characterized in that, The method includes: The videos in the video set to be processed are grouped to obtain multiple video pairs; For any of the video pairs, the first video and the second video contained in the video pair are calibrated using multiple preset calibration models to obtain multiple calibration information; the calibration information includes first calibration information for the first video and second calibration information for the second video; Based on the first calibration information and the second calibration information in each of the calibration information, the uncertainty rate of the video pair is determined; the uncertainty rate is used to characterize the difficulty of calibrating the video pair; Video pairs whose uncertainty rates meet the calibration conditions are identified as target video pairs, and a target video set is obtained based on the target video pairs; Determining the uncertainty rate of the video pair based on the first calibration information and the second calibration information in each of the calibration information includes: When the absolute value of the information value of each of the calibration information is greater than a specified threshold, the calibration information that includes the first calibration information being greater than the second calibration information is taken as the first calibration result, and the calibration information that includes the first calibration information not being greater than the second calibration information is taken as the second calibration result; the information value of the calibration information is the difference between the first calibration information and the second calibration information in the calibration information; The number of the first calibration results is determined as the first quantity, and the number of the second calibration results is determined as the second quantity; The uncertainty rate of the video pair is determined based on the first quantity, the second quantity, and the total quantity of the first quantity and the second quantity.
2. The method according to claim 1, characterized in that, The calibration condition is that the uncertainty rate is equal to 1 / 2.
3. The method according to claim 1, characterized in that, The videos in the set of videos to be processed are grouped to obtain multiple video pairs, including: Identify the video attributes of each video contained in the video set to be processed; Based on the video attributes of each video, two videos with the same video attributes are grouped into the same group to obtain the multiple video pairs.
4. The method according to claim 1, characterized in that, The process involves using multiple preset calibration models to calibrate the first and second videos contained in the video pair, respectively, to obtain multiple calibration information, including: The first video and the second video are respectively input into the plurality of preset calibration models, so that the plurality of preset calibration models calibrate the first video and the second video respectively; The output data of each of the multiple preset calibration models are obtained and used as the multiple calibration information.
5. The method according to claim 1, characterized in that, The multiple preset calibration models are generated based on multiple data subsets; the method further includes: Sampling is performed on the specified video set according to different sampling strategies; For any of the sampling strategies, a subset of videos sampled according to the sampling strategy is determined as a data subset, thus obtaining the plurality of data subsets.
6. The method according to claim 5, characterized in that, The different sampling strategies have different sampling locations and / or sampling numbers.
7. A video processing apparatus, characterized in that, The device includes: The grouping module is configured to group the videos contained in the video set to be processed, resulting in multiple video pairs. The calibration module is configured to perform calibration on any given video pair using multiple preset calibration models, respectively, for the first video and the second video contained in the video pair, to obtain multiple calibration information; the calibration information includes first calibration information for the first video and second calibration information for the second video; An uncertainty rate determination module is configured to determine the uncertainty rate of the video pair based on the first calibration information and the second calibration information in each of the calibration information; the uncertainty rate is used to characterize the difficulty of calibrating the video pair; The target determination module is configured to determine video pairs whose uncertainty rates meet calibration conditions as target video pairs, so as to obtain a target video set based on the target video pairs; Specifically, the uncertainty rate determination module is configured to execute: When the absolute value of the information value of each of the calibration information is greater than a specified threshold, the calibration information that includes the first calibration information being greater than the second calibration information is taken as the first calibration result, and the calibration information that includes the first calibration information not being greater than the second calibration information is taken as the second calibration result; the information value of the calibration information is the difference between the first calibration information and the second calibration information in the calibration information; The number of the first calibration results is determined as the first quantity, and the number of the second calibration results is determined as the second quantity; The uncertainty rate of the video pair is determined based on the first quantity, the second quantity, and the total quantity of the first quantity and the second quantity.
8. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device performs the method as described in any one of claims 1 to 6.