Image processing method and device, product and equipment
By augmenting the ternary image groups between video frames, augmenting image groups are generated to train the interpolation model, the problem of poor interpolation effect in the prior art is solved, and more efficient and accurate interpolation processing between video frames is achieved.
Patent Information
- Application Number
- CN202510096638.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The prior art is difficult to effectively perform excellent interpolation between video frames of videos through the interpolation model, especially in the case of diverse content of video frames.
By acquiring the ternary image group and performing augmentation processing based on the augmentation strategy set, an augmentation image group is generated to train the interpolation model. The method includes judging M types of motion augmentation strategies, and performing augmentation processing on the ternary image group according to the strategy, and generating an augmentation image group for training the interpolation model.
It improves the training effect of the interpolation model, improves the effect of interpolation between video frames, and enhances the robustness and accuracy of the interpolation model.
Smart Images

Figure CN119996602A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of video frame insertion, and in particular to an image processing method, device, product and equipment. Background Art
[0002] Frame interpolation refers to generating one or more new video frames between two consecutive video frames in a video to improve the smoothness of the video. In actual application scenarios, interpolation between video frames in a video can be achieved through a trained interpolation model. However, since the image content of the two video frames that need to be interpolated in a video is usually diverse, it is very challenging to use a trained interpolation model to perform excellent interpolation between the two video frames. Based on this, how to use an interpolation model to achieve excellent interpolation between consecutive video frames in a video is currently a hot issue. Summary of the invention
[0003] The present application provides an image processing method, apparatus, product and device, which can improve the training effect of the interpolation model, thereby improving the effect of interpolating between video frames of a video.
[0004] On the one hand, the present application provides an image processing method, the method comprising:
[0005] Acquire a ternary image group, where the ternary image group includes three video frames sampled from a sample video;
[0006] Based on the augmentation strategy set, the ternary image group is augmented to determine, and augmentation indication information for the ternary image group is obtained, wherein the augmentation strategy set includes N motion augmentation strategies, each motion augmentation strategy is used to indicate augmentation processing from a motion feature dimension of the ternary image group, and the augmentation indication information includes M motion augmentation strategies determined from the N motion augmentation strategies, where N is a positive integer, M is a non-negative integer, and M is less than or equal to N;
[0007] Acquiring an augmented image group based on M motion augmentation strategies and the ternary image group, wherein the method for acquiring the augmented image group includes: if M is greater than 0, augmenting the ternary image group from M motion feature dimensions of the ternary image group according to instructions of the M motion augmentation strategies to obtain the augmented image group; if M is equal to 0, using the ternary image group as the augmented image group;
[0008] The interpolation model is trained by using the augmented image group to obtain a trained interpolation model, and the trained interpolation model is used to perform interpolation processing between video frames of the video.
[0009] In one implementation, acquiring an augmented image group based on M motion augmentation strategies and a ternary image group includes:
[0010] If the M motion augmentation strategies are serial, then according to the serial order of the M motion augmentation strategies, the ternary image group is augmented in sequence from the M motion feature dimensions indicated by the M motion augmentation strategies to obtain an augmented image group;
[0011] If the M motion augmentation strategies are parallel, the ternary image groups are augmented respectively from the motion feature dimensions indicated by the M motion augmentation strategies to obtain M augmented image groups corresponding to the M motion augmentation strategies.
[0012] In one embodiment, the above method further comprises:
[0013] Acquire a first video frame and a second video frame in a video;
[0014] Calling the trained interpolation model to generate a newly added video frame between the first video frame and the second video frame based on the first video frame and the second video frame;
[0015] A newly added video frame is inserted between the first video frame and the second video frame of the video.
[0016] On one hand, the present application provides an image processing device, which includes:
[0017] An acquisition module, used for acquiring a ternary image group, where the ternary image group includes three video frames sampled from a sample video;
[0018] a judgment module, configured to perform augmentation judgment on the ternary image group based on an augmentation strategy set, and obtain augmentation indication information for the ternary image group, wherein the augmentation strategy set includes N motion augmentation strategies, each motion augmentation strategy is used to indicate augmentation processing from a motion feature dimension of the ternary image group, and the augmentation indication information includes M motion augmentation strategies obtained by judgment from the N motion augmentation strategies, where N is a positive integer, M is a non-negative integer, and M is less than or equal to N;
[0019] an augmentation module, configured to obtain an augmented image group based on M motion augmentation strategies and the ternary image group, wherein the method for obtaining the augmented image group comprises: if M is greater than 0, augmenting the ternary image group from M motion feature dimensions of the ternary image group according to instructions of the M motion augmentation strategies to obtain the augmented image group; and if M is equal to 0, using the ternary image group as the augmented image group;
[0020] The training module is used to train the interpolation model using the augmented image group to obtain the trained interpolation model, and the trained interpolation model is used to perform interpolation processing between video frames of the video.
[0021] In one implementation, the M motion augmentation strategies include a brightness augmentation strategy, and the motion feature dimension indicated by the brightness augmentation strategy is a brightness motion feature dimension;
[0022] The augmentation module obtains the augmented image group based on M motion augmentation strategies and the ternary image group, including:
[0023] Acquire a preset brightness variation range, and determine a first brightness variation amplitude for the ternary image group within the brightness variation range;
[0024] The ternary image group is augmented from a brightness motion feature dimension based on the first brightness change amplitude to obtain an augmented image group.
[0025] In one implementation, the ternary image group includes, in sequence, a first image at a first moment, a second image at a second moment, and a third image at a third moment, wherein the second moment is between the first moment and the third moment; and the augmentation module performs augmentation processing on the ternary image group from a brightness feature dimension based on the first brightness change amplitude to obtain an augmented image group, including:
[0026] Using the first image as the first image after brightness enhancement;
[0027] Calculating a second brightness change amplitude for the second image based on the second moment and the first brightness change amplitude, and performing weighted processing on the color pixel value of each pixel point in the second image using the second brightness change amplitude to obtain a second image after brightness enhancement;
[0028] Using the first brightness variation amplitude to perform weighted processing on the color pixel value of each pixel point in the third image, so as to obtain a third image with brightness enhancement;
[0029] The first image after brightness enhancement, the second image after brightness enhancement, and the third image after brightness enhancement are taken as an augmented image group.
[0030] In one implementation, the augmentation module determines the first brightness change amplitude for the ternary image group within the brightness change range, including:
[0031] Generate a first random number within the brightness variation range;
[0032] The generated first random number is used as the first brightness change amplitude for the ternary image group.
[0033] In one implementation, the ternary image group includes a first image, a second image, and a third image in sequence, the M motion augmentation strategies include a subtitle augmentation strategy, and the motion feature dimension indicated by the subtitle augmentation strategy is a subtitle motion feature dimension;
[0034] The augmentation module obtains the augmented image group based on M motion augmentation strategies and the ternary image group, including:
[0035] Selecting a target subtitle changing method for the ternary image group from a plurality of preset subtitle changing methods;
[0036] According to the target subtitle change mode, the first image, the second image and the third image are augmented from the subtitle motion feature dimension to obtain an augmented image group.
[0037] In one implementation, if the target subtitle change mode is a subtitle unchanged mode, the augmentation module augments the first image, the second image, and the third image from the subtitle motion feature dimension according to the target subtitle change mode to obtain an augmented image group, including:
[0038] Generate a first subtitle, and obtain a first image position among the first image, the second image, and the third image;
[0039] Adding the first subtitle at the position of the first image in the first image, the second image and the third image respectively, to obtain the first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation;
[0040] The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are taken as an augmented image group.
[0041] In one implementation, if the target subtitle change mode is a subtitle disappearance mode, the augmentation module augments the first image, the second image, and the third image from the subtitle motion feature dimension according to the target subtitle change mode to obtain an augmented image group, including:
[0042] Generate a first subtitle, and obtain a first image position in the first image and the second image;
[0043] Adding the first subtitle at the position of the first image in the first image and the second image respectively to obtain the first image after subtitle augmentation and the second image after subtitle augmentation;
[0044] using the third image as the third image after the captions are augmented;
[0045] The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are taken as an augmented image group.
[0046] In one implementation, if the target subtitle change mode is a subtitle appearance mode, the augmentation module augments the first image, the second image, and the third image from the subtitle motion feature dimension according to the target subtitle change mode to obtain an augmented image group, including:
[0047] The first image and the second image are used as the first image after subtitle augmentation and the second image after subtitle augmentation respectively;
[0048] Generate a first subtitle, obtain the first image position in the third image, add the first subtitle at the first image position in the third image, and obtain a third image after the subtitle is augmented;
[0049] The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are taken as an augmented image group.
[0050] In one implementation, if the target subtitle change mode is a subtitle switching mode, the augmentation module augments the first image, the second image, and the third image from the subtitle motion feature dimension according to the target subtitle change mode to obtain an augmented image group, including:
[0051] Generate a first subtitle and a second subtitle, and obtain a first image position in the first image, the second image, and the third image;
[0052] Adding the first subtitle at the position of the first image in the first image and the second image respectively to obtain the first image after subtitle augmentation and the second image after subtitle augmentation;
[0053] Adding a second subtitle at the position of the first image in the third image to obtain a third image with the subtitles augmented;
[0054] The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are taken as an augmented image group.
[0055] In one implementation, the ternary image group sequentially includes a first image at a first moment, a second image at a second moment, and a third image at a third moment, the second moment is between the first moment and the third moment, the M motion augmentation strategies include a text augmentation strategy, and the motion feature dimension indicated by the text augmentation strategy is a text motion feature dimension;
[0056] The augmentation module obtains the augmented image group based on M motion augmentation strategies and the ternary image group, including:
[0057] Selecting a target text motion mode for the ternary image group from a plurality of preset text motion modes;
[0058] A target text to be added is generated, and the first image, the second image, and the third image are augmented according to the target text motion mode and based on the target text from the text motion feature dimension to obtain an augmented image group.
[0059] In one implementation, if the target text motion mode is a text translation mode, the augmentation module augments the first image, the second image, and the third image according to the target text motion mode based on the target text from the text motion feature dimension to obtain an augmented image group, including:
[0060] Obtaining a text displacement corresponding to the target text and a second image position in the first image;
[0061] Adding the target text at the position of the second image in the first image to obtain the first image after text augmentation;
[0062] Calculating a third image position in the second image based on the second image position, the second moment, and the text displacement, and adding the target text at the third image position in the second image to obtain a second image after text augmentation;
[0063] Calculating a fourth image position in the third image based on the second image position and the text displacement, and adding target text at the fourth image position in the third image to obtain a third image after text augmentation;
[0064] The first image after text augmentation, the second image after text augmentation, and the third image after text augmentation are taken as an augmented image group.
[0065] In one implementation, if the target text motion mode is a text zoom mode, the augmentation module augments the first image, the second image, and the third image according to the target text motion mode based on the target text from the text motion feature dimension to obtain an augmented image group, including:
[0066] acquiring a fifth image position among the first image, the second image, and the third image;
[0067] Adding the target text at the fifth image position in the first image to obtain the first image after text augmentation;
[0068] Acquire a first zoom scale for the target text, and calculate a second zoom scale for the target text based on the first zoom scale and the second moment;
[0069] The target text is scaled using the second scaling scale to obtain the first scaled text, and the target text is scaled using the first scaling scale to obtain the second scaled text;
[0070] Adding the first scaled text at the fifth image position in the second image to obtain the text-augmented second image, and adding the second scaled text at the fifth image position in the third image to obtain the text-augmented third image;
[0071] The first image after text augmentation, the second image after text augmentation, and the third image after text augmentation are taken as an augmented image group.
[0072] In one implementation, if the target text motion mode is a text rotation mode, the augmentation module augments the first image, the second image, and the third image according to the target text motion mode based on the target text from the text motion feature dimension to obtain an augmented image group, including:
[0073] acquiring a sixth image position among the first image, the second image, and the third image;
[0074] Adding target text at the sixth image position in the first image, the second image and the third image respectively, to obtain a first image after adding text, a second image after adding text and a third image after adding text;
[0075] Obtaining a first rotation angle for the target text, and rotating the target text in the first image after the text is added by the first rotation angle to obtain the first image after the text is augmented;
[0076] Acquire a second rotation angle for the target text, calculate a third rotation angle for the target text based on the first rotation angle, the second rotation angle and the second moment, and rotate the target text in the second image after the text is added by the third rotation angle to obtain the second image after the text is augmented;
[0077] Calculating a fourth rotation angle for the target text based on the first rotation angle and the second rotation angle, and rotating the target text in the third image after the text is added by the fourth rotation angle to obtain the third image after the text is augmented;
[0078] The first image after text augmentation, the second image after text augmentation, and the third image after text augmentation are taken as an augmented image group.
[0079] In one implementation, the augmentation module acquires the augmented image group based on M motion augmentation strategies and the ternary image group, including:
[0080] If the M motion augmentation strategies are serial, then according to the serial order of the M motion augmentation strategies, the ternary image group is augmented in sequence from the M motion feature dimensions indicated by the M motion augmentation strategies to obtain an augmented image group;
[0081] If the M motion augmentation strategies are parallel, the ternary image groups are augmented respectively from the motion feature dimensions indicated by the M motion augmentation strategies to obtain M augmented image groups corresponding to the M motion augmentation strategies.
[0082] In one implementation, the judgment module performs augmentation judgment on the ternary image group based on the augmentation strategy set to obtain augmentation indication information for the ternary image group, including:
[0083] Obtain the augmentation judgment probability corresponding to each motion augmentation strategy in the augmentation strategy set;
[0084] Perform augmentation judgment on each motion augmentation strategy based on the augmentation judgment probability corresponding to each motion augmentation strategy, and obtain the judgment result corresponding to each motion augmentation strategy. The judgment result corresponding to any motion augmentation strategy is an adopted result or a non-adopted result.
[0085] The augmentation indication information is determined by determining that the motion augmentation strategy of the determined result is an adopted result, and the M types of motion augmentation strategies include the motion augmentation strategy of the determined result being an adopted result.
[0086] In one implementation, any one of the N motion augmentation strategies is a target motion augmentation strategy; the judgment module performs augmentation judgment on each motion augmentation strategy based on the augmentation judgment probability corresponding to each motion augmentation strategy, and a method of obtaining a judgment result corresponding to each motion augmentation strategy includes:
[0087] Obtain the total probability interval and the probability sub-interval corresponding to the augmented judgment probability corresponding to the target motion augmentation strategy within the total probability interval;
[0088] generating a second random number within the total probability interval;
[0089] If the second random number is within the probability subinterval, determining that the judgment result corresponding to the target motion augmentation strategy is the adopted result;
[0090] If the second random number is not within the probability subinterval, the judgment result corresponding to the target motion augmentation strategy is determined to be a non-adopted result.
[0091] In one implementation, the augmented image group includes a first augmented image, a second augmented image, and a third augmented image in sequence; and the training module uses the augmented image group to train the interpolation model to obtain the trained interpolation model, including:
[0092] Calling the interpolation model to generate a newly added image between the first augmented image and the third augmented image based on the first augmented image and the third augmented image;
[0093] Obtaining a generation loss of the interpolation model for the new image based on the new image and the second augmented image, where the generation loss is used to reflect the image difference between the new image and the second augmented image;
[0094] The model parameters of the interpolation model are corrected by generating losses to obtain the trained interpolation model.
[0095] In one embodiment, the image processing device further includes a frame insertion module, which is used to:
[0096] Acquire a first video frame and a second video frame in a video;
[0097] Calling the trained interpolation model to generate a newly added video frame between the first video frame and the second video frame based on the first video frame and the second video frame;
[0098] A newly added video frame is inserted between the first video frame and the second video frame of the video.
[0099] In one aspect, the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes a method in one aspect of the present application.
[0100] In one aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor executes the method in the above aspect.
[0101] According to one aspect of the present application, a computer program product is provided, the computer program product comprising a computer program, the computer program being stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the method provided in various optional modes such as the above-mentioned one aspect.
[0102] The present application can obtain a ternary image group, which includes three video frames sampled from a sample video; and can perform augmentation judgment on the ternary image group based on an augmentation strategy set to obtain augmentation indication information for the ternary image group, wherein the augmentation strategy set includes N motion augmentation strategies, each motion augmentation strategy is used to indicate augmentation processing from a motion feature dimension of the ternary image group, and the augmentation indication information includes M motion augmentation strategies judged from the N motion augmentation strategies, where N is a positive integer, M is a non-negative integer, and M is less than or equal to N; However, an augmented image group can be obtained based on M motion augmentation strategies and a ternary image group, wherein the method for obtaining the augmented image group includes: if M is greater than 0, the ternary image group is augmented from M motion feature dimensions of the ternary image group according to instructions of the M motion augmentation strategies to obtain an augmented image group; if M is equal to 0, the ternary image group is used as the augmented image group; and the augmented image group can be used to train an interpolation model to obtain a trained interpolation model, and the trained interpolation model is used to perform interpolation processing between video frames of the video. It can be seen that the method proposed in the present application can perform augmentation processing on the ternary image group used to train the interpolation model through N motion feature dimensions indicated by N motion augmentation strategies, thereby taking into account the change modes of each image in the ternary image group in multiple motion feature dimensions, and obtaining an augmented image group with rich change modes between images. Therefore, by training the interpolation model through the augmented image group obtained by augmentation, the robustness and accuracy of the trained interpolation model can be improved, and thus, the interpolation model obtained by training can also achieve excellent interpolation processing between video frames of the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0103] In order to more clearly illustrate the technical solutions in the present application or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0104] Figure 1 It is a structural diagram of a network architecture for video frame insertion provided by an embodiment of the present application;
[0105] Figure 2 It is a scene diagram of a training interpolation model provided in an embodiment of the present application;
[0106] Figure 3 It is a flowchart of an image processing method provided in an embodiment of the present application;
[0107] Figure 4 is a schematic diagram of another scene of training an interpolation model provided in an embodiment of the present application;
[0108] Figure 5 It is a schematic diagram of a process of augmenting a ternary image group using a brightness augmentation strategy provided by an embodiment of the present application;
[0109] Figure 6 This is a schematic diagram of the effect of augmenting a ternary image group using a brightness augmentation strategy provided by an embodiment of the present application. Figure 1 ;
[0110] Figure 7 This is a schematic diagram of the effect of augmenting a ternary image group using a brightness augmentation strategy provided by an embodiment of the present application. Figure 2 ;
[0111] Figure 8 It is a schematic diagram of a process of augmenting a ternary image group using a subtitle augmentation strategy provided by an embodiment of the present application;
[0112] Fig. 9 This is a schematic diagram of the effect of augmenting a ternary image group using a subtitle augmentation strategy provided by an embodiment of the present application. Figure 1 ;
[0113] Fig.10 This is a schematic diagram of the effect of augmenting a ternary image group using a subtitle augmentation strategy provided by an embodiment of the present application. Figure 2 ;
[0114] Fig.11 This is a schematic diagram of the effect of augmenting a ternary image group using a subtitle augmentation strategy provided by an embodiment of the present application. Figure 3 ;
[0115] Fig.12 This is a schematic diagram of the effect of augmenting a ternary image group using a subtitle augmentation strategy provided by an embodiment of the present application. Figure 4 ;
[0116] Fig.13 It is a schematic diagram of a process of augmenting a ternary image group using a text augmentation strategy provided by an embodiment of the present application;
[0117] Fig.14 This is a schematic diagram of the effect of augmenting a ternary image group using a text augmentation strategy provided in an embodiment of the present application. Figure 1 ;
[0118] Fig.15 This is a schematic diagram of the effect of augmenting a ternary image group using a text augmentation strategy provided in an embodiment of the present application. Figure 2 ;
[0119] Fig.16This is a schematic diagram of the effect of augmenting a ternary image group using a text augmentation strategy provided in an embodiment of the present application. Figure 3 ;
[0120] Fig.17 is a structural schematic diagram of an image processing device provided in an embodiment of the present application;
[0121] Fig.18 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0122] The following will be combined with the drawings in this application to clearly and completely describe the technical solutions in this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0123] All data collected in this application (such as ternary image groups, augmentation strategy sets, interpolation models and other related data) are collected with the consent and authorization of the owner of the data (such as users, institutions or enterprises), and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards in the relevant regions.
[0124] Here, the relevant technical concepts involved in this application are described:
[0125] EMA model: It is an efficient and advanced video interpolation method that can efficiently interpolate videos based on motion and appearance information extracted by inter-frame attention.
[0126] Frame interpolation: refers to generating one or more new video frames between two consecutive video frames in a video to improve the smoothness of the video.
[0127] See also Figure 1 , Figure 1 Schematic diagram of a network architecture for video frame insertion provided by an embodiment of the present application. Figure 1 As shown, the network architecture may include a server 200 and a terminal device cluster, and the terminal device cluster may include one or more terminal devices, and the number of terminal devices is not limited here. Figure 1 As shown, the multiple terminal devices may specifically include terminal device 1, terminal device 2, terminal device 3, ..., terminal device n; Figure 1 As shown, terminal device 1, terminal device 2, terminal device 3, ..., terminal device n can all be connected to the server 200 through a network, so that each terminal device can exchange data with the server 200 through the network connection.
[0128] like Figure 1 The server 200 shown can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (content distribution network), and big data and artificial intelligence platforms. The terminal device can be: a smart phone, a tablet computer, a laptop, a desktop computer, a smart TV, a car terminal, a smart home, and other smart terminals. The following takes the communication between the terminal device 1 and the server 200 as an example to describe the specific embodiments of the present application.
[0129] The server 200 can train an interpolation model, and can perform interpolation processing on a video that needs to be interpolated through the trained interpolation model, thereby obtaining an interpolated video. The terminal device 1 can have a video program, which can be used to play a video, such as an application, a web program, or a small program. The server 200 can send the interpolated video to the terminal device 1, so that the video program in the terminal device 1 can play the received interpolated video for the user to watch. The process of the server 200 training the interpolation model can be described as follows.
[0130] Please also see Figure 2 , Figure 2 is a schematic diagram of a scene for training an interpolation model provided in an embodiment of the present application. Figure 2 As shown above Figure 1 The server 200 in the method can obtain a ternary image group, which can be an original sample for training an interpolation model. The ternary image group can include three video frames sampled from a sample video, such as the three video frames can include a first image, a second image, and a third image in sequence. The server 200 can also obtain an augmentation strategy set, which can include N motion augmentation strategies for augmenting the ternary image group (such as including motion augmentation strategy 1 to motion augmentation strategy N here), where N is a positive integer. A motion augmentation strategy is used to indicate augmentation processing from a motion feature dimension of the ternary image group.
[0131] The server 200 may perform augmentation judgment on the ternary image group through the augmentation strategy set to obtain augmentation indication information for the ternary image group, and the augmentation indication information may include M motion augmentation strategies (such as including motion augmentation strategy 1 to motion augmentation strategy M here) obtained from the augmentation strategy set, where M is a non-negative integer and M is less than or equal to N. The server 200 may obtain an augmented image group through the M motion augmentation strategies and the ternary image group, including: if M is greater than 0, the server 200 may perform augmentation processing on the ternary image group from the M motion feature dimensions of the ternary image group according to the instructions of the M motion augmentation strategies to obtain an augmented image group of the ternary image group, and if M is equal to 0, the server 200 may directly use the ternary image group as its augmented image group. The augmented image group may include an augmented first image, an augmented second image, and an augmented third image.
[0132] The server 200 can use the augmented image group to train the interpolation model to obtain a trained interpolation model. The trained interpolation model can be used to perform interpolation processing between consecutive video frames of the video, so that a video with more coherent video images can be obtained through interpolation processing.
[0133] By adopting the method provided by the present application, the augmentation processing of the ternary image group can be realized through various motion augmentation strategies to obtain an augmented image group. The interpolation model obtained by training the augmented image group can be more suitable for interpolating between continuous video frames whose video images have changes in various motion feature dimensions, thereby improving the effect of interpolating the video.
[0134] See also Figure 3 , Figure 3 : is a flowchart of an image processing method provided by an embodiment of the present application. The execution subject in the embodiment of the present application may be an image processing device (referred to as a processing device), which may be a computer device or a computer device cluster composed of multiple computer devices, which may be a server, a terminal device, or other devices, and the present application does not limit this. Figure 3 As shown, the method may include:
[0135] Step S101 , obtaining a ternary image group, where the ternary image group includes three video frames sampled from a sample video.
[0136] Specifically, the processing device may obtain a ternary image group, which may be a sample for training an interpolation model (which may be an EMA model). There may be many ternary image groups (i.e., there may be many samples), and the specific number of ternary image groups may be determined according to actual application scenarios. Since the principles for processing each ternary image group are independent and the same, the number of ternary image groups will not be emphasized in the following process.
[0137] The ternary image group may include three images, and the three images may be three video frames sampled from the sample video, that is, the ternary image group may include three video frames sampled from the sample video, and the arrangement order of the three video frames in the ternary image group is the arrangement order of the three video frames in the sample video. There may be many sample videos, and the type of the sample video is not limited, and may be any video type, such as animation type, real-life type, etc. The three images in the ternary image group may be three continuous video frames in the sample video, or may be three discontinuous video frames in the sample video. If the three images in the ternary image group are three discontinuous video frames in the sample video, the sampling interval between the three video frames (such as the number of video frames in the interval) will not be too large, such as being less than or equal to the set sampling interval threshold.
[0138] Among them, the three images included in the ternary image group in sequence can be called the first image, the second image and the third image, that is, the ternary image group can include the first image, the second image and the third image in sequence. The first image can be at a first moment, the second image can be at a second moment, and the third image can be at a third moment, the second moment is between the first moment and the third moment, and the first moment, the second moment and the third moment can be the moments used by the present application to identify the arrangement order between the first image, the second image and the third image. Among them, the image sizes (i.e., image sizes) of the first image, the second image and the third image can all be the same, which is specifically embodied in that in the first image, the second image and the third image, the number of rows of pixels can be equal, and the number of columns of pixels can also be equal.
[0139] For example, the first moment may be moment 0, the third moment may be moment 1, and the second moment may be a moment between moment 0 and moment 1, that is, the second moment may be a moment between 0 and 1. The second moment may be represented as moment t, and t may be a value between 0 and 1 (but t is not equal to 0 or 1, such as t may be equal to 0.2, 0.3, 0.5 or 0.7, etc.), and the value of t may be used in the subsequent augmentation process of the ternary image group.
[0140] The specific value of the second moment t can be determined according to the sampling interval of the three images in the ternary image group in the sample video, and each ternary image group can carry its corresponding second moment t. The second moment t can be used to reflect the size relationship (such as proportional relationship) between the sampling interval from the first image to the second image and the sampling interval from the second image to the third image. Among them, each video frame in the sample video can have its own serial number, such as the serial number of the first video frame in the sample video can be 1, the serial number of the second video frame can be 2, the serial number of the third video frame can be 3, and so on.
[0141] For example, the first image, the second image and the third image in the ternary image group are obtained by sampling the first video frame, the third video frame and the seventh video frame in the sample video respectively, that is, the first image can be the first video frame, the second image can be the third video frame, and the third image can be the seventh video frame. Then, the second moment t can be (3-1) / (7-1) equal to 1 / 3, that is, the second moment t is equal to: the sampling interval from the second image to the first image (the sequence number 3 of the second image minus the sequence number 1 of the first image is equal to 2) divided by the sampling interval from the third image to the first image (the sequence number 7 of the third image minus the sequence number 1 of the first image is equal to 6).
[0142] For another example, the first image, the second image and the third image in the ternary image group are obtained by sampling the 1st video frame, the 4th video frame and the 7th video frame in the sample video respectively, that is, the first image can be the 1st video frame, the second image can be the 4th video frame, and the third image can be the 7th video frame. Then, the second moment t can be (4-1) / (7-1) which is equal to 1 / 2, that is, the second moment t is equal to: the sampling interval from the second image to the first image (the serial number 4 of the second image minus the serial number 1 of the first image is equal to 3) divided by the sampling interval from the third image to the first image (the serial number 7 of the third image minus the serial number 1 of the first image is equal to 6).
[0143] For another example, the first image, the second image and the third image in the ternary image group are obtained by sampling the first video frame, the second video frame and the third video frame in the sample video respectively, that is, the first image can be the first video frame, the second image can be the second video frame, and the third image can be the third video frame. Then, the second moment t can be (2-1) / (3-1) which is equal to 1 / 2, that is, the second moment t is equal to: the sampling interval from the second image to the first image (the sequence number of the second image 2 minus the sequence number of the first image 1 is equal to 1) divided by the sampling interval from the third image to the first image (the sequence number of the third image 3 minus the sequence number of the first image 1 is equal to 2).
[0144] As can be seen from the above, the three images in the ternary image group are not necessarily three consecutive video frames in the sample video, but there is a sequence between the three images, specifically including: the first image in the sample video is the video frame before the second image, and the second image is the video frame before the third image. Different values of t generated can make the effect of augmentation processing on the ternary image group different. For details, please refer to the relevant description in the subsequent embodiments. Therefore, the first image can also be expressed as I0, and the second image can be expressed as I t , the third image can represent I1.
[0145] Step S102, based on the augmentation strategy set, the ternary image group is augmented and judged to obtain augmentation indication information for the ternary image group, the augmentation strategy set includes N motion augmentation strategies, each motion augmentation strategy is used to indicate augmentation processing from a motion feature dimension of the ternary image group, and the augmentation indication information includes M motion augmentation strategies judged from the N motion augmentation strategies, N is a positive integer, M is a non-negative integer, and M is less than or equal to N.
[0146] Specifically, the processing device may perform augmentation judgment on the ternary image group through an augmentation strategy set to obtain augmentation indication information for the ternary image group. The augmentation strategy set may include N motion augmentation strategies, where N is a positive integer, and each motion augmentation strategy may be used to indicate augmentation processing from a motion feature dimension of the ternary image group, that is, a motion augmentation strategy may be used to indicate augmentation processing from a motion feature dimension of the ternary image group, and the augmentation processing may refer to augmentation processing of the ternary image group. The motion feature dimension may refer to the dimension of motion features between image frames of each image in the ternary image group.
[0147] For example, the N motion augmentation strategies may include but are not limited to: brightness augmentation strategy, subtitle augmentation strategy, and text augmentation strategy. The motion feature dimension indicated by the brightness augmentation strategy may be a brightness motion feature dimension, that is, the brightness augmentation strategy may be used to augment the ternary image group from the feature dimension of the brightness change (i.e., brightness motion) between the images in the ternary image group.
[0148] The motion feature dimension indicated by the above-mentioned subtitle augmentation strategy may be a subtitle motion feature dimension, that is, the subtitle augmentation strategy may be used to augment the ternary image group from the feature dimension of the subtitle change (ie, subtitle motion) between the images in the ternary image group.
[0149] The motion feature dimension indicated by the above text augmentation strategy may be a text motion feature dimension, that is, the text augmentation strategy may be used to augment the ternary image group from the feature dimension of text changes (ie, text motion) between the images in the ternary image group.
[0150] There may be no text and / or subtitle in each image in the above ternary image map.
[0151] And the augmentation judgment of the ternary image group through the augmentation strategy set can be to judge whether to use each motion augmentation strategy in the augmentation strategy set to perform augmentation processing on the ternary image group. Therefore, the augmentation indication information for the ternary image group may include M motion augmentation strategies, and the M motion augmentation strategies may include motion augmentation strategies that are judged from the above-mentioned N motion augmentation strategies and are needed to perform augmentation processing on the ternary image group. M is a non-negative integer, and M is less than or equal to N, that is, M can be equal to 0 or an integer greater than 0. The motion augmentation strategies other than the M motion augmentation strategies in the above-mentioned N motion augmentation strategies are the motion augmentation strategies that are judged not to be used for augmentation processing on the ternary image group. It can be understood that when M is equal to 0, it indicates that it is judged that none of the above-mentioned N motion augmentation strategies are used to perform augmentation processing on the ternary image group.
[0152] For example, in one embodiment, the present application can perform augmentation judgment on the ternary image group through the above-mentioned augmentation strategy set in a probability-based manner. The augmentation judgment process may include: the processing device can obtain the augmentation judgment probability corresponding to each motion augmentation strategy in the augmentation strategy set, and one motion augmentation strategy can correspond to one augmentation judgment probability. The augmentation judgment probability corresponding to one motion augmentation strategy is the probability of judging whether to use the motion augmentation strategy to perform augmentation processing on the ternary image group. Among them, the augmentation judgment probability corresponding to each motion augmentation strategy can be the same or different, and the augmentation judgment probability corresponding to each motion augmentation strategy can be flexibly set according to actual application requirements. For example, the augmentation judgment probability corresponding to the above-mentioned brightness augmentation strategy can be set to 5%, the augmentation judgment probability corresponding to the above-mentioned subtitle augmentation strategy can be set to 10%, and the augmentation judgment probability corresponding to the above-mentioned text augmentation strategy can also be set to 10%.
[0153] Therefore, the processing device can perform augmentation judgment on each motion augmentation strategy respectively according to the augmentation judgment probability corresponding to each motion augmentation strategy to obtain the judgment result corresponding to each motion augmentation strategy. One motion augmentation strategy can correspond to one judgment result, and the judgment result corresponding to any motion augmentation strategy can be an adopted result or a non-adopted result. If the judgment result corresponding to a motion augmentation strategy is an adopted result, it indicates that it is judged that the motion augmentation strategy needs to be adopted to perform augmentation processing on the ternary image group; if the judgment result corresponding to a motion augmentation strategy is a non-adopted result, it indicates that it is judged that the motion augmentation strategy does not need to be adopted to perform augmentation processing on the ternary image group.
[0154] The processing may determine the augmentation indication information for the ternary image group by using the motion augmentation strategy whose corresponding judgment result is the adopted result among the N motion augmentation strategies, and the M motion augmentation strategies in the augmentation indication information may include the motion augmentation strategy whose judgment result is the adopted result.
[0155] Among them, any one of the N motion augmentation strategies can be called a target motion augmentation strategy. Since each motion augmentation strategy is augmented by the augmentation judgment probability corresponding to each motion augmentation strategy, the principle of obtaining the judgment result corresponding to each motion augmentation strategy is the same and independent.
[0156] Therefore, the following will take the process of performing augmentation judgment on the target motion augmentation strategy through the augmentation judgment probability corresponding to the target motion augmentation strategy to obtain the judgment result corresponding to the target motion augmentation strategy as an example for specific explanation, as described below.
[0157] The processing device can obtain the total probability interval, and obtain the probability sub-interval corresponding to the augmented judgment probability corresponding to the target motion augmentation strategy within the total probability interval. The total probability interval can be an interval from the minimum value to the maximum value that the augmented judgment probability can take, the minimum value can be equal to 0, and the maximum value can be equal to 1. Therefore, the total probability interval can be the interval [0, 1], or the total probability interval can also be expressed as an interval of 0 to 100%. The probability sub-interval corresponding to the augmented judgment probability corresponding to the target motion augmentation strategy can be the interval from the minimum value to the augmented judgment probability corresponding to the target motion augmentation strategy within the total probability interval. For example, if the augmented judgment probability corresponding to the target motion augmentation strategy is equal to 10%, then the probability sub-interval corresponding to the augmented judgment probability corresponding to the target motion augmentation strategy can be the interval of 0 to 10% within 0 to 100%, or it can also be expressed as the interval [0, 0.1] within the total probability interval [0, 1].
[0158] The processing device can generate a random number within the total probability interval, and the random number can be referred to as a second random number. If the second random number is within the probability sub-interval corresponding to the augmented judgment probability corresponding to the target motion augmentation strategy, it can be determined that the judgment result corresponding to the target motion augmentation strategy is the above-mentioned adopted result. Conversely, if the second random number is not within (i.e., not within) the probability sub-interval corresponding to the augmented judgment probability corresponding to the target motion augmentation strategy, but is within the interval other than the probability sub-interval in the total probability interval, it can be determined that the judgment result corresponding to the target motion augmentation strategy is the above-mentioned non-adopted result. Among them, since the probability that the random number generated within the total probability interval is equal to any value within the total probability interval is the same, the present application can reflect the probability of augmented judgment of the motion augmentation strategy by generating random numbers.
[0159] Alternatively, the probability sub-interval corresponding to the augmented judgment probability corresponding to the target motion augmentation strategy may not be the interval from the minimum value that the above-mentioned augmented judgment probability can take to the augmented judgment probability corresponding to the target motion augmentation strategy, but may be any continuous sub-interval within the total probability interval and the interval length (the interval length may be equal to the value of the maximum value within the interval minus the minimum value) equal to the augmented judgment probability corresponding to the target motion augmentation strategy.
[0160] The processing device can independently perform augmentation judgment on each ternary image group according to the principles described above to obtain augmentation indication information for each ternary image group. The augmentation indication information of different ternary image groups can be the same or different, that is, the M motion augmentation strategies used for augmentation processing on each ternary image group can be the same or different, and M corresponding to different ternary image groups (that is, the number of motion augmentation strategies included in the augmentation indication information of different ternary image groups) can also be equal or different.
[0161] Step S103, obtaining an augmented image group based on M motion augmentation strategies and the ternary image group, the method of obtaining the augmented image group includes: if M is greater than 0, then according to the instructions of the M motion augmentation strategies, the ternary image group is augmented from the M motion feature dimensions of the ternary image group to obtain the augmented image group; if M is equal to 0, the ternary image group is used as the augmented image group.
[0162] Specifically, the processing device can obtain the augmented image group corresponding to the ternary image group through the M motion augmentation strategies included in the ternary image group and its augmentation indication information. Since there can be many ternary image groups, the processing device can independently obtain the augmented image group corresponding to each ternary image group through the respective augmentation indication information of each ternary image group. The following is a specific description using the process of obtaining the augmented image group corresponding to a single ternary image group as an example.
[0163] The N motion augmentation strategies in the augmentation strategy set may be serial or parallel. If the N motion augmentation strategies are serial, the M motion augmentation strategies included in the augmentation indication information may also be serial; and if the N motion augmentation strategies are parallel, the M motion augmentation strategies may also be parallel.
[0164] If the M motion augmentation strategies are serial, then the M motion augmentation strategies can have a serial order, the latter motion augmentation strategy can be superimposed on the basis of the former motion augmentation strategy, and the serial order between the M motion augmentation strategies (that is, the order of precedence between the M motion augmentation strategies) can be the same as the serial order of the M motion augmentation strategies in the above-mentioned N motion augmentation strategies (that is, the order of precedence of the M motion augmentation strategies in the N motion augmentation strategies). For example, the N motion augmentation strategies may include a brightness augmentation strategy, a subtitle augmentation strategy, and a text augmentation strategy in series, and the M motion augmentation strategies include the brightness augmentation strategy and the text augmentation strategy, then the M motion augmentation strategies may include a brightness augmentation strategy and a text augmentation strategy in series; for another example, if the M motion augmentation strategies include the subtitle augmentation strategy and the text augmentation strategy, then the M motion augmentation strategies may include a subtitle augmentation strategy and a text augmentation strategy in series; for another example, if the M motion augmentation strategies include the brightness augmentation strategy, the subtitle augmentation strategy and the text augmentation strategy, then the M motion augmentation strategies may include a brightness augmentation strategy, a subtitle augmentation strategy, and a text augmentation strategy in series.
[0165] In this case, an augmented image group corresponding to the ternary image group is obtained through the M motion augmentation strategies. For example, the processing device can perform augmentation processing on the ternary image group in sequence from the M motion feature dimensions indicated by the M motion augmentation strategies according to the serial order between the M motion augmentation strategies, so as to obtain an augmented image group corresponding to the ternary image group.
[0166] For example, the M motion augmentation strategies include the above-mentioned brightness augmentation strategy, subtitle augmentation strategy and text augmentation strategy in series, and the processing device can adopt the brightness augmentation strategy to perform augmentation processing on the ternary image group to obtain a ternary image group after brightness augmentation; the processing device can also adopt the subtitle augmentation strategy to perform augmentation processing on the ternary image group after brightness augmentation to obtain a ternary image group after subtitle augmentation; and the processing device can adopt the text augmentation strategy to perform augmentation processing on the ternary image group after subtitle augmentation to obtain an augmented image group that finally corresponds to the ternary image group.
[0167] If the M motion augmentation strategies are parallel, the processing device can obtain M augmented image groups corresponding to the ternary image group through the M motion augmentation strategies, and one of the M motion augmentation strategies is used to obtain an augmented image group of the ternary image group, that is, in this case, the ternary image group can be augmented by the M motion augmentation strategies to obtain M augmented image groups corresponding to the ternary image group (when M is not equal to 0). If the processing device can perform augmentation processing on the ternary image group respectively according to the motion feature dimensions indicated by the M motion augmentation strategies, the M augmented image groups corresponding to the M motion augmentation strategies can be obtained, and the M augmented image groups are the M augmented image groups corresponding to the ternary image group.
[0168] For example, the M motion augmentation strategies include the above-mentioned brightness augmentation strategy and subtitle augmentation strategy in parallel. The processing device can adopt the brightness augmentation strategy to perform augmentation processing on the ternary image group, and obtain an augmented image group corresponding to the ternary image group. The processing device can adopt the subtitle augmentation strategy to perform augmentation processing on the ternary image group, and obtain an augmented image group corresponding to the ternary image group, thereby obtaining two augmented image groups corresponding to the ternary image group.
[0169] Through the above process, it can be known that if M is greater than 0, the ternary image group can be augmented from the M motion feature dimensions of the ternary image group according to the instructions of the M motion augmentation strategies to obtain an augmented image group corresponding to the ternary image group.
[0170] If M is equal to 0, it indicates that there is no need to adopt any motion augmentation strategy to augment the ternary image group. Therefore, the ternary image group can be directly used as the augmented image group. That is, in this case, the augmented image group can be the original ternary image group.
[0171] The specific implementation process of augmenting the ternary image group using the above brightness augmentation strategy, subtitle augmentation strategy and text augmentation strategy can be found in the following Figure 5 , Figure 8 as well as Fig.13 Related description in .
[0172] The processing device may perform augmentation processing on each ternary image group according to the principle described above, so as to obtain an augmented image group corresponding to each ternary image group.
[0173] Step S104: The frame interpolation model is trained using the augmented image group to obtain a trained frame interpolation model. The trained frame interpolation model is used to perform frame interpolation processing between video frames of the video.
[0174] Specifically, the processing device can use the augmented image group obtained above to train the interpolation model to obtain a trained interpolation model, and the trained interpolation model is a trained model that can be used for interpolation processing between consecutive frames of the video. In addition, in the process of training the interpolation model, the principle of processing each augmented image group can be the same.
[0175] In one embodiment, a processing device uses an augmented image group to train an interpolation model to obtain a trained interpolation model, which may include: the augmented image group may include a first augmented image, a second augmented image, and a third augmented image in sequence, the first augmented image may be obtained by augmenting the first image in the ternary image group, the second augmented image may be obtained by augmenting the second image in the ternary image group, and the third augmented image may be obtained by augmenting the third image in the ternary image group, that is, the first image corresponds to the first augmented image, the second image corresponds to the second augmented image, and the third image corresponds to the third augmented image.
[0176] The processing device may call the interpolation model to generate a newly added image between the first augmented image and the third augmented image in the augmented image group, wherein the newly added image is an image predicted and generated by the interpolation model for interpolation between the first augmented image and the third augmented image. The processing device may obtain the generation loss of the interpolation model for the newly added image through the newly added image and the second augmented image in the augmented image group, and the generation loss may be used to reflect the generation deviation of the interpolation model for the newly added image.
[0177] For example, the generation loss may be the MAE (mean absolute error) or MSE (mean square error) between the newly added image and the second augmented image, or may be any other type of loss that can be used to reflect the difference between images. The generation loss may be used to reflect the image difference between the newly added image and the second augmented image. For example, the larger the generation loss, the greater the image difference between the newly added image and the second augmented image. Conversely, the smaller the generation loss, the smaller the image difference between the newly added image and the second augmented image.
[0178] The processing device can correct the model parameters of the interpolation model through the generation loss. When the correction of the model parameters of the interpolation model is completed, the interpolation model with the corrected model parameters can be used as the above-mentioned trained interpolation model. For example, the processing device can perform multiple rounds of iterative training on the interpolation model according to the principle described above (i.e., perform multiple rounds of iterative correction on the model parameters of the interpolation model). When the number of iterative trainings on the interpolation model is equal to the set number threshold, or the model parameters of the interpolation model are corrected to a convergence state, it can be considered that the correction of the model parameters of the interpolation model is completed. Among them, the goal of correcting the model parameters of the interpolation model through the generation loss can be to correct the model parameters of the interpolation model so that the generation loss can approach a minimum value (such as approaching 0).
[0179] See also Figure 4 , Figure 4 is a schematic diagram of another scene of training an interpolation model provided in an embodiment of the present application. Figure 4 As shown, the sample set may include multiple ternary image groups, specifically ternary image groups 1 to ternary image groups K, and the second moments (i.e., moments t) corresponding to the second images in different ternary image groups may be different. For example, the second moment corresponding to the second image in ternary image group 1 may be moment t1, and the second moment corresponding to the second image in ternary image group 2 may be moment t2.
[0180] The processing device may perform a round of training on the interpolation model through a batch of ternary image groups, and the batch of ternary image groups may include some ternary image groups in the sample set. The processing device may divide the ternary image groups in the sample set into multiple batches of ternary image groups to perform multiple rounds of iterative training on the interpolation model. There may be overlapping ternary image groups in different batches of ternary image groups, or there may be no overlapping ternary image groups in different batches of ternary image groups, which may be determined according to the actual application scenario. Here, the various motion augmentation strategies in the above-mentioned augmentation strategy set may be serial, and the augmentation strategy set may include a brightness augmentation strategy (corresponding to the brightness gradient augmentation in the figure), a subtitle augmentation strategy (corresponding to the subtitle change augmentation in the figure), and a text augmentation strategy (corresponding to the text motion augmentation in the figure, also referred to as text motion augmentation).
[0181] During a round of training of the interpolation model, for each ternary image group in the current batch, the processing device can judge whether to perform brightness gradient augmentation according to the probability (such as augmentation judgment probability) according to the above principle. After obtaining the judgment result corresponding to the brightness augmentation strategy, it can continue to judge whether to perform subtitle change augmentation according to the probability. After obtaining the judgment result corresponding to the subtitle augmentation strategy, it can continue to judge whether to perform text motion augmentation according to the probability, and obtain the judgment result corresponding to the text augmentation strategy. Among them, the judgment result corresponding to any motion augmentation strategy is the above-mentioned adoption result (that is, it is judged that the corresponding motion augmentation strategy should be adopted) or the non-adoption result (that is, it is judged that the corresponding motion augmentation strategy should not be adopted). Through the judgment results of various motion augmentation strategies, the processing device can obtain the respective augmentation indication information for each ternary image group in the current batch.
[0182] The processing device can perform augmentation processing on each ternary image group in the current batch according to the augmentation indication information of each ternary image group, so as to obtain the augmented image group corresponding to each ternary image group. The processing device can train the interpolation model through the augmented image groups corresponding to each ternary image group in the current batch, so as to obtain the trained interpolation model. For example, for an augmented image group, the processing device can call the interpolation model to generate and output the generated t time frame (such as the above-mentioned newly added image) through the moment t, the 0 time frame (i.e., the first image after augmentation) and the 1 time frame (i.e., the third image after augmentation) in the augmented image group. Therefore, through the generated t time frame and the t time frame in the augmented image group, the generation loss of the interpolation model for the generated t time frame can be calculated, and the generation loss can be the generation loss corresponding to the augmented image group (i.e., the loss function, such as MAE or MSE).
[0183] The processing device can add up (i.e., sum) the generation losses corresponding to the augmented image groups of each of the ternary image groups in the current batch, and can obtain the generation loss (which can be called the total generation loss) generated by the augmented image groups of each of the ternary image groups in the current batch. The processing device can perform backpropagation through the total generation loss to correct the model parameters of the interpolation model to obtain the trained interpolation model. The goal of correcting the model parameters of the interpolation model through the total generation loss is also to correct the model parameters of the interpolation model so that the total generation loss tends to be minimized (e.g., tends to 0).
[0184] After obtaining the trained interpolation model, the trained interpolation model can be used to perform interpolation processing between video frames of a video (such as between consecutive video frames), as described below.
[0185] For example, the processing device may obtain a first video frame and a second video frame that are continuous in a video, that is, the first video frame and the second video frame may be any two continuous video frames that need to be interpolated in the video. The video may be any video that needs to be interpolated, such as a short video, a video of a film or TV series, or a news video, etc.
[0186] The processing device may call the trained interpolation model to generate a newly added video frame between the first video frame and the second video frame through the first video frame and the second video frame, wherein the principle of the trained interpolation model generating the newly added video frame through the first video frame and the second video frame may be the same as the principle of the interpolation model generating the newly added image through the first augmented image and the second augmented image. The processing device may insert the newly added video frame between the first video frame and the second video frame of the video, i.e., implement interpolation processing of the video.
[0187] The method proposed in the present application can perform augmentation processing on the ternary image group used for training the interpolation model through N motion feature dimensions indicated by N motion augmentation strategies, thereby taking into account the change mode of each image in the ternary image group in multiple motion feature dimensions, and obtaining an augmented image group with rich change modes between images. Therefore, by training the interpolation model through the augmented image group obtained by augmentation, the robustness and accuracy of the trained interpolation model can be improved, and thus, the trained interpolation model can also achieve interpolation processing with excellent effect between video frames of the video.
[0188] See also Figure 5 , Figure 5 1 is a flow chart of an embodiment of the present application providing a method for augmenting a ternary image group using a brightness augmentation strategy. The embodiment of the present application may describe a process of directly augmenting a ternary image group using the brightness augmentation strategy to obtain an augmented image group when the M motion augmentation strategies determined above are parallel. If the M motion augmentation strategies are serial, the implementation process of augmenting an image group using the brightness augmentation strategy is similar. Figure 5 As shown, the process may include:
[0189] Step S201 : obtaining a preset brightness variation range, and determining a first brightness variation amplitude for the ternary image group within the brightness variation range.
[0190] Specifically, the M motion augmentation strategies obtained by the above judgment may include a brightness augmentation strategy. The embodiment of the present application specifically describes the principle of using the brightness augmentation strategy to directly perform augmentation processing on the original ternary image group. For example, when the M motion augmentation strategies are serial, and the brightness augmentation strategy is the first motion augmentation strategy in the series among the M motion augmentation strategies (i.e., the motion augmentation strategy in the front of the series), the brightness augmentation strategy can be used to directly perform augmentation processing on the ternary image group; or, when the M motion augmentation strategies are parallel, the brightness augmentation strategy can also be used to directly perform augmentation processing on the ternary image group.
[0191] However, if the M motion augmentation strategies are serial, and the brightness augmentation strategy is not the first serial motion augmentation strategy among the M motion augmentation strategies, then the brightness augmentation strategy is not used to directly augment the original ternary image group. Instead, the brightness augmentation strategy is used to augment the ternary image group augmented by the previous serial motion augmentation strategy.
[0192] For example, if the M motion augmentation strategies include a serial subtitle augmentation strategy and a brightness augmentation strategy, the processing device can use the subtitle augmentation strategy to perform augmentation processing on the original ternary image group to obtain a ternary image group after subtitle augmentation; in this case, the processing device uses the brightness augmentation strategy to perform augmentation processing on the ternary image group after subtitle augmentation to obtain a ternary image group after brightness augmentation. In this case, the ternary image group after brightness augmentation is the augmented image group corresponding to the ternary image group, and the ternary image group after subtitle augmentation is the ternary image group obtained by augmentation processing by the previous serial motion augmentation strategy (i.e., the subtitle augmentation strategy) of the brightness augmentation strategy.
[0193] The principle of using the above brightness augmentation strategy to augment the ternary image group augmented by the previous serial motion augmentation strategy is the same as the principle of using the brightness augmentation strategy to directly augment the original ternary image group as described in the following process of the embodiment of the present application. Therefore, the embodiment of the present application specifically describes the process of directly augmenting the original ternary image group by using the brightness augmentation strategy, so as to exemplify the specific implementation method of the brightness augmentation strategy, as described in the following content.
[0194] The processing device can obtain a preset brightness variation range, which is a range used to gradually adjust the brightness between images in the ternary image group under the brightness augmentation strategy. For example, the brightness variation range can be set to [0.2, 1.8], or the brightness variation range can also be set to other ranges according to actual application requirements.
[0195] The processing device may determine a first brightness change amplitude for the ternary image group within the brightness change range. The first brightness change amplitude may be an amplitude for changing and adjusting the brightness of the third image at the third moment in the ternary image group.
[0196] In one embodiment, the processing device may randomly determine the first brightness variation range within the brightness variation range. For example, the processing device may generate a random number within the brightness variation range, and the random number may be referred to as a first random number. The processing device may use the generated first random number as the first brightness variation range for the ternary image group, and the first brightness variation range may be a numerical value.
[0197] For example, the processing device may use a random algorithm (which may be any algorithm for generating random numbers) to generate a first random number within the brightness variation range, wherein when the random algorithm is used to generate random numbers within the brightness variation range, the probability that the generated random number is any value within the brightness variation range is the same. Therefore, in this way, it is possible to achieve random and equal probability determination of the first brightness variation range within the brightness variation range, and the principle and effect of generating random numbers (such as the second random number within the above-mentioned total probability interval) in other steps of the present application may also be similar to the principle and effect of generating the first random number here.
[0198] Step S202 , performing augmentation processing on the ternary image group from a brightness motion feature dimension based on the first brightness change amplitude to obtain an augmented image group.
[0199] Specifically, the processing device may use the first brightness variation amplitude determined above to adjust the brightness of the third image in the ternary image group. The third image after the brightness adjustment here may be referred to as the third image after brightness enhancement.
[0200] For example, the processing device can use the first brightness change amplitude to perform weighted processing on the color pixel value of each pixel in the third image to obtain the third image after brightness enhancement. For example, the processing device can multiply the color pixel value of each pixel in the third image by the first brightness change amplitude to obtain the third image after brightness enhancement.
[0201] Among them, the color pixel value of the pixel point can be an RGB pixel value (R represents the red color channel, G represents the green color channel, and B represents the blue color channel), and the RGB pixel value can include the pixel value on the R color channel, the pixel value on the G color channel, and the pixel value on the B color channel. Therefore, multiplying the color pixel value of the pixel point by the first brightness change amplitude can refer to multiplying the pixel value on the R color channel, the pixel value on the G color channel, and the pixel value on the B color channel of the pixel point by the first brightness change amplitude respectively. It should be noted that the brightness of the pixel point can be adjusted by adjusting the color pixel value of the pixel point.
[0202] Through the above process, the first brightness change amplitude can be understood as a weighted value for weighting the brightness of the third image. When the first brightness change amplitude is less than 1, the brightness of the third image becomes darker; when the first brightness change amplitude is equal to 1, the brightness of the third image remains unchanged; when the first brightness change amplitude is greater than 1, the brightness of the third image becomes brighter.
[0203] The present application can implement the augmentation processing of the brightness gradient between each image in the ternary image group through the brightness augmentation strategy. Therefore, the processing device can calculate the brightness change amplitude for the second image in the ternary image group through the second moment (i.e., moment t) and the first brightness change amplitude, and the brightness change amplitude for the second image can be called the second brightness change amplitude, and the second brightness change amplitude can be the amplitude used to adjust the brightness of the second image.
[0204] For example, the first brightness variation amplitude can be recorded as α, and the processing device can calculate the value of 1+t×(α-1) as the second brightness variation amplitude for the second image. The processing device can use the second brightness variation amplitude to adjust the brightness of the second image, and the second image after brightness adjustment here can be referred to as the second image after brightness enhancement. Similarly, the processing device can use the second brightness variation amplitude to perform weighted processing on the color pixel value of each pixel point in the second image to obtain the second image after brightness enhancement. For example, the processing device can multiply the color pixel value of each pixel point in the second image by the second brightness variation amplitude to obtain the second image after brightness enhancement.
[0205] Furthermore, the processing device may keep the brightness of the first image in the ternary image group unchanged, and directly use the first image as the first image after brightness enhancement.
[0206] Through the above formula 1+t×(α-1), the amplitude of brightness adjustment for the second image is smaller than the amplitude of brightness adjustment for the third image, and on the basis of brightness adjustment for the second image and the third image, the brightness of the first image can be kept unchanged, thereby achieving a gradual brightness effect among the first image after brightness enhancement, the second image after brightness enhancement, and the third image after brightness enhancement.
[0207] In the case where the M motion augmentation strategies are parallel, the processing device can directly use the brightness-augmented first image, the brightness-augmented second image, and the brightness-augmented third image as the augmented image group corresponding to the ternary image group, that is, the brightness-augmented first image, the brightness-augmented second image, and the brightness-augmented third image can constitute the augmented image group. The brightness-augmented first image is the augmented first image, the brightness-augmented second image is the augmented second image, and the brightness-augmented third image is the augmented third image.
[0208] See also Figure 6 and Figure 7 , Figure 6 This is a schematic diagram of the effect of augmenting a ternary image group using a brightness augmentation strategy provided by an embodiment of the present application. Figure 1 , Figure 7 This is a schematic diagram of the effect of augmenting a ternary image group using a brightness augmentation strategy provided by an embodiment of the present application. Figure 2 .like Figure 6 As shown, in the augmented image group obtained by augmenting the ternary image group, the image brightness of the first image after brightness augmentation, the second image after brightness augmentation, and the third image after brightness augmentation gradually darkens. In this case, the first brightness variation amplitude α may be less than 1. Figure 7 As shown, in the augmented image group obtained by augmenting the ternary image group, the image brightness of the first image after brightness enhancement, the second image after brightness enhancement and the third image after brightness enhancement gradually brightens. In this case, the above-mentioned first brightness change amplitude α can be greater than 1.
[0209] By adopting the above process of the present application, the brightness augmentation strategy can be used to implement the augmentation processing of the brightness gradient between each image in the ternary image group. By training the interpolation model through the ternary image group after the brightness gradient augmentation processing, the interpolation model can be made more adaptable to the interpolation between consecutive video frames with sudden changes in brightness, and the interpolation effect between consecutive video frames with sudden changes in brightness is improved, so that the brightness of the video frames inserted by the present application will not be unevenly distributed, the texture will not be unnatural, and there will be no strange feeling, but will present a normal gradient effect.
[0210] See also Figure 8 , Figure 8 1 is a flow chart of augmenting a ternary image group using a subtitle augmentation strategy provided by an embodiment of the present application. In the embodiment of the present application, it may be described that when the M motion augmentation strategies obtained by the above judgment are parallel, the process of directly augmenting the ternary image group using the subtitle augmentation strategy to obtain an augmented image group is described. If the M motion augmentation strategies are serial, the implementation process of augmenting the image group using the subtitle augmentation strategy is the same. Figure 8 As shown, the process may include:
[0211] Step S301 : selecting a target subtitle changing method for a ternary image group from a plurality of preset subtitle changing methods.
[0212] Specifically, the M motion augmentation strategies obtained by the above judgment may include a subtitle augmentation strategy. The embodiment of the present application specifically describes the principle of using the subtitle augmentation strategy to directly perform augmentation processing on the original ternary image group. For example, when the M motion augmentation strategies are serial, and the subtitle augmentation strategy is the first motion augmentation strategy in the series among the M motion augmentation strategies (i.e., the motion augmentation strategy in the front of the series), the subtitle augmentation strategy can be used to directly perform augmentation processing on the ternary image group; or, when the M motion augmentation strategies are parallel, the subtitle augmentation strategy can also be used to directly perform augmentation processing on the ternary image group.
[0213] However, if the M motion augmentation strategies are serial, and the subtitle augmentation strategy is not the first serial motion augmentation strategy among the M motion augmentation strategies, then the subtitle augmentation strategy is not used to directly augment the original ternary image group. Instead, the subtitle augmentation strategy is used to augment the ternary image group augmented by the previous serial motion augmentation strategy.
[0214] For example, if the M motion augmentation strategies include a brightness augmentation strategy and a subtitle augmentation strategy in series, the processing device can use the brightness augmentation strategy to perform augmentation processing on the original ternary image group to obtain a ternary image group after brightness augmentation; in this case, the processing device uses the subtitle augmentation strategy to perform augmentation processing on the ternary image group after brightness augmentation to obtain a ternary image group after subtitle augmentation. In this case, the ternary image group after subtitle augmentation is the augmented image group corresponding to the ternary image group, and the ternary image group after brightness augmentation is the ternary image group obtained by augmentation processing by the previous serial motion augmentation strategy (i.e., brightness augmentation strategy) of the subtitle augmentation strategy.
[0215] The principle of using the above-mentioned subtitle augmentation strategy to augment the ternary image group augmented by the previous serial motion augmentation strategy is the same as the principle of using the subtitle augmentation strategy to directly augment the original ternary image group as described in the following process of the embodiment of the present application. Therefore, the embodiment of the present application specifically describes the process of directly augmenting the original ternary image group by using the subtitle augmentation strategy, so as to exemplify the specific implementation method of the subtitle augmentation strategy, as described in the following content.
[0216] The subtitle augmentation strategy of the present application may preset a plurality of subtitle change modes for augmenting the subtitle changes of the ternary image group. Such a plurality of subtitle change modes include, but are not limited to: a subtitle unchanged mode, a subtitle disappearing mode, a subtitle appearing mode, and a subtitle switching mode. The subtitle unchanged mode may refer to a mode in which the subtitles on each image of the ternary image group are the same and do not change; the subtitle disappearing mode may refer to a mode in which the subtitles on each image of the ternary image group change from presence to absence; the subtitle appearing mode may refer to a mode in which the subtitles on each image of the ternary image group change from absence to presence; the subtitle switching mode may refer to a mode in which the subtitles on each image of the ternary image group switch and change the subtitle content.
[0217] The processing device may select a target subtitle changing mode for the ternary image group from the preset multiple subtitle changing modes. Since the multiple subtitle changing modes are mutually exclusive and will not occur at the same time, the target subtitle changing mode may be any one of the multiple subtitle changing modes. For example, the processing device may randomly select a subtitle changing mode from the multiple subtitle changing modes as the target subtitle changing mode.
[0218] Alternatively, the processing device may also set a selection probability for each subtitle change mode, and the sum of the selection probabilities corresponding to the multiple subtitle change modes may be equal to 1, so that the processing device may select a subtitle change mode from the multiple subtitle change modes as the target subtitle change mode according to the selection probability of each subtitle change mode. For example, the selection probability corresponding to the subtitle unchanged mode may be set to 40%, and the selection probabilities corresponding to the subtitle disappearance mode, the subtitle appearance mode, and the subtitle switching mode may all be set to 20%.
[0219] Correspondingly, the processing device can allocate the probability sub-intervals corresponding to each subtitle changing mode in the total probability interval of 0 to 100%, such as the probability sub-interval corresponding to the subtitle unchanged mode can be 0 to 40%, the probability sub-interval corresponding to the subtitle disappearing mode can be 40% to 60%, the probability sub-interval corresponding to the subtitle appearing mode can be 60% to 80%, and the probability sub-interval corresponding to the subtitle switching mode can be 80% to 100%. The processing device can generate a random number in the total probability interval of 0 to 100%, which can be called a third random number. Which subtitle changing mode the third random number is in can be used as the target subtitle changing mode. For example, if the third random number is equal to 50%, the subtitle disappearing mode corresponding to the probability sub-interval of 40% to 60% where the 50% is located can be used as the target subtitle changing mode.
[0220] Step S302: perform augmentation processing on the first image, the second image and the third image from the subtitle motion feature dimension according to the target subtitle change mode to obtain an augmented image group.
[0221] Specifically, the processing device can perform augmentation processing on the first image, the second image and the third image in the ternary image group from the subtitle motion feature dimension according to the target subtitle change mode to obtain an augmented image group corresponding to the ternary image group, as described below.
[0222] If the target subtitle change mode is the above-mentioned subtitle unchanged mode, the processing device can generate a subtitle, and the generated subtitle can be called a first subtitle. The first subtitle can be a subtitle randomly generated by the processing device. For example, the number of subtitle characters, subtitle content (such as the specific characters contained in the subtitle), subtitle font, subtitle size, subtitle direction (such as vertical or horizontal), subtitle brightness, subtitle stroke width, and / or subtitle stroke brightness of the first subtitle can all be random.
[0223] Among them, optional ranges can also be set for the above-mentioned subtitle size (also known as subtitle size), subtitle direction, subtitle brightness, subtitle stroke width, subtitle stroke brightness and other subtitle attributes according to actual needs. For example, the optional range for subtitle size can be [20, 120], and the unit can be pixel (pixel). The optional range for subtitle direction can include from top to bottom (i.e. vertical) and from left to right (i.e. horizontal). The optional range for subtitle brightness can be [220, 255], which belongs to the range within the standard dynamic range, and the overall range of the standard dynamic range can be [0, 255]. The optional range for subtitle stroke width can be [0, 6], and the unit can also be pixel (pixel). The optional range for subtitle stroke brightness can be [0, 32], which also belongs to the range within the standard dynamic range. The present application can flexibly set the optional ranges of various subtitle attributes according to actual conditions, and can randomly determine the various subtitle attributes of the subtitles to be generated within the optional ranges of various subtitle attributes.
[0224] Since the target subtitle change mode is a subtitle unchanged mode, the processing device can obtain an image position in the first image, the second image, and the third image, and the image position can be referred to as the first image position. For example, the first image position can be represented as (x1, y1), where x1 can be the row number where the pixel point is located, and y1 can be the column number where the pixel point is located. The first image position can be an image position randomly determined in the first image, the second image, and the third image, and the first image, the second image, and the third image all have the first image position.
[0225] The processing device can add the first subtitle to the first image, the second image, and the third image at the position of the first image, respectively, to obtain the first image after subtitle augmentation, the second image after subtitle augmentation, and the third image after subtitle augmentation. The first image after subtitle augmentation can be the first image after the first subtitle is added to the position of the first image, the second image after subtitle augmentation can be the second image after the first subtitle is added to the position of the first image, and the third image after subtitle augmentation can be the third image after the first subtitle is added to the position of the first image. In this way, the positions of the subtitles (i.e., the first subtitles) contained in the first image after subtitle augmentation, the second image after subtitle augmentation, and the third image after subtitle augmentation and the subtitle content are consistent and do not change, thereby achieving the effect of unchanged subtitles.
[0226] The first subtitle may occupy a rectangular area in the image, the length of the rectangular area may be the length of the first subtitle, and the width of the rectangular area may be the width of the first subtitle. When the first subtitle is added to the first image, the second image, and the third image, the position of the upper left corner of the first subtitle (such as the upper left corner of the rectangular area occupied by the first subtitle) may be at the first image position, that is, the first image position may be the position of the upper left corner of the added first subtitle, and the same applies to subsequent subtitle additions to the image.
[0227] When the above-mentioned M motion augmentation strategies are in parallel, the processing device can directly regard the first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation as an augmented image group, that is, the first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation can constitute an augmented image group.
[0228] If the target subtitle change mode is the above-mentioned subtitle disappearance mode, the processing device may also randomly generate a subtitle, and the subtitle may also be referred to as the first subtitle, and the method of generating the subtitle is the same as the method of generating the above-mentioned first subtitle. The processing device may obtain an image position in the first image and the second image, and may also refer to the image position as the first image position, and the first image position may be an image position randomly obtained in the first image and the second image.
[0229] The processing device can add the first subtitle at the position of the first image in the first image and the second image respectively to obtain the first image after subtitle augmentation and the second image after subtitle augmentation, wherein the first image after subtitle augmentation is the first image after the first subtitle is added at the position of the first image, and the second image after subtitle augmentation is the second image after the first subtitle is added at the position of the first image. The content and position of the subtitles (such as the first subtitles) contained in the first image after subtitle augmentation and the second image after subtitle augmentation are consistent and do not change.
[0230] And, the processing device can directly use the original third image as the third image after subtitle augmentation. Similarly, when the above-mentioned M motion augmentation strategies are parallel, the processing device can directly use the first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation as an augmented image group, that is, the first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation can constitute an augmented image group. In the augmented image group, subtitles exist in the first image after subtitle augmentation and the second image after subtitle augmentation, but subtitles do not exist in the third image after subtitle augmentation, thereby achieving the effect of subtitle disappearance (that is, subtitles disappear from existence).
[0231] If the target subtitle change mode is the above subtitle appearance mode, the processing device can also randomly generate a subtitle, and the subtitle can also be referred to as the first subtitle. The method of generating the subtitle is the same as the method of generating the above first subtitle. The processing device can obtain an image position in the third image, and the image position can also be referred to as the first image position. The first image position can be an image position randomly obtained in the third image.
[0232] The processing device may add the first subtitle at the position of the first image in the third image to obtain a third image with subtitle augmentation, and the third image with subtitle augmentation is the third image after the first subtitle is added at the position of the first image. Furthermore, the processing device may directly use the original first image and the second image as the first image with subtitle augmentation and the second image with subtitle augmentation, respectively.
[0233] Similarly, when the above-mentioned M motion augmentation strategies are in parallel, the processing device can directly use the first image after subtitle augmentation, the second image after subtitle augmentation, and the third image after subtitle augmentation as an augmented image group. In the augmented image group, there are no subtitles in the first image after subtitle augmentation and the second image after subtitle augmentation, but there are subtitles in the third image after subtitle augmentation, so as to achieve the effect of subtitle appearance (i.e., subtitles appear from nothing).
[0234] If the target subtitle change mode is the above subtitle switching mode, the processing device can generate two random subtitles. The two subtitles can be generated in the same way as the first subtitle. One of the two subtitles can be called the first subtitle, and the other subtitle can be called the second subtitle. The first subtitle and the second subtitle can be different subtitles.
[0235] The processing device may obtain an image position in the first image, the second image, and the third image, and the image position may also be referred to as the first image position, and the first image position may be an image position randomly obtained in the first image, the second image, and the third image. The processing device may add a first subtitle to the first image position in the first image and the second image, respectively, and may obtain a first image augmented with subtitles and a second image augmented with subtitles.
[0236] In addition, the processing device can add a second subtitle at the position of the first image in the third image, and a third image with augmented subtitles can be obtained. Similarly, when the above M motion augmentation strategies are parallel, the processing device can directly use the first image with augmented subtitles, the second image with augmented subtitles, and the third image with augmented subtitles as the augmented image group. In this augmented image group, the subtitles in the first image and the second image are the same, both being the first subtitle, while the subtitles in the first image and the second image are different from those in the third image. The subtitle in the third image is the second subtitle, thus achieving the effect of subtitle switching (i.e., switching from the first subtitle to the second subtitle).
[0237] Please refer to Figures 9 to 12 , Fig. 9 which is a schematic diagram of the effect of augmenting a ternary image group using a subtitle augmentation strategy provided by an embodiment of the present application Figure 1 , Fig.10 which is a schematic diagram of the effect of augmenting a ternary image group using a subtitle augmentation strategy provided by an embodiment of the present application Figure 2 , Fig.11 which is a schematic diagram of the effect of augmenting a ternary image group using a subtitle augmentation strategy provided by an embodiment of the present application Figure 3 , Fig.12 which is a schematic diagram of the effect of augmenting a ternary image group using a subtitle augmentation strategy provided by an embodiment of the present application Figure 4 .
[0238] As Fig. 9 shown, after augmenting the ternary image group using the subtitle unchanged method, in the first image with augmented subtitles, the second image with augmented subtitles, and the third image with augmented subtitles in the obtained augmented image group, the same subtitle "Anci" exists at the same image position. And as Fig.10 shown, after augmenting the ternary image group using the subtitle disappearance method, in the first image with augmented subtitles and the second image with augmented subtitles in the obtained augmented image group, the same subtitle "Adui" exists at the same image position, while there is no subtitle in the third image with augmented subtitles.
[0239] Again, as Fig.11 shown, after augmenting the ternary image group using the subtitle appearance method, there is no subtitle in the first image with augmented subtitles and the second image with augmented subtitles in the obtained augmented image group, while there is a subtitle "Duansong" in the third image with augmented subtitles. As Fig.12As shown, after augmenting the ternary image group by means of subtitle switching, in the first image with augmented subtitles and the second image with augmented subtitles in the obtained augmented image group, there is the same subtitle "Sifang En" at the same image position, while there is a subtitle "Bici Lvhua" in the third image with augmented subtitles, and this subtitle "Bici Lvhua" is different from the subtitle "Sifang En" in the first image with augmented subtitles and the second image with augmented subtitles.
[0240] By adopting the above process of the present application, augmentation processing of various subtitle changes among the images in the ternary image group can be achieved through a subtitle augmentation strategy. By training an interpolation model with the ternary image group after the augmentation processing of subtitle changes, the interpolation model can be made more adaptable to interpolating between consecutive video frames with subtitle changes, improving the interpolation effect between consecutive video frames with subtitle changes, so that the subtitles in the inserted video frames will not have serious afterimages and distortions and will not present an unnatural and strange feeling.
[0241] Please refer to Fig.13 , Fig.13 FIG. is a schematic flowchart of a process for augmenting a ternary image group by adopting a text augmentation strategy provided by an embodiment of the present application. In the embodiment of the present application, it may describe the process of directly augmenting the ternary image group by adopting this text augmentation strategy when it is determined that the M motion augmentation strategies are parallel, and if the M motion augmentation strategies are serial, the implementation process of augmenting the image group by adopting this text augmentation strategy is the same. As Fig.13 shown, the process may include:
[0242] Step S401, select a target text motion mode for the ternary image group from a plurality of preset text motion modes.
[0243] Specifically, the M motion augmentation strategies obtained by the above determination may include a text augmentation strategy, and the embodiment of the present application specifically describes the principle of directly augmenting the original ternary image group by adopting this text augmentation strategy. For example, when the M motion augmentation strategies are serial and this text augmentation strategy is the first motion augmentation strategy in the serial M motion augmentation strategies (i.e., the motion augmentation strategy at the front of the serial), this text augmentation strategy can be used to directly augment the ternary image group; or when the M motion augmentation strategies are parallel, this text augmentation strategy can also be used to directly augment the ternary image group.
[0244] However, if the M motion augmentation strategies are serial, and the text augmentation strategy is not the first serial motion augmentation strategy among the M motion augmentation strategies, then the text augmentation strategy is not used to directly augment the original ternary image group. Instead, the text augmentation strategy is used to augment the ternary image group augmented by the previous serial motion augmentation strategy.
[0245] For example, if the M motion augmentation strategies include a brightness augmentation strategy and a text augmentation strategy in series, the processing device can use the brightness augmentation strategy to perform augmentation processing on the original ternary image group to obtain a ternary image group after brightness augmentation; in this case, the processing device uses the text augmentation strategy to perform augmentation processing on the ternary image group after brightness augmentation to obtain a ternary image group after text augmentation. In this case, the ternary image group after text augmentation is the augmented image group corresponding to the ternary image group, and the ternary image group after brightness augmentation is the ternary image group obtained by augmentation processing by the previous serial motion augmentation strategy (i.e., brightness augmentation strategy) of the text augmentation strategy.
[0246] The principle of using the above text augmentation strategy to augment the ternary image group augmented by the previous serial motion augmentation strategy is the same as the principle of using the text augmentation strategy to directly augment the original ternary image group as described in the following process of the embodiment of the present application. Therefore, the embodiment of the present application specifically describes the process of directly augmenting the original ternary image group using the text augmentation strategy, so as to exemplify the specific implementation method of the text augmentation strategy, as described in the following content.
[0247] The text augmentation strategy of the present application may preset a variety of text movement modes for augmenting the text movement of the ternary image group. For example, the multiple text movement modes include but are not limited to: text translation mode, text scaling mode, and text rotation mode. The text translation mode may refer to a mode in which the text between each image is displayed in translation; the text scaling mode may refer to a mode in which the text between each image is displayed in scaling; and the text rotation mode may refer to a mode in which the text between each image is displayed in rotation.
[0248] The processing device may select a target text motion mode for the ternary image group from the preset multiple text motion modes. Since the multiple text motion modes are not mutually exclusive and may occur simultaneously, the target text motion mode may be at least one (i.e., one or more) of the multiple text motion modes. For example, the processing device may randomly select at least one text motion mode from the multiple text motion modes as the target text motion mode.
[0249] Alternatively, the processing device may also set a selection probability for each text motion mode, and the sum of the selection probabilities corresponding to the multiple text motion modes may not be equal to 1. The processing device may select at least one text motion mode from the multiple text motion modes as the target text motion mode according to the selection probability of each text motion mode. For example, the selection probabilities corresponding to the text translation mode, the text scaling mode, and the text rotation mode may all be set to 50%.
[0250] Correspondingly, the processing device can set a probability sub-interval corresponding to each text motion mode in three probability total intervals of 0 to 100%, and one text motion mode corresponds to one probability total interval, such as the probability sub-interval corresponding to the text translation mode can be an interval of 0 to 50% in the corresponding probability total interval, the probability sub-interval corresponding to the text scaling mode can be an interval of 0 to 50% in the corresponding probability total interval, and the probability sub-interval corresponding to the text rotation mode can be an interval of 0 to 50% in the corresponding probability total interval. The processing device can generate a random number in the probability total interval corresponding to each text motion mode. If the random number generated in the probability total interval corresponding to a text motion mode is within the probability sub-interval corresponding to the text motion mode, the text motion mode can be used as the target text motion mode, and if the random number generated in the probability total interval corresponding to a text motion mode is not within the probability sub-interval corresponding to the text motion mode, the text motion mode can be not used as the target text motion mode.
[0251] However, if the random numbers generated within the total probability interval corresponding to various text motion modes are not within the probability sub-interval corresponding to various text motion modes, that is, no text motion mode is selected through the selection probabilities corresponding to various text motion modes, then at least one text motion mode can be randomly selected from the multiple text motion modes as the target text motion mode.
[0252] Step S402 , generating a target text to be added, and performing augmentation processing on the first image, the second image and the third image based on the target text motion feature dimension according to the target text motion mode to obtain an augmented image group.
[0253] Specifically, the processing device can generate the text to be added, and the text to be added can be referred to as the target text. The target text can be a text randomly generated by the processing device, such as the number of characters, text content (such as the specific text contained), text font (such as text font), text size (such as text size, which can also be referred to as text size), text direction (such as vertical or horizontal, etc.), text brightness, text stroke width and text stroke brightness of the target text can all be random.
[0254] Alternatively, the present application may also set the respective optional ranges of various text attributes according to actual needs, and may randomly determine various text attributes within the respective optional ranges of various text attributes. For example, in the present application, the optional range of text size may be [20, 120], and the unit may be pixels; the optional range of text direction may include from top to bottom (i.e. vertical) and from left to right (i.e. horizontal); the optional range of text brightness may be [220, 255], which may be a range within the standard dynamic range, the optional range of text stroke width may be [0, 6], and the optional range of text stroke brightness may be [0, 32].
[0255] The processing device can perform augmentation processing on the first image, the second image and the third image in the ternary image group from the text motion feature dimension according to the target text motion mode through the target text generated above to obtain an augmented image group, as described below.
[0256] In the following process, the specific process of augmenting the ternary image group when the target text motion mode is any text motion mode is described respectively, so as to illustrate the implementation principle of various text motion modes. However, it should be noted that the text can be translated, scaled, and rotated at the same time, so the above-mentioned multiple text motion modes can be performed at the same time, that is, the target text motion mode can include the multiple text motion modes at the same time.
[0257] When the target text motion mode includes at least two of the multiple text motion modes, the order of augmentation processing of the ternary image group by sequentially superimposing each target text motion mode can be arbitrary, or a fixed order can be set, and the present application does not impose any restrictions on this. In this case, the text processed when the ternary image group is augmented by each target text motion mode can be the same text (such as the target text), and the latter target text motion mode can be a continuation of the augmentation processing of the ternary image group by the former target text motion mode, as described below. The following first exemplarily describes the implementation principles of each of the various text motion modes.
[0258] If the target text movement mode is text translation, the processing device can obtain the text displacement (m x , m y ), m x Indicates the displacement of the target text in the horizontal direction (can be the x-axis (horizontal and vertical) direction), m y Indicates the displacement of the target text in the vertical direction (which can be the y-axis (vertical) direction). The text displacement (m x , m y ) can be the displacement of the target text from the first image to the third image, m x and my Each can be positive or negative.
[0259] The processing device can obtain an image position in the first image, and the image position can be called the second image position. The second image position can be an image position randomly determined in the first image. The second image position can be expressed as (x2, y2), where x2 is the number of rows of pixels in the second image position in the horizontal direction (which can be the horizontal coordinate of the second image position), and y2 is the number of columns of pixels in the second image position in the vertical direction (which can be the vertical coordinate of the second image position).
[0260] The processing device may add the target text at the second image position in the first image to obtain the first image after text augmentation. The second image position may also be the position of the target text at the upper left corner in the first image after text augmentation. The target text may also occupy a rectangular area in the first image after text augmentation. The length of the rectangular area is the length of the target text, the width of the rectangular area is the width of the target text, and the upper left corner of the rectangular area may be at (i.e. located at) the second image position. The same is true for adding text to the image later.
[0261] The processing device can also calculate the third image position in the second image through the second image position, the second time t and the text displacement. The third image position can be (x2+t×m x ,y2+t×m y ), t is a positive number less than 1, the x2+t×m x That is, the horizontal coordinate of the third image position (which can be the row number of the pixel point at the third image position in the second image), the y2+t×m y That is, the ordinate of the third image position (which may be the column number of the pixel point at the third image position in the second image). The processing device may add the target text at the third image position in the second image to obtain the second image after text augmentation.
[0262] Furthermore, the processing device can also calculate the fourth image position in the third image by using the second image position and the text displacement. The fourth image position can be (x2+m x ,y2+m y ), the x2+m x That is, the horizontal coordinate of the fourth image position (which can be the row number of the pixel point at the fourth image position in the third image), the y2+m y That is, the ordinate of the fourth image position (which may be the column number of the pixel point at the fourth image position in the third image). The processing device may add the target text at the fourth image position in the third image to obtain the third image after text augmentation.
[0263] In one implementation, in order to ensure that part or all of the target text is displayed within the image range after translation, x2+t×m can be defined. x and x2+m x Both are greater than 0 and less than the width of the image, y2+t×m y and y2+m y are greater than 0 and less than the height of the image, and x2+m x Add the width of the target text, and x2+t×m x Plus the width of the target text must be greater than 0, y2+m y Add the height of the target text and y2+t×m y The height of the target text must be greater than 0. Based on the above conditions, the processing device can randomly determine the text displacement (m x , m y ).
[0264] The processing device can use the first image after text augmentation, the second image after text augmentation, and the third image after text augmentation as an augmented image group, that is, the first image after text augmentation, the second image after text augmentation, and the third image after text augmentation can constitute an augmented image group. In the augmented image group, the position of the target text in each image is a sequentially shifted effect.
[0265] If the target text movement mode is the above-mentioned text scaling mode, the processing device can obtain the fifth image position in the first image, the second image, and the third image, and the fifth image position can also be an image position randomly determined in the first image, the second image, and the third image. The processing device can directly add the target text at the fifth image position in the first image to obtain the first image after text augmentation.
[0266] The processing device may also obtain a first zoom scale for the target text, which may be a ratio or multiple of zooming in and out of the target text, and may be a zoom scale of the target text in the third image. If the first zoom scale is less than 1, the target text is reduced by using the first zoom scale, and if the first zoom scale is greater than 1, the target text is enlarged by using the first zoom scale. The first zoom scale may be denoted as s.
[0267] The processing device can calculate the second scaling scale of the target text by the first scaling scale and the second time t, and the second scaling scale can be the scaling scale of the target text in the second image. The processing device can use the second scaling scale to scale the target text to obtain the first scaling text, and the processing device can use the first scaling scale to scale the target text to obtain the second scaling text.
[0268] The processing device may add the first scaled text at the fifth image position in the second image to obtain the second image after text augmentation. The processing device may also add the second scaled text at the fifth image position in the third image to obtain the third image after text augmentation. The processing device may use the first image after text augmentation, the second image after text augmentation, and the third image after text augmentation as an augmented image group, and the text added to each image in the augmented image group is the effect of sequential scaling.
[0269] If the target text movement mode is the above-mentioned text rotation mode, the processing device can obtain the sixth image position in the first image, the second image and the third image, and the sixth image position can also be an image position randomly determined in the first image, the second image and the third image. The processing device can add the target text in a uniform direction at the sixth image position in the first image, the second image and the third image to obtain the first image after adding the text, the second image after adding the text, and the third image after adding the text. For example, the target text can be added uniformly from top to bottom (vertical direction) or from left to right (horizontal direction) at the sixth image position in the first image, the second image and the third image.
[0270] The processing device can obtain the rotation center of the target text in the first image after adding text, the second image after adding text, and the third image after adding text, and can subsequently rotate the target text in the first image after adding text, the second image after adding text, and the third image after adding text based on the rotation center. For example, in the first image after adding text, the second image after adding text, and the third image after adding text, the target text can occupy a rectangular area, the height of the rectangular area is the height of the target text, the width of the rectangular area is the width of the target text, and the position of the upper left corner of the rectangular area can be the sixth image position. Therefore, a pixel point can be arbitrarily selected in the rectangular area as the rotation center of the target text.
[0271] The processing device can obtain a first rotation angle for the target text, which can be the rotation angle of the target text in the first image. The first rotation angle can be an angle randomly determined within a first angle range, such as the first angle range can be set to [0, 360].
[0272] Therefore, the processing device can rotate the target text in the first image after adding text by the first rotation angle through the above-mentioned rotation center, and obtain the first image after text augmentation. The rotation angle of the target text in the first image after text augmentation is the first rotation angle.
[0273] The processing device can also obtain a second rotation angle for the target text, and the second rotation angle is used to determine the angle at which the target text is gradually rotated in the second image and the third image based on the first rotation angle. The second rotation angle can be an angle randomly determined within a second angle range. For example, in order to ensure the effect of rotating the text, the second angle range can be set to [0, 180] or [0, 60]. The second angle range can be smaller than the first angle range, so that the audience can perceive the direction of the text rotation, such as clockwise rotation or counterclockwise rotation.
[0274] The processing device can calculate a third rotation angle for the target text through the first rotation angle, the second rotation angle and the second time t, and the third rotation angle can be the rotation angle of the target text in the second image. The first rotation angle can be recorded as r0, and the second rotation angle can be recorded as r. For example, the third rotation angle can be equal to r0+t×r.
[0275] The processing device can rotate the target text in the second image after adding text by the third rotation angle through the above-mentioned rotation center to obtain the second image after text augmentation, and the rotation angle of the target text in the second image after text augmentation is the third rotation angle.
[0276] The processing device can also calculate a fourth rotation angle for the target text by using the first rotation angle and the second rotation angle. The fourth rotation angle can be the rotation angle of the target text in the third image. For example, the third rotation angle can be equal to r0+r. The processing device can rotate the target text in the third image after adding text by the fourth rotation angle through the rotation center to obtain the third image after text augmentation. The rotation angle of the target text in the third image after text augmentation is the fourth rotation angle.
[0277] The processing device may use the first image after text augmentation, the second image after text augmentation, and the third image after text augmentation as an augmented image group, and the target text in each image of the augmented image group is gradually rotated. When rotating the target text in the first image after adding text, the second image after adding text, and the third image after adding text, the target text may be uniformly rotated clockwise or uniformly rotated counterclockwise.
[0278] Through the above process, the implementation principle of the text translation mode, the text scaling mode and the text rotation mode is described. If the target text motion mode includes at least two (i.e. more than one) text motion modes, the at least two text motion modes may have a serial order, and the serial order may be random or fixed. The processing device may perform augmentation processing on the ternary image group in sequence using the at least two text motion modes according to the serial order between the at least two text motion modes, so as to obtain the augmented image group corresponding to the ternary image group. In this process, the latter text motion mode may be augmented on the basis of the augmentation processing of the target text by the former text motion mode, and then augmented in a superimposed manner. Therefore, it is only necessary to add the target text when the ternary image group is augmented by the first text motion mode in the series, and it is not necessary to add the target text when the augmentation processing is performed by the subsequent text motion modes in the series. When the augmentation processing is performed by the last text motion mode in the series, the augmented image group can be obtained.
[0279] For example, if the target text movement mode includes a text translation mode, a text scaling mode, and a text rotation mode in series, the processing device can perform augmentation processing on the ternary image group by the text translation mode to obtain a ternary image group after translation augmentation, and the target text can be added to each image of the ternary image group after translation augmentation according to the implementation principle of the above-mentioned text translation mode. The processing device can directly perform augmentation processing (here, scaling processing) on the target text added to each image of the ternary image group after translation augmentation by the text scaling mode to obtain a ternary image group after scaling augmentation. In this process, the principle of scaling the target text added to each image of the ternary image group after translation augmentation is the same as the principle of scaling the target text added at the fifth image position of each image in the ternary image group, except that it is no longer necessary to repeatedly add the target text here. The processing device can also directly perform augmentation processing (here, rotation processing) on the target text already in the scaled and augmented ternary image group by text rotation, and obtain the final augmented image group. In this process, the principle of rotating the target text already in the scaled and augmented ternary image group is the same as the principle of rotating the target text added at the sixth image position of each image in the ternary image group, except that there is no need to repeatedly add the target text here.
[0280] For another example, if the target text movement mode includes a text scaling mode and a text rotation mode in series, the processing device can perform augmentation processing on the ternary image group by the text scaling mode to obtain a scaled and augmented ternary image group, and each image of the scaled and augmented ternary image group can be added with the target text according to the implementation principle of the above text scaling mode. The processing device can directly perform augmentation processing (here, rotation processing) on the target text added to the scaled and augmented ternary image group by the text rotation mode to obtain the final augmented image group, and there is no need to repeatedly add the target text in this process.
[0281] For example, if the target text movement mode includes a text rotation mode and a text translation mode in series, the processing device can perform augmentation processing on the ternary image group by the text rotation mode to obtain a rotated and augmented ternary image group, and each image of the rotated and augmented ternary image group can be added with the target text according to the implementation principle of the above text rotation mode. The processing device can directly perform augmentation processing (here, translation processing) on the target text added to the rotated and augmented ternary image group by the text translation mode to obtain the final augmented image group, and there is no need to repeatedly add the target text in this process.
[0282] Alternatively, in one embodiment, if there are at least two target text motion modes, each target text motion mode can also operate on different texts. In this case, the at least two target text motion modes can be independently performed on the ternary image group, and when each target text motion mode is performed, the target text corresponding to each target text motion mode can be generated, and the target text corresponding to each target text motion mode can be added to the ternary image group according to the respective implementation principles of each text motion mode.
[0283] Among them, the target texts corresponding to different target text movement modes can all be randomly generated, the target texts corresponding to different target text movement modes can be different, and the added positions of the target texts corresponding to different target text movement modes in each image of the ternary image group (i.e., the image position where the added target text is located) can be different. In this case, it can be understood that the various target text movement modes are parallel, and the order of performing each target text movement mode can be arbitrary.
[0284] For example, if at least two target text movement modes include a text translation mode and a text scaling mode, the ternary image group can be augmented by the text translation mode to obtain a ternary image group after translation augmentation, and the target text corresponding to the text translation mode can be added to the ternary image group according to the implementation principle of the above-mentioned text translation mode. The processing device can also perform augmentation processing on the ternary image group after translation augmentation by the text scaling mode to obtain a final augmented image group, and the augmented image group is not only added with the target text corresponding to the text translation mode, but also with the target text corresponding to the text scaling mode. That is, in the process of augmenting the ternary image group after translation augmentation by the text scaling mode, the target text corresponding to the text scaling mode will also be added to the ternary image group after translation augmentation.
[0285] See also Figure 14 to Figure 16 , Fig.14 This is a schematic diagram of the effect of augmenting a ternary image group using a text augmentation strategy provided in an embodiment of the present application. Figure 1 , Fig.15 This is a schematic diagram of the effect of augmenting a ternary image group using a text augmentation strategy provided in an embodiment of the present application. Figure 2 , Fig.16 This is a schematic diagram of the effect of augmenting a ternary image group using a text augmentation strategy provided in an embodiment of the present application. Figure 3 .
[0286] like Fig.14As shown, in the augmented image group obtained by augmenting the ternary image group using the text translation method, the first image after text augmentation, the second image after text augmentation, and the third image after text augmentation contain the same text "stone cave", but the position of the text "stone cave" in the first image after text augmentation, the second image after text augmentation, and the third image after text augmentation is gradually translated diagonally upward and to the right within the image.
[0287] For another example Fig.15 As shown, in the augmented image group obtained by augmenting the ternary image group using the text scaling method, the first image after text augmentation, the second image after text augmentation, and the third image after text augmentation contain the same text "Aohuafang", but the text "Aohuafang" is gradually enlarged in the first image after text augmentation, the second image after text augmentation, and the third image after text augmentation.
[0288] And, as Fig.16 shown, in the augmented image group obtained by augmenting the ternary image group using the text rotation method, the first image after text augmentation, the second image after text augmentation, and the third image after text augmentation contain the same text "Kangnan Rhododendron A", but the text "Kangnan Rhododendron A" is gradually rotated counterclockwise by different angles in the first image after text augmentation, the second image after text augmentation, and the third image after text augmentation.
[0289] The above process of the present application can achieve the augmentation processing of various text motions between the images in the ternary image group through the text augmentation strategy. By training the interpolation model with the ternary image group after the augmentation processing of the text motion, the interpolation model can be made more adaptable to the interpolation between consecutive video frames with text motion, improving the effect of interpolating between consecutive video frames with text motion, so that the text in the inserted video frames will not have overlapping, ghosting, or blurring.
[0290] After augmenting the ternary image group using the above series of methods provided by the present application, the interpolation effect in various common scenarios (such as scenes with gradual brightness change, scenes with subtitle change, and scenes with text motion) can be greatly improved, making the trained interpolation model have higher robustness and practical value.
[0291] Please refer to Fig.17 , Fig.17 which is a schematic structural diagram of an image processing device provided by an embodiment of the present application. As Fig.17 shown, the image processing device 170 may include: an acquisition module 1701, a judgment module 1702, an augmentation module 1703, and a training module 1704.
[0292] An acquisition module 1701 is used to acquire a ternary image group, where the ternary image group includes three video frames sampled from a sample video;
[0293] A judgment module 1702 is used to perform augmentation judgment on the ternary image group based on the augmentation strategy set to obtain augmentation indication information for the ternary image group, wherein the augmentation strategy set includes N motion augmentation strategies, each motion augmentation strategy is used to indicate augmentation processing from a motion feature dimension of the ternary image group, and the augmentation indication information includes M motion augmentation strategies obtained by judgment from the N motion augmentation strategies, where N is a positive integer, M is a non-negative integer, and M is less than or equal to N;
[0294] The augmentation module 1703 is used to obtain an augmented image group based on M motion augmentation strategies and the ternary image group, and the method of obtaining the augmented image group includes: if M is greater than 0, then according to the instructions of the M motion augmentation strategies, the ternary image group is augmented from the M motion feature dimensions of the ternary image group to obtain the augmented image group; if M is equal to 0, then the ternary image group is used as the augmented image group;
[0295] The training module 1704 is used to train the interpolation model using the augmented image group to obtain a trained interpolation model, and the trained interpolation model is used to perform interpolation processing between video frames of the video.
[0296] In one implementation, the M motion augmentation strategies include a brightness augmentation strategy, and the motion feature dimension indicated by the brightness augmentation strategy is a brightness motion feature dimension;
[0297] The augmentation module 1703 acquires the augmented image group based on the M motion augmentation strategies and the ternary image group, including:
[0298] Acquire a preset brightness variation range, and determine a first brightness variation amplitude for the ternary image group within the brightness variation range;
[0299] The ternary image group is augmented from a brightness motion feature dimension based on the first brightness change amplitude to obtain an augmented image group.
[0300] In one implementation, the ternary image group includes, in sequence, a first image at a first moment, a second image at a second moment, and a third image at a third moment, where the second moment is between the first moment and the third moment; the augmentation module 1703 performs augmentation processing on the ternary image group from a brightness feature dimension based on the first brightness change amplitude to obtain an augmented image group, including:
[0301] Using the first image as the first image after brightness enhancement;
[0302] Calculating a second brightness change amplitude for the second image based on the second moment and the first brightness change amplitude, and performing weighted processing on the color pixel value of each pixel point in the second image using the second brightness change amplitude to obtain a second image after brightness enhancement;
[0303] Using the first brightness variation amplitude to perform weighted processing on the color pixel value of each pixel point in the third image, so as to obtain a third image with brightness enhancement;
[0304] The first image after brightness enhancement, the second image after brightness enhancement, and the third image after brightness enhancement are taken as an augmented image group.
[0305] In one implementation, the augmentation module 1703 determines the first brightness change amplitude for the ternary image group within the brightness change range, including:
[0306] Generate a first random number within the brightness variation range;
[0307] The generated first random number is used as the first brightness change amplitude for the ternary image group.
[0308] In one implementation, the ternary image group includes a first image, a second image, and a third image in sequence, the M motion augmentation strategies include a subtitle augmentation strategy, and the motion feature dimension indicated by the subtitle augmentation strategy is a subtitle motion feature dimension;
[0309] The augmentation module 1703 acquires the augmented image group based on the M motion augmentation strategies and the ternary image group, including:
[0310] Selecting a target subtitle changing method for the ternary image group from a plurality of preset subtitle changing methods;
[0311] According to the target subtitle change mode, the first image, the second image and the third image are augmented from the subtitle motion feature dimension to obtain an augmented image group.
[0312] In one implementation, if the target subtitle change mode is a subtitle unchanged mode, the augmentation module 1703 augments the first image, the second image, and the third image from the subtitle motion feature dimension according to the target subtitle change mode to obtain an augmented image group, including:
[0313] Generate a first subtitle, and obtain a first image position among the first image, the second image, and the third image;
[0314] Adding the first subtitle at the position of the first image in the first image, the second image and the third image respectively, to obtain the first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation;
[0315] The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are taken as an augmented image group.
[0316] In one implementation, if the target subtitle change mode is a subtitle disappearance mode, the augmentation module 1703 augments the first image, the second image, and the third image from the subtitle motion feature dimension according to the target subtitle change mode to obtain an augmented image group, including:
[0317] Generate a first subtitle, and obtain a first image position in the first image and the second image;
[0318] Adding the first subtitle at the position of the first image in the first image and the second image respectively to obtain the first image after subtitle augmentation and the second image after subtitle augmentation;
[0319] using the third image as the third image after the captions are augmented;
[0320] The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are taken as an augmented image group.
[0321] In one implementation, if the target subtitle change mode is a subtitle appearance mode, the augmentation module 1703 augments the first image, the second image, and the third image from the subtitle motion feature dimension according to the target subtitle change mode to obtain an augmented image group, including:
[0322] The first image and the second image are used as the first image after subtitle augmentation and the second image after subtitle augmentation respectively;
[0323] Generate a first subtitle, obtain the first image position in the third image, add the first subtitle at the first image position in the third image, and obtain a third image after the subtitle is augmented;
[0324] The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are taken as an augmented image group.
[0325] In one implementation, if the target subtitle change mode is a subtitle switching mode, the augmentation module 1703 performs augmentation processing on the first image, the second image, and the third image from the subtitle motion feature dimension according to the target subtitle change mode to obtain an augmented image group, including:
[0326] Generate a first subtitle and a second subtitle, and obtain a first image position in the first image, the second image, and the third image;
[0327] Adding the first subtitle at the position of the first image in the first image and the second image respectively to obtain the first image after subtitle augmentation and the second image after subtitle augmentation;
[0328] Adding a second subtitle at the position of the first image in the third image to obtain a third image with the subtitles augmented;
[0329] The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are taken as an augmented image group.
[0330] In one implementation, the ternary image group sequentially includes a first image at a first moment, a second image at a second moment, and a third image at a third moment, the second moment is between the first moment and the third moment, the M motion augmentation strategies include a text augmentation strategy, and the motion feature dimension indicated by the text augmentation strategy is a text motion feature dimension;
[0331] The augmentation module 1703 acquires the augmented image group based on the M motion augmentation strategies and the ternary image group, including:
[0332] Selecting a target text motion mode for the ternary image group from a plurality of preset text motion modes;
[0333] A target text to be added is generated, and the first image, the second image, and the third image are augmented according to the target text motion mode and based on the target text from the text motion feature dimension to obtain an augmented image group.
[0334] In one implementation, if the target text motion mode is a text translation mode, the augmentation module 1703 augments the first image, the second image, and the third image according to the target text motion mode based on the target text from the text motion feature dimension to obtain an augmented image group, including:
[0335] Obtaining a text displacement corresponding to the target text and a second image position in the first image;
[0336] Adding the target text at the position of the second image in the first image to obtain the first image after text augmentation;
[0337] Calculating a third image position in the second image based on the second image position, the second moment, and the text displacement, and adding the target text at the third image position in the second image to obtain a second image after text augmentation;
[0338] Calculating a fourth image position in the third image based on the second image position and the text displacement, and adding target text at the fourth image position in the third image to obtain a third image after text augmentation;
[0339] The first image after text augmentation, the second image after text augmentation, and the third image after text augmentation are taken as an augmented image group.
[0340] In one implementation, if the target text motion mode is a text zoom mode, the augmentation module 1703 augments the first image, the second image, and the third image according to the target text motion mode based on the target text from the text motion feature dimension to obtain an augmented image group, including:
[0341] acquiring a fifth image position among the first image, the second image, and the third image;
[0342] Adding the target text at the fifth image position in the first image to obtain the first image after text augmentation;
[0343] Acquire a first zoom scale for the target text, and calculate a second zoom scale for the target text based on the first zoom scale and the second moment;
[0344] The target text is scaled using the second scaling scale to obtain the first scaled text, and the target text is scaled using the first scaling scale to obtain the second scaled text;
[0345] Adding the first scaled text at the fifth image position in the second image to obtain the text-augmented second image, and adding the second scaled text at the fifth image position in the third image to obtain the text-augmented third image;
[0346] The first image after text augmentation, the second image after text augmentation, and the third image after text augmentation are taken as an augmented image group.
[0347] In one implementation, if the target text motion mode is a text rotation mode, the augmentation module 1703 augments the first image, the second image, and the third image according to the target text motion mode based on the target text from the text motion feature dimension to obtain an augmented image group, including:
[0348] acquiring a sixth image position among the first image, the second image, and the third image;
[0349] Adding target text at the sixth image position in the first image, the second image and the third image respectively, to obtain a first image after adding text, a second image after adding text and a third image after adding text;
[0350] Obtaining a first rotation angle for the target text, and rotating the target text in the first image after the text is added by the first rotation angle to obtain the first image after the text is augmented;
[0351] Acquire a second rotation angle for the target text, calculate a third rotation angle for the target text based on the first rotation angle, the second rotation angle and the second moment, and rotate the target text in the second image after the text is added by the third rotation angle to obtain the second image after the text is augmented;
[0352] Calculating a fourth rotation angle for the target text based on the first rotation angle and the second rotation angle, and rotating the target text in the third image after the text is added by the fourth rotation angle to obtain the third image after the text is augmented;
[0353] The first image after text augmentation, the second image after text augmentation, and the third image after text augmentation are taken as an augmented image group.
[0354] In one implementation, the augmentation module 1703 acquires the augmented image group based on M motion augmentation strategies and the ternary image group, including:
[0355] If the M motion augmentation strategies are serial, then according to the serial order of the M motion augmentation strategies, the ternary image group is augmented in sequence from the M motion feature dimensions indicated by the M motion augmentation strategies to obtain an augmented image group;
[0356] If the M motion augmentation strategies are parallel, the ternary image groups are augmented respectively from the motion feature dimensions indicated by the M motion augmentation strategies to obtain M augmented image groups corresponding to the M motion augmentation strategies.
[0357] In one implementation, the determination module 1702 performs augmentation determination on the ternary image group based on the augmentation strategy set to obtain augmentation indication information for the ternary image group, including:
[0358] Obtain the augmentation judgment probability corresponding to each motion augmentation strategy in the augmentation strategy set;
[0359] Performing augmentation judgment on each motion augmentation strategy based on the augmentation judgment probability corresponding to each motion augmentation strategy, and obtaining the judgment result corresponding to each motion augmentation strategy. The judgment result corresponding to any motion augmentation strategy is an adopted result or a non-adopted result.
[0360] The augmentation indication information is determined by determining that the motion augmentation strategy of the determined result is an adopted result, and the M types of motion augmentation strategies include the motion augmentation strategy of the determined result being an adopted result.
[0361] In one implementation, any one of the N motion augmentation strategies is a target motion augmentation strategy; the judgment module 1702 performs augmentation judgment on each motion augmentation strategy based on the augmentation judgment probability corresponding to each motion augmentation strategy, and obtains the judgment result corresponding to each motion augmentation strategy, including:
[0362] Obtain the total probability interval and the probability sub-interval corresponding to the augmented judgment probability corresponding to the target motion augmentation strategy within the total probability interval;
[0363] generating a second random number within the total probability interval;
[0364] If the second random number is within the probability subinterval, determining that the judgment result corresponding to the target motion augmentation strategy is the adopted result;
[0365] If the second random number is not within the probability subinterval, the judgment result corresponding to the target motion augmentation strategy is determined to be a non-adopted result.
[0366] In one implementation, the augmented image group includes a first augmented image, a second augmented image, and a third augmented image in sequence; the training module 1704 uses the augmented image group to train the interpolation model to obtain the trained interpolation model, including:
[0367] Calling the interpolation model to generate a newly added image between the first augmented image and the third augmented image based on the first augmented image and the third augmented image;
[0368] Obtaining a generation loss of the interpolation model for the new image based on the new image and the second augmented image, where the generation loss is used to reflect the image difference between the new image and the second augmented image;
[0369] The model parameters of the interpolation model are corrected by generating losses to obtain the trained interpolation model.
[0370] In one implementation, the image processing device 170 further includes a frame insertion module 1705, and the frame insertion module 1705 is used to:
[0371] Acquire a first video frame and a second video frame in a video;
[0372] Calling the trained interpolation model to generate a newly added video frame between the first video frame and the second video frame based on the first video frame and the second video frame;
[0373] A newly added video frame is inserted between the first video frame and the second video frame of the video.
[0374] According to one embodiment of the present application, Figure 3 The steps involved in the image processing method shown can be represented by Fig.17 The various modules in the image processing device 170 shown in the figure are executed. For example, Figure 3 The step S101 shown in FIG. 1 can be performed by Fig.17 The acquisition module 1701 in is executed, Figure 3 The step S102 shown in FIG. 1 can be performed by Fig.17 The judgment module 1702 is used to execute; Figure 3 The step S103 shown in FIG. 1 can be performed by Fig.17 The augmentation module 1703 in the embodiment is used to execute, Figure 3 The step S104 shown in FIG. 1 can be performed by Fig.17 The training module 1704 in is used to execute.
[0375] The present application can obtain a ternary image group, which includes three video frames sampled from a sample video; and can perform augmentation judgment on the ternary image group based on an augmentation strategy set to obtain augmentation indication information for the ternary image group, wherein the augmentation strategy set includes N motion augmentation strategies, each motion augmentation strategy is used to indicate augmentation processing from a motion feature dimension of the ternary image group, and the augmentation indication information includes M motion augmentation strategies judged from the N motion augmentation strategies, where N is a positive integer, M is a non-negative integer, and M is less than or equal to N; However, an augmented image group can be obtained based on M motion augmentation strategies and a ternary image group, wherein the method for obtaining the augmented image group includes: if M is greater than 0, the ternary image group is augmented from M motion feature dimensions of the ternary image group according to instructions of the M motion augmentation strategies to obtain an augmented image group; if M is equal to 0, the ternary image group is used as the augmented image group; and the augmented image group can be used to train an interpolation model to obtain a trained interpolation model, and the trained interpolation model is used to perform interpolation processing between video frames of the video. It can be seen that the method proposed in the present application can perform augmentation processing on the ternary image group used to train the interpolation model through N motion feature dimensions indicated by N motion augmentation strategies, thereby taking into account the change modes of each image in the ternary image group in multiple motion feature dimensions, and obtaining an augmented image group with rich change modes between images. Therefore, by training the interpolation model through the augmented image group obtained by augmentation, the robustness and accuracy of the trained interpolation model can be improved, and thus, the interpolation model obtained by training can also achieve excellent interpolation processing between video frames of the video.
[0376] According to one embodiment of the present application, Fig.17 The various modules in the image processing device 170 shown can be separately or all combined into one or several units to constitute, or one (some) of the units can be further divided into multiple smaller sub-units in function, and the same operation can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In practical applications, the functions of a module can also be implemented by multiple units, or the functions of multiple modules can be implemented by one unit. In other embodiments of the present application, the image processing device 170 may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0377] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0378] According to one embodiment of the present application, a computer program capable of executing the steps involved in the corresponding methods shown in the embodiments of the present application can be run on a general-purpose computer device (the computer device may include processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM)) to construct the following. Fig.17 The image processing apparatus 170 shown in FIG. The computer program may be recorded on a computer-readable recording medium, and may be loaded into the computer device through the computer-readable recording medium and run therein.
[0379] See also Fig.18 , Fig.18 Schematic diagram of the structure of a computer device provided in an embodiment of the present application. Fig.18 As shown, the computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, in some embodiments, the computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 1005 may optionally be at least one storage device located away from the aforementioned processor 1001. As Fig.18 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a device control application program.
[0380] exist Fig.18In the computer device 1000 shown, the network interface 1004 can provide a network communication function; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:
[0381] Acquire a ternary image group, where the ternary image group includes three video frames sampled from a sample video;
[0382] Based on the augmentation strategy set, the ternary image group is augmented to determine, and augmentation indication information for the ternary image group is obtained, wherein the augmentation strategy set includes N motion augmentation strategies, each motion augmentation strategy is used to indicate augmentation processing from a motion feature dimension of the ternary image group, and the augmentation indication information includes M motion augmentation strategies determined from the N motion augmentation strategies, where N is a positive integer, M is a non-negative integer, and M is less than or equal to N;
[0383] Acquiring an augmented image group based on M motion augmentation strategies and the ternary image group, wherein the method for acquiring the augmented image group includes: if M is greater than 0, augmenting the ternary image group from M motion feature dimensions of the ternary image group according to instructions of the M motion augmentation strategies to obtain the augmented image group; if M is equal to 0, using the ternary image group as the augmented image group;
[0384] The interpolation model is trained by using the augmented image group to obtain a trained interpolation model, and the trained interpolation model is used to perform interpolation processing between video frames of the video.
[0385] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the description of the above-mentioned image processing method in each embodiment of the present application, and can also execute the above-mentioned Fig.17 The description of the image processing device 170 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated here either.
[0386] In addition, it should be pointed out here that: the present application also provides a computer-readable storage medium, and a computer program is stored in the computer-readable storage medium. When the processor executes the computer program, the description of the image processing method in each embodiment of the present application can be executed. Therefore, it will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated. For technical details not disclosed in the computer storage medium embodiment involved in the present application, please refer to the description of the method embodiment of the present application.
[0387] As an example, the above-mentioned computer program may be deployed on one computer device for execution, or deployed on multiple computer devices located at one location for execution, or executed on multiple computer devices distributed at multiple locations and interconnected by a communication network. Multiple computer devices distributed at multiple locations and interconnected by a communication network may constitute a blockchain network.
[0388] The computer-readable storage medium may be an internal storage unit of the computer device, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of the computer device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0389] The present application provides a computer program product, which includes a computer program, which is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the description of the above-mentioned image processing method in each embodiment of the present application, and therefore, it will not be repeated here. In addition, the description of the beneficial effects of adopting the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in the present application, please refer to the description of the method embodiment of the present application.
[0390] The terms "first", "second", etc. in the description, claims, and drawings of the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, device, product, or equipment that includes a series of steps or units is not limited to the listed steps or modules, but optionally includes steps or modules that are not listed, or optionally includes other step units inherent to these processes, methods, devices, products, or equipment.
[0391] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0392] The above disclosure is only the preferred embodiment of the present application, which certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a ternary image group, wherein the ternary image group includes three video frames sampled from a sample video; Performing augmentation judgment on the ternary image group based on an augmentation strategy set to obtain augmentation indication information for the ternary image group, wherein the augmentation strategy set includes N motion augmentation strategies, each motion augmentation strategy is used to indicate augmentation processing from a motion feature dimension of the ternary image group, and the augmentation indication information includes M motion augmentation strategies judged from the N motion augmentation strategies, where N is a positive integer, M is a non-negative integer, and M is less than or equal to N; Acquiring an augmented image group based on the M motion augmentation strategies and the ternary image group, wherein the method for acquiring the augmented image group comprises: if M is greater than 0, augmenting the ternary image group from M motion feature dimensions of the ternary image group according to instructions of the M motion augmentation strategies to obtain the augmented image group; if M is equal to 0, using the ternary image group as the augmented image group; The augmented image group is used to train the interpolation model to obtain a trained interpolation model, and the trained interpolation model is used to perform interpolation processing between video frames of the video.
2. The method according to claim 1, characterized in that The M motion augmentation strategies include a brightness augmentation strategy, and the motion feature dimension indicated by the brightness augmentation strategy is a brightness motion feature dimension; The acquiring the augmented image group based on the M motion augmentation strategies and the ternary image group includes: Acquire a preset brightness variation range, and determine a first brightness variation amplitude for the ternary image group within the brightness variation range; The ternary image group is augmented from the brightness motion feature dimension based on the first brightness change amplitude to obtain the augmented image group.
3. The method according to claim 2, characterized in that The ternary image group includes, in sequence, a first image at a first moment, a second image at a second moment, and a third image at a third moment, wherein the second moment is between the first moment and the third moment; and the augmenting process is performed on the ternary image group from the brightness feature dimension based on the first brightness change amplitude to obtain the augmented image group, including: Using the first image as the first image after brightness enhancement; Calculating a second brightness change amplitude for the second image based on the second moment and the first brightness change amplitude, and performing weighted processing on a color pixel value of each pixel point in the second image using the second brightness change amplitude to obtain a second image after brightness enhancement; Using the first brightness variation amplitude to perform weighted processing on the color pixel value of each pixel point in the third image to obtain a third image with brightness enhanced; The first image after brightness enhancement, the second image after brightness enhancement, and the third image after brightness enhancement are used as the augmented image group.
4. The method according to claim 2, characterized in that Determining a first brightness change amplitude for the ternary image group within the brightness change range includes: Generate a first random number within the brightness variation range; The generated first random number is used as the first brightness change amplitude for the ternary image group.
5. The method according to claim 1, characterized in that The ternary image group includes a first image, a second image, and a third image in sequence, the M motion augmentation strategies include a subtitle augmentation strategy, and the motion feature dimension indicated by the subtitle augmentation strategy is a subtitle motion feature dimension; The acquiring the augmented image group based on the M motion augmentation strategies and the ternary image group includes: Selecting a target subtitle changing mode for the ternary image group from a plurality of preset subtitle changing modes; The first image, the second image and the third image are augmented according to the target subtitle change mode and from the subtitle motion feature dimension to obtain the augmented image group.
6. The method according to claim 5, characterized in that If the target subtitle change mode is a subtitle unchanged mode, the first image, the second image and the third image are augmented from the subtitle motion feature dimension according to the target subtitle change mode to obtain the augmented image group, including: Generate a first subtitle, and obtain a first image position in the first image, the second image, and the third image; adding the first subtitles at the positions of the first image in the first image, the second image, and the third image respectively, to obtain a first image augmented with subtitles, a second image augmented with subtitles, and a third image augmented with subtitles; The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are used as the augmented image group.
7. The method according to claim 5, characterized in that If the target subtitle change mode is a subtitle disappearance mode, the first image, the second image, and the third image are augmented from the subtitle motion feature dimension according to the target subtitle change mode to obtain the augmented image group, including: Generate a first subtitle, and obtain a first image position in the first image and the second image; Adding the first subtitle at the position of the first image in the first image and the second image respectively to obtain a first image augmented with subtitles and a second image augmented with subtitles; Using the third image as the third image after the captions are augmented; The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are used as the augmented image group.
8. The method according to claim 5, characterized in that If the target subtitle change mode is a subtitle appearance mode, the first image, the second image, and the third image are augmented from the subtitle motion feature dimension according to the target subtitle change mode to obtain the augmented image group, including: Using the first image and the second image as the first image after subtitle augmentation and the second image after subtitle augmentation respectively; generating a first subtitle, obtaining a first image position in the third image, and adding the first subtitle at the first image position in the third image to obtain a third image after subtitle augmentation; The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are used as the augmented image group.
9. The method according to claim 5, characterized in that If the target subtitle change mode is a subtitle switching mode, the first image, the second image, and the third image are augmented from the subtitle motion feature dimension according to the target subtitle change mode to obtain the augmented image group, including: Generate a first subtitle and a second subtitle, and obtain a first image position in the first image, the second image, and the third image; Adding the first subtitle at the position of the first image in the first image and the second image respectively to obtain a first image augmented with subtitles and a second image augmented with subtitles; adding the second subtitle at the position of the first image in the third image to obtain a third image with subtitle augmentation; The first image after subtitle augmentation, the second image after subtitle augmentation and the third image after subtitle augmentation are used as the augmented image group.
10. The method according to claim 1, characterized in that The ternary image group sequentially includes a first image at a first moment, a second image at a second moment, and a third image at a third moment, the second moment is between the first moment and the third moment, the M motion augmentation strategies include a text augmentation strategy, and the motion feature dimension indicated by the text augmentation strategy is a text motion feature dimension; The acquiring the augmented image group based on the M motion augmentation strategies and the ternary image group includes: Selecting a target text motion mode for the ternary image group from a plurality of preset text motion modes; Generate a target text to be added, and perform augmentation processing on the first image, the second image and the third image based on the target text from the text motion feature dimension according to the target text motion mode to obtain the augmented image group.
11. The method according to claim 10, characterized in that If the target text movement mode is a text translation mode, the first image, the second image, and the third image are augmented according to the target text movement mode based on the target text from the text movement feature dimension to obtain the augmented image group, including: Acquire a text displacement corresponding to the target text and a second image position in the first image; Adding the target text at the position of the second image in the first image to obtain a first image after text augmentation; Calculating a third image position in the second image based on the second image position, the second moment and the text displacement, and adding the target text at the third image position in the second image to obtain a second image after text augmentation; Calculating a fourth image position in the third image based on the second image position and the text displacement, and adding the target text at the fourth image position in the third image to obtain a third image after text augmentation; The first image after text augmentation, the second image after text augmentation and the third image after text augmentation are used as the augmented image group.
12. The method according to claim 10, characterized in that If the target text motion mode is a text zoom mode, the first image, the second image, and the third image are augmented according to the target text motion mode based on the target text from the text motion feature dimension to obtain the augmented image group, including: Acquire a fifth image position among the first image, the second image and the third image; Adding the target text at the fifth image position in the first image to obtain a first image after text augmentation; Acquire a first zoom scale for the target text, and calculate a second zoom scale for the target text based on the first zoom scale and the second moment; Scaling the target text using the second scaling scale to obtain a first scaled text, and scaling the target text using the first scaling scale to obtain a second scaled text; adding the first scaled text at the fifth image position in the second image to obtain a text-augmented second image, and adding the second scaled text at the fifth image position in the third image to obtain a text-augmented third image; The first image after text augmentation, the second image after text augmentation and the third image after text augmentation are used as the augmented image group.
13. The method according to claim 10, characterized in that If the target text movement mode is a text rotation mode, the first image, the second image and the third image are augmented according to the target text movement mode based on the target text from the text movement feature dimension to obtain the augmented image group, including: Acquire a sixth image position among the first image, the second image, and the third image; Adding the target text at the sixth image position in the first image, the second image, and the third image respectively to obtain a first image after adding the text, a second image after adding the text, and a third image after adding the text; Acquire a first rotation angle for the target text, and rotate the target text in the first image after the text is added by the first rotation angle to obtain a first image after the text is augmented; Acquire a second rotation angle for the target text, calculate a third rotation angle for the target text based on the first rotation angle, the second rotation angle and the second moment, and rotate the target text in the second image after adding text by the third rotation angle to obtain a second image after text augmentation; Calculating a fourth rotation angle for the target text based on the first rotation angle and the second rotation angle, and rotating the target text in the third image after the text is added by the fourth rotation angle to obtain a third image after the text is augmented; The first image after text augmentation, the second image after text augmentation and the third image after text augmentation are used as the augmented image group.
14. The method according to claim 1, wherein: The performing augmentation judgment on the ternary image group based on the augmentation strategy set to obtain augmentation indication information for the ternary image group includes: Obtaining the augmentation judgment probability corresponding to each motion augmentation strategy in the augmentation strategy set; Performing augmentation judgment on each motion augmentation strategy based on the augmentation judgment probability corresponding to each motion augmentation strategy, and obtaining a judgment result corresponding to each motion augmentation strategy, wherein the judgment result corresponding to any motion augmentation strategy is an adopted result or a non-adopted result; The augmentation indication information is determined by determining that a judgment result is the motion augmentation strategy of the adopted result, and the M types of motion augmentation strategies include the motion augmentation strategy of which the judgment result is the adopted result.
15. The method according to claim 14, characterized in that Any one of the N motion augmentation strategies is a target motion augmentation strategy; performing augmentation judgment on each motion augmentation strategy based on the augmentation judgment probability corresponding to each motion augmentation strategy, and obtaining a judgment result corresponding to each motion augmentation strategy, includes: Obtaining a total probability interval and a probability sub-interval corresponding to the augmented judgment probability corresponding to the target motion augmentation strategy within the total probability interval; generating a second random number within the total probability interval; If the second random number is within the probability subinterval, determining that the judgment result corresponding to the target motion augmentation strategy is the adopted result; If the second random number is not within the probability sub-interval, it is determined that the judgment result corresponding to the target motion augmentation strategy is the non-adoption result.
16. The method according to claim 1, wherein: The augmented image group includes a first augmented image, a second augmented image and a third augmented image in sequence; and the augmented image group is used to train the interpolation model to obtain the trained interpolation model, including: Calling the interpolation model to generate a newly added image between the first augmented image and the third augmented image based on the first augmented image and the third augmented image; Acquire, based on the newly added image and the second augmented image, a generation loss of the interpolation model for the newly added image, wherein the generation loss is used to reflect an image difference between the newly added image and the second augmented image; The model parameters of the interpolation model are corrected by using the generation loss to obtain the trained interpolation model.
17. An image processing device, characterized in that: The method comprises: An acquisition module, used for acquiring a ternary image group, wherein the ternary image group includes three video frames sampled from a sample video; a judgment module, configured to perform augmentation judgment on the ternary image group based on an augmentation strategy set, and obtain augmentation indication information for the ternary image group, wherein the augmentation strategy set includes N motion augmentation strategies, each of which is used to indicate augmentation processing from a motion feature dimension of the ternary image group, and the augmentation indication information includes M motion augmentation strategies obtained by judgment from the N motion augmentation strategies, where N is a positive integer, M is a non-negative integer, and M is less than or equal to N; an augmentation module, configured to obtain an augmented image group based on the M motion augmentation strategies and the ternary image group, wherein the method for obtaining the augmented image group comprises: if M is greater than 0, augmenting the ternary image group from the M motion feature dimensions of the ternary image group according to instructions of the M motion augmentation strategies to obtain the augmented image group; and if M is equal to 0, using the ternary image group as the augmented image group; The training module is used to train the interpolation model using the augmented image group to obtain the trained interpolation model, and the trained interpolation model is used to perform interpolation processing between video frames of the video.
18. A computer program product, characterized in that The invention comprises a computer program, which implements the steps of the method according to any one of claims 1 to 16 when being executed by a processor.
19. A computer device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 16.
20. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Training data augmentation method and device, equipment and storage medium
CN114548229A
Video frame insertion method and device, equipment and storage medium
CN115942045A
Lightweight depth video frame interpolation method based on attention mechanism
CN116320246A
Predictive interpolation of a video signal
US20040017852A1
Method and system for selecting augmentation strategy for image data
WO2021164228A1