Methods, apparatus, media, and devices for processing a plurality of image sequences

By processing multiple image sequences and using the splitting and merging of RGB and Alpha channels, video gifts are automatically synthesized, solving the problems of monotonous and costly gift effects on live streaming platforms, and achieving unique gift effects and efficient resource utilization.

CN115187763BActive Publication Date: 2026-05-29SHANGHAI BILIBILI TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI BILIBILI TECH CO LTD
Filing Date
2022-07-07
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Gift effects on live streaming platforms are implemented using video, resulting in limited effects and the need for manually customized transparent videos, which increases costs and resource consumption.

Method used

By processing multiple image sequences, video gifts are automatically synthesized. By splitting and merging RGB and Alpha channels, image format conversion and compositing are achieved, reducing the processing cost of multi-video layer compositing.

Benefits of technology

It enables unique gift effects for multiple users, reduces labor costs, and improves processing efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187763B_ABST
    Figure CN115187763B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, device, computer program product, non-transitory computer-readable storage medium and electronic device for processing a plurality of image sequences. The method comprises: performing the following merging processing on a plurality of images with the same sequence number in the plurality of image sequences: decoding an image in a current image sequence to obtain a current to-be-combined image; and merging the current to-be-combined image and a previous combined image to obtain a current combined image, wherein the first combined image is obtained by: decoding an image in a first image sequence to obtain a first to-be-combined image; decoding an image in a second image sequence to obtain a second to-be-combined image; and merging the second to-be-combined image and the first to-be-combined image to obtain the first combined image. According to the embodiments provided by the present disclosure, a video gift can be automatically combined, and the labor cost of manual processing of the combination of multiple video layers is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of image and video processing technology, and more specifically to a method, apparatus, computer program product, non-transitory computer-readable storage medium, and electronic device for processing multiple image sequences. Background Technology

[0002] This section is intended to introduce aspects of the art that may relate to the various aspects of this disclosure described below and / or claimed. It is believed that this section will help provide background information to facilitate a better understanding of the various aspects of this disclosure. Therefore, it should be understood that these descriptions should be interpreted in this context and not as an admission of prior art.

[0003] Gifts on live streaming platforms can be created using vector animations. However, vector animations consume more computing resources during playback, and their gift effects are not as rich as those of videos. Therefore, gifts on live streaming platforms are increasingly being created using videos. Since gift effects on live streaming platforms need to be transparent, while videos are usually opaque, professional video editing software (such as Adobe) is needed to customize transparent video gift effects.

[0004] However, in this approach, the gift effect is relatively simple, and all users on the live streaming platform receive the same gift effect. Summary of the Invention

[0005] The purpose of this disclosure is to provide a method, apparatus, computer program product, non-transitory computer-readable storage medium, and electronic device for processing multiple image sequences to avoid the manual costs of composite processing of multiple video layers.

[0006] According to a first aspect of this disclosure, a method for processing multiple image sequences is provided. The method includes: performing the following merging process on multiple images having the same sequence number in the multiple image sequences: decoding the images in the current image sequence to obtain a current image to be synthesized; merging the current image to be synthesized with a previous synthesized image to obtain a current synthesized image, wherein the first synthesized image is obtained by: decoding the images in the first image sequence to obtain a first image to be synthesized; decoding the images in the second image sequence to obtain a second image to be synthesized; merging the second image to be synthesized with the first image to be synthesized to obtain a first synthesized image.

[0007] According to a second aspect of this disclosure, an apparatus for processing multiple image sequences is provided, comprising: a processing module configured to perform the following merging process on multiple images having the same sequence number in the multiple image sequences: decoding the images in the current image sequence to obtain a current image to be synthesized; merging the current image to be synthesized with a previous synthesized image to obtain a current synthesized image, wherein the first synthesized image is obtained by: decoding the images in the first image sequence to obtain a first image to be synthesized; decoding the images in the second image sequence to obtain a second image to be synthesized; merging the second image to be synthesized with the first image to be synthesized to obtain a first synthesized image.

[0008] According to a third aspect of this disclosure, a computer program product is provided, including program code instructions that, when executed by a computer, cause the computer to perform the method described according to a first aspect of this disclosure.

[0009] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described according to a first aspect of this disclosure.

[0010] According to a fifth aspect of this disclosure, an electronic device is provided, comprising: a processor; a memory in electronic communication with the processor; and instructions stored in the memory and executable by the processor to cause the electronic device to perform the method according to a first aspect of this disclosure.

[0011] According to the embodiments provided in this disclosure, video gifts can be synthesized automatically, avoiding the manual costs of compositing multiple video layers.

[0012] It should be understood that the content described in this section is not intended to identify key or essential features of the claimed invention, nor is it intended to be used alone to determine the scope of the claimed invention. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In all drawings, the same reference numerals refer to similar but not necessarily the same elements.

[0014] Figure 1An example of splitting the RGB and Alpha channels of an image according to this disclosure is shown.

[0015] Figure 2 A flowchart is shown illustrating a method for merging the first frame of n (n is an integer greater than 1) image sequences according to an embodiment of the present disclosure.

[0016] Figure 3 An example of a positional parameter associated with the positions of a first region and a second region in an image, according to an embodiment of the present disclosure, is shown.

[0017] Figure 4 Examples of normalized first and second region images according to embodiments of this disclosure are shown.

[0018] Figure 5 An example of multiple layers according to an embodiment of this disclosure is shown.

[0019] Figure 6 An example of an output image according to an embodiment of this disclosure is shown.

[0020] Figure 7 A flowchart illustrating a method for processing multiple image sequences according to some embodiments of the present disclosure is shown.

[0021] Figure 8 An exemplary block diagram of an apparatus for processing multiple images according to embodiments of the present disclosure is shown.

[0022] Figure 9 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown.

[0023] Specific implementation method

[0024] The present disclosure will be described more fully below with reference to the accompanying drawings. However, the present disclosure may be embodied in many alternative forms and should not be construed as limited to the embodiments described herein. Therefore, although the present disclosure is readily adaptable to various modifications and alternatives, specific embodiments thereof are shown by way of example in the accompanying drawings and will be described in detail herein. However, it should be understood that this is not intended to limit the present disclosure to the specific forms disclosed, but rather, the present disclosure covers all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure as defined by the claims.

[0025] It should be understood that although various elements may be described herein using terms such as first, second, etc., these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the teachings of this disclosure.

[0026] This document describes several examples using block diagrams and / or flowcharts, where each block represents a section of circuitry, modular blocks, or code comprising one or more executable instructions for implementing a specified logical function. It should also be noted that in other implementations, the functions described in the blocks may occur in a different order. For example, depending on the function involved, two blocks shown consecutively may actually execute substantially simultaneously, or these blocks may sometimes execute in reverse order.

[0027] The phrases “according to…example” or “in…example” used in this document mean that a particular feature, structure, or characteristic described in connection with the example can be included in at least one implementation of this disclosure. The phrases “according to…example” or “in…example” appearing in different places throughout this document do not necessarily refer to the same example, nor are they necessarily separate or alternative examples that are mutually exclusive with other examples.

[0028] RGB color mode is an industry-standard color model. RGB represents the red, green, and blue channels, and these colors are mixed and superimposed. In the RGB encoding method, each color can be represented by three variables to indicate the intensity of red, green, and blue. YUV (also known as YCrCb) is a color encoding method used in modern television systems. In modern color television systems, three-tube color cameras or color CCD cameras are typically used to capture images. The acquired color image signals are then separated, amplified, and corrected separately to obtain RGB values. These are then processed by a matrix transformation circuit to obtain the luminance signal Y and two color difference signals RY (U) and BY (V). In the YUV encoding method, "Y" represents luminance (or luminance), which is the grayscale value, while "U" and "V" represent chrominance (or chroma), which describe the color and saturation of the image and are used to specify the color of a pixel.

[0029] Video frames can be composed of multiple different channels. For example, an RGB video frame consists of a red (R) channel, a green (G) channel, and a blue (B) channel, while a YUV video frame consists of a luminance channel (Y), a first chroma channel (U), and a second chroma channel (V). In this document, the terms "image," "picture," "frame," "picture," and "video frame" are used interchangeably, and "layer," "image sequence," and "video" are used interchangeably.

[0030] The alpha channel (or alpha channel) refers to the transparency and semi-transparency of an image. For example, in a bitmap using 16 bits per pixel, each pixel can be represented by 5 bits for red, 5 bits for green, 5 bits for blue, and the last bit is the alpha. In this case, it either represents transparency or not, because the alpha bit only has two possible representations: 0 or 1. As another example, in a bitmap using 32 bits, red, green, blue, and the alpha channel are each represented by 8 bits. In this case, the alpha channel can not only represent transparency or opacity, but also 256 levels of semi-transparency, because the 8 bits of the alpha channel can have 256 different possible representations.

[0031] In this disclosure, the effect of image transparency can be achieved by splitting the RGB channel data and Alpha channel data of an image into different locations in a non-transparent video. Figure 1 An example of splitting the RGB and alpha channels of an image according to this disclosure is shown. Figure 1 As shown, the non-transparent video frame 10 includes rectangular regions 20 and 30, where the RGB channel data of the image is placed in rectangular region 20, and the alpha channel data of the image is placed in region 30. Figure 1 In the example, the RGB value of the pure black pixel at rectangular area 20 is [0,0,0], and the RGB value of the pure white pixel is [255,255,255]. In this example, the data of the image's alpha channel is mapped to black and white data. For example, if the value of the image's alpha channel at a certain pixel is [0], then that pixel can be mapped to an RGB value of [0,0,0], which is pure black; if the value of the image's alpha channel at a certain pixel is

[255] , then that pixel can be mapped to an RGB value of [255,255,255], which is pure white.

[0032] The following describes a method for processing multiple image sequences according to an embodiment of this disclosure, based on the RGB channel data and Alpha channel data of the split images.

[0033] Figure 2 A flowchart illustrating a method for merging the first frame of n image sequences (n being an integer greater than 1) according to an embodiment of this disclosure is shown. Figure 2 As shown, the method 200 includes:

[0034] Decode the first frame in the first image sequence to obtain the first image to be synthesized;

[0035] Decode the first frame in the second image sequence to obtain the second image to be synthesized;

[0036] The second image to be synthesized and the first image to be synthesized are merged to obtain the first synthesized image;

[0037] Decode the first frame of the third image sequence to obtain the third image to be synthesized;

[0038] The third image to be synthesized is merged with the first image to be synthesized to obtain the second image;

[0039] Similarly, the first frame in the nth image sequence is decoded to obtain the nth image to be synthesized;

[0040] The nth image to be synthesized and the (n-2)th synthesized image are merged to obtain the (n-1)th synthesized image.

[0041] In this example, after merging the first frame of n image sequences using the above method, the (n-1)th composite image is the merged image.

[0042] In this example, the method for merging the second, third, ..., mth frames (where m is an integer greater than 1) in a sequence of n images is the same as that for the first frame.

[0043] The method for processing multiple image sequences provided in the embodiments of this disclosure can automatically synthesize video gifts, avoiding the manual cost of synthesizing multiple video layers.

[0044] exist Figure 2 In the example, decoding one frame from the image sequence yields the following image to be synthesized:

[0045] Based on the color information and transparency information of the image, the image is converted from the first format to a second format having the color channel and transparency channel.

[0046] In this example, the image or frame may include a first region and a second region. The first region may contain the image's color information, and the second region may contain the image's transparency information. The image may be a first format with color channels. In this example, the image's color information may be non-alpha channel data, such as RGB channel data or YUV channel data, and the image's transparency information may be information corresponding to alpha channel data, such as mapping alpha channel data to RGB channel data. As an example, Figure 1The rectangular area at position 20 contains the image's color information, and the rectangular area at position 30 contains the image's transparency information. In the context of a live streaming platform, the color information in this example could be the color of the gift, and the transparency information could be the gift's transparency information.

[0047] In this example, the first and second regions can be located anywhere in the image. As an example, Figure 1 Rectangular region 20 can be an example of the first region, and rectangular region 30 can be an example of the second region.

[0048] In this example, the color channel can be any non-Alpha channel image channel, such as RGB or YUV channels, while the transparency channel is the Alpha channel described above. In this example, the image format can be represented by the way the pixels in the image are represented. A first-format image can represent pixels using color channels, while a second-format image can represent pixels using both color channels and transparency channels. For example, if the pixel value in the image is [0,0,0] (representing a value of 0 for all three RGB channels), then the image is in the first format; if the pixel value in the image is [0,0,0,1] (representing a value of 0 for all three RGB channels and a value of 1 for the Alpha channel), then the image is in the second format.

[0049] In this example, the image's color and transparency information can be used to convert the image from the first format to the second format. The following section combines... Figure 1 The steps for converting image formats are explained. For example... Figure 1 As shown, if the RGB value of a certain pixel in the image at the rectangular area 20 is [0,0,0] (i.e., pure black), and the RGB value of the transparency information corresponding to that pixel is [0,0,0] (i.e., pure black), then the value of that pixel can be converted to [0,0,0,0] (representing that the values ​​of the three RGB channels are 0 and the value of the Alpha channel is 0), thereby converting the format of that pixel from the first format to the second format.

[0050] In this example, multiple images in a first format can be converted into images in a second format according to a preset order. In this example, the preset order can be a user-specified order. For example, given images A, B, and C in the first format, the format conversion can be performed sequentially in the order A→B→C, or in the order A→C→B. Therefore, multiple images in the first format can be converted in various different orders.

[0051] In this example, the preset order of the n image sequences (e.g., the first image sequence, the second image sequence, etc.) can be a user-specified order, allowing multiple images in the second format to be merged sequentially. In this example, merging images can also be a merging of multiple image channels, such as merging the alpha channels of multiple images.

[0052] The method for processing multiple image sequences provided in the embodiments of this disclosure can synthesize multiple video layers into multiple different videos by using a user-defined synthesis order, thereby synthesizing different gift effects for each user.

[0053] exist Figure 2 In the example, converting the image from the first format to a second format having the color channel and the transparency channel based on the color information and transparency information of the image may include:

[0054] Step S2022: Based on the positional parameters associated with the positions of the first region and the second region in the image, the image is segmented into a first region image and a second region image.

[0055] Figure 3 An example of positional parameters associated with the positions of the first and second regions in an image, according to an embodiment of this disclosure, is shown. Figure 3 As shown, the left half of the image is the first region of this disclosure, and the upper right half is the second region of this disclosure. The coordinates of the top left pixel of the image are (X=0, Y=0). The position parameters of the first region are a rectangular area with a height (H) of 1280 and a width (W) of 720, starting from the zero point (X=0, Y=0); the position parameters of the second region are a rectangular area with a height (H) of 640 and a width (W) of 360, starting from (X=724, Y=0). In this example, the image can be segmented according to the position parameters of the first and second regions, for example, the first and second regions can be cut out from the image respectively. In this example, after cutting out the first region from the image, the first region image can be obtained, and after cutting out the second region from the image, the second region image can be obtained.

[0056] Step S2024: Normalize the size of the first region image and the second region image.

[0057] In this example, resampling can normalize the data sizes of the first and second region images. In some optional examples, the sizes of the first and second region images can be normalized using an image scaling algorithm. In this example, the image scaling algorithm can include linear interpolation, bicubic interpolation, Lanczos interpolation, and new edge-directed interpolation. The linear interpolation algorithm (Bilinear algorithm) performs interpolation in two directions at once, interpolating through four adjacent pixels to obtain the desired pixel. The bicubic interpolation algorithm (Bicubic algorithm) considers not only the influence of the gray values ​​of the four directly adjacent pixels but also the influence of their gray value change rates, thus producing smoother edges than the linear interpolation algorithm. The new edge-directed interpolation algorithm derives the prediction coefficients for optimal linear MMSE prediction by utilizing local covariance properties.

[0058] Step S2026: Map the color information contained in the normalized first region image and the transparency information contained in the normalized second region image to the pixels of the color channel and the transparency channel of the image to be synthesized.

[0059] Figure 4 Examples of normalized first and second region images according to embodiments of this disclosure are shown. Figure 4 As shown, the image on the left is the normalized first region image, and the image on the right is the normalized second region image. The first and second region images have the same size. In this example, the color information in the first region image and the transparency information in the second region image can be mapped to pixels in the color channels and pixels in the transparency channels. For example, if the pixel value at coordinates (X=6, Y=6) in the first region image is [0,0,0] (representing that the values ​​of the RGB channels are 0), and the pixel value at the same coordinate in the second region image is [0,0,0] (representing that the values ​​of the RGB channels are 0), then after mapping, the pixel value at that coordinate in the image is [0,0,0,0] (representing that the values ​​of the RGB channels are 0, and the value of the Alpha channel is 0). The mapped RGB channel pixel value at that coordinate is the pixel value of the first region image at that coordinate, and the mapped Alpha channel pixel value is the mapped value of the second region image at that coordinate.

[0060] The method for processing multiple image sequences provided in the embodiments of this disclosure simplifies the process of converting image formats and improves processing efficiency by performing image segmentation and normalization operations.

[0061] In some embodiments, the plurality of image sequences includes an image sequence belonging to the background layer and an image sequence belonging to the foreground layer, and, Figure 2 The current image to be composited is the foreground layer, and the previous composited image is the background layer. In this example, the image sequence belonging to the background layer can be the bottommost layer among multiple layers, and the image sequence belonging to the foreground layer can be the topmost layer among multiple layers. Figure 5 An example of multiple layers according to an embodiment of this disclosure is shown. For example... Figure 5 As shown, the background video layer includes sequence frames A1, A2, A3, ..., An; the second video layer includes sequence frames B1, B2, B3, ..., Bn; and the third video layer includes C1, C2, C3, ..., Cn. The second and third video layers belong to the foreground video layer and are respectively the middle and upper layers. In this example, the frames of each video layer at time t can be obtained, namely frames A1, B1, and C1, and then frames A1, B1, and C1 are sequentially converted from the first format to the second format. After processing the video frames of each layer at time t, the frames of each video layer at the next time t' (i.e., frames A2, B2, and C2) can be obtained, and then frames A2, B2, and C2 are sequentially converted from the first format to the second format, thereby realizing pipelined processing and reducing the memory usage of the processing.

[0062] In other embodiments, the current image to be synthesized and the previous image to be synthesized can be merged according to the following equation:

[0063] α0=α a +α b (1-α a Equation 1

[0064]

[0065] Among them, C a C represents the pixels of the color channel of the current image to be synthesized. b For the pixels of the color channel of the previous synthesized image, α a α is the pixel of the transparency channel of the current image to be synthesized. b α0 is the pixel of the transparency channel of the previous composite image, α0 is the pixel of the transparency channel of the current composite image, and C0 is the pixel of the color channel of the current composite image.

[0066] By using the above equation to merge images, the portability of the method for processing multiple image sequences provided according to embodiments of this disclosure across the toolchain can be improved. For example, the merging of multiple images can be achieved using the FFmpeg tool. FFmpeg is open-source free software that can perform recording, conversion, and streaming functions for various audio and video formats. When merging multiple images using the above equation, this can be implemented using the FFmpeg tool.

[0067] In other embodiments, the pixels of the color channels of multiple images are pre-multiplied pixels, and the current image to be synthesized and the previous image to be synthesized can be merged according to the following equation:

[0068] α0=α a +α b (1-α a Equation 3

[0069] c0 = c a +c b (1-α a Equation 4

[0070] Among them, c a c represents the pixels of the color channel of the current image to be synthesized. b For the pixels of the color channel of the previous synthesized image, α a α is the pixel of the transparency channel of the current image to be synthesized. b Let α0 be the pixel of the transparency channel of the previous synthesized image, c0 be the pixel of the transparency channel of the current synthesized image, and c0 be the pixel of the color channel of the current synthesized image. Premultiplication, also known as alpha premultiplication, refers to performing alpha pre-calculation on the pixels of the color channels of the image according to the following equation:

[0071] c i =α i C i Equation 5

[0072] Among them, C i For the pixels of the color channels of the image before premultiplication, α i c represents the number of pixels in the image's transparency channel. iLet i = 1, 2, ..., N (where N is the total number of pixels in the image's color channels after premultiplication). In image merging, if the pixels in an image's color channels are premultiplied pixels, merging multiple images according to Equations 1 and 2 (i.e., processing them as non-premultiplied pixels) will result in a darker merged image, affecting the processing quality. Distinguishing between premultiplied and non-premultiplied images can improve image quality.

[0073] In some embodiments, the plurality of image sequences may include an image sequence belonging to the background layer and an image sequence belonging to the foreground layer, and, Figure 2 The current image to be synthesized is the background layer, and the previous synthesized image is the foreground layer. Continue combining... Figure 5 To illustrate, in this example, the frames of each video layer at time t can be obtained, namely frames C1, B1, and A1. Then, frames C1, B1, and A1 are sequentially converted from the first format to the second format. After processing the video frames of each layer at time t, the frames of each video layer at the next time t' (i.e., frames C2, B2, and A2) can be obtained, and then frames C2, B2, and A2 are sequentially converted from the first format to the second format, thereby achieving pipelined processing and reducing the memory usage of the processing.

[0074] In some examples, Figure 2 The method further includes: converting the merged image into an output image including the first region and the second region based on the color channel and the transparency channel of the merged image.

[0075] In this example, the first region may contain the color information of the output image, and the second region may contain the transparency information of the output image. The output image is in the first format. In this example, the pixels of the color channels of the merged image can be mapped to the color information of the output image, and the transparency channels of the merged image can be mapped to the transparency information of the output image. For example, if the pixel value at a certain coordinate in the merged image is [0,0,0,0] (representing that the values ​​of the RGB channels are 0 and the value of the Alpha channel is 0), then the RGB value at that coordinate can be mapped to the corresponding pixel value [0,0,0] (representing that the values ​​of the RGB channels are 0) in the first region of the output image, and the Alpha value at that coordinate can be mapped to the corresponding pixel value [0,0,0] (representing that the values ​​of the RGB channels are 0) in the second region of the output image. In this example, the position parameters of the first and second regions in the output image can be the same as the position parameters of the first and second regions of the first format image in step S602. In this example, after obtaining the output image, it can be processed by line prediction, transformation, quantization, entropy coding, inverse quantization, inverse transformation, reconstruction, filtering, etc., to output an encoded stream. The method for processing multiple image sequences provided in this embodiment can obtain an output image with the same format as the input image by performing an inverse transformation on the merged image, thereby improving the automation level of the processing.

[0076] In some embodiments, converting the merged image into an output image including the first region and the second region based on the color channel and the transparency channel of the merged image may include:

[0077] Step S6062: Map the pixels of the color channel and the pixels of the transparency channel of the image obtained after merging to a first image containing color information and a second image containing transparency information, respectively.

[0078] In this example, both the first and second images are in the first format and are the same size. For instance, if the pixel value at (X=6, Y=6) in the merged image is [0,0,0,0] (representing that the values ​​of the RGB channels are 0 and the Alpha channel is 0), then the RGB value at that coordinate can be mapped to the pixel value [0,0,0] (representing that the values ​​of the RGB channels are 0) in the first image at (X=6, Y=6), and the Alpha value at that coordinate can be mapped to the pixel value [0,0,0] (representing that the values ​​of the RGB channels are 0) in the second image at (X=6, Y=6).

[0079] Step S6064: Based on the position parameters associated with the positions of the first region and the second region in the image, convert the first image and the second image into the first region and the second region of the output image, respectively.

[0080] In this example, the position parameters of the first and second regions in the output image can be the same as the position parameters associated with the positions of the first and second regions in the image (input image).

[0081] In other embodiments, step S6064 described above may include:

[0082] Step S60642: Based on the positional parameters associated with the positions of the first region and the second region in the image, resample the first image and / or the second image to obtain the first region image and the second region image.

[0083] In this example, the first region image may contain color information of the output image, and the second region image may contain transparency information of the output image. In some optional examples, the first image and / or the second image may be resampled according to an image scaling algorithm. In this example, the image scaling algorithm may include linear interpolation, quadratic cubic interpolation, Lanczos algorithm, new edge-guided interpolation, etc.

[0084] Step S60644: Draw the first region image and the second region image onto the same canvas to obtain the output image, wherein the first region image corresponds to the first region and the second region image corresponds to the second region.

[0085] Figure 6 An example of an output image according to an embodiment of this disclosure is shown. Figure 6 As shown, on the black canvas, the first region image 70 can be drawn onto the specified area, and the second region image 80 can be drawn onto the specified area.

[0086] In other embodiments, the first format in the above embodiments of this disclosure is a red-green-blue (RGB) pixel format, and the second format is a red-green-blue Alpha (RGBA) pixel format. In still other embodiments, the second region in the above embodiments of this disclosure is formed by mapping the image's transparency information to the luminance channel (Y) in the luminance-chrominance (YUV) pixel format. By mapping the image's transparency information to the luminance channel in YUV, space can be saved and data compression efficiency improved when compressing video frame data.

[0087] For the purpose of illustrating these embodiments, Figure 7 A flowchart illustrating a method for processing multiple image sequences according to these embodiments is shown. Figure 7As shown, the image can be input in the order of background first, then foreground. Then, based on the position parameters of the first and second regions in the image, the graphic data in each region is extracted to obtain RGB and Alpha regions. Next, the sizes of the RGB and Alpha regions are normalized through resampling. Then, the data of the Alpha region is mapped to the Alpha channel and merged with the RGB channel mapped from the RGB region data to obtain RGBA data. Next, it is determined whether the input image is a background layer. If it is, it is further determined whether it is the last layer. If it is not a background layer, it is merged with the previous RGBA data result (i.e., image merging in this disclosure). If the input image is the last layer, the merged RGBA data (i.e., the merged image in this disclosure) can be obtained; otherwise, the RGBA data of the current layer is output and the next layer of input image is processed (i.e., returning to the initial "input" step). Then, the merged RGBA data is input and its channels are separated to obtain the RGB channel image (i.e., the first image in this disclosure) and the Alpha channel image (i.e., the second image in this disclosure). Next, the alpha channel image is resampled, and the alpha channel data is mapped to black and white data. Then, it is drawn onto the canvas according to a specified format (which can be determined according to the position parameters in this disclosure). Finally, the resulting output image is encoded, and multiple images of the next frame are processed.

[0088] according to Figure 7 The example shown demonstrates a method for processing multiple image sequences that enables pipelined processing, eliminates the need for local storage of intermediate files, has low memory consumption during processing, and is easily portable to existing toolchains.

[0089] The method for processing multiple image sequences provided in this disclosure has the advantage of strong portability, and the method can be implemented in the FFmpeg toolkit.

[0090] Figure 8 An exemplary block diagram of an apparatus for processing multiple image sequences according to embodiments of the present disclosure is shown. Figure 8As shown, the device 800 includes a processing module 801 configured to perform the following merging process on multiple images with the same sequence number in the plurality of image sequences: decoding the images in the current image sequence to obtain a current image to be synthesized; merging the current image to be synthesized with a previous synthesized image to obtain a current synthesized image, wherein the first synthesized image is obtained by the following method: decoding the images in the first image sequence to obtain a first image to be synthesized; decoding the images in the second image sequence to obtain a second image to be synthesized; merging the second image to be synthesized with the first image to be synthesized to obtain a first synthesized image.

[0091] The apparatus for processing multiple images provided according to embodiments of this disclosure can synthesize multiple video layers into multiple different videos by using a user-defined synthesis order, thereby synthesizing different gift effects for each user.

[0092] The operations, features, and advantages described above for method 200 also apply to device 800 and its included modules. For the sake of brevity, some operations, features, and advantages will not be repeated here.

[0093] In some examples, the image includes a first region and a second region, the first region containing color information of the image and the second region containing transparency information of the image, the image being in a first format with color channels, and the processing module 801 further includes: a first conversion module configured to convert the image from the first format to a second format having the color channels and transparency channels based on the color information and transparency information of the image.

[0094] In some examples, the first conversion module includes: a segmentation module configured to segment the image into a first region image and a second region image based on position parameters associated with the positions of the first region and the second region in the image, wherein the first region image contains the color information of the image, the second region image contains the transparency information of the image, and both the first region image and the second region image are in the first format; a normalization module configured to normalize the size of the first region image and the second region image; and a first mapping module configured to map the color information contained in the normalized first region image and the transparency information contained in the normalized second region image to pixels of the color channel and the transparency channel of the image to be synthesized, respectively.

[0095] In some examples, the plurality of image sequences includes an image sequence belonging to the background layer and an image sequence belonging to the foreground layer, and the current image to be synthesized is the foreground layer, and the previous synthesized image is the background layer.

[0096] In some examples, the plurality of image sequences includes an image sequence belonging to a background layer and an image sequence belonging to a foreground layer, and the current image to be synthesized is the background layer, and the previous synthesized image is the foreground layer.

[0097] In some examples, merging the current image to be synthesized and the previous synthesized image to obtain the current synthesized image includes merging the current image to be synthesized and the previous synthesized image according to the following equation:

[0098]

[0099] Among them, C a C represents the pixels of the color channel of the current image to be synthesized. b For the pixels of the color channel of the previous synthesized image, α a α is the pixel of the transparency channel of the current image to be synthesized. b α0 is the pixel of the transparency channel of the previous composite image, C0 is the pixel of the transparency channel of the current composite image, and C0 is the pixel of the color channel of the current composite image.

[0100] In some examples, the pixels of the color channels of the plurality of images are pre-multiplied pixels, and merging the current image to be synthesized and the previous synthesized image to obtain the current synthesized image includes:

[0101] The current image to be synthesized and the previous image to be synthesized are merged according to the following equation:

[0102] α0=α a +α b (1-α a ), c0 = c a +c b (1-α a )

[0103] Among them, c a c represents the pixels of the color channel of the current image to be synthesized. b For the pixels of the color channel of the previous synthesized image, α a α is the pixel of the transparency channel of the current image to be synthesized. bα0 is the pixel of the transparency channel of the previous synthesized image, c0 is the pixel of the transparency channel of the current synthesized image, and c0 is the pixel of the color channel of the current synthesized image.

[0104] In some examples, the normalization module is further configured to normalize the sizes of the first region image and the second region image according to an image scaling algorithm, wherein the image scaling algorithm includes at least one of the following: a linear interpolation algorithm; a quadratic cubic interpolation algorithm; a Lanczos algorithm; and a new edge-guided interpolation algorithm.

[0105] In some examples, the apparatus 900 further includes a second conversion module configured to convert the merged image into an output image including the first region and the second region based on the color channel and the transparency channel of the merged image, wherein the first region contains color information of the output image, the second region contains transparency information of the output image, and the output image is in the first format.

[0106] In some examples, the second conversion module includes: a second mapping module configured to map pixels of the color channel and pixels of the alpha channel of the merged image to a first image containing color information and a second image containing alpha information, respectively, wherein the first image and the second image are both in the first format; and a third conversion module configured to convert the first image and the second image into the first region and the second region of the output image, respectively, based on position parameters associated with the positions of the first region and the second region in the image.

[0107] In some examples, the third conversion module includes: a resampling module configured to resample the first image and / or the second image based on position parameters associated with the positions of the first region and the second region in the image, to obtain a first region image and a second region image, wherein the first region image contains the color information of the output image and the second region image contains the transparency information of the output image; and a drawing module configured to draw the first region image and the second region image onto the same canvas to obtain the output image, wherein the first region image corresponds to the first region and the second region image corresponds to the second region.

[0108] In some examples, the first format is a red-green-blue (RGB) pixel format, and the second format is a red-green-blue Alpha (RGBA) pixel format.

[0109] In some examples, the second region is formed by mapping the transparency information of the image to a target channel in the luminance-chrominance (YUV) pixel format.

[0110] In some examples, the target channel is the luminance channel.

[0111] According to another aspect of this disclosure, a computer program product is provided, including program code instructions that, when executed by a computer, cause the computer to perform the method described above.

[0112] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method according to the above description.

[0113] According to another aspect of this disclosure, an electronic device is provided, comprising: a processor; a memory in electronic communication with the processor; and instructions stored in the memory and executable by the processor to cause the electronic device to perform the method described above.

[0114] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. See also Figure 9 The present invention describes a structural block diagram of an electronic device 90 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein. Figure 9As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 may also store various programs and data required for the operation of the device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904. Multiple components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, optical disk, etc.; and a communication unit 909, such as a network card, modem, wireless transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0115] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as methods for processing multiple image sequences. For example, in some embodiments, the method for processing multiple image sequences may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the methods for processing multiple image sequences described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured in any other suitable manner (e.g., by means of firmware) to perform a method of processing multiple image sequences.

[0116] The various illustrative logics, logic blocks, modules, circuits, and algorithmic processes described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. The interchangeability of hardware and software has been generally described in terms of functionality and illustrated in the aforementioned illustrative components, blocks, modules, circuits, and processes. Whether this functionality is implemented in hardware or software depends on the specific application and the design constraints on the overall system.

[0117] Hardware and data processing apparatuses for implementing the various illustrative logics, logic blocks, modules, and circuits described in conjunction with the aspects disclosed herein may be implemented or performed by general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor or any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. In some aspects, specific processes and methods may be performed by circuits specific to a given function.

[0118] In one or more aspects, the described functionality can be implemented in hardware, digital electronic circuits, computer software, firmware (including the structures disclosed in this specification and their equivalents) or any combination thereof. The aspects of the subject matter described in this specification can also be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing apparatus or for controlling the operation of a data processing apparatus.

[0119] If implemented in software, the functionality can be stored or transferred as one or more instructions or code onto a computer-readable medium. The processes of the methods or algorithms disclosed herein can be implemented in a processor-executable software module that may reside on a computer-readable medium. Computer-readable media include computer storage media and communication media, including any medium capable of transferring a computer program from one place to another. Storage media can be any available medium accessible to a computer. By way of example and not limitation, this computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to store the required program code in the form of instructions or data structures and is accessible to a computer. Furthermore, any connection can be properly referred to as a computer-readable medium. The disks and discs used herein include high-density optical discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically copy data, while discs optically copy data using lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operation of a method or algorithm may be one or any combination or set of code and instructions on a machine-readable and computer-readable medium, which may be incorporated into a computer program product.

[0120] The various embodiments in this disclosure are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments are basically similar to the method embodiments, so the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the method embodiments.

Claims

1. A method for processing multiple image sequences, comprising: Multiple images with the same sequence number from the multiple image sequences are merged as follows: The images in the current image sequence are decoded to obtain the current image to be synthesized; The current image to be synthesized is merged with the previous synthesized image to obtain the current synthesized image. The first synthesized image was obtained using the following method: Decode the images in the first image sequence to obtain the first image to be synthesized; Decode the images in the second image sequence to obtain the second image to be synthesized; The second image to be synthesized and the first image to be synthesized are merged to obtain the first synthesized image; Decoding the images in the current image sequence to obtain the current image to be synthesized includes: converting the image from a first format to a second format with color channels and transparency channels based on the color information and transparency information of the image.

2. The method according to claim 1, wherein, The image includes a first region and a second region. The first region contains the color information of the image, and the second region contains the transparency information of the image. The image is in a first format with color channels.

3. The method according to claim 2, wherein, The process of converting the image from the first format to a second format having the color channel and the transparency channel based on the color information and the transparency information of the image includes: Based on positional parameters associated with the positions of the first and second regions in the image, the image is segmented into a first region image and a second region image, wherein, The first region image contains the color information of the image, and the second region image contains the transparency information of the image. Both the first region image and the second region image are in the first format. Normalize the sizes of the first region image and the second region image; and The color information contained in the normalized first region image and the transparency information contained in the normalized second region image are respectively mapped to the pixels of the color channel and the pixels of the transparency channel of the image to be synthesized.

4. The method according to claim 2 or 3, wherein, The plurality of image sequences includes image sequences belonging to the background layer and image sequences belonging to the foreground layer, and, The current image to be synthesized is the foreground layer, and the previous synthesized image is the background layer.

5. The method according to claim 1 or 3, wherein, The plurality of image sequences includes image sequences belonging to the background layer and image sequences belonging to the foreground layer, and, The current image to be synthesized is the background layer, and the previous synthesized image is the foreground layer.

6. The method according to claim 4, wherein, The step of merging the current image to be synthesized and the previous synthesized image to obtain the current synthesized image includes: The current image to be synthesized and the previous image to be synthesized are merged according to the following equation: , , in, C a The pixels of the color channel of the current image to be synthesized. C b The pixels of the color channel of the previous synthesized image. α a The pixels of the transparency channel of the current image to be synthesized. α b The pixel of the transparency channel of the previous synthesized image. α 0 represents the pixel value of the transparency channel in the currently synthesized image. C 0 represents the pixel of the color channel of the currently synthesized image.

7. The method according to claim 4, wherein, The pixels of the color channels of the plurality of images are pre-multiplied pixels, and, The current image to be synthesized and the previous synthesized image are merged to obtain the current synthesized image, which includes: The current image to be synthesized and the previous image to be synthesized are merged according to the following equation: α0=α a +a b (1-a a ),c0=c a +c b (1-a a ) Among them, c a c represents the pixels of the color channel of the current image to be synthesized. b For the pixels of the color channel of the previous synthesized image, α a α is the pixel of the transparency channel of the current image to be synthesized. b α0 is the pixel of the transparency channel of the previous synthesized image, c0 is the pixel of the transparency channel of the current synthesized image, and c0 is the pixel of the color channel of the current synthesized image.

8. The method according to claim 3, wherein, The normalization of the sizes of the first region image and the second region image includes: The sizes of the first region image and the second region image are normalized according to an image scaling algorithm, wherein the image scaling algorithm includes at least one of the following: Linear interpolation algorithm; Quadratic cubic interpolation algorithm; Lanczos algorithm; and A new edge-guided interpolation algorithm.

9. The method according to claim 2, further comprising: Based on the color channels and transparency channels of the merged image, the merged image is converted into an output image including the first region and the second region. The first region contains the color information of the output image, the second region contains the transparency information of the output image, and the output image is in the first format.

10. The method according to claim 9, wherein, The step of converting the merged image into an output image including the first region and the second region based on the color channel and the transparency channel of the merged image includes: The pixels of the color channel and the pixels of the alpha channel of the merged image are respectively mapped to a first image containing color information and a second image containing alpha information, wherein both the first image and the second image are in the first format; and Based on positional parameters associated with the positions of the first region and the second region in the image, the first image and the second image are respectively converted into the first region and the second region of the output image.

11. The method according to claim 10, wherein, The step of converting the first image and the second image into the first and second regions of the output image based on position parameters associated with the positions of the first region and the second region in the image includes: Based on positional parameters associated with the positions of the first and second regions in the image, the first and / or second images are resampled to obtain a first region image and a second region image. Wherein, the first region image contains the color information of the output image, and the second region image contains the transparency information of the output image; and The first region image and the second region image are drawn onto the same canvas to obtain the output image, wherein the first region image corresponds to the first region and the second region image corresponds to the second region.

12. The method according to claim 2, wherein, The first format is a red-green-blue (RGB) pixel format, and the second format is a red-green-blue Alpha (RGBA) pixel format.

13. The method according to claim 2, wherein, The second region is formed by mapping the transparency information of the image to a target channel in the luminance-chrominance (YUV) pixel format.

14. The method according to claim 13, wherein, The target channel is the luminance channel.

15. An apparatus for processing multiple image sequences, comprising: The processing module is configured to perform the following merging process on multiple images with the same sequence number in the multiple image sequences: The images in the current image sequence are decoded to obtain the current image to be synthesized; The current image to be synthesized is merged with the previous synthesized image to obtain the current synthesized image. The first synthesized image was obtained using the following method: Decode the images in the first image sequence to obtain the first image to be synthesized; Decode the images in the second image sequence to obtain the second image to be synthesized; The second image to be synthesized and the first image to be synthesized are merged to obtain the first synthesized image; Decoding the images in the current image sequence to obtain the current image to be synthesized includes: converting the image from a first format to a second format with color channels and transparency channels based on the color information and transparency information of the image.

16. A computer program product comprising program code instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 14.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 14.

18. An electronic device comprising: processor, A memory that communicates electronically with the processor; as well as Instructions, which are stored in the memory and can be executed by the processor, to cause the electronic device to perform the method according to any one of claims 1 to 14.