Image processing methods, apparatus, devices and storage media

By using hardware decoding modules, such as DSP chips, to decode and render image frames in live streaming, the rendering latency issues of virtual avatar transformation and virtual gift display have been resolved, resulting in faster decoding speeds and smoother display effects.

CN116781967BActive Publication Date: 2025-11-14BEIJING LIUJIANFANG TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310722493.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-16
Publication Date
2025-11-14
Estimated Expiration
2043-06-16

AI Technical Summary

Technical Problem

How to effectively reduce the rendering latency of virtual avatar transformation and virtual gift display in live streaming to improve user experience.

Method used

A hardware decoding module is used to decode and render the target image frame. The hardware decoding module, such as a DSP chip, is used to quickly decode and render the image frame, thereby improving the decoding speed and reducing the rendering latency.

Benefits of technology

The use of a hardware decoding module significantly improved image decoding speed, reduced rendering latency, enhanced the smoothness of virtual avatar customization and virtual gift display, and improved the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116781967B_ABST
    Figure CN116781967B_ABST
Patent Text Reader

Abstract

This invention provides an image processing method, apparatus, device, and storage medium. The method includes: a terminal device, in response to a user-triggered selection operation, determining a target image frame in a video file. The video file includes at least one image frame obtained by video encoding at least one original image, and the target image frame is any one of the at least one image frame. Each original image contains a virtual object displayed in a live streaming interface. Then, the terminal device decodes the target image frame using a hardware decoding module. Finally, the terminal device renders the decoding result. As can be seen, the terminal device uses a hardware decoding module to decode the target image frame in the video file, thus improving the image decoding speed and further reducing image rendering latency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image rendering technology, and in particular to an image processing method, apparatus, device, and storage medium. Background Technology

[0002] With the rapid development of the online live streaming industry, live streaming has become a daily form of entertainment for many users. Users can be either viewers or streamers. For example, viewers can send virtual gifts to streamers in the live stream room, and both streamers and viewers can change their virtual avatars to better facilitate interaction between them.

[0003] During the process of virtual avatars changing outfits, the virtual objects such as clothing or facial features are generally displayed as images; while when viewers send virtual gifts to streamers, the virtual gifts are generally displayed as animations. Therefore, how to effectively reduce rendering latency so that users can see the virtual avatar's transformation or virtual gifts in real time has become an urgent problem to be solved. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide an image processing method, apparatus, device, and storage medium to improve the speed of image decoding and reduce rendering latency.

[0005] In a first aspect, embodiments of the present invention provide an image processing method, comprising:

[0006] In response to a user-triggered selection operation, a target image frame is determined in a video file, the video file including at least one image frame obtained by video encoding at least one original image, the target image frame being any one of the at least one image frames, and any original image containing a virtual object displayed in the live streaming interface;

[0007] The target image frame is decoded using a hardware decoding module;

[0008] Rendering and decoding results.

[0009] In a second aspect, embodiments of the present invention provide an image processing apparatus, comprising:

[0010] The determination module is used to determine a target image frame in a video file in response to a selection operation triggered by a user. The video file includes at least one image frame obtained by video encoding at least one original image. The target image frame is any one of the at least one image frames. Any original image contains a virtual object displayed in the live broadcast interface.

[0011] A decoding module is used to decode the target image frame using a hardware decoding module;

[0012] The rendering module is used to render the decoded results.

[0013] Thirdly, embodiments of the present invention provide an electronic device including a processor and a memory, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the image processing method described in the first aspect. The electronic device may also include a communication interface for communicating with other devices or communication networks.

[0014] Fourthly, embodiments of the present invention provide a non-transitory machine-readable storage medium storing executable code, which, when executed by a processor of an electronic device, enables the processor to at least implement the image processing method as described in the first aspect.

[0015] The image processing method provided in this invention involves a terminal device determining a target image frame in a video file in response to a user-triggered selection operation. The video file includes at least one image frame obtained by video encoding at least one original image, and the target image frame is any one of the at least one image frame. Each original image contains a virtual object displayed on the live streaming interface. The terminal device then decodes the target image frame using a hardware decoding module. Finally, the terminal device renders the decoding result. As can be seen, the terminal device uses a hardware decoding module to decode the target image frame in the video file, thus improving the image decoding speed and further reducing image rendering latency. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A flowchart of an image processing method provided in an embodiment of the present invention;

[0018] Figure 2 A schematic diagram of the structure of a video file provided in an embodiment of the present invention;

[0019] Figure 3 A flowchart of another image processing method provided in an embodiment of the present invention;

[0020] Figure 4 This is a structural diagram of a live streaming scenario provided by an embodiment of the present invention;

[0021] Figure 5 A flowchart illustrating yet another image processing method provided in an embodiment of the present invention;

[0022] Figure 6 This is a schematic diagram of the structure of a target image frame provided in an embodiment of the present invention;

[0023] Figure 7 This is a schematic diagram of another live streaming scenario provided by an embodiment of the present invention;

[0024] Figure 8 This is a schematic diagram of the structure of an image processing device provided in an embodiment of the present invention;

[0025] Figure 9 To and Figure 8 The illustrated embodiment provides a schematic diagram of the electronic device corresponding to the image processing apparatus. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” used in the embodiments of this invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. “Multiple” generally includes at least two, but does not exclude the inclusion of at least one.

[0028] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0029] Depending on the context, the words “if” or “suppose” as used here can be interpreted as “when” or “in response to determination” or “in response to identification.” Similarly, depending on the context, the phrases “if determination” or “if identification (of the condition or event of the statement)” can be interpreted as “when determination” or “in response to determination” or “when identification (of the condition or event of the statement)” or “in response to identification (of the condition or event of the statement).”

[0030] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a product or system comprising a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a product or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the product or system that includes said element.

[0031] Before describing the image processing methods provided in the following embodiments of the present invention in detail, the relevant concepts involved in the following embodiments can also be explained:

[0032] Keyframes: These frames can be decoded independently of other frames. They contain all the information of the video and help the decoder quickly locate a specific frame in the video, thereby improving the video encoding efficiency.

[0033] Hardware decoding: This involves using hardware, specifically a chip, to decode data. This contrasts with software decoding, which uses the CPU to decode data. The advantages of hardware decoding are faster decoding speed and lower power consumption.

[0034] H.264 is a video coding standard jointly proposed by the International Organization for Standardization (ISO) and the International Telecommunication Union (ITU).

[0035] H.265 is a new video coding standard developed after H.264.

[0036] Based on the above description, some embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Where there is no conflict between the embodiments, the following embodiments and features can be combined with each other. Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.

[0037] Figure 1 This is a flowchart illustrating an image processing method provided in an embodiment of the present invention. The image processing method provided in this embodiment can be executed by any terminal device with a live streaming application installed. The terminal device can be a smartphone, tablet, laptop, etc. Figure 1 As shown, the method includes the following steps:

[0038] S101, in response to a user-triggered selection operation, a target image frame is determined in a video file, the video file including at least one image frame obtained by video encoding at least one original image, the target image frame being any one of the at least one image frame, and any original image containing a virtual object displayed in the live broadcast interface.

[0039] In response to a selection operation triggered by the user's interface on the live streaming platform, the terminal device may optionally determine a first identifier of the virtual object selected by the user. Then, it determines an image frame in the video file that has a second identifier matching this first identifier; this image frame is the target image frame. The virtual object contained in the target image frame is the same as the virtual object selected by the user on the interface. Essentially, the image frame is a rendered texture obtained after texture mapping of the virtual object.

[0040] When a user wants to customize their virtual avatar, the virtual object they choose can be the image material used to create that avatar; the essence of image material is texture material. In this case, the user can be either a streamer or a viewer. The image material for the virtual avatar can include at least one of facial features, hairstyle, and clothing. When a user wants to send a gift, the virtual object they choose can also be a virtual gift sent to the streamer. In this case, the user is a viewer.

[0041] Optionally, the video file can be pre-stored locally on the terminal device. The video file may include at least one image frame obtained by video encoding at least one original image. Any original image may contain a virtual object displayed on the live streaming interface. Optionally, the virtual object may be a texture rendered from a 3D model or the result of rendering a 2D sequence of frame animations. Optionally, the original image may be a Portable Network Graphic (PNG). Specifically, the terminal device may encode each original image according to a preset encoding standard to obtain an encoded image frame. The preset encoding standard may be H.264 or H.265, etc. Then, the encoded image frames from each original image are stitched together to obtain the video file.

[0042] In practice, virtual avatars are typically represented as 3D models. For each type of image material that constitutes a virtual avatar, the live streaming platform can provide at least one; for example, the platform might offer 40 eyebrow shapes and 20 hairstyles. Each image material has a corresponding video file. When any image material is updated, the corresponding video file can be directly modified, facilitating image material management. In practice, multiple virtual gifts provided by the live streaming platform can collectively form a single video file.

[0043] For example, suppose an original image contains a virtual object, such as eyebrows, which is an image element needed to construct a user's virtual avatar. If a live streaming platform offers 40 different eyebrow shapes for users to choose from, the terminal device can pre-encode these 40 original images containing different eyebrow shapes to obtain 40 image frames, and further, a video file containing these 40 image frames. This video file composed of 40 eyebrow shapes can also be combined with... Figure 2 understand.

[0044] After a user selects their desired eyebrow shape (shape 1) on the live streaming platform's interface, the terminal device responds to this selection by obtaining a first identifier corresponding to eyebrow shape 1. Then, the terminal device determines the corresponding video file based on the first identifier, and further identifies the target image frame within that video file that contains a second identifier matching the first identifier. This target image frame contains the eyebrow shape 1 selected by the user.

[0045] In practice, virtual gifts are typically represented as 2D sequence frame animations. Similarly, suppose a live streaming platform offers three virtual gifts for users to choose from. Optionally, virtual gift 1 could be a starry sky, virtual gift 2 could be a paper crane, and virtual gift 3 could be a dandelion. The terminal device can then encode these three original images to obtain three image frames, and further, generate a video file containing these three image frames.

[0046] When a user selects the desired virtual gift 1, namely "Starry Sky," on the interface provided by the live streaming platform, the terminal device responds to the user's selection of virtual gift 1 by obtaining a first identifier corresponding to virtual gift 1. Then, the terminal device can determine the target image frame, which contains a second identifier matching the first identifier, within the video file corresponding to the virtual gift. This target image frame contains the virtual gift 1 selected by the user.

[0047] S102 uses a hardware decoding module to decode the target image frame.

[0048] S103, rendering the decoding result.

[0049] Next, the terminal device can immediately decode the target image frame using its built-in hardware decoding module, and the hardware decoding module uses the vertex data and texture information contained in the decoding result to complete the rendering of the decoded result. The hardware decoding module can be a dedicated microprocessor, such as a Digital Signal Processing (DSP) chip.

[0050] Optionally, in order to improve the decoding speed of the target video frame, the terminal device can also set each image frame in the video file as a keyframe so that each image frame can be decoded independently.

[0051] In this embodiment, the terminal device, in response to a user-triggered selection operation, determines a target image frame in the video file. The video file includes at least one image frame obtained by video encoding at least one original image, and the target image frame is any one of the at least one image frame. Each original image contains a virtual object displayed on the live streaming interface. The terminal device then decodes the target image frame using a hardware decoding module and renders the decoding result.

[0052] Furthermore, in the traditional process of decoding multiple raw images using a CPU, the CPU incurs performance loss by decoding each raw image, i.e., increasing latency. However, in this embodiment of the invention, a hardware decoding module, such as a DSP chip, can quickly decode the target image frames in the video file, resulting in minimal latency during the decoding process. Moreover, the decoding result can be directly rendered by the hardware decoding module, meaning the decoding result is transmitted within the same hardware, thereby reducing image rendering latency.

[0053] Figure 1 As mentioned in the embodiments, any virtual object in the original image may include image materials needed to form the virtual avatar of the anchor or viewer, or it may include virtual gifts. Different virtual objects correspond to different display methods; for example, image materials needed for a virtual avatar are usually displayed as images, while virtual gifts are usually displayed as animations. Therefore, different image processing methods can be used for different virtual objects to improve their display effect.

[0054] When the virtual object is an image resource Figure 3 A flowchart of another image processing method provided in an embodiment of the present invention is shown below. Figure 3 As shown, the method may include the following steps:

[0055] S201, in response to the user's selection operation of the target image material, determine the target index corresponding to the target image material in the header data of the video file.

[0056] S202, in the video file, the image frame pointed to by the target index is determined as the target image frame.

[0057] The terminal device can respond to a user's selection of a target image material by determining the target index corresponding to the target image material in the header data of the video file. Optionally, there can be a preset correspondence between the material identifier and the index. When the user selects a target image material, the terminal device can first determine the video file corresponding to the type of the target image material based on its first identifier. Then, based on the aforementioned preset correspondence, it can directly determine the target index corresponding to the first identifier of the target image material in the header data of the video file. The image frame pointed to by this target index is the target image frame. It is easy to understand that the target image frame is any one of at least one image frame included in the video file.

[0058] To generate a video file, the terminal device can encode each original image according to a preset encoding standard to obtain an encoded image frame. Then, the encoded image frames from each original image are stitched together to obtain the video file. In this embodiment, the virtual objects contained in the original images can be any type of image material required to construct the virtual avatar of the broadcaster or viewer.

[0059] Optionally, the generated video file may specifically include two parts: header data and image frames. The header data of the video file may include a second identifier for each image frame in the video file, and may also include an index for each image frame. The second identifier reflects the content of the image material contained in the image frame; for example, the second identifier may indicate that the image frame contains "eyebrow 1" or "hair 5," etc. The index reflects the position of the image frame within the entire video file.

[0060] S203 uses a hardware decoding module to decode the target image frame.

[0061] S204, rendering the decoding result.

[0062] For the specific implementation process of steps S203 to 204 above, please refer to [link / reference]. Figure 1 The specific descriptions of the relevant steps in the illustrated embodiments will not be repeated here.

[0063] In this embodiment, in response to the user's selection of the target image material, the terminal device first determines the target index corresponding to the target image material in the header data of the video file. Then, the terminal device identifies the image frame pointed to by the target index as the target image frame in the video file. Next, the terminal device decodes the target image frame using a hardware decoding module and renders the decoding result.

[0064] In the above method, on the one hand, the terminal device can quickly locate the target image frame pointed to by the target index from the video file by using the preset correspondence between the material identifier and the index that it has stored in advance. On the other hand, the terminal device uses a hardware decoding module to decode the target image frame in the video file, thus improving the image decoding speed and further reducing the image rendering latency, making the user's virtual avatar dressing process, i.e., the image rendering process, smoother, thereby improving the user experience.

[0065] Furthermore, any content not described in detail in this embodiment, as well as the technical effects that can be achieved, can be found in the relevant descriptions of the above embodiments, and will not be repeated here.

[0066] For ease of understanding, such as Figure 4 As shown, it can be combined with the online live streaming scenario to... Figure 3 The specific implementation process of the image processing method provided in the illustrated embodiment is explained by way of example.

[0067] In a scenario where a user is conducting a live online broadcast, the user can be a broadcaster, and the terminal device can be a smartphone. When the broadcaster wants to change their virtual avatar, for example, to change their eyebrow shape, they can click on the alternative eyebrow shapes displayed on the smartphone's interface (i.e., the image material in the above embodiment) to select the target image material, i.e., eyebrow shape 1 from the alternative eyebrow shapes.

[0068] Then, the smartphone can respond to the broadcaster's selection of eyebrow shape 1, determine the first identifier of eyebrow shape 1, and then determine the video file corresponding to the eyebrow shape based on this first identifier. Then, based on the preset correspondence between the eyebrow shape identifier and the index, the target index matching the first identifier is determined from the header data of the video file corresponding to the eyebrow shape. The image frame pointed to by the target index is the target image frame containing eyebrow shape 1.

[0069] The video file corresponding to each eyebrow shape may include header data and 40 image frames corresponding to 40 different eyebrow shapes. The header data of the video file may include a second identifier for each of the 40 image frames, and may also include an index for each of the 40 image frames. The second identifier reflects the eyebrow shape contained in each of the 40 image frames, and the index reflects the position of each of the 40 image frames within the entire video file.

[0070] Ultimately, the smartphone can use its built-in hardware decoding module to decode the target image frames in the video file and render the decoding results, so that the virtual image of the anchor with eyebrow shape 1 can be displayed on the smartphone.

[0071] and Figure 3 Similarly, in the illustrated embodiment, when any virtual object in any original image includes a virtual gift, the following approach can be used. Figure 5 The method shown is used for image processing. Then... Figure 5 A flowchart of another image processing method provided in an embodiment of the present invention, the method may include the following steps:

[0072] S301, in response to the user's selection of the target virtual gift, determines the target index corresponding to the target virtual gift in the header data of the video file.

[0073] S302, in the video file, the image frame pointed to by the target index is determined as the target image frame.

[0074] The terminal device can respond to the user's selection of a target virtual gift, determine the video file corresponding to the virtual gift, and then determine the target index corresponding to the target virtual gift in the header data of the video file. Optionally, when the user selects a target virtual gift, the terminal device can directly determine the target index corresponding to the first identifier of the target virtual gift based on a preset correspondence between virtual gift identifiers and indexes. The image frame pointed to by this target index is the target image frame. It is easy to understand that the target image frame is any one of at least one image frame included in the video file.

[0075] To generate a video file, the terminal device can encode each original image according to a preset encoding standard to obtain an encoded image frame. Then, the encoded image frames from each original image are stitched together to obtain the video file. In this embodiment, the virtual object contained in the original image is a virtual gift.

[0076] In practice, virtual gifts are usually displayed in the form of animation. Any original image can contain multiple sub-images reflecting different animation states of the virtual gift. Rendering these sub-images sequentially allows the virtual gift to be displayed in animated form. The virtual gifts in these sub-images are typically irregularly shaped, such as starry skies, origami cranes, or dandelions.

[0077] Terminal devices can use texture packing technology to combine multiple sub-images with different image materials into a single original image. Similarly, the image frame obtained by video encoding the original image also contains multiple sub-image frames, which can reflect different animation states of the virtual gift.

[0078] S303, according to the order information and position information recorded in the target index, decode the sub-image frames in the target image frame. The order information reflects the rendering order of the sub-image frames, and the position information reflects the position of the sub-image frames in the target image frame.

[0079] S304, rendering the decoding result.

[0080] Next, the terminal device can decode the sub-image frames in the target image frame sequentially according to the order and position information recorded in the target index, using its built-in hardware decoding module. Finally, the terminal device renders the decoding results. The order information reflects the rendering order of the sub-image frames, and the position information reflects the position of the sub-image frames within the target image frame.

[0081] For example, assuming the virtual gift is a starry sky, the target image frame includes three sub-image frames reflecting different animation states of the starry sky. The structure of the target image frame, i.e., the starry sky, can be combined with... Figure 6 Understood. At this point, the target index corresponding to the target image frame can record the rendering order of the three sub-image frames and their respective positions within the target image frame. The terminal device can then decode the three sub-image frames sequentially using its built-in hardware decoding module according to the target index, and render the decoded results of the three sub-image frames according to their rendering order. Rendering in this order will produce a starry sky with dynamic effects.

[0082] In this embodiment, in response to the user's selection of a target virtual gift, the terminal device determines the target index corresponding to the target virtual gift in the header data of the video file. Then, the terminal device identifies the image frame pointed to by the target index as the target image frame in the video file. Next, the terminal device decodes the sub-image frames in the target image frame using its built-in hardware decoding module, according to the order and position information recorded in the target index. Finally, the terminal device renders the decoding result.

[0083] In the above method, on the one hand, the terminal device can quickly locate the target image frame pointed to by the target index from the video file by using the preset correspondence between the virtual gift identifier and the index stored in advance. On the other hand, the terminal device uses a hardware decoding module to decode the target image frame in the video file. Therefore, it can improve the image decoding speed and further reduce the image rendering latency, making the process of the user sending a virtual gift or switching the currently displayed virtual gift, i.e., the image rendering process, faster. In other words, it can make the animation effect displayed by the terminal device smoother, thereby improving the user experience.

[0084] Furthermore, any content not described in detail in this embodiment, as well as the technical effects that can be achieved, can be found in the relevant descriptions of the above embodiments, and will not be repeated here.

[0085] For ease of understanding, such as Figure 7 As shown, it can be combined with the online live streaming scenario to Figure 5 The specific implementation process of the image processing method provided in the illustrated embodiment is explained by way of example.

[0086] In a scenario where users are live-streaming online, the user can be a viewer, and the terminal device can be a smartphone. When a viewer wants to send a virtual gift to the streamer, they can select the target virtual gift (virtual gift 1) from the list of available virtual gifts displayed on their smartphone's interface (as described in the above embodiment). The available virtual gifts can include three types: virtual gift 1 can be a starry sky, virtual gift 2 can be a paper crane, and virtual gift 3 can be a dandelion.

[0087] Then, in response to the viewer's selection of virtual gift 1, the smartphone can determine the first identifier of the virtual gift, then determine the corresponding video file based on the first identifier, and then determine the target index matching the first identifier from the header data of the video file corresponding to the virtual gift based on the preset correspondence between the virtual gift identifier and the target index. The image frame pointed to by the target index is the target image frame containing virtual gift 1.

[0088] The video file corresponding to the virtual gift may include header data and three image frames corresponding to the three types of virtual gifts. The header data of the video file may include a second identifier for each of the three image frames, and may also include an index for each of the three image frames. The second identifier is used to reflect the type of virtual gift contained in each of the three image frames, and the index is used to reflect the position of each of the three image frames in the entire video file.

[0089] Following the above Figure 5 In the example shown, the target image frame is assumed to include three sub-image frames reflecting different animation states of the starry sky. In this case, the target index corresponding to the target image frame can record the rendering order of the three sub-image frames and their respective positions within the target image frame. The smartphone can then decode the three sub-image frames sequentially using its built-in hardware decoding module according to the target index, and render the corresponding decoding results according to the rendering order of the three sub-image frames. The smartphone can then display virtual gift 1, i.e., the starry sky.

[0090] The following will describe in detail one or more embodiments of the device collaboration apparatus of the present invention. Those skilled in the art will understand that these device collaboration apparatuses can all be configured using commercially available hardware components through the steps taught in this solution.

[0091] Figure 8 A schematic diagram of the structure of an image processing device provided in an embodiment of the present invention is shown below. Figure 8 As shown, the device includes:

[0092] The determination module 11 is used to determine a target image frame in a video file in response to a selection operation triggered by a user. The video file includes at least one image frame obtained by video encoding at least one original image. The target image frame is any one of the at least one image frames. Any original image contains a virtual object displayed in the live broadcast interface.

[0093] Decoding module 12 is used to decode the target image frame with the aid of a hardware decoding module.

[0094] Rendering module 13 is used to render the decoding results.

[0095] The virtual objects in any original image include image materials necessary to construct the virtual avatar of the broadcaster or viewer. The virtual objects in any original image also include virtual gifts, and any image frame obtained by video encoding the original image includes multiple sub-image frames reflecting different animation states of the virtual gifts.

[0096] Optionally, the determining module 11 is configured to, in response to the user's selection operation of the target image material, determine the target index corresponding to the target image material in the header data of the video file; and determine the image frame pointed to by the target index as the target image frame in the video file.

[0097] Optionally, the determining module 11 is used to determine the target index corresponding to the target image material based on a preset correspondence between the material identifier and the index.

[0098] Optionally, the determining module 11 is configured to, in response to a user's selection of a target virtual gift, determine the target index corresponding to the target virtual gift in the header data of the video file; and determine the image frame pointed to by the target index as the target image frame in the video file.

[0099] The decoding module 12 is used to decode the sub-image frames in the target image frame according to the order information and position information recorded in the target index. The order information reflects the rendering order of the sub-image frames, and the position information reflects the position of the sub-image frames in the target image frame.

[0100] Optionally, the device further includes: an original image generation module 14, used to generate any one of the original images, wherein the any one of the original images includes multiple sub-images reflecting different animation states of the virtual gift.

[0101] The video file generation module 15 is used to generate the video file based on multiple original images.

[0102] Figure 8 The device shown can perform Figures 1 to 7 For the methods shown in the embodiments, the parts not described in detail in this embodiment can be referred to the following: Figures 1 to 7 The relevant descriptions of the illustrated embodiments are provided below. For the execution process and technical effects of this technical solution, please refer to [link / reference]. Figures 1 to 7 The descriptions in the illustrated embodiments will not be repeated here.

[0103] In one possible design, the image processing device can be structured as an electronic device, such as... Figure 9 As shown, the electronic device may include a processor 21 and a memory 22. The memory 22 is used to store data supporting the electronic device in performing the above-described actions. Figures 1 to 7 The image processing method program provided in the illustrated embodiment is configured by the processor 21 to execute the program stored in the memory 22.

[0104] The program includes one or more computer instructions, wherein when the one or more computer instructions are executed by the processor 21, they can perform the following steps:

[0105] In response to a user-triggered selection operation, a target image frame is determined in a video file, the video file including at least one image frame obtained by video encoding at least one original image, the target image frame being any one of the at least one image frames, and any original image containing a virtual object displayed in the live streaming interface;

[0106] The target image frame is decoded using a hardware decoding module;

[0107] Rendering and decoding results.

[0108] Optionally, the processor 21 is further configured to perform the aforementioned Figures 1 to 7 All or part of the steps in the illustrated embodiments.

[0109] The structure of the electronic device may also include a communication interface 23 for the electronic device to communicate with other devices or communication networks.

[0110] Furthermore, embodiments of the present invention provide a non-transitory machine-readable storage medium for storing computer software instructions used in the aforementioned electronic device, which includes instructions for executing the above-mentioned... Figures 1 to 7 The procedure involved in the image processing method in the illustrated embodiment.

[0111] This invention also provides a computer program product, which includes computer program instructions that are read and executed by a processor to perform the above-described... Figures 1 to 7 The image processing method in the illustrated embodiment.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image processing method, characterized in that, include: In response to a user-triggered selection operation, a first identifier of the virtual object selected by the user is determined, and a video file corresponding to the virtual object is determined based on the first identifier. In the video file, an image frame with a second identifier that matches the first identifier is determined as the target image frame. The video file includes at least one image frame obtained by video encoding at least one original image. The at least one image frame is a keyframe. The target image frame is any one of the at least one image frame. Any original image contains the virtual object displayed in the live broadcast interface. The target image frame is decoded using a hardware decoding module; The hardware decoding module is a digital signal processing (DSP) chip. The hardware decoding module utilizes the vertex data and texture information contained in the decoding result to complete the rendering of the decoding result.

2. The method according to claim 1, characterized in that, The virtual objects in any of the original images include the image materials required to constitute the virtual avatar of the anchor or user.

3. The method according to claim 2, characterized in that, The step of responding to a user-triggered selection operation by determining the video file corresponding to the virtual object based on the first identifier, and determining the target image frame within the video file, includes: In response to the user's selection of a target image material, the video file corresponding to the target image material is determined according to the first identifier of the target image material, and the target index corresponding to the target image material is determined in the header data of the video file; The image frame pointed to by the target index in the video file is determined as the target image frame.

4. The method according to claim 3, characterized in that, Determining the target index corresponding to the target image material in the header data of the video file includes: Based on the preset correspondence between the material identifier and the index, the target index corresponding to the first identifier of the target image material is determined in the header data of the video file.

5. The method according to claim 1, characterized in that, The virtual object in any original image includes a virtual gift, and any image frame obtained by video encoding the original image includes multiple sub-image frames reflecting different animation states of the virtual gift.

6. The method according to claim 5, characterized in that, The step of determining the target image frame in the video file in response to a user-triggered selection operation includes: In response to the user's selection of a target virtual gift, the target index corresponding to the target virtual gift is determined in the header data of the video file; In the video file, the image frame pointed to by the target index is determined as the target image frame; Decoding the target image frame using a hardware decoding module includes: According to the order information and position information recorded in the target index, the sub-image frames in the target image frame are decoded. The order information reflects the rendering order of the sub-image frames, and the position information reflects the position of the sub-image frames in the target image frame.

7. The method according to claim 6, characterized in that, The method further includes: Generate any one of the original images, wherein the original image comprises multiple sub-images reflecting different animation states of the virtual gift; The video file is generated from multiple original images.

8. An image processing apparatus, characterized in that, include: The determination module is used to respond to a selection operation triggered by the user, determine a first identifier of the virtual object selected by the user, determine the video file corresponding to the virtual object according to the first identifier, and determine an image frame with a second identifier that matches the first identifier in the video file as a target image frame. The video file includes at least one image frame obtained by video encoding at least one original image. The at least one image frame is a keyframe. The target image frame is any one of the at least one image frame. Any original image contains the virtual object displayed in the live broadcast interface. A decoding module is used to decode the target image frame using a hardware decoding module; The hardware decoding module is a digital signal processing (DSP) chip. The rendering module is used to render the decoding result by utilizing the vertex data and texture information contained in the decoding result with the help of the hardware decoding module.

9. An electronic device, characterized in that, When the computer instructions stored in the electronic device are executed by one or more processors, the one or more processors cause the processors to perform at least the following actions: In response to a user-triggered selection operation, a first identifier of the virtual object selected by the user is determined, and a video file corresponding to the virtual object is determined based on the first identifier. In the video file, an image frame with a second identifier that matches the first identifier is determined as the target image frame. The video file includes at least one image frame obtained by video encoding at least one original image. The at least one image frame is a keyframe. The target image frame is any one of the at least one image frame. Any original image contains the virtual object displayed in the live broadcast interface. The target image frame is decoded using a hardware decoding module; The hardware decoding module is a digital signal processing (DSP) chip. The hardware decoding module utilizes the vertex data and texture information contained in the decoding result to complete the rendering of the decoding result.

10. A non-transitory machine-readable storage medium, characterized in that, The non-transitory machine-readable storage medium stores executable code that, when executed by a processor of an electronic device, causes the processor to perform the image processing method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video processing method and video processing device

    CN106385591A

  • Eyebrow shape transformation method and device, client, server and storage medium

    CN114581294A

  • Virtual gift special effect playing method and device, equipment and medium

    CN116193151A