Image processing method and device, storage medium and electronic equipment

By acquiring and analyzing the positional changes of target elements in continuous images, determining their relative positions, and switching layers, the problems of low processing efficiency and poor 3D effects in existing technologies are solved, and efficient 3D effect generation is achieved.

CN121937685APending Publication Date: 2026-04-28MIGU MUSIC CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MIGU MUSIC CO LTD
Filing Date
2025-12-03
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, the need to manually process images where the digital human overlaps with the background border results in low processing efficiency and poor 3D effects in complex scenes.

Method used

By acquiring the positional changes of target elements in continuous images, the relative position of the target elements and the background image is determined, and the display layers of the target elements and the background image are switched in the switchable frame interval to generate continuous target images.

Benefits of technology

It improves the efficiency of layer switching, ensures the enhancement of 3D effects, and achieves accurate display of target elements when the background image borders overlap.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937685A_ABST
    Figure CN121937685A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, a storage medium and electronic equipment, and relates to the technical field of image processing, and the method comprises the steps: obtaining the position change condition of a target element in a continuous image; determining a relative position of the target element and a background image containing a frame according to the position change condition; determining a switchable frame interval corresponding to the target element from the continuous image based on the relative position; and switching the target element and the display layer of the background image in the switchable frame interval to generate a target continuous image. Compared with the prior art, the method has the advantages that the target element can be displayed in the corresponding target layer under the condition that the target element coincides with the frame of the background image, so that the three-dimensional effect of the generated target continuous image can be ensured, and the processing efficiency of layer switching is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image processing method, apparatus, storage medium and electronic device. Background Technology

[0002] Naked-eye 3D video is a technology that allows users to view video content with a 3D effect on a regular screen without wearing 3D glasses or a virtual reality (VR) headset. It achieves a stereoscopic visual effect perceived naturally by the human eye through specific display devices and multi-view rendering technology.

[0003] Currently, the main method involves acquiring an image where the digital human overlaps with the background border, and then manually placing the background border from that image below the digital human material. This allows the digital human to cover the background border as it moves, creating a 3D effect of the digital human moving out of the background frame.

[0004] However, this image processing method requires manual processing to identify images where the digital human overlaps with the background border. If there are many complex scenes that need to be processed, the processing efficiency will be low, and the resulting video may even have poor 3D effects. Summary of the Invention

[0005] In view of this, this application provides an image processing method, apparatus, storage medium and electronic device, the main purpose of which is to improve the technical problem that the existing technology requires manual processing of images in which the digital human overlaps with the background border. If there are a large number of complex scenes that need to be processed, the processing efficiency will be low, and even the three-dimensional effect of the obtained video will be poor.

[0006] In a first aspect, this application provides an image processing method, comprising: To obtain the positional changes of target elements in a continuous image; The relative position of the target element and the background image containing the border is determined based on the positional changes. Based on the relative position, determine the switchable frame interval corresponding to the target element from the continuous images; The display layers of the target element and the background image are switched within the switchable frame interval to generate a continuous target image.

[0007] Secondly, this application provides an image processing apparatus, comprising: The acquisition module is configured to acquire the positional changes of target elements in continuous images; The determination module is configured to determine the relative position of the target element and the background image containing the border based on the position change. The determining module is also configured to determine the switchable frame interval corresponding to the target element from the continuous images based on the relative position; The switching module is configured to switch the display layers of the target element and the background image within the switchable frame interval to generate a continuous target image.

[0008] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image processing method described in the first aspect.

[0009] Fourthly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the image processing method described in the first aspect.

[0010] By employing the above technical solutions, this application provides an image processing method, apparatus, storage medium, and electronic device. Compared with existing technologies, this application obtains the positional changes of target elements in a continuous image; determines the relative position of the target element and the background image containing the border based on the positional changes, enabling the application to analyze the positional changes and accurately determine the relative position of the target element and the background image containing the border; determines the switchable frame interval corresponding to the target element from the continuous image based on the relative position; and switches the display layers of the target element and the background image within the switchable frame interval to generate a continuous target image. This allows the application to determine the frame interval where layer switching is possible based on the continuous image, and then switch the layer where the target element is located within the frame interval, i.e., switch the layer where the target element is located to the upper or lower layer of the background image. This enables the target element to be displayed on the corresponding target layer when the border of the target element and the background image overlaps, thereby ensuring the three-dimensional effect of the generated continuous target image and improving the processing efficiency of layer switching. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1A schematic flowchart of an image processing method provided in an embodiment of this application is shown; Figure 2 A schematic flowchart of an image processing method provided in an embodiment of this application is shown; Figure 3 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 4 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 5 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 6 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 7 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 8 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 9 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 10 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 11 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 12 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 13 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 14 A schematic diagram illustrating an example provided in an embodiment of this application is shown; Figure 15 This illustration shows a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application; Figure 16 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0014] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0015] To address the technical problem in existing technologies that require manual processing to identify images where the digital human overlaps with the background border, leading to low processing efficiency and even poor 3D quality in the resulting video, this embodiment provides an image processing method, such as... Figure 1As shown, the method includes: Step 101: Obtain the positional changes of the target element in the continuous image.

[0016] In the embodiments of this application, continuous images can be images arranged continuously in time or space order, used to represent dynamic processes, motion changes, or continuous scenes; for example, continuous images in the embodiments of this application can be videos, animations, etc.

[0017] In some examples, the elements contained in an image can be the basic components of all visual content in the image, which together constitute the overall information and visual effect of the image. These elements can be static, dynamic, semantic, or structured; it should be noted that the target element in the embodiments of this application can specifically be an element that needs to be combined with a background image containing a border; for example, the target element can include, but is not limited to, a digital human image, a target object, a target task, etc.

[0018] In this embodiment, the position change can be the position change of the target element in the continuous image; for example, if the continuous image includes continuous image 1, continuous image 2, continuous image 3, and continuous image 4, the change can include the change of the position of the target element in continuous image 1 to the position of the target element in continuous image 2, the change of the position of the target element in continuous image 2 to the position of the target element in continuous image 3, and the change of the position of the target element in continuous image 3 to the position of the target element in continuous image 4.

[0019] Step 102: Determine the relative position of the target element and the background image containing the border based on the position change.

[0020] In this application embodiment, the background image including the border can be an image with an opaque border or a background that is transparent except for the border. For example, it can be a "door" shaped background, a "return" shaped background, etc.

[0021] As an alternative approach, the relative position of the target element and the background image containing the border can be determined based on the positional changes. For example, the relative position of the target element and the background image containing the border can be that the target element is in the upper left part of the background image, or it can be that the target element is in the lower right part of the background image, and so on. No further examples will be given here.

[0022] Step 103: Determine the switchable frame interval corresponding to the target element from the continuous images based on the relative position.

[0023] In some examples, the switchable frame range can be the frame range in a continuous image where layer switching is possible; for example, based on the relative position of the border and the target element, if it is determined that images 3-15 in a continuous image can be layer switched, then frames 3-15 in the continuous image can be the switchable frame range; if it is determined that images 6-9 in a continuous image can be layer switched, then frames 6-9 in the target continuous image can be the switchable frame range, and so on, without further examples here.

[0024] Step 104: Switch the display layers of the target element and the background image within the switchable frame interval to generate a continuous target image.

[0025] In this embodiment, the display layer is a technique for managing multiple image elements in layers. Each layer can be edited independently, and its transparency, overlay method, etc., can be controlled to ultimately combine them into a complete image or interface. It should be noted that the display layer in this embodiment can be a display layer between the target element and the background image.

[0026] In some examples, the target element can be displayed on top of the background image or on the bottom of the background image based on the switchable frame interval.

[0027] Compared with existing technologies, this embodiment obtains the positional changes of target elements in a continuous image; determines the relative position of the target element and the background image containing the border based on the positional changes, enabling this embodiment to analyze the positional changes and accurately determine the relative position of the target element and the background image containing the border; determines the switchable frame interval corresponding to the target element from the continuous image based on the relative position; and switches the display layers of the target element and the background image within the switchable frame interval to generate a continuous target image. This embodiment can determine the frame interval where layer switching is possible based on the continuous image, and then switch the layer where the target element is located within the frame interval, that is, switch the layer where the target element is located to the upper layer or the lower layer of the background image. This allows the target element to be displayed on the corresponding target layer when the border of the target element and the background image overlaps, thereby ensuring the three-dimensional effect of the generated continuous target image and improving the processing efficiency of layer switching.

[0028] As a refinement and extension of the above embodiments, when performing the task of "determining the switchable frame interval corresponding to the target element from the continuous images based on relative position", the following methods can be used, but are not limited to them, such as... Figure 2 As shown, the method includes: Step 201: Divide the continuous images into multiple continuous image groups based on their relative positions.

[0029] For example, taking a digital human as the target element in a continuous image, for a newly created digital human video, each frame of data needs to be recorded synchronously, including: time, distance to the screen (i.e., the display screen in this embodiment), and coordinates of the rectangular area formed by the digital human. This information is stored in the Supplemental Enhancement Information (SEI) area encoded in H264, and the final encapsulation format is a video with a transparent background channel, such as MOV. For existing digital human videos, the distance from the digital human to the screen (i.e., the square area occupied by the digital human) can be calculated based on the size of the digital human image in each frame (i.e., the square area occupied by the digital human), with the distance being the closest to the screen when it is at its maximum. At the same time, the digital human information is supplementarily saved in the SEI area of ​​the H264 encoded video for use in subsequent positioning and matching.

[0030] In this embodiment, the video (i.e., the continuous images in this embodiment) can be logically segmented based on a fixed number of background widths (i.e., the background images in this embodiment). For example, such as... Figure 3 As shown, the video (i.e., the continuous image in this application embodiment) can be logically segmented based on the width of three backgrounds (i.e., the background image in this application embodiment).

[0031] It should be noted that the continuous image can be divided according to the width of 3 backgrounds (i.e., the background image in this embodiment of the application) or according to other reference conditions. The specific division method is not specifically limited in this embodiment of the application.

[0032] Step 202: For a target image group in multiple consecutive image groups, determine at least one set of target image sequences in the target image group in which the border of the background image does not coincide with the target element.

[0033] In this embodiment, the relative position of the target element and the background image can determine whether the border of the background image coincides with the target element. For example, based on the example in step 202, based on the relative position of the digital human (i.e., the target element in this embodiment) video (i.e., the continuous images in this embodiment) and the background border, it can be calculated whether the digital human (i.e., the target element in this embodiment) in each frame is within the inner border of the background. Figure 4 As shown, the second and fourth frames indicate that the digital person at that location (i.e., the target element in this embodiment) is contained within the border and does not overlap with the border.

[0034] As an optional method, all frames are processed repeatedly to obtain, such as Figure 5 The curve shown indicates that at least one image sequence in which the border does not coincide with the target element can be identified. Figure 5The interval corresponding to the thick gray line segment.

[0035] Step 203: Determine the frame intervals corresponding to at least one set of target image sequences as the target switchable frame intervals corresponding to the target image group.

[0036] In some examples, the image sequence in which the border of the background image does not coincide with the target element can be the frame interval to be switched. For example, if images 6-10 in the target image group are image sequences in which the border of the background image does not coincide with the target element, then images 6-10 can be the target switchable frame interval corresponding to the target image group.

[0037] For example, if images 2-8 in the target image group are a sequence of images in which the border of the background image does not coincide with the target element, then images 2-8 can be the target switchable frame interval corresponding to the target image group.

[0038] Step 204: Determine the target switchable frame intervals corresponding to multiple consecutive image groups as the switchable frame intervals corresponding to the target element.

[0039] Optionally, when performing the action of "switching the display layers of the target element and the background image in the switchable frame interval to generate a target continuous image", the following methods can be used, but are not limited to: determining the background binding points corresponding to multiple continuous image groups based on their relative positions; switching the display layers of the target element and the background image in the target switchable frame interval for the target image group in the multiple continuous image groups; compositing the switched target image group based on the target background binding points corresponding to the target image group; and generating a target continuous image based on the continuous images composed of multiple continuous image groups.

[0040] In this embodiment of the application, the target continuous image can be an image obtained by combining a continuous image with a background image containing a border. It should be noted that each frame of the target continuous image contains a background image and a target element.

[0041] For example, if image 3 in a continuous image can be determined as the point of combination with the background image based on its relative position, then the continuous image and the background image containing the border can be composited based on image 3; if image 5 in a continuous image can be determined as the point of combination with the background image based on its relative position, then the continuous image and the background image containing the border can be composited based on image 5, and so on, without further examples here.

[0042] Optionally, when performing the action of "switching the display layers of the target element and the background image in the target switchable frame interval for a target image group in multiple consecutive image groups", the following methods can be used, but are not limited to: determining the target display layer of the target element in the target switchable frame interval based on the relative position; and switching the display layer of the target element to the target display layer in the switchable frame interval.

[0043] In some examples, the target display layer can be the layer that the target element needs to be switched to. For instance, if the target display layer is the lower layer, the target element can be switched to the lower layer of the background image, and if the target display layer is the upper layer, the target element can be switched to the upper layer of the background image.

[0044] Optionally, when performing the "determine the target display layer of the target element in the target switchable frame interval based on relative position", the following methods can be used, but are not limited to these, including: determining the position change trend of the switchable frame interval based on relative position; if the change trend is determined to be approaching the display screen, determining the upper layer of the background image as the target display layer; if the change trend is determined to be moving away from the display screen, determining the lower layer of the background image as the target display layer.

[0045] For example, if the position change trend is a falling edge (i.e., the change trend in this application embodiment is towards the display screen), then the digital human (i.e., the target element in this application embodiment) is at the bottom layer and the border background is at the top layer; if the position change trend is a rising edge (i.e., the change trend in this application embodiment is away from the display screen), then the digital human (i.e., the target element in this application embodiment) is at the top layer and the border background is at the bottom layer.

[0046] It should be noted that the digital human (i.e., the target element in this embodiment) is switched to the foreground when the falling edge occurs (i.e., the trend of change in this embodiment is towards the display screen), and the digital human (i.e., the target element in this embodiment) is switched to the background when the rising edge occurs (i.e., the trend of change in this embodiment is away from the display screen). For specific switching frames, they can be selected according to a strategy or the position can be randomly selected by default. The selection of specific switching frames in the switchable frame range is not specifically limited in this embodiment.

[0047] For example, such as Figure 6 As shown, based on the identified switchable foreground and background location regions (i.e., the switchable frame intervals in this embodiment), a time point can be randomly selected to segment the foreground and background; correspondingly, when the next rising edge is encountered (i.e., the trend of change in this embodiment is moving away from the display screen), it means that the digital human (i.e., the target element in this embodiment) has moved into the background border, thus it can be as follows: Figure 7As shown, in the next switchable area (i.e., the switchable frame interval in this embodiment), a time point is randomly selected again, and the foreground and background are switched.

[0048] Optionally, when performing the "determining the background junction point corresponding to multiple consecutive image groups based on relative position", the following methods can be used, but are not limited to these: determining the distance information between the target element and the display screen based on the relative position, where the display screen is the screen that displays the consecutive images; determining the target frame image that meets the distance condition from the consecutive images based on the distance information, and determining the target frame image as the junction point between the consecutive images and the background image.

[0049] As an alternative approach, the screen (i.e., the display screen in the embodiments of this application) can specifically be the screen of a device used to display multiple consecutive frames of images, such as a mobile phone screen, a tablet screen, a monitor screen, etc.

[0050] In the embodiments of this application, such as Figure 8 As shown, the screen can be placed vertically (equivalent to placing the phone vertically). The Z-axis can represent the distance from the digital human (i.e., the continuous image containing the target element in this embodiment) to the screen (i.e., the display screen in this embodiment). For example, if the phone is placed on a table and the user is looking at the screen, the X-axis can be the width of the phone, and the Z-axis can be the depth of field from the eye to the phone. The X and Z axes form the area of ​​the desktop on the back of the phone.

[0051] For example, such as Figure 9 As shown, based on the troughs of the movement trajectory of the digital human (i.e., the target element in this embodiment) in the video (i.e., the continuous image in this embodiment), m2 is determined to be the point closest to the screen. Based on the positional changes, the trajectory of the target element in the continuous image can be generated, specifically as follows: Figure 10 As shown, when the digital human (i.e. the target element in this application embodiment) has too large an activity range in the video (i.e. the continuous image in this application embodiment), the digital human (i.e. the target element in this application embodiment) can be logically divided into multiple motion areas (i.e. multiple continuous image groups in this application embodiment), the troughs (m2, m5, m6) of the distance between the digital human (i.e. the target element in this application embodiment) and the screen are calculated, and the first trough position m2 is selected as the reference position, and the background image is tiled.

[0052] For example, within a single motion area of ​​the digital human (i.e., the target element in this embodiment), the position of the digital human (i.e., the target element in this embodiment) closest to the screen (i.e., the position with the fewest Z-axis coordinates) can be found as the point of combination for compositing the digital human (i.e., the target element in this embodiment) video (i.e., the continuous image in this embodiment) and the background (i.e., the background image in this embodiment), that is, the position where the digital human (i.e., the target element in this embodiment) goes out of bounding box (i.e., the border in this embodiment). This process is to find the digital human (i.e., the target element in this embodiment) to cover the background border to the greatest extent, that is, to maximize the out-of-bounds effect. It should be noted that determining the point of combination before compositing the continuous image with the background image can make the naked-eye 3D effect more obvious.

[0053] As an alternative approach, within a single motion area of ​​the digital human (i.e., the target element in this embodiment), the position furthest from the screen (i.e., the position with the largest Z-axis coordinate) of the digital human (i.e., the target element in this embodiment) can be found as the point of combination for compositing the digital human (i.e., the target element in this embodiment) video (i.e., the continuous image in this embodiment) and the background (i.e., the background image in this embodiment), that is, the position where the digital human (i.e., the target element in this embodiment) exits the frame (i.e., the border in this embodiment) from the lower layer of the background image. This process involves finding the background border to cover the digital human (i.e., the target element in this embodiment) to the greatest extent possible, i.e., maximizing the out-of-frame effect. It should be noted that determining the point of combination before compositing the continuous image with the background image can make the naked-eye 3D effect more obvious.

[0054] It should be noted that during the synthesis process, if the digital human (i.e., the target element in this embodiment) and the background may not overlap (i.e., the digital human (i.e., the target element in this embodiment) is entirely within the background border), dynamic adjustments can be made according to the following, but not limited to: if the user accepts a decrease in the clarity of the digital human (i.e., the target element in this embodiment), the size of the digital human (i.e., the target element in this embodiment) video (i.e., the continuous image in this embodiment) is enlarged according to the background border ratio; otherwise, the position of the digital human (i.e., the target element in this embodiment) video (i.e., the continuous image in this embodiment) is shifted downwards, with the bottom of the digital human (i.e., the target element in this embodiment) closest to the screen shifted downwards to the middle position of the background border. After adjustment, video positioning and matching are re-executed.

[0055] In this embodiment of the application, for a video within an active area (i.e., a continuous group of images in this embodiment), when the digital human (i.e., the target element in this embodiment) remains outside the background border for a relatively long period of time, for example, Figure 11 As shown, videos longer than 1 second will be directly cropped. Specifically, the video within the dotted frame that remains outside the border for 3 seconds is invisible to the digital human; therefore, only the first 1 second is retained, and the remaining 2 seconds are cropped and removed during compositing. Similarly, videos in other activity areas are also cropped according to this principle. Figure 12 As shown, the video frames from B to C will be cropped.

[0056] In some examples, since there are multiple local active regions (i.e., consecutive image groups in this embodiment), it is necessary to reset the background border position according to the current active region (i.e., consecutive image groups in this embodiment). Therefore, after playing one active region (i.e., consecutive image groups in this embodiment), it is necessary to transition to the next region (i.e., consecutive image groups in this embodiment); for example, such as Figure 13 As shown, the background border moves from position m2 to position m5; in the first active area (i.e., the continuous image group in this embodiment), the background is placed at position m2, starting from point 0, and the video composite ends at point B (the last exposed position of this scene); in the second active area (i.e., the continuous image group in this embodiment), the background is placed at position m5, starting from point C, and the video composite ends at the last exposed position of this scene; if there are subsequent active areas (i.e., the continuous image group in this embodiment), this process continues to complete the video composite for each area (i.e., the continuous image group in this embodiment); the videos (i.e., the continuous images in this embodiment) are connected pairwise by custom transitions, which can be specified by parameters. The default transition is a switching between darkening and brightening of the video.

[0057] As an example, this application also provides an example, such as Figure 14As shown, the system can automatically switch between the digital human (i.e., the target element in this embodiment) and the background foreground and background to create a naked-eye 3D effect of entering and exiting the frame. Specifically, it includes: Step 1, Material Preparation: A background image with a border (portrait mode), and a video of the digital human (i.e., the target element in this embodiment) (landscape mode). Step 2, Material Positioning: Dynamically calculate the position of the digital human (i.e., the target element in this embodiment) exiting the frame. At this position, the border background and the digital human (i.e., the target element in this embodiment) video are aligned. Step 3, Switching Timing Calculation: Calculate the positions of all video frames embedded in the border of the digital human (i.e., the target element in this embodiment), which serve as the timing for the switching between the background border and the digital human foreground and background, and dynamically select them according to a strategy. Step 4, If the switching timing does not meet the requirements (the overlapping point and non-overlapping point of the digital human and background cannot be found simultaneously), readjust the material size according to the strategy, and then proceed with the positioning in Step 2. Step 5, Video Compositing: Based on the selected foreground and background switching time points, cut and composite the background and the digital human (i.e., the target element in this embodiment) (portrait mode cropping).

[0058] It should be noted that in this embodiment, the bordered background image and the digital human video (i.e., the continuous image in this embodiment) are layered and superimposed. Through dynamic calculation that separates the background border from the digital human (i.e., the target element in this embodiment), the upper and lower layers are automatically adjusted to achieve a 3D effect of the digital human (i.e., the target element in this embodiment) entering and exiting the frame during movement, without the need for external devices. Furthermore, for large-scene digital human (i.e., the target element in this embodiment) videos, partitioned processing and synthesis are supported. Ultimately, given any bordered door background image and a widescreen motion video of the digital human, a 3D effect of the digital human (i.e., the target element in this embodiment) entering and exiting the door (i.e., the border in this embodiment) is automatically synthesized.

[0059] Existing technical solutions only involve static production. For each instance of a digital human entering or exiting a background border (which acts like a door frame), the hierarchical relationship between the digital human and the border needs to be adjusted. During editing, the digital human video and background image need to be manually separated according to their corresponding levels. At a selected time point (when there is no overlap), the levels of the two materials need to be manually adjusted; then, the levels need to be adjusted again for the next entry and exit from the frame. For complex digital human movements, manual operation is extremely tedious and time-consuming. To address this high complexity, this application's embodiment can employ strategies to achieve matching such as video scaling or movement. For large-scale activities of the digital human (i.e., the target element in this application's embodiment), the activity is divided into zones, and the zones are used to synthesize videos. Finally, the videos are synthesized using transitions, avoiding situations where a fixed background position causes the digital human (i.e., the target element in this application's embodiment) to be out of the scene, thus improving the success rate of video production. 3. Automated switching between foreground and background. The switching between foreground and background is determined based on the falling edge (i.e., the trend in this application's embodiment is towards the display screen) or rising edge (i.e., the trend in this application's embodiment is away from the display screen) of the digital human's (i.e., the target element in this application's embodiment) movement.

[0060] Compared with existing technologies, this embodiment obtains the positional changes of target elements in a continuous image; determines the relative position of the target element and the background image containing the border based on the positional changes, enabling this embodiment to analyze the positional changes and accurately determine the relative position of the target element and the background image containing the border; determines the switchable frame interval corresponding to the target element from the continuous image based on the relative position; and switches the display layers of the target element and the background image within the switchable frame interval to generate a continuous target image. This embodiment can determine the frame interval where layer switching is possible based on the continuous image, and then switch the layer where the target element is located within the frame interval, that is, switch the layer where the target element is located to the upper layer or the lower layer of the background image. This allows the target element to be displayed on the corresponding target layer when the border of the target element and the background image overlaps, thereby ensuring the three-dimensional effect of the generated continuous target image and improving the processing efficiency of layer switching.

[0061] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides an image processing device, such as... Figure 15 As shown, the device includes: an acquisition module 31, a determination module 32, and a switching module 33.

[0062] The acquisition module 31 is configured to acquire the positional changes of target elements in a continuous image; The determining module 32 is configured to determine the relative position of the target element and the background image containing the border based on the position change. The determining module 32 is further configured to determine the switchable frame interval corresponding to the target element from the continuous images based on the relative position; The switching module 33 is configured to switch the display layers of the target element and the background image in the switchable frame interval to generate a continuous target image.

[0063] In some examples of this embodiment, the determining module 32 is specifically configured to divide the continuous image into multiple continuous image groups based on the relative position; for a target image group in the multiple continuous image groups, determine at least one set of target image sequences from the target image group in which the border of the background image does not coincide with the target element; determine the frame interval corresponding to the at least one set of target image sequences as the target switchable frame interval corresponding to the target image group; and determine the target switchable frame interval corresponding to the multiple continuous image groups as the switchable frame interval corresponding to the target element.

[0064] In some examples of this embodiment, the switching module 33 is specifically configured to: determine the background binding points corresponding to the plurality of consecutive image groups based on the relative positions; switch the display layers of the target element and the background image in the target switchable frame interval for the target image group in the plurality of consecutive image groups; synthesize the switched target image group based on the target background binding points corresponding to the target image group; and generate the target continuous image based on the continuous images synthesized from the plurality of consecutive image groups.

[0065] In some examples of this embodiment, the switching module 33 is further configured to determine the target display layer of the target element in the target switchable frame interval based on the relative position; and to switch the display layer of the target element to the target display layer in the switchable frame interval.

[0066] In some examples of this embodiment, the switching module 33 is further configured to determine the position change trend of the switchable frame interval based on the relative position; if the change trend is determined to be approaching the display screen, the upper layer of the background image is determined as the target display layer; if the change trend is determined to be moving away from the display screen, the lower layer of the background image is determined as the target display layer.

[0067] In some examples of this embodiment, the switching module 33 is further configured to determine the distance information between the target element and the display screen based on the relative position, wherein the display screen is the screen for displaying the continuous image; determine a target frame image that meets the distance condition from the continuous image based on the distance information, and determine the target frame image as the junction point of the continuous image and the background image.

[0068] It should be noted that other corresponding descriptions of the functional units involved in the image processing apparatus provided in this embodiment can be found in [reference needed]. Figure 1 and Figure 2 The corresponding descriptions in [the document] will not be repeated here.

[0069] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method shown.

[0070] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0071] like Figure 16 The diagram shown is a hardware structure schematic of an electronic device according to the present invention, comprising: At least one processor 401; and, A memory 402 is communicatively connected to at least one of the processors 401; wherein, The memory 402 stores instructions that can be executed by at least one of the processors to enable the at least one of the processors to perform the image processing method as described above.

[0072] Figure 16 Take a processor 401 as an example.

[0073] The electronic device may also include an input device 403 and a display device 404.

[0074] The processor 401, memory 402, input device 403, and display device 404 can be connected via a bus or other means. Figure 16 Taking the example of a connection between China and Israel via a bus.

[0075] Memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the image processing method in the embodiments of this application, for example, Figure 1 and Figure 2 The method flow is shown. The processor 401 executes various functional applications and data processing by running non-volatile software programs, instructions, and modules stored in the memory 402, thereby implementing the image processing method in the above embodiments.

[0076] Memory 402 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created according to the use of the image processing method, etc. Furthermore, memory 402 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 402 may optionally include memory remotely located relative to processor 401, and these remote memories may be connected to the apparatus performing the image processing method via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0077] The input device 403 can receive user clicks and generate signal inputs related to user settings and function control of the image processing method. The display device 404 may include a display screen or other display device.

[0078] The one or more modules are stored in the memory 402, and when run by the one or more processors 401, they execute the image processing method in any of the above method embodiments.

[0079] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.

[0080] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0081] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0082] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware. By applying the solution of this embodiment, compared with the existing technology, this embodiment obtains the position change of the target element in the continuous image; determines the relative position of the target element and the background image containing the border based on the position change, so that this embodiment can analyze based on the position change and accurately analyze the relative position of the target element and the background image containing the border; determines the switchable frame interval corresponding to the target element from the continuous image based on the relative position; switches the display layer of the target element and the background image in the switchable frame interval to generate the target continuous image, so that this embodiment can determine the frame interval in which the layer can be switched based on the continuous image, and then switch the layer where the target element is located in the frame interval content, that is, the layer where the target element is located can be switched to the upper layer or the lower layer of the background image, so that the target element can be displayed in the corresponding target layer when the border of the target element and the background image overlaps, thereby ensuring the three-dimensional effect of the generated target continuous image and improving the processing efficiency of layer switching.

[0083] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0084] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. An image processing method, characterized in that, include: To obtain the positional changes of target elements in a continuous image; The relative position of the target element and the background image containing the border is determined based on the positional changes. Based on the relative position, determine the switchable frame interval corresponding to the target element from the continuous images; The display layers of the target element and the background image are switched within the switchable frame interval to generate a continuous target image.

2. The method according to claim 1, characterized in that, Determining the switchable frame interval corresponding to the target element from the continuous images based on the relative position includes: The continuous images are divided into multiple continuous image groups based on their relative positions; For a target image group in the plurality of consecutive image groups, determine at least one set of target image sequences from the target image group in which the border of the background image does not coincide with the target element; The frame intervals corresponding to the at least one set of target image sequences are determined as the target switchable frame intervals corresponding to the target image group; The target switchable frame intervals corresponding to the multiple consecutive image groups are determined as the switchable frame intervals corresponding to the target element.

3. The method according to claim 2, characterized in that, The step of switching the display layers of the target element and the background image within the switchable frame interval to generate a continuous target image includes: Based on the relative positions, determine the background binding points corresponding to the plurality of consecutive image groups respectively; For the target image group in the plurality of consecutive image groups, the display layers of the target element and the background image are switched in the target switchable frame interval; The switched target image group is synthesized based on the target background binding points corresponding to the target image group; The target continuous image is generated based on the continuous images synthesized from the multiple continuous image groups respectively.

4. The method according to claim 3, characterized in that, The step of switching the display layers of the target element and the background image within the target switchable frame interval for the target image group in the plurality of consecutive image groups includes: Based on the relative position, determine the target display layer of the target element in the target switchable frame interval; In the switchable frame interval, the display layer of the target element is switched to the target display layer.

5. The method according to claim 4, characterized in that, Determining the target display layer of the target element within the target switchable frame interval based on the relative position includes: The positional change trend of the switchable frame interval is determined based on the relative position; If the trend of change is determined to be approaching the display screen, the upper layer of the background image is determined as the target display layer; If the trend of change is determined to be moving away from the display screen, the lower layer of the background image is determined as the target display layer.

6. The method according to claim 3, characterized in that, Determining the background binding points corresponding to the plurality of consecutive image groups based on the relative positions includes: The distance information between the target element and the display screen is determined based on the relative position, wherein the display screen is the screen that displays the continuous image; Based on the distance information, a target frame image that meets the distance condition is determined from the continuous images, and the target frame image is determined as the junction point of the continuous images and the background image.

7. An image processing apparatus, characterized in that, include: The acquisition module is configured to acquire the positional changes of target elements in continuous images; The determination module is configured to determine the relative position of the target element and the background image containing the border based on the position change. The determining module is also configured to determine the switchable frame interval corresponding to the target element from the continuous images based on the relative position; The switching module is configured to switch the display layers of the target element and the background image within the switchable frame interval to generate a continuous target image.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.