Generation method and generation apparatus for animation, and electronic device
By determining the moving direction of the mirror according to the content of the picture, the poor animation display effect caused by the fixed moving direction of the mirror in the prior art is solved, and high-quality animations are generated and the user's visual experience is enhanced.
Patent Information
- Application Number
- PCT/CN2024/104283
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-07-08
- Publication Date
- 2025-05-08
AI Technical Summary
When generating an animation with the display effect of the mirror, the movement direction of the mirror is fixed and the image content cannot be flexibly matched, resulting in poor animation display effect.
By obtaining the displayed content of the picture, the movement direction of the moving mirror is determined, so that the movement direction of the moving mirror is related to the content of the picture, thereby generating a harmonious high-quality animation of the moving mirror.
It enhances the picture and immersion of the animation, improves the display effect of the animation, and provides users with a coherent immersive experience.
Smart Images

Figure CN2024104283_08052025_PF_FP_ABST
Abstract
Description
Animation generation method, generation device and electronic equipment
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on October 30, 2023, with application number 202311435092.1 and application name “Animation generation method, generation device and electronic device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of image processing technology, and in particular to an animation generation method, generation device and electronic equipment. Background Art
[0003] With the continuous advancement of terminal technology, more and more users are using images to record their lives. Some users also want to create animations of their favorite photos for display. To meet this need, technology for generating animations from images has emerged. Users can combine multiple images from their local gallery to create an animation, which can then be shared on social media platforms. This animation can also be used as a dynamic wallpaper for the desktop or screensaver, broadening the source of dynamic wallpapers for terminals.
[0004] When generating animations from images, the terminal device can perform panning to enhance the presentation of the images in the animation. For example, the images in the animation can be displayed from left to right, from top to bottom, or zoomed in and out.
[0005] Currently, when a terminal device generates an animation with a camera movement effect, it can process the images in the animation according to the preset camera movement direction. For example, if the preset camera movement direction is from left to right, the images in the generated animation will be displayed according to the left-to-right camera movement direction. However, animations often include multiple images, and current technology can only display these multiple images according to a fixed camera movement direction. This lacks flexibility in the use of camera movement and results in poor animation display effects.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide an animation generation method, generation device and electronic device, which can determine the movement direction of the camera according to the display content of the image, can generate high-quality animation with harmonious camera movement, and improve the display effect of the animation.
[0008] In a first aspect, a method for generating an animation is provided, comprising: obtaining a first picture; determining a first camera movement direction based on the display content of the first picture; and generating the animation based on the first camera movement direction and the first picture, wherein the first camera movement direction is the camera movement direction corresponding to the animation.
[0009] According to the animation generation method provided by the embodiment of the present application, the camera movement direction (referred to as the first camera movement direction) can be first determined based on the display content of the image (referred to as the first image), and then the animation can be generated based on the first camera movement direction and the first image, that is, the camera is displayed according to the first camera movement direction to generate the animation. Because the camera movement direction is determined based on the display content of the image, the camera movement direction and the display content of the image can be associated and matched, and the camera movement can be flexibly used according to the image content, thereby enhancing the visual sense and immersion of the animation, and can generate high-quality animations with harmonious camera movements, thereby improving the display effect of the animation.
[0010] Because the movement direction of the camera is correlated with the displayed content of the image, the use of the camera is more harmonious and smooth, the quality of the generated animation is higher, and the picture sense is stronger, thereby improving the display effect of the animation, providing users with a coherent immersive experience, and improving the user's visual experience.
[0011] Optionally, the first picture may be captured by the electronic device itself, or may come from the Internet (eg, a social platform).
[0012] Optionally, the first image can also be obtained from other animations, which can be, for example, images or videos in GIF format. For example, the first image can be obtained by taking a screenshot or extracting frames from a video. This application does not limit the specific source of the first image.
[0013] Optionally, the embodiments of the present application do not limit the specific method of "determining the first camera direction based on the image content." For example, the subject in the first image can be identified, and the first camera direction can be determined based on the type of the subject. The subject here can be, for example, a human, an animal, a plant, a mountain, a river, a toy, jewelry, fruit, a building, furniture, the sun, the moon, a vehicle, etc., but is not limited thereto. The first camera direction can also be determined based on a more specific type of identified subject. For example, animals can be categorized as cats, dogs, lions, elephants, or birds, and the first camera direction corresponding to different types of animals may be different. And / or, the first camera direction can also be determined based on factors such as the number of subjects, their position in the image, and the proportion of the entire image. The first camera direction can also be determined based on factors such as the layout and color tone of the image. In addition, if the subject is a person, the first camera direction can also be determined based on factors such as the person's expression, movement, and gestures.
[0014] Optionally, when an animation is generated by using multiple images, the "animation" in "the first camera movement direction is the camera movement direction corresponding to the animation" can also be understood or replaced by an animation clip. For example, the multiple images include a first image and a second image, the animation clip corresponding to the first image is the first animation clip, the animation clip corresponding to the second image is the second animation clip, the first camera movement direction is the camera movement direction corresponding to the first animation clip, and the second camera movement direction is the camera movement direction corresponding to the second animation clip. The first animation clip and the second animation clip are spliced together to form a complete animation.
[0015] In a possible implementation, determining the first camera movement direction according to the display content of the first image includes: identifying the display content of the first image to obtain a first orientation of the target object; and determining the first camera movement direction according to the first orientation.
[0016] Determining the camera direction through the directional information of the target object in the image can make the determined camera movement direction more closely matched with the image content, making the camera movement more harmonious and smooth, enhancing the visual and immersive feeling of the animation, and being conducive to generating high-quality animations with harmonious camera movement, thereby improving the display effect of the animation and increasing its playability and fun.
[0017] Optionally, the target object can be any object in the first picture, such as a person, an animal, a plant (such as a tree), a scenery (such as a mountain, a river, etc.), a building (such as a house, a bridge, etc.), a road, a toy, jewelry, fruit, furniture, the sun, the moon, a vehicle, etc., but is not limited thereto.
[0018] Optionally, the target object may be the subject of the photograph or may not be the subject of the photograph. For example, the target object may be a river or a path in the shooting background, which is not limited in this application.
[0019] Optionally, the display content of the first picture can be understood according to an image understanding neural network model to obtain or determine the first orientation.
[0020] In a possible implementation, before identifying the display content of the first image to obtain a first pointing of the target object, the method further includes: identifying the display content of the first image to obtain multiple objects; and determining the target object based on the multiple objects.
[0021] That is to say, in some scenarios, when recognizing the first image, multiple objects (such as multiple people) may be recognized. In this case, the target object can be determined based on the multiple objects first, and then the first direction corresponding to the target object can be identified. The above settings can enable the generation method provided by this application to meet the usage requirements in more scenarios, or enable the generation method provided by this application to have a wider scope of application. This application does not specifically limit the specific implementation method of determining the target object based on multiple objects.
[0022] In one possible implementation, determining the target object based on the multiple objects includes: determining the subject among the multiple objects as the target object; or, determining the target object from the multiple objects based on preset priority information; or, determining the target object from the multiple objects based on the proportion of the screen occupied by the multiple objects; or, randomly selecting the multiple objects to obtain the target object.
[0023] That is to say, the electronic device understands the content of the first picture and may identify multiple objects. At this time, the electronic device can determine the aforementioned target object based on the multiple objects according to preset logical rules. Optionally, the electronic device can determine the subject (such as a person or an animal) among the multiple objects as the target object. Optionally, the electronic device can determine the object with the highest priority as the target object based on preset priority information. For example, different objects such as people, rivers, trees, buildings, roads, etc. have different priorities, and the person with the highest priority can be determined as the target object. Optionally, the electronic device can also determine the object with the largest proportion of the picture among the multiple objects as the target object. For example, if there are multiple people in the picture, the person with the largest proportion of the picture can be determined as the target object. Optionally, the multiple objects can also be randomly extracted, and the extracted object can be determined as the target object.
[0024] In a possible implementation, determining the first camera movement direction according to the first direction includes: determining a direction along the first direction as the first camera movement direction, or determining a direction opposite to the first direction as the first camera movement direction.
[0025] By determining the direction along or against the first pointing direction as the first camera movement direction, the method is simple and easy to implement, and can better correlate and match the first camera movement direction with the image content, thereby enhancing the visual sense and immersion of the animation and improving the user's visual experience.
[0026] Alternatively, the first direction may be a sight line direction from the lower left to the upper right, and the first camera movement direction is a direction opposite (opposite) to the sight line direction. For example, the first camera movement direction may be from right to left, from top to bottom, or from top right to bottom left (i.e., a combination of the two directions from right to left and from top to bottom). This application does not limit this. The situations in other directions are similar and will not be repeated here.
[0027] Alternatively, the direction along the first pointing direction may be determined as the first camera movement direction. For a line of sight from the lower left to the upper right, the first camera movement direction may be from left to right, from bottom to top, or from the lower left to the upper right (i.e., a combination of the two directions from left to right and from bottom to top). This application does not limit this. The situations in other directions are similar and will not be described in detail here.
[0028] In a possible implementation, the first direction includes a face orientation, a sight line direction, a moving direction, or an extending direction.
[0029] For example, the first direction may be the face orientation or sight direction of the subject (eg, a person or animal), the moving direction of the person, animal, or object, or the extending direction of a path or river.
[0030] For example, the target object may be a person, and the facial orientation or gaze direction may be determined through a facial feature recognition (detection) algorithm (such as a deep neural network model).
[0031] Optionally, the target object here is any object in the first image, and may or may not be the subject of the photograph. For example, the target object may be an object in the background of the subject of the photograph. For example, the target object may be a person, animal, or movable non-living object (such as a balloon, white cloud, or car) in the first image. In this case, the first direction may be the direction of movement of the person, animal, or object. For another example, the target object may be a river or a path. In this case, the first direction may be the extension direction of the river or path, etc., but is not limited thereto.
[0032] In a possible implementation, the first direction includes at least one of from front to back, from back to front, from left to right, from right to left, from top to bottom, and from bottom to top.
[0033] Optionally, the above directions can be defined with reference to the lens of the camera movement. Among them, from left to right, from right to left, from top to bottom, and from bottom to top are all plane movements within the picture (paper surface), the four directions of up, down, left, and right are the directions of up, down, left, and right of the picture seen by the user's naked eye, and the front and back directions are the directions perpendicular to the picture (paper surface), the direction close to the paper surface is forward, and the direction away from the paper surface is backward. Therefore, the above-mentioned direction from front to back is the direction away from the picture (paper surface), and from back to front is the direction gradually approaching the picture (paper surface).
[0034] In a possible implementation, determining the first camera movement direction according to the display content of the first image includes: identifying the display content of the first image to obtain multiple directions; and determining the first camera movement direction according to the multiple directions.
[0035] That is to say, in some scenarios, when recognizing the first image, multiple first orientations (such as multiple sight lines or face orientations) may be recognized. In this case, the first camera movement direction can be determined based on the multiple orientations. The above settings can enable the generation method provided in this application to meet the usage requirements in more scenarios, or enable the generation method provided in this application to have a wider range of adaptability. This application does not specifically limit the specific implementation method of determining the first camera movement direction based on multiple orientations.
[0036] In a possible implementation, determining the first camera movement direction based on the multiple directions includes: determining a first direction from the multiple directions based on preset priority information, and determining the first camera movement direction based on the first direction; or, determining a first direction from the multiple directions based on other display information in the display content of the first picture, and determining the first camera movement direction based on the first direction, wherein the other display information is information other than the multiple directions; or, determining the first camera movement direction based on the multiple directions and a preset artificial intelligence model.
[0037] That is to say, the electronic device understands the content of the first picture and may identify multiple directions. At this time, the electronic device can determine the aforementioned first direction of movement of the mirror based on the multiple directions according to preset logical rules. For example, the electronic device can use the highest priority direction (such as face orientation or line of sight direction) to determine the first direction of movement of the mirror based on preset priority information. For another example, the electronic device can also combine other information about the content of the picture (such as the type of subject, layout, proportion of the screen occupied, etc.) to determine the first direction of movement from multiple directions, and then determine the first direction of movement of the mirror based on the first direction. For another example, the multiple directions can also be used as input parameters to jointly determine the first direction of movement of the mirror according to preset rules or algorithms (such as AI models). This application does not impose any restrictions on the specific implementation method.
[0038] Optionally, the multiple directions include at least one of a face orientation, a line of sight direction, a movement direction, or an extension direction. For example, the multiple directions include multiple face orientations, or multiple line of sight directions, or at least one face orientation (line of sight direction) and the movement direction of at least one person or object. Alternatively, the multiple directions may include the extension direction of a water flow or a path.
[0039] In a possible implementation, the generation method further includes: obtaining a depth map of the first image; and generating the animation according to the first camera movement direction and the first image includes: generating the animation according to the first camera movement direction, the depth map and the first image.
[0040] Through the above settings, the generated animation has a 3D effect. Under different camera angles, the relative size and relative position (distance) of different objects in the picture can change, thereby enhancing the visual sense and immersiveness of the animation, improving the display effect of the animation, and thus improving the user's viewing experience.
[0041] Alternatively, the depth map can be acquired using a stereo camera or a time-of-flight (TOF) camera. That is, while the camera is acquiring the first image, it also acquires the depth map corresponding to the first image. In other words, the first image itself contains depth map information.
[0042] In a possible implementation, obtaining the depth map of the first image includes: performing monocular depth estimation on the first image to obtain the depth map.
[0043] In one possible implementation, generating the animation based on the first camera movement direction, the depth map and the first picture includes: performing depth layering on the depth map to identify a depth edge area; performing color filling on the depth edge area to obtain an edge background image; generating a 3D point cloud based on the first picture, the depth map and the edge background image; and generating the animation based on the first camera movement direction and the 3D point cloud.
[0044] Optionally, the 3D point cloud may include color information at different positions. In this case, the animation may be generated directly according to the first camera movement direction and the 3D point cloud.
[0045] Optionally, when the 3D point cloud does not include the color information, the animation may be generated according to the first camera movement direction, the 3D point cloud and the first image (color information).
[0046] In a possible implementation manner, the generating method further includes: displaying the animation.
[0047] Optionally, the animation here can be an animation effect (dynamic effect) rendered and displayed (temporarily output) in real time by the electronic device during display. The temporarily output animation effect can also be generated or saved as a video file or GIF dynamic image for playback the next time it is displayed, or for users to share on social media.
[0048] In a possible implementation, the animation is displayed as a dynamic wallpaper, thereby broadening the sources of dynamic wallpapers for the terminal.
[0049] In a possible implementation, the animation is a video or a graphics interchange format (GIF) picture, but is not limited thereto. For example, the animation may also be a dynamic picture in other formats.
[0050] In a second aspect, an animation generation device is provided, comprising: an acquisition unit for acquiring a first picture; a determination unit for determining a first camera movement direction based on the display content of the first picture; and a generation unit for generating the animation based on the first camera movement direction and the first picture, wherein the first camera movement direction is the camera movement direction corresponding to the animation.
[0051] In a possible implementation, the determining unit is specifically configured to: identify display content of the first image to obtain a first orientation of the target object; and determine the first camera movement direction according to the first orientation.
[0052] In a possible implementation, before identifying the display content of the first image to obtain a first pointing of the target object, the determination unit is specifically used to: identify the display content of the first image to obtain multiple objects; and determine the target object based on the multiple objects.
[0053] In one possible implementation, the determination unit is specifically used to: determine the subject among the multiple objects as the target object; or, determine the target object from the multiple objects based on preset priority information; or, determine the target object from the multiple objects based on the proportion of the screen occupied by the multiple objects; or, randomly select the multiple objects to obtain the target object.
[0054] In a possible implementation manner, the determining unit is specifically configured to: determine a direction along the first pointing direction as the first mirror movement direction, or determine a direction opposite to the first pointing direction as the first mirror movement direction.
[0055] In a possible implementation, the first direction includes a face orientation, a sight line direction, a moving direction, or an extending direction.
[0056] In a possible implementation, the first direction includes at least one of from front to back, from back to front, from left to right, from right to left, from top to bottom, and from bottom to top.
[0057] In a possible implementation, the determining unit is specifically configured to: identify display content of the first image to obtain multiple directions; and determine the first camera movement direction according to the multiple directions.
[0058] In a possible implementation, the acquisition unit is further configured to: acquire a depth map of the first image; and the generation unit is specifically configured to: generate the animation according to the first camera movement direction, the depth map, and the first image.
[0059] In a possible implementation manner, the acquisition unit is specifically configured to perform monocular depth estimation on the first image to acquire the depth map.
[0060] In one possible implementation, the generation unit is specifically used to: perform depth layering on the depth map to identify depth edge areas; perform color filling on the depth edge areas to obtain an edge background image; generate a 3D point cloud based on the first image, the depth map, and the edge background image; and generate the animation based on the first camera direction and the 3D point cloud.
[0061] In a possible implementation, the generating device further includes: a display unit, configured to display the animation.
[0062] In a possible implementation, the animation is displayed as a dynamic wallpaper.
[0063] In a possible implementation, the animation is a video or a picture in GIF format.
[0064] In a third aspect, an electronic device is provided, comprising: a memory storing instructions; and a processor, wherein when the instructions are executed by the processor, the electronic device executes the method provided by any possible implementation of the first aspect.
[0065] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is run on a computer, the computer is caused to execute the method provided by any possible implementation of the first aspect.
[0066] In a fifth aspect, a computer program product is provided, characterized in that it includes: computer program code, which, when the computer program code is run on a computer, enables the computer to execute the method provided by any possible implementation of the first aspect.
[0067] In a sixth aspect, a chip is provided, comprising: a processor for calling and running a computer program from a memory, so that an electronic device equipped with the chip executes the method provided by any possible implementation of the first aspect.
[0068] It can be understood that the animation generation device provided by the second aspect, the electronic device provided by the third aspect, the computer-readable storage medium provided by the fourth aspect, the computer program product provided by the fifth aspect, and the chip provided by the sixth aspect are all used to execute the method provided by the first aspect. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] FIG1 is a schematic diagram of a zoom lens provided in an embodiment of the present application.
[0070] FIG2 is a schematic diagram of a zoom lens provided in an embodiment of the present application.
[0071] FIG3 is a schematic diagram of a panning lens provided in an embodiment of the present application.
[0072] FIG4 is a schematic diagram of a panning lens provided in an embodiment of the present application.
[0073] FIG5 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0074] FIG6 is a block diagram of the software structure of the electronic device provided in an embodiment of the present application.
[0075] FIG7 is a schematic diagram of human-computer interaction in an application scenario provided by an embodiment of the present application.
[0076] FIG8 is a schematic diagram of the playback process of the animation generated in the scenario of FIG7 .
[0077] FIG9 is a schematic diagram of human-computer interaction in another application scenario provided by an embodiment of the present application.
[0078] FIG. 10 is a schematic diagram of an example of the playback process of the first picture of the animation generated in the scenario of FIG. 9 .
[0079] FIG11 is a schematic diagram of another example of the playback process of the first picture.
[0080] FIG12 is a schematic diagram of another example of the playback process of the first picture.
[0081] FIG. 13 is a schematic diagram of the playback process of the second picture of the animation generated in the scenario of FIG. 9 .
[0082] FIG. 14 is a diagram illustrating multiple examples of first orientations of a first picture.
[0083] FIG15 is a schematic diagram of another example of the playback process of the first picture.
[0084] FIG16 is a schematic diagram of a process of generating a 3D animation based on a first image.
[0085] FIG17 is a flow chart of the animation generation method provided in an embodiment of the present application.
[0086] FIG18 is a schematic block diagram of an animation generating device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0087] The technical solutions of this application will be described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, rather than all the embodiments.
[0088] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0089] The term "comprising" herein indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof. The terms "comprising", "including", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized. In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, unless otherwise stated, "multiple" means two or more.
[0090] The term "and / or" in this document simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Additionally, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0091] With the continuous development of terminal technology, more and more users are beginning to use pictures to record their lives. For some satisfactory pictures, such as landscape photos, users also want to create animations for playback and display. To meet this demand, technologies for generating animations from pictures (such as slideshows) have emerged. Users can combine multiple pictures in the local gallery to create an animation, and then share it on social platforms. In some cases, it is also possible to use a single picture to generate an animation, which can also be used as a dynamic wallpaper for the desktop or screensaver, thereby broadening the source of dynamic wallpapers for terminals.
[0092] Camera movement generally refers to motion shots. In video shooting, camera movement is an important narrative form that reflects creative agility and artistic value. Appropriate use of camera movement in videos helps portray characters, establish their characteristics, and establish the atmosphere of a scene. It also promotes the narrative. Different camera movements can control the rhythm of the narrative, creating different visual experiences and psychological implications. When generating animations from images, terminal devices can use camera movement to display the images to enhance the presentation of the images in the video. For example, images can be displayed from left to right, top to bottom, or zoomed in and out.
[0093] To improve video quality, the industry has developed various camera techniques. Common camera techniques include push-in, pull-out, pan, pan, and follow. These techniques can be used to zoom in or out, move the subject closer or further away, and translate or rotate, thereby enhancing the video's atmosphere and emotion.
[0094] To facilitate understanding of the technical solution of this application, the following first introduces the various camera movement types (methods) involved in the embodiments of this application in conjunction with the accompanying drawings. Figure 1 is a schematic diagram of a push shot, Figure 2 is a schematic diagram of a pull shot, Figure 3 is a schematic diagram of a pan shot, and Figure 4 is a schematic diagram of a pan shot.
[0095] A push shot is a shot that moves slowly or rapidly forward while the subject remains in the same position. The shot also changes from a distant shot to a full shot, a medium shot, a close-up, or even a close-up. This can be seen as a gradually decreasing field of view (FOV) or increasing focal length. The primary function of this shot is to highlight the subject, gradually focusing the viewer's attention and intensifying their visual experience, creating a state of scrutiny. For example, when photographing a person, as the camera moves forward, the person's proportion within the frame gradually increases, making them stand out more clearly.
[0096] Figure 1 shows a martial artist practicing Tai Chi, with the martial artist being the subject. The relative sizes of the martial artist in the different mobile phone images in Figure 1 indicate that as the camera moves from back to front, the area occupied by the martial artist in the entire frame gradually increases, transitioning from a distant shot to a medium shot and finally a close shot. This sequence of shots can be considered a push shot.
[0097] A pull shot moves in the opposite direction of a push shot. The camera moves from near to far, moving backwards away from the subject. The framing range expands, the subject shrinks, and the distance from the viewer gradually increases. The image shifts from a small number to a larger one, from a small part to a whole. In terms of shot size, it shifts from a close-up or a close-up to a medium shot, then to a panoramic or long shot. This can be seen as a gradually increasing field of view (FOV) or a decreasing focal length. The primary purpose of a pull shot is to convey the character's surroundings and enhance the image's atmosphere.
[0098] Figure 2 shows a martial artist practicing Tai Chi, with the martial artist being the subject. The relative sizes of the martial artist in the different mobile phone images in Figure 2 indicate that as the camera zooms out from front to back, the area occupied by the martial artist in the entire frame gradually decreases, transitioning from a close shot to a medium shot and then to a distant shot. This sequence of shots can be considered a zoom shot.
[0099] A dolly is similar to a push-pull camera, but the motion path differs. While a push-pull camera moves the lens forward and backward, a dolly is a unidirectional movement along a fixed path. "Move" can be understood as parallel movement, and the direction of movement can be horizontal, vertical, or at a certain angle (i.e., both horizontally and vertically). The movement path is usually a straight line.
[0100] Horizontal lens movement can expand the horizontal field of view. Moving the lens left and right horizontally can create a scene similar to how people move and explore in real life. For example, horizontal movement can be used to capture vast natural scenery. Vertical lens movement can expand the vertical field of view, such as when capturing tall subjects like buildings or mountains.
[0101] As shown in Figure 3, for panning, the camera can move from left to right, right to left, bottom to top, top to bottom, or a combination of two of the four directions. For example, when the camera moves from left to right, the warrior moves to the left relative to the entire phone screen, and the warrior's size within the screen remains unchanged. If other objects are present, the warrior's relative position to those objects remains unchanged.
[0102] A panning shot is similar to a dolly shot, but when filming, it involves rotating the phone or camera while stationary, moving the lens in an arc. For example, if you stand still, holding your phone and shooting from left to right, the phone's movement will be in an arc; the key point is that it remains stationary. This "panning" gradually reveals the scene in front of the camera, making the image more immersive. It's ideal for capturing continuous action or large-scale scenery, like a person's gaze scanning the subject in a certain direction. Panning can represent a subject's eyes, taking in everything around them. It plays a unique role in describing space and introducing an environment. Side-to-side panning is often used to introduce grand scenes, while vertical panning is often used to showcase the grandeur and precariousness of tall objects.
[0103] As shown in Figure 4, panning can be done from left to right, right to left, bottom to top, top to bottom, or any combination of two of these directions. For example, when panning from left to right, the warrior's position and size relative to other objects in the image will change.
[0104] Push shots, pull shots, tracking shots, and pan shots are the basic types of camera movements. In actual applications, camera movements may be composed of multiple movements described above. In other words, a camera movement may include multiple of the four types mentioned above. For example, it may include both push shots and tracking shots, or both pull shots and tracking shots, or both tracking shots and pan shots, or both push shots, tracking shots, and pan shots, etc., thereby achieving a better display effect for the captured video.
[0105] Table 1 shows the lens movement directions corresponding to different types of lens movements, where the movements include movement and rotation. As shown in Table 1, in the embodiment of the present application, the movement (movement) direction of the push lens can be understood as the lens gradually approaching the paper in the direction perpendicular to the paper, that is, from back to front. The movement (movement) direction of the pull lens can be understood as the lens gradually moving away from the paper in the direction perpendicular to the paper, that is, from front to back. The movement (movement) direction of the shift lens can be understood as the lens making horizontal, vertical or tilted movements in a plane parallel to the paper, that is, the movement direction of the shift lens includes from left to right, from right to left, from bottom to top and from top to bottom, and a combination of two of the above four directions. The movement (shaking) direction of the pan lens can be understood as the lens performing horizontal rotation, vertical rotation or tilted rotation with a fixed point away from the paper as the center of the circle, that is, the shaking direction of the pan lens includes from left to right, from right to left, from bottom to top and from top to bottom, and a combination of two of the above four directions.
[0106] Table 1: Lens movement directions corresponding to different types of camera movements
[0107] Currently, when a terminal device generates an animation with a camera movement display effect based on a picture, the terminal device can process the picture in the animation according to a pre-set camera movement direction. For example, if the pre-set camera movement direction is from left to right, the picture will be displayed in the generated animation according to the camera movement direction from left to right. For another example, if the pre-set camera movement direction is from upper right to lower left, the picture will be displayed in the generated animation according to the camera movement direction from upper right to lower left. However, multiple pictures are usually added to the animation, and the content displayed in different pictures is different, and may even differ greatly. The current technology can only display the multiple pictures according to a fixed and unique camera movement direction. The use of camera movement is not flexible enough, resulting in poor animation display effect.
[0108] In view of this, an embodiment of the present application provides a method for generating an animation, which can be, for example, a video or a picture in GIF format. According to the generation method, the movement direction of the camera (recorded as the first camera direction) can be first determined based on the display content of the picture (recorded as the first picture), and then the animation is generated based on the first camera direction and the first picture, that is, the camera is displayed on the first picture according to the first camera direction, thereby generating the animation. Since the movement direction of the camera is determined based on the display content of the picture, the movement direction of the camera and the display content of the picture can be associated and matched, and the camera can be flexibly used according to the content of the picture, thereby enhancing the visual sense and immersion of the animation, and being able to generate a high-quality animation with harmonious camera movement, thereby improving the display effect of the animation.
[0109] Furthermore, in embodiments of the present application, a preset algorithm (e.g., an artificial intelligence (AI) model) can be used to understand the displayed content of the first image and identify at least one directional information in the image (referred to as a first directional information). The first directional information can be, for example, the facial orientation, line of sight, or movement direction of a subject (e.g., a person or animal), or the extension direction of a path or river. The first directional information can include, for example, at least one of front to back, back to front, left to right, right to left, top to bottom, and bottom to top.
[0110] Then, the direction of camera movement (i.e., the first camera movement direction) is determined based on the first direction. For example, the direction along the first direction can be determined as the first camera movement direction, or the direction opposite to the first direction can be determined as the first camera movement direction, but it is not limited to this. Then, an animation can be generated according to the first camera movement direction and the first picture. Since the matching degree between the camera movement direction and the picture content is higher, the use of the camera movement is more harmonious and smooth, so that the quality of the generated animation is also higher and the picture sense is stronger, thereby improving the display effect of the animation, providing users with a coherent immersive experience, and improving the user's visual experience.
[0111] The animation generation method provided in the embodiments of the present application can be applied to electronic devices or a separate application that implements the animation generation method of the present application. For example, the generation method can be applied to electronic devices such as mobile phones, tablet computers, cameras, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). The embodiments of the present application do not impose any restrictions on the specific type of electronic device.
[0112] For example, FIG5 is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of the present application. As shown in FIG5, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0113] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0114] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0115] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0116] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0117] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0118] The I2C interface is a bidirectional synchronous serial bus consisting of a serial data line (SDA) and a serial clock line (SCL). The I2S interface can be used for audio communication. The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. The UART interface is a universal serial data bus used for asynchronous communication; this bus can be a bidirectional communication bus that converts the data to be transmitted between serial and parallel communication. The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display 194 and the camera 193; MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). The GPIO interface can be configured through software and can be configured as either a control signal or a data signal. The USB interface 130 is an interface that complies with USB standards and specifications, and can specifically be a Mini USB interface, a Micro USB interface, a USB Type-C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, or to transmit data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as AR devices.
[0119] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0120] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also provide power to the electronic device via the power management module 141.
[0121] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.
[0122] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0123] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0124] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0125] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0126] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0127] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with a network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time division-synchronous code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GNSS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0128] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0129] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a Micro LED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0130] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.
[0131] The ISP processes data fed back by camera 193. Camera 193 is used to capture still images or videos. The digital signal processor processes digital signals, including digital image signals and other digital signals. The video codec compresses or decompresses digital video.
[0132] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.
[0133] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, and at least one application required for a function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.). In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0134] The external memory interface 120 can be used to connect an external memory, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0135] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0136] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be set in the processor 110, or some functional modules of the audio module 170 can be set in the processor 110. The speaker 170A, also known as the "speaker", is used to convert audio electrical signals into sound signals. The receiver 170B, also known as the "earpiece", is used to convert audio electrical signals into sound signals. The microphone 170C, also known as the "microphone" or "microphone", is used to convert sound signals into electrical signals. The headphone jack 170D is used to connect wired headphones. The headphone jack 170D can be a USB interface 130, or it can be a 3.5mm open mobile terminal platform (OMTP) standard interface, or a standard interface of the Cellular Telecommunications Industry Association of the USA (CTIA).
[0137] The buttons 190 include a power button, a volume button, etc. The buttons 190 can be mechanical buttons. They can also be touch buttons. The electronic device 100 can receive button inputs and generate key signal inputs related to the user settings and function controls of the electronic device 100. The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts or for touch vibration feedback. The indicator 192 can be an indicator light that can be used to indicate the charging status, power changes, messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to and separated from the electronic device 100 by inserting it into the SIM card interface 195 or removing it from the SIM card interface 195. The electronic device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc.
[0138] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to illustrate the software structure of the electronic device 100.
[0139] FIG6 is a block diagram of the software structure of the electronic device 100 according to an embodiment of the present application.
[0140] As shown in Figure 6, a software system using a layered architecture is divided into several layers, each with distinct roles and divisions of labor. Layers communicate with each other via software interfaces. In some embodiments, the software system can be divided into four layers: from top to bottom, the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0141] The application layer may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and short message.
[0142] The application framework layer provides an application programming interface (API) and a programming framework for applications in the application layer. The application framework layer may include some predefined functions.
[0143] For example, the application framework layer includes the window manager, content provider, view system, telephony manager, resource manager, and notification manager.
[0144] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, and take screenshots.
[0145] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, and phone books.
[0146] The view system includes visual controls, such as controls for displaying text and images. The view system can be used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon can include a view for displaying text and a view for displaying images.
[0147] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (connected or hung up).
[0148] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, and video files.
[0149] The Notification Manager allows applications to display notification information in the status bar. It can be used to convey notification-type messages and can disappear automatically after a short stay without user interaction. For example, the Notification Manager is used to notify downloads and message reminders. The Notification Manager can also manage notifications that appear in the status bar at the top of the system in the form of icons or scrolling text, such as notifications from applications running in the background. The Notification Manager can also manage notifications that appear on the screen in the form of dialog windows, such as prompting text messages in the status bar, emitting alert sounds, vibrating electronic devices, and flashing indicator lights.
[0150] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.
[0151] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0152] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine performs functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0153] The system library can include multiple functional modules, such as a surface manager, a media library, a 3D graphics processing library (such as the open graphics library for embedded systems (OpenGL ES)) and a 2D graphics engine (such as the skia graphics library (SGL)).
[0154] The surface manager is used to manage the display subsystem and provide the fusion of 2D layers and 3D layers for multiple applications.
[0155] The media library supports playback and recording of multiple audio and video formats, as well as still image files. It supports a variety of audio and video codecs, such as MPEG4, H.264, Moving Picture Experts Group Audio Layer III (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR), Joint Photographic Experts Group (JPG), and Portable Network Graphics (PNG).
[0156] The 3D graphics processing library can be used to implement 3D graphics drawing, image rendering, compositing and layer processing.
[0157] A 2D graphics engine is a drawing engine for 2D drawings.
[0158] The kernel layer is the layer between hardware and software. The kernel layer can include driver modules such as display driver, camera driver, audio driver, and sensor driver.
[0159] In an embodiment of the present application, the user opens the camera application, and the camera application of the application layer in Figure 6 starts the shooting function and sends instructions to the kernel layer to mobilize the camera driver, sensor driver and display driver of the kernel layer, so that the electronic device 100 can start the camera to capture images. During the camera image capture process, light is transmitted to the image sensor through the camera, and the image sensor performs photoelectric conversion on the light signal and converts it into an image visible to the user's naked eye. The output image data is passed to the system library in Figure 6 in the form of a data stream. The three-dimensional graphics processing library and the image processing library implement drawing, image rendering, synthesis and layer processing, etc. to generate a display layer; the surface manager performs fusion processing on the display layer, etc., and passes it to the content provider, window manager and view system of the application framework layer to control the display of the display interface. Finally, the image is displayed in the image preview area of the camera application or the display screen of the electronic device 100.
[0160] Furthermore, when the user is satisfied with one or more photos they have taken, they can use them to generate an animation, such as a video or animated GIF. The image processing library takes the photos and interprets their content, determining the direction of camera movement. The animation is then generated by displaying the photos in that direction.
[0161] The surface manager acquires and processes animation data. Specifically, it can also provide layer synthesis services. A 3D graphics processing library or 2D graphics engine can draw and render the synthesized layers, and then send them to the display screen for display. For playback windows (such as video playback windows) displayed on the display screen, data refresh, layer synthesis, drawing, rendering, and display transmission can be performed based on the animation data to keep the display content of the playback window updated in real time.
[0162] Taking an electronic device (specifically a mobile phone) having the structure shown in Figures 5 and 6 as an example, the following will continue to introduce the animation generation method provided by the embodiment of the present application in combination with specific application scenarios. Figure 7 is a schematic diagram of human-computer interaction in an application scenario provided by an embodiment of the present application. As shown in Figure 7, as a possible application scenario, a user can generate an animation through a static picture (i.e., the first picture) and use the animation as a desktop wallpaper and / or lock screen wallpaper.
[0163] Specifically, as shown in part (a) of Figure 7, the user can click or touch the "Settings" button on the phone's desktop to enter the settings page. The settings page, shown in part (b) of Figure 7, includes options such as WLAN, Bluetooth, Mobile Network, Satellite Network, HyperTerminal, More Connections, Desktop and Personalization, and Display and Brightness. The user clicks on the "Desktop and Personalization" option. Furthermore, the Desktop and Personalization page, shown in part (c) of Figure 7, includes options such as Theme, Wallpaper, Screen Off Display, Magazine Lock Screen, and Icons. The user clicks on the "Wallpaper" option to enter the wallpaper settings page.
[0164] In the wallpaper page shown in part (d) of Figure 7, which includes options such as My Wallpapers, Recommended Wallpapers, and Natural Wallpapers, the user clicks the "Select from Gallery" option in My Wallpapers to enter the Gallery page. In the All Photos page shown in part (e) of Figure 7, the user selects a favorite photo to use as wallpaper. The user clicks on a selected photo to automatically enter the Set Wallpaper page. In the Set Wallpaper page shown in part (f) of Figure 7, the phone can convert the 2D static photo selected in the previous step into a 3D animation under the user's operation, and further set the animation as the desktop wallpaper and / or lock screen wallpaper for use.
[0165] Specifically, as shown in part (f) of Figure 7, the photo selected in the previous step is displayed on the wallpaper setting page. The user can choose to set the photo as an ordinary static wallpaper, or choose to set the photo as a dynamic wallpaper. When the user determines that the photo needs to be used as a dynamic wallpaper, the "dynamic wallpaper" option in the page can be checked, and then the "apply" button can be clicked. In response to the user's click operation, the mobile phone converts the photo into an animation according to the animation generation method provided in the embodiment of the present application, and then displays and applies it. For example, after the user clicks the "apply" button, the photo is automatically converted into an animation, and a new page (not shown in the figure) pops up. In the new page, the user can choose to set the generated animation as a dynamic wallpaper for the desktop, or as a dynamic wallpaper for the lock screen, or as a dynamic wallpaper for both the desktop and the lock screen.
[0166] The mobile phone in Figure 7 can set the selected static image to animation according to the following logic:
[0167] The phone receives a picture selected by the user, which could be, for example, a photo of a woman. Using a preset deep neural network model, the phone interprets the image's content and determines that the woman's face or gaze is facing from front to back (i.e., from inside the paper to outside, or from inside the screen to outside). Based on preset logic, the phone then determines the camera movement direction, which is the opposite of the face's direction—from back to front (i.e., from outside the paper to inside, or the direction the user is looking at the screen)—as the direction of the camera movement. This camera movement direction is the direction of a zoom shot. Therefore, a zoom shot can be applied to the image, i.e., the camera moves from back to front to display the image, thereby generating an animation.
[0168] Figure 8 is a schematic diagram of the playback process of the animation generated in the scene of Figure 7, showing three frames of the animation. Comparing the three frames shown in parts (a), (b), and (c) of Figure 8, it can be seen that as time progresses, the camera gradually approaches the woman in the picture, and the proportion of the woman in the frame gradually increases. The perspective changes from a distant shot to a medium shot, a close shot, and even a close-up.
[0169] According to the embodiments shown in Figures 7 and 8, since the movement direction of the camera is determined according to the facial orientation or line of sight of the characters in the picture, the use of the camera is more harmonious and smooth, and the camera can be flexibly used according to the content of the picture, thereby enhancing the visual sense and immersiveness of the animation, and being able to generate high-quality animations with harmonious camera movements, thereby improving the display effect of the animation and thereby improving the user experience.
[0170] It is worth mentioning that, in the embodiments shown in Figures 7 and 8, dynamic wallpaper is generated by only one picture, and in other implementations, dynamic wallpaper can also be generated by multiple pictures. At this time, for each picture, the mirror movement direction is determined according to the above logic, and the corresponding animation segment is generated according to the picture and its own corresponding mirror movement direction. Multiple pictures correspond to multiple animation segments, and the multiple animation segments are spliced together to obtain a complete animation, that is, a dynamic wallpaper generated by multiple pictures. It is easy to understand that the (corresponding) mirror movement directions used by different pictures can be different, which makes the application more flexible in the use of mirror movement, makes the mirror movement direction in the animation match the picture more highly, and thus improves the display effect of the animation.
[0171] FIG9 is a schematic diagram of human-computer interaction in another application scenario provided by an embodiment of the present application. As shown in FIG9 , as another possible application scenario, a user can generate an animation from a static image and share the animation on a social platform or send it to friends and family for viewing. The animation can be, for example, a video or a dynamic image in GIF format, but is not limited thereto.
[0172] As shown in Figure 9 (a), when a user is satisfied with the images captured with their camera and / or obtained online, they can create an animation based on these images. At this point, the user can click the "One-click Movie" or "Short Video Editing" button on the desktop. This article uses the "One-click Movie" button as an example. After clicking the "One-click Movie" button, the user enters the local gallery page.
[0173] As shown in part (b) of Figure 9, the user first selects multiple pictures on the local map library page and then clicks the "One-click to Create" button. In response to the click operation, the mobile phone generates an animation based on the multiple pictures selected by the user.
[0174] As shown in part (c) of Figure 9, once the animation is generated, the animation playback page automatically enters and automatically plays. If the user is satisfied with the animation, they can click the Export button to export and save the animation locally. The user can also share the animation on social platforms such as Weibo or WeChat Moments, or send it to friends and family via instant messaging software.
[0175] The mobile phone in Figure 9 can set the selected static image to animation according to the following logic:
[0176] The mobile phone obtains two pictures selected by the user. For example, the two pictures are the first picture shown in FIG10 and the second picture shown in FIG13 , wherein the first picture includes an animal doll and the second picture includes a dancer in the woods.
[0177] The mobile phone first understands the content of the first picture through a preset deep neural network model, determines that the animal doll's line of sight is towards the upper right (looking at the apple on the tree), and then, according to the preset logic, determines the direction opposite to the line of sight - from right to left as the camera movement direction. The camera movement direction can be the movement direction of a pan or tilt lens, so a pan and / or tilt lens can be applied to the picture, that is, the camera moves and / or shakes from right to left to display the first picture, thereby generating an animation clip corresponding to the first picture.
[0178] Figure 10 is a schematic diagram illustrating the playback process of an example of a first image, showing three frames of an animation clip corresponding to the first image. A panning and / or swaying camera is applied to the first image. Comparing the three frames shown in parts (a), (b), and (c) of Figure 10, it can be seen that, as time progresses, the camera gradually pans and / or sways from right to left, and the animal figure gradually moves from left to right within the frame.
[0179] Optionally, the animal doll's line of sight is towards the upper right, or the direction opposite to the line of sight - from top to bottom - can be determined as the camera movement direction according to preset logic. The first picture can be subjected to a panning and / or shaking shot, that is, the camera is moved and / or shaken from top to bottom to display the first picture, thereby generating an animation clip corresponding to the first picture.
[0180] Figure 11 is a schematic diagram illustrating another example of the playback process of a first image, showing three frames of an animation clip corresponding to the first image. A panning and / or swaying camera is applied to the first image. Comparing the three frames shown in parts (a), (b), and (c) of Figure 11, it can be seen that as time progresses, the camera gradually pans and / or sways from top to bottom, and the animal figure gradually moves from bottom to top within the frame.
[0181] Optionally, the sight line direction of the animal doll is from the lower left to the upper right, or the direction opposite to the sight line direction - from the upper right to the lower left - can be determined as the camera movement direction according to preset logic. The first picture can be subjected to a panning and / or shaking shot, that is, the camera is moved and / or shaken from the upper right to the lower left to display the first picture, thereby generating an animation clip corresponding to the first picture.
[0182] Figure 12 is a schematic diagram illustrating another example of the playback process of a first image, showing three frames of an animation clip corresponding to the first image. A panning and / or swaying camera is applied to the first image. Comparing the three frames shown in parts (a), (b), and (c) of Figure 12, it can be seen that, as time progresses, the camera gradually pans and / or sways from the upper right to the lower left, and the animal figure gradually moves from the lower left to the upper right within the frame.
[0183] That is, for the line of sight direction from the lower left to the upper right, the direction opposite (opposite) to the line of sight direction may include from right to left, from top to bottom, or from the upper right to the lower left (i.e., a combination of the two directions from right to left and from top to bottom), and this application does not limit this. The situations in other directions are similar and will not be repeated here.
[0184] Alternatively, the direction along the line of sight may be determined as the camera movement direction. For a line of sight from the lower left to the upper right, the direction along the line of sight may include a direction from left to right, from bottom to top, or from the lower left to the upper right (i.e., a combination of the two directions from left to right and from bottom to top). This application does not limit this. The situations in other directions are similar and will not be described in detail here.
[0185] The specific process of generating an animation clip from the second image is similar to that of the first image. The mobile phone continues to understand the content of the second image through a preset deep neural network model, determines that the dancer's line of sight in the image is from the right front to the left back (with the direction of the camera movement as a reference), and then, according to the preset logic, determines the direction opposite to the line of sight - from left to right - as the camera movement direction. The camera movement direction can be the movement direction of a pan or tilt lens, so a pan and / or tilt lens can be applied to the second image, that is, the camera moves and / or shakes from left to right to display the second image, thereby generating an animation clip corresponding to the second image.
[0186] Figure 13 is a schematic diagram of the playback process of the second image. Figure 13 shows three frames of the animation clip corresponding to the second image. A panning and / or swaying camera is applied to the second image. Comparing the three frames shown in parts (a), (b), and (c) of Figure 13, it can be seen that as time progresses, the camera gradually moves and / or pans from left to right, and the dancer gradually moves from right to left within the frame.
[0187] Optionally, the camera direction corresponding to the second picture can also be from back to front. In this case, a push-in shot can be applied to the second picture, that is, the camera gradually approaches the dancer in the picture, and the proportion of the dancer in the picture gradually increases.
[0188] Optionally, the camera movement direction corresponding to the second picture can also include from back to front and from left to right. In this case, not only a push shot is applied to the second picture, but also a pan shot and / or a shake shot can be applied to the second picture at the same time.
[0189] The animation clip corresponding to the first image and the animation clip corresponding to the second image are sequentially spliced together to create a complete animation. The more images the user selects, the more animation clips need to be spliced together, which means the duration of the complete animation is also longer. The user can change the order of multiple animation clips by adjusting the order of the images.
[0190] In an embodiment of the present application, the display content of the first image can be identified to obtain directional information (recorded as the first direction), and then the movement of the lens is determined based on the first direction. The first direction here is the directional information in the identified first image. The first direction can be, for example, the face orientation and sight direction of the shooting subject (such as a person or an animal) in the aforementioned embodiment, or the moving direction of a person, animal or object, or the extension direction of a path or river. The first direction can, for example, include at least one of from front to back, from back to front, from left to right, from right to left, from top to bottom, and from bottom to top. The first direction here is further illustrated with examples. Figure 14 is a schematic diagram of multiple examples of the first direction of the first image.
[0191] The first picture shown in part (a) of Figure 14 includes a road. The extension direction of the road in the picture is from the upper left to the lower right. Therefore, the extension direction can be determined as the first orientation, and the movement direction of the camera can be further determined based on the first orientation. Optionally, the movement direction of the camera can be opposite to (facing) the extension direction, for example, from right to left, from bottom to top, or from bottom right to upper left. Optionally, the movement direction of the camera can be along (along) the extension direction, for example, from left to right, from top to bottom, or from upper left to lower right.
[0192] The first picture shown in part (b) of Figure 14 includes a river, and the extension direction of the river in the picture is roughly from top to bottom, so this direction can be determined as the first direction, and the moving direction of the camera can be further determined based on the first direction. Optionally, the moving direction of the camera can be against the extension direction, for example, from bottom to top. Optionally, the moving direction of the camera can be along the extension direction, for example, from top to bottom. Optionally, the river has a left-right circling in the picture, and the moving direction of the camera can also include swinging or reciprocating in the left-right direction. That is to say, in an embodiment of the present application, the moving direction of the camera (i.e., the first camera direction) can include a specific direction, and can also include regular or irregular reciprocating swinging in a certain direction, and the present application does not limit this.
[0193] The first picture shown in part (c) of Figure 14 includes a standing boy with one hand on his waist and the fingers of the other hand pointing to the left. The boy's line of sight is from front to back (i.e., from the inside of the paper to the outside of the paper). In other words, the first picture shown in part (c) of Figure 14 includes multiple directions. At this time, the movement direction of the camera can be determined according to any one of the directions. For example, according to the direction of the fingers in the picture, the movement direction of the camera can be from left to right or from right to left. Alternatively, the movement direction of the camera can be from back to front (pushing the camera) or from front to back (pulling the camera) according to the direction of sight or the direction of face. Which direction to use to determine the movement direction of the camera can be determined according to a preset priority, for example, the direction of face or sight has the highest priority, or the movement direction of the camera can be determined in combination with multiple directions.
[0194] The first image shown in part (d) of Figure 14 shows a man running. Understanding the content of this first image reveals that the man's facial orientation, gaze direction, and movement direction (i.e., running direction) are all from left to right. Therefore, the first orientation can be determined to be from left to right, and the camera movement direction can be from left to right or from right to left.
[0195] The first image shown in part (e) of Figure 14 includes a boy riding a bicycle from right to left, and the boy's line of sight is from front to back (i.e., from the inside of the paper to the outside of the paper). In other words, the first image shown in part (e) of Figure 14 also includes multiple directions. In this embodiment, based on the priority of the directions, the face direction or the line of sight direction can be preferentially used as the first direction (to determine the movement direction of the camera). In this case, the movement direction of the camera can be from back to front (push the camera) or from front to back (pull the camera).
[0196] The first picture shown in part (f) of Figure 14 includes an animal doll and a balloon. By understanding the content of the first picture, it can be known that the balloon is floating from bottom to top, and the animal doll's line of sight is looking towards the upper right direction at the floating balloon. In other words, the first picture shown in part (f) of Figure 14 also includes multiple directions. In this embodiment, the moving direction of the camera can be determined based on information such as multiple directions (i.e., the floating direction of the balloon and the line of sight of the animal doll). For example, the moving direction of the camera can be from top to bottom or from bottom to top, etc. In addition, the line of sight direction can also be used to determine the moving direction of the camera. The choice of the specific method depends on the internal implementation of the mobile phone or the pre-set rules, and this application does not make any special restrictions on this.
[0197] FIG15 is a schematic diagram illustrating another example of the playback process of the first image. Understanding the content of the first image shown in FIG15 reveals that the image includes two children, whose faces and gazes are facing different directions. Both children are running simultaneously and in the same direction, from left to right. The camera's movement direction can be determined based on the running direction. For example, the camera's movement direction can be from right to left or from left to right.
[0198] Figure 15 shows three frames of the animation clip corresponding to the first image. A panning and / or swaying camera can be applied to the first image. Comparing the three frames shown in parts (a), (b), and (c) of Figure 15, it can be seen that as time progresses, the camera gradually pans and / or sways from left to right, and the child gradually moves from right to left within the image.
[0199] In recent years, with the continuous development of image processing technology, it has become possible to generate 3D animations from ordinary 2D static images. 3D images have better display effects and are gaining popularity among more and more users. Furthermore, in the embodiment shown in FIG15 , a depth map of the first image can also be obtained, and the animation can be generated based on the determined camera movement direction, the first image, and the depth map of the first image.
[0200] Through the above settings, the generated animation has a 3D effect. Under different camera angles, the relative size and relative position (distance) of different objects in the picture can change, thereby enhancing the visual sense and immersiveness of the animation, improving the display effect of the animation, and thus improving the user's viewing experience.
[0201] Optionally, monocular depth estimation may be performed on the first image using a deep neural network model to obtain the depth map.
[0202] Alternatively, the depth map can be acquired using a stereo camera or a time-of-flight (TOF) camera. That is, while the camera is acquiring the first image, it also acquires the depth map corresponding to the first image. In other words, the first image itself contains depth map information.
[0203] Figure 16 is a flow chart of generating a 3D animation based on a first image. As shown in Figure 16, as a specific example, the 3D animation can be generated by following the steps below:
[0204] Step 1: Obtain the first image, use the deep neural network model to perform monocular depth estimation on the first image to obtain a depth map, then segment the depth discontinuous areas, that is, perform depth layering on the depth map, and then identify the depth edge areas.
[0205] For example, performing depth layering on the depth map may include dividing the depth map into a foreground and a background, and identifying a depth edge region may include identifying a hole region in the background.
[0206] Step 2: Fill the depth edge region with color to obtain an edge background image. For example, the depth edge region (e.g., a hole region in the background) can be filled using a depth filling neural network model and a color filling neural network model to obtain an edge background image.
[0207] Step 3: Generate a 3D point cloud based on the first image, the depth map, and the edge background map. The 3D point cloud here is also commonly referred to as a 3D mesh or a 3D feature point cloud.
[0208] Step 4: The image understanding neural network model is used to understand the display content of the first image to determine the visual orientation (i.e., the direction of sight), and then the camera movement direction is determined based on the visual orientation.
[0209] Step 5: Generate an animation based on the camera movement direction, 3D point cloud, and first image determined in the previous step. For example, the camera movement control module controls the camera to move in the direction determined in the previous step. During this movement, as the camera's perspective changes, the image is rendered from the new perspective, generating a 3D animation from the 2D static image.
[0210] Step 6: The generated 3D animation is displayed on the electronic device, and the user watches the 3D animation. If the animation effect is satisfactory, the 3D animation can be set as a dynamic wallpaper, shared on social platforms, or sent to friends and family for viewing.
[0211] Based on the aforementioned multiple embodiments, the present application further provides an animation generation method, which can be applied to the electronic device 100 provided in the aforementioned embodiments, or to a chip or chip system within the electronic device 100. The following is a method embodiment provided by the present application, which corresponds to the above embodiments. For example, in the following description, the first camera movement direction corresponds to or is equivalent to the camera movement direction mentioned above.
[0212] FIG17 is a flow chart of a method 200 for generating an animation according to an embodiment of the present application. As shown in FIG17 , the method 200 includes the following steps:
[0213] Step 210: The electronic device obtains a first picture.
[0214] In step 220 , the electronic device determines a first camera movement direction according to display content of the first image.
[0215] Step 230: The electronic device generates an animation according to the first camera movement direction and the first picture, wherein the first camera movement direction is a camera movement direction corresponding to the animation.
[0216] Specifically, the electronic device obtains the first picture selected by the user, and uses a preset AI model (such as a neural network model) to understand the display content of the first picture, and determines the first camera direction according to the result of the understanding. Then, an animation is generated based on the first camera direction and the first picture. The first camera direction here refers to the direction of camera movement, and the camera can display the first picture according to the first camera direction, thereby generating the animation. Since the camera movement direction is determined based on the display content of the picture, the camera movement direction and the display content of the picture can be correlated and matched, and the camera movement can be flexibly used according to the content of the picture, thereby enhancing the picture sense and immersion of the animation, and generating a high-quality animation with harmonious camera movement, thereby improving the display effect of the animation.
[0217] The embodiments of the present application do not limit the specific method of "determining the first camera direction based on the image content." For example, the subject in the first image can be identified, and the first camera direction can be determined based on the type of the subject. The subject here can be, for example, a human, an animal, a plant, a mountain, a river, a toy, jewelry, fruit, a building, furniture, the sun, the moon, a vehicle, etc., but is not limited to these. The first camera direction can also be determined based on a more specific type of identified subject. For example, animals can be categorized as cats, dogs, lions, elephants, or birds, and the first camera direction corresponding to different types of animals may be different. And / or, the first camera direction can also be determined based on factors such as the number of subjects, their position in the image, and the proportion of the entire image. In addition, the first camera direction can also be determined based on factors such as the layout and color tone of the image. Furthermore, if the subject is a person, the first camera direction can also be determined based on factors such as the person's expression, movement, and gestures.
[0218] Optionally, when an animation is generated by using multiple images, the “animation” in “the first camera movement direction is the camera movement direction corresponding to the animation” can also be understood or replaced by an animation clip. For example, the multiple images include a first image and a second image, the animation clip corresponding to the first image is the first animation clip, the animation clip corresponding to the second image is the second animation clip, the first camera movement direction is the camera movement direction corresponding to the first animation clip, and the second camera movement direction is the camera movement direction corresponding to the second animation clip. The first animation clip and the second animation clip are spliced together to form a complete animation.
[0219] Furthermore, in the embodiment of the present application, in step 220, the electronic device determines the first camera movement direction according to the display content of the first image, which specifically includes:
[0220] In step 221 , the electronic device identifies the display content of the first image to obtain a first direction of the target object.
[0221] In step 222 , the electronic device determines the first mirror movement direction according to the first orientation.
[0222] Specifically, the display content of the first image can be identified to obtain a first orientation, and then the movement of the camera can be determined based on the first orientation. The first orientation here is the orientation information within the identified first image. The first orientation can be, for example, the facial orientation or sight direction of the subject (e.g., a person or animal) in the aforementioned embodiment, or the direction of movement of a person, animal, or object, or the extension direction of a path or river. The first orientation can, for example, include at least one of front to back, back to front, left to right, right to left, top to bottom, and bottom to top.
[0223] Furthermore, in an embodiment of the present application, the first direction may be the direction information of the target object in the image, that is, the display content of the first image may be identified to obtain the first direction of the target object, and then the movement of the camera is determined according to the first direction. The target object here may be any object in the first image, such as a person (such as the aforementioned lady, dancer, or little boy), an animal (such as the aforementioned animal doll), a plant (such as a tree), a scene (such as the aforementioned river, etc.), a building (such as a house, a bridge, etc.), a road, a toy, jewelry, fruit, furniture, the sun, the moon, a vehicle, etc., but is not limited thereto. The target object may be the subject of the photograph or not, for example, the target object may be a river or a path in the shooting background, and the present application does not limit this.
[0224] Optionally, the first image may contain multiple objects, and the target object may be one of the multiple objects. For example, the display content of the first image may be recognized to obtain multiple objects, and then the target object may be determined based on the multiple objects.
[0225] That is to say, in some scenarios, when recognizing the first image, multiple objects (such as multiple people) may be recognized. In this case, the target object can be determined based on the multiple objects first, and then the first direction corresponding to the target object can be identified. The above settings can enable the generation method provided by this application to meet the usage requirements in more scenarios, or enable the generation method provided by this application to have a wider scope of application. This application does not specifically limit the specific implementation method of determining the target object based on multiple objects.
[0226] Optionally, determining the target object based on the multiple objects includes: determining the subject among the multiple objects as the target object; or, determining the target object from the multiple objects based on preset priority information; or, determining the target object from the multiple objects based on the proportion of the screen occupied by the multiple objects; or, randomly selecting the multiple objects to obtain the target object.
[0227] That is to say, the electronic device understands the content of the first picture and may identify multiple objects. At this time, the electronic device can determine the aforementioned target object based on the multiple objects according to preset logical rules. Optionally, the electronic device can determine the subject (such as a person or an animal) among the multiple objects as the target object. Optionally, the electronic device can determine the object with the highest priority as the target object based on preset priority information. For example, different objects such as people, rivers, trees, buildings, roads, etc. have different priorities, and the person with the highest priority can be determined as the target object. Optionally, the electronic device can also determine the object with the largest proportion of the picture among the multiple objects as the target object. For example, if there are multiple people in the picture, the person with the largest proportion of the picture can be determined as the target object. Optionally, the multiple objects can also be randomly extracted, and the extracted object can be determined as the target object.
[0228] Optionally, the electronic device determines the first mirror movement direction according to the first direction, and may determine the direction along the first direction as the first mirror movement direction, or may determine the direction opposite to the first direction as the first mirror movement direction, but is not limited thereto.
[0229] By determining the direction along or against the first pointing direction as the first camera movement direction, the method is simple and easy to implement, and can better correlate and match the first camera movement direction with the image content, thereby enhancing the visual sense and immersion of the animation and improving the user's visual experience.
[0230] Optionally, in other implementations, in step 220, the electronic device determines the first camera movement direction according to the display content of the first image, specifically including:
[0231] The electronic device identifies the display content of the first image to obtain multiple directions;
[0232] The electronic device determines the aforementioned first mirror movement direction according to the multiple directions.
[0233] That is to say, the electronic device understands the content of the first picture and may identify multiple directions. At this time, the electronic device can determine the aforementioned first direction of movement of the mirror based on the multiple directions according to preset logical rules. For example, the electronic device can use the highest priority direction (such as face orientation or line of sight direction) to determine the first direction of movement of the mirror based on preset priority information. For another example, the electronic device can also combine other information about the content of the picture (such as the type of subject, layout, proportion of the screen occupied, etc.) to determine the first direction of movement from multiple directions, and then determine the first direction of movement of the mirror based on the first direction. For another example, the multiple directions can also be used as input parameters to jointly determine the first direction of movement of the mirror according to preset rules or algorithms (such as AI models). This application does not impose any restrictions on the specific implementation method.
[0234] Furthermore, in the embodiment of the present application, the generation method 200 further includes: the electronic device obtains a depth map of the first image. Step 230, the electronic device generates an animation according to the first camera movement direction and the first image, specifically including:
[0235] The electronic device generates an animation according to the first camera movement direction, the depth map and the first picture.
[0236] Through the above settings, the generated animation has a 3D effect. Under different camera angles, the relative size and relative position (distance) of different objects in the picture can change, thereby enhancing the visual sense and immersiveness of the animation, improving the display effect of the animation, and thus improving the user's viewing experience.
[0237] Optionally, monocular depth estimation may be performed on the first image using a deep neural network model to obtain the depth map.
[0238] Alternatively, the depth map can be acquired using a stereo camera or a time-of-flight (TOF) camera. That is, while the camera is acquiring the first image, it also acquires the depth map corresponding to the first image. In other words, the first image itself contains depth map information.
[0239] Furthermore, the electronic device generates an animation according to the first camera movement direction, the depth map, and the first image, specifically including:
[0240] The electronic device performs depth layering on the depth map to identify depth edge areas.
[0241] The electronic device fills the depth edge area with color to obtain an edge background image.
[0242] The electronic device generates a 3D point cloud based on the first image, the depth map and the edge background map.
[0243] The electronic device generates the animation according to the first camera movement direction and the 3D point cloud.
[0244] Optionally, the 3D point cloud may include color information at different positions. In this case, the animation may be generated directly according to the first camera movement direction and the 3D point cloud.
[0245] Optionally, when the 3D point cloud does not include the color information, the animation may be generated according to the first camera movement direction, the 3D point cloud and the first image (color information).
[0246] Optionally, the generating method 200 further includes: displaying the animation.
[0247] Optionally, the animation here can be an animation effect (dynamic effect) rendered and displayed (temporarily output) in real time by the electronic device during display. The temporarily output animation effect can also be generated or saved as a video file or GIF dynamic image for playback the next time it is displayed, or for users to share on social media.
[0248] Optionally, the animation is displayed as a dynamic wallpaper, thereby broadening the sources of dynamic wallpapers for the terminal.
[0249] Optionally, the animation is a video or a GIF image.
[0250] Figure 18 is a schematic block diagram of an animation generation device 300 provided in an embodiment of the present application. The generation device 300 can be an electronic device in an embodiment of the present application (e.g., the electronic device 100 or mobile phone in the aforementioned embodiment), or a processor or chip within an electronic device. As shown in Figure 18, the generation device 300 includes an acquisition unit 310, a determination unit 320, a generation unit 330, and a display unit 340.
[0251] The acquiring unit 310 is configured to acquire a first image.
[0252] The determining unit 320 is configured to determine a first camera movement direction according to the display content of the first image.
[0253] The generating unit 330 is configured to generate an animation according to the first camera movement direction and the first picture, wherein the first camera movement direction is a camera movement direction corresponding to the animation.
[0254] Optionally, the determining unit 320 is specifically configured to: identify display content of the first image to obtain a first orientation of the target object; and determine the first camera movement direction according to the first orientation.
[0255] Optionally, before identifying the display content of the first image to obtain the first pointing of the target object, the determination unit 320 is specifically used to: identify the display content of the first image to obtain multiple objects; and determine the target object based on the multiple objects.
[0256] Optionally, the determination unit 320 is specifically used to: determine the subject among the multiple objects as the target object; or, determine the target object from the multiple objects based on preset priority information; or, determine the target object from the multiple objects based on the proportion of the screen occupied by the multiple objects; or, randomly select the multiple objects to obtain the target object.
[0257] Optionally, the determining unit 320 is specifically configured to: determine a direction along the first pointing direction as the first mirror movement direction, or determine a direction opposite to the first pointing direction as the first mirror movement direction.
[0258] Optionally, the first direction includes face orientation, sight direction, movement direction or extension direction.
[0259] Optionally, the first direction includes at least one of from front to back, from back to front, from left to right, from right to left, from top to bottom, and from bottom to top.
[0260] Optionally, the determining unit 320 is specifically configured to: identify display content of the first image to obtain multiple directions; and determine the first camera movement direction according to the multiple directions.
[0261] Optionally, the determination unit 320 is specifically used to: determine a first direction from the multiple directions according to preset priority information, and determine the first camera movement direction according to the first direction; or, determine a first direction from the multiple directions according to other display information in the display content of the first picture, and determine the first camera movement direction according to the first direction, wherein the other display information is information other than the multiple directions; or, determine the first camera movement direction based on the multiple directions and a preset artificial intelligence model.
[0262] Optionally, the acquisition unit 310 is further configured to: acquire a depth map of the first image. The generation unit 330 is specifically configured to: generate the animation according to the first camera movement direction, the depth map and the first image.
[0263] Optionally, the acquiring unit 310 is specifically configured to perform monocular depth estimation on the first image to acquire the depth map.
[0264] Optionally, the generation unit 330 is specifically used to: perform depth layering on the depth map to identify depth edge areas; perform color filling on the depth edge areas to obtain an edge background image; generate a 3D point cloud based on the first image, the depth map and the edge background image; and generate the animation based on the first camera direction and the 3D point cloud.
[0265] Optionally, the 3D point cloud may include color information at different positions. In this case, the animation may be generated directly according to the first camera movement direction and the 3D point cloud.
[0266] Optionally, when the 3D point cloud does not include the color information, the animation may be generated according to the first camera movement direction, the 3D point cloud and the first image (color information).
[0267] Optionally, the generating device 300 further includes: a display unit 340, configured to display the animation.
[0268] Optionally, the animation is displayed as a dynamic wallpaper.
[0269] Optionally, the animation is a video or a picture in GIF format.
[0270] The generating device 300 corresponds to the electronic device in the aforementioned method embodiment, the acquiring unit 310 is used to execute the aforementioned step 210, the determining unit 320 is used to execute the aforementioned step 220 (for example, including step 221 and step 222), and the generating unit 330 is used to execute the aforementioned step 230.
[0271] The present application also provides an electronic device, including: a memory storing instructions; and a processor, which, when the instructions are executed by the processor, causes the electronic device to perform the method or steps provided in any of the aforementioned embodiments. The electronic device may be, for example, the electronic device 100 or the mobile phone in the aforementioned embodiments.
[0272] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is run on a computer, the computer is caused to execute the method or steps provided in any of the aforementioned embodiments.
[0273] An embodiment of the present application further provides a computer program product, comprising: computer program code, which, when executed on a computer, enables the computer to execute the method or steps provided in any of the aforementioned embodiments.
[0274] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store computer-executable instructions, and when the device is running, the processor can execute the computer-executable instructions stored in the memory to enable the chip to execute the methods in the above-mentioned method embodiments.
[0275] Among them, the electronic device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0276] Through the description of the above implementation methods, technical personnel in the relevant field can understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0277] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0278] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and other media that can store program code.
[0279] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for generating an animation, characterized in that: include: Get the first picture; Determining a first camera movement direction according to display content of the first picture; The animation is generated according to the first camera movement direction and the first picture, wherein the first camera movement direction is a moving direction of the camera corresponding to the animation.
2. The generation method according to claim 1, characterized in that: The determining the first camera movement direction according to the display content of the first picture includes: Identify the display content of the first picture to obtain a first direction of the target object; The first camera movement direction is determined according to the first orientation.
3. The generation method according to claim 2, characterized in that: Before identifying the display content of the first picture to obtain a first direction of the target object, the method further includes: Recognize the display content of the first picture to obtain multiple objects; The target object is determined according to the multiple objects.
4. The generation method according to claim 3, characterized in that: The determining the target object according to the multiple objects includes: determining a subject among the multiple objects as the target object; or, determining a target object from the plurality of objects according to preset priority information; or, determining a target object from the multiple objects according to the screen ratios occupied by the multiple objects; or, The multiple objects are randomly selected to obtain the target object.
5. The generation method according to any one of claims 2 to 4, characterized in that: The determining the first mirror movement direction according to the first orientation includes: The direction along the first pointing direction is determined as the first camera movement direction, or, A direction opposite to the first pointing direction is determined as the first camera movement direction.
6. The generation method according to any one of claims 2 to 5, characterized in that: The first orientation includes a face orientation, a sight line direction, a moving direction or an extending direction.
7. The generation method according to any one of claims 2 to 6, characterized in that: The first direction includes at least one of from front to back, from back to front, from left to right, from right to left, from top to bottom, and from bottom to top.
8. The generation method according to any one of claims 1 to 7, characterized in that: The generating method further comprises: Obtaining a depth map of the first image; Generating the animation according to the first camera movement direction and the first picture includes: The animation is generated according to the first camera movement direction, the depth map and the first picture.
9. The generation method according to any one of claims 1 to 8, characterized in that: The generating method further comprises: The animation is displayed.
10. The generation method according to any one of claims 1 to 9, characterized in that: The animation is displayed as a dynamic wallpaper.
11. The generation method according to any one of claims 1 to 10, characterized in that: The animation is a video or a picture in GIF format.
12. An animation generating device, characterized in that: include: An acquiring unit, configured to acquire a first picture; a determining unit, configured to determine a first camera movement direction according to display content of the first picture; A generating unit is used to generate the animation according to the first camera movement direction and the first picture, wherein the first camera movement direction is a moving direction of the camera corresponding to the animation.
13. The generating device according to claim 12, characterized in that: The determining unit is specifically used for: Identify the display content of the first picture to obtain a first direction of the target object; The first camera movement direction is determined according to the first orientation.
14. The generating device according to claim 13, characterized in that: Before identifying the display content of the first picture to obtain the first direction of the target object, the determining unit is specifically used to: Recognize the display content of the first picture to obtain multiple objects; The target object is determined according to the multiple objects.
15. The generating device according to claim 14, characterized in that: The determining unit is specifically used for: determining a subject among the multiple objects as the target object; or, determining a target object from the plurality of objects according to preset priority information; or, Determining a target object from the multiple objects according to the proportions of the screen occupied by the multiple objects; or, The multiple objects are randomly selected to obtain the target object.
16. The generating device according to any one of claims 13 to 15, characterized in that: The determining unit is specifically used for: The direction along the first pointing direction is determined as the first camera movement direction, or, A direction opposite to the first pointing direction is determined as the first camera movement direction.
17. The generating device according to any one of claims 13 to 16, characterized in that: The first orientation includes a face orientation, a sight line direction, a moving direction or an extending direction.
18. The generating device according to any one of claims 13 to 17, characterized in that: The first direction includes at least one of from front to back, from back to front, from left to right, from right to left, from top to bottom, and from bottom to top.
19. The generating device according to any one of claims 12 to 18, characterized in that: The acquisition unit is also used for: Obtaining a depth map of the first image; The generating unit is specifically used for: The animation is generated according to the first camera movement direction, the depth map and the first picture.
20. The generating device according to any one of claims 12 to 19, characterized in that: The generating device further comprises: A display unit is used to display the animation.
21. The generating device according to any one of claims 12 to 20, characterized in that: The animation is displayed as a dynamic wallpaper.
22. The generating device according to any one of claims 12 to 21, characterized in that: The animation is a video or a picture in GIF format.
23. An electronic device, characterized in that: include: A memory storing instructions; The processor, when the instructions are executed by the processor, causes the electronic device to execute the method as described in any one of claims 1-11.
24. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed on a computer, the computer is enabled to execute the method according to any one of claims 1 to 11.
25. A computer program product, characterized in that include: A computer program code, when the computer program code is run on a computer, causes the computer to execute the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Animation generation method and device and electronic equipment
CN119919544A
Picture dynamic processing method, device and terminal equipment
CN103473799A
Fully automatic dynamic articulated model calibration
CN103608844A
Lens animation generating method and system
CN105976416A
Picture display method and device, electronic equipment and storage medium
CN116501227A