Image generation method and device

By dividing the shooting area in the shooting interface and allowing the user to drag the image to the editing area to generate the target image, the problem of cumbersome operation in the prior art is solved, and efficient image generation is achieved.

CN120416656APending Publication Date: 2025-08-01VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510536187.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, users need to manually intercept and splice part of the areas of multiple images to generate target images, which is cumbersome and inefficient.

Method used

An image generation method and device are provided, which simplifies the operation process by dividing a molecule shooting area in the shooting interface, receiving user input to display the captured image, and allowing the user to drag the image to the image editing area to generate a target image.

Benefits of technology

Eliminates the user to manually crop and stitch the images, and generates target images directly in the shooting and editing areas, improving operational efficiency and simplifying user operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416656A_ABST
    Figure CN120416656A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and device, and belongs to the technical field of electronic equipment. The method comprises the steps that first input of a user to at least one sub-shooting area in a shooting area of a shooting interface is received, each sub-shooting area displays a part of multimedia files of a to-be-shot multimedia file, and the part of multimedia files, displayed in each sub-shooting area, of the to-be-shot multimedia file do not have an overlapping area; in response to the first input, at least one shot image is displayed, one shot image is displayed in one sub-shooting area, and the shot image in one sub-shooting area is the shot image of a part of the multimedia file displayed in the sub-shooting area; receiving a second input that the user drags at least one shot image to an image editing area of the shooting interface; and in response to the second input, generating a target image based on the at least one shot image in the image editing area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of electronic devices, and particularly relates to an image generation method and apparatus. Background Art

[0002] With the rapid development of electronic device technology, people's demand for images captured by electronic devices is also increasing. To obtain the images required by users, users may need to perform various forms of editing on the captured images. For example, if a user only wants a part of the image area of each of multiple images, the user needs to manually intercept each image and then splice the intercepted image areas to obtain the new image required by the user, which is a cumbersome operation. Summary of the Invention

[0003] The purpose of the embodiments of this application is to provide an image generation method and apparatus, which can simplify user operations, quickly intercept partial areas of multiple images, and obtain the target image required by the user based on the intercepted image areas.

[0004] In a first aspect, the embodiments of this application provide an image generation method, which includes:

[0005] Receiving a first input from a user for at least one sub - shooting area in the shooting area of a shooting interface, where each sub - shooting area respectively displays a part of a multimedia file to be shot, and there is no overlapping area for the parts of the multimedia file to be shot displayed in each sub - shooting area;

[0006] In response to the first input, displaying at least one shooting image, where one shooting image is displayed in one sub - shooting area, and the shooting image displayed in one sub - shooting area is an image obtained by shooting the part of the multimedia file displayed in the sub - shooting area;

[0007] Receiving a second input from the user to drag the at least one shooting image to an image editing area of the shooting interface;

[0008] In response to the second input, generating a target image based on the at least one shooting image in the image editing area.

[0009] In a second aspect, the embodiments of this application provide an image generation apparatus, which includes:

[0010] A receiving module, configured to receive a first input from a user for at least one sub - shooting area in the shooting area of a shooting interface, where each sub - shooting area respectively displays a part of a multimedia file to be shot, and there is no overlapping area for the parts of the multimedia file to be shot displayed in each sub - shooting area;

[0011] A display module, configured to display at least one captured image in response to the first input, where one captured image is displayed within one sub-capturing area, and the captured image displayed within one sub-capturing area is an image obtained by capturing a part of the multimedia file displayed within the sub-capturing area;

[0012] The receiving module is further configured to receive a second input from the user to drag the at least one captured image to an image editing area of the capturing interface;

[0013] An image generation module, configured to generate a target image based on the at least one captured image within the image editing area in response to the second input.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0016] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0017] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0018] In an embodiment of the present application, when partial multimedia files of a multimedia file to be captured are respectively displayed in each sub - capture area of the capture area on the capture interface, and there is no overlapping area among the partial multimedia files of the multimedia file to be captured displayed in each sub - capture area, by responding to a first input from the user on at least one sub - capture area, captured images of the partial multimedia files displayed therein can be respectively shown in at least one sub - capture area where the first input is executed. In this way, partial image areas of the multimedia file to be captured can be directly obtained without the user cropping the multimedia file to be captured, simplifying the user operation. Then, when a second input from the user to drag at least one captured image to the image editing area of the capture interface is received, a target image can be directly generated based on at least one captured image in the image editing area, without the user manually splicing at least one captured image in the image editing area, simplifying the user operation and improving the efficiency of generating the target image. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a schematic flowchart of an image generation method provided by some embodiments of the present application;

[0020] Figure 2 is one of the schematic diagrams of a multimedia file to be captured provided by some embodiments of the present application;

[0021] Figure 3 is one of the schematic diagrams of a capture interface provided by some embodiments of the present application;

[0022] Figure 4 is another schematic diagram of a capture interface provided by some embodiments of the present application;

[0023] Figure 5 is another schematic diagram of a multimedia file to be captured provided by some embodiments of the present application;

[0024] Figure 6 is another schematic diagram of a capture interface provided by some embodiments of the present application;

[0025] Figure 7 is another schematic diagram of a capture interface provided by some embodiments of the present application;

[0026] Figure 8 is another schematic diagram of a capture interface provided by some embodiments of the present application;

[0027] Figure 9 is another schematic diagram of a capture interface provided by some embodiments of the present application;

[0028] Figure 10 is another schematic diagram of a capture interface provided by some embodiments of the present application;

[0029] Figure 11 It is the eighth schematic diagram of the shooting interface provided by some embodiments of the present application;

[0030] Figure 12 It is the schematic structural diagram of the image generation device shown by some embodiments of the present application;

[0031] Figure 13 It is the schematic structural diagram of the electronic device shown by some embodiments of the present application;

[0032] Figure 14 It is the schematic hardware structure diagram of the electronic device shown by some embodiments of the present application. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope protected by the present application.

[0034] The terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or N. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / " generally means that the associated objects before and after are in an "or" relationship.

[0035] The terms used in the implementation manner part of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application.

[0036] Next, the terms related to the embodiments of the present invention will be explained.

[0037] The identifier in the present application is used to indicate information such as text, symbols, images, etc., and can use controls or other containers as the carrier for displaying information, including but not limited to text identifiers, symbol identifiers, and image identifiers.

[0038] The technical solution of the embodiment of the present application can be applied to scenarios where it is necessary to intercept partial image areas of images corresponding to multiple multimedia files, and then stitch the image areas of each intercepted image together to form a new image required by the user. For example, when a user is doing exercises on the first and second pages of a workbook assigned by a teacher, he or she gets one question at the top of the first page and two questions in the middle of the second page wrong. At this time, the user needs to intercept the wrong question on the first page and the two wrong questions on the second page and organize them into a collection of wrong questions. For another example, a user wants to shoot a video in a park. The video may involve several parts: buildings, a boy and a girl playing, landscapes, and a boy playing by himself. At this time, the user only wants to stitch together the content of a boy and a girl playing and the landscapes in the video into one image.

[0039] The image generation method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0040] Figure 1 This is a flow chart of an image generation method provided in an embodiment of the present application. The execution subject of the image generation method can be an electronic device, which can be but is not limited to a personal computer (PC), a smart phone, a tablet computer or a personal digital assistant (PDA).

[0041] It should be noted that the electronic device can be a single-screen electronic device or a folding-screen electronic device, which is not limited in the embodiments of this application.

[0042] like Figure 1 As shown, the image generation method provided in the embodiment of the present application may include steps 110 to 140.

[0043] Step 110: Receive a first input from a user on at least one sub-shooting area in a shooting area of a shooting interface.

[0044] The shooting interface may be an interface for shooting multimedia files to be shot, for example, the shooting interface may be a shooting interface of a camera application.

[0045] It should be noted that the shooting interface is different from the shooting interface of the existing camera application. The shooting interface may include a shooting area for shooting images and an image editing area for editing images.

[0046] It should be noted that the aforementioned shooting interface may be a component service provided by the operating system to support the display of the shooting interface within the operating system-managed interface. The component service may also support the display of the shooting area and image editing area within the shooting interface. The operating system may specifically be Linux. The component service may be built using, but is not limited to, the Spring Framework, SpringBoot Framework, or ToySpring Framework.

[0047] The multimedia file to be photographed may be a file to be photographed, for example, a multimedia file to be photographed, for example, an image or video to be photographed.

[0048] In one example, reference Figure 2 When the user is doing the exercises on the first and second pages of the exercise book assigned by the teacher, he or she makes mistakes on the first question on the first page 21 and the second and third questions on the second page 22. At this time, the user needs to extract the first question on the first page 21 and the second and third questions on the second page 22 and organize them into a collection of wrong questions. If the user plans to first take a picture of the first page 21 of the exercise book to form a picture, and then take a picture of the second page 22 of the exercise book to form a second picture, the multimedia file to be taken will be the image corresponding to the first page 21 of the exercise book and the image corresponding to the second page 22 of the exercise book.

[0049] In another example, a user wants to shoot a video in a park. The video may involve several parts: buildings, a boy and a girl playing, landscape, and a boy playing alone. At this time, the user only wants to splice the content of the boy and the girl playing and the landscape in the video into an image, and the multimedia file to be shot is the video to be shot.

[0050] The shooting area may be an area for shooting the multimedia file to be shot. The shooting area may include at least one sub-shooting area, each sub-shooting area displays a portion of the multimedia file to be shot, and the portions of the multimedia file to be shot displayed in each sub-shooting area do not overlap.

[0051] That is, when capturing a multimedia file, the file can be divided into multiple parts, and then a portion of the multimedia file can be displayed in each sub-capturing area. For example, if the multimedia file is an image, the image can be divided into multiple image areas, each of which can be displayed in a different sub-capturing area. If the multimedia file is a video, the video can be divided into multiple video frames, each of which can be displayed in a different sub-capturing area.

[0052] It should be noted that the number of sub - shooting areas in the shooting area is determined based on the multimedia file to be shot. That is, when shooting each multimedia file to be shot, the number of sub - shooting areas in the shooting area is different. For example, for a multimedia file to be shot that is an image full of text information, the number of sub - shooting areas can be determined based on the number of paragraphs of the text information in the image. For example, the number of sub - shooting areas is the same as the number of paragraphs of the text information in the image. In the case where the multimedia file to be shot is an image containing tables, pictures, and text information, the number of sub - shooting areas can be determined based on the number of tables, the number of pictures, and the number of paragraphs of text information. For example, the number of sub - shooting areas can be the sum of the number of tables, the number of pictures, and the number of paragraphs of text information. In the case where the multimedia file to be shot is a video, the number of sub - shooting areas can be determined based on the number of video frames included in the video. For example, the number of sub - shooting areas can be the same as the number of video frames included in the video.

[0053] Continuing to refer to the first example above, taking the multimedia file to be shot as an image of the first page 21 as an example, as Figure 2 shown, the first page 21 includes the title "Exercises for the First Unit" and two questions. Among them, the first question involves the text information part 211 of the question and the attached drawing 212 of the question. The second question involves the text information part 213 of the question. The second page 22 includes three questions, among which are the text information part 221 of the first question, the text information part 222 of the second question, and the text information part 223 of the third question in the second page 22. Then when shooting the first page 21, the title and the two questions on the first page 21 can be divided according to the pictures and the paragraphs of text information, and are divided into a total of 4 parts. One part is the text information part "Exercises for the First Unit" 210 of the title, one part is the text information part 211 of the first question, another part is the attached drawing part 212 of the first question, and there is also one part which is the text information part 213 of the second question.

[0054] Taking the electronic device as a single - screen electronic device as an example, when shooting the first page 21, referring to Figure 3 , Figure 3It is a schematic diagram of the display of the shooting interface. In the shooting interface 30, there is a shooting area 31. According to the division in the first page 21, the shooting area 31 can also be divided into 4 parts, that is, the shooting area 31 includes 4 sub-shooting areas, namely sub-shooting area 310, sub-shooting area 311, sub-shooting area 312, and sub-shooting area 313. Then, in the sub-shooting area 310, a preview image of the text information part of the title "Exercises for the First Unit" is displayed; in the sub-shooting area 311, a preview image of the text information part of the first question is displayed; in the sub-shooting area 312, a preview image of the drawing part of the first question is displayed; and in the sub-shooting area 313, a preview image of the text information part of the second question is displayed.

[0055] In another example, taking a video to be shot as an example of the multimedia file to be shot, it should be noted that at this time, the user has not finished shooting the video, and it is only the shooting preview stage of the video. Therefore, this video can be called a preview video, and the video frames in this preview video are preview video frames.

[0056] This preview video contains a total of 4 preview video frames. Among them, the first preview video frame is a video frame of a building, the second preview video frame is a video frame of a group photo of two people, the third preview video frame is a video frame of a scenery, and the fourth preview video frame is a video frame of a single boy. Then, when shooting the above video, refer to Figure 4 , Figure 4 It is a schematic diagram of the display of the shooting interface. In the shooting interface 40, there is a shooting area 41. According to the division of the video to be shot, the shooting area 41 can also be divided into 4 parts, that is, the shooting area 41 includes 4 sub-shooting areas, namely sub-shooting area 411, sub-shooting area 412, sub-shooting area 413, and sub-shooting area 414. Then, in the sub-shooting area 411, a preview video frame image of a building is displayed; in the sub-shooting area 412, a preview video frame image of a group photo of two people is displayed; in the sub-shooting area 413, a preview video frame image of a scenery is displayed; and in the sub-shooting area 414, a preview video frame image of a single boy is displayed.

[0057] The first input may be an input by the user to at least one sub - shooting area in the shooting area of the shooting interface. The above - mentioned first input is used to display the shooting images of partial multimedia files within at least one sub - shooting area, and the first input may be a first operation. Exemplarily, the above - mentioned first input includes, but is not limited to: the touch input of the user to at least one sub - shooting area in the shooting area of the shooting interface through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a slide gesture, a drag gesture, a pressure - recognition gesture, a long - press gesture, an area - change gesture, a double - press gesture, and a double - click gesture; the click input in the embodiments of the present application may be a single - click input, a double - click input, or a click input of any number of times, etc., and may also be a long - press input or a short - press input. For example, the above - mentioned first input may be: the touch input of the user to at least one sub - shooting area in the shooting area of the shooting interface. For example, the above - mentioned first input may be: the click input of the user to at least one sub - shooting area in the shooting area of the shooting interface.

[0058] Continue to refer to Figure 3 , if the user wants the first question on the first page 21, the user can click on the sub - shooting area that displays the preview image of the first question, that is, the user can click on sub - shooting area 311 and sub - shooting area 312.

[0059] In another example, continue to refer to Figure 4 , if the user wants the video - frame content of a two - person group photo and the video - frame content of a landscape, the user can click on the sub - shooting area 412 that displays the preview video - frame image of the two - person group photo, and the sub - shooting area 413 that displays the preview video - frame image of the landscape.

[0060] It should be noted that when the multimedia file to be shot is a video, the preview images of the video frames displayed in the sub - shooting area may be arranged according to the preview shooting order of the video, that is, the video - frame content previewed first is displayed in the front - located sub - shooting area, and the video - frame content previewed later is displayed in the rear - located sub - shooting area.

[0061] It should be noted that it is also possible to directly shoot a video in the shooting area first, and then display each video frame of the shot video in the sub - shooting areas of the shooting area. In this way, compared with displaying preview video frames in the sub - shooting area, the efficiency of the video frames displayed in the sub - shooting area when directly displaying the shot video frames in the sub - shooting area is higher. However, whether to shoot the video first or directly display the preview video frames in the sub - shooting area can be set according to user needs and is not limited in the embodiments of the present application.

[0062] In addition, it should be noted that in the case where a video has been directly and successfully shot within the shooting area in advance, the shot video is stored in the memory rather than on the hardware disk of the electronic device.

[0063] Step 120: In response to a first input, display at least one captured image.

[0064] Among them, one captured image is displayed within one sub-shooting area, and the captured image displayed within one sub-shooting area is the captured image of the partial multimedia file displayed within that sub-shooting area.

[0065] In some embodiments of the present application, after the user performs a first input on a certain sub-shooting area, in response to this first input, the preview image within the sub-shooting area where the first input is performed can be shot to obtain the captured image of this preview image, and this captured image is displayed within that sub-shooting area.

[0066] Continue to refer to Figure 3 , after the user clicks on sub-shooting area 311 and sub-shooting area 312, the preview image of the text information part of the first question displayed within sub-shooting area 311 can be shot to obtain the captured image of the text information part of the first question, and this captured image is also displayed within sub-shooting area 311, and the preview image of the attached drawing part of the first question displayed within sub-shooting area 312 is also shot to obtain the captured image of the attached drawing part of the first question, and this captured image is also displayed within sub-shooting area 312.

[0067] Continue to refer to Figure 4 , after the user clicks on sub-shooting area 412 and sub-shooting area 413, the preview images within sub-shooting area 412 and sub-shooting area 413 can be shot respectively to obtain a captured image of two people taking a group photo, and this captured image is displayed within sub-shooting area 412, and a captured image of the scenery is obtained and this captured image is displayed within sub-shooting area 413.

[0068] It should be noted that after shooting the preview image displayed within at least one sub-shooting area to obtain a captured image, this captured image can be saved in the memory of the electronic device. In this way, the captured image does not occupy the hardware disk space of the electronic device. In addition, since only the preview image within the sub-shooting area where the first input is performed is shot rather than the entire multimedia file to be shot, in the memory of the electronic device, it is not necessary to occupy the memory of the entire multimedia file to be shot, and only the memory corresponding to the captured image of the preview image within the sub-shooting area where the first input is performed needs to be occupied.

[0069] It should be noted that it is known which part of the multimedia file each sub - shooting area corresponds to in the multimedia file to be shot. Therefore, after the user performs the first input on a certain sub - shooting area, according to the range of the part of the multimedia file to be shot corresponding to the sub - shooting area, the part of the multimedia file to be shot corresponding to the sub - shooting area can be shot to obtain a captured image of the part of the multimedia file to be shot corresponding to the sub - shooting area.

[0070] Continuing to refer to the above example, as Figure 5 shown, Figure 5 is a schematic diagram of the multimedia file to be shot. Taking the first page in Figure 2 as an example, the shooting area of the first page 21 is from pixel point (a, b) to pixel point (c, d), and the picture 212 part of the first page 21 is from pixel point (a1, b1) to pixel point (c1, d1). For example, the shooting area of the first page 21 is from pixel point (0, 0) to pixel point (200, 300), and the picture 212 part of the first page 21 is from pixel point (70, 80) to pixel point (170, 130).

[0071] Normally, shooting the first page 21 will obtain a photo of 200 * 300 pixel points. If the user wants to intercept the part of the attached drawing of the first question in the first page 21, after the user clicks on the sub - shooting area corresponding to the part of the attached drawing of the first question, the part of the attached drawing of the first question in the first page 21 can be intercepted, that is, the part from pixel point (70, 80) to pixel point (170, 130) in the first page 21 can be intercepted to obtain a captured image of this part.

[0072] In this way, when saving the captured image, only the image of the part from pixel point (70, 80) to pixel point (170, 130) is saved, and the entire image of the first page 21 from pixel point (0, 0) to pixel point (200, 300) is not saved.

[0073] Step 130: Receive a second input from the user to drag at least one captured image to the image editing area of the shooting interface.

[0074] Among them, the image editing area can be an area for editing images, such as Figure 3 the area 32 in Figure 4 and the area 42 in

[0075] It should be noted that in the case where the electronic device is a single-screen electronic device, the electronic device can be split-screened. One part of the split screen displays the shooting area of the shooting interface, and the other part of the split screen displays the image editing area of the shooting interface. For example, the electronic device can be divided into upper and lower half split screens. The upper half split screen displays the shooting area of the shooting interface, and the lower half split screen displays the image editing area of the shooting interface, as Figure 3 and Figure 4 shown.

[0076] In the case where the electronic device is a foldable screen electronic device, for example, the electronic device is a left and right dual-screen electronic device, the shooting area of the shooting interface can be displayed on one screen of the electronic device, and the image editing area of the shooting interface can be displayed on the other screen. For example, the shooting area of the shooting interface can be displayed on the left screen, and the image editing area of the shooting interface can be displayed on the right screen, so as to increase the user's operation space and avoid user misoperation.

[0077] The second input can be the input of the user dragging at least one captured image to the image editing area of the shooting interface. The above second input is used to generate a target image based on at least one captured image in the image editing area. The second input can be a second operation. Exemplarily, the above second input includes but is not limited to: the touch input of the user dragging at least one captured image to the image editing area of the shooting interface through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements, and the embodiments of the present invention do not make limitations. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, etc., and can also be a long press input or a short press input. For example, the above second input can be: the touch input of the user dragging at least one captured image to the image editing area of the shooting interface. For example, the above second input can be: the drag input of the user dragging at least one captured image to the image editing area of the shooting interface.

[0078] Step 140, in response to the second input, generate a target image based on at least one captured image in the image editing area.

[0079] Among them, the target image can be an image generated based on at least one captured image dragged into the image editing area.

[0080] In some embodiments of the present application, by responding to the user dragging at least one captured image into the image editing area, a target image can be generated based on at least one captured image in the image editing area.

[0081] It should be noted that after the target image is generated, the target image can be saved in the album application, that is, saved in the hardware disk of the electronic device.

[0082] In some embodiments of the present application, the image editing area includes image generation controls, such as Figure 3 the "Save" control 33 in Figure 4 and the "Save" control 43 in

[0083] To improve the generation efficiency of the target image, step 140 may specifically include:

[0084] In response to the second input, at least one captured image is displayed in the image editing area;

[0085] Receive the user's third input to the image generation control;

[0086] In response to the third input, at least one captured image in the image editing area is spliced to generate a target image.

[0087] Among them, the third input may be the user's input to the image generation control in the image editing area. The above third input is used to splice at least one captured image in the image editing area to generate a target image. The third input may be a third operation. Exemplarily, the above third input includes, but is not limited to: the user's touch input to the image generation control in the image editing area through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements, and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, and may also be a long press input or a short press input. For example, the above third input may be: the user's touch input to the image generation control in the image editing area. For example, the above third input may be: the user's click input to the image generation control in the image editing area.

[0088] In some embodiments of the present application, after the user drags at least one captured image to the image editing area, at least one captured image can be displayed in the image editing area, and then in response to the user's third input to the image generation control in the image editing area, at least one captured image displayed in the image editing area can be spliced according to the display position of at least one captured image in the image editing area, and the target image can be obtained. Specifically, when generating the target image, the captured image with a forward display position in the image editing area is displayed at a forward position in the target image.

[0089] Continuing to refer to the above example, as Figure 6 shown, the user can drag the captured image 61 of the text information part of the first question displayed in the sub-capturing area 311 and the captured image 62 of the attached drawing of the first question displayed in the sub-capturing area 312 to the image editing area 32, and then the captured image 61 and the captured image 62 can be displayed in the image editing area 32. Then the user clicks Figure 6 the "Save" control 63 in the image editing area 32 in

[0090] Continuing to refer to the above example, as Figure 7 shown, the user can drag the captured image 71 of the group photo of two people displayed in the sub-capturing area 412 and the captured image 72 of the scenery displayed in the sub-capturing area 413 to the image editing area 42, and then the captured image 71 and the captured image 72 can be displayed in the image editing area 32. Then the user clicks Figure 7 the "Save" control 73 in the image editing area 32 in

[0091] It should be noted that after dragging at least one captured image from the capturing area to the image editing area, the user can also adjust at least one captured image displayed in the captured image, including but not limited to adjusting the display position and size of at least one captured image. In this way, at least one captured image can be spliced in sequence according to the display position of at least one captured image adjusted by the user to obtain the target image.

[0092] In the embodiments of the present application, in response to the second input, at least one captured image is displayed in the image editing area, and then in response to the third input of the user to the image generation control, at least one captured image in the image editing area can be directly spliced to generate the target image, without the user manually splicing at least one generated captured image, which simplifies the user operation and improves the generation efficiency of the target image.

[0093] In some implementations of the present application, the image editing area may further include an insertion control for inserting the description information of the captured image, and the insertion control may include at least one of an audio control and a text input box. As Figure 6 the audio control 65 and the text input box 66 in

[0094] Before receiving the third input of the user, the above-mentioned method may further include:

[0095] Receive a fourth input from the user to drag the insertion control into the first preset range area of the first captured image;

[0096] In response to the fourth input, display the insertion control within the first preset range area;

[0097] Receive a fifth input from the user for the insertion control;

[0098] In response to the fifth input, display an information identifier for indicating the description information of the first captured image;

[0099] The stitching of at least one captured image within the image editing area to obtain a target image may include:

[0100] Stitch other captured images and the first captured image with the information identifier to obtain a target image.

[0101] Wherein, the first captured image may be at least one of the at least one captured image displayed within the image editing area, that is, the at least one captured image may include the first captured image. For example, the first captured image may be the captured image 61 as in Figure 6 .

[0102] The first preset range area may be the preset range area of the first captured image. For example, it may be within a certain distance below the first captured image, such as within 1 millimeter below the first captured image.

[0103] The fourth input may be the input of the user dragging the insertion control into the first preset range area of the first captured image. The above fourth input is used to display the insertion control within the first preset range area, and the fourth input may be a fourth operation. Exemplarily, the above fourth input includes but is not limited to: the touch input of the user dragging the insertion control into the first preset range area of the first captured image through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, and may also be a long press input or a short press input. For example, the above fourth input may be: the touch input of the user dragging the insertion control into the first preset range area of the first captured image. For example, the above fourth input may be: the drag input of the user dragging the insertion control into the first preset range area of the first captured image.

[0104] The fifth input may be the user's input to the insertion control. The above-mentioned fifth input is used to display an information identifier for indicating the description information of the first captured image. The fifth input may be a fifth operation. Exemplarily, the above-mentioned fifth input includes, but is not limited to: the user's touch input to the insertion control through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above-mentioned fifth input may be: the user's touch input to the insertion control. For example, the above-mentioned fifth input may be: the user's long press input or filling input to the insertion control.

[0105] The information identifier may be an identifier for indicating the description information of the first captured image.

[0106] Other captured images may be captured images in the image editing area other than the first captured image. For example, when the first captured image is the captured image 61 as shown in Figure 6 the other captured image is the captured image 62.

[0107] In some embodiments of the present application, by responding to a fourth input in which the user drags the insertion control to a first preset range area of the first captured image, the insertion control can be displayed within the first preset range area, and then by responding to the user's fifth input to the insertion control, an information identifier for indicating the description information of the first captured image can be displayed. In this way, when at least one captured image is stitched, other captured images and the first captured image with the information identifier can be stitched to obtain a target image.

[0108] In the embodiments of the present application, by adding corresponding description information to the captured images in the image editing area where the insertion control can be located, the obtained target image can include the description information of the captured images, so that the user can intuitively view the information described by the captured images and can intuitively know the content described by the captured images according to the description information.

[0109] In some embodiments of the present application, in order to enhance the diversity of the description information of the first captured image, the display of the information identifier for indicating the description information of the first captured image may include:

[0110] When the insertion control includes an audio control, displaying a voice progress bar for indicating the description information of the first captured image;

[0111] When the inserted control includes a text input box, first text information for indicating description information of the first captured image is displayed within the text input box.

[0112] Wherein, the voice progress bar can be used to indicate the voice information describing the first captured image.

[0113] The first text information can be text information for describing the first captured image.

[0114] In some embodiments of the present application, when the inserted control includes an audio control, in response to a fifth input of the user to the inserted control, a voice progress bar can be displayed. The voice progress bar can be used to indicate the voice information describing the first captured image. Then, when stitching the captured images in the image editing area, the voice progress bar is stitched together.

[0115] Continuing to refer to Figure 6 , taking the inserted control as the audio control 65 and the first captured image as the captured image 61 as an example, the user can drag the audio control 65 below the captured image 61. In this way, as Figure 8 shown, the audio control 65 is displayed below the captured image 61.

[0116] Taking the fifth input as a long press on the audio control 65 displayed below the captured image 61 as an example, when the user long presses the audio control 65 displayed below the captured image 61, voice information describing the captured image 61 can be recorded, such as "the first question on page 1 of the exercise book". Then, a voice progress bar 81 is displayed. The voice progress bar 81 can be used to indicate the voice information "the first question on page 1 of the exercise book". Then, after the user clicks the "save" control 82, the captured image 62 and the captured image 61 with the voice progress bar 81 can be stitched together to obtain the target image 83.

[0117] In some embodiments of the present application, when the inserted control includes a text input box, in response to a fifth input of the user to the text input box, such as inputting the first text information for describing the first captured image in the text input, the first text information for indicating the description information of the first captured image can be directly displayed within the text input box.

[0118] Continuing to refer to Figure 6 , taking the inserted control as the text input box 66 and the first captured image as the captured image 61 as an example, the user can drag the text input box 66 below the captured image 61. In this way, as Figure 9 shown, the text input box 66 is displayed below the captured image 61.

[0119] Taking the fifth input as an example of filling in the text input box 66 displayed below the captured image 61, the user can input the first text information "This is the first question on page 1 of the math workbook" in the text input box 66 displayed below the captured image 61. Then, the first text information "the first question on page 1 of the workbook" can be displayed in the input box 66. After that, when the user clicks the "Save" control 92, the captured image 62 and the captured image 61 of the input box 66 with the first text information "the first question on page 1 of the workbook" can be spliced together to obtain the target image 93.

[0120] In an embodiment of the present application, when the inserted control includes an audio control, a voice progress bar for indicating the description information of the first captured image can be displayed, so that voice description information can be added to the first captured image. When the inserted control includes a text input box, the first text information for indicating the description information of the first captured image can be displayed in the text input box, so that text description information can be added to the first captured image. In this way, according to the user's needs, different forms of description information can be added to the first captured image, improving the diversity of the description information of the first captured image.

[0121] In some embodiments of the present application, in order to improve the flexibility of the user to view the description information of the first captured image, when the inserted control includes an audio control, after the voice progress bar for indicating the description information of the first captured image is displayed, the above-mentioned method may further include:

[0122] Receiving a seventh input from the user to the voice progress bar;

[0123] In response to the seventh input, converting the voice information into third text information;

[0124] Displaying the third text information within a first preset range area;

[0125] The splicing of the other captured image and the first captured image with the information identifier to obtain the target image may include:

[0126] Splicing the other captured image and the first captured image with the third text information to obtain the target image.

[0127] Among them, the seventh input may be the user's input to the voice progress bar. The above-mentioned seventh input is used to convert voice information into third text information, and the seventh input may be a seventh operation. Exemplarily, the above-mentioned seventh input includes but is not limited to: the user's touch input to the voice progress bar through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a slide gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above-mentioned seventh input may be: the user's touch input to the voice progress bar. For example, the above-mentioned seventh input may be: the user's long press input to the voice progress bar.

[0128] The third text information may be the text information converted from the voice information indicated by the voice progress bar.

[0129] In some embodiments of the present application, if the user wants to print the final target image, but it is inconvenient to view the description information of the first captured image on the voice progress bar, the voice information indicated by the voice progress bar can be converted into third text information by responding to the seventh input of the user to the voice progress bar, and then the third text information can be directly displayed in the first preset range area. In this way, when the captured images in the image editing area are stitched together, the converted third text information is stitched together.

[0130] Continue to refer to Figure 8 , after the voice progress bar 81 is displayed, the user can long press the voice progress bar 81, and then as Figure 10 shown, convert the voice information "the first question on page 1 of the exercise book" indicated by the voice progress bar 81 into the third text information "the first question on page 1 of the exercise book" 100. Then the user clicks the "Save" control 101, and the captured image 62 and the captured image 61 with the third text information "the first question on page 1 of the exercise book" 100 can be stitched together to obtain the target image 102.

[0131] In the embodiments of the present application, according to the user's needs, by responding to the seventh input to the voice progress bar, the voice information indicated by the voice progress bar can be converted into third text information. In this way, the user can intuitively know the content described by the first captured image according to the third text information. In this way, when it is inconvenient for the user to listen to the voice information, the voice information can be converted into third text information, which improves the flexibility of the user to view the description information of the first captured image.

[0132] In some embodiments of the present application, in order to improve the viewing accuracy of the description information of the first captured image, when the inserted control includes an audio control, after the voice progress bar indicating the description information of the first captured image is displayed, the method described above may further include:

[0133] Receiving an eighth input from the user to the voice progress bar;

[0134] In response to the eighth input, playing the voice information describing the first captured image indicated by the voice progress bar.

[0135] Wherein, the eighth input may be an input from the user to the voice progress bar, and the above-mentioned eighth input is used to play the voice information describing the first captured image indicated by the voice progress bar, and the eighth input may be an eighth operation. Exemplarily, the above-mentioned eighth input includes, but is not limited to: a touch input of the user to the voice progress bar through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements, and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above-mentioned eighth input may be: a touch input of the user to the voice progress bar. For example, the above-mentioned eighth input may be: a click input of the user to the voice progress bar.

[0136] In some embodiments of the present application, when the user needs to play the voice information describing the captured image, by clicking on the voice progress bar, the voice information can be played.

[0137] Continuing to refer to Figure 8 , after the voice progress bar 81 is displayed, the user can click on the voice progress bar 81 to play the voice information "Question 1 on Page 1 of the workbook".

[0138] In the embodiments of the present application, by responding to the eighth input of the user to the voice progress bar, the voice information describing the first captured image indicated by the voice progress bar can be played, so that the user can intuitively understand the voice information describing the first captured image, and avoid the problem of incorrect text information conversion due to accents and other issues when converting the voice information into text information, thereby improving the viewing accuracy of the description information of the first captured image.

[0139] In some embodiments of the present application, in order to improve the flexibility of adding text description information to the captured images in the image editing area, after displaying at least one captured image in response to the first input, the method described above may further include:

[0140] Identify the text information in the second captured image to obtain the second text information of the second captured image;

[0141] The splicing of at least one captured image in the image editing area to obtain a target image may include:

[0142] Receive a sixth input from the user to drag the second text information into a second preset range area of the third captured image;

[0143] In response to the sixth input, splice the other captured images and the third captured image with the second text information to obtain a target image.

[0144] Among them, the second captured image may be any one of the at least one captured image displayed in at least one captured area.

[0145] The second text information may be the text information obtained after identifying the text information in the second captured image.

[0146] The third captured image may be any one of the at least one captured image in the image editing area.

[0147] The second preset range area may be a preset range area of the third captured image, for example, it may be within a certain distance below the third captured image, such as within 1 millimeter below the third captured image.

[0148] The sixth input may be an input where the user drags the second text information into a second preset range area of the third captured image. The above sixth input is used to splice other captured images and the third captured image with the second text information to obtain a target image. The sixth input may be a sixth operation. Exemplarily, the above sixth input includes but is not limited to: a touch input where the user drags the second text information into the second preset range area of the third captured image through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a slide gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above sixth input may be: a touch input where the user drags the second text information into the second preset range area of the third captured image. For example, the above sixth input may be: a drag input where the user drags the second text information into the second preset range area of the third captured image.

[0149] The other captured images may be the captured images in the image editing area except the third captured image.

[0150] In some embodiments of the present application, when adding text description information to the captured images in the image editing area, in addition to inputting text information in the text input box as described above, the text information in the captured images in the capture area may also be recognized to obtain the second text information of the captured image. Then, in response to the operation where the user drags the captured image to the third captured image in the image editing area, the recognized second text information of the captured image may be used as the description information of the third captured image and displayed in the second preset range area of the third captured image. Then, when splicing at least one captured image, the other captured images and the third captured image with the second text information may be spliced to obtain a target image.

[0151] Continue to refer to Figure 6 , taking the captured image 60 with the text information "Exercises for the First Unit" displayed in the sub-capture area 310 as the second captured image and the captured image 61 in the image editing area as the third captured image as an example. After recognizing the text information "Exercises for the First Unit" in the captured image 60 displayed in the sub-capture area 310, the text information "Exercises for the First Unit" is obtained. Then, as Figure 11As shown, the user can drag the captured image 60 in the sub - capture area 310 and drag it below the captured image 61. In this way, the text information "Exercises for the First Unit" can be displayed below the captured image 61. Then, after the user clicks the "Save" control 111, the captured image 62 and the captured image 61 with the text information "Exercises for the First Unit" can be spliced together to obtain the target image 112.

[0152] In the embodiments of the present application, by recognizing the text information in the captured images within the capture area to obtain second text information, in response to an eighth input in which the user drags the second text information to a third captured image, text description information can be added to the third captured image, thus enhancing the flexibility of adding text description information to the captured images within the image editing area.

[0153] The image generation method provided by the embodiments of the present application may have an image generation device as the execution subject. In the embodiments of the present application, taking the image generation device executing the image generation method as an example, the image generation device provided by the embodiments of the present application is described.

[0154] Figure 12 It is a schematic structural diagram of an image generation device shown according to an exemplary embodiment. As Figure 12 shown, the image generation device 1200 may include:

[0155] A receiving module 1210, configured to receive a first input from the user for at least one sub - capture area in the capture area of the capture interface. Each sub - capture area respectively displays a partial multimedia file of the multimedia file to be captured, and there is no overlapping area for the partial multimedia files of the multimedia file to be captured displayed in each sub - capture area;

[0156] A display module 1220, configured to display at least one captured image in response to the first input. One captured image is displayed within one sub - capture area, and the captured image within one sub - capture area is the captured image of the partial multimedia file displayed in the sub - capture area;

[0157] The receiving module 1210 is further configured to receive a second input from the user for dragging the at least one captured image to the image editing area of the capture interface;

[0158] An image generation module 1230, configured to generate a target image based on the at least one captured image within the image editing area in response to the second input.

[0159] In an embodiment of the present application, in each sub - shooting area of the shooting area of the shooting interface, partial multimedia files of the multimedia file to be shot are respectively displayed. And when there is no overlapping area among the partial multimedia files of the multimedia file to be shot displayed in each sub - shooting area, by responding to a first input of the user to at least one sub - shooting area, shooting images of the partial multimedia files displayed therein can be respectively displayed in at least one sub - shooting area where the first input is executed. In this way, partial image areas of the multimedia file to be shot can be directly obtained without the user cropping the multimedia file to be shot, which simplifies the user operation. Then, when receiving a second input from the user to drag at least one shooting image to the image editing area of the shooting interface, a target image can be directly generated based on at least one shooting image in the image editing area, without the user manually stitching at least one shooting image in the image editing area, which simplifies the user operation and improves the efficiency of target image generation.

[0160] In some embodiments of the present application, the image editing area includes an image generation control; a display module 1220, which is further configured to, in response to the second input, display the at least one shooting image in the image editing area;

[0161] A receiving module 1210, which is further configured to receive a third input of the user to the image generation control;

[0162] An image generation module 1230, specifically configured to, in response to the third input, splice the at least one shooting image in the image editing area to generate a target image.

[0163] In some embodiments of the present application, the image editing area further includes an insertion control for inserting description information of a shooting image, and the insertion control includes at least one of an audio control and a text input box;

[0164] The receiving module 1210 is further configured to, before receiving the third input of the user, receive a fourth input of the user to drag the insertion control to a first preset range area of a first shooting image, and the at least one shooting image includes the first shooting image;

[0165] The display module 1220 is further configured to, in response to the fourth input, display the insertion control in the first preset range area;

[0166] The receiving module 1210 is further configured to receive a fifth input of the user to the insertion control;

[0167] The display module 1220 is further configured to, in response to the fifth input, display an information identifier for indicating the description information of the first shooting image;

[0168] The image generation module 1230 is specifically configured to splice other captured images and the first captured image having the information identifier to obtain a target image, where the other captured images are the captured images in the image editing area except the first captured image.

[0169] In some embodiments of the present application, the display module 1220 is specifically configured to:

[0170] When the insertion control includes the audio control, display a voice progress bar for indicating the description information of the first captured image, where the voice progress bar is used to indicate the voice information describing the first captured image;

[0171] When the insertion control includes the text input box, display first text information for indicating the description information of the first captured image within the text input box.

[0172] In some embodiments of the present application, after displaying at least one captured image in response to the first input, the device further includes:

[0173] An identification module, configured to identify the text information in the second captured image to obtain the second text information of the second captured image, where the second captured image is any one of the at least one captured images displayed in the at least one capture area;

[0174] The receiving module 1210 is further configured to receive a sixth input in which the user drags the second text information to a second preset range area of a third captured image, where the third captured image is any one of the at least one captured images in the image editing area;

[0175] The image generation module 1230 is specifically configured to respond to the sixth input, splice other captured images and the third captured image having the second text information to obtain a target image, where the other captured images are the captured images in the image editing area except the third captured image.

[0176] The image generation device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0177] The image generation device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0178] The image generation device provided in the embodiments of the present application can implement Figure 1 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.

[0179] Optionally, as Figure 13 shown, the embodiments of the present application further provide an electronic device 1300, including a processor 1301 and a memory 1302. A program or instruction that can run on the processor 1301 is stored on the memory 1302. When the program or instruction is executed by the processor 1301, it implements each step of the above-mentioned image generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0180] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0181] Figure 14 A schematic diagram of the hardware structure of an electronic device for implementing the embodiments of the present application.

[0182] The electronic device 1400 includes, but is not limited to, components such as a radio frequency unit 1401, a network module 1402, an audio output unit 1403, an input unit 1404, a sensor 1405, a display unit 1406, a user input unit 1407, an interface unit 1408, a memory 1409, and a processor 1410, etc.

[0183] Those skilled in the art can understand that the electronic device 1400 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1410 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 14 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0184] Among them, the user input unit 1407 is used to receive a first input from the user for at least one sub - shooting area in the shooting area of the shooting interface. Each sub - shooting area respectively displays a part of the multimedia file to be shot, and there is no overlapping area for the parts of the multimedia file to be shot displayed in each sub - shooting area;

[0185] The display unit 1406 is used to display at least one shooting image in response to the first input. One shooting image is displayed in one sub - shooting area, and the shooting image displayed in one sub - shooting area is the shooting image of the part of the multimedia file displayed in the sub - shooting area;

[0186] The user input unit 1407 is further used to receive a second input from the user for dragging the at least one shooting image to the image editing area of the shooting interface;

[0187] The processor 1410 is used to generate a target image based on the at least one shooting image in the image editing area in response to the second input.

[0188] In this way, in each sub-shooting area of the shooting area on the shooting interface, partial multimedia files of the multimedia file to be shot are respectively displayed, and when there is no overlapping area among the partial multimedia files of the multimedia file to be shot displayed in each sub-shooting area, by responding to the first input of the user to at least one sub-shooting area, shooting images of the partial multimedia files displayed therein can be respectively displayed in at least one sub-shooting area where the first input is executed. In this way, partial image areas of the multimedia file to be shot can be directly obtained without the user cropping the multimedia file to be shot, simplifying the user operation. Then, when receiving the second input of the user dragging at least one shooting image to the image editing area of the shooting interface, a target image can be directly generated based on at least one shooting image in the image editing area, without the user manually splicing at least one shooting image in the image editing area, simplifying the user operation and improving the target image generation efficiency.

[0189] Optionally, the image editing area includes an image generation control; the display unit 1406 is further configured to, in response to the second input, display the at least one shooting image in the image editing area;

[0190] The user input unit 1407 is further configured to receive a third input of the user to the image generation control;

[0191] The processor 1410 is further configured to, in response to the third input, splice the at least one shooting image in the image editing area to generate a target image.

[0192] In this way, in response to the second input, at least one shooting image is displayed in the image editing area, and then in response to the third input of the user to the image generation control, the at least one shooting image in the image editing area can be directly spliced to generate a target image, without the user manually splicing the at least one generated shooting image, simplifying the user operation and improving the generation efficiency of the target image.

[0193] Optionally, the image editing area further includes an insertion control for inserting description information of the shooting image, and the insertion control includes at least one of an audio control and a text input box; the user input unit 1407 is further configured to receive a fourth input of the user dragging the insertion control to a first preset range area of the first shooting image, and the at least one shooting image includes the first shooting image;

[0194] The display unit 1406 is further configured to, in response to the fourth input, display the insertion control in the first preset range area;

[0195] The user input unit 1407 is further configured to receive a fifth input of the user to the insertion control;

[0196] The display unit 1406 is further configured to, in response to the fifth input, display an information identifier for indicating description information of the first captured image.

[0197] The processor 1410 is further configured to splice other captured images and the first captured image having the information identifier to obtain a target image, where the other captured images are captured images in the image editing area except the first captured image.

[0198] In this way, by adding corresponding description information to the captured images in the image editing area through the insertion control, the obtained target image can include the description information of the captured images, so that the user can intuitively view the information described by the captured images and can intuitively know the content described by the captured images according to the description information.

[0199] Optionally, the display unit 1406 is further configured to, when the insertion control includes the audio control, display a voice progress bar for indicating the description information of the first captured image, where the voice progress bar is used to indicate the voice information describing the first captured image; when the insertion control includes the text input box, display first text information for indicating the description information of the first captured image in the text input box.

[0200] In this way, when the insertion control includes the audio control, a voice progress bar for indicating the description information of the first captured image can be displayed, so that voice description information can be added to the first captured image. When the insertion control includes the text input box, first text information for indicating the description information of the first captured image can be displayed in the text input box, so that text description information can be added to the first captured image. In this way, different forms of description information can be added to the first captured image according to the user's needs, improving the diversity of the description information of the first captured image.

[0201] Optionally, after displaying at least one captured image in response to the first input, the processor 1410 is further configured to recognize text information in a second captured image to obtain second text information of the second captured image, where the second captured image is any one of the at least one captured images displayed in the at least one capture area.

[0202] The user input unit 1407 is further configured to receive a sixth input in which the user drags the second text information to a second preset range area of a third captured image, where the third captured image is any one of the at least one captured images in the image editing area.

[0203] The processor 1410 is further configured to, in response to the sixth input, splice other captured images and the third captured image having the second text information to obtain a target image, where the other captured images are the captured images in the image editing area except the third captured image.

[0204] In this way, by recognizing the text information in the captured images within the capture area and in response to the eighth input in which the user drags the text information to the third captured image, text description information can be added to the third captured image, thereby enhancing the flexibility of adding text description information to the captured images in the image editing area.

[0205] It should be understood that, in the embodiment of the present application, the input unit 1404 may include a Graphics Processing Unit (GPU) 14041 and a microphone 14042. The graphics processor 14041 processes the image data of still pictures or videos obtained by an image capture device (such as a color camera) in a video capture mode or an image capture mode. The display unit 1406 may include a display panel 14061, and the display panel 14061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1407 includes at least one of a touch panel 14071 and other input devices 14072. The touch panel 14071 is also referred to as a touch screen. The touch panel 14071 may include two parts: a touch detection device and a touch controller. The other input devices 14072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0206] The memory 1409 can be used to store software programs and various data. The memory 1409 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1409 can include volatile memory or non-volatile memory, or the memory 1409 can include both volatile and non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1409 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memory.

[0207] The processor 1410 may include one or more processing units; optionally, the processor 1410 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1410 either.

[0208] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above embodiment of the image generation method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0209] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks or optical discs, etc.

[0210] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above image generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0211] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system or system-on-chip, etc.

[0212] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above image generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0213] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0214] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0215] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. An image generation method, characterized in that, The method includes: Receiving a first input from a user for at least one sub - shooting area in a shooting area of a shooting interface, each of the sub - shooting areas respectively displays a partial multimedia file of a multimedia file to be shot, and there is no overlapping area for the partial multimedia files of the multimedia file to be shot displayed in each of the sub - shooting areas; In response to the first input, displaying at least one shooting image, one shooting image is displayed in one sub - shooting area, and the shooting image in one sub - shooting area is a shooting image of the partial multimedia file displayed in the sub - shooting area; Receiving a second input from the user to drag the at least one shooting image to an image editing area of the shooting interface; In response to the second input, generating a target image based on the at least one shooting image in the image editing area.

2. The method according to claim 1, characterized in that The image editing area includes an image generation control; The generating a target image based on the at least one shooting image in the image editing area in response to the second input includes: In response to the second input, displaying the at least one shooting image in the image editing area; Receiving a third input from the user for the image generation control; In response to the third input, splicing the at least one shooting image in the image editing area to generate a target image.

3. The method according to claim 2, characterized in that, The image editing area further includes an insertion control for inserting description information of a shooting image, and the insertion control includes at least one of an audio control and a text input box; Before receiving the third input from the user, the method further includes: Receiving a fourth input from the user to drag the insertion control to a first preset range area of a first shooting image, and the at least one shooting image includes the first shooting image; In response to the fourth input, displaying the insertion control in the first preset range area; Receiving a fifth input from the user for the insertion control; In response to the fifth input, displaying an information identifier for indicating the description information of the first shooting image; The splicing the at least one shooting image in the image editing area to obtain a target image includes: Splicing other shooting images and the first shooting image with the information identifier to obtain a target image, and the other shooting images are shooting images in the image editing area except the first shooting image.

4. The method according to claim 3, characterized in that, The displaying an information identifier for indicating the description information of the first shooting image includes: In the case where the insertion control includes the audio control, displaying a voice progress bar for indicating the description information of the first shooting image, and the voice progress bar is used to indicate voice information describing the first shooting image; In the case where the insertion control includes the text input box, displaying first text information for indicating the description information of the first shooting image in the text input box.

5. The method according to claim 2, characterized in that, After displaying at least one shooting image in response to the first input, the method further includes: Identify the text information in the second captured image to obtain the second text information of the second captured image, where the second captured image is any one of the at least one captured image displayed within the at least one captured area; The splicing of the at least one captured image within the image editing area to obtain a target image includes: Receiving a sixth input from the user to drag the text information into a second preset range area of a third captured image, where the third captured image is any one of the at least one captured image within the image editing area; In response to the sixth input, splice other captured images and the third captured image with the second text information to obtain a target image, where the other captured images are the captured images in the image editing area except the third captured image.

6. An image generation device, characterized in that, The device includes: A receiving module, configured to receive a first input from the user for at least one sub-captured area in the captured area of the capture interface, where each sub-captured area respectively displays a part of a multimedia file to be captured, and there is no overlapping area for the parts of the multimedia file to be captured displayed within each sub-captured area; A display module, configured to display at least one captured image in response to the first input, where one captured image is displayed within one sub-captured area, and the captured image within one sub-captured area is the captured image of the part of the multimedia file displayed within the sub-captured area; The receiving module is further configured to receive a second input from the user to drag the at least one captured image to the image editing area of the capture interface; An image generation module, configured to generate a target image based on the at least one captured image within the image editing area in response to the second input.

7. The device according to claim 6, characterized in that, The image editing area includes an image generation control; The display module is further configured to display the at least one captured image within the image editing area in response to the second input; The receiving module is further configured to receive a third input from the user for the image generation control; The image generation module is specifically configured to splice the at least one captured image within the image editing area in response to the third input to generate a target image.

8. The device according to claim 7, characterized in that The image editing area further includes an insertion control for inserting description information of a captured image, and the insertion control includes at least one of an audio control and a text input box; The receiving module is further configured to receive a fourth input from the user to drag the insertion control into a first preset range area of a first captured image before receiving the third input from the user, where the at least one captured image includes the first captured image; The display module is further configured to display the insertion control within the first preset range area in response to the fourth input; The receiving module is further configured to receive a fifth input from the user for the insertion control; The display module is further configured to display an information identifier for indicating the description information of the first captured image in response to the fifth input; The image generation module is specifically configured to splice other captured images and the first captured image with the information identifier to obtain a target image, where the other captured images are the captured images in the image editing area except the first captured image.

9. The device according to claim 8, wherein The display module is specifically configured to: When the insertion control includes the audio control, display a voice progress bar for indicating the description information of the first captured image, where the voice progress bar is used to indicate the voice information describing the first captured image; When the insertion control includes the text input box, display first text information for indicating the description information of the first captured image in the text input box.

10. The device according to claim 7, characterized in that, After displaying at least one captured image in response to the first input, the apparatus further includes: An identification module, configured to identify the text information in the second captured image to obtain second text information of the second captured image, where the second captured image is any one of the at least one captured images displayed in the at least one capture area; The receiving module is further configured to receive a sixth input in which the user drags the second text information to a second preset range area of a third captured image, where the third captured image is any one of the at least one captured images in the image editing area; The image generation module is specifically configured to respond to the sixth input, splice other captured images and the third captured image with the text information to obtain a target image, where the other captured images are the captured images in the image editing area except the third captured image.