Interaction method and device, electronic equipment and storage medium
By receiving identification information and text prompts from multiple reference images in the video generation application, the problem of only being able to generate videos from a single image in the prior art is solved, and the flexibility and accuracy of generating videos from multiple images are achieved.
Patent Information
- Application Number
- CN202510994076.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-18
AI Technical Summary
In existing technologies, video generation applications can typically only generate videos based on a single reference image or the first and last frames, which cannot meet users' needs for generation based on multiple reference images.
An interactive method is provided, which receives the identification information of multiple reference images and text prompts through a guided input area. The object description phrases correspond to the reference images, and the image identification information is interspersed in the text prompts to generate a video.
It enables video generation based on multiple reference images, meeting users' needs for complex scene construction and refined content control, and improving the accuracy and flexibility of video generation.
Smart Images

Figure CN120897092A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of video production, and particularly relates to an interaction method and device, an electronic device, and a storage medium. BACKGROUND
[0002] With the rapid development of science and technology, video generation technology has become one of the important achievements in the field of digitalization today. Through a video generation model, the technology automatically generates video content that meets expectations according to reference images and / or text prompt words, bringing new innovation opportunities to fields such as film and television creation and advertisement design.
[0003] In related technologies, an application program supporting a video generation function usually only provides a function of generating a video based on a single reference image, or generates a video given a first frame image and a last frame image. However, in practice, users often want to generate a video based on multiple reference images, and therefore how to assist users in generating a video based on multiple reference images is a problem that needs to be solved at present. SUMMARY
[0004] To solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an interaction method, device, electronic device, and storage medium.
[0005] In a first aspect, the present disclosure provides an interaction method, comprising:
[0006] displaying a video generation configuration page, the video generation configuration page comprising a guide sentence input area, a first reference image adding option, and a first generation option;
[0007] receiving a first reference image through the first reference image adding option, and displaying identification information of the received first reference image in the guide sentence input area; the number of the first reference images is multiple;
[0008] receiving a text prompt word through the guide sentence input area, and displaying the text prompt word in the guide sentence input area; the text prompt word comprises an object description word group, the object description word group corresponds to the first reference image; the object description word group is used to describe an object in the first reference image corresponding thereto; the identification information of the first reference image is interspersed in the text prompt word; the number of characters between the position of the object description word group in the text prompt word and the position of the identification information of the first reference image corresponding thereto in the text prompt word is less than a preset character number;
[0009] generating a first video based on the content in the guide sentence input area in response to a selection operation on the first generation option.
[0010] In a second aspect, the present disclosure provides an interaction device, comprising:
[0011] A first display module configured to display a video generation configuration page, the video generation configuration page comprising a guide sentence input area, a first reference image adding option, and a first generation option;
[0012] A first collection module configured to receive a first reference image through the first reference image adding option and display identification information of the received first reference image in the guide sentence input area; the number of the first reference images is plural;
[0013] A second collection module configured to receive a text prompt through the guide sentence input area and display the text prompt in the guide sentence input area; the text prompt comprises an object description phrase group, the object description phrase group corresponding to the first reference image; the object description phrase group is used to describe an object in the first reference image corresponding thereto; the identification information of the first reference image is interspersed in the text prompt; the number of characters between the position of the object description phrase group in the text prompt and the position of the identification information of the first reference image corresponding thereto in the text prompt is less than a preset character number;
[0014] A generation module configured to generate a first video based on the content in the guide sentence input area in response to a selection operation on the first generation option.
[0015] In a third aspect, the present disclosure provides an electronic device, comprising:
[0016] One or more processors;
[0017] A storage device configured to store one or more programs;
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the interaction method as described above.
[0019] In a fourth aspect, the present disclosure provides a computer readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the interaction method as described above.
[0020] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:
[0021] The technical scheme provided by the embodiment of the present disclosure receives a first reference image through the first reference image adding option, and displays identification information of the received first reference image in the guide language input area; the number of the first reference images is multiple; receives a text prompt word through the guide language input area, and displays the text prompt word in the guide language input area; the text prompt word includes an object description group, the object description group corresponds to the first reference image; the object description group is used to describe the object in the first reference image corresponding thereto; the identification information of the first reference image is interspersed in the text prompt word; the number of characters between the position of the object description group in the text prompt word and the position of the identification information of the first reference image corresponding thereto in the text prompt word is less than a preset character number; and a first video is generated based on the content in the guide language input area in response to a selection operation on the first generation option. The essence is to give an interactive link for video generation based on multiple first reference images, which can meet the demand of the user for generating a video based on multiple first reference images. BRIEF DESCRIPTION OF DRAWINGS
[0022] The accompanying drawings, which are incorporated herein and constitute part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.
[0024] Figure 1 A flowchart of an interactive method provided by the embodiment of the present disclosure;
[0025] Figures 2-11 Schematic diagrams of several pages in the embodiment of the present disclosure;
[0026] Figure 12 A structural schematic diagram of an interactive device in the embodiment of the present disclosure;
[0027] Figure 13 A structural schematic diagram of an electronic device in the embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained based on these drawings without creative labor.
[0029] Many particular details are set forth in the following description in order to provide a thorough understanding of the present disclosure. However, the present disclosure can be practiced according to other embodiments that depart from the specific details disclosed herein. It is understood that the examples described in this specification are only examples and are not intended to limit the scope, applicability, or configuration of the present disclosure in any way.
[0030] Figure 1 A flowchart of an interaction method provided by an embodiment of the present disclosure is provided. The embodiment can be applied to the case of interaction in a client. The method can be executed by an interaction device, which can be implemented in software and / or hardware. The device can be configured in an electronic device, such as a terminal, for example, a smartphone, a palm computer, a tablet computer, a wearable device with a display screen, a desktop computer, a notebook computer, an all-in-one computer, a smart home device, and the like. Alternatively, the embodiment can be applied to the case of interaction in a server. The method can be executed by an interaction device, which can be implemented in software and / or hardware. The device can be configured in an electronic device, such as a server.
[0031] As shown in Figure 1 , the method specifically can include:
[0032] S110, a video generation configuration page is displayed, and the video generation configuration page includes a guide language input area, a first reference image adding option, and a first generation option.
[0033] The video generation configuration page may, for example, be a page for collecting necessary information required for generating a video. In practice, the present application does not limit the information specifically required by the video generation configuration page. Exemplarily, the video generation configuration page can be used to collect guide language, version information of a video generation model used for generating a video, screen motion speed of a generated video, a panning mode, a format, and the like.
[0034] The guide language may, for example, be information used for input into a video generation model to control a video generation process. In the present application, the guide language displayed in the guide language input area is guide language used for generating a first video, that is, guide language corresponding to the first video. The guide language corresponding to the first video includes a text prompt word and a first reference image.
[0035] The guide language input area may, for example, be an area for collecting a text prompt word and displaying currently collected guide language. In some scenarios, a user can perform text input in the guide language input area.
[0036] The first reference image adding option may, for example, be an option for adding a first reference image. The first reference image can be understood as guide image that needs to be input into a video generation model and is part of the guide language. When the first reference image is input into the video generation model, the first reference image will participate in the control of the video generation process.
[0037] The first generation option may be, for example, an option for triggering the video generation model to run. When the first generation option is selected, it means that the user approves the video generation based on the information collected on the current video generation configuration page.
[0038] For example, referring to Figure 2 The video generation configuration page includes a prompt word input area, a first reference image adding option, and a first generation option.
[0039] In S120, a first reference image is received through the first reference image adding option, and identification information of the received first reference image is displayed in the guide sentence input area; the number of first reference images is multiple.
[0040] The identification information of the first reference image may be, for example, information for distinguishing different first reference images. For example, the identification information of the first reference image may include a thumbnail, a name, or a number of the first reference image, etc. Further, if the identification information of the first reference image includes a number of the first reference image, the number of the first reference image may be determined based on the arrangement order of the reference image in the guide sentence input area.
[0041] Receiving a certain first reference image means that the first reference image is used as part of the guide sentence.
[0042] The purpose of "displaying the identification information of the received first reference image in the guide sentence input area" is to visually display the identification information of the first reference image so that the user can clearly know which images are currently used as the first reference image.
[0043] There are various implementation methods for this step, which are not limited in the present application. For example, the implementation method of this step may include: in response to the selection operation of the first reference image adding option, displaying an image selection page; the image selection page includes multiple images; in response to the selection operation of a target image on the image selection page, the target image is used as a first reference image, and the identification information of the first reference image is displayed in the guide sentence input area.
[0044] The image selection page may be, for example, a page for providing multiple images for user selection. The target image may be, for example, an image selected by the user on the image selection page. For example, the image selection page may be an album page.
[0045] S130, receive the text prompt word through the guidance language input area, and display the text prompt word in the guidance language input area; the text prompt word comprises an object description word group, the object description word group corresponds to the first reference image; the object description word group is used to describe the object in the first reference image corresponding thereto; the identification information of the first reference image is interspersed in the text prompt word; the number of characters between the position of the object description word group in the text prompt word and the position of the identification information of the first reference image corresponding thereto in the text prompt word is less than a preset character number.
[0046] The object may be, for example, a thing in the first reference image, or an environment created by the first reference image, etc. If the object is a thing in the first reference image, the object may be, for example, a person, an animal, a plant, a building or an object, etc. If the object is an environment created by the first reference image, the object may be, for example, a bedroom environment, a cinema environment, a coffee shop environment, a forest environment, a morning environment or an evening environment, etc.
[0047] The object description word group may be, for example, a word group in the text prompt word, which corresponds to a specific first reference image and is used to describe the object or the characteristics of the object in the first reference image corresponding thereto, so that the video generation model "understands" which object in the first reference image the user wants to introduce into the newly generated video; or so that the video generation model "understands" which characteristics of which object in the first reference image the user wants to introduce into the newly generated video.
[0048] The identification information of the first reference image is interspersed in the text prompt word, for example, it may mean that in the guidance language input area, the identification information of the first reference image and the text prompt word are mixed together, instead of further dividing the guidance language input area into two non-overlapping sub-areas, one sub-area is used to display the identification information of the first reference image, and the other sub-area is used to display the text prompt word.
[0049] The preset character quantity may be a parameter preset for determining whether the object description phrase and the first reference image have a corresponding relationship. The application does not limit the specific value of the preset character quantity. If the character quantity between the position of the object description phrase in the text prompt phrase and the position of the first reference image in the text prompt phrase is less than the preset character quantity, it is considered that the two have a corresponding relationship, otherwise, it is considered that the two do not have a corresponding relationship. Therefore, the character quantity between the position of the object description phrase in the text prompt phrase and the position of the identification information of the first reference image corresponding to the object description phrase in the text prompt phrase is less than the preset character quantity, for example, it can be said that the position of the object description phrase in the text prompt phrase and the position of the identification information of the first reference image corresponding to the object description phrase in the text prompt phrase are close, so that according to the position in the text prompt phrase, the corresponding relationship between the object description phrase and the first reference image can be clearly distinguished. It should be emphasized that in practice, the text prompt phrase can be understood as a sentence, and the position of the object description phrase in the text prompt phrase and the position of the identification information of the first reference image corresponding to the object description phrase in the text prompt phrase are close, that is, the character quantity between the two is less, and in some scenarios, it can be expressed as being in the same sentence. Again, the "close" here refers to the character interval of the object description phrase and the first reference image in the text prompt phrase, not the physical distance between the two.
[0050] For example, referring to Figure 3 , the user can add image 1, image 2 and image 3 as the first reference image in sequence through the first reference image adding option, and display the identification information of image 1, the identification information of image 2 and the identification information of image 3 in the guide sentence input area. The identification information of image 1 includes a thumbnail of image 1 and a number of image 1. In Figure 3 , the "No. Figure 1 " after the thumbnail of image 1 represents the number of image 1, which is determined based on the arrangement order of the reference image in the guide sentence input area.
[0051] Continuing to refer to Figure 3The prompt word input area further includes a text prompt word, which is the text remaining in the information displayed by the prompt word input area except for the identification information of the first reference image, and specifically is "a kitten with patterned fur and a puppy in a park chase". The text prompt word includes three object description phrases "a kitten with patterned fur", "a puppy" and "a park". Among them, "a kitten with patterned fur" corresponds to image 1, which describes the object in image 1, in other words, image 1 includes a kitten with patterned fur. In an actual use scenario, since the user wants to input "a kitten with patterned fur" in image 1 into the video generation model to control the generation of the video, the user sets image 1 as the first reference image by means of the first reference image adding option, displays the identification information of image 1 in the prompt word input area, and inputs the object description phrase "a kitten with patterned fur" corresponding to image 1 in the guide sentence input area. The cases of other object description phrases are similar. Details are not described here.
[0052] Continuing to refer to Figure 3 In the prompt word input area, the identification information of the first reference image is interspersed in the text prompt word, and the number of characters between the identification information of image 1 and "a kitten with patterned fur" is less than the preset number of characters. Based on the content of the prompt word input area, it can be unambiguously obtained that image 1 corresponds to "a kitten with patterned fur". The correspondence between other first reference images and object description phrases is similar. Details are not described here.
[0053] It should be noted that in actual application, the user can construct the guide sentence through various interactive modes. For example, the user can first set multiple first reference images at a time through the first reference image adding option, and the electronic device displays the identification information of each first reference image in the guide sentence input area; then the user can input part of the text prompt word by moving the input cursor between the first reference image identification information until the guide sentence with mixed arrangement of the identification information of the first reference image and the text prompt word is formed. Alternatively, the user can first input the complete text prompt word in the guide sentence input area at a time, and then set the first reference image through the first reference image adding option by positioning the input cursor to the appropriate position in the text prompt word, so as to form the guide sentence with mixed arrangement of the identification information of the first reference image and the text prompt word. In addition, the user can also add one first reference image through the first reference image adding option, and then input part of the text prompt word; add another first reference image through the first reference image adding option, and continue to input part of the text prompt word; and so on, to gradually construct the guide sentence with mixed arrangement of the identification information of the first reference image and the text prompt word.
[0054] S140, in response to the selection operation on the first generation option, generating a first video based on the content in the guide sentence input area.
[0055] The content in the guidance language input area is the guidance language. Optionally, in response to the selection operation on the first generation option, the content in the guidance language input area is input to the video generation model to obtain a first video. The first video is a video generated based on the plurality of first reference images.
[0056] The above technical solution receives a first reference image through the first reference image adding option, and displays identification information of the received first reference image in the guidance language input area; the number of the first reference images is a plurality; receives a text prompt word through the guidance language input area, and displays the text prompt word in the guidance language input area; the text prompt word includes an object description word group, the object description word group corresponds to the first reference image; the object description word group is used to describe an object in the first reference image corresponding thereto; the identification information of the first reference image is interspersed in the text prompt word; the number of characters between the position of the object description word group in the text prompt word and the position of the identification information of the first reference image corresponding thereto in the text prompt word is less than a preset character number; in response to the selection operation on the first generation option, a first video is generated based on the content in the guidance language input area. The essence is to give an interactive link for video generation based on a plurality of first reference images, which can meet the user's demand for generating a video based on a plurality of first reference images.
[0057] It should be noted that in practice, one technical difficulty of generating a video based on a plurality of first reference images is how to organize the guidance language so that the user can more accurately express the influence and / or action mode of the plurality of first reference images input by the user on the video generation process. The above technical solution sets the text prompt word to include an object description word group, the object description word group corresponds to the first reference image; the object description word group is used to describe an object in the first reference image corresponding thereto; the identification information of the first reference image is interspersed in the text prompt word; the number of characters between the position of the object description word group in the text prompt word and the position of the identification information of the first reference image corresponding thereto in the text prompt word is less than a preset character number; it gives an organization mode of the guidance language, which can help the user to accurately express the influence and / or action mode of the plurality of first reference images input by the user on the video generation process, so as to make the generated first video meet the user's demand for complex scene construction and fine content control.
[0058] On the basis of the above technical solution, optionally, after S130 and before S140, the method can further include: in response to a first selection operation on a target first reference image in the guidance language input area, displaying a first replacement option; in response to a selection operation on the first replacement option, replacing the target first reference image and replacing the identification information of the target first reference image in the guidance language input area.
[0059] The target first reference image may be, for example, a first reference image selected by the user. The first replacement option may be, for example, an option corresponding to the target first reference image, for replacing the target first reference image. Replacing the target first reference image may mean, for example, no longer using the image currently serving as the target first reference image, but instead using another image to replace the image as a new target first reference image, which will be subsequently input into the video generation model for generating the first video.
[0060] For example, referring to Figure 4 If the user clicks the identification information of image 2 in the guide sentence input area to take image 2 as the target first reference image, the operation of the user clicking the identification information of image 2 in the guide sentence input area is regarded as a selection operation of the target first reference image in the guide sentence input area, and a first replacement option corresponding to image 2 (i.e. Figure 4 the “Replace” option above the identification information of image 2 in the guide sentence input area) is displayed. If the user continues to click the “Replace” option, an image selection page can be displayed; the image selection page includes multiple images; in response to a selection operation of a target image in the image selection page, the target image is used to replace Figure 2 . Assuming that the target image is image 4, the identification information of image 4 will replace the identification information of image 2 and be displayed in the guide sentence input area. Subsequently, image 4 will be input into the video generation model, but image 2 will not be input into the video generation model.
[0061] By setting that in response to a first selection operation of a target first reference image in a guide sentence input area, a first replacement option is displayed; and in response to a selection operation of the first replacement option, the target first reference image is replaced, and the identification information of the target first reference image in the guide sentence input area is replaced. The essence is to allow the user to replace the first reference image, thereby meeting the user's demand for modifying the guide sentence that has been input in the guide sentence input area.
[0062] On the basis of the above technical solution, optionally, after S130 and before S140, the method can further include: in response to a second selection operation of a target first reference image in a guide sentence input area, displaying a deletion option; and in response to a selection operation of the deletion option, deleting the target first reference image and deleting the identification information of the target first reference image from the guide sentence input area.
[0063] The target first reference image may be, for example, a first reference image selected by the user. The first selection operation and the second selection operation can be the same interactive operation or different interactive operations, which are not limited in the present application. The interactive operation here can be a clicking operation, a sliding operation, a dragging operation, or a long-pressing operation, etc.
[0064] The deletion option may be, for example, an option corresponding to the target first reference image, for deleting the target first reference image. Deleting the target first reference image may be, for example, no longer using the current target first reference image, and the current target first reference image will not be input into the video generation model in the future to generate the first video.
[0065] For example, if the first selection operation and the second selection operation are the same interactive operation, which is a click operation. If the user clicks the identification information of image 2 in the guide sentence input area, image 2 is taken as the target first reference image, and the first replacement option (i.e., the “replacement” option in Figure 4 Figure 4 The user can click the “deletion” option to delete image 2. After deleting image 2, the identification information of image 2 will no longer be displayed in the guide sentence input area, and image 2 will not be input into the video generation model in the future.
[0066] By setting that in response to the second selection operation on the target first reference image in the guide sentence input area, the deletion option is displayed; in response to the selection operation on the deletion option, the target first reference image is deleted, and the identification information of the target first reference image is deleted from the guide sentence input area. The essence is to allow the user to delete the first reference image, which meets the user's demand for modifying the guide sentence that has been input in the guide sentence input area.
[0067] On the basis of the above technical solutions, the method may further include: in response to an operation of publishing the second video as a template, displaying a publishing configuration page; the second video includes the first video; the publishing configuration page includes identification information of the first reference image used by the first video and a publishing option; in response to a third selection operation on the target first reference image in the publishing configuration page, displaying a constraint condition input box corresponding to the first target slot; the first target slot is a material slot represented by the target first reference image; and the content received through the constraint condition input box is taken as a constraint condition corresponding to the first target slot; and in response to a selection operation on the publishing option, publishing the second video as a template based on the correspondence between the first target slot and the constraint condition.
[0068] The second video includes the first video, for example, may mean that the second video includes one or more video segments, and the first video is one of the video segments constituting the second video. When the second video includes multiple video segments, the multiple video segments are organized together by at least one of the following ways: sequentially spliced in a certain track according to time sequence, and layered picture effect is realized by multi-track superposition.
[0069] The second video is published as a template, which means that the second video is converted into a reusable standardized creative framework for the template user to quickly apply and generate a video of similar style. Here, the "template" is essentially a semi-finished product content containing preset editing logic, visual effects and structural layout, and through the open replaceable area (such as materials, text, music), the threshold of creation of others is reduced. The "replaceable area" is also called material slot.
[0070] The publishing configuration page may be, for example, a page for configuring basic information, material slots and the like of the template.
[0071] The publishing option may be, for example, an option for triggering the publishing of the second video in the form of a template. When the publishing option is selected, it means that the user agrees to publish the template based on the information configured on the current publishing configuration page.
[0072] The target first reference image may be, for example, a first reference image selected by the user. The third selection operation may be the same interactive operation as the first selection operation or the second selection operation mentioned above, or may be a different interactive operation, which is not limited in the present application. The interactive operation may be a click operation, a sliding operation, a dragging operation, or a long press operation, and the like.
[0073] When the second video is published as a template, since the second video includes the first video, the generation of the first video uses the target first reference image, and the second video can be understood as an example of this template, and the target first reference image is the material of the second video. The first target slot may be, for example, the position or vacancy occupied by the target first reference image in the template. In some scenarios, it can also be understood that the target first reference image fills the first target slot.
[0074] In some scenarios, when the second video template is used, if the target first reference image is replaced, the image used to replace the target first reference image will be inserted into the material slot represented by the target first reference image (i.e. the first target slot). In other words, the image used to replace the target first reference image will fill the material slot represented by the target first reference image (i.e. the first target slot).
[0075] The constraint condition corresponding to the first target slot may be, for example, information for prompting conditions that need to be met by the image used to replace the target first reference image. Only when the image used to replace the target first reference image meets the constraint condition corresponding to the first target slot, the new video produced based on the template will have better visual effects. The constraint condition corresponding to the first target slot input box is used to receive the constraint condition corresponding to the first target slot.
[0076] Since the guide speech used for generating the first video requires the object description phrase to correspond to the first reference image, when a new image is used to replace the first reference image in the template application stage, if the new image is not selected properly, the correspondence between the new image and the object description phrase can be invalid. For example, in the guide speech used for generating the first video, an object description phrase is "puppy", and the first reference image corresponding to the first object description phrase includes a puppy. If an image including a kitten is used to replace the first reference image, the object description phrase "puppy" and the kitten in the new first reference image do not match, i.e., the correspondence between the new image and the object description phrase is invalid. The occurrence of this situation can cause the visual effect of the video generated based on the new image to be poor. By configuring the constraint condition, the template user can be reminded to select an image that will not cause the correspondence between the new image and the object description phrase to be invalid. Alternatively, by configuring the constraint condition, the template user can be reminded to select an image that will cause the new image to have a high degree of matching with the object description phrase.
[0077] Optionally, in some scenarios, the second video further includes other video segments in addition to the first video, and the publishing configuration page further includes identification information of the other video segments.
[0078] For example, it is assumed that the first video is generated based on the guide speech in Figure 3 . The second video includes video segment 1 and the first video, where the video segment 1 is obtained by recording through a terminal and is different from the first video in the generation manner. It is assumed that the "publish as template" option is included in the page for displaying the second video. If a user clicks the "publish as template" option, see Figure 5 , a publishing configuration page is displayed. In the publishing configuration page, the basic information of the template (such as the theme of the template, the characteristics of the video that can be made by the template, and the cover image of the template) and the material slot can be configured.
[0079] In Figure 5 , the publishing configuration page includes the identification information of the video segment 1 and the identification information of the first reference images (i.e., image 1, image 2, and image 3) used to obtain the first video. In the publishing configuration page, the "modify" option corresponding to the identification information of each first reference image is further included. If a user selects a certain "modify" option, it is considered that the first reference image corresponding to the "modify" option is selected. The first reference image corresponding to the "modify" option is the target first reference image. For example, continuing to refer to Figure 5 , if the user clicks the "modify" option corresponding to the identification information of image 2, image 2 is used as the target first reference image, see Figure 6corresponding to image 2. In this example, since the object description phrase corresponding to image 2 in the guide speech of the first video is "puppy", the user can fill in "puppy" or "dog" or "animal" or the like in the constraint condition input box corresponding to image 2, and the word filled in by the user in the constraint condition input box corresponding to image 2 will be used as the constraint condition of the material slot represented by image 2. After the second video is published as a template, the constraint condition of the material slot represented by image 2 will prompt the template user to replace image 2 with an image that meets the constraint condition of the material slot represented by image 2, so that the resulting video has a better visual effect.
[0080] Therefore, by setting, in response to a third selection operation of a target first reference image on a publishing configuration page, a constraint condition input box corresponding to a first target slot; the first target slot is a material slot represented by the target first reference image; and receiving content through the constraint condition input box as a constraint condition corresponding to the first target slot; the essence is to allow the configuration of the constraint condition corresponding to the material slot during the template publishing process, so as to guide the template user to correctly select an image for replacing the target first reference image through the constraint condition in the future, and thus ensure that a video with a better visual effect can be finally generated using the template.
[0081] Further, before publishing the second video as a template based on the correspondence between the first target slot and the constraint condition, the method can further include: in response to a fourth selection operation of the target first reference image on the publishing configuration page, setting the state of the first target slot to an allowed replacement state.
[0082] The fourth selection operation can be the same interactive operation as the first selection operation, the second selection operation, or the third selection operation mentioned above, or can be a different interactive operation, which is not limited in the present application. The interactive operation can be a click operation, a sliding operation, a dragging operation, or a long-press operation, etc.
[0083] The state of the material slot can include an allowed replacement state and a restricted replacement state. The state of the first target slot being in the allowed replacement state means that the template user is allowed to fill the first target slot with other images during the application of the second video template. The state of the first target slot being in the restricted replacement state means that the template user is not allowed to fill the first target slot with other images during the application of the second video template. If the state of a certain material slot is in the restricted replacement state, the image or video used to fill the material slot during the creation of the second video can be used during the creation of a new video using the second video template.
[0084] Exemplarily, in the second video, the first target slot is represented by image 2, and the second target slot is represented by image 3. Figure 5In the above technical solutions, the checkmark in the upper right corner of the identification information of the video clip 1 means that the material slot represented by the video clip 1 is in the replaceable state. The absence of a checkmark in the upper right corner of the identification information of the image 1 means that the material slot represented by the image 1 is in the replaceable state. Similarly, in the above technical solutions, the checkmark in the upper right corner of the identification information of the image 2 and the image 3 means that the material slots represented by the image 2 and the image 3 are in the replaceable state. Figure 5
[0085] In some scenarios, certain or a first reference image has a critical impact on the generation of a first video, for example, it can determine the overall style, main content or key visual features of the video. In this case, there are often strict requirements for the quality, content or format of the first reference image. In this case, by setting the state of the first target slot to the replaceable state in response to the selection operation of the target first reference image on the publishing configuration page, the essence is to allow the template publisher to set which first reference image can be replaced and which first reference image cannot be replaced according to the characteristics of the template made by the template publisher in the template publishing configuration stage, thereby avoiding the situation that the output quality of the video is reduced due to the replacement of a relatively important first reference image in the template use stage.
[0086] On the basis of the above technical solutions, the method can further include: in response to an application operation of a third video template, displaying a first slot filling page; the third video includes the first video; the first video corresponds to a plurality of first material slots; the first slot filling page includes identification information of the first material slot and a second generation option; in response to a configuration operation of a target first material slot, configuring a second reference image for filling the target first material slot; in response to a selection operation of the second generation option, obtaining a fourth video based on the second reference image, a corresponding relationship between the second reference image and the first material slot, and the third video template; the fourth video includes a fifth video generated based on the second reference image.
[0087] The third video template can be a standardized template for secondary creation generated based on the content and editing logic of the third video. In other words, the third video template can be a reusable template framework that fixes the editing structure (such as video clip order, transition effect, caption style, music rhythm, etc.) of the third video, retains the core design logic of the third video, and opens part of the material slot (such as pictures, video clips, text content, etc.) for replacement by other users.
[0088] The third video includes the first video, for example, can mean that the third video includes one or more video segments, and the first video is one of the video segments constituting the third video. The first video is a video generated based on the method of generating a video based on a plurality of first reference images. When the third video includes a plurality of video segments, the plurality of video segments are organized together by at least one of the following ways: spliced in time sequence on a certain track, and layered picture effect is realized by multi-track superposition.
[0089] The application operation of the third video template can be, for example, an operation for triggering video creation based on the third video template. An important operation of video creation based on the third video template is to configure the material for filling the material slot in the third video template which is in the replaceable state.
[0090] The first slot filling page can be, for example, a page for assisting the template user to configure the filling image or video for the material slot in the third video template which is in the replaceable state.
[0091] Since in the third video, the first video is generated based on a plurality of first reference images. In the third material template, each first reference image represents a first material slot.
[0092] The identification information of the first material slot can be, for example, information for distinguishing different material slots.
[0093] The second generation option can be, for example, an option for triggering video generation based on the information configured by the current first slot filling page and the third video template. When the second generation option is selected, it means that the user approves the information configured by the current first slot filling page.
[0094] The target first material slot can be, for example, the selected first material slot, or the first material slot in the selected state. The configuration operation on the target first material slot can be, for example, an operation for indicating filling a specific image into the target first material slot. In some scenarios, when it is detected that a first material slot is in the selected state and a selection operation on an image in the album is received, the first material slot is determined as the target first material slot, and the selection operation on the image in the album is taken as the configuration operation on the target first material slot.
[0095] The image specified by the configuration operation for filling the target first material slot is the second reference image. The essence of "configuring the second reference image for filling the target first material slot" is to fill the target first material slot with that image, in other words, the essence of "configuring the second reference image for filling the target first material slot" is to establish a correspondence between the target first material slot and the second reference image.
[0096] The fourth video may be, for example, a result of re-creating a video based on the third video template.
[0097] In some scenarios, since the third video includes the first video, the creation process of the third video can be divided into two processes. The first process is the process of generating the first video based on the plurality of first reference images. The second process is the process of organizing the first video together with other materials (such as video clips, audio, text, special effects, etc.). Based on this, the third video template can be understood as a nested result of the first template and the second template. The first template is a reusable template framework that solidifies the generation process of the first video. The second template is a reusable template framework that solidifies the process of organizing the first video together with other materials (such as video clips, audio, text, special effects, etc.).
[0098] Since the fourth video is obtained based on the third video template, the fifth video in the fourth video can be regarded as a video obtained based on the second reference image, the correspondence between the second reference image and the first material slot, and the first template. In other words, the fifth video corresponds to the first video. The reason why the fifth video and the first video are different videos is that the first video uses the first reference image in the generation process, while the fifth video uses the second reference image in the generation process.
[0099] Optionally, in response to the target first material slot being in the selected state, the first slot filling page displays the constraint condition corresponding to the target first material slot. The purpose of this setting is to ensure that the template user can select a suitable image as the second reference image to fill the target first material slot under the prompting action of the constraint condition, thereby ensuring that the final generated fourth video has a better visual effect.
[0100] Optionally, the first slot filling page further includes a plurality of images or videos to be selected and available for filling the material slot.
[0101] For example, assume that the third video includes two video clips, video clip 2 and the first video, and the first video is generated based on three first reference images, image 1, image 2, and image 3. When creating a video based on the third video template, refer to Figure 7 , the first slot filling page is displayed, which includes the identification information of three first material slots and the identification information of one second material slot. The second material slot is the material slot represented by the video clip 2, the first material slot 1 is the material slot represented by the image 1, the first material slot 2 is the material slot represented by the image 2, and the first material slot 3 is the material slot represented by the image 3. Assume that the user clicks the identification information of the first material slot 2, taking the first material slot 2 as the target first material slot, refer to Figure 7The first material slot 2 is in a selected state, and the first slot filling page displays a constraint condition corresponding to the first material slot 2. The constraint condition is "please use an image including a puppy". Figure 7 In some embodiments, the first slot filling page further includes a plurality of selectable images that can be used to fill the material slot, such as image a to image i. Suppose the user then clicks on image b, and it is determined that image b is used to fill the first material slot 2, i.e., a corresponding relationship is established between image b and the first material slot 2. Figure 7 The "next step" option in the first slot filling page is a second generation option. If the user clicks on the "next step" option in the first slot filling page, a fourth video is generated based on the information configured in the first slot filling page and the third video template. Figure 7
[0102] In some embodiments, the method further includes: in response to an application operation on the third video template, displaying a first slot filling page; the third video includes a first video; the first video corresponds to a plurality of first material slots; the first slot filling page includes identification information of the first material slot and a second generation option; and in response to a configuration operation on a target first material slot, configuring a second reference image for filling the target first material slot. In essence, in the application stage of the third video template, the material slot represented by the first reference image used to generate the first video is opened to the third video template user, so that the third video template user can replace the first reference image and generate a fifth video corresponding to the first video that meets their own needs, thereby meeting the personalized use needs of the user.
[0103] Further, the method can further include: displaying a video display page, the video display page including the fourth video; the video display page further including identification information of the fifth video; in response to a selection operation on the fifth video in the video display page, displaying a second replacement option corresponding to the fifth video; and in response to a selection operation on the second replacement option, displaying a second slot filling page, the second slot filling page including identification information of the first material slot and a constraint condition corresponding to the first material slot.
[0104] The video display page can be, for example, a page for displaying the fourth video. The video display page including the fourth video can be, for example, that the video display page includes a video playback area and a playback control option. The playback control option can be, for example, an option for adjusting the playback. Illustratively, the playback control option can include one or more of the following: play, pause, fast forward, fast backward, and full screen playback.
[0105] The video display page can further include identification information of other video segments constituting the fourth video except the fifth video.
[0106] The selection operation on the fifth video in the video display page may be, for example, a selection operation on identification information of the fifth video in the video display page. The second replacement option may be, for example, an option for triggering replacement of the second reference image used to generate the fifth video.
[0107] The second slot filling page may be, for example, a page for assisting the user in replacing the second reference image used to generate the fifth video. Since the second reference image used to generate the fifth video occupies the first material slot, or in other words, since the second reference image has a correspondence relationship with the first material slot, the second reference image is used to generate the fifth video. Therefore, the second slot filling page includes identification information of the first material slot and constraint conditions corresponding to the first material slot. Optionally, the second slot filling page further includes a plurality of to-be-selected images or videos that can be used to fill the material slot.
[0108] The identification information of the first material slot and the constraint conditions corresponding to the first material slot are displayed on the second slot filling page, which aims to assist the template user in reconfiguring the reference image used to fill the first material slot for the first material slot.
[0109] Further, a third reference image used to fill a target first material slot may be configured in response to a configuration operation on the target first material slot in the second slot filling page, that is, the correspondence relationship between the second reference image and the target first material slot is released, and a correspondence relationship between the third reference image and the target first material slot is established.
[0110] For example, it is assumed that the third video includes two video clips, namely video clip 2 and the first video, and the first video is generated based on three first reference images, namely image 1, image 2, and image 3. When video creation is performed based on the third video template, a video obtained is a fourth video. The fourth video includes video clip 3 and a fifth video, wherein video clip 3 corresponds to video clip 2, and the fifth video corresponds to the first video. The fifth video is generated based on three second reference images, namely image a, image b, and image c. Referring to Figure 8 , a video display page is displayed, which includes a video playing area for playing the fourth video, identification information of video clip 3, and identification information of the fifth video. If a user clicks the identification information of the fifth video, referring to Figure 9 , a second replacement option corresponding to the fifth video (i.e., the “Replace” option in Figure 9 ) is displayed. If the user further clicks the second replacement option corresponding to the fifth video (i.e., the “Replace” option in Figure 9 ), referring to Figure 10 , a second slot filling page is displayed. The second slot filling page includes identification information of the first material slot and constraint conditions corresponding to the first material slot. InFigure 10 In the embodiment, since the identification information of the first material slot 2 is in the selected state, the second slot filling page displays the constraint condition corresponding to the first material slot 2. Referring to Figure 10 , the second slot filling page further includes a plurality of candidate images that can be used to fill the first material slot. On the basis of Figure 10 , the user can select a target first material slot on the second slot filling page, and select a third reference image for filling the target first material slot.
[0111] By setting, in response to the selection operation of the fifth video on the video display page, the second replacement option corresponding to the fifth video is displayed; and in response to the selection operation of the second replacement option, the second slot filling page is displayed, and the second slot filling page includes the identification information of the first material slot and the constraint condition corresponding to the first material slot. The essence is that after the fourth video is generated, the template user is allowed to replace the image used to fill the first material slot. When the second reference image used to fill the first material slot is replaced by the third reference image, it means that the fifth video will be regenerated, and it also means that the fourth video will be updated. Such a setting can meet the user's demand for editing and updating the fourth video in the case that the user is not satisfied with the fourth video.
[0112] In some embodiments, on the basis of the above technical solution, the method can further include: in response to the selection operation of the fifth video on the video display page, a guide sentence modification option is displayed; in response to the selection operation of the guide sentence modification option, a guide sentence corresponding to the fifth video is displayed; the guide sentence includes identification information of the second reference image and a text prompt word; in response to the replacement operation of the second reference image in the guide sentence, the second reference image is replaced, and the identification information of the second reference image in the guide sentence is replaced. Further, the method can further include: in response to the modification operation of the text prompt word in the guide sentence, the text prompt word in the guide sentence is modified.
[0113] The guide sentence modification option may, for example, be an option for modifying the guide sentence.
[0114] The first video is generated based on the guide sentence, and the guide sentence includes the first reference image and the text prompt word. Since the fifth video corresponds to the first video, it is assumed that in order to generate the fifth video, at least part of the first reference image used to generate the first video is replaced by the second reference image, but the text prompt information is not modified, and the guide sentence corresponding to the fifth video includes the first reference image and the text prompt information in the guide sentence corresponding to the first video which are not replaced, and the second reference image used to replace the first reference image. For example, it is assumed Figure 3The guide language given in the middle is the guide language corresponding to the first video. To generate the fifth video, image 2 is replaced by image b, image 3 is replaced by image c, image 1 is not replaced, and the text prompt information is not modified. The guide language corresponding to the fifth video includes image 1, image b, image c, and the text prompt information. The text prompt information in the guide language corresponding to the fifth video is the same as the text prompt information in the guide language corresponding to the first video.
[0115] The organization of the guide language corresponding to the fifth video is the same as the organization of the guide language corresponding to the first video, that is, the identification information of the reference image is interspersed in the text prompt word, and the number of characters between the position of the object description phrase in the text prompt word and the position of the identification information of the reference image corresponding thereto in the text prompt word is less than the preset number of characters.
[0116] For example, referring to Figure 8 , if the user clicks the identification information of the fifth video, referring to Figure 9 , the guide language modification option corresponding to the fifth video (that is, the "modify special effect" option in Figure 9 ) is displayed. If the user further clicks the guide language modification option corresponding to the fifth video (that is, the "modify special effect" option in Figure 9 ), referring to Figure 11 , the guide language corresponding to the fifth video is displayed. Assuming that the user wants to replace image b, the user can select the identification information of image b in the guide language. Alternatively, after selecting the identification information of image b, an image selection page can be invoked to select an image for replacing image b. When image b is replaced, the identification information of image b in the guide language will also be updated synchronously. In addition, the user can perform a modification operation on the text prompt word in the guide language based on Figure 11 .
[0117] By setting that in response to a selection operation on the guide language modification option, the guide language corresponding to the fifth video is displayed; the guide language includes the identification information of the second reference image and the text prompt word; in response to a replacement operation on the second reference image in the guide language, the second reference image in the guide language is replaced; and in response to a modification operation on the text prompt word in the guide language, the text prompt word in the guide language is modified. The essence is to display the guide language used to generate the fifth video to the template user. This can help the template user better understand the reason for the poor visual effect of the fifth video from the perspective of the correspondence between the object description phrase and the second reference image, so as to reconfigure the reference image or the text prompt information to improve the matching degree between the object description phrase having the correspondence and the second reference image, and thus optimize the effect of video creation using the third video template.
[0118] It should be noted that in the technical solutions provided in the present application, the selection operation of a specific option or identification information can be realized through direct physical contact of the user (such as touch screen operation), indirect physical input device (such as mouse, touchpad), or non-contact input means (such as gesture recognition, voice control), and the like. If it is realized through direct physical contact of the user (such as touch screen operation), indirect physical input device (such as mouse, touchpad), it can include but is not limited to the following several user interaction modes: click operation, sliding operation, dragging operation, or long press operation, and the like. In addition, according to different application scenarios and technical implementation conditions, the selection operation can also be extended to use more advanced interaction technologies, such as eye tracking, brain-computer interface, and other emerging human-computer interaction methods.
[0119] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, and the like of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0120] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will need to obtain and use the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the electronic device, application program, server or storage medium, and the like software or hardware that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0121] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0122] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0123] It should be noted that for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.
[0124] Figure 12Fig. 1 is a structural schematic diagram of an interactive device according to an embodiment of the present disclosure. The interactive device provided by the embodiment of the present disclosure can be configured in a client or in a server. Referring to Fig. 1, the interactive device includes a first display module 310, a first collection module 320, a second collection module 330 and a generation module 340. Figure 12 The interactive device specifically includes:
[0125] The first display module 310 is configured to display a video generation configuration page, wherein the video generation configuration page includes a guide language input area, a first reference image adding option and a first generation option.
[0126] The first collection module 320 is configured to receive a first reference image through the first reference image adding option and display identification information of the received first reference image in the guide language input area; the number of the first reference images is multiple.
[0127] The second collection module 330 is configured to receive a text prompt through the guide language input area and display the text prompt in the guide language input area; the text prompt includes an object description phrase group, the object description phrase group corresponds to the first reference image; the object description phrase group is used to describe an object in the first reference image corresponding thereto; the identification information of the first reference image is interspersed in the text prompt; the number of characters between the position of the object description phrase group in the text prompt and the position of the identification information of the first reference image corresponding thereto in the text prompt is less than a preset character number.
[0128] The generation module 340 is configured to generate a first video based on the content in the guide language input area in response to a selection operation on the first generation option.
[0129] Further, the device further includes a first modification module configured to:
[0130] display a first replacement option in response to a first selection operation on a target first reference image in the guide language input area;
[0131] replace the target first reference image and replace the identification information of the target first reference image in the guide language input area in response to a selection operation on the first replacement option.
[0132] Further, the device further includes a second modification module configured to:
[0133] display a deletion option in response to a second selection operation on a target first reference image in the guide language input area;
[0134] delete the target first reference image and delete the identification information of the target first reference image from the guide language input area in response to a selection operation on the deletion option.
[0135] Further, the apparatus further comprises a publishing module configured to:
[0136] In response to an operation of publishing the second video as a template, display a publishing configuration page; the second video comprises the first video; the publishing configuration page comprises identification information of the first reference image used by the first video and a publishing option;
[0137] In response to a third selection operation of a target first reference image on the publishing configuration page, display a constraint condition input box corresponding to a first target slot; the first target slot is a material slot represented by the target first reference image;
[0138] Receive content through the constraint condition input box as a constraint condition corresponding to the first target slot;
[0139] In response to a selection operation on the publishing option, publish the second video as a template based on the correspondence between the first target slot and the constraint condition.
[0140] Further, the publishing module is further configured to:
[0141] In response to a fourth selection operation of a target first reference image on the publishing configuration page before publishing the second video as a template based on the correspondence between the first target slot and the constraint condition in response to a selection operation on the publishing option, set the state of the first target slot to an allowed replacement state.
[0142] Further, the apparatus further comprises an application module configured to:
[0143] In response to an application operation of a third video template, display a first slot filling page; the third video comprises the first video; the first video corresponds to a plurality of first material slots; the first slot filling page comprises identification information of the first material slot and a second generation option;
[0144] In response to a configuration operation of a target first material slot, configure a second reference image for filling the target first material slot;
[0145] In response to a selection operation on the second generation option, generate a fourth video based on the second reference image, the correspondence between the second reference image and the first material slot, and the third video template; the fourth video comprises a fifth video generated based on the second reference image.
[0146] Further, the application module is further configured to:
[0147] In response to the target first material slot being in a selected state, a first slot filling page is displayed, the first slot filling page displaying a constraint condition corresponding to the target first material slot.
[0148] Further, the application module is further configured to:
[0149] display a video display page, the video display page including a fourth video; the video display page further including identification information of the fifth video;
[0150] In response to a selection operation of the fifth video on the video display page, a second replacement option corresponding to the fifth video is displayed.
[0151] In response to a selection operation of the second replacement option, a second slot filling page is displayed, the second slot filling page including identification information of the first material slot and a constraint condition corresponding to the first material slot.
[0152] Further, the application module is further configured to:
[0153] In response to a selection operation of the fifth video on the video display page, a guide sentence modification option is displayed.
[0154] In response to a selection operation of the guide sentence modification option, a guide sentence corresponding to the fifth video is displayed; the guide sentence including identification information of a second reference image and a text prompt word.
[0155] In response to a replacement operation of the second reference image in the guide sentence, the second reference image is replaced, and the identification information of the second reference image in the guide sentence is replaced.
[0156] Further, the application module is further configured to:
[0157] In response to a modification operation of the text prompt word in the guide sentence, the text prompt word in the guide sentence is modified.
[0158] The interaction device provided by the embodiments of the present disclosure can perform the steps performed by the client or the server in the interaction method provided by the embodiments of the present disclosure, and has the execution steps and beneficial effects, which will not be repeated here.
[0159] A first reference diagram 13 is shown below, which shows a structural schematic diagram suitable for implementing the electronic device 1000 in the embodiments of the present disclosure. The electronic device 1000 in the embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle terminal (e.g., a car navigation terminal), a wearable electronic device, and the like, and a stationary terminal such as a digital TV, a desktop computer, a smart home device, and the like. Figure 13 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0160] As shown in Figure 13 The electronic device 1000 can include a processing device (e.g., a central processing unit, a graphic processing unit, etc.) 1001 that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003 to implement the interaction method of the embodiments as described in the present disclosure. Various programs and information required for the operation of the electronic device 1000 are also stored in the RAM 1003. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0161] Generally, the following devices can be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 1008 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 1009. The communication device 1009 can allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange information. Although Figure 13 The electronic device 1000 having various devices is shown, but it should be understood that all the devices shown are not required to be implemented or provided. More or less devices can be alternatively implemented or provided.
[0162] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts, thereby implementing the interaction methods as described above. In such embodiments, the computer program can be downloaded and installed from a network by the communication device 1009, or installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.
[0163] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include an information signal in a baseband or as part of a carrier wave that carries computer-readable program code. Such a propagated signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code contained in a computer-readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, optical fiber, RF, etc., or any suitable combination of the foregoing.
[0164] In some embodiments, the client, server can communicate using any known or later developed network protocols, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital information communication (e.g., communication networks). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any known or later developed network.
[0165] The computer readable medium described above can be included in the electronic device described above; or can exist separately, without being assembled into the electronic device.
[0166] The computer readable medium described above carries one or more programs, which when executed by the electronic device, cause the electronic device to:
[0167] display a video generation configuration page, the video generation configuration page including a guide sentence input area, a first reference image adding option, and a first generation option;
[0168] receive a first reference image through the first reference image adding option, and display identification information of the received first reference image in the guide sentence input area; the number of the first reference images is multiple;
[0169] receive a text prompt through the guide sentence input area, and display the text prompt in the guide sentence input area; the text prompt includes an object description phrase, the object description phrase corresponding to the first reference image; the object description phrase is used to describe an object in the first reference image corresponding thereto; the identification information of the first reference image is interspersed in the text prompt; the number of characters between the position of the object description phrase in the text prompt and the position of the identification information of the first reference image corresponding thereto in the text prompt is less than a preset number of characters;
[0170] In response to a selection operation on the first generation option, generate a first video based on the content in the guide sentence input area.
[0171] Optionally, when the one or more programs are executed by the electronic device, the electronic device can further perform other steps described in the above embodiments.
[0172] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0173] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0174] The units described in the embodiments of the present disclosure can be implemented by hardware, software, or a combination of hardware and software. In some cases, the names of the units do not constitute a limitation on the units themselves.
[0175] The functions described in this specification can be implemented in part or in whole through one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0176] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores the program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more of an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0177] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, comprising:
[0178] one or more processors;
[0179] a memory for storing one or more programs;
[0180] when the one or more programs are executed by the one or more processors, the one or more processors implement any of the interaction methods as provided by the present disclosure.
[0181] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements any of the interaction methods as provided by the present disclosure.
[0182] The embodiments of the present disclosure also provide a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the interaction method as described above is implemented.
[0183] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0184] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An interaction method, characterized in that, include: The video generation configuration page is displayed, which includes a guide input area, a first reference image addition option, and a first generation option. The system receives a first reference image by adding an option to the first reference image, and displays the identification information of the received first reference image in the guide input area; the number of first reference images is multiple. The system receives and displays text prompts in the guidance input area. The text prompts include object description phrases that correspond to the first reference image. These object description phrases describe objects in the corresponding first reference image. Identification information of the first reference image is interspersed within the text prompts. The number of characters between the position of the object description phrase in the text prompts and the position of the identification information of the corresponding first reference image in the text prompts is less than a preset number of characters. In response to the selection of the first generation option, a first video is generated based on the content in the guide input area.
2. The method according to claim 1, characterized in that, The method further includes: In response to a first selection operation on the target first reference image in the guide input area, a first replacement option is displayed; In response to the selection of the first replacement option, the target first reference image is replaced, and the identification information of the target first reference image in the guide input area is replaced.
3. The method according to claim 1, characterized in that, The method further includes: In response to a second selection operation on the target first reference image in the guide input area, a deletion option is displayed; In response to the selection of the deletion option, the target first reference image is deleted, and the identification information of the target first reference image is deleted from the guidance input area.
4. The method according to claim 1, characterized in that, The method further includes: In response to the operation of publishing the second video as a template, a publishing configuration page is displayed; the second video includes the first video; the publishing configuration page includes identification information of the first reference image used by the first video and publishing options; In response to the third selection operation of the target first reference image on the publishing configuration page, a constraint input box corresponding to the first target slot is displayed; the first target slot is the material slot represented by the target first reference image. The content received through the constraint input box is used as the constraint condition corresponding to the first target slot. In response to the selection of the publishing option, the second video is published as a template based on the correspondence between the first target slot and the constraint conditions.
5. The method according to claim 4, characterized in that, In response to the selection of the publishing option, before publishing the second video as a template based on the correspondence between the first target slot and the constraints, the method further includes: In response to the fourth selection operation of the target first reference image on the release configuration page, the status of the first target slot is set to allow replacement.
6. The method according to claim 1, characterized in that, The method further includes: In response to the application operation of the third video template, the first slot filling page is displayed; the third video includes the first video; the first video corresponds to multiple first material slots; the first slot filling page includes the identification information of the first material slots and a second generation option; In response to the configuration operation of the target first material slot, a second reference image is configured to fill the target first material slot; In response to the selection of the second generation option, a fourth video is obtained based on the second reference image, the correspondence between the second reference image and the first material slot, and the third video template; the fourth video includes a fifth video, which is generated based on the second reference image.
7. The method according to claim 6, characterized in that, The method further includes: In response to the target first material slot being selected, the constraint conditions corresponding to the target first material slot are displayed on the first slot filling page.
8. The method according to claim 6, characterized in that, The method further includes: A video display page is provided, which includes a fourth video; the video display page also includes the identification information of the fifth video. In response to the selection operation of the fifth video on the video display page, a second replacement option corresponding to the fifth video is displayed; In response to the selection of the second replacement option, a second slot filling page is displayed, which includes the identification information of the first material slot and the constraints corresponding to the first material slot.
9. The method according to claim 8, characterized in that, The method further includes: In response to the selection of the fifth video on the video display page, an option to modify the introductory text is displayed; In response to the selection of the option to modify the introductory text, an introductory text corresponding to the fifth video is displayed; the introductory text includes identification information of the second reference image and text prompts; In response to the replacement operation of the second reference image in the prompt, the second reference image is replaced, and the identification information of the second reference image in the prompt is replaced.
10. The method according to claim 9, characterized in that, The method further includes: In response to a modification operation on the text prompt in the guidance, the text prompt in the guidance is modified.
11. An interactive device, characterized in that, include: The first display module is used to display the video generation configuration page, which includes a guide input area, a first reference image addition option, and a first generation option. The first collection module is used to receive the first reference image through the option to add the first reference image, and to display the identification information of the received first reference image in the guide input area; the number of the first reference images is multiple; The second collection module is used to receive text prompts through the guidance input area and display the text prompts in the guidance input area; the text prompts include object description phrases, which correspond to the first reference image; the object description phrases are used to describe the objects in the corresponding first reference image; the identification information of the first reference image is interspersed in the text prompts; the number of characters between the position of the object description phrase in the text prompts and the position of the identification information of the corresponding first reference image in the text prompts is less than a preset number of characters; A generation module is configured to generate a first video based on the content in the guide input area in response to a selection operation of the first generation option.
12. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-10.
Citation Information
Patent Citations
Text-based image generation method and device, electronic equipment and storage medium
CN117493599A
Video generation method, device and equipment, computer readable storage medium and product
CN118413717A
Video generation method and device, equipment and storage medium
CN118842959A
Multimedia resource generation method and device, electronic equipment and storage medium
CN120264087A
Interactive media data generation method and device, electronic equipment and storage medium
CN120296183A
Cited By
Video processing method and device, equipment, storage medium and product
CN121665056A