Content generation method and apparatus, electronic device, and storage medium
By determining the target perspective and using a content generation model to process image media to generate perspective effects, the problem of low video generation efficiency and difficulty in meeting user expectations in existing technologies has been solved, achieving efficient and realistic video generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-12-24
- Publication Date
- 2026-08-04
AI Technical Summary
Existing AI-based video generation methods require users to input complex descriptive terms, resulting in low video generation efficiency and quality that fails to meet user expectations.
By determining the target perspective, the content generation model is used to process image media to generate perspective effects content, simplifying user input and directly generating effects videos with the perspective as input.
It improves video generation efficiency, resulting in more consistent and authentic video content with higher quality.
Smart Images

Figure CN119850788B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence-generated content technology, and in particular to a content generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, interactive image generation based on artificial intelligence (AI) technology has become an increasingly popular content creation method. Users can input descriptive words and use the capabilities of artificial intelligence models to generate corresponding videos, which greatly improves the efficiency of content creation.
[0003] In existing technologies, AI model-based video generation technology requires users to describe the video content using complex descriptive words in order to generate the corresponding video. However, the design of descriptive words requires users to have a high level of experience in model operation, and it is difficult for ordinary users to quickly achieve reasonable and accurate input of descriptive words.
[0004] Therefore, existing AI-based solutions for creating special effects videos suffer from low video generation efficiency and difficulty in achieving user-expected video quality. Summary of the Invention
[0005] This disclosure provides a content generation method, apparatus, electronic device, and storage medium to overcome the problems of low video generation efficiency and difficulty in achieving user expectations in video quality.
[0006] In a first aspect, embodiments of this disclosure provide a content generation method, including:
[0007] In response to user operation, a target perspective is determined, which is used to characterize the observation perspective in a target state; an image medium is acquired, and a content generation model is invoked to process the image medium based on the target perspective to generate perspective effect content, wherein the perspective effect content has effect video footage, which is used to demonstrate the visual effect of observing the image medium from the target perspective.
[0008] Secondly, embodiments of this disclosure provide a content generation apparatus, including:
[0009] An interaction module is used to respond to user operations and determine a target perspective, which is used to characterize the observation perspective in a target state;
[0010] The generation module is used to acquire image media and call the content generation model to process the image media based on the target perspective to generate perspective effect content. The perspective effect content includes effect video footage, which is used to demonstrate the visual effect of observing the image media from the target perspective.
[0011] Thirdly, embodiments of this disclosure provide an electronic device, including: a processor and a memory;
[0012] The memory stores computer-executed instructions;
[0013] The processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the content generation method as described in the first aspect and various possible designs of the first aspect.
[0014] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the content generation method described in the first aspect and various possible designs of the first aspect.
[0015] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the content generation method described in the first aspect and various possible designs of the first aspect.
[0016] The content generation method, apparatus, electronic device, and storage medium provided in this embodiment determine a target perspective in response to user operations. This target perspective characterizes the observation perspective under a target state. An image medium is acquired, and a content generation model is invoked to process the image medium based on the target perspective, generating perspective effect content. This perspective effect content includes special effects video frames, which demonstrate the visual effects of observing the image medium from the target perspective. By responding to user operations, a target perspective characterizing the observation perspective of a target object is determined. Then, using the target perspective as input, an image generation model is invoked to process the image medium, generating perspective effect content. This perspective effect content includes the visual effects of observing the image medium from the target object's perspective, thus achieving video generation based on the perspective dimension. This process eliminates the need for users to input overly complex descriptive terms, effectively improving video generation efficiency. Furthermore, the generated perspective effect content is a video generated by the content generation model based on its understanding of the target object's observation perspective, resulting in better content consistency, higher realism, and improved video quality. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is an application scenario diagram of the content generation method provided in the embodiments of this disclosure;
[0019] Figure 2 Flowchart of the content generation method provided in the embodiments of this disclosure Figure 1 ;
[0020] Figure 3 A schematic diagram of a first page provided for an embodiment of this disclosure;
[0021] Figure 4 for Figure 2 A flowchart of a specific implementation of step S101 in the illustrated embodiment;
[0022] Figure 5 A schematic diagram of a special effects content page provided in an embodiment of this disclosure;
[0023] Figure 6 for Figure 2 A flowchart of another specific implementation of step S101 in the illustrated embodiment;
[0024] Figure 7 A schematic diagram of a second page provided for an embodiment of this disclosure;
[0025] Figure 8 for Figure 6 A flowchart illustrating the specific implementation of step S1013 in the illustrated embodiment;
[0026] Figure 9 A schematic diagram illustrating the generation of perspective effects content provided in an embodiment of this disclosure;
[0027] Figure 10 Flowchart of the content generation method provided in the embodiments of this disclosure Figure 2 ;
[0028] Figure 11 Structural block diagram of the content generation apparatus provided in the embodiments of this disclosure
[0029] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure;
[0030] Figure 13 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0032] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0033] The application scenarios of the embodiments of this disclosure are explained below:
[0034] The content generation method provided in this disclosure can be applied to applications (APPs) with video generation and video production functions, such as short video applications and video editing applications. More specifically, it can be applied to application scenarios such as AI model-based image-generated video, adding special effects to videos, and video continuation (video-to-video). The execution subject of this embodiment can be a terminal device running the aforementioned application with video generation function, a server deploying the server corresponding to the aforementioned application, or other electronic devices that perform similar functions. Specifically, when the execution subject is a terminal device, the terminal device executes the method provided in this embodiment by running the aforementioned application; when the execution subject is a server, the server of the aforementioned application with video generation function can run partially or entirely on the server, and the method provided in this embodiment is executed on the server side, while the terminal device runs the client of the application. The communication between the server and the terminal device is based on server-client communication, thereby enabling the terminal device to obtain the execution result of the method provided in this embodiment and display it as needed.
[0035] In some embodiments, the terminal device or server can implement the video generation method provided in this disclosure by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be program-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be local applications, i.e., programs that need to be installed in the operating system to run, or mini-programs embedded in any app, i.e., programs that run in a browser environment. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin; the specific implementation can be configured as needed. Furthermore, in implementing the video generation method provided in this disclosure, the video generation device can execute the method by running computer-executable instructions or computer programs set locally, or by calling computer-executable instructions or computer programs set in an external server. In some embodiments, the server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communication, cloud database, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. Among these, cloud services may be interactive processing services that can be invoked by terminal devices.
[0036] Figure 1 This is an application scenario diagram of the content generation method provided in the embodiments of this disclosure, with reference to... Figure 1As shown in the diagram, taking a terminal device as an example, a target application with video generation capabilities runs on the terminal device. The user can then input the image media to be processed and descriptive terms through this application. For example, the image media might be a picture, and the descriptive terms might be a description of the content of the special effects video to be generated. Next, based on the image media and descriptive terms, the terminal device calls a content generation model implemented using Artificial Intelligence Generated Content (AIGC) technology to generate a special effects video or image that combines the image content and the content described by the descriptive terms. For example, as shown in the diagram, the image media is a portrait photo P1, and the descriptive term is "make the person in the picture wear glasses." Based on this image media and descriptive terms, a special effects video V1 is generated. The content of this special effects video is the process of the "person" in the aforementioned portrait photo putting on glasses. This example illustrates the process of model-based image-generated video. In another possible application scenario, the image media input by the user can also be a video to be processed. The terminal device calls the content generation model based on the video to be processed and the descriptive words to generate a special effects video that combines the video content of the video to be processed and the content described by the descriptive words. That is, the process of generating video based on the model.
[0037] In existing technologies, in the aforementioned application scenarios of video generation based on AIGC technology, users are typically required to describe the video content using complex descriptive words in order to generate the corresponding video. For example, users need to describe the visual style, content changes, etc. However, the design of descriptive words requires users to have a high level of experience in model operation. Ordinary users find it difficult to quickly and accurately input reasonable descriptive words, resulting in low video generation efficiency and video quality that fails to meet user expectations.
[0038] This disclosure provides a content generation method to solve the above-mentioned problems.
[0039] refer to Figure 2 , Figure 2 Flowchart of the content generation method provided in the embodiments of this disclosure Figure 1 The method of this embodiment can be applied to terminal devices, and the content generation method includes:
[0040] Step S101: In response to user operation, determine the target viewpoint, which is used to characterize the observation viewpoint in the target state.
[0041] Step S102: Obtain the image media and call the content generation model to process the image media based on the target perspective to generate perspective effect content. The perspective effect content includes effect video footage, which is used to demonstrate the visual effects of observing the image media from the target perspective.
[0042] For example, refer to Figure 1 The illustrated application scenario diagram shows a terminal device running a target application, such as a video editing app, and displaying an interactive interface through this application. The user also receives user input through this interface. The specific implementation of the user operation is determined by the design of the target application's interactive interface. In one possible implementation, the target application's interactive interface includes a first page, and the user operation includes a selection operation on the first page. Before step S101, the following steps are also included:
[0043] Step S100A: Display the first page, which displays special effects components corresponding to at least two viewing angles.
[0044] Accordingly, the specific implementation of step S101 includes: in response to the selection operation of the target effect component in the first page, determining the viewing angle corresponding to the target effect component as the target viewing angle.
[0045] For example, the first page can be the homepage for perspective effects. Within this first page, multiple effect components are set up, each corresponding to a viewing perspective, to provide users with viewing perspectives to choose from in different states. In one possible implementation, the target perspective includes the viewing perspective of the target object, and the effect video footage is used to demonstrate the visual effects of observing image media from the viewing perspective of the target object. Figure 3 A schematic diagram of a first page provided for an embodiment of this disclosure, with reference to... Figure 3As shown, for example, the first page includes special effects components #1, #2, and #3. Special effects component #1 corresponds to the observation perspective of "the person who ate the poisonous mushroom," special effects component #2 corresponds to the observation perspective of "the alien," and special effects component #3 corresponds to the observation perspective of "the fly." Each special effects component is titled with the name of the corresponding observed object and the fixed descriptive phrase "the world as seen," i.e., the title of special effects component #1 is "the world as seen by the person who ate the poisonous mushroom," the title of special effects component #2 is "the world as seen by the alien," and the title of special effects component #3 is "the world as seen by the fly." The state corresponding to the observation perspective refers to the characteristics of the observation perspective. Therefore, the target perspective can be the observation perspective of the target observed object, such as "the perspective of the person who ate the poisonous mushroom." Alternatively, it can be a perspective that does not limit the target observed object but only limits the observation state, such as "the perspective of someone who ate the poisonous mushroom." Taking the observation perspective corresponding to effect component #1 as an example, the state corresponding to this observation perspective is "eating poisonous mushrooms," and its observation perspective is "human perspective." After combination, it serves as the target perspective, namely "the perspective of the person who ate the poisonous mushrooms." Furthermore, when a user performs a selection operation on the first page, for example, selecting "effect component #1" as the target effect component, the terminal device, based on this selection operation, uses the target perspective corresponding to "effect component #1," i.e., "the perspective of the person who ate the poisonous mushrooms," as the target perspective used in subsequent steps.
[0046] Furthermore, in one possible implementation, the selection operation includes a first selection operation and a second selection operation; that is, the selection operation consists of two operation steps: the first selection operation and the second selection operation. Correspondingly, as shown below... Figure 4 As shown, the specific implementation of step S101 includes:
[0047] Step S1011: In response to the second selection operation for the target effect component, display the effect content page corresponding to the target effect component. The effect content page displays at least two effect videos generated based on the viewing angle corresponding to the target effect component, and a confirmation component is set in the effect content page.
[0048] Step S1012: In response to the first selection operation for the confirmed component, the viewing angle corresponding to the target effect component is determined as the target viewing angle.
[0049] For example, after displaying the first page, firstly, in response to the second selection operation for the target effect component, the effect content page corresponding to the target effect component is displayed. The effect content page displays multiple effect videos; specifically, it displays the first frame, video cover, and other images of multiple effect videos. By further clicking on the images, the selected effect video can be played. This effect video is the effect video generated based on the viewing angle corresponding to the target effect component selected by the second selection operation.
[0050] Figure 5 This is a schematic diagram of a special effects content page provided in an embodiment of the present disclosure, such as... Figure 5 As shown, for reference Figure 3 In the illustrated embodiment, after the user performs a second selection operation (e.g., a click operation) and triggers the special effects component #3, the user enters the special effects content page corresponding to the special effects component #3. The special effects content page displays special effects videos generated based on the viewing angle corresponding to the target special effects component, such as special effects video A, special effects video B, special effects video C, and special effects video D shown in the figure. Optionally, for each special effects video, the user information corresponding to the special effects video is also displayed, such as the posting user name (represented by User_1, User_2, etc. in the figure), the number of likes, the number of reposts, etc. Furthermore, the aforementioned special effects video refers to the special effects video generated by the user based on the observation perspective corresponding to the target special effects component. More specifically, the observation perspective corresponding to the special effects component #3 is, for example, "the fly's perspective." As shown in the figure, the special effects video (the special effects video frame, i.e., a frame in the special effects video) generated based on the observation perspective presents the user-uploaded image media in a way that simulates "the fly's perspective" (the figure uses "hexagons" to simulate the "compound eyes" effect of a fly), thereby enabling the special effects video (the special effects video frame) to express the visual effect of observing the image media from "the fly's perspective." Afterwards, the user confirms the observation perspective corresponding to the target special effects component as the target perspective, i.e., "the fly's perspective," by clicking the confirmation component set in the special effects content page (first selection operation). In this embodiment, after the target effect component is triggered, the corresponding effect content page is first displayed, and the effect content page displays various visual effects of the effect video generated by the viewing angle corresponding to the target effect component. This realizes the aggregation of user-created content, improves the efficiency of information display, and enables users to match the target viewpoint of interest more quickly.
[0051] In this embodiment, a scheme for determining the target perspective is provided. That is, the user is given a choice of the pre-generated observation perspectives, and a target perspective is determined from the pre-generated observation perspectives based on the user's choice. This process does not require the user to input descriptive words, so it can improve the interaction efficiency, reduce the operation requirements of the user, and improve the generation efficiency of special effects videos.
[0052] Furthermore, in one possible implementation, a preview area is provided within the special effects component, and during or after the execution of step S100A, the following is also included:
[0053] The preview area of each special effects component on the first page displays recommended special effects videos corresponding to the special effects components. The recommended special effects videos are determined by sorting the special effects videos according to their video attributes, which are generated based on the viewing perspective corresponding to the special effects component. The video attributes include the video generation time and / or the video access popularity.
[0054] Among them, reference Figure 4 , Figure 5 As described in the illustrated embodiment, the preview area is used to display the special effects videos (the special effects video footage) in the special effects content page corresponding to the special effects component. Through this preview area, users can observe the content in the special effects content page without entering it, thereby improving interaction efficiency. Specifically, the preview area displays recommended special effects videos corresponding to the special effects component, that is, recommended special effects videos within the special effects content page. These recommended special effects videos are, for example, the top N special effects videos within the special effects content page, sorted by video attributes. Video attributes could be, for example, video generation time, meaning the N most recently generated special effects videos within the special effects content page are displayed in the preview area; or, attributes could be, for example, video popularity, determined by one or more of views, shares, and favorites, meaning the N special effects videos with the highest popularity within the special effects content page are displayed in the preview area.
[0055] In this embodiment, by configuring the preview area of the special effects component, the special effects component can display recommended special effects videos in the special effects content page through the preview area, thereby further improving the interaction efficiency.
[0056] In another possible implementation, user operations include input operations, specifically operations for inputting text. That is, the user determines the target perspective by inputting descriptive text, thereby achieving the generation of special effects video based on the target perspective, i.e., content generation. Specifically, before step S101, the following steps are also included:
[0057] Step S100B: Display the second page, which has a text input component.
[0058] Accordingly, such as Figure 6 As shown, in another possible implementation, the specific implementation of step S101 includes:
[0059] Step S1013: In response to an input operation on the text input component in the second page, generate target text, which is used to determine the target observation object.
[0060] Step S1014: Determine the target perspective based on the target text.
[0061] For example, in another implementation, before step S101, a second page is first displayed. This second page is used to receive text input by the user. By displaying adaptation prompts on the second page, the user is guided to input relevant text, thereby enabling the terminal device to generate target text based on the user's input. The target text can be the object of observation from the target perspective. Figure 7 A schematic diagram of a second page provided in an embodiment of this disclosure, with reference to... Figure 7 As shown, for example, the second page can be accessed from the first page. For instance, as shown in the figure, the first page contains a "Create Perspective" control. Clicking this control redirects the user to the second page, which contains a fixed prompt, such as "I want to observe the world from the perspective of _____". Then, based on the prompt, the user enters "alien" (the target text) in the text input component (e.g., the underlined "_____" position shown in the figure). After the user clicks the "OK" button, the terminal device determines the corresponding target perspective based on this text, which is "alien perspective". Subsequent processing steps are then performed based on this target perspective until a corresponding video or image of the input image media observed from an alien perspective is generated, i.e., perspective effect content.
[0062] Furthermore, in one possible implementation, the input operation includes a first trigger operation and a second trigger operation. The second page is configured with a random generation component for randomly generating random objects with complete semantics. Accordingly, such as... Figure 8 As shown, the specific implementation of step S1013 includes:
[0063] Step S1013-1: In response to the first triggering operation of the random generation component on the second page, randomly generated text is displayed in the text input component on the second page. The randomly generated text is used to characterize a randomly generated observation object.
[0064] Step S1013-2: In response to the second trigger operation for the second page, determine the randomly generated text as the target text.
[0065] For example, after a user triggers the random generation component of the second page through the first trigger operation, the terminal device will display randomly generated text in the text input component of the second page. The randomly generated text may be, for example, "elephant," "fly," or "alien." This randomly generated text can be a word randomly selected from a pre-generated lexicon used to store the names of observed objects, or it can be a word randomly generated by a large language model that can represent the observed object. Further, the randomly generated text can consist of at least two parts: the first part is the text representing the name of the observed object, such as "elephant" or "fly," and the second part is the limiting text describing the state of the standard observed object, such as "moving" or "ate a poisonous mushroom." The specific generation method of the randomly generated text can be set as needed, and no specific restrictions are imposed here. Then, in response to the second trigger operation for the second page, refer to... Figure 7 As shown, the second triggering operation is, for example, a click operation on the "Confirm" button on the second page, which determines the randomly generated text as the target text and executes subsequent steps.
[0066] Furthermore, after obtaining and determining the target viewpoint through the interaction process in step S101, the terminal device further acquires image media, which can be images or videos. This image media can be stored locally on the terminal device or in the cloud, and is obtained in response to the user's media selection operation. Furthermore, the image media can be a single frame of an image or a segment of video, or it can be multiple frames of images or multiple segments of video. When the image media is multiple frames of images or multiple segments of video, the terminal device can acquire these multiple frames of images or multiple segments of video from the local or cloud media library all at once, or it can acquire them in several installments in response to multiple media selection operations by the user. The specific implementation method can be set as needed.
[0067] Subsequently, the terminal device processes the aforementioned image media and target perspective by invoking the content generation model. That is, using the image media and target perspective as input to the content generation model, and leveraging the video generation capabilities of the video generation module, it generates perspective effects content with the perspective effect of observing the image media from the viewpoint of the target object. In this embodiment, the content generation model possesses semantic understanding, reasoning, and video generation capabilities. It can understand the meaning of the input target perspective, infer the visual effect of viewing an object from that perspective, and generate corresponding visual effects videos based on this visual effect and the image media. The content generation model can be deployed locally on the terminal device or in the cloud. Specifically, the video generation module has the ability to understand the target perspective and add visual effects to the target image based on the target perspective. When the image media is video, the video generation module adds the aforementioned perspective effects to at least one video frame (usually all or most video frames) and / or generates new video frames with the aforementioned perspective effects (i.e., video continuation), thereby enabling the processed video to represent the viewpoint of the target object. When the image medium is a picture, the video generation module will generate more pictures based on that picture, thus forming a video (i.e., image-generated video). The generated pictures will have the visual effect of being viewed from the perspective of the target object, so that the processed video has the ability to represent the perspective of the target object.
[0068] Meanwhile, the perspective effects content generated based on the above steps has at least one of the following target features: target camera movement features, target lens angle, and target art style features; wherein, the target camera movement features are used to characterize the camera movement pattern of the video frame of the perspective effects content; the target lens angle characterizes the lens angle of the video frame of the perspective effects content; the target art style features characterize the art style of the video frame of the perspective effects content; the target features are determined based on the target state or target observation object corresponding to the target perspective.
[0069] Figure 9 This is a schematic diagram illustrating the generation of viewpoint effects content provided in an embodiment of this disclosure, such as... Figure 9As shown, firstly, based on the interaction between the terminal device and the user, the terminal device acquires image media and a target perspective. The target perspective can be represented by text, such as text T1 shown in the figure. More specifically, the content of text T1 is, for example, "the fly's perspective," while the image media is, for example, image P1. Next, the terminal device inputs the aforementioned image P1 and text T1 into a content generation model. This video generation model includes an inference sub-model and an image generation sub-model. After processing by the inference sub-model and the image generation sub-model respectively, a perspective effect video V1 (i.e., perspective effect content) is output. Referring to the figure, firstly, the content features of the effect video frame in perspective effect video V1 match the target perspective (i.e., "the fly's perspective"), that is, simulating "the fly's perspective" to observe the content of image P1 (refer to video frame P2 in perspective effect video V1 in the figure, where a "hexagon" is used to represent the fly's "compound eye" visual effect). Secondly, optionally, the perspective effect video V1 can also have target visual style characteristics that match the target observation perspective, such as a "high contrast" style, to simulate and represent the world as seen from the target perspective. Thirdly, optionally, during playback, the perspective effect video V1 can also have target camera movement characteristics and target lens angles. For example, during playback, the perspective effect video V1 may present a visual effect of rapid left-right camera movement to simulate the behavioral characteristics of the target observed object; and it may also observe objects from a top-down perspective to simulate the observation point position of the target observed object. For example, when the observation point of the target observed object is relatively high (e.g., the target observed object is an "elephant," and the target perspective is "the elephant's perspective"), the perspective effect video V1 may present a visual effect of observing objects in image P1 from a "top-down" perspective.
[0070] Of course, it is understandable that in other possible implementations, the image medium can also be video, such as video V0. Similar to the above process, after inputting text T1 and video V0 into the content generation model, the content generation model will generate the corresponding perspective effect video (i.e., perspective effect content) after understanding and reasoning based on the content of text T1. The specific implementation process will not be elaborated here.
[0071] By inputting the target perspective into the content generation model, the model processes image media based on its understanding of the target perspective, thereby generating self-consistent, smooth, and realistic perspective effects content. This enables control over multiple features such as video frames and camera movement. It is essentially a result-oriented AI video generation technology that uses "perspective" to represent "effects," thus enabling the expression of complex video effects without requiring user control over the changes in the generated video, thereby improving video generation efficiency and video quality.
[0072] In this embodiment, a target perspective is determined in response to user operation. The target perspective represents the observation perspective in a target state. Image media is acquired, and a content generation model is invoked to process the image media based on the target perspective, generating perspective effect content. This perspective effect content includes special effects video frames, which demonstrate the visual effects of observing the image media from the target perspective. By responding to user operation, a target perspective representing the observation perspective of the target object is determined. Then, using the target perspective as input, the image generation model is invoked to process the image media, generating perspective effect content. This perspective effect content includes the visual effects of observing the image media from the target object's observation perspective, thus realizing video generation based on the perspective dimension. This process does not require users to input overly complex descriptive words, thus effectively improving the efficiency of video generation. Furthermore, the generated perspective effect content is a video generated by the content generation model based on its own understanding of the target object's observation perspective, resulting in better content consistency, higher realism, and improved video quality.
[0073] refer to Figure 10 , Figure 10 Flowchart of the content generation method provided in the embodiments of this disclosure Figure 2 This embodiment is in Figure 2 Based on the illustrated embodiment, the interaction process is further refined, and the content generation method includes:
[0074] Step S201: Display the second page, which has a text input component.
[0075] Step S202: In response to an input operation on the text input component in the second page, generate target text, which is used to determine the target observation object.
[0076] Step S203: Generate a first prompt word based on the semantics of the target text. The first prompt word is used to characterize the perspective features of the observation perspective.
[0077] Step S204: Process the first cue word using a pre-trained large language model to generate an observation object name. The observation object name is used to indicate the observation object with the perspective features of the observation perspective indicated by the first cue word.
[0078] Step S205: Determine the target perspective based on the name of the object being observed.
[0079] For example, the second page is a text input page, see reference. Figure 2In the embodiment shown, after the terminal device displays the second page, the user inputs target text by applying an input operation to the text input component on the second page. The target text is information that characterizes the features of the target object. For example, the content of the target text is "a flying perspective that can see the distant horizon". At this time, the target text can characterize some features of the target object, but does not directly indicate a specific target object.
[0080] In this scenario, the terminal device, for example, invokes a large language model to perform semantic analysis on the target text and generates a corresponding first prompt word based on the semantics of the target text. This first prompt word represents the perspective features of the observation viewpoint. For example, the first prompt word generated for the target text is "has a high-altitude perspective, can move quickly in the air,...". Then, a pre-trained large language model processes the first prompt word to predict the observed object with the aforementioned perspective features, generating the object's name, such as "drone" or "bird." Furthermore, the target perspective is determined, such as "drone perspective" or "bird perspective."
[0081] In this embodiment, by combining a large language model to process the target file input by the user, a target perspective recommended by the model is generated. This allows users to generate perspective effects content without being limited to using existing or common perspectives, greatly improving the diversity and flexibility of perspectives and enhancing video creation and interaction efficiency.
[0082] Step S206: Generate a second prompt word based on the target perspective. The second prompt word is used to instruct the content generation model to generate video based on the observation perspective of the target object.
[0083] Step S207: Generate a special effects video by combining the second prompt word and the image media input content into a model.
[0084] Furthermore, after obtaining the target perspective, one possible implementation is to directly input the target perspective (text) and media data into the content generation model, and utilize the content generation model's own reasoning and video generation capabilities to generate corresponding perspective effects content (i.e., Figure 2(The implementation method in the illustrated embodiment). In another possible implementation, a second prompt word can be generated first based on the target perspective. Then, the second prompt word is input into the content generation model to guide the video generation mode to generate the corresponding video. One possible implementation is that the second prompt word can be descriptive text generated by analyzing the characteristics of the target observation object corresponding to the target perspective. For example, the target perspective is "the perspective of a fly," and the corresponding target observation object is "a fly." After processing it based on a large language model, the corresponding descriptive text is generated as "a flying insect that moves quickly, has an irregular flight trajectory, has compound eyes, and a wide field of vision...". Then, based on the above descriptive text and the corresponding prompt word template, a second prompt word is generated to instruct the content generation model to infer the visual effects of the target perspective based on the content of the above descriptive text, and thus generate the corresponding special effects video. The solution in this embodiment is equivalent to decentralizing the reasoning ability of the content generation model to the outside, using other language models as a supplement to the reasoning ability, thereby improving the performance of the content generation model, reducing the training cost of the content generation model, and ultimately improving the quality of the generated video.
[0085] Furthermore, in one possible implementation, the second prompt word includes a first segment and at least one of a second and a third segment. The first segment characterizes the target observation object, the second segment characterizes the height and / or pitch angle of the observation point of the target observation object, and the third segment characterizes the motion pattern of the observation point and / or the motion pattern of the observation angle of the target observation object. For example, the content of the second prompt word is "a flying insect with compound eyes, moving at a height of less than 10 meters, with an irregular flight trajectory," where "a flying insect with compound eyes" is the first segment; "moving at a height of less than 10 meters" is the second segment; and "irregular flight trajectory" is the third segment. By using a second prompt word input content generation model composed of one or more of the above segment words, the aim is to control the visual style of the video frame (first segment), the camera angle of the video frame (second segment), and the camera motion pattern of the video frame (third segment) of the viewpoint effect content, thereby enabling the generated viewpoint effect content to possess target characteristics matching the target observation object. This enables complex video effects expressions using perspective effects, and the process requires no user control over the changes in the generated video, thus improving video generation efficiency and quality. The specific content and representation of the target features are... Figure 2 The embodiments shown have already been described, and will not be repeated here.
[0086] Corresponding to the content generation method in the above embodiments, Figure 11This is a structural block diagram of a content generation apparatus provided in an embodiment of this disclosure. The method described in the above embodiments can be executed by this content generation apparatus, which can be implemented by software and / or hardware, and can be integrated into an electronic device with certain data processing capabilities. The electronic device may include, but is not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities such as desktop computers and supercomputers.
[0087] For ease of explanation, only the parts relevant to embodiments of this disclosure are shown. (Refer to...) Figure 11 The content generation device 3 includes:
[0088] Interaction module 31 is used to respond to user operations and determine the target perspective, which is used to characterize the observation perspective in the target state;
[0089] The generation module 32 is used to acquire image media and call the content generation model to process the image media based on the target perspective to generate perspective effect content. The perspective effect content includes effect video footage, which is used to demonstrate the visual effects of observing the image media from the target perspective.
[0090] According to one or more embodiments of this disclosure, the user operation includes a selection operation, and the interaction module 31 is further configured to: display a first page, on which at least two special effect components corresponding to viewing angles are displayed; when the interaction module 31 determines the target viewing angle in response to the user operation, it is specifically configured to: in response to the selection operation of the target special effect component on the first page, determine the viewing angle corresponding to the target special effect component as the target viewing angle.
[0091] According to one or more embodiments of this disclosure, the selection operation includes a first selection operation and a second selection operation. When the interaction module 31 determines the viewing angle corresponding to the target effect component as the target viewing angle in response to the selection operation of the target effect component in the first page, it is specifically used to: in response to the second selection operation of the target effect component, display the effect content page corresponding to the target effect component, the effect content page displays at least two effect videos generated based on the viewing angle corresponding to the target effect component, and the effect content page is provided with a confirmation component; in response to the first selection operation of the confirmation component, determine the viewing angle corresponding to the target effect component as the target viewing angle.
[0092] According to one or more embodiments of this disclosure, a preview area is provided within the special effects component. The interaction module 31 is further configured to: display recommended special effects videos corresponding to the special effects components in the preview area of each special effects component on the first page; wherein, the recommended special effects videos are videos determined by sorting the video attributes of each special effects video among the special effects videos generated based on the viewing angle corresponding to the special effects component; the video attributes include the video generation time and / or the video access popularity.
[0093] According to one or more embodiments of this disclosure, user operations include input operations; the interaction module 31 is further configured to: display a second page, the second page being provided with a text input component; when the interaction module 31 determines a target perspective in response to a user operation, it is specifically configured to: generate target text in response to an input operation on the text input component in the second page, the target text being used to determine the target observation object; and determine the target perspective based on the target text.
[0094] According to one or more embodiments of this disclosure, when determining the target perspective based on the target text, the interaction module 31 is specifically configured to: generate a first prompt word based on the semantics of the target text, the first prompt word being used to characterize the perspective features of the observation perspective; process the first prompt word through a pre-trained large language model to generate an observation object name, the observation object name being used to indicate an observation object having the perspective features of the observation perspective indicated by the first prompt word; and determine the target perspective based on the observation object name.
[0095] According to one or more embodiments of this disclosure, the input operation includes a first trigger operation and a second trigger operation. A random generation component is configured in the second page. When the interaction module 31 generates target text in response to an input operation on the text input component in the second page, it is specifically used to: display the randomly generated text in the text input component of the second page in response to the first trigger operation on the random generation component of the second page, wherein the randomly generated text is used to characterize the randomly generated observation object; and determine the randomly generated text as the target text in response to the second trigger operation on the second page.
[0096] According to one or more embodiments of this disclosure, the viewpoint effect content has at least one of the following target features: target camera movement features, target lens angle, and target art style features; wherein, the target camera movement features are used to characterize the camera movement pattern of the video frame of the viewpoint effect content; the target lens angle characterizes the lens angle of the video frame of the viewpoint effect content; the target art style features characterize the art style of the video frame of the viewpoint effect content; the target features are determined based on the target state or target observation object corresponding to the target viewpoint.
[0097] According to one or more embodiments of this disclosure, the generation module 32 is specifically used to: obtain a second prompt word corresponding to the target viewpoint, the second prompt word being used to instruct the content generation model to generate a video based on the viewpoint of the target object; input the second prompt word and image media into the content generation model to generate a special effects video.
[0098] According to one or more embodiments of this disclosure, the second prompt word includes a first segment and at least one of a second segment and a third segment, wherein the first segment is used to characterize the target observation object, the second segment is used to characterize the height and / or pitch angle of the observation point of the target observation object, and the third segment is used to characterize the motion law of the observation point of the target observation object and / or the motion law of the observation angle.
[0099] The interaction module 31 and the generation module 32 are connected. The content generation device 3 provided in this embodiment can execute the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.
[0100] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, such as... Figure 12 As shown, the electronic device 4 includes:
[0101] Processor 41, and memory 42 communicatively connected to processor 41;
[0102] Memory 42 stores instructions executed by the computer;
[0103] The processor 41 executes computer execution instructions stored in the memory 42 to achieve, for example, Figures 2-10 The content generation method in the illustrated embodiment.
[0104] Optionally, the processor 41 and the memory 42 are connected via a bus 43.
[0105] For relevant instructions, please refer to the corresponding text. Figures 2-10 The relevant descriptions and effects of the steps in the corresponding embodiments are understood, and will not be elaborated on here.
[0106] This disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement this disclosure. Figures 2-10 The content generation provided by any of the corresponding embodiments.
[0107] This disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements this disclosure. Figures 2-10 The content generation method provided in any of the corresponding embodiments.
[0108] To implement the above embodiments, this disclosure also provides an electronic device.
[0109] refer to Figure 13 The diagram illustrates a structural schematic of an electronic device 900 suitable for implementing embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 13 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0110] like Figure 13 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0111] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 13 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0112] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0113] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0114] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0115] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0116] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0117] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0118] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units or modules do not necessarily limit the specific unit itself.
[0119] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0120] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0121] In a first aspect, according to one or more embodiments of this disclosure, a content generation method is provided, comprising:
[0122] In response to user operation, a target perspective is determined, which is used to characterize the observation perspective in a target state; an image medium is acquired, and a content generation model is invoked to process the image medium based on the target perspective to generate perspective effect content, wherein the perspective effect content has effect video footage, which is used to demonstrate the visual effect of observing the image medium from the target perspective.
[0123] According to one or more embodiments of this disclosure, the user operation includes a selection operation, and the method further includes: displaying a first page, wherein the first page displays at least two special effect components corresponding to viewing angles; the step of determining the target viewing angle in response to the user operation includes: in response to a selection operation on a target special effect component in the first page, determining the viewing angle corresponding to the target special effect component as the target viewing angle.
[0124] According to one or more embodiments of this disclosure, the selection operation includes a first selection operation and a second selection operation. The step of determining the viewing angle corresponding to the target effect component as the target viewing angle in response to the selection operation for the target effect component on the first page includes: in response to the second selection operation for the target effect component, displaying an effect content page corresponding to the target effect component, the effect content page displaying at least two effect videos generated based on the viewing angle corresponding to the target effect component, and the effect content page having a confirmation component; and in response to the first selection operation for the confirmation component, determining the viewing angle corresponding to the target effect component as the target viewing angle.
[0125] According to one or more embodiments of this disclosure, the special effects component is provided with a preview area, and the method further includes: displaying recommended special effects videos corresponding to the special effects component in the preview area of each special effects component on the first page; wherein, the recommended special effects videos are videos determined by sorting the special effects videos according to the video attributes of each special effects video among the special effects videos generated based on the viewing angle corresponding to the special effects component; the video attributes include video generation time and / or video access popularity.
[0126] According to one or more embodiments of this disclosure, the user operation includes an input operation; the method further includes: displaying a second page, the second page being provided with a text input component; the step of determining a target perspective in response to the user operation includes: generating target text in response to an input operation on the text input component in the second page, the target text being used to determine the target observation object; and determining a target perspective based on the target text.
[0127] According to one or more embodiments of this disclosure, determining a target perspective based on the target text includes: generating a first cue word based on the semantics of the target text, the first cue word being used to characterize the perspective features of the observation perspective; processing the first cue word through a pre-trained large language model to generate an observation object name, the observation object name being used to indicate an observation object having the perspective features of the observation perspective indicated by the first cue word; and determining the target perspective based on the observation object name.
[0128] According to one or more embodiments of this disclosure, the input operation includes a first trigger operation and a second trigger operation. A random generation component is configured within the second page. Generating target text in response to an input operation on the text input component within the second page includes: displaying randomly generated text within the text input component of the second page in response to the first trigger operation on the random generation component of the second page, the randomly generated text being used to characterize a randomly generated observation object; and determining the randomly generated text as the target text in response to the second trigger operation on the second page.
[0129] According to one or more embodiments of this disclosure, the perspective effect content has at least one of the following target features: target camera movement features, target lens angle, and target art style features; wherein, the target camera movement features are used to characterize the camera movement pattern of the video frame of the perspective effect content; the target lens angle characterizes the lens angle of the video frame of the perspective effect content; the target art style features characterize the art style of the video frame of the perspective effect content; the target features are determined based on the target state or target observation object corresponding to the target perspective.
[0130] According to one or more embodiments of this disclosure, invoking a content generation model to process the image media based on the target viewpoint and generate viewpoint effect content includes: obtaining a second prompt word corresponding to the target viewpoint, the second prompt word being used to instruct the content generation model to generate a video based on the observation viewpoint of the target object; inputting the second prompt word and the image media into the content generation model to generate the effect video.
[0131] According to one or more embodiments of this disclosure, the second prompt word includes a first segmentation word, and at least one of a second segmentation word and a third segmentation word, wherein the first segmentation word is used to characterize the target observation object, the second segmentation word is used to characterize the height and / or pitch angle of the observation point of the target observation object, and the third segmentation word is used to characterize the motion law of the observation point of the target observation object and / or the motion law of the observation angle.
[0132] Secondly, according to one or more embodiments of this disclosure, a content generation apparatus is provided, comprising:
[0133] An interaction module is used to respond to user operations and determine a target perspective, which is used to characterize the observation perspective in a target state;
[0134] The generation module is used to acquire image media and call the content generation model to process the image media based on the target perspective to generate perspective effect content. The perspective effect content includes effect video footage, which is used to demonstrate the visual effect of observing the image media from the target perspective.
[0135] According to one or more embodiments of this disclosure, the user operation includes a selection operation, and the interaction module is further configured to: display a first page, wherein the first page displays special effect components corresponding to at least two viewing angles; when the interaction module determines the target viewing angle in response to the user operation, it is specifically configured to: in response to the selection operation of the target special effect component in the first page, determine the viewing angle corresponding to the target special effect component as the target viewing angle.
[0136] According to one or more embodiments of this disclosure, the selection operation includes a first selection operation and a second selection operation. When the interaction module determines the viewing angle corresponding to the target effect component as the target viewing angle in response to the selection operation of the target effect component in the first page, it is specifically used to: in response to the second selection operation of the target effect component, display the effect content page corresponding to the target effect component, the effect content page displays at least two effect videos generated based on the viewing angle corresponding to the target effect component, and the effect content page is provided with a confirmation component; in response to the first selection operation of the confirmation component, determine the viewing angle corresponding to the target effect component as the target viewing angle.
[0137] According to one or more embodiments of this disclosure, the special effects component is provided with a preview area, and the interaction module is further configured to: display recommended special effects videos corresponding to the special effects component in the preview area of each special effects component on the first page; wherein, the recommended special effects videos are videos determined by sorting the special effects videos according to the video attributes of each special effects video among the special effects videos generated based on the viewing perspective corresponding to the special effects component; the video attributes include video generation time and / or video access popularity.
[0138] According to one or more embodiments of this disclosure, the user operation includes an input operation; the interaction module is further configured to: display a second page, the second page being provided with a text input component; when the interaction module determines a target perspective in response to a user operation, it is specifically configured to: generate target text in response to an input operation on the text input component in the second page, the target text being used to determine the target observation object; and determine the target perspective based on the target text.
[0139] According to one or more embodiments of this disclosure, when the interaction module determines the target perspective based on the target text, it is specifically configured to: generate a first prompt word based on the semantics of the target text, wherein the first prompt word is used to characterize the perspective features of the observation perspective; process the first prompt word through a pre-trained large language model to generate an observation object name, wherein the observation object name is used to indicate an observation object having the perspective features of the observation perspective indicated by the first prompt word; and determine the target perspective based on the observation object name.
[0140] According to one or more embodiments of this disclosure, the input operation includes a first trigger operation and a second trigger operation. A random generation component is configured within the second page. When the interaction module generates target text in response to an input operation on the text input component within the second page, it specifically performs the following: in response to the first trigger operation on the random generation component of the second page, displays randomly generated text within the text input component of the second page, the randomly generated text being used to characterize a randomly generated observation object; and in response to the second trigger operation on the second page, determines the randomly generated text as the target text.
[0141] According to one or more embodiments of this disclosure, the perspective effect content has at least one of the following target features: target camera movement features, target lens angle, and target art style features; wherein, the target camera movement features are used to characterize the camera movement pattern of the video frame of the perspective effect content; the target lens angle characterizes the lens angle of the video frame of the perspective effect content; the target art style features characterize the art style of the video frame of the perspective effect content; the target features are determined based on the target state or target observation object corresponding to the target perspective.
[0142] According to one or more embodiments of this disclosure, the generation module is specifically configured to: obtain a second prompt word corresponding to the target viewpoint, the second prompt word being used to instruct the content generation model to generate a video based on the viewpoint of the target observation object; input the second prompt word and the image media into the content generation model to generate the special effects video.
[0143] According to one or more embodiments of this disclosure, the second prompt word includes a first segmentation word, and at least one of a second segmentation word and a third segmentation word, wherein the first segmentation word is used to characterize the target observation object, the second segmentation word is used to characterize the height and / or pitch angle of the observation point of the target observation object, and the third segmentation word is used to characterize the motion law of the observation point of the target observation object and / or the motion law of the observation angle.
[0144] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0145] The memory stores computer-executed instructions;
[0146] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the content generation method as described in the first aspect and various possible designs of the first aspect.
[0147] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, and when a processor executes the computer-executable instructions, the content generation method described in the first aspect and various possible designs of the first aspect is implemented.
[0148] Fifthly, according to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the content generation method as described in the first aspect and various possible designs of the first aspect.
[0149] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0150] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0151] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A content generation method, characterized in that, include: In response to user operation, a target perspective is determined, which is used to characterize the observation perspective in a target state; wherein, the state corresponding to the observation perspective refers to the characteristics of the observation perspective; the target perspective includes a fictional observation perspective determined based on the target text input by the user; The system acquires an image medium, which may be a picture or a video; and acquires a second prompt word corresponding to the target viewpoint. The second prompt word and the image medium are then input into a content generation model to generate viewpoint effect video content. The second prompt word instructs the content generation model to generate a video based on the viewpoint of the target object. The second prompt word is generated according to the target viewpoint. The second prompt word includes a first word segment and at least one of a second word segment and a third word segment. The first word segment represents the target object, the second word segment represents the height and / or pitch angle of the target object's observation point, and the third word segment represents the motion law of the target object's observation point and / or the motion law of its observation angle. If the image medium is a picture, then a video content with perspective effects containing multiple frames is generated; if the image medium is a video, then perspective effects are added to the video to generate new video content with perspective effects; wherein, the content generation model has semantic understanding ability, reasoning ability, and video generation ability. The perspective effect video content has special effects video footage, which is used to demonstrate the visual effect of viewing the image media from the target perspective; the perspective effect video content has at least one of the following target features: target camera movement features, target lens angle, and target art style features; The target camera movement features are used to characterize the camera movement patterns of the video frames in the video content with the perspective effects. The target lens angle represents the lens angle of the video frame in the video content of the perspective effect video content; The target art style features characterize the visual style of the video frame in the video content with the perspective effects; The target features are determined based on the target state or target observation object corresponding to the target perspective.
2. The method according to claim 1, characterized in that, The user operation includes a selection operation, and the method further includes: The first page is displayed, which contains special effects components corresponding to at least two viewing angles. The process of determining the target viewpoint in response to user actions includes: In response to the selection operation of the target effect component within the first page, the viewing angle corresponding to the target effect component is determined as the target viewing angle.
3. The method according to claim 2, characterized in that, The selection operation includes a first selection operation and a second selection operation. The step of responding to the selection operation on the target effect component within the first page by determining the viewing angle corresponding to the target effect component as the target viewing angle includes: In response to a second selection operation for a target effect component, an effect content page corresponding to the target effect component is displayed. The effect content page displays at least two effect videos generated based on the viewing angle corresponding to the target effect component. A confirmation component is provided in the effect content page. In response to the first selection operation for the confirmed component, the viewing angle corresponding to the target effect component is determined as the target viewing angle.
4. The method according to claim 2, characterized in that, The special effects component includes a preview area, and the method further includes: In the preview area of each of the special effects components on the first page, recommended special effects videos corresponding to the special effects components are displayed; The recommended special effects video is a video determined by sorting the special effects videos according to their video attributes from the special effects videos generated based on the viewing angle corresponding to the special effects component. The video attributes include the video generation time and / or the video access popularity.
5. The method according to claim 1, characterized in that, The target perspective includes the perspective of the target object being observed; the user operation includes input operations; the method further includes: Display a second page, which includes a text input component; The process of determining the target viewpoint in response to user actions includes: In response to an input operation on a text input component within the second page, target text is generated, which is used to determine the target observation object; Determine the target perspective based on the target text.
6. The method according to claim 5, characterized in that, Based on the target text, determine the target perspective, including: Based on the semantics of the target text, a first prompt word is generated, which is used to characterize the perspective features of the observation perspective; The first prompt word is processed by a pre-trained model to generate an observation object name, which is used to indicate an observation object with the viewpoint features indicated by the first prompt word. Determine the target perspective based on the name of the observed object.
7. The method according to claim 5, characterized in that, The input operation includes a first trigger operation and a second trigger operation. The second page is configured with a random generation component. The step of generating target text in response to an input operation on the text input component within the second page includes: In response to a first trigger operation on the randomized component of the second page, randomly generated text is displayed in the text input component of the second page, the randomly generated text being used to characterize the randomly generated observation object; In response to a second triggering operation on the second page, the randomly generated text is determined as the target text.
8. A content generation apparatus, characterized in that, include: An interaction module is used to respond to user operations and determine a target perspective, which is used to characterize the observation perspective in a target state; wherein, the state corresponding to the observation perspective refers to the characteristics of the observation perspective; the target perspective includes a fictional observation perspective determined based on the target text input by the user; A generation module is used to acquire image media, which may be a picture or a video; and to acquire a second prompt word corresponding to the target viewpoint, inputting the second prompt word and the image media into a content generation model to generate viewpoint effect video content; wherein, the second prompt word is used to instruct the content generation model to generate video based on the observation viewpoint of the target object, and the second prompt word is generated according to the target viewpoint; the second prompt word includes a first word segment, and at least one of a second word segment and a third word segment, the first word segment being used to characterize the target object, the second word segmenting the height and / or pitch angle of the observation point of the target object, and the third word segmenting the motion law of the observation point and / or the motion law of the observation angle of the target object; if the image media is a picture, then viewpoint effect video content containing multiple frames is generated; if the... If the image media is video, then perspective effects are added to the video to generate new perspective special effects video content; wherein, the content generation model has semantic understanding ability, reasoning ability, and video generation ability; the perspective special effects video content has special effects video frames, which are used to demonstrate the visual effects of observing the image media from the target perspective; the perspective special effects video content has at least one of the following target features: target camera movement features, target lens angle, and target style features; wherein, the target camera movement features are used to characterize the camera movement rules of the video frames of the perspective special effects video content; the target lens angle characterizes the lens angle of the video frames of the perspective special effects video content; the target style features characterize the visual style of the video frames of the perspective special effects video content; the target features are determined based on the target state or target observation object corresponding to the target perspective.
9. An electronic device, characterized in that, include: Processor and memory; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the content generation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the content generation method as described in any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the content generation method as described in any one of claims 1 to 7.