Reference object processing method and device, visual data generation method and device, equipment, storage medium and program product
By generating reference objects containing reference images and attribute information, the problems of low quality of visual data generation and waste of resources in the prior art are solved, and higher quality video generation and resource conservation are achieved.
Patent Information
- Application Number
- CN202510343574.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-27
AI Technical Summary
Existing video generation technology is difficult to achieve the expected results, the generated visual data is of poor quality, and each generation task requires a large amount of computing and storage resources.
By generating a reference object containing reference image and attribute information, and generating expected visual data based on the reference object, the generation quality is improved and resources are saved.
It improves the quality of visual data generation, expands the application scenarios of reference objects, and saves computing resources and storage resources.
Smart Images

Figure CN120215786A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of video generation, and in particular, to a method for processing a reference object, a method for generating visual data, a device, a device, a storage medium, and a program product. Background Art
[0002] The development of deep learning has laid a foundation for technologies such as text-to-image and image-to-video. Through deep learning models, text and image information can be effectively processed to achieve the mapping from text descriptions to image content, and the mapping from images to video content. People are also keen on integrating their favorite concepts or characters into the generated works through model fine-tuning.
[0003] However, related technologies usually directly generate the final desired visual data (such as images or videos) based on the input images and texts through a pre-trained visual data generation model. However, the generated visual data often fails to meet people's expectations, and the quality of visual data generation is poor; moreover, for each visual data generation task, it is necessary to collect input data such as images and texts and perform corresponding processing on the input data. Therefore, it will greatly waste the computing resources and storage resources of the visual data generation system. Summary of the Invention
[0004] Embodiments of this application provide a method for processing a reference object, a method for generating visual data, a device, a device, a storage medium, and a program product. By generating a reference object including a reference image and attribute information, and then generating the expected visual data based on the reference object, it can not only improve the quality of visual data generation, but also expand the application scenarios of the reference object because the reference object can be applied to different visual data generation tasks, saving the computing resources and storage resources of the visual data generation system.
[0005] The technical solution of the embodiments of this application is implemented as follows:
[0006] Embodiments of this application provide a method for processing a reference object. The method includes: in response to a trigger operation on a reference object creation button, displaying a reference object creation interface; the reference object creation interface includes an interactive reference image upload area and an interactive attribute information area; in response to an image upload operation performed in the reference image upload area, displaying the uploaded reference image in the reference image upload area; displaying the attribute information of the reference image in the attribute information area; in response to a trigger operation on a confirmation button, generating a reference object including the reference image and the attribute information; the reference object is used to provide reference information when generating visual data.
[0007] An embodiment of the present application provides a reference object processing device, including: a first display module, configured to display a reference object creation interface in response to a trigger operation on a reference object creation button; the reference object creation interface includes an interactive reference image upload area and an interactive attribute information area; a second display module, configured to display an uploaded reference image in the reference image upload area in response to an image upload operation performed in the reference image upload area; a third display module, configured to display attribute information of the reference image in the attribute information area; a generation module, configured to generate a reference object including the reference image and the attribute information in response to a trigger operation on a confirmation button; the reference object is used to provide reference information when generating visual data.
[0008] In the above solution, the third display module is further configured to: when displaying the reference image, display, in the attribute information area, attribute information generated after performing image recognition on the reference image; or, in response to an information input operation performed in the attribute information area, display the input attribute information in the attribute information area.
[0009] In the above solution, the generated attribute information includes at least one of the following: a visual effect type text and an image content description text; the third display module is further configured to: display the visual effect type text and the image content description text generated after performing image recognition on the reference image.
[0010] In the above solution, the third display module is further configured to: after displaying the attribute information generated after performing image recognition on the reference image, in response to a rewrite operation on the generated attribute information, display the rewritten attribute information in the attribute information area.
[0011] In the above solution, the attribute information includes a tag of the reference image; the third display module is further configured to: in response to a tag input operation performed in the attribute information area, display the input tag in the attribute information area; or, in response to a tag selection operation performed in the attribute information area, display the selected tag in the attribute information area.
[0012] In the above solution, the first display module is further configured to: display a reference object display interface in response to a query operation on a reference object library performed in a visual data generation interface; wherein, the reference object display interface includes thumbnails of reference objects in the reference object library and the reference object creation button; the reference object library includes at least one type of reference object; in response to a trigger operation on the reference object creation button in the reference object display interface, display the reference object creation interface.
[0013] In the above solution, the first display module is further configured to: in response to a query operation on the reference object library performed in the data processing interface, display a reference object display interface; wherein, the reference object display interface includes thumbnails of reference objects in the reference object library and the reference object creation button; the reference object library includes at least one type of reference object; in response to a trigger operation on the reference object creation button in the reference object display interface, display the reference object creation interface.
[0014] In the above solution, the first display module is further configured to: in response to a query operation on the reference object corresponding to the target tag, display the reference object display interface corresponding to the target tag; wherein, the reference object display interface includes thumbnails of reference objects with the target tag and the reference object creation button; in response to a trigger operation on the reference object creation button in the reference object display interface, display the reference object creation interface; wherein, in the attribute information area of the reference object creation interface, the target tag is displayed.
[0015] In the above solution, the generation module is further configured to: after generating a reference object including the reference image and the attribute information, in response to a processing operation on the reference image, generate a reference object including at least the processed reference image; in response to an editing operation on the attribute information, generate a reference object including at least the edited attribute information.
[0016] An embodiment of the present application provides a visual data generation method, the method includes: in response to a visual data generation request, obtain at least one target reference object from a reference object library, and obtain input text; wherein, the reference objects in the reference object library are generated by using the above reference object processing method; based on the input text and the target reference object, generate target visual data; wherein, the target visual data includes information of the target reference object.
[0017] An embodiment of the present application provides a visual data generation device, the device includes: an acquisition module, configured to obtain at least one target reference object from a reference object library in response to a visual data generation request, and obtain input text; wherein, the reference objects in the reference object library are generated by using the above reference object processing method; a target visual data generation module, configured to generate target visual data based on the input text and the target reference object; wherein, the target visual data includes information of the target reference object.
[0018] An embodiment of the present application provides an electronic device, where the electronic device includes: a memory for storing computer-executable instructions or computer programs; a processor for implementing the reference object processing method or the visual data generation method provided by the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.
[0019] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which are used to implement the reference object processing method or the visual data generation method provided by the embodiment of the present application when executed by a processor.
[0020] An embodiment of the present application provides a computer program product including a computer program or computer-executable instructions, where when the computer program or computer-executable instructions are executed by a processor, the reference object processing method or the visual data generation method provided by the embodiment of the present application is implemented.
[0021] The embodiment of the present application has the following beneficial effects:
[0022] Before generating visual data, a reference object is generated first. The reference object includes a reference image and attribute information. In the process of generating the reference object, first, in response to a trigger operation on the reference object creation button, a reference object creation interface is displayed; then, in response to an image upload operation performed in the reference image upload area on the reference object creation interface, the uploaded reference image is displayed in the reference image upload area, and the attribute information of the reference image is displayed in the attribute information area on the reference object creation interface; finally, in response to a trigger operation on the confirmation button, a reference object including the reference image and attribute information is generated, and the reference object is used to provide reference information when generating visual data. In this way, by generating a reference object including a reference image and attribute information, and then generating the expected visual data based on the reference object, not only can the generation quality of the visual data be improved, but also because the reference object can be applied to different visual data generation tasks, the application scenario of the reference object is expanded, and the computing resources and storage resources of the visual data generation system are saved. Description of the Drawings
[0023] Figure 1 is an optional flowchart of the reference object processing method provided by the embodiment of the present application;
[0024] Figure 2 is an implementation flowchart of displaying the reference object creation interface provided by the embodiment of the present application;
[0025] Figure 3 is another implementation flowchart of displaying the reference object creation interface provided by the embodiment of the present application;
[0026] Figure 4 It is another schematic diagram of the implementation process for displaying the reference object creation interface provided by the embodiments of the present application;
[0027] Figure 5 It is another optional schematic diagram of the process for the reference object processing method provided by the embodiments of the present application;
[0028] Figure 6 It is yet another optional schematic diagram of the process for the reference object processing method provided by the embodiments of the present application;
[0029] Figure 7 It is an optional schematic diagram of the process for the visual data generation method provided by the embodiments of the present application;
[0030] Figure 8 It is another optional schematic diagram of the process for the visual data generation method provided by the embodiments of the present application;
[0031] Figure 9 It is the functional interface diagram of the visual data generation method provided by the embodiments of the present application;
[0032] Figure 10 It is a startup interface diagram of the main body creation process provided by the embodiments of the present application;
[0033] Figure 11 It is a reference object creation interface diagram provided by the embodiments of the present application;
[0034] Figure 12 It is the product home page interface diagram provided by the embodiments of the present application;
[0035] Figure 13 It is another startup interface diagram of the main body creation process provided by the embodiments of the present application;
[0036] Figure 14 It is another reference object creation interface diagram provided by the embodiments of the present application;
[0037] Figure 15 It is a schematic diagram of inputting labels in the reference object creation interface provided by the embodiments of the present application;
[0038] Figure 16 It is a schematic diagram of selecting labels in the reference object creation interface provided by the embodiments of the present application;
[0039] Figure 17 It is the interface diagram for searching for the already created main body provided by the embodiments of the present application;
[0040] Figure 18 It is a schematic diagram of automatically substituting labels provided by the embodiments of the present application;
[0041] Figure 19It is a schematic diagram for automatically generating the style and description of the uploaded image provided by an embodiment of the present application;
[0042] Figure 20 It is a schematic diagram for entering the main body editing interface provided by an embodiment of the present application;
[0043] Figure 21 It is a diagram of the main body editing interface provided by an embodiment of the present application;
[0044] Figure 22 It is a diagram of the interface for modifying the style and description provided by an embodiment of the present application;
[0045] Figure 23 It is a block diagram of a reference object-based processing device provided by an embodiment of the present application;
[0046] Figure 24 It is a block diagram of a visual data generation device provided by an embodiment of the present application;
[0047] Figure 25 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0048] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0049] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and they can be combined with each other without conflict.
[0050] In the following description, the terms "first\second\third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0051] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of that module or unit.
[0052] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those skilled in the art to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0053] In the embodiments of the present application, when collecting and processing relevant data in practical applications, the informed consent or separate consent of the personal information subject should be obtained strictly in accordance with the requirements of relevant laws and regulations, and subsequent data use and processing should be carried out within the scope authorized by laws and regulations and the personal information subject.
[0054] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0055] 1) In response to: used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more executed operations can be real-time or can have a set delay; in the absence of special instructions, there is no restriction on the execution order of the multiple executed operations.
[0056] 2) Human-computer interaction interface, an interface for providing human-computer interaction functions / an interface for displaying image information.
[0057] For example, a graphical user interface (GUI) is presented, such as an augmented reality (AR) interface, a virtual reality (VR) interface, a voice user interface (VUI), an interactive projection interface (using projection technology to display information on a flat surface), an eye movement detection interface (an interface controlled by detecting the user's line of sight), a holographic interface (a three-dimensional hologram formed by projecting an image through holographic projection technology, allowing a stereoscopic image to be seen without wearing special glasses), a multimodal interface (an interactive interface that combines multiple interaction methods such as touch, vision, and hearing), a brain-machine interface (BMI) interface, etc.
[0058] 3) Prompt: It is a text instruction or question input by the user to the model, used to guide the model to generate a specific type of response or content.
[0059] To solve the problem in related technologies that when the generation quality of visual data is poor, the visual data generation task is large, and there is duplicate content in the visual data generation task, which causes a large waste of computing resources and storage resources, the embodiments of the present application provide a method for processing reference objects and a method for generating visual data. This method forms a specific product, such as a visual data generation application, and provides a visual data generation function (which can also be referred to as the "reference video generation" function in the embodiments of the present application) in this visual data generation application. The reference information such as the characters, props, or scenes referred to in the reference video generation function can be saved into the reference object library. In the reference video generation function, one or more reference objects that have been created can be selected from the reference object library, and then based on the selected reference objects and the input text (this input text is used to indicate the actions or display effects of the selected reference objects in the final visual data (such as a video)).
[0060] The reference object processing method provided in the embodiment of the present application provides an implementation process of how to generate a reference object in a reference object library, and the visual data generation method provided in the embodiment of the present application provides an implementation process of how to generate the final expected visual data based on the reference object in the reference object library. Specifically, in the reference object processing method provided in the embodiment of the present application, first, in response to a trigger operation on a reference object creation button, a reference object creation interface is displayed; the reference object creation interface includes an interactive reference image upload area and an interactive attribute information area; then, in response to an image upload operation performed in the reference image upload area, the uploaded reference image is displayed in the reference image upload area; and, the attribute information of the reference image is displayed in the attribute information area; finally, in response to a trigger operation on a confirmation button, a reference object containing a reference image and attribute information is generated; the reference object is used to provide reference information when generating visual data. In the visual data generation method provided in the embodiment of the present application, first, in response to a visual data generation request, at least one target reference object is obtained from a reference object library, and input text is obtained; wherein the reference object in the reference object library is generated using the aforementioned reference object processing method; then, based on the input text and the target reference object, target visual data is generated; wherein the target visual data includes information about the target reference object. In this way, by generating a reference object containing a reference image and attribute information, and thereby generating expected visual data based on the reference object, not only can the quality of visual data generation be improved, but also because the reference object can be applied to different visual data generation tasks, the application scenarios of the reference object are expanded, saving computing resources and storage resources of the visual data generation system.
[0061] Below, the reference object processing method provided in the embodiment of the present application is first described.
[0062] The reference object processing method provided in the embodiment of the present application can be applied to electronic devices such as laptops, tablet computers, and desktop computers, that is, the electronic device used to implement the reference object processing method can be implemented as a terminal, and the embodiment of the present application does not impose any restrictions on the specific type of electronic device.
[0063] The reference object processing method provided in the embodiment of the present application is described in detail below with reference to the accompanying drawings.
[0064] Figure 1 This is an optional flow chart of the reference object processing method provided in the embodiment of the present application. The method can be applied to an electronic device, which can be a terminal. That is, the reference object processing method of each embodiment of the present application can be executed by the terminal, or can also be executed through interaction between a server and a terminal. The following will take the electronic device as an example for exemplary description. Figure 1 As shown, the method includes the following steps S101 to S104:
[0065] Step S101: The terminal responds to a trigger operation for creating a reference object button and displays a reference object creation interface.
[0066] A visual data generation application can be installed on the terminal. The visual data generation application provides a reference object generation function. Through this reference object generation function, a reference object can be generated and stored in a reference object library. In this way, when visual data is generated subsequently, visual data can be generated based on the reference objects in the reference object library. That is to say, the information of the reference objects in the reference object library can be referred to for visual data generation.
[0067] In the embodiments of the present application, before generating a reference object, a reference object creation button can be displayed on the client interface of the video data generation application. The reference object creation button can be displayed on any interface of the visual data generation application. For example, it can be displayed on the main interface of the visual data generation application, or it can also be displayed on the function interface of the reference object generation function, or it can also be displayed on the function interface of the visual data generation function.
[0068] The reference object creation interface can include an interactive reference image upload area and an interactive attribute information area. The reference image upload area includes an interactive reference image upload button. In this reference image upload area, the user can interact with the reference image upload button to independently upload a reference image. For example, the user can select a reference image from the local album of the terminal and upload it, or can query and upload the downloaded reference images from the currently logged-in account of the visual data generation application, or can also query and upload reference images from the application background storage space of the visual data generation application.
[0069] The attribute information area includes an interactive attribute information button. In this attribute information area, the user can interact with the attribute information button to input attribute information by himself. In the embodiments of the present application, the attribute information includes but is not limited to at least one of the following: the name of the reference image, the visual effect type text, the image content description text, tags and other information.
[0070] In one implementation, see Figure 2 , Figure 2 It shows that in step S101, in response to a trigger operation for the reference object creation button, displaying the reference object creation interface can be implemented through the following step S1011 and step S1012:
[0071] Step S1011: The terminal responds to a query operation for the reference object library executed on the visual data generation interface and displays a reference object display interface.
[0072] Here, the visual data generation interface includes a query button for the reference object library. The user performs a query operation by clicking this query button. After the terminal receives the user's query operation, it switches from the visual data generation interface to the reference object display interface. The reference object display interface includes thumbnails of the reference objects in the reference object library and reference object creation buttons. The reference object library includes at least one type of reference object.
[0073] A thumbnail is a reduced version of an image, usually used for quick preview, navigation, or presenting an overview of the main content. The size of a thumbnail is usually much smaller than the representative image of the original reference object. For example, the size of a thumbnail can include 128x128 pixels, 256x256 pixels, etc. The smaller size makes the thumbnail load quickly and is suitable for quick display in the reference object display interface. Although the thumbnail size is small, the thumbnail usually retains the key visual features of the original content, such as the face of a person, the main elements of a scene, etc., so that the user can quickly identify the content. The thumbnail in the embodiment of the present application can be a static image (such as JPEG, PNG format), a dynamic image (such as GIF, WebP animation format), or even a key frame of a video. The representative image of a reference object is an image containing the key information of the reference object, for example, it can be an image including a person.
[0074] The thumbnails of the reference objects include, but are not limited to, the following types: cropped thumbnails, fitted thumbnails, textured thumbnails, custom-sized thumbnails, animated thumbnails, and square thumbnails.
[0075] The characteristic of the cropped thumbnail is that the representative image of the reference object is cropped to the aspect ratio of the thumbnail. Unless the aspect ratio of the thumbnail rectangle is exactly the same as that of the representative image of the reference object, some content of the representative image of the reference object will be cropped to ensure that the entire thumbnail rectangle is filled with the representative image data of the reference object. The position of the cropping rectangle in the representative image of the reference object can be controlled by attributes such as horizontal alignment and vertical alignment. The cropped thumbnail is suitable for scenarios where the main core part of the reference object needs to be highlighted, such as the display of a person's head or a key prop. The characteristic of the fitted thumbnail is that the entire representative image of the reference object is adjusted into the thumbnail rectangle. The representative image of the reference object will not be cropped, but there may be blank areas. The blank areas can be filled with the background color. The position of the reference object in the thumbnail can be controlled by the horizontal and vertical alignment attributes. The fitted thumbnail is suitable for scenarios where the complete content of the reference object needs to be retained, such as the display of a complete scene or the full view of a prop. The characteristic of the textured thumbnail is that the representative image of the reference object is cropped according to the resolution. The representative image of the reference object is scaled by a certain ratio. If the resulting image is larger than the thumbnail, the image will be cropped to fit; if the scaled image is smaller than the thumbnail, the remaining area will be filled with the background color. The textured thumbnail is suitable for scenarios where the image quality needs to be maintained at different resolutions, such as the display of thumbnails on different devices. The characteristic of the custom-sized thumbnail is that the user can specify the size of the thumbnail, including the width and height. The thumbnail can be generated according to the specified size while maintaining the aspect ratio of the original representative image of the reference object. The custom-sized thumbnail is suitable for scenarios where different display devices or platforms need to be adapted. The characteristic of the animated thumbnail is that an animated thumbnail is generated using an animation format (such as WebM or WebP), which can display the key actions or scenes of the video. The animated thumbnail is suitable for scenarios where user attention needs to be attracted. The characteristic of the square thumbnail is that the image is cropped into a square, which is suitable for scenarios where a unified format is required.
[0076] The reference object creation button is used to generate a reference object creation request when the user interacts with the reference object creation button, thereby triggering the reference object creation process.
[0077] Step S1012, the terminal displays a reference object creation interface in response to a trigger operation on the reference object creation button in the reference object display interface.
[0078] Here, when the user clicks the reference object creation button in the reference object display interface, it switches to the reference object creation interface.
[0079] In the embodiments of the present application, a query operation on the reference object library is performed in the visual data generation interface, thereby entering the reference object display interface. Then, by performing a triggering operation on the reference object creation button in this reference object display interface, the reference object creation interface is entered, and relevant operations for creating a reference object are performed in this reference object creation interface to create a reference object. In this way, through the implementation process from step S1011 to step S1012, when the user is in the visual data generation interface, their thinking has focused on the core task of video creation. The creation of reference objects in this scenario is closely linked to the final visual data generation. For example, when the user is conceiving a video about outdoor adventures and clicks on "My Reference Object" in the reference object creation interface, the user can naturally incorporate the ideas in their mind about the image of an explorer (i.e., the reference object) into the creation process. Thus, in the subsequent reference object creation process, a reference object that conforms to the video theme can be created, such as a reference object of an explorer wearing mountaineering equipment. This close scenario association enables the user to more efficiently match the reference object with the video content and reduces the cost of switching thinking.
[0080] In addition, in the implementation process from step S1011 to step S1012, the operation process has good coherence. Because for the user, from selecting elements (such as people, props, scenes, etc.) in the reference object to saving the generated reference object into the reference object library and then generating a video based on the reference object, it is a coherent operation process. In this process, the user can directly click on "My Reference Object" in the visual data generation interface to create a reference object, avoiding frequent jumps between different interfaces. For example, when the user browses a food-making video and finds that the image of the chef in it is very suitable for the video style they want, they can directly create a similar chef as a reference object in this reference object display interface and immediately use this reference object to generate the video content they want, and the whole process is completed in one go, improving the user experience.
[0081] In another implementation, refer to Figure 3 , Figure 3 which shows that in step S101, in response to the triggering operation on the reference object creation button, the reference object creation interface is displayed, and it can also be implemented through the following step S1013 and step S1014:
[0082] Step S1013, the terminal responds to the query operation on the reference object library performed in the data processing interface and displays the reference object display interface.
[0083] Here, the data processing interface can be an interface for implementing at least one data processing function. For example, the data processing interface can be the main interface or the home page of a visual data generation application. On the data processing interface, operation buttons for video generation functions can be provided, and operation buttons for video pushing functions can also be provided, etc. The data processing interface includes a query button for the reference object library. The user performs a query operation by clicking the query button. After the terminal receives the user's query operation, it switches from the data processing interface to the reference object display interface. The reference object display interface includes thumbnails of reference objects in the reference object library and a reference object creation button. The reference object library includes at least one type of reference object.
[0084] Step S1014, the terminal displays a reference object creation interface in response to a trigger operation on the reference object creation button in the reference object display interface.
[0085] In the embodiments of the present application, through the query operation on the reference object library performed on the data processing interface, the reference object display interface is entered, and in the reference object display interface, by performing a trigger operation on the reference object creation button, the reference object creation interface is further entered, so as to perform related operations for creating a reference object in the reference object creation interface to create a reference object. In this way, through the implementation process of steps S1013 to S1014, global management of the visual data generation application can be realized. Because the home page of the visual data generation application is usually a more macroscopic perspective, users can manage all reference objects here. By clicking "My Reference Objects" on the home page to create a new reference object, users can globally plan the reference object as an independent resource. For example, users may want to create a series of reference objects of different styles of characters for multiple different video projects. Creating reference objects on the home page can allow them to enter the reference object display interface more clearly and quickly, so as to see information such as the quantity and type of reference objects they have created, facilitating them to classify and organize the resources of reference objects, and thus better overall planning of their creative resources.
[0086] In addition, the implementation process of steps S1013 to S1014 can also facilitate the reuse and expansion of reference objects. Since the home page is the starting interface after the user enters the product, after the user enters and creates a reference object at this position, the reference object can be conveniently applied to different functional modules. For example, if a user creates a reference object of a cute cartoon animal, in addition to being used for video generation in the visual data generation interface, this reference object of the cartoon animal can also be reused in other possible functional modules of the product, such as dynamic wallpaper production and emoji generation. This way of creating reference objects from the home page provides a wider space for the expansion of the application of reference objects for users, maximizing the value of reference objects.
[0087] In yet another implementation, referring to Figure 4 , Figure 4 it shows that in response to a trigger operation for creating a button for a reference object in step S101, a reference object creation interface is displayed, and it can also be implemented through the following step S1015 and step S1016:
[0088] Step S1015, the terminal responds to a query operation for a reference object corresponding to a target label, and displays a reference object display interface corresponding to the target label.
[0089] Here, the target label can be any label selected by the user. For example, the target label can be any one of person, animal, cartoon, plant, anime, book, building, scenery, etc. In the embodiments of the present application, the user can select the target label in the reference object query interface (the reference object query interface can be a visual data generation interface or an interface entered by clicking the reference object query button in the visual data generation interface, or a data processing interface or an interface entered by clicking the reference object query button in the data processing interface), and perform a query operation on the reference object of the target label, requesting to query the reference object in the reference object library that has the target label. Of course, the reference object query interface can also be the reference object display interface. At the same time, in the reference object display interface, multiple labels corresponding to all the existing reference objects in the current reference object library can also be displayed. Then, the user can select the target label from the displayed multiple labels, and the terminal will display the reference object corresponding to the target label, that is, display the reference object display interface corresponding to the target label.
[0090] The reference object display interface includes a thumbnail of the reference object with the target label and a reference object creation button.
[0091] Step S1016, the terminal responds to a trigger operation for the reference object creation button in the reference object display interface, and displays a reference object creation interface; wherein, in the attribute information area of the reference object creation interface, the target label is displayed.
[0092] In the embodiments of the present application, by selecting a target tag, a reference object display interface corresponding to the target tag is displayed. In this way, the search efficiency can be improved, and the reference object with the target tag can be quickly located. Because when a large number of reference objects are stored in the reference object library, it is difficult for users to quickly find the specific reference object they need. Through the tag filtering function, users can quickly locate relevant reference objects according to the tags. For example, users add the "ancient style" tag to some reference objects. When making a video in the ancient style, they only need to select the "ancient style" tag to immediately find all relevant reference objects, without browsing each reference object in the reference object library one by one. Moreover, it can greatly save time. In the case of a large number of reference objects, this filtering function can significantly reduce the time for users to find the required reference objects and improve work efficiency. Especially in some creative scenarios where reference objects need to be frequently switched, quickly finding the appropriate reference object can help users create more smoothly.
[0093] In addition, by managing the reference objects in the reference object library through tags, the systematicness of reference object management can be enhanced, making the classification clear, because tags can help users classify and manage reference objects. Users can add tags to reference objects according to different dimensions such as the style, use, and scenario of the reference objects, so as to divide the reference objects in the reference object library into different categories. For example, reference objects can be divided into major categories such as "characters", "props", and "scenes", and then more detailed tags can be added under each major category according to specific styles or functions, such as "modern characters", "ancient characters", "indoor scenes", "outdoor scenes", etc. This classification method makes the reference object library more orderly, and users can manage and search for reference objects more systematically. Moreover, it is also convenient for maintenance and update. When users need to add new reference objects or modify existing reference objects, tags can help users quickly find the categories where the relevant reference objects are located. For example, when a user newly creates a reference object of a "cyberpunk style" character, it can be classified into the corresponding category by adding the "cyberpunk" tag to maintain the classification consistency of the reference object library.
[0094] Meanwhile, in the embodiment of the present application, by selecting a target tag to display the reference object display interface corresponding to the target tag, when it is necessary to update the attributes or styles of a certain type of reference object, all relevant reference objects can also be quickly found through the tags for batch operations. In addition, by filtering the corresponding reference objects through tags, the flexibility and diversity of creation can be improved. The tag filtering function allows users to more conveniently explore different categories of reference objects, thereby inspiring new creative inspirations. For example, the user originally only planned to make a video using "animal" reference objects, but when filtering the "animal" tag, unexpectedly found an animal reference object with the "fantasy" tag, which may inspire the user to try making a fantasy-style animal adventure video. Through tag filtering, users can more easily combine reference objects of different categories or styles to create more diverse works. And since different creative projects may require different types of reference objects, through tag filtering, users can quickly find the appropriate reference objects according to the specific requirements of the current project. For example, when making a promotional video about technology products, the user can select tags such as "technology" and "modern" to quickly find reference objects related to the technology theme, such as futuristic character images and high-tech props, so as to better meet the creative needs. Of course, users can also add personalized tags to reference objects according to their own creative habits and preferences, thereby creating a classification system of reference object libraries that meets personal needs. This personalized method makes the reference object library more in line with the user's usage habits, improving the user's satisfaction and loyalty to the product. For new users, the tag filtering function is intuitive and easy to understand, and they can quickly get started without complex operations. Users do not need to spend a lot of time learning how to find the required reference objects in the huge reference object library, reducing the learning cost of the product and improving the usability of the product. In a creative project with team collaboration, tags can help team members classify and search for reference objects according to a unified standard. For example, the team can pre-agree on the tag system for reference objects, and each member adds tags according to this system when creating reference objects. In this way, when other members need to search for reference objects, they can quickly find the required reference objects through the tags, avoiding confusion caused by inconsistent classification of reference objects.
[0095] Step S102, the terminal responds to the image upload operation performed in the reference image upload area and displays the uploaded reference image in the reference image upload area.
[0096] In the embodiment of the present application, the reference image can be an image in any format. For example, it can be any one or more of JPEG (or JPG), PNG, GIF, WebP, TIFF, BMP, SVG, HEIF (or HEIC), RAW, PSD, etc.
[0097] In the embodiments of the present application, performing an image upload operation in the reference image upload area can be that the user clicks on the reference image upload area or the image upload button, triggers the file selection box, and selects a local image file for upload. This method is simple and intuitive, suitable for most users, especially those who are not familiar with advanced operations. Or, performing an image upload operation in the reference image upload area can also be that the user drags the image file from the local to the specified reference image upload area, and the system automatically captures the file and starts the upload. This method is convenient for operation and suitable for scenarios where multiple pictures need to be uploaded quickly, improving the user experience. Or, performing an image upload operation in the reference image upload area can also be that the user copies the image file to the clipboard, then moves the mouse to the reference image upload area and performs a paste operation. This method is suitable for scenarios where images need to be quickly copied from other places and is flexible in operation.
[0098] After the user selects or uploads an image, a preview image of the reference image can be immediately displayed in the reference image upload area, allowing the user to instantly see the uploaded image effect and avoid uploading incorrect images. In the embodiments of the present application, the user is allowed to select multiple image files for upload at one time and can perform batch operations during the upload process, such as adding descriptions and deleting unnecessary pictures. A progress bar can also be displayed during the upload process to let the user know the upload progress.
[0099] Step S103, the terminal displays the attribute information of the reference image in the attribute information area.
[0100] In some embodiments, the step S103 of displaying the attribute information of the reference image in the attribute information area can be implemented in any of the following ways:
[0101] Way 1: When the reference image is displayed, the attribute information generated after performing image recognition on the reference image is displayed in the attribute information area.
[0102] The image recognition function can be used to perform image recognition on the reference image to obtain the attribute information of the reference image. For example, a pre-trained image recognition model can be called to perform image recognition on the reference image.
[0103] In the embodiments of the present application, after the user uploads the reference image, they can immediately see the attribute information of the reference image, which provides instant feedback to the user and helps them better understand the content of the reference image.
[0104] In some embodiments, the generated attribute information includes at least one of the following: visual effect type text and image content description text. Correspondingly, in the embodiments of the present application, displaying the attribute information generated after performing image recognition on the reference image can be displaying the visual effect type text and image content description text generated after performing image recognition on the reference image.
[0105] In some embodiments, after displaying the attribute information generated after image recognition of a reference image, the rewritten attribute information may also be displayed in the attribute information area in response to a rewrite operation on the generated attribute information.
[0106] The rewrite operation may involve rewriting multiple aspects of the attribute information generated by image recognition, including rewriting the content of various aspects such as tags, descriptive text, keywords, metadata, sentiment descriptions, scene descriptions, subject descriptions, language and cultural adaptability, review and compliance, and technical details. By allowing users to rewrite this content, the accuracy, relevance, personalization, and user experience of the content can be significantly improved, while meeting diverse application scenarios and requirements.
[0107] Here, a tag is a short description of the image content, usually used for classification and retrieval. Rewriting tags can be correcting incorrect tags, refining tag classification, and adding personalized tags. For example, image recognition may mislabel a "Corgi" as a "Husky", and the user can rewrite it to the correct tag. Another example is that the user can refine "animal" to "pet dog" or "wild animal", or the user can add more personalized tags according to their own needs, such as "my pet", "travel souvenir", etc. The descriptive text is a detailed description of the image content, usually used to provide more abundant information. The user can adjust the descriptive text to make it more accurately reflect the image content. For example, rewrite "a person at the seaside" to "a young man surfing at the seaside". The user can also add more details to make the description more vivid. For example, rewrite "scenery" to "the golden beach and the blue sea under the sunset"; the user can also adjust the style of the description according to the need, such as changing from a formal style to an oral style, or from a concise style to a literary style. Keywords are words used for search engine optimization or content retrieval. The user can add or modify keywords to make them more in line with the search habits of the target audience. For example, change "food" to more specific keywords such as "pasta", "pizza", etc.; the user can also add keywords related to the image content to improve the retrieval efficiency of the content. Metadata is additional information about the reference image, such as the shooting time, shooting device, and shooting location, etc. The user can add missing metadata, such as the shooting location or time. The user can also modify incorrect metadata, such as the wrong shooting date or location. The user can also add personalized metadata, such as the name of the shooter or the background story of the shooting. The emotional description is a description of the emotion or atmosphere conveyed by the image. The user can adjust the emotional description to make it more in line with the emotion conveyed by the image. For example, change "happy" to "energetic happiness"; the user can also adjust the atmosphere of the description according to the need, such as changing "calm" to "peaceful and serene". The scene description is a description of the scene or background shown in the image. The user can refine "indoor" to more specific scenes such as "living room", "kitchen", etc. The user can also add more background information, such as "inside an ancient castle" or "on the bustling urban street". The subject description is a description of the main object or person in the image. The user can refine the description of the subject, such as changing "a person" to "a lady in a red dress", and the user can also adjust the description of the relationship between the subjects, such as changing "two people" to "a couple". Language and Cultural Adaptation describes the language and cultural background of the descriptive text. The user can translate the descriptive text into other languages to adapt to users with different language backgrounds; the user can also adjust the descriptive text according to different cultural backgrounds to make it more in line with the habits and preferences of local users.Review and Compliance ensure that the description text complies with platform rules and social norms. By filtering inappropriate content, users can delete or modify inappropriate words or content in the description; users can also adjust the description text to conform to the brand positioning and values. Technical Details include the technical parameters of the reference image, such as resolution, format, shooting parameters, etc. Users can add missing technical details, such as the resolution or shooting parameters of the image; users can also modify incorrect technical parameters, such as image format or resolution.
[0108] In the embodiments of the present application, users are allowed to rewrite the attribute information generated by image recognition, which not only improves the accuracy and relevance of the content, but also enhances the user experience and sense of participation, optimizes content management and retrieval, adapts to diverse application scenarios, improves the scalability and flexibility of the content, promotes the sharing and dissemination of the content, supports content review and management, and improves the accessibility of the content. This functional design enables visual data generation applications to better meet user needs, improve content quality, and enhance user stickiness, thus standing out in the highly competitive market.
[0109] Method 2: In response to an information input operation performed in the attribute information area, display the input attribute information in the attribute information area.
[0110] In some embodiments, the attribute information includes the label of the reference image. Accordingly, in the embodiments of the present application, in response to an information input operation performed in the attribute information area, displaying the input attribute information in the attribute information area may be in response to a label input operation performed in the attribute information area, and displaying the input label in the attribute information area; or, it may also be in response to a label selection operation performed in the attribute information area, and displaying the selected label in the attribute information area.
[0111] Step S104, the terminal generates a reference object including the reference image and the attribute information in response to a trigger operation on the confirmation button.
[0112] In the embodiments of the present application, the reference object is used to provide reference information when generating visual data. Among them, in the existing related technologies, if a user needs to include the same object in multiple visual data (a visual large model generates one visual data through one visual data generation task), the user can only upload the same image or text information in each visual data generation task. However, due to the model hallucination of the visual large model and the randomness of AIGC content generation, it is impossible to well maintain the consistency of the same object in the visual data generated in different tasks. In the embodiments of the present application, in the process of generating visual data based on the reference object, due to multiple generated reference objects, the user selects one or more target reference objects, and then the visual data generation model can quickly identify the relevant information of the target reference object selected by the user, thereby ensuring the consistency of the reference object in different visual data generation tasks.
[0113] Furthermore, in the model training or model fine-tuning stage, the visual data generation model in the embodiments of the present application is adapted to the reference object processing model in the embodiments of the present application. In this stage, based on a small amount of training data, the training data includes training reference images, training attribute information, and first training visual data. First, the training reference object is generated through the reference object processing model, and then the visual data generation model generates second training visual data based on the training reference images, training attribute information, and training reference object. Based on the error loss between the second training visual data and the first training visual data, the parameters of the reference object processing model are adjusted. Therefore, for the same reference object generated by the reference object processing model, the visual data generation model can perform more accurate identification, further ensuring the consistency of the unified reference object in different visual data.
[0114] Furthermore, the generated reference object and its corresponding reference image and attribute information are stored in a structured manner, which is convenient for the visual data generation model to quickly and accurately identify the relevant information of each reference object.
[0115] In some embodiments, after generating the reference object including the reference image and attribute information in step S104, the terminal can also generate a reference object including at least the processed reference image in response to a processing operation on the reference image; or, the terminal can also generate a reference object including at least the edited attribute information in response to an editing operation on the attribute information.
[0116] Here, the processing operations for the reference images can be deletion operations, replacement operations, addition operations, modification operations, etc. For example, one or more reference images in the reference object can be deleted through a deletion operation; one or more reference images in the reference object can be replaced through a replacement operation; one or more new reference images can be added to the reference object through an addition operation; the local content of one or more reference images in the reference object can be modified through a modification operation, such as segmenting the reference image, including the segmented partial images, or retouching the reference image to modify one or more objects in the reference image.
[0117] The editing operations for the attribute information can be rewriting operations, addition operations, deletion operations, etc. For example, the content of the attribute information can be rewritten through a rewriting operation; new content can be added to the attribute information through an addition operation; part or all of the content in the attribute information can be deleted through the deletion operation.
[0118] In the embodiments of the present application, users are allowed to process the reference images and attribute information after generating the reference object, which can not only improve the accuracy and relevance of the content, but also correct errors or inaccurate content by modifying the reference images and attribute information. For example, if image recognition wrongly labels a "Corgi" as a "Husky", the user can manually modify the label. Moreover, the user can also adjust the reference images and attribute information according to their own needs and preferences to make the reference images and attribute information more in line with specific usage scenarios or styles. Allowing users to edit the reference images and attribute information gives users more control and autonomy. Instead of passively accepting the content generated by the system, users can actively participate in the optimization and customization of the content, thereby enhancing the satisfaction and loyalty of users to the visual data generation application, and further increasing the activity and user stickiness of the visual data generation application.
[0119] The reference object processing method provided by the embodiments of the present application generates a reference object before generating visual data. The reference object includes a reference image and attribute information. During the process of generating the reference object, first, in response to a trigger operation on the reference object creation button, a reference object creation interface is displayed; then, in response to an image upload operation performed in the reference image upload area on the reference object creation interface, the uploaded reference image is displayed in the reference image upload area, and the attribute information of the reference image is displayed in the attribute information area on the reference object creation interface; finally, in response to a trigger operation on the confirmation button, a reference object including the reference image and attribute information is generated, and the reference object is used to provide reference information when generating visual data. In this way, by generating a reference object including a reference image and attribute information, and then generating the expected visual data based on the reference object, not only can the generation quality of the visual data be improved, but also because the reference object can be applied to different visual data generation tasks, the application scenario of the reference object is expanded, and the computing resources and storage resources of the visual data generation system are saved.
[0120] Figure 5 is another optional flowchart of the reference object processing method provided by the embodiments of the present application, as Figure 5 shown, the method includes the following steps S201 to step S213:
[0121] Step S201, the terminal responds to a query operation on the reference object library performed on the visual data generation interface and displays a reference object display interface.
[0122] Here, the reference object display interface includes thumbnails of reference objects in the reference object library and a reference object creation button. The reference object library includes at least one type of reference object.
[0123] Step S202, the terminal responds to a trigger operation on the reference object creation button in the reference object display interface and displays a reference object creation interface.
[0124] Here, the reference object creation interface includes an interactive reference image upload area and an interactive attribute information area.
[0125] Step S203, the terminal responds to a query operation on the reference object library performed on the data processing interface and displays a reference object display interface.
[0126] Here, the reference object display interface includes thumbnails of reference objects in the reference object library and a reference object creation button. The reference object library includes at least one type of reference object.
[0127] Step S204, the terminal responds to a trigger operation on the reference object creation button in the reference object display interface and displays a reference object creation interface.
[0128] Step S205: The terminal responds to a query operation for the reference object corresponding to the target tag and displays a reference object display interface corresponding to the target tag.
[0129] The reference object display interface includes a thumbnail of the reference object with the target tag and a reference object creation button.
[0130] Step S206: The terminal responds to a trigger operation on the reference object creation button in the reference object display interface and displays a reference object creation interface; wherein, in the attribute information area of the reference object creation interface, the target tag is displayed.
[0131] Step S207: The terminal responds to an image upload operation performed in the reference image upload area and displays the uploaded reference image in the reference image upload area.
[0132] Step S208: When the terminal displays the reference image, in the attribute information area, it displays the attribute information generated after performing image recognition on the reference image.
[0133] Here, the terminal can display the text of the visual effect type and the text description of the image content generated after performing image recognition on the reference image.
[0134] Step S209: The terminal responds to a rewrite operation on the generated attribute information and displays the rewritten attribute information in the attribute information area.
[0135] Step S210: The terminal responds to an information input operation performed in the attribute information area and displays the input attribute information in the attribute information area.
[0136] Here, the terminal can respond to a tag input operation performed in the attribute information area and display the input tag in the attribute information area; or, the terminal can also respond to a tag selection operation performed in the attribute information area and display the selected tag in the attribute information area.
[0137] Step S211: The terminal responds to a trigger operation on the confirmation button and generates a reference object including the reference image and the attribute information; the reference object is used to provide reference information when generating visual data.
[0138] Step S212: The terminal responds to a processing operation on the reference image and generates a reference object including at least the processed reference image.
[0139] Step S213: The terminal responds to an editing operation on the attribute information and generates a reference object including at least the edited attribute information.
[0140] The reference object processing method provided by the embodiments of this application is such that the terminal generates a reference object before generating visual data. The reference object includes a reference image and attribute information. During the process of generating the reference object, by generating a reference object that includes the reference image and attribute information, expected visual data can be generated based on this reference object. This not only improves the generation quality of the visual data, but also, since the reference object can be applied to different visual data generation tasks, the application scenarios of the reference object are extended, saving the computing resources and storage resources of the visual data generation system.
[0141] Figure 6 is another optional flowchart of the reference object processing method provided by the embodiments of this application. As Figure 6 shown, the method includes the following steps S301 to step S317:
[0142] Step S301, the terminal receives a trigger operation from the user for the reference object creation button.
[0143] Step S302, in response to the trigger operation for the reference object creation button, the terminal sends a page refresh request to the server.
[0144] Here, after the terminal receives the trigger operation from the user for the reference object creation button, it can, in response to the trigger operation for the reference object creation button, automatically generate a page refresh request and send the page refresh request to the server.
[0145] The terminal can send a page refresh request to the server through protocols such as the HyperText Transfer Protocol (HTTP) or WebSocket to request the interface data of the reference object creation interface. The page refresh request can include some necessary parameters, such as: timestamp, user identity information, and other parameters.
[0146] Step S303, the server, in response to the page refresh request, generates the interface data of the reference object creation interface.
[0147] After receiving the page refresh request, the server can parse the parameters in the page refresh request and query the latest interface data of the reference object creation interface from the database according to these parameters. After the query is completed, the server can encapsulate the query result into data in JSON format and prepare to return it to the terminal.
[0148] Step S304, the server sends the interface data to the terminal.
[0149] Here, the server returns the queried interface data to the terminal as the response result. The response result may include the following content: status code (indicating whether the request is successfully responded), response header (including some metadata such as content type, etc.), and response body (including the actual data, which can be in JSON format).
[0150] Step S305, the terminal creates an interface based on the interface data to display the reference object.
[0151] After receiving the response result from the server, the terminal parses the JSON data in the response body to obtain the interface data, and renders the interface data onto the current page to obtain the interface for creating the reference object. If the returned interface data is empty, the terminal may prompt the user with messages such as "Interface loading failed" or "Interface loading timed out".
[0152] The interface for creating the reference object includes an interactive reference image upload area and an interactive attribute information area.
[0153] Step S306, the terminal receives the image upload operation performed by the user in the reference image upload area.
[0154] Step S307, in response to the image upload operation, the terminal obtains the reference image uploaded by the user and displays the uploaded reference image in the reference image upload area.
[0155] Step S308, the terminal sends the reference image to the server.
[0156] Step S309, the server recognizes the reference image to obtain the attribute information of the reference image.
[0157] In the embodiments of the present application, a suitable image recognition model can be selected to recognize the reference image. The image recognition models here include: Convolutional Neural Network (CNN), Residual Network (ResNet), and pre-trained models. Among them, the convolutional neural network is suitable for processing image data and can automatically extract image features; the residual network solves the problem of gradient disappearance in the training of deep networks by introducing residual connections and is suitable for complex image recognition tasks; pre-trained models such as VGG, MobileNet, etc. These models have been pre-trained on large-scale datasets and can be directly used or fine-tuned.
[0158] Step S310, the server sends the attribute information of the reference image to the terminal.
[0159] Step S311, the terminal displays the attribute information of the reference image in the attribute information area.
[0160] Step S312, the terminal receives a triggering operation by the user on the confirmation button.
[0161] Step S313, the terminal generates a reference object generation request in response to the triggering operation.
[0162] Step S314, the terminal sends the reference object generation request to the server.
[0163] Here, after receiving the triggering operation by the user on the confirmation button, the terminal can, in response to the triggering operation on the confirmation button, automatically generate a reference object generation request and send the reference object generation request to the server.
[0164] The terminal can also send a reference object generation request to the server through protocols such as HTTP or Web Socket, requesting the server to generate a new reference object based on the uploaded reference image and the generated attribute information.
[0165] Step S315, the server generates a reference object including the reference image and the attribute information in response to the reference object generation request.
[0166] The reference object in the embodiment of the present application is used to provide reference information when generating visual data.
[0167] Step S316, the server stores the reference object in the reference object library.
[0168] In the embodiment of the present application, the newly generated reference object can be stored in the reference object library. The reference object library here can be a pre-constructed reference object library.
[0169] When constructing the reference object library, the reference object structure can be defined first, and the basic attributes of the reference object are determined, such as name, type (person, scene, prop, etc.), label, description text, upload time, etc. Then, the storage format of the reference object is designed. For example, the information of each reference object is stored in the form of a JSON object. When storing each reference object, a unique identifier (ID) can be assigned to each reference object for subsequent retrieval and management. An index can also be constructed for the reference object, specifically according to attributes such as the label and description text of the reference object to facilitate quick retrieval.
[0170] In some embodiments, a backup library can also be constructed, and the data in the reference object library can be synchronized regularly, and the data in the reference object library is backed up to this backup library to ensure the consistency of the data in the two libraries. And through data backup, data loss can be prevented.
[0171] Step S317, the server sends the generated reference object to the terminal.
[0172] Based on the reference object processing method provided in the above embodiments, next, the visual data generation method provided in the embodiments of the present application will be described.
[0173] Figure 7 FIG. 4 is an optional flowchart of the visual data generation method provided in the embodiments of the present application. This method can be applied to an electronic device, which can be a terminal or a server. That is, the visual data generation method in the embodiments of the present application can be executed by a terminal, or by a server, or by interaction between a server and a terminal. Hereinafter, an example will be given with the electronic device being a server. As Figure 7 shown, the method includes the following steps S401 to S402:
[0174] Step S401, the server, in response to a visual data generation request, obtains at least one target reference object from a reference object library, and obtains an input text.
[0175] Here, the reference objects in the reference object library are generated by using the above reference object processing method.
[0176] The input text can be a prompt text input by the user on the visual data generation interface. The input text can include the following contents from multiple angles: theme description, scene details, main body features, action description, emotional expression, style and tone, time information, technical details.
[0177] For example, the theme description can be "A young man is surfing by the sea"; the scene details can be "On the golden beach at sunset, the waves gently lap the shore"; the main body features can be "A young lady in a red dress with long flowing hair"; the action description can be "She is running cheerfully with a big smile"; the emotional expression can be "The picture is filled with a relaxed and pleasant atmosphere"; the style and tone can be "The picture style is realistic and the tone is warm"; the time information can be "On a sunny spring morning"; the technical details can be "The video resolution is 1080p and the frame rate is 30fps".
[0178] In some embodiments, on the visual data generation interface of the terminal, prompt information of the input text can also be displayed to prompt the user to describe the input text from multiple angles such as the above theme description, scene details, main body features, action description, emotional expression, style and tone, time information, technical details.
[0179] Step S402, the server generates target visual data based on the input text and the target reference object; wherein, the target visual data includes information of the target reference object.
[0180] Here, the target visual data can be an image or a video. Taking a video as an example here, the server can call a video generation model to automatically generate a target video based on the input text and the target reference object. For example, after the user inputs text, the visual data generation application will automatically generate a video script and generate a video.
[0181] In the embodiments of the present application, after the user inputs text and selects a target reference object, the video generation model can parse the content of the input text, extract keywords, and generate a target video based on the keywords of the input text and the target reference object.
[0182] After generating the target video, the user can further edit the generated target video, such as adding subtitles, adjusting the background music, modifying the lens switching, etc. After completion of the editing, the target video can be exported and shared on social media or other platforms.
[0183] In some embodiments, relevant reference objects can also be matched from the reference object library based on the input text of the user, so as to generate a target video based on the input text and the matched reference objects.
[0184] The visual data generation method provided by the embodiments of the present application generates expected target visual data based on pre-generated reference objects, and the reference objects include reference images and attribute information. In this way, in the process of generating the reference objects, by generating reference objects including reference images and attribute information, and then generating expected visual data based on the reference objects, not only can the generation quality of the visual data be improved, but also because the reference objects can be applied to different visual data generation tasks, the application scenarios of the reference objects are expanded, and the computing resources and storage resources of the visual data generation system are saved.
[0185] Figure 8 is another optional process schematic diagram of the visual data generation method provided by the embodiments of the present application. As Figure 8 shown, the method includes the following steps S501 to step S507:
[0186] Step S501, the terminal receives a visual data generation operation input by the user.
[0187] Step S502, the terminal generates a visual data generation request in response to the visual data generation operation.
[0188] Step S503, the terminal sends the visual data generation request to the server.
[0189] Here, after receiving the visual data generation operation input by the user, the terminal can, in response to the visual data generation operation, automatically generate a visual data generation request and send the visual data generation request to the server.
[0190] The terminal can send a visual data generation request to the server through protocols such as HTTP or Web Socket to request the generation of target visual data.
[0191] In step S504, the server, in response to the visual data generation request, obtains at least one target reference object from the reference object library and obtains the input text.
[0192] The reference objects in the reference object library are generated by using the above-mentioned reference object processing method.
[0193] In step S505, the server generates target visual data based on the input text and the target reference object.
[0194] The target visual data includes information of the target reference object.
[0195] In step S506, the server sends the target visual data to the terminal.
[0196] In step S507, the terminal displays the target visual data on the current interface.
[0197] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.
[0198] The embodiments of the present application provide a reference object processing method and a visual data generation method. The reference object in this method can be called a "subject". The embodiments of the present application will be described by taking the visual data as a video as an example.
[0199] See Figure 9 , Figure 9 is a functional interface diagram of the visual data generation method provided by the embodiments of the present application. The user can save the characters, props, or scenes referred to in the function of using the reference to generate a video 901 into the subject library. In the function of using the reference to generate a video 901, one or more subjects that have been created can be selected from the subject library of my subjects 902, and then the desired video can be generated based on the selected subjects and the input text 903 (the input text is used to indicate the actions or display effects of the selected subjects in the final video).
[0200] Next, the subject creation process (i.e., the above-mentioned reference object processing method) provided by the embodiments of the present application will be described.
[0201] In the embodiments of the present application, when starting the function of creating a subject and clicking to create a subject, there can be at least two paths. Among them, in the first path, Figure 10 is the start interface diagram of the subject creation process provided by the embodiments of the present application. As Figure 10As shown, click on my subject 902 in the reference live video 901, then you can select the created subject to generate a video, or you can choose to create a subject. When choosing to create a subject, you can click on the plus button 1001 in Figure 10 (i.e., the above-mentioned reference object creation button). After clicking the plus button 1001, you will enter the reference object creation interface as shown in Figure 11 . In the reference object creation interface shown in Figure 11 , there is a display of an interactive reference image upload area 1101 and an interactive attribute information area 1102.
[0202] In the second path, as shown in the product home page interface diagram in Figure 12 , after clicking on my subject 1201 on the product home page, you will enter the interface as shown in Figure 13 . Then select to create a subject 1301 (i.e., the above-mentioned reference object creation button), and then you will enter the reference object creation interface as shown in Figure 14 . In the reference object creation interface shown in Figure 14 , there is a display of an interactive reference image upload area 1401 and an interactive attribute information area 1402. It should be noted that the reference object creation interface entered through the above first path can be exactly the same as the reference object creation interface entered through the second path.
[0203] After entering the reference object creation interface, the user can input subject-related information: at least upload at least one picture of a subject (such as a person, an object, or a scene) (i.e., the above-mentioned reference image), and input a "subject name" and "tags". Among them, the specific tags can be selected from historical creations or entered independently. As shown in the reference object creation interface in Figure 15 , one reference image 1501 has been input, the "subject name" is "Male Baby", and tags can also be input. Refer to Figure 16 , tags can also be selected from historical creation 1601.
[0204] In some embodiments, it is also possible to search for created subjects in my subject library through specific tags. As shown in Figure 17 , through the "person" tag, two created subjects 1701 are found. Further, at this time, if you choose to create a subject, the selected tags can be automatically substituted. As shown in Figure 18 , the automatically substituted tag 1801 is "person", and there is no need to input it again.
[0205] In the embodiments of the present application, relevant information of the subject can also be automatically generated. Based on the pictures input by the user, through a multimodal understanding model, the "style" of the subject to be created (i.e., the above-mentioned visual effect type text) and the "description" of the subject (i.e., the above-mentioned image content description text) can be automatically generated. Among them, the "style" and the "description" of the subject can be manually edited and modified. For example, Figure 19 as shown, the automatically generated "style" 1901 of the uploaded image is "realistic", and the "description" 1902 of the uploaded image is "A cute boy, about three or four years old, wearing a white T-shirt with a cartoon pattern. His hair is curly and fluffy, with an innocent expression on his face. He is holding a cookie in his hand and seems to be about to eat it."
[0206] It should be noted that in the process of generating the "style" and the "description" of the subject to be created, it is a very important step to confirm that the information described by the subject is correct. At this step, the user can correct errors to ensure that the information described by the multimodal understanding model is correct, so as to ensure that the generated result better meets the user's expectations.
[0207] After the subject is successfully created, the created subject can also be edited. For example, the uploaded picture can be deleted, a new picture can be added, and the name, label, style, and description of the subject can be edited. As Figure 20 shown, the edit button displayed on the subject 2001 can be clicked to edit the subject 2001, or the delete button displayed on the subject 2001 can be clicked to delete the subject 2001. After clicking the edit button, an edit interface as Figure 21 shown can be entered. In this edit interface, the delete button 2101 can also be clicked to delete the previously uploaded picture, the plus button 2102 can be clicked to add a new picture, and the information such as the name, label, style, and description of the subject can be modified. Among them, Figure 21 shows the interface for modifying the name and label of the subject, Figure 22 shows the interface for modifying the style and description.
[0208] The embodiments of the present application can not only improve the user's creation efficiency, and the user does not need to upload each time when using the subject, but also, according to the current evaluation results, it can better ensure the feature stability of the subject and improve the video generation effect.
[0209] Based on the reference object processing method described in the above embodiments, Figure 23The following shows a structural block diagram of a reference object processing device provided by an embodiment of the present application. The reference object processing device 100 may be a device in an electronic device (such as a server). The reference object processing device can be implemented in software, and it can be software in the form of programs and plugins, etc., including the following software modules: A first display module 101, configured to display a reference object creation interface in response to a trigger operation for creating a button for the reference object; the reference object creation interface includes an interactive reference image upload area and an interactive attribute information area; A second display module 102, configured to display the uploaded reference image in the reference image upload area in response to an image upload operation performed in the reference image upload area; A third display module 103, configured to display the attribute information of the reference image in the attribute information area; A generation module 104, configured to generate a reference object including the reference image and the attribute information in response to a trigger operation for a confirmation button; the reference object is used to provide reference information when generating visual data.
[0210] In some embodiments, the third display module 103 is further configured to: when displaying the reference image, display the attribute information generated after image recognition of the reference image in the attribute information area; or, in response to an information input operation performed in the attribute information area, display the input attribute information in the attribute information area.
[0211] In some embodiments, the generated attribute information includes at least one of the following: a visual effect type text and an image content description text; the third display module 103 is further configured to: display the visual effect type text and the image content description text generated after image recognition of the reference image.
[0212] In some embodiments, the third display module 103 is further configured to: after displaying the attribute information generated after image recognition of the reference image, display the rewritten attribute information in the attribute information area in response to a rewrite operation for the generated attribute information.
[0213] In some embodiments, the attribute information includes a tag of the reference image; the third display module 103 is further configured to: display the input tag in the attribute information area in response to a tag input operation performed in the attribute information area; or, display the selected tag in the attribute information area in response to a tag selection operation performed in the attribute information area.
[0214] In some embodiments, the first display module 101 is further configured to: in response to a query operation on the reference object library performed in the visual data generation interface, display a reference object display interface; wherein, the reference object display interface includes thumbnails of reference objects in the reference object library and the reference object creation button; the reference object library includes at least one type of reference object; in response to a trigger operation on the reference object creation button in the reference object display interface, display the reference object creation interface.
[0215] In some embodiments, the first display module 101 is further configured to: in response to a query operation on the reference object library performed in the data processing interface, display a reference object display interface; wherein, the reference object display interface includes thumbnails of reference objects in the reference object library and the reference object creation button; the reference object library includes at least one type of reference object; in response to a trigger operation on the reference object creation button in the reference object display interface, display the reference object creation interface.
[0216] In some embodiments, the first display module 101 is further configured to: in response to a query operation on the reference object corresponding to the target label, display a reference object display interface corresponding to the target label; wherein, the reference object display interface includes thumbnails of reference objects with the target label and the reference object creation button; in response to a trigger operation on the reference object creation button in the reference object display interface, display the reference object creation interface; wherein, in the attribute information area of the reference object creation interface, the target label is displayed.
[0217] In some embodiments, the generation module 104 is further configured to: after generating a reference object including the reference image and the attribute information, in response to a processing operation on the reference image, generate a reference object including at least the processed reference image; in response to an editing operation on the attribute information, generate a reference object including at least the edited attribute information.
[0218] Based on the visual data generation method described in the above embodiments, Figure 24The structure block diagram of a visual data generation device provided by an embodiment of the present application is shown. The visual data generation device 200 may be a device in an electronic device (such as a server). The visual data generation device can be implemented in software, and it can be software in the form of programs and plugins, including the following software modules: an acquisition module 201, configured to obtain at least one target reference object from a reference object library in response to a visual data generation request, and obtain input text; wherein, the reference objects in the reference object library are generated by using the above-mentioned reference object processing method; a target visual data generation module 202, configured to generate target visual data based on the input text and the target reference object; wherein, the target visual data includes information of the target reference object.
[0219] It should be noted that the description of the device in the embodiment of the present application is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment, so it will not be elaborated here. For the technical details not disclosed in this device embodiment, please refer to the description of the method embodiment of the present application for understanding.
[0220] An embodiment of the present application provides an electronic device, Figure 25 which is a schematic structural diagram of the electronic device provided by the embodiment of the present application. As Figure 25 shown, the electronic device 130 includes: at least one processor 131 ( Figure 12 only one is shown herein), a memory 132, and computer-executable instructions 133 stored in the memory 132 and executable on at least one processor 131. When the processor 131 executes the executable instructions 133, the steps in any of the above-mentioned reference object processing methods or visual data generation method embodiments are implemented.
[0221] The electronic device may include but is not limited to the processor 131 and the memory 132. Those skilled in the art can understand that Figure 25 this is only an example of the electronic device 130, and does not constitute a limitation on the electronic device 130. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0222] The processor 131 may be a central processing unit (CPU). The processor 131 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0223] In some embodiments, the memory 132 may be an internal storage unit of the electronic device 130, such as the hard disk or memory of the electronic device 130. In some other embodiments, the memory 132 may also be an external storage device of the electronic device 130, such as a plug-in hard disk equipped on the electronic device 130, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 132 may also include both the internal storage unit and the external storage device of the electronic device 130. The memory 132 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of a computer program. The memory 132 may also be used to temporarily store data that has been output or is to be output.
[0224] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions. The computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the reference object processing method or the visual data generation method described above in the embodiments of the present application.
[0225] An embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions or a computer program are stored. When the computer-executable instructions or the computer program are executed by a processor, the processor will be caused to execute the reference object processing method or the visual data generation method provided in the embodiments of the present application. For example, Figure 1 the reference object processing method shown in Figure 7 or the visual data generation method shown in
[0226] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.
[0227] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0228] As an example, the computer-executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program under discussion, or stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code).
[0229] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed at multiple locations and interconnected by a communication network.
[0230] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are all included in the protection scope of the present application.
Claims
1. A reference object processing method, characterized in that: The method comprises: In response to a trigger operation on a reference object creation button, a reference object creation interface is displayed; the reference object creation interface includes an interactive reference image upload area and an interactive attribute information area; In response to an image upload operation performed in the reference image upload area, displaying an uploaded reference image in the reference image upload area; Displaying the attribute information of the reference image in the attribute information area; In response to a trigger operation on a confirmation button, a reference object including the reference image and the attribute information is generated; the reference object is used to provide reference information when generating visual data.
2. The method according to claim 1, characterized in that The displaying the attribute information of the reference image in the attribute information area includes: When displaying the reference image, in the attribute information area, attribute information generated after image recognition is performed on the reference image is displayed; or, In response to an information input operation performed in the attribute information area, the input attribute information is displayed in the attribute information area.
3. The method according to claim 2, characterized in that The generated attribute information includes at least one of the following: a visual effect type text and an image content description text; the display of the attribute information generated after performing image recognition on the reference image includes: Display the visual effect type text and image content description text generated after image recognition is performed on the reference image.
4. The method according to claim 2, characterized in that: After displaying the attribute information generated after performing image recognition on the reference image, the method further includes: In response to a rewriting operation on the generated attribute information, the rewritten attribute information is displayed in the attribute information area.
5. The method according to claim 2, characterized in that: The attribute information includes a label of the reference image; The step of displaying the inputted attribute information in the attribute information area in response to the information input operation performed in the attribute information area comprises: In response to a tag input operation performed in the attribute information area, displaying the input tag in the attribute information area; or, In response to a tag selection operation performed in the attribute information area, the selected tag is displayed in the attribute information area.
6. The method according to any one of claims 1 to 5, characterized in that: The step of displaying a reference object creation interface in response to a trigger operation on a reference object creation button includes: In response to a query operation on a reference object library executed on a visual data generation interface, a reference object display interface is displayed; wherein the reference object display interface includes thumbnails of reference objects in the reference object library and a reference object creation button; the reference object library includes at least one type of reference object; In response to a triggering operation on the reference object creation button in the reference object display interface, the reference object creation interface is displayed.
7. The method according to any one of claims 1 to 5, characterized in that: The step of displaying a reference object creation interface in response to a trigger operation on a reference object creation button includes: In response to a query operation on a reference object library executed on a data processing interface, a reference object display interface is displayed; wherein the reference object display interface includes thumbnails of reference objects in the reference object library and a reference object creation button; the reference object library includes at least one type of reference object; In response to a triggering operation on the reference object creation button in the reference object display interface, the reference object creation interface is displayed.
8. The method according to any one of claims 1 to 5, characterized in that: The step of displaying a reference object creation interface in response to a trigger operation on a reference object creation button includes: In response to a query operation for a reference object corresponding to a target tag, displaying a reference object display interface corresponding to the target tag; wherein the reference object display interface includes a thumbnail of the reference object having the target tag and a reference object creation button; In response to a triggering operation on the reference object creation button in the reference object display interface, the reference object creation interface is displayed; wherein the target tag is displayed in the attribute information area of the reference object creation interface.
9. The method according to any one of claims 1 to 5, characterized in that: After generating a reference object including the reference image and the attribute information, the method further includes: In response to the processing operation on the reference image, generating a reference object including at least a processed reference image; In response to an editing operation on the attribute information, a reference object including at least the edited attribute information is generated.
10. A method for generating visual data, characterized in that: The method comprises: In response to a visual data generation request, obtaining at least one target reference object from a reference object library, and obtaining input text; wherein the reference objects in the reference object library are generated using the reference object processing method according to any one of claims 1 to 9; Based on the input text and the target reference object, target visual data is generated; wherein the target visual data includes information of the target reference object.
11. A reference object processing device, characterized in that: The device comprises: A first display module, configured to display a reference object creation interface in response to a trigger operation on a reference object creation button; the reference object creation interface includes an interactive reference image upload area and an interactive attribute information area; a second display module, configured to display the uploaded reference image in the reference image upload area in response to an image upload operation performed in the reference image upload area; A third display module, used for displaying the attribute information of the reference image in the attribute information area; A generation module is used to generate a reference object including the reference image and the attribute information in response to a trigger operation on a confirmation button; the reference object is used to provide reference information when generating visual data.
12. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions or computer programs; A processor, for implementing the reference object processing method described in any one of claims 1 to 9, or implementing the visual data generating method described in claim 10, when executing the computer executable instructions or computer program stored in the memory.
13. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the reference object processing method described in any one of claims 1 to 9 is implemented, or the visual data generating method described in claim 10 is implemented.
14. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the reference object processing method described in any one of claims 1 to 9 is implemented, or the visual data generating method described in claim 10 is implemented.