Data processing method and device and electronic equipment
By combining the descriptive information of multimedia data and user input information in the generative model and adjusting the video material data in real time, the problem of low user engagement on existing short video platforms is solved, and more efficient video generation and improved user experience are achieved.
Patent Information
- Application Number
- CN202510719804.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
AI Technical Summary
The video editing function of existing short video platforms mainly relies on post-processing of native videos, which has low user participation and insufficient interactivity, resulting in a poor user experience during the video generation process.
By determining the description information set of multiple images in the multimedia data and inputting the user input information and the description information set into the generative model, the material data of the target video is generated, allowing the user to adjust the timestamp, audio time period, text data, background audio, cover data and transition data in the material data in real time to generate the target video.
It improves the user's participation and real-time control ability in the video generation process, reduces the adjustment and regeneration process, and improves video generation efficiency and user experience.
Smart Images

Figure CN120658922A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a data processing method, device, and electronic device. Background Art
[0002] Artificial Intelligence Generated Content (AIGC) technology is rapidly developing in the short video production space. It can automatically generate creative video content based on text, image, or voice input, significantly lowering the barrier to entry for video production. Users can quickly create personalized short videos without requiring professional editing skills, significantly improving content creation efficiency. However, the video editing capabilities of existing short video platforms primarily rely on post-processing of native videos, and user engagement and interactivity during the video generation process are low. Summary of the Invention
[0003] The present disclosure provides a data processing method, device, and electronic device to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of the present disclosure, there is provided a data processing method, comprising:
[0005] Determining a set of description information corresponding to a plurality of images in the multimedia data;
[0006] Inputting the user input information and the description information set into a generative model, so that the generative model obtains material data corresponding to the target video based on the user input information and the description information set;
[0007] The target video is generated based on the material data.
[0008] In the above solution, generating the target video based on the material data includes:
[0009] Output the video to be confirmed;
[0010] adjusting at least one of a timestamp of at least one target image, a time period of audio, text data, background audio, cover data, and transition data in the material data based on first feedback information from a user regarding the video to be confirmed, to obtain adjusted material data corresponding to the target video;
[0011] The target video is generated based on the adjusted material data corresponding to the target video.
[0012] In the above solution, if the multimedia data is a video, then determining a set of description information corresponding to a plurality of images in the multimedia data includes at least one of the following:
[0013] Sampling video frames in a video to obtain a plurality of images; determining a set of description information corresponding to the plurality of images based on a timestamp of each image in the video and a text description and keywords corresponding to each image;
[0014] Segmenting the audio data in the video to obtain multiple audio segments; determining a set of description information corresponding to the multiple images based on a time period and / or text content corresponding to each audio segment;
[0015] Alternatively, if the multimedia data is an image, determining a set of description information corresponding to a plurality of images in the multimedia data includes:
[0016] Determine a text description and / or keyword corresponding to each image, and determine a description information set corresponding to the multiple images.
[0017] In the above solution, the step of inputting the received user input information and the description information set into the generative model so that the generative model processes the multimedia data based on the user input information and the description information set to obtain material data corresponding to the target video includes at least one of the following:
[0018] Determining, based on the generative model, at least one image corresponding to the text description and / or keyword corresponding to the first key information included in the user input information in the description information set as the target image corresponding to the target video; determining a timestamp corresponding to the target image as the material data corresponding to the target video;
[0019] Acquire, based on the generative model, a text description and / or keyword corresponding to the first key information in the description information set; determine, based on the text description and / or keyword, text data corresponding to the target video; and determine that the text data is material data corresponding to the target video;
[0020] Acquire a text description and / or keyword corresponding to the first key information in the description information set based on the generative model; determine background audio corresponding to the target video based on the text description and / or keyword; and identify the background audio as material data corresponding to the target video;
[0021] Acquire a text description and / or keyword corresponding to the first key information in the description information set based on the generative model; determine cover data corresponding to the target video based on the text description and / or keyword; and determine that the cover data is material data corresponding to the target video;
[0022] Determining transition data of the target video corresponding to the first key information based on the generative model; determining an identifier of the transition data as material data corresponding to the target video;
[0023] Based on the generative model, the text content corresponding to the first key information in the description information set is obtained; based on the text content, the target audio corresponding to the target video is determined; and the target audio is determined to be the material data corresponding to the target video.
[0024] In the above solution, the at least one image corresponding to the text description and / or keyword in the description information set that matches the first key information included in the user input information is determined based on the generative model as the target image corresponding to the target video, including at least one of the following:
[0025] In response to the first key information including content amount indication information, the generative model determines, based on the text description of each image, at least one image whose content amount satisfies the content amount indication information, as a target image corresponding to the target video;
[0026] In response to the first key information including emotion indication information, the generative model determines, based on keywords or text descriptions of each image, at least one image that satisfies the emotion indication information as a target image corresponding to the target video;
[0027] In response to the first key information including target indication information, the generative model determines, based on the keywords or text description of each image, at least one image including the target indication information as the target image corresponding to the target video.
[0028] In the above solution, after obtaining the material data corresponding to the target video, the method further includes:
[0029] Obtaining second feedback information from the user regarding the material data corresponding to the target video;
[0030] The generative model is enabled to adjust at least one of the timestamp of the target image, the time period of the target audio, the text data, the background audio, the cover data and the transition data in the material data based on the second feedback information to obtain the adjusted material data corresponding to the target video.
[0031] In the above solution, the generative model adjusts at least one of the timestamp of the target image, the time period of the target audio, the text data, the background audio, the cover data, and the transition data in the material data based on the second feedback information to obtain the adjusted material data corresponding to the target video, including at least one of the following:
[0032] The generative model determines second key information included in the second feedback information; and at least one of deleting a timestamp of at least one target image corresponding to the second key information in the material data, adding a timestamp of at least one image corresponding to the second key information in the description information set to the material information, and adjusting an order of the timestamps of at least one target image corresponding to the second key information in the material data;
[0033] The generative model updates the copy data in the material information based on the second feedback information;
[0034] The generative model replaces or edits the background audio data in the material information based on the second feedback information;
[0035] The generative model replaces or edits the cover data in the material information based on the second feedback information;
[0036] The generative model replaces or edits the transition data in the material information based on the second feedback information;
[0037] The generative model determines the second key information included in the second feedback information; deletes at least one time period of the target audio corresponding to the second key information in the material data, adds at least one time period of the target audio corresponding to the second key information in the description information set to the material information, and adjusts the playback position of at least one time period of the target audio corresponding to the second key information in the material data.
[0038] In the above solution, inputting the material data corresponding to the target video into the generative model to obtain the target video includes at least one of the following steps performed by the generative model:
[0039] Acquire the target image based on the timestamp of each target image in the material data, and determine the order in which the target images appear in the target video;
[0040] Determining the playback position of the target audio in the target video based on the time period of the target audio and the timestamp of the target image in the material data; or determining the playback position of the target audio in the target video based on the indicated position in the material data;
[0041] Using the cover data in the material data as the cover of the target video;
[0042] Using the background audio data in the material data as the background audio of the target video;
[0043] Using the text data in the material data as the text corresponding to the target video;
[0044] The transition data in the material data is used as a transition picture between two target images in the target video.
[0045] According to a second aspect of the present disclosure, there is provided a data processing device, comprising:
[0046] An identification module, configured to determine a set of description information corresponding to a plurality of images in multimedia data;
[0047] a processing module, configured to input the user input information and the descriptive information set into a generative model, so that the generative model obtains material data corresponding to the target video based on the user input information and the descriptive information set;
[0048] A synthesis module is used to generate the target video based on the material data.
[0049] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0050] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the present disclosure.
[0051] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings, in which several embodiments of the present disclosure are shown by way of example and not limitation, wherein:
[0053] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts.
[0054] Figure 1 A first optional flow chart of the data processing method provided by an embodiment of the present disclosure is shown;
[0055] Figure 2 A second optional flow chart of the data processing method provided by the embodiment of the present disclosure is shown;
[0056] Figure 3 A third optional flow chart of the data processing method provided by the embodiment of the present disclosure is shown;
[0057] Figure 4A fourth optional flow chart of the data processing method provided by an embodiment of the present disclosure is shown;
[0058] Figure 5 A fifth optional flow chart of the data processing method provided in an embodiment of the present disclosure is shown;
[0059] Figure 6 A sixth optional flow chart of the data processing method provided in an embodiment of the present disclosure is shown;
[0060] Figure 7 An architectural diagram illustrating a data processing method provided by an embodiment of the present disclosure is shown;
[0061] Figure 8 An optional structural diagram of a data processing device provided by an embodiment of the present disclosure is shown;
[0062] Figure 9 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0063] To make the purposes, features, and advantages of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative work shall fall within the scope of protection of the present disclosure.
[0064] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0065] In the following description, the terms "first\second" are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.
[0066] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by those skilled in the art in the art of this disclosure. The terms used in this disclosure are only for the purpose of describing the embodiments of this disclosure and are not intended to limit this disclosure.
[0067] It should be understood that in the various embodiments of the present disclosure, the size of the serial number of each implementation process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.
[0068] Figure 1 A first optional flow chart of the data processing method provided by an embodiment of the present disclosure is shown, and will be explained according to each part.
[0069] Step S101: Determine a description information set corresponding to multiple images in multimedia data.
[0070] In some embodiments, the multimedia data may include video or image; the description information set may include at least one of a text description, text keywords and a timestamp for at least one image in the multimedia data; wherein the timestamp includes the moment when the image corresponding to the text description is played in the video when the multimedia data is a video; or the order of the images corresponding to the text description in the multimedia data when the multimedia data is an image.
[0071] In some embodiments, the text description includes a description of the content presented by the image; and the text keywords include keywords obtained from the text description.
[0072] Step S102: inputting the user input information and the description information set into a generative model, so that the generative model obtains material data corresponding to the target video based on the user input information and the description information set.
[0073] In some embodiments, the user input information may include a prompt word of the generative model, which is used to instruct the generative model to determine the material data of the target video based on the descriptive information set. Optionally, a prompt word template can be pre-set to enable the user to determine the user input information based on the prompt word template and actual needs.
[0074] In specific implementation, the user input information may include information input before the material data is first generated; it may also include feedback information based on the material data after the generative model outputs the material data; it may also include feedback information based on the target video after the target video is generated.
[0075] In some embodiments, the material data may include at least one of a timestamp of at least one target image, a time period of audio, text data, background audio, cover data, and transition data.
[0076] In some embodiments, the timestamp of the at least one target image and the time period of the audio can be determined by the generative model based on matching the user input information with the description information set; at least one of the copy data, background audio, cover data and transition data can be determined based on the user input information, based on the text description and / or keywords of the at least one target image, or based on at least one of the text description, keywords or text content included in the description information set.
[0077] Step S103: Generate the target video based on the material data.
[0078] In some embodiments, at least one target image and audio are obtained from multimedia data based on the timestamp of at least one target image and the time period of the audio included in the material data; and the target video is generated based on at least two of the at least one target image, audio, text data, background audio, cover data, and transition data.
[0079] Among them, the target image is the video frame played or displayed in the target video; the audio is the sound played in the target video, the copy data includes the introduction corresponding to the target video, the copy data can also include the copy corresponding to each video in the target video, the background audio includes the background sound when the target video is played, the transition data includes the transition filter of at least one video frame in the target video, and the transition data can also include the transition animation or transition special effects between any two adjacent video frames in the target video.
[0080] In this way, through the data processing method provided by the embodiment of the present disclosure, the image frames in the native video (such as multimedia data) can be directly processed at the frame level. In the process of generating the target video, the user input information is interacted with the generative model to improve user participation, increase the real-time control and dynamic adjustment capabilities of the generation process, so that the generated target video can meet user needs to the greatest extent, reduce the adjustment and regeneration process, and improve video generation efficiency and user experience.
[0081] Figure 2 A second optional flow chart of the data processing method provided in an embodiment of the present disclosure is shown, and will be explained according to each step.
[0082] Step S201: In response to the multimedia data being a video, a description information set is determined based on the video.
[0083] In some embodiments, step S201 may be implemented based on a generative model. Optionally, the module of the generative model implementing step S201 may be the same as or different from the module of the generative model implementing steps S203 to S204.
[0084] In some embodiments, the video is input into the generative model, which samples video frames in the video to obtain multiple images. The generative model processes each of the multiple images, obtains text descriptions and keywords corresponding to each image, and determines the playback time of each image in the video as a timestamp. The text descriptions, keywords, and timestamps corresponding to all images are determined to constitute the description information set. The sampling frequency can be set according to actual needs.
[0085] The generative model processes each of the multiple images to obtain a text description and keywords corresponding to each image. This may include the generative model separately identifying the content presented in each of the multiple images and determining text that objectively describes the content presented in each image as the text description corresponding to the image. The keywords are at least one of key words, emotions, environments, and scenes extracted from the text description.
[0086] For example, if any frame in a video contains a red mouse, the text description corresponding to the image might include: "The image background is a portion of a black desktop, with a red mouse on the desktop. The red mouse is located near the center on the right side of the image. The red mouse has a streamlined, compact design and a smooth surface." Furthermore, keywords might include: "mouse," "desktop," and "red."
[0087] For example, if any frame in a video includes a little girl in a bush of flowers, the text description corresponding to the image might include: "A little girl in a yellow dress is sitting in a bush of roses. The little girl is in the center of the image and has a smiling expression." Keywords might include: "bush of flowers," "little girl," and "happy."
[0088] In some embodiments, the generative model can also process the audio data in the video, specifically, it can segment the audio data to obtain multiple audio segments; the generative model identifies the content and emotion expressed by each audio segment, determines the text content corresponding to each audio segment; and determines the time period of each audio segment based on the playback time period of each audio segment in the video. The content and emotion expressed by each audio segment can include voice-to-text processing of the audio, and the obtained text is the text content; or, the generative model determines the emotion of the audio based on the music features, tone features, and key sound features (such as crying, laughing, shouting, etc.) in the audio, and determines the emotion-related description as the text content corresponding to the audio segment.
[0089] When segmenting, you can segment by time, with each segment having the same audio duration; you can also segment by content, for example, if the audio data includes a conversation between multiple people, then the continuous sound output by each person is a segment; or you can segment by audio waveform, with each sentence or paragraph being a segment.
[0090] Alternatively, in step S202 , in response to the multimedia data being a plurality of images, a description information set is determined based on the plurality of images.
[0091] In some embodiments, the generative model processes each of the multiple images, obtains a text description and keywords corresponding to each image, and determines the playback time of each of the multiple images in the video as a timestamp; determines that the text descriptions, keywords and timestamps corresponding to all images constitute the description information set.
[0092] The specific processing method is the same as the processing method of the at least one image data sampled in step S201, and will not be repeated here.
[0093] Step S203: input the user input information and the description information set into the generative model, so that the generative model obtains material data corresponding to the target video based on the user input information and the description information set.
[0094] The specific steps of step S203 are the same as those of step S102 and will not be repeated here.
[0095] Step S204: Generate the target video based on the material data.
[0096] The specific steps of step S204 are the same as those of step S103 and will not be repeated here.
[0097] In this way, through the data processing method provided by the embodiment of the present disclosure, a target video can be generated based on a video or image, and image-level processing can be achieved. In the process of generating the target video, user input information is interacted with the generative model to enhance user participation, increase the real-time control and dynamic adjustment capabilities of the generation process, so that the generated target video can meet user needs to the greatest extent, reduce the adjustment and regeneration process, and improve video generation efficiency and user experience.
[0098] Figure 3 A third optional flow chart of the data processing method provided in an embodiment of the present disclosure is shown, and will be explained according to each step.
[0099] Figure 3 The figure shows the specific steps involved in step S102 or steps S202 and S203 for obtaining the material data corresponding to the target video based on the user input information and the description information set. The material data includes at least one of the timestamp of at least one target image, text data, background audio identifier, cover data, transition data identifier, and target audio. Determining the material data corresponding to the target video includes:
[0100] Step S301: determining a target image based on a description information set and user input information.
[0101] In some embodiments, the generative model determines that at least one image corresponding to the text description and / or keyword corresponding to the first key information included in the user input information in the description information set is the target image corresponding to the target video; and determines that the timestamp corresponding to the at least one image is the material data corresponding to the target video.
[0102] In specific implementation, the carrier can obtain first key information from the user input information, match the first key information with text descriptions and / or keywords corresponding to multiple images, and determine that at least one image corresponding to the successfully matched text description and / or keyword is the target image corresponding to the target video.
[0103] In some embodiments, the first key information may include at least one of content information and time information, namely, target indication information, used to indicate the images containing the target to be screened. The content information includes the content contained in the images to be matched, such as colors, people, clothing, buildings, animals, plants, tools, natural landscapes, etc. For example, if the user input information includes extracting scenery from a video to generate a new video, the first key information may include "scenery," and the generative model determines at least one image whose text description or keyword includes "scenery" as the target image. Alternatively, if the user input information includes selecting a scene of a person wearing purple clothes from a video to generate a new video, the first key information may include "purple" and "people," and the generative model determines at least one image whose text description or keyword includes "purple" and "people" as the target image. Alternatively, if the user input information includes editing a scene of a goal from a video to generate a new video, the first key information may include "goal," and the generative model determines at least one image whose text description or keyword includes "goal" as the target image, as well as at least one image between a seconds before and b seconds after the at least one image.
[0104] In other embodiments, the user input information may not include specific content information and time information, that is, the user input information only includes generation instructions, such as: "Extract the wonderful parts of this video to generate a new video", then the first key information can be content quantity indication information.
[0105] During specific implementation, the generative model selects images containing a large amount of content from the multimedia data as target images. Specifically, the generative model determines, based on the textual description of each image, at least one image whose content volume satisfies the content volume indication information, as the target image corresponding to the target video. The generative model can determine the content volume based on the number of characters in the textual description, i.e., the more characters in the textual description, the more content volume; or the generative model can determine the content volume based on the number of objects in the textual description, i.e., the more objects in the textual description, the more content volume; or the generative model can determine the content volume based on the action information included in the textual description, i.e., the more relevant descriptions of the action information in the textual description, the more content volume.
[0106] In some further embodiments, the user input information may include emotion-indicating information, such as "extract the crying part of a video to generate a new video." In this case, the first key information may be the emotion-indicating information. The generative model uses "crying" as the key information to match the text description and keywords, and determines that at least one image containing "crying" is the target image.
[0107] In some embodiments, after determining the target image, the timestamp corresponding to the target image is determined as the material data corresponding to the target video, that is, the target image is represented by the timestamp in the material data.
[0108] Step S302: Determine the text data.
[0109] In some embodiments, the generative model can determine the text data of the target video based on the text description and keywords of at least one target image; it can also determine the text data of the target video based on user input information and the text descriptions of all images in the description information set.
[0110] In a specific implementation, the generative model determines text descriptions of multiple images used to generate the copy data, and then generates the copy data based on the text descriptions. The copy data includes at least one of an introduction to the target video, the copy of any image in the target video, and the copy of the target video.
[0111] Step S303: Determine background audio.
[0112] In some embodiments, the generative model can determine the emotional attributes of the target video based on the text description of each of the multiple target images, and match it in the background audio database based on the emotional attributes, and determine that the successfully matched background audio is the background audio data corresponding to the target video.
[0113] In other embodiments, the generative model retrieves background audio data indicated in the user input information from the background audio database, determines that the background audio is background audio data corresponding to the target video, and identifies the background audio as material data corresponding to the target video; wherein the background audio identifier may include at least one of the background audio name, ID, album, composer, and performer.
[0114] Step S304: determine the cover data.
[0115] In some embodiments, the generative model can obtain the target image with the largest content or the most representative content from the at least one target image as the cover data; or, the generative model generates cover data based on the at least one target image; or, the generative model uses the image specified in the user input information as the cover data; or, the generative module processes the at least one target image based on the user input information to generate cover data.
[0116] In some embodiments, if the generative model determines that any target image is cover data, the target image is represented by a timestamp of the target image in the material data.
[0117] Step S305: determine transition data.
[0118] In some embodiments, the generative model can determine the transition data (such as transition animation, transition effects or transition filters) between the two adjacent target images based on the text descriptions of the two adjacent target images; or, the carrier determines that the transition data indicated by the user input information is the transition data of the target video.
[0119] Step S306: determine the target audio.
[0120] In some embodiments, the generative model retrieves textual content corresponding to the first key information in the descriptive information set and determines, based on the textual content, a target audio corresponding to the target video. Alternatively, the generative model determines that the audio indicated in the user input data is the target audio corresponding to the target video.
[0121] It should be noted that the execution order of the above steps S301 to S306 is not fixed, and whether a step is executed can be determined according to actual needs. For example, if the multimedia data is an image, step S306 does not need to be executed.
[0122] In this way, through the data processing method provided by the embodiment of the present disclosure, the generative model can determine the material data required to generate the target video based on the description information set and user input information, laying the foundation for the subsequent generation of the target video.
[0123] Figure 4 A fourth optional flow chart of the data processing method provided by an embodiment of the present disclosure is shown, and will be explained according to each part.
[0124] Step S401: Determine a description information set corresponding to multiple images in multimedia data.
[0125] The specific steps of step S401 are the same as those of step S101 or step S201 to step S202, and will not be repeated here.
[0126] Step S402: input the user input information and the description information set into a generative model, so that the generative model obtains material data corresponding to the target video based on the user input information and the description information set.
[0127] The specific steps of step S402 are the same as step S102 or at least one of steps S301 to S306, and will not be repeated here.
[0128] Step S403: Adjust the material data based on the second feedback information to obtain adjusted material data.
[0129] In some embodiments, after the generative model outputs the material data based on steps S401 to S402, it can also obtain second feedback information from the user regarding the material data corresponding to the target video, and adjust the material data based on the second feedback information, specifically including adjusting at least one of the timestamp of the target image, the time period of the target audio, the text data, the background audio, the cover data, and the transition data in the material data to obtain the adjusted material data corresponding to the target video.
[0130] In some embodiments, the generative model determines the second key information included in the second feedback information; deletes the timestamp of at least one target image corresponding to the second key information in the material data, adds the timestamp of at least one image corresponding to the second key information in the description information set to the material information, and adjusts the order of the timestamps of at least one target image corresponding to the second key information in the material data.
[0131] In some embodiments, the generative model updates the copy data in the material information based on the second feedback information; specifically, it may include updating the copy data of a specified target image, updating the copy data of a target video cover, or updating one of the copy data of a specified time interval in the target video.
[0132] In some embodiments, the generative model replaces or edits the background audio data in the material information based on the second feedback information. For example, the generative model re-acquires the background audio from a background audio database based on the second feedback information and replaces the current background audio; or edits the background audio, such as clipping, splicing, or adjusting the playback order of different sub-audios within the background audio.
[0133] In some embodiments, the generative model replaces or edits the cover data in the material information based on the second feedback information, which may specifically include reconfirming the cover data or regenerating the cover data.
[0134] In some embodiments, the generative model replaces or edits the transition data in the material information based on the second feedback information, which may specifically include replacing the transition data between the two specified target images according to the instructions of the second feedback information, or deleting the transition data between the two specified target images according to the instructions of the second feedback information, or adding transition data between the two specified target images according to the instructions of the second feedback information.
[0135] In some embodiments, the generative model determines the second key information included in the second feedback information; deletes at least one of the time period of the target audio corresponding to the second key information in the material data, adds the time period of at least one audio corresponding to the second key information in the description information set to the material information, and adjusts the playback position of at least one of the time period of the target audio corresponding to the second key information in the material data.
[0136] In some embodiments, step S402 to step S403 may be repeated until confirmation indication information for the last output material data is received, and it is determined that the last output material data can be used to generate the target video.
[0137] Step S404: Generate the target video based on the material data.
[0138] In some embodiments, the generative model, or other program, generates the target video based on the footage data.
[0139] In specific implementation, it includes at least one of the following: based on the timestamp of each target image in the material data, obtaining the target image and determining the order in which the target images appear in the target video; based on the time period of the target audio in the material data and the timestamp of the target image, determining the playback position of the target audio in the target video; or, determining the playback position of the target audio in the target video based on the indicated position in the material data; using the cover data in the material data as the cover of the target video; using the background audio data in the material data as the background audio of the target video; using the copy data in the material data as the copy corresponding to the target video, the copy may include the cover copy, the copy of at least one image in the target video or the copy corresponding to the target video; using the transition data in the material data as the transition screen between two target images in the target video.
[0140] In this way, through the data processing method provided by the embodiment of the present disclosure, a target video can be generated based on a video or image, and image-level processing can be achieved. In the process of generating material data and a target video, user input information (first feedback information and / or second feedback information) is used to interact with the generative model, thereby improving user participation, increasing the real-time control and dynamic adjustment capabilities of the generation process, and enabling the generated target video to meet user needs to the greatest extent, reducing the adjustment and regeneration processes, and improving video generation efficiency and user experience.
[0141] Figure 5 A fifth optional flow chart of the data processing method provided in an embodiment of the present disclosure is shown, and will be explained according to each step.
[0142] In some embodiments, after generating the target video based on step S103, step S204, or step S404, the method may further include:
[0143] Step S501: output the video to be confirmed.
[0144] In some embodiments, after determining the material data corresponding to the target video, the generative model can synthesize a video to be confirmed; the video to be confirmed can be a sub-video of the first n seconds of the target video, or a sub-video of the middle m seconds of the target video; if it is a sub-video of the middle m seconds, the video to be confirmed can be a key sub-video, that is, a sub-video that includes key content or emotional fluctuations.
[0145] In other embodiments, the video to be confirmed may also be a target video corresponding to the material data.
[0146] Step S502: Adjust the material data based on the first feedback information.
[0147] In some embodiments, the generative model receives first feedback information based on the video to be confirmed; adjusts the source data based on the first feedback information, specifically adjusting at least one of the timestamp of at least one target image, the time period of the audio, text data, background audio, cover data, and transition data in the source data, to obtain adjusted source data corresponding to the target video; and generates the target video based on the adjusted source data corresponding to the target video. The specific adjustment steps are the same as those in step S403 and are not repeated here.
[0148] In some optional embodiments, the first feedback information may be determined based on a prompt word template and user needs, that is, the user fills the feedback information into the prompt word template to obtain the first feedback information.
[0149] Optionally, the generative model may continue to repeat steps S501 to S502, that is, output a new video to be confirmed based on the adjusted material data, and receive third feedback information from the user; the third feedback information may include adjustment information or confirmation information. If the third feedback information includes adjustment information, the material data will continue to be updated and a new video to be confirmed will be output based on the material data until the received feedback information includes confirmation information, and the target video will be output based on the final material data.
[0150] In this way, through the data processing method provided by the embodiment of the present disclosure, after generating the video to be confirmed (sub-video or the entire target video), the material data can be adjusted based on the feedback information input by the user, and then a new video to be confirmed can be output based on the adjusted material data, so that the generated video can meet user needs to the greatest extent, reduce the adjustment and regeneration process, and improve video generation efficiency and user experience.
[0151] Figure 6 FIG. 6 shows a sixth optional flow chart of the data processing method provided in an embodiment of the present disclosure. Figure 7 The following is an architecture diagram of the data processing method provided by the embodiment of the present disclosure, which will be explained according to each step.
[0152] like Figure 7As shown, the generative model shown in the present disclosure includes two modules, namely LLM-A and LLM-B, wherein LLM-A is used to extract text descriptions and keywords from images and obtain text content from audio, that is, the step of determining a description information set based on multimedia data, that is, executing steps S101, S201, steps S301 to S302, and step S401; the LLM-B is used to determine the material data of the target video based on user input information and the description information set, that is, steps S102, S203, steps S301 to S306, steps S402 to S403, and steps S501 to S502. Optional steps S103, S204, and S404, that is, the process of generating a target video based on material data, can be implemented by LLM-B or through a program.
[0153] Step S601: Determine a generative model.
[0154] In some embodiments, AIPC is used as the hardware basis and is equipped with a localized multimodal model (LLM); based on the corresponding optimized prompt project (Prompt-A), a generative model LLM-A is defined; based on the corresponding optimized prompt project (Prompt-B), a generative model LLM-B is defined.
[0155] Among them, Prompt-A and Prompt-B can be modularized and added or modified according to the requirements of the application scenario; various image filters, video effects, etc. can be implemented by the generative model LLM-B or program.
[0156] Step S602: Determine multiple images based on the multimedia data provided by the user.
[0157] In some embodiments, multimedia data provided by the user is input into the generative model LLM-A, and frame processing is performed, which may specifically include completing the sampling task based on multiple downsampling schemes to achieve frame division and storing the sampled images.
[0158] Step S603: Identify the multiple images and obtain a description information set.
[0159] In some embodiments, the plurality of images are identified based on the generative model LLM-A, and a text description, keywords, and timestamp corresponding to each image are determined.
[0160] Optionally, the generative model LLM-A recognizes each image in turn, obtains the recognition content (Context-C) of the image texture, and stores it in the description information set. The recognition content may include a timestamp, keywords, and text description.
[0161] like Figure 7As shown in the figure, 4 images are sampled from the video and input into the generative model LLM-A in sequence to obtain the recognition content of each image. Specifically, it includes:
[0162] "The red mouse is located near the center on the right side of the image. It has an ergonomic, streamlined, compact design with a smooth, arc-shaped shape. The timestamp is 2948039109."
[0163] "The mouse adopts a deep red color scheme, which has a strong visual appeal. Its surface treatment process and material texture show the characteristics of a high-end product. The timestamp is 2948039110."
[0164] "A red mouse is placed on the desk, to the right of the person wearing a white shirt and a red lanyard. The timestamp is 2948039111."
[0165] "The mouse features a red body with a black scroll wheel and button assembly. Its connecting cable is clearly visible on the left side of the image. The timestamp is 2948039112."
[0166] Step S604: Conduct an interactive dialogue.
[0167] In some embodiments, the description information set is used as the context and combined with the user input information to enable the generative model LLM-B to complete the reasoning task, which may specifically include removing duplicate images or content, arranging the order in which each image appears in the target video, arranging the context, and generating material data based on the user-defined adjustment content in the user input information and the content of the prompt project (Prompt-B); determining the generation scheme of the target video based on the prompt project (Prompt-B) or the user input information, and determining the target image, effects, background audio, transition special effects, text, clips, etc. included in the target video based on the generation scheme to obtain the material data.
[0168] like Figure 7 As shown, the material data may include:
[0169] “Summary of the contents (.yaml)”
[0170] simple:1jpg,2jpg,11jpg (basic configuration items: 1.jpg, 2.jpg, 11.jpg)
[0171] metadata
[0172] Some Functions be used:
[0173] Step S605: Generate a target video based on the material data.
[0174] In some embodiments, based on the target image, timestamp and other dimensional information included in the material data, at least one sampled and stored image is reselected to generate a target video.
[0175] In some optional embodiments, after the target video is generated, feedback information may be received, and the material data may be adjusted based on the feedback information to generate the target video again.
[0176] In other embodiments, the LLM identifies the sampled video frames and determines a text description and a timestamp corresponding to each image.
[0177] Segment the audio data in the video and determine the text content and time period corresponding to each audio segment.
[0178] Based on the text description, timestamp, text content, and time period, a description information set is determined. Combined with the prompt word (Prompt-B), the generative model generates the source data. The target video is generated based on the source data.
[0179] The prompts could include "You are an AI assistant that can analyze cases based on screen text and audio text, summarize content, select background music from documents, and select screen footage to edit a 15-second video."
[0180] Material data can include:
[0181] “Theme: What the story is about…
[0182] Screen clips:
[0183] 00:00:00-00:01:00
[0184] 00:01:16-00:01:35 02:11:00
[0186] Background music: 'fad.mp3'
[0187] Filter effect: 'ddd'"
[0188] In this way, through the data processing method provided by the embodiment of the present disclosure, a target video can be generated based on a video or image, and image-level processing can be achieved. In the process of generating material data and a target video, user input information (first feedback information and / or second feedback information) is used to interact with the generative model, thereby improving user participation, increasing the real-time control and dynamic adjustment capabilities of the generation process, and enabling the generated target video to meet user needs to the greatest extent, reducing the adjustment and regeneration processes, and improving video generation efficiency and user experience.
[0189] Figure 8 An optional structural diagram of a data processing device provided by an embodiment of the present disclosure is shown, and will be explained according to each part.
[0190] In some embodiments, the data processing device 700 includes an identification module 701 , a processing module 702 and a synthesis module 703 .
[0191] The identification module 701 is used to determine a set of description information corresponding to multiple images in the multimedia data;
[0192] The processing module 702 is configured to input the user input information and the description information set into a generative model, so that the generative model obtains material data corresponding to the target video based on the user input information and the description information set;
[0193] The synthesis module 703 is configured to generate the target video based on the material data.
[0194] The processing module 702 is further configured to output the video to be confirmed;
[0195] adjusting at least one of a timestamp of at least one target image, a time period of audio, text data, background audio, cover data, and transition data in the material data based on first feedback information from a user regarding the video to be confirmed, to obtain adjusted material data corresponding to the target video;
[0196] The target video is generated based on the adjusted material data corresponding to the target video.
[0197] The identification module 701 is specifically configured to:
[0198] Sampling video frames in a video to obtain a plurality of images; determining a set of description information corresponding to the plurality of images based on a timestamp of each image in the video and a text description and keywords corresponding to each image;
[0199] The audio data in the video is segmented to obtain multiple audio segments; and a description information set corresponding to the multiple images is determined based on the time period and / or text content corresponding to each audio segment.
[0200] If the multimedia data is an image, the recognition module 701 is specifically configured to determine a text description and keywords corresponding to each image, and determine a description information set corresponding to the multiple images.
[0201] The processing module 702 is specifically configured to:
[0202] Determining, based on the generative model, at least one image corresponding to the text description and / or keyword corresponding to the first key information included in the user input information in the description information set as a target image corresponding to the target video; and determining a timestamp corresponding to the at least one image as material data corresponding to the target video;
[0203] Acquire, based on the generative model, a text description and / or keyword corresponding to the first key information in the description information set; determine, based on the text description and / or keyword, text data corresponding to the target video; and determine that the text data is material data corresponding to the target video;
[0204] Acquire a text description and / or keyword corresponding to the first key information in the description information set based on the generative model; determine background audio corresponding to the target video based on the text description and / or keyword; and identify the background audio as material data corresponding to the target video;
[0205] Acquire a text description and / or keyword corresponding to the first key information in the description information set based on the generative model; determine cover data corresponding to the target video based on the text description and / or keyword; and determine that the cover data is material data corresponding to the target video;
[0206] Determining transition data of the target video corresponding to the first key information based on the generative model; determining an identifier of the transition data as material data corresponding to the target video;
[0207] Based on the generative model, text content corresponding to the first key information in the description information set is obtained; based on the text content, a target audio corresponding to the target video is determined.
[0208] The identification module 701 is specifically configured to:
[0209] In response to the first key information including content amount indication information, the generative model determines, based on the text description of each image, at least one image whose content amount satisfies the content amount indication information, as a target image corresponding to the target video;
[0210] In response to the first key information including emotion indication information, the generative model determines, based on keywords or text descriptions of each image, at least one image that satisfies the emotion indication information as a target image corresponding to the target video;
[0211] In response to the first key information including target indication information, the generative model determines, based on the keywords or text description of each image, at least one image including the target indication information as the target image corresponding to the target video.
[0212] The identification module 701 is specifically configured to match the text description and / or keywords in a background audio database and determine that the successfully matched background audio is the background audio data corresponding to the target video.
[0213] In some embodiments, the user input information includes second feedback information regarding the material data corresponding to the target video, and the identification module 701 is further configured to obtain the second feedback information regarding the material data corresponding to the target video after obtaining the material data corresponding to the target video;
[0214] The generative model is enabled to adjust at least one of the timestamp of the target image, the time period of the target audio, the text data, the background audio, the cover data and the transition data in the material data based on the second feedback information to obtain the adjusted material data corresponding to the target video.
[0215] The identification module 701 is specifically used for one of the following:
[0216] The generative model determines second key information included in the second feedback information; and at least one of deleting a timestamp of at least one target image corresponding to the second key information in the material data, adding a timestamp of at least one image corresponding to the second key information in the description information set to the material information, and adjusting an order of the timestamps of at least one target image corresponding to the second key information in the material data;
[0217] The generative model updates the copy data in the material information based on the second feedback information;
[0218] The generative model replaces or edits the background audio data in the material information based on the second feedback information;
[0219] The generative model replaces or edits the cover data in the material information based on the second feedback information;
[0220] The generative model replaces or edits the transition data in the material information based on the second feedback information;
[0221] The generative model determines the second key information included in the second feedback information; deletes at least one time period of the target audio corresponding to the second key information in the material data, adds at least one time period of the target audio corresponding to the second key information in the description information set to the material information, and adjusts the playback position of at least one time period of the target audio corresponding to the second key information in the material data.
[0222] The synthesis module 703 is specifically configured to implement one of the following steps:
[0223] Acquire the target image based on the timestamp of each target image in the material data, and determine the order in which the target images appear in the target video;
[0224] Determining the playback position of the target audio in the target video based on the time period of the target audio and the timestamp of the target image in the material data; or determining the playback position of the target audio in the target video based on the indicated position in the material data;
[0225] Using the cover data in the material data as the cover of the target video;
[0226] Using the background audio data in the material data as the background audio of the target video;
[0227] Using the text data in the material data as the text corresponding to the target video;
[0228] The transition data in the material data is used as a transition picture between two target images in the target video.
[0229] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0230] Figure 9 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0231] like Figure 9As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0232] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0233] The computing unit 801 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the data processing method in any other appropriate manner (e.g., by means of firmware).
[0234] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0235] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0236] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0237] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0238] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0239] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0240] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0241] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means two or more, unless otherwise specifically defined.
[0242] The above description is merely a specific embodiment of the present disclosure, but the scope of protection of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this disclosure should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.
Claims
1. A data processing method, comprising: Determining a set of description information corresponding to a plurality of images in the multimedia data; Inputting the user input information and the description information set into a generative model, so that the generative model obtains material data corresponding to the target video based on the user input information and the description information set; The target video is generated based on the material data.
2. The method according to claim 1, wherein generating the target video based on the material data comprises: Output the video to be confirmed; adjusting at least one of a timestamp of at least one target image, a time period of audio, text data, background audio, cover data, and transition data in the material data based on first feedback information from a user regarding the video to be confirmed, to obtain adjusted material data corresponding to the target video; The target video is generated based on the adjusted material data corresponding to the target video.
3. The method according to claim 1, wherein if the multimedia data is a video, determining a set of description information corresponding to a plurality of images in the multimedia data comprises at least one of the following: Sampling video frames in a video to obtain a plurality of images; determining a set of description information corresponding to the plurality of images based on a timestamp of each image in the video and a text description and keywords corresponding to each image; Segmenting the audio data in the video to obtain multiple audio segments; Determining a set of description information corresponding to the plurality of images based on a time period and / or text content corresponding to each audio segment; Alternatively, if the multimedia data is an image, determining a set of description information corresponding to a plurality of images in the multimedia data includes: Determine a text description and / or keyword corresponding to each image, and determine a description information set corresponding to the multiple images.
4. The method according to claim 1, wherein the step of inputting the received user input information and the description information set into a generative model so that the generative model processes the multimedia data based on the user input information and the description information set to obtain material data corresponding to the target video comprises at least one of the following: Determining, based on the generative model, at least one image corresponding to the text description and / or keyword corresponding to the first key information included in the user input information in the description information set as the target image corresponding to the target video; determining a timestamp corresponding to the target image as the material data corresponding to the target video; Acquire, based on the generative model, a text description and / or keyword corresponding to the first key information in the description information set; determine, based on the text description and / or keyword, copywriting data corresponding to the target video; Determining that the copy data is material data corresponding to the target video; Acquire a text description and / or keyword corresponding to the first key information in the description information set based on the generative model; determine background audio corresponding to the target video based on the text description and / or keyword; Determining that the identifier of the background audio is material data corresponding to the target video; Acquire a text description and / or keyword corresponding to the first key information in the description information set based on the generative model; determine cover data corresponding to the target video based on the text description and / or keyword; and determine that the cover data is material data corresponding to the target video; Determining transition data of the target video corresponding to the first key information based on the generative model; determining an identifier of the transition data as material data corresponding to the target video; Based on the generative model, the text content corresponding to the first key information in the description information set is obtained; based on the text content, the target audio corresponding to the target video is determined; and the target audio is determined to be the material data corresponding to the target video.
5. The method according to claim 4, wherein determining, based on the generative model, at least one image corresponding to a text description and / or keyword in the description information set that matches the first key information included in the user input information as the target image corresponding to the target video comprises at least one of the following: In response to the first key information including content amount indication information, the generative model determines, based on the text description of each image, at least one image whose content amount satisfies the content amount indication information, as a target image corresponding to the target video; In response to the first key information including emotion indication information, the generative model determines, based on keywords or text descriptions of each image, at least one image that satisfies the emotion indication information as a target image corresponding to the target video; In response to the first key information including target indication information, the generative model determines, based on the keywords or text description of each image, at least one image including the target indication information as the target image corresponding to the target video.
6. The method according to claim 1, after obtaining the material data corresponding to the target video, the method further comprises: Afterwards, the method further includes: Obtaining second feedback information from the user regarding the material data corresponding to the target video; The generative model is enabled to adjust at least one of the timestamp of the target image, the time period of the target audio, the text data, the background audio, the cover data and the transition data in the material data based on the second feedback information to obtain the adjusted material data corresponding to the target video.
7. The method according to claim 6, wherein the generative model adjusts at least one of the timestamp of the target image, the time period of the target audio, the text data, the background audio, the cover data, and the transition data in the material data based on the second feedback information to obtain the adjusted material data corresponding to the target video, including at least one of the following: The generative model determines second key information included in the second feedback information; and at least one of deleting a timestamp of at least one target image corresponding to the second key information in the material data, adding a timestamp of at least one image corresponding to the second key information in the description information set to the material information, and adjusting an order of the timestamps of at least one target image corresponding to the second key information in the material data; The generative model updates the copy data in the material information based on the second feedback information; The generative model replaces or edits the background audio data in the material information based on the second feedback information; The generative model replaces or edits the cover data in the material information based on the second feedback information; The generative model replaces or edits the transition data in the material information based on the second feedback information; The generative model determines the second key information included in the second feedback information; deletes at least one time period of the target audio corresponding to the second key information in the material data, adds at least one time period of the target audio corresponding to the second key information in the description information set to the material information, and adjusts the playback position of at least one time period of the target audio corresponding to the second key information in the material data.
8. The method according to claim 1 or 7, wherein inputting the material data corresponding to the target video into the generative model to obtain the target video comprises the generative model performing at least one of the following steps: Acquire the target image based on the timestamp of each target image in the material data, and determine the order in which the target images appear in the target video; Determining a playback position of the target audio in the target video based on a time period of the target audio and a timestamp of the target image in the material data; Alternatively, determining the playback position of the target audio in the target video based on the indicated position in the material data; Using the cover data in the material data as the cover of the target video; Using the background audio data in the material data as the background audio of the target video; Using the text data in the material data as the text corresponding to the target video; The transition data in the material data is used as a transition picture between two target images in the target video.
9. A data processing device comprising: An identification module, configured to determine a set of description information corresponding to a plurality of images in multimedia data; a processing module, configured to input the user input information and the descriptive information set into a generative model, so that the generative model obtains material data corresponding to the target video based on the user input information and the descriptive information set; A synthesis module is used to generate the target video based on the material data.
10. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 8.