Special-effects processing method and apparatus, and electronic device and storage medium

By obtaining and processing special effect description data, generating and displaying special effect effect data, the problems of complexity and user limitations of the special effect processing process in the existing technology are solved, and a simplified special effect processing process and a better user experience are achieved.

WO2025092671A1PCT designated stage expired Publication Date: 2025-05-08BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/127844
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-10-28
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

In the prior art, the special effects processing process is complicated and time-consuming. Especially for users with limited operations such as visual impairment, it is difficult to collect original images and select special effects props, which leads to difficulties for users to perform the special effects processing process and affects the user experience.

Method used

By obtaining special effect description data, including content description data such as text and audio, in response to user triggering operations, the corresponding special effect effect data, including special effect images and image sound effects, realize special effect processing based on special effect description data.

Benefits of technology

The special effects processing process is simplified, and the special effects processing can be achieved through simple interactive operations, which is more applicable and improves the special effects editing experience, especially for users with limited operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024127844_08052025_PF_FP_ABST
    Figure CN2024127844_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a special-effects processing method and apparatus, and an electronic device and a storage medium. The method comprises: in response to a first trigger operation, acquiring special-effects description data, wherein the special-effects description data at least comprises content description data, and the content description data comprises content description text and / or content description audio; and in response to a second trigger operation inputted for the special-effects description data, displaying at least one kind of special-effects effect data corresponding to the special-effects description data, wherein the special-effects effect data comprises a special-effects image corresponding to the special-effects description data and an image sound effect corresponding to the special-effects image. The technical solution of the embodiments of the present disclosure realizes the effects of generating a corresponding special-effects effect on the basis of special-effects description data and visually displaying the special-effects effect, and can be realized by means of a simple interaction operation, has a wider applicability, and improves the special-effects editing experience.
Need to check novelty before this filing date? Find Prior Art

Description

Special effects processing method, device, electronic device and storage medium

[0001] This application claims priority to the Chinese invention patent application entitled “Special effects processing method, device, electronic device and storage medium” filed on October 31, 2023, with application number 202311436350.8. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present disclosure relate to the field of special effect processing technology, and in particular to a special effect processing method, device, electronic device, and storage medium. Background Art

[0003] In the context of image processing or video production, the use of special effects props is very popular among users. Users can use the selected special effects props to process images or videos, so that the resulting special effects image presents the special effects corresponding to the special effects props.

[0004] In related technologies, when performing special effects processing, the original image to be processed is usually obtained. Then, based on the triggered and selected special effects props or special effects elements, the original image can be edited and processed so that the special effects image obtained after processing presents the corresponding special effects effect. This special effects processing method requires multiple interactive operations to complete, and the special effects processing process is relatively complex and laborious. In particular, for users with limited operations such as visual impairments, it may be difficult to capture the original image and select special effects props. As a result, the user's execution of special effects processing has certain difficulties, which makes the application of special effects props have certain user limitations, thereby affecting the user experience of special effects props.

[0005] Summary of the Invention

[0006] The present disclosure provides a special effect processing method, device, electronic device and storage medium to achieve the effect of generating corresponding special effect effects based on special effect description data and performing visual display.

[0007] In a first aspect, an embodiment of the present disclosure provides a special effects processing method, the method comprising:

[0008] In response to a first trigger operation, obtaining special effect description data, wherein the special effect description data at least includes content description data, and the content description data includes content description text and / or content description audio;

[0009] In response to a second trigger operation input for the special effect description data, at least one special effect data corresponding to the special effect description data is displayed, wherein the special effect data includes a special effect image corresponding to the special effect description data and an image sound effect corresponding to the special effect image.

[0010] In a second aspect, an embodiment of the present disclosure further provides a special effects processing device, the device comprising:

[0011] a special effect description acquisition module, configured to acquire special effect description data in response to a first trigger operation, wherein the special effect description data at least includes content description data, and the content description data includes content description text and / or content description audio;

[0012] A special effect display module is used to display at least one special effect data corresponding to the special effect description data in response to a second trigger operation input for the special effect description data, wherein the special effect data includes a special effect image corresponding to the special effect description data and an image sound effect corresponding to the special effect image.

[0013] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:

[0014] one or more processors;

[0015] a storage device for storing one or more programs,

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the special effects processing method as described in any one of the embodiments of the present disclosure.

[0017] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the special effects processing method as described in any one of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0019] FIG1 is a schematic flow chart of a special effects processing method provided by an embodiment of the present disclosure;

[0020] FIG2 is a flow chart of another special effects processing method provided by an embodiment of the present disclosure;

[0021] FIG3 is a flow chart of another special effects processing method provided by an embodiment of the present disclosure;

[0022] FIG4 is a flow chart of another special effects processing method provided by an embodiment of the present disclosure;

[0023] FIG5 is a schematic diagram of an interface of a special effects processing method provided by an embodiment of the present disclosure;

[0024] FIG6 is a schematic structural diagram of a special effects processing device provided by an embodiment of the present disclosure;

[0025] FIG7 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0027] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0028] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0029] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0030] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0031] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0032] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0033] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0034] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0035] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0036] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0037] Before introducing this technical solution, an example application scenario can be first described. This technical solution can be applied to any special effects processing scenario. For example, when performing special effects processing, the original image to be processed is typically acquired. Then, the original image can be edited based on the triggered and selected special effects props or special effects elements. During the editing process, the special effects editing results need to be confirmed through a special effects preview image so that the final special effects image can present the corresponding special effects. This special effects processing method requires the user to perform multiple interactive operations before completion, making the entire special effects processing process relatively complex. Furthermore, for users with visual impairments (such as the blind, color-weak, or color-blind), it may be impossible to acquire the original image and select special effects props, resulting in the user being unable to perform the special effects processing process, which poses certain user limitations. In this case, based on the technical solution of the embodiments of the present disclosure, when performing special effects processing, special effects description data can be acquired. This description data can be data that visually describes the special effects generated after the special effects processing. Furthermore, the special effects description data can be special effects description text and / or special effects description audio. Furthermore, corresponding special effect data can be generated based on the special effect description data, and the generated special effect data can be displayed on the display interface. This achieves the effect of generating corresponding special effects based on special effect description data and visually displaying them, and this can be achieved through simple interactive operations, with wider applicability and an enhanced special effect editing experience.

[0038] Before introducing the present technical solution, it should be noted that the device for executing the special effects processing method provided by the embodiment of the present disclosure can be integrated into application software that supports the special effects processing function, and the software can be installed in an electronic device. Optionally, the electronic device can be a mobile terminal or a PC. The application software can be a type of software for image / video / text processing. The specific application software will not be described here one by one, as long as image / video / text processing can be achieved. It can also be a specially developed application program, which is integrated into the software for achieving special effects processing, or integrated into the corresponding page, and the user can achieve special effects processing through the page integrated on the PC.

[0039] Figure 1 is a flow chart of a special effects processing method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to situations where corresponding special effects are generated and displayed based on the description data of the special effects. The method can be executed by a special effects processing device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, a PC or a server, etc.

[0040] As shown in FIG1 , the method of this embodiment may specifically include:

[0041] S110 : In response to a first trigger operation, obtain special effect description data.

[0042] In the disclosed embodiments, the first trigger operation can be understood as an operation for acquiring special effect description data, i.e., an operation that, when triggered, can acquire the special effect description data to be processed. In the disclosed embodiments, a control for triggering data acquisition can be pre-set in application software or applications that support special effect video processing. When a user triggers this control, it is determined that the first trigger operation has been detected, and a response to the first trigger operation is performed. Thus, the special effect description data is acquired.

[0043] In the embodiment of the present disclosure, special effect description data can be understood as data that provides a visual description of the special effect generated after special effect processing, or as description data that matches the user's special effect processing requirements. In the embodiment of the present disclosure, the special effect description data includes at least content description data. The content description data can be understood as data that provides a visual description of the content displayed by the special effect generated after special effect processing. In the embodiment of the present disclosure, the content description data can include forward description data and / or reverse description data, that is, the content description data can include only forward description data, only reverse description data, or both forward description data and direction description data, which is not specifically limited in the embodiment of the present disclosure. Forward description data can be data used to indicate image content that is desired to be included in the special effect image. In other words, the description content corresponding to the forward description data can be image content that is desired to appear in the special effect image. Reverse description data can be data used to suppress image content that appears in the special effect image. In other words, the description content corresponding to the reverse description data can be image content that is desired not to be displayed in the special effect image.

[0044] It should be noted that the content description data may include description data in various forms. Optionally, the content description data includes content description text and / or content description audio. The content description text may be understood as content description data presented on the interface in text form. It should be noted that the content description text may be a text consisting of one or more sentences, or a text consisting of one or more keywords, etc., and the embodiments of the present disclosure do not specifically limit this. The content description audio may be understood as content description data presented on the interface in audio form, and the data may store audio information of the user's description of the special effects content. The content description audio may be an audio consisting of one or more sentences, or a text consisting of one or more keywords, etc., and the embodiments of the present disclosure do not specifically limit this.

[0045] It should be noted that the special effect description data may also include other descriptive data in addition to the content description data. Optionally, it may include at least one desired special effect style of the special effect data, the desired number of special effect data to be generated, and a correlation index. In actual applications, the other descriptive data in the special effect description data may be system default data or user-defined data, and this is not specifically limited in the embodiments of the present disclosure.

[0046] In the disclosed embodiments, a data collection control can be pre-set. Furthermore, upon detecting a trigger operation on the data collection control, corresponding special effect description data can be collected. Optionally, the data collection control can include a content collection control, a special effect style editing control, a generated quantity editing control, and a correlation index editing control. Specifically, for different forms of content description data, the content collection control can also include a text collection control and / or an audio collection control.

[0047] In actual applications, corresponding text can be input based on the edit box displayed in the terminal device display interface. Furthermore, when the trigger operation of the text collection control is detected, the text input in the edit box can be used as content description text, and the content description text can be used as content description data in the special effect description data, and the content description data can also be displayed in the corresponding edit box; or, when the user triggers the audio collection control, the user's audio information can be collected. Furthermore, the collected audio information can be used as content description audio, and the content description audio can be used as content description data in the special effect description data. At the same time, the content description audio can also be analyzed and processed, thereby identifying the text corresponding to the content description audio, and the identified text can be displayed in the corresponding edit box.

[0048] In actual applications, when the display interface also includes a special effect style editing control, a generation quantity editing control, and a correlation index editing control, a trigger operation can be input for any of the above controls to obtain corresponding special effect description data. Exemplarily, when a trigger operation for a special effect style editing control is detected, at least one special effect style to be selected can be displayed in the display interface. Furthermore, when a selection trigger operation for any special effect style is detected, the special effect style corresponding to the selection trigger operation can be determined, and the special effect style can be used as the expected special effect style of the special effect effect data in the special effect description data. When an edit trigger operation for the generation quantity editing control is detected, the generation quantity displayed on the generation quantity editing control can be used as the expected generation quantity of the special effect effect data in the special effect description data. When an edit trigger operation for the correlation index editing control is detected, the correlation index displayed on the correlation index editing control can be used as the correlation index in the special effect description data.

[0049] S120 : In response to a second trigger operation inputted for special effect description data, display at least one special effect data corresponding to the special effect description data.

[0050] In the disclosed embodiments, the second trigger operation can be understood as a special effect processing trigger operation, i.e., an operation that, when triggered, generates and displays corresponding special effect data based on the acquired special effect description data. In the disclosed embodiments, a special effect processing control can be pre-set on the display interface. Then, when a trigger operation is detected for the control, it can be determined that a second trigger operation has been detected, and in response to the second trigger operation, at least one special effect data corresponding to the special effect description data can be displayed.

[0051] In the disclosed embodiments, special effect data may be understood as data representing a special effect corresponding to the special effect description data. The special effect data may include a special effect image corresponding to the special effect description data and an image and sound effect corresponding to the special effect image. The at least one special effect data may be understood as at least one special effect image and an image and sound effect corresponding to each special effect image.

[0052] In an embodiment of the present disclosure, a special effects image may be an image presenting a special effects effect, and the presented special effects effect matches the special effects effect represented by the special effects description data. The image sound effect may be a special effects sound effect that matches the special effects style presented by the special effects image. For example, if the special effects style corresponding to the special effects image is a comic style, the image sound effect may be a special effects sound effect corresponding to the comic style. Alternatively, the image sound effect may also be a special effects sound effect that matches the image content presented by the special effects image. For example, if the special effects image includes a pet, the image sound effect may be the cry of the pet. Alternatively, the image sound effect may also be text audio that describes the image content of the special effects image, etc.

[0053] In practical applications, after obtaining the special effect description data, a second trigger operation can be input for the special effect description data. Furthermore, in response to the second trigger operation, the special effect description data can be processed to determine the special effect description text corresponding to the special effect description data. Furthermore, at least one special effect image can be generated based on the special effect description text. Subsequently, the image and sound effects corresponding to each special effect image can be determined. Consequently, special effect data corresponding to the special effect description data can be generated based on each special effect image and the corresponding image and sound effects, and the generated special effect data can be displayed on a display interface.

[0054] It should be noted that the special effect description data may also include at least one desired special effect style, a desired number of special effect data to be generated, and at least one of a correlation index. In practical applications, the special effect description data including different descriptive data may have different special effect data generation processes. The following describes the special effect data generation processes for different situations.

[0055] Optionally, the special effect description data also includes at least one desired special effect style for the special effect data. The desired special effect style includes an image style and / or a sound effect style. Accordingly, displaying at least one special effect data corresponding to the special effect description data includes: determining a to-be-used image style and sound effect style based on the desired special effect style; determining a special effect image corresponding to the content description data based on the to-be-used image style; determining an image and sound effect corresponding to the special effect image based on the to-be-used sound effect style; generating special effect data based on the special effect image and the image and sound effect; and displaying the special effect data.

[0056] In the disclosed embodiment, the desired special effect style can be understood as the special effect style presented by the desired special effect data. The image style can be understood as the special effect style presented by the special effect image. Optionally, the image style can include comic style, cyberpunk style, ink painting style, and realistic style. The sound effect style can be understood as the special effect style presented by the image sound effect. Optionally, the sound effect style can include comic sound effects, broadcast sound effects, and electronic sound effects. The image style to be used can be the image style based on which the special effect image is determined. The sound effect style to be used can be the sound effect style based on which the image sound effect is determined.

[0057] In practical applications, the desired special effects style includes an image style and / or a sound effect style. That is, the desired special effects style may include an image style but not a sound effect style; or may include a sound effect style but not an image style; or may include both an image style and a sound effect style. Regardless of whether the desired special effects style includes both image and sound effect styles, or only includes one of the image and sound effect styles, it is necessary to determine the image and sound effect styles to be used based on the desired special effects style. Therefore, for different desired special effects styles, the method for determining the corresponding image style and sound effect style to be used may be different. The following will explain the method for determining the image and sound effect styles to be used based on the desired special effects style in different situations.

[0058] Case 1: When the desired special effect style includes an image style but does not include a sound effect style, the included image style is used as the image style to be used, and the sound effect style to be used is determined based on the image style to be used.

[0059] In the disclosed embodiments, multiple image styles and multiple sound effect styles can be pre-determined. Furthermore, a mapping relationship between the image styles and the sound effect styles can be established. The image styles and the sound effect styles can then be stored in corresponding material libraries, and the mapping relationship can also be stored in the terminal device. This allows the corresponding sound effect style or image style to be determined based on the mapping relationship when the image style or sound effect style is determined.

[0060] In actual applications, if the desired special effects style includes an image style but does not include an audio style, the image style included in the desired special effects style can be used as the image style to be used. Furthermore, the audio style corresponding to the image style to be used can be determined based on a pre-set mapping relationship, and the determined audio style can be used as the audio style to be used.

[0061] Second case: when the desired special effect style includes a sound effect style but does not include an image style, the included sound effect style is used as the sound effect style to be used, and the image style to be used is determined based on the sound effect style to be used.

[0062] In actual applications, if the desired special effects style includes a sound effect style but does not include an image style, the sound effect style included in the desired special effects style can be used as the sound effect style to be used. Furthermore, the image style corresponding to the sound effect style to be used can be determined based on a pre-set mapping relationship, and the determined image style can be used as the image style to be used.

[0063] The third case: when the desired special effect style includes an image style and a sound effect style, the included image style is used as the image style to be used, and the included sound effect style is used as the sound effect style to be used.

[0064] In actual applications, when the desired special effects style includes both image style and sound effect style, the image style included in the desired special effects style can be used as the image style to be used, and the sound effect style included in the desired special effects style can be used as the sound effect style to be used.

[0065] It should be noted that the image style and sound effect style included in the desired special effect style may be the special effect style determined after the user's selection trigger operation during the special effect description data acquisition stage. Alternatively, the image style and sound effect style included in the desired special effect style may be the special effect style set by the system by default during the special effect description data acquisition stage.

[0066] In actual applications, after determining the image style and sound effect style to be used based on the desired special effect style, the image style to be used and the content description data can be processed according to a preset image stylization method. Furthermore, the special effect image corresponding to the content description data can be determined. The preset image stylization method can be any stylization processing method. Optionally, the special effect image can be determined based on a pre-trained image stylization model. Exemplarily, multiple image stylization models can be deployed in the terminal device in advance, each image stylization model corresponds to an image style, and a style identifier corresponding to the corresponding image style is configured for each image stylization model. After determining the image style to be used, the corresponding image stylization model can be called according to the style identifier corresponding to the image style to be used. Furthermore, the content description data can be processed according to the image stylization model to obtain the corresponding special effect image.

[0067] Furthermore, the sound effect style and special effect image to be used can be processed according to a preset sound effect stylization method. Furthermore, the image sound effect corresponding to the special effect image can be determined. The special effect image and the image sound effect can then be associated to obtain special effect data, which can then be displayed on a display interface.

[0068] Optionally, the special effect description data further includes an expected number of special effect data to be generated. Accordingly, displaying at least one special effect data corresponding to the special effect description data includes: generating the special effect data corresponding to the special effect description data based on the expected number of special effect data to be generated, and displaying the special effect data.

[0069] In the disclosed embodiment, the expected number of generated images can be understood as the number of special effect data that are expected to be generated in the end. For example, assuming that the expected number of generated images is 4, the special effect data that are finally generated can be four special effect images and the image and sound effects corresponding to each special effect image.

[0070] In actual applications, when the special effects description data includes the expected number of special effects data to be generated, the special effects description data can be processed based on the expected number of generation so that the number of special effects data finally generated matches the expected number of generation, and the generated special effects data can be displayed on the display interface.

[0071] It should be noted that when the expected number of generated images is a value greater than 1, that is, when multiple special effect images are ultimately generated, the image content of each special effect image can match at least part of the description data in the content description data; the image style corresponding to each special effect image can match the image style to be used.

[0072] Optionally, the special effect description data further includes a correlation index. Accordingly, displaying at least one special effect data corresponding to the special effect description data includes: generating special effect data corresponding to the special effect description data based on each correlation index, and displaying the special effect data.

[0073] In an embodiment of the present disclosure, a correlation index is used to indicate the degree of correlation between special effect data and content description data. The correlation index may be represented by a numerical value, that is, the degree of correlation between special effect data and content description data may be indicated by the size of the numerical value. Exemplarily, the correlation index may be represented by the interval [0,1], wherein a correlation index of 0 may indicate that there is no correlation between the special effect data and the content description data; a correlation index of 1 may indicate that the degree of correlation between the special effect data and the content description data is strongly correlated. In an embodiment of the present disclosure, the correlation index may refer to the degree of correlation between the image content of the special effect image in the special effect data and the content description data in the special effect description data, or may refer to the degree of correlation between the image style of the special effect image and / or the sound effect style of the image sound effect in the special effect data and the expected special effect style in the special effect description data, etc.

[0074] In practical applications, if the acquired special effect description data includes a correlation index, the special effect description data can be processed based on the correlation index so that the degree of correlation between the finally generated special effect data and the special effect description data matches the correlation index. Thus, special effect data can be obtained and displayed on a display interface.

[0075] The technical solution of the disclosed embodiment obtains special effect description data in response to a first trigger operation, thereby providing a data basis for subsequent special effect processing. Furthermore, in response to a second trigger operation input for the special effect description data, at least one special effect data corresponding to the special effect description data is displayed, thereby solving the problems in related technologies such as the special effect processing process being relatively complicated, labor-intensive, and having certain user limitations. It realizes the effect of generating corresponding special effect effects based on the special effect description data and visually displaying them, and can be achieved through simple interactive operations, thus having wider applicability and improving the special effect editing experience.

[0076] FIG2 is a flow chart of another special effects processing method provided by an embodiment of the present disclosure. The technical solution of this embodiment is based on the above-described embodiment. After obtaining the special effects description data, a text feature vector corresponding to the special effects description data can be determined. Furthermore, a special effects image is generated based on the text feature vector, and an image and sound effect corresponding to the special effects image is determined. Thus, special effects data can be generated and displayed based on the special effects image and the image and sound effect. Technical features that are identical or similar to those of the above-described embodiments are not repeated here.

[0077] As shown in FIG2 , the method of this embodiment may specifically include:

[0078] S210: In response to a first trigger operation, obtain special effect description data.

[0079] S220 : In response to a second trigger operation inputted for the special effect description data, determine a text to be processed corresponding to the special effect description data, and determine a first text feature vector corresponding to the text to be processed.

[0080] In the disclosed embodiments, the special effect description data includes at least content description data, and other descriptive data included in the special effect description data, such as the desired special effect style, the desired number of generated effects, and the correlation index, is data represented in other forms. Therefore, the to-be-processed text corresponding to the special effect description data can be the to-be-processed text corresponding to the content description data. The to-be-processed text can be understood as content description data represented in textual form.

[0081] In the embodiment of the present disclosure, the content description data includes forward description data and reverse description data. Accordingly, the text to be processed may include forward description text corresponding to the forward description data and reverse description text corresponding to the reverse description data. It should be noted that, when the content description data includes content description text, the content description text can be used as at least part of the text in the text to be processed; when the content description data includes content description audio, the content description audio can be recognized and processed, and the recognized text can be used as at least part of the text in the text to be processed. The first text feature vector can be a vector obtained after feature extraction of the text to be processed, which is used to characterize its feature information.

[0082] In practical applications, upon detecting a second triggering operation in response to input of special effect description data, the second triggering operation can be responded to, and the content description data in the special effect description data can be determined. Furthermore, the content description data can be processed to convert it into text data, and the resulting text can be used as the text to be processed. Furthermore, feature extraction can be performed on the text to be processed, and the resulting feature information can be used as a first text vector corresponding to the text to be processed.

[0083] It should be noted that the first text feature vector corresponds to the text to be processed, which in turn corresponds to the content description data. If the content description data includes both forward description data and reverse description data, the first text feature vector may include both the forward text vector and the reverse text vector; if the text to be processed includes both forward description data and reverse description data, the first text feature vector may include both the forward text vector and the reverse text vector; if the text to be processed includes both reverse description data and forward description data, the first text feature vector may include both the reverse text vector and the forward text vector.

[0084] Optionally, when the content description data includes forward description data and reverse description data, determining the text to be processed corresponding to the special effect description data, and determining the first text feature vector corresponding to the text to be processed, includes: determining the forward description text corresponding to the forward description data, and determining the forward text vector corresponding to the forward description text; determining the reverse description text corresponding to the reverse description data, and determining the reverse text vector corresponding to the reverse description text; determining the first text feature vector based on the forward text vector and the reverse text vector.

[0085] In practical applications, when content description data is obtained and it is determined that the content description data includes forward description data and reverse description data, the content description data can be divided into forward description data and reverse description data. Further, it is determined whether the forward description data includes forward content description audio. If so, the forward content description audio can be subjected to audio information recognition processing, and the recognized text can be used as the forward description text. At this time, if the forward description data also includes forward content description text, the forward content description text can be directly used as the forward description text. Furthermore, the forward description text corresponding to the forward description data can be obtained. Afterwards, text feature extraction can be performed on the forward description text according to a preset feature extraction method to obtain a forward text vector corresponding to the forward description text.

[0086] Furthermore, it is determined whether the reverse description data includes reverse content description audio. If so, audio information recognition processing can be performed on the reverse content description audio, and the recognized text can be used as the reverse description text. In this case, if the reverse description data also includes reverse content description text, the reverse content description text can be directly used as the reverse description text. Furthermore, the reverse description text corresponding to the reverse description data can be obtained. Afterwards, text feature extraction can be performed on the reverse description text according to a preset feature extraction method to obtain a reverse text vector corresponding to the reverse description text.

[0087] Afterwards, the forward text vector and the reverse text vector can be concatenated to obtain a first text feature vector corresponding to the text to be processed. The advantage of this arrangement is that it ensures that the first text feature vector includes the effects of both the forward text vector and the reverse text vector, thereby improving the matching degree between the final generated special effect data and the special effect description data.

[0088] It should be noted that the preset feature extraction method can be any method. Optionally, it can be based on the text encoder to extract features of the forward description text, and use the features output by the text encoder as the forward text vector.

[0089] In actual application, there are also cases where the content description data only includes forward description data or only includes reverse description data. These two cases are described below.

[0090] Optionally, the content description data includes forward description data and does not include reverse description data. Determining the text to be processed corresponding to the special effect description data and determining a first text feature vector corresponding to the text to be processed includes: determining the forward description text corresponding to the forward description data, determining a forward text vector corresponding to the forward description text, and using the forward text vector as the first text feature vector.

[0091] Optionally, the content description data includes reverse description data but does not include forward description data. Determining the text to be processed corresponding to the special effect description data and determining a first text feature vector corresponding to the text to be processed includes: determining reverse description text corresponding to the reverse description data, determining a reverse text vector corresponding to the reverse description text, and using the reverse text vector as the first text feature vector.

[0092] S230 : Generate at least one special effect image based on the first text feature vector, and determine the image and sound effects corresponding to each special effect image.

[0093] In the embodiment of the present disclosure, after obtaining the first text feature vector, at least one special effect image can be generated based on the first text feature vector, wherein each special effect image can include at least part of the content in the first text vector.

[0094] In practical applications, the first text feature vector can be processed according to a preset image generation method to obtain at least one special effect image corresponding to the first text vector. The preset image generation method can be any image generation method, and optionally, the special effect image can be generated based on an image generation model.

[0095] In an embodiment of the present disclosure, optionally, at least one special effect image is generated based on the first text feature vector, including: inputting the first text feature vector and the first Gaussian noise vector into at least one first image generation model, and obtaining the special effect image output by each first image generation model.

[0096] In the embodiment of the present disclosure, the first Gaussian noise vector can be any randomly determined Gaussian noise vector. The first image generation model can be understood as a diffusion model that takes a text vector and a Gaussian noise vector as input objects to process the text vector and the Gaussian noise vector. At least one first image generation model can be an image generation model for generating images with different special effects. Exemplarily, the first image generation model can be a stable diffusion model. Stable Diffusion is an image generation model based on a diffusion process that can generate high-quality, high-resolution images. It is the latest diffusion model. The core idea of ​​Stable Diffusion is to gradually approach the real image by continuously adjusting the implicit representation of the image. Its specific implementation method is that it can include a forward diffusion process and a reverse denoising process. Exemplarily, in the forward process, given a data point x0 sampled from a real data distribution (such as the data distribution to which a natural image belongs), a small amount of Gaussian noise is added to the sample in T steps to generate a series of noise samples: x1, x2,…, x T , so that x T Close to Gaussian distribution. In the reverse denoising process, sampling from Gaussian noise, the above process is reversed, and from x T Initially, the neural network is trained to continuously remove Gaussian noise until the sample is restored to the natural image x0.

[0097] In an embodiment of the present disclosure, the first image generation model is obtained by training a pre-established first diffusion model based on a first sample feature vector corresponding to the first sample description text and a desired effect image corresponding to the first sample description text.

[0098] It should be noted that before applying the first image generation model provided in the embodiments of this disclosure, the pre-established first diffusion model must first be trained. Before training the model, multiple training samples can be constructed to train the model based on these training samples. To improve the accuracy of the first image generation model, as many and diverse training samples as possible can be constructed.

[0099] Optionally, the training process of the first diffusion model may be as follows: obtaining multiple training samples, wherein the training samples may include a first sample feature vector corresponding to the first sample description text and an expected effect image corresponding to the first sample description text; for each training sample, inputting the first sample feature vector in the training sample into the first diffusion model, so as to gradually add Gaussian noise to the first sample feature vector based on the forward diffusion module in the first diffusion model, so that the first sample feature vector after adding noise is close to a Gaussian distribution. Furthermore, the first sample feature vector after adding noise may be gradually denoised based on the reverse denoising module in the first diffusion model to obtain an actual output image; determining a loss value based on the actual output image and the expected effect image in the training sample; modifying the model parameters in the first diffusion model based on the loss value, and taking the convergence of the loss function in the first diffusion model as the training target to obtain a first image generation model.

[0100] It should be noted that whether it is one first image generation model or multiple first image generation models, they can all be obtained by training based on the above process, and the embodiments of the present disclosure will not be described in detail here.

[0101] In practical applications, when the first text feature vector is obtained, the first Gaussian noise vector can be randomly determined. Furthermore, the first text feature vector and the first Gaussian noise vector can be input into at least one first image generation model. Furthermore, for each first image generation model, the first text feature vector and the first Gaussian noise vector can be processed based on the first image generation model to obtain the special effects image output by the first image generation model. Furthermore, after each first image generation model outputs a special effects image, the special effects image output by each first image generation model can be obtained, and the obtained special effects image can be used as the special effects image corresponding to the special effects description data. The advantage of this setting is that the special effects image is generated based on the diffusion model, which improves the generation quality and generation efficiency of the special effects image. Furthermore, the special effects image finally obtained can meet the user's special effects processing needs.

[0102] Furthermore, after obtaining at least one special effect image, the image and sound effects corresponding to each special effect image can be determined separately. It should be noted that for users with visual impairments (such as blind people, color-weak people, or color-blind people), they may not be able to see the special effect images displayed in the display interface. Therefore, in the embodiment of the present disclosure, in addition to the image and sound effects determined based on the sound effect style, the image and sound effects may also include text audio corresponding to the image description text, so that users with visual impairments can understand the image content included in the special effect image through the text audio.

[0103] In practical applications, when an image sound effect includes text audio corresponding to an image description text, when determining the image sound effect corresponding to the special effect image, the image description text corresponding to the special effect image can be first determined. Furthermore, the image description text can be processed to obtain the text audio corresponding to the image description text.

[0104] Optionally, the image sound effects corresponding to each special effect image are determined respectively, including: for each special effect image, determining the image description text corresponding to the special effect image, and generating text audio corresponding to the image description text.

[0105] In the disclosed embodiments, the image description text can be understood as text that vividly describes the image content included in the special effect image. For example, the image description text corresponding to the special effect image can be "A little boy in black sits in the shade of a tree. In the distance, a flying bird appears particularly dazzling in the afternoon sun."

[0106] In the embodiment of the present disclosure, the text audio can be understood as a multimedia data stream storing sound content. At the same time, the sound content stored in this data stream is sound information corresponding to the image description text.

[0107] It should be noted that after obtaining the image description text, in addition to being used as the basis for generating text audio, the obtained image description text can also be displayed in the display area corresponding to the special effect image, thereby achieving the effect of corresponding display of the special effect image and the image description text, thereby enhancing the richness and diversity of the special effect data.

[0108] Optionally, determining the image description text corresponding to the special effect image includes: determining image feature data corresponding to the special effect image, converting the image feature data into image text features, generating image description text based on the image text features, and displaying the image description text corresponding to the special effect image.

[0109] In the embodiments of the present disclosure, the image feature data can be understood as data obtained by extracting features from the special effect image and used to represent its feature information. The image text feature can be a text feature corresponding to the image feature data.

[0110] In practical applications, features of the special effect image can be extracted to obtain image feature data corresponding to the special effect image. Furthermore, the image feature data can be converted into text information to obtain image text features corresponding to the image feature data. Thereafter, the image text features can be processed, and the text obtained after processing can be used as image description text corresponding to the special effect image. Furthermore, the image description text can be displayed at a preset position in the special effect image display area. Thus, the image description text can be displayed corresponding to the special effect image. The preset position can be any position in the special effect image display area, and optionally, it can be the bottom of the display area.

[0111] It should be noted that the process of determining the image description text corresponding to the special effect image can be implemented according to a preset image-to-text method. Among them, the preset image-to-text method can be any method for converting an image into text, and optionally, it can be implemented based on an image-to-text model. Exemplarily, after obtaining the special effect image, the special effect image can be input into an image encoder to perform image feature extraction on the special effect image based on the image encoder to obtain image features corresponding to the special effect image. Further, the image features can be input into a query module to convert the image features into text features based on the query module, and output the text features. Afterwards, the text features can be input into a large language model (LLM) to perform text information recognition on the text features based on the large language model to obtain text output, which can be used as the image description text corresponding to the special effect image.

[0112] Furthermore, the image description text can be processed according to a preset text-to-sound method to obtain the text audio corresponding to the image description text. The preset text-to-sound method can be any method for converting text into audio, and can optionally be based on a diffusion model.

[0113] It should be noted that audio waves in the air are the vibrations of air molecules generated by sound propagating through the air. An audio waveform represents the time-varying vibrational displacement, with its intensity represented by amplitude. Essentially, an audio waveform is a mixture of frequencies, so audio analysis typically begins by converting the raw waveform in the time domain into a spectrogram in the time-frequency domain. Specifically, the audio is segmented into windows, and then the short-time Fourier transform (STFT) is used to calculate the amplitude of each frequency. The STFT is repeated for each window along the time axis, resulting in a two-dimensional complex-valued image, where the X-axis represents frames (time) and the Y-axis represents frequency. The complex values ​​in this complex-valued image can be further converted to absolute values ​​(or amplitudes) and phases. Due to the wide frequency range of sound, the Mel-scale is often used to convert the spectrogram (amplitude map) into a Mel-spectrogram. Mel-spectrograms are widely used in many audio-related tasks and can be viewed as single-channel images, where the height of the image represents the frequency domain and the width represents the time domain. Based on this, in the technical solution provided in the embodiment of the present disclosure, the audio waveform can be converted into a mel-spectrogram transform, and then the mel-spectrogram can be synthesized in the manner of a synthetic image using a diffusion model. Finally, the mel-spectrogram can be converted back into an audio waveform using a neural network vocoder to realize the sound synthesis.

[0114] Optionally, generating text audio corresponding to the image description text includes: determining a second text feature vector corresponding to the image description text, inputting the second text feature vector and a second Gaussian noise vector into a second image generation model to obtain a target mel-spectrogram corresponding to the image description text; converting the target mel-spectrogram into audio data, and using the converted audio data as the text audio corresponding to the image description text.

[0115] In an embodiment of the present disclosure, the second text feature vector may be a vector obtained after feature extraction of the image description text, used to characterize its feature information. The second Gaussian noise vector may be an arbitrary Gaussian noise vector determined randomly. The second image generation model may be understood as a diffusion model that takes the text vector and the Gaussian noise vector as input objects to process the text vector and the Gaussian noise vector. The second image generation model is obtained by training a pre-established second diffusion model based on the second sample feature vector of the second sample description text and the expected mel-spectrogram corresponding to the second sample description text. The second sample description text may be any content description text. The second sample feature vector may be a vector obtained after feature extraction of the second sample description text, used to characterize its feature information. The expected mel-spectrogram may be a true mel-spectrogram corresponding to the second sample description text.

[0116] It should be noted that before applying the second image generation model provided in the embodiments of this disclosure, the pre-established second diffusion model must first be trained. Before training the model, multiple training samples can be constructed to train the model based on these training samples. To improve the accuracy of the second image generation model, as many and diverse training samples as possible can be constructed.

[0117] Optionally, the training process of the second diffusion model may be as follows: obtaining multiple training samples, wherein the training samples may include a second sample feature vector corresponding to the second sample description text and an expected mel-spectrogram corresponding to the second sample description text; for each training sample, inputting the second sample feature vector in the training sample into the second diffusion model, and gradually adding Gaussian noise to the second sample feature vector based on the forward diffusion module in the second diffusion model, so that the second sample feature vector after adding noise is close to a Gaussian distribution. Furthermore, the second sample feature vector after adding noise can be gradually denoised based on the reverse denoising module in the second diffusion model to obtain an actual mel-spectrogram; determining a loss value based on the actual mel-spectrogram and the expected mel-spectrogram in the training sample; modifying the model parameters in the second diffusion model based on the loss value, and taking the convergence of the loss function in the second diffusion model as the training goal to obtain a second image generation model.

[0118] In practical applications, after obtaining the image description text, text feature extraction can be performed on the image description text according to a preset feature extraction method, and the extracted feature information can be used as a second text feature vector. The preset feature extraction method can be any method, and optionally, it can be implemented based on a neural network model (such as a text encoder). Exemplarily, the image description text is input into a text encoder to perform feature extraction on the image description text based on the text encoder, and the features obtained after the extraction can be used as a second text feature vector corresponding to the image description text.

[0119] Furthermore, a second Gaussian noise vector can be randomly determined, and the second text feature vector and the second Gaussian noise vector can be input into a second image generation model to process the second text feature vector and the second Gaussian noise vector based on the second image generation model. The image output by the second image generation model can be used as a target mel-spectrogram corresponding to the image description text.

[0120] It should be noted that since the corresponding audio information cannot be played through the mel-spectrogram, it is also necessary to convert the target mel-spectrogram into an audio waveform to determine the corresponding text audio based on the audio waveform. In addition, since the target mel-spectrogram is generated based on the second image generation model, the conversion of the target mel-spectrogram into audio data can be implemented based on a neural network model (e.g., a vocoder).

[0121] In practical applications, after obtaining the target mel-spectrogram, the target mel-spectrogram can be input into a vocoder. The vocoder can then process the target mel-spectrogram to convert it into audio data. The converted audio data can then be used as text audio corresponding to the image description text.

[0122] S240 , generating special effect data corresponding to the special effect description data based on the special effect image and the image and sound effects corresponding to the special effect image, and displaying the special effect data.

[0123] In practical applications, after obtaining at least one special effect image and the image and sound effects corresponding to each special effect image, the special effect image and the corresponding image and sound effects can be combined for each special effect image to obtain special effect data corresponding to the special effect description data. Furthermore, after obtaining at least one special effect data, each special effect data can be displayed on a display interface.

[0124] It should be noted that when displaying special effect data, the special effect images in the special effect data can be directly displayed on the display interface. For image and sound effects in the special effect data, the image and sound effects corresponding to the special effect images can be played directly while the special effect images are displayed; or, when a play trigger operation for the image and sound effects is detected, the corresponding image and sound effects can be played in response to the play trigger operation.

[0125] The technical solution of the embodiment of the present disclosure obtains special effect description data in response to a first trigger operation, and further, in response to a second trigger operation input for the special effect description data, determines the text to be processed corresponding to the special effect description data, and determines the first text feature vector corresponding to the text to be processed. Thereafter, at least one special effect image is generated based on the first text feature vector, and the image sound effect corresponding to each special effect image is determined respectively. Finally, special effect effect data corresponding to the special effect description data is generated based on the special effect image and the image sound effect corresponding to the special effect image, and the special effect effect data is displayed, thereby realizing the effects of generating special effect images based on text and generating special effect sound effects based on special effect images, enhancing the diversity and richness of the special effect editing process, and improving the flexibility of the special effect editing process.

[0126] Figure 3 is a flow chart illustrating another special effects processing method provided by an embodiment of the present disclosure. The technical solution of this embodiment, based on the above embodiment, can determine target special effects data in response to a selection trigger operation for the special effects data after displaying at least one special effects data. Furthermore, the target special effects data can be processed in response to a third trigger operation for the target special effects data. Technical features identical or similar to those of the above embodiment are not further described here.

[0127] As shown in FIG3 , the method of this embodiment may specifically include:

[0128] S310: In response to a first trigger operation, obtain special effect description data.

[0129] S320: In response to a second trigger operation inputted for special effect description data, display at least one special effect data corresponding to the special effect description data.

[0130] S330 : In response to a selection triggering operation on the displayed special effect data, generate target special effect data based on the selected special effect data.

[0131] In the embodiment of the present disclosure, the displayed special effect data can be set to a selectable state in advance. Furthermore, when the special effect data is the selected special effect data, the special effect data is subsequently processed.

[0132] In the embodiment of the present disclosure, the selection trigger operation for the displayed special effect data can be any operation acting on the special effect data. Optionally, a click operation on the special effect data. Exemplarily, when it is detected that the user inputs a click operation on the displayed special effect data through an input device or a touch point, the special effect data can be used as the selected special effect data. Furthermore, the selected special effect data can be subsequently processed. Among them, the click operation can be a single click operation or a multiple click operation (such as a double-click operation, etc.). In order to facilitate user operations, the selection trigger operation for the special effect data can also be that when it is detected that the user's pause time on the special effect data based on the input device or touch point reaches a preset time length, the special effect data can be used as the selected special effect data. Furthermore, the special effect data can be subsequently processed.

[0133] In actual applications, after at least one generated special effect data is displayed on a display interface, the user can select from the displayed special effect data. Furthermore, upon detecting a user input trigger operation for a special effect data selection via an input device or a touch point, the system can respond to the selection trigger operation, determine the selected special effect data, and generate target special effect data based on the selected special effect data.

[0134] It should be noted that the selected special effect data may be one or more types. If the selected special effect data is one type, the selected special effect data may be used as the target special effect data. If the selected special effect data is multiple types, the selected special effect data may be combined together, and the data obtained by the combination may be used as the target special effect data.

[0135] S340 : In response to a third triggering operation on the target special effect data, the target special effect data is saved to a preset storage space and / or the target special effect data is published to a target display platform.

[0136] In the embodiment of the present disclosure, the third trigger operation can be understood as a special effect data editing trigger operation, that is, an operation of performing subsequent editing on the generated target special effect data after being triggered. Optionally, the third trigger operation may include a data storage trigger operation and / or a data release trigger operation, etc. In the embodiment of the present disclosure, the preset storage space may be a storage space in the application software; or, it may be a storage space in the local terminal; or, it may be a storage space in other external devices, etc. The target display platform can be understood as a platform that can display special effect data. The target display platform can be any platform, optionally, the platform to which the application software containing the special effect props belongs or the platform to which other application software belongs, etc.

[0137] In the embodiment of the present disclosure, a special effect data editing control can be pre-set on the display interface, and the user can input a trigger operation into the special effect data editing control to subsequently edit the generated target special effect data. Optionally, the special effect data editing control can include a data storage control and / or a data publishing control.

[0138] In actual applications, when a trigger operation is detected for a data storage control, the trigger operation can be responded to by popping up a storage setting interface including a storage path setting item on the display interface. The user can set various setting items on the storage setting interface through the trigger operation. Then, when the storage setting trigger operation is detected, the trigger operation is responded to and the target special effect data is stored in the preset storage space.

[0139] In actual applications, if a trigger operation is detected for a data publishing control, the trigger operation can be responded to by popping up a display interface including multiple candidate display platforms on the display interface, allowing the user to select from the multiple candidate display platforms through the trigger operation. Furthermore, if a user selection trigger operation is detected for any candidate display platform, the selected candidate display platform can be selected as the target display platform, and the target special effect data can be published to the target display platform.

[0140] The technical solution of the embodiment of the present disclosure obtains special effect description data in response to a first trigger operation, and further, in response to a second trigger operation input for the special effect description data, displays at least one special effect data corresponding to the special effect description data, and then, in response to a selection trigger operation for the displayed special effect data, generates target special effect data based on the selected special effect data, and in response to a third trigger operation for the target special effect data, saves the target special effect data to a preset storage space and / or publishes the target special effect data to a target display platform, thereby achieving the effect of subsequent processing of the displayed special effect data, meeting the user's demand for personalized editing of the special effect data, and improving the user's experience of using special effect props.

[0141] Figure 4 is a flow chart illustrating another special effects processing method provided by an embodiment of the present disclosure. This embodiment, based on the above embodiment, presents a sound effect playback control corresponding to each image and sound effect when the displayed special effects data includes two or more. Furthermore, the corresponding image and sound effect can be played in response to a control trigger operation acting on the sound effect playback control. Technical features identical or similar to those of the above embodiment are not further detailed here.

[0142] As shown in FIG4 , the method of this embodiment may specifically include:

[0143] S410: In response to a first trigger operation, obtain special effect description data.

[0144] S420: In response to a second trigger operation inputted for special effect description data, display at least one special effect data corresponding to the special effect description data.

[0145] S430: When the special effect data includes two or more image and sound effects, display the sound effect playback controls corresponding to each image and sound effect.

[0146] In the embodiment of the present disclosure, the sound effect playback control can be understood as a control that can play the corresponding sound effect after being triggered. Optionally, the sound effect playback control can include a background sound effect playback control and / or a text audio playback control, etc.

[0147] It should be noted that when the special effect data includes two or more image and sound effects, for example, it may include image background sound effects corresponding to the sound effect style to be used, and text audio corresponding to the image description text. If the special effect data is directly displayed at the same time as it is generated, multiple image and sound effects may be played simultaneously, thereby affecting the playback quality of the image and sound effects and reducing the special effect display quality of the special effect data.

[0148] Based on this, in the technical solution of the embodiment of the present disclosure, after the special effect data is generated, the audio playback control corresponding to each image sound effect in the special effect data can be preset. Furthermore, in the case where the special effect data includes two or more image sound effects, when the special effect data is displayed, the sound effect playback control corresponding to each image sound effect can be displayed separately, so that when a trigger operation for any sound effect playback control is detected, the trigger operation is responded to and the corresponding image sound effect is played. Exemplarily, in the case where the sound effect playback control includes a background sound effect playback control and a text audio playback control, the background sound effect playback control can be displayed at any position on the special effect image, and the text audio playback control can be displayed at any position in the display area corresponding to the image description text.

[0149] S440: In response to a control triggering operation acting on a sound effect playing control, play the image sound effect corresponding to the triggered sound effect playing control.

[0150] In the disclosed embodiments, the control triggering operation may be any operation acting on the sound effect playback control. Optionally, the click operation acting on the sound effect playback control may be, for example, a single click operation or a multiple click operation (e.g., a double-click operation). To facilitate user operation, the control triggering operation acting on the sound effect playback control may also be that when it is detected that the user pauses on the sound effect playback control based on the input device or touch point for a preset duration, the sound effect playback control may be used as the triggered sound effect playback control.

[0151] In actual applications, after displaying the sound effect playback controls corresponding to each image and sound effect, the user can trigger the displayed sound effect playback controls to play the corresponding image and sound effect. Furthermore, upon detecting a control triggering operation on a sound effect playback control, the system can respond to the triggering operation, using the control as the triggered sound effect playback control, and determining the image and sound effect corresponding to the control. This can then be played.

[0152] It should be noted that for each displayed sound effect playback control, when the image sound effect corresponding to the triggered sound effect playback control is playing, the triggered sound effect playback control can be switched to a sound effect pause control and displayed. Furthermore, when a trigger operation for the sound effect pause control is detected, the trigger operation is responded to and the image sound effect corresponding to the triggered sound effect pause control is paused.

[0153] It should also be noted that during the process of playing the image sound effect corresponding to the triggered sound effect playback control, if a control trigger operation acting on other sound effect playback controls is detected, the currently playing image sound effect can be switched to the image sound effect corresponding to the sound effect playback control and played.

[0154] Exemplarily, Figure 5 can be a schematic diagram of the interface of the special effects processing method provided by the embodiment of the present disclosure. As shown in Figure 5, the area on the left side of the interface can be a special effects description data configuration area, and the selection style display area in Figure 5 can be used as the desired special effects style configuration area. For example, when it is detected that the mouse cursor (arrow in the figure) inputs a trigger operation on the "Style 1" control, "Style 1" can be used as the desired special effects style. The prompt word display area in Figure 5 can be used as a configuration area for forward description data, and the user can enter forward description data in the display area through an input device (such as a mouse or microphone). The shielded word display area in Figure 5 can be used as a configuration area for reverse description data, and the user can enter reverse description data in the display area through an input device. The quantity display area in Figure 5 can be used as a configuration area for the desired generated quantity, and the user can configure the desired generated quantity through the slider control in the display area. The correlation display area in Figure 5 can be used as a configuration area for the correlation index, and the user can configure the correlation index through the slider control in the display area.

[0155] Furthermore, when a trigger operation is detected for the "Done" control shown in Figure 5 , the trigger operation is responded to by displaying at least one special effect data corresponding to the special effect description data in the right area of ​​the interface. Continuing with Figure 5 , the right area of ​​the interface can be used as a display area for special effect data. This display area currently displays six types of special effect data, each of which includes a special effect image and a corresponding image sound effect. Furthermore, each type of special effect data includes two image sound effects. This display area may include a special effect image 51, image description text 52, a sound effect playback control 53 for the image background sound effect, and a sound effect playback control 54 for text audio. In actual application, when a trigger operation is detected with the mouse cursor (arrow in the figure) on the sound effect playback control 53, the image background sound effect corresponding to the sound effect playback control 53 is played. When a trigger operation is detected with the mouse cursor (arrow in the figure) on the sound effect playback control 54, the text audio corresponding to the sound effect playback control 54 is played. Furthermore, the display area for special effect data also includes a "Save to Result" control. In actual application, when the mouse cursor (arrow in the figure) is detected to input a trigger operation on the "Save to Result" control, the selected special effect data can be saved to the preset storage space.

[0156] The technical solution of the disclosed embodiment obtains special effect description data in response to a first trigger operation, and further, in response to a second trigger operation inputted for the special effect description data, displays at least one special effect data corresponding to the special effect description data. Thereafter, when the special effect data includes two or more image and sound effects, the sound effect playback control corresponding to each image and sound effect is displayed respectively. In response to a control trigger operation acting on the sound effect playback control, the image and sound effect corresponding to the triggered sound effect playback control is played. This achieves the effect of selectively playing image and sound effects when the special effect data includes two or more image and sound effects. This improves the display flexibility of the special effect data and enhances the special effect display quality of the special effect data.

[0157] The technical solution of the disclosed embodiment obtains special effect description data in response to a first trigger operation, thereby providing a data basis for subsequent special effect processing. Furthermore, in response to a second trigger operation input for the special effect description data, at least one special effect data corresponding to the special effect description data is displayed, thereby solving the problems in related technologies such as the special effect processing process being relatively complicated, labor-intensive, and having certain user limitations. It realizes the effect of generating corresponding special effect effects based on the special effect description data and visually displaying them, and can be achieved through simple interactive operations, thus having wider applicability and improving the special effect editing experience.

[0158] FIG6 is a structural diagram of a special effect processing device provided by an embodiment of the present disclosure. As shown in FIG6 , the device includes: a special effect description acquisition module 610 and a special effect display module 620 .

[0159] Among them, the special effect description acquisition module 610 is used to obtain special effect description data in response to a first trigger operation, wherein the special effect description data at least includes content description data, and the content description data includes content description text and / or content description audio; the special effect display module 620 is used to display at least one special effect data corresponding to the special effect description data in response to a second trigger operation input for the special effect description data, wherein the special effect data includes a special effect image corresponding to the special effect description data and an image sound effect corresponding to the special effect image.

[0160] On the basis of the above-mentioned optional technical solutions, optionally, the special effect display module 620 includes: a text feature vector determination unit, a special effect image generation unit and a special effect data generation unit.

[0161] a text feature vector determining unit, configured to determine a text to be processed corresponding to the special effect description data, and to determine a first text feature vector corresponding to the text to be processed;

[0162] a special effect image generating unit, configured to generate at least one special effect image corresponding to the special effect description data based on the first text feature vector, and respectively determine an image and sound effect corresponding to each special effect image;

[0163] The special effect data generating unit is configured to generate special effect data corresponding to the special effect description data based on the special effect image and the image and sound effects corresponding to the special effect image.

[0164] Based on the above optional technical solutions, optionally, the content description data includes forward description data and reverse description data, and the text to be processed includes forward description text corresponding to the forward description data and reverse description text corresponding to the reverse description data; the forward description data is used to indicate image content that is desired to be included in the special effect image, and the reverse description data is used to suppress image content that appears in the special effect image;

[0165] The text feature vector determination unit includes: a forward text vector determination subunit, a reverse text vector determination subunit and a text feature vector determination subunit.

[0166] A forward text vector determining subunit, configured to determine a forward description text corresponding to the forward description data, and to determine a forward text vector corresponding to the forward description text;

[0167] a reverse text vector determining subunit, configured to determine a reverse description text corresponding to the reverse description data, and determine a reverse text vector corresponding to the reverse description text;

[0168] The text feature vector determining subunit is configured to determine a first text feature vector based on the forward text vector and the reverse text vector.

[0169] On the basis of the above-mentioned optional technical solutions, optionally, the special effect image generation unit includes: a special effect image generation subunit;

[0170] The special effect image generation subunit is used to input the first text feature vector and the first Gaussian noise vector into at least one first image generation model, and obtain the special effect image corresponding to the special effect description data output by each first image generation model, wherein the first image generation model is trained based on a pre-established first diffusion model based on the first sample feature vector corresponding to the first sample description text and the expected effect image corresponding to the first sample description text.

[0171] On the basis of the above-mentioned optional technical solutions, optionally, the special effect image generation unit further includes: an image and sound effect generation subunit;

[0172] The image and sound effect generation subunit is used to determine, for each special effect image, an image description text corresponding to the special effect image, and generate text audio corresponding to the image description text.

[0173] Based on the above-mentioned optional technical solutions, optionally, the image sound effect generation subunit is specifically used to determine a second text feature vector corresponding to the image description text, input the second text feature vector and the second Gaussian noise vector into a second image generation model, and obtain a target mel-spectrogram corresponding to the image description text, wherein the second image generation model is trained on a pre-established second diffusion model based on the second sample feature vector of the second sample description text and the expected mel-spectrogram corresponding to the second sample description text; convert the target mel-spectrogram into audio data, and use the converted audio data as text audio corresponding to the image description text.

[0174] Based on the above-mentioned optional technical solutions, optionally, the image and sound effect generation subunit is specifically used to determine the image feature data corresponding to the special effect image, convert the image feature data into image text features, generate image description text based on the image text features, and display the image description text corresponding to the special effect image.

[0175] Based on the above optional technical solutions, optionally, the special effect description data further includes at least one desired special effect style of the special effect data, and the desired special effect style includes an image style and / or a sound effect style;

[0176] The special effect display module 620 includes: a style determination unit, a special effect image determination unit, an image and sound effect determination unit, and a first special effect generation unit.

[0177] a style determining unit, configured to determine an image style and a sound effect style to be used based on the desired special effect style;

[0178] a special effect image determining unit, configured to determine a special effect image corresponding to the content description data based on the image style to be used;

[0179] an image and sound effect determining unit, configured to determine an image and sound effect corresponding to the special effect image based on the sound effect style to be used;

[0180] The first special effect generation unit is used to generate special effect data based on the special effect image and the image sound effect, and display the special effect data.

[0181] Based on the above optional technical solutions, optionally, the special effect description data further includes the expected number of special effect data to be generated:

[0182] The special effect display module 620 further includes: a second special effect generation unit.

[0183] The second special effect generation unit is configured to generate special effect data corresponding to the content description data based on each of the expected generation quantities, and display the special effect data.

[0184] Based on the above optional technical solutions, optionally, the special effect description data further includes a correlation index, and the correlation index is used to indicate the degree of correlation between the special effect data and the content description data;

[0185] The special effect display module 620 further includes: a third special effect generation unit.

[0186] The third special effect generation unit is used to generate special effect data corresponding to the special effect description data based on each of the correlation indicators, and display the special effect data.

[0187] On the basis of the above-mentioned optional technical solutions, optionally, the device further includes: a target special effect data generating module and a target special effect data storing module.

[0188] a target special effect data generating module, configured to generate target special effect data corresponding to the special effect description data based on the selected special effect data in response to a selection trigger operation on the displayed special effect data after the at least one special effect data corresponding to the special effect description data is displayed;

[0189] The target special effect data saving module is used to save the target special effect data to a preset storage space and / or publish the target special effect data to a target display platform in response to a third trigger operation on the target special effect data.

[0190] On the basis of the above-mentioned optional technical solutions, optionally, the device further includes: a control display module and an image and sound effect playback module.

[0191] a control display module for, after displaying at least one special effect data corresponding to the special effect description data, respectively displaying a sound effect playback control corresponding to each of the image and sound effects when the special effect data includes two or more image and sound effects corresponding to the special effect image;

[0192] The image and sound effect playing module is used to respond to a control triggering operation acting on the sound effect playing control and play the image and sound effect corresponding to the triggered sound effect playing control.

[0193] The technical solution of the disclosed embodiment obtains special effect description data in response to a first trigger operation, thereby providing a data basis for subsequent special effect processing. Furthermore, in response to a second trigger operation input for the special effect description data, at least one special effect data corresponding to the special effect description data is displayed, thereby solving the problems in related technologies such as the special effect processing process being relatively complicated, labor-intensive, and having certain user limitations. It realizes the effect of generating corresponding special effect effects based on the special effect description data and visually displaying them, and can be achieved through simple interactive operations, thus having wider applicability and improving the special effect editing experience.

[0194] The special effects processing device provided by the embodiments of the present disclosure can execute the special effects processing method provided by any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution method.

[0195] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0196] FIG7 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Referring to FIG7 , a schematic diagram of the structure of an electronic device (such as a terminal device or server in FIG7 ) 700 suitable for implementing an embodiment of the present disclosure is shown below. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in FIG7 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0197] As shown in FIG7 , the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the electronic device 700 are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An edit / output (I / O) interface 705 is also connected to the bus 704.

[0198] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 7 shows the electronic device 700 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may be implemented or present instead.

[0199] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0200] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0201] The electronic device provided by the embodiment of the present disclosure and the special effects processing method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in the embodiment of the present disclosure, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0202] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the special effect processing method provided in the above embodiment is implemented.

[0203] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0204] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0205] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0206] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: in response to a first trigger operation, obtains special effect description data, wherein the special effect description data at least includes content description data, and the content description data includes content description text and / or content description audio; in response to a second trigger operation input for the special effect description data, displays at least one special effect effect data corresponding to the special effect description data, wherein the special effect effect data includes a special effect image corresponding to the special effect description data and an image sound effect corresponding to the special effect image.

[0207] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0208] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0209] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0210] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0211] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0212] According to one or more embodiments of the present disclosure, [Example 1] provides a special effects processing method, including:

[0213] In response to a first trigger operation, obtaining special effect description data, wherein the special effect description data at least includes content description data, and the content description data includes content description text and / or content description audio;

[0214] In response to a second trigger operation input for the special effect description data, at least one special effect data corresponding to the special effect description data is displayed, wherein the special effect data includes a special effect image corresponding to the special effect description data and an image sound effect corresponding to the special effect image.

[0215] According to one or more embodiments of the present disclosure, [Example 2] provides the method of Example 1, further comprising:

[0216] Optionally, the display of at least one special effect data corresponding to the special effect description data includes: determining the text to be processed corresponding to the special effect description data, and determining a first text feature vector corresponding to the text to be processed; generating at least one special effect image based on the first text feature vector, and respectively determining the image sound effect corresponding to each of the special effect images; generating special effect data corresponding to the special effect description data based on the special effect image and the image sound effect corresponding to the special effect image.

[0217] According to one or more embodiments of the present disclosure, [Example 3] provides the method of Example 2, further comprising:

[0218] Optionally, the content description data includes forward description data and reverse description data, and the text to be processed includes forward description text corresponding to the forward description data and reverse description text corresponding to the reverse description data; the forward description data is used to indicate image content expected to be included in the special effect image, and the reverse description data is used to suppress image content appearing in the special effect image; determining the text to be processed corresponding to the special effect description data and determining the first text feature vector corresponding to the text to be processed include: determining the forward description text corresponding to the forward description data and determining the forward text vector corresponding to the forward description text; determining the reverse description text corresponding to the reverse description data and determining the reverse text vector corresponding to the reverse description text; determining the first text feature vector based on the forward text vector and the reverse text vector.

[0219] According to one or more embodiments of the present disclosure, [Example 4] provides the method of Example 2, further comprising:

[0220] Optionally, at least one special effect image is generated based on the text to be processed, including: inputting the first text feature vector and the first Gaussian noise vector into at least one first image generation model, and obtaining special effect images output by each first image generation model respectively, wherein the first image generation model is trained on a pre-established first diffusion model based on the first sample feature vector corresponding to the first sample description text and the expected effect image corresponding to the first sample description text.

[0221] According to one or more embodiments of the present disclosure, [Example 5] provides the method of Example 2, further comprising:

[0222] Optionally, respectively determining the image sound effect corresponding to each of the special effect images includes: for each of the special effect images, determining an image description text corresponding to the special effect image, and generating text audio corresponding to the image description text.

[0223] According to one or more embodiments of the present disclosure, [Example 6] provides the method of Example 5, further comprising:

[0224] Optionally, generating text audio corresponding to the image description text includes: determining a second text feature vector corresponding to the image description text, inputting the second text feature vector and a second Gaussian noise vector into a second image generation model to obtain a target mel-spectrogram corresponding to the image description text, wherein the second image generation model is trained on a pre-established second diffusion model based on the second sample feature vector of the second sample description text and the expected mel-spectrogram corresponding to the second sample description text; converting the target mel-spectrogram into audio data, and using the converted audio data as the text audio corresponding to the image description text.

[0225] According to one or more embodiments of the present disclosure, [Example 7] provides the method of Example 2, further comprising:

[0226] Optionally, determining the image description text corresponding to the special effect image includes: determining image feature data corresponding to the special effect image, converting the image feature data into image text features, generating image description text based on the image text features, and displaying the image description text corresponding to the special effect image.

[0227] According to one or more embodiments of the present disclosure, [Example 8] provides the method of Example 1, further comprising:

[0228] Optionally, the special effect description data also includes at least one expected special effect style of special effect data, and the expected special effect style includes an image style and / or a sound effect style; the display of at least one special effect data corresponding to the special effect description data includes: determining the image style and sound effect style to be used based on the expected special effect style; determining the special effect image corresponding to the content description data based on the image style to be used; determining the image sound effect corresponding to the special effect image based on the sound effect style to be used; generating special effect data based on the special effect image and the image sound effect, and displaying the special effect data.

[0229] According to one or more embodiments of the present disclosure, [Example 9] provides the method of Example 1, further comprising:

[0230] Optionally, the special effect description data also includes an expected generation quantity of special effect data; the display of at least one special effect data corresponding to the special effect description data includes: generating special effect data corresponding to the special effect description data based on each expected generation quantity, and displaying the special effect data.

[0231] According to one or more embodiments of the present disclosure, [Example 10] provides the method of Example 1, further comprising:

[0232] Optionally, the special effect description data also includes a correlation index, which is used to indicate the degree of correlation between the special effect data and the content description data; the display of at least one special effect data corresponding to the special effect description data includes: generating special effect data corresponding to the special effect description data based on each of the correlation indicators, and displaying the special effect data.

[0233] According to one or more embodiments of the present disclosure, [Example 11] provides the method of Example 1, further comprising:

[0234] Optionally, after the display of at least one special effect data corresponding to the special effect description data, it also includes: in response to a selection trigger operation for the displayed special effect data, generating target special effect data based on the selected special effect data; in response to a third trigger operation for the target special effect data, saving the target special effect data to a preset storage space and / or publishing the target special effect data to a target display platform.

[0235] According to one or more embodiments of the present disclosure, [Example 12] provides the method of Example 1, further comprising:

[0236] Optionally, after displaying at least one special effect data corresponding to the special effect description data, it also includes: when the special effect data includes two or more of the image sound effects, displaying the sound effect playback controls corresponding to each of the image sound effects respectively; in response to a control trigger operation acting on the sound effect playback control, playing the image sound effect corresponding to the triggered sound effect playback control.

[0237] According to one or more embodiments of the present disclosure, [Example 13] provides a special effects processing device, including:

[0238] a special effect description acquisition module, configured to acquire special effect description data in response to a first trigger operation, wherein the special effect description data at least includes content description data, and the content description data includes content description text and / or content description audio;

[0239] A special effect display module is used to display at least one special effect data corresponding to the special effect description data in response to a second trigger operation input for the special effect description data, wherein the special effect data includes a special effect image corresponding to the special effect description data and an image sound effect corresponding to the special effect image.

[0240] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0241] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0242] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A special effects processing method, comprising: In response to a first trigger operation, obtaining special effect description data, wherein the special effect description data at least includes content description data, and the content description data includes content description text and / or content description audio; In response to a second trigger operation input for the special effect description data, at least one special effect data corresponding to the special effect description data is displayed, wherein the special effect data includes a special effect image corresponding to the special effect description data and an image sound effect corresponding to the special effect image.

2. The special effect processing method according to claim 1, wherein the displaying of at least one special effect data corresponding to the special effect description data comprises: Determining a text to be processed corresponding to the special effect description data, and determining a first text feature vector corresponding to the text to be processed; generating at least one special effect image based on the first text feature vector, and determining the image and sound effects corresponding to each of the special effect images; Special effect data corresponding to the special effect description data is generated based on the special effect image and the image sound effect corresponding to the special effect image, and the special effect data is displayed.

3. The special effect processing method according to claim 2, wherein the content description data includes forward description data and reverse description data, and the text to be processed includes forward description text corresponding to the forward description data and reverse description text corresponding to the reverse description data; the forward description data is used to indicate image content expected to be included in the special effect image, and the reverse description data is used to suppress image content appearing in the special effect image; The step of determining the text to be processed corresponding to the special effect description data and determining a first text feature vector corresponding to the text to be processed includes: Determine a forward description text corresponding to the forward description data, and determine a forward text vector corresponding to the forward description text; Determine the reverse description text corresponding to the reverse description data, and determine the reverse text vector corresponding to the reverse description text; A first text feature vector is determined based on the forward text vector and the reverse text vector.

4. The special effect processing method according to claim 2, wherein generating at least one special effect image based on the first text feature vector comprises: The first text feature vector and the first Gaussian noise vector are input into at least one first image generation model to obtain special effect images output by each of the first image generation models, wherein the first image generation model is trained based on a pre-established first diffusion model based on the first sample feature vector corresponding to the first sample description text and the expected effect image corresponding to the first sample description text.

5. The special effect processing method according to claim 2, wherein the image sound effect comprises text audio corresponding to the image description text; and the step of respectively determining the image sound effect corresponding to each of the special effect images comprises: For each of the special effect images, an image description text corresponding to the special effect image is determined, and a text audio corresponding to the image description text is generated.

6. The special effect processing method according to claim 5, wherein the generating of the text audio corresponding to the image description text comprises: Determine a second text feature vector corresponding to the image description text, input the second text feature vector and a second Gaussian noise vector into a second image generation model, and obtain a target mel-spectrogram corresponding to the image description text, wherein the second image generation model is trained on a pre-established second diffusion model based on a second sample feature vector of the second sample description text and an expected mel-spectrogram corresponding to the second sample description text; The target mel-spectrogram is converted into audio data, and the converted audio data is used as text audio corresponding to the image description text.

7. The special effect processing method according to claim 2, wherein the determining the image description text corresponding to the special effect image comprises: Determine image feature data corresponding to the special effect image, convert the image feature data into image text features, generate image description text based on the image text features, and display the image description text corresponding to the special effect image.

8. The special effect processing method according to claim 1, wherein the special effect description data further comprises at least one desired special effect style of the special effect data, and the desired special effect style comprises an image style and / or a sound effect style; The displaying of at least one special effect data corresponding to the special effect description data includes: Determining an image style and a sound effect style to be used based on the desired special effect style; Determining a special effect image corresponding to the content description data based on the image style to be used; Determining an image sound effect corresponding to the special effect image based on the sound effect style to be used; Special effect data is generated based on the special effect image and the image sound effect, and the special effect data is displayed.

9. The special effect processing method according to claim 1, wherein the special effect description data further includes an expected number of special effect data to be generated, including: The displaying of at least one special effect data corresponding to the special effect description data includes: Special effect data corresponding to the content description data is generated based on each of the expected generation quantities, and the special effect data is displayed.

10. The special effect processing method according to claim 1, wherein the special effect description data further includes a correlation index, and the correlation index is used to indicate the degree of correlation between the special effect data and the content description data, including: The displaying of at least one special effect data corresponding to the special effect description data includes: Special effect data corresponding to the special effect description data is generated based on each of the correlation indicators, and the special effect data is displayed.

11. The special effect processing method according to claim 1, wherein after displaying at least one special effect data corresponding to the special effect description data, it further comprises: In response to a selection triggering operation on the displayed special effect data, generating target special effect data corresponding to the special effect description data based on the selected special effect data; In response to a third trigger operation on the target special effect data, the target special effect data is saved to a preset storage space and / or the target special effect data is published to a target display platform.

12. The special effect processing method according to claim 1, wherein after displaying at least one special effect data corresponding to the special effect description data, it further comprises: In the case where the special effect data includes two or more image and sound effects corresponding to the special effect image, respectively displaying a sound effect playback control corresponding to each of the image and sound effects; In response to a control triggering operation acting on the sound effect playing control, the image sound effect corresponding to the triggered sound effect playing control is played.

13. A special effects processing device, comprising: A special effect description acquisition module, configured to acquire special effect description data in response to a first trigger operation, wherein the special effect description data at least includes content description data, and the content description data includes content description text and / or content description audio; A special effect display module is used to display at least one special effect data corresponding to the special effect description data in response to a second trigger operation input for the special effect description data, wherein the special effect data includes a special effect image corresponding to the special effect description data and an image sound effect corresponding to the special effect image.

14. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the special effect processing method as described in any one of claims 1-12.

15. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to execute the special effect processing method according to any one of claims 1 to 12 when executed by a computer processor.

Citation Information

Patent Citations

  • Image-processing device, method, and program

    CN103155536A

  • Text-based virtual object animation generation method and device, storage medium and terminal

    CN112184858A

  • Image generation and diffusion model training method, electronic equipment and storage medium

    CN116450873A

  • Media content generation method and device, electronic equipment and storage medium

    CN116886989A

  • Special effect processing method and device, electronic equipment and storage medium

    CN117440206A