Visual content generation method and device, equipment, medium and program product

By extracting and fusing low-frequency and high-frequency components of visual content, the clarity of visual content is optimized, the problem of detail distortion during the generation process is solved, and the generation effect of visual content is improved.

CN120897103APending Publication Date: 2025-11-04BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510998922.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

In existing visual content generation methods, the details in the generated visual results are easily distorted, affecting the generation effect.

Method used

By acquiring the first and second visual content generated by the model, blurring is performed on each to extract low-frequency and high-frequency components. These components are then combined to generate a third visual content with higher clarity. Finally, content fusion is performed to generate a fourth visual content to optimize clarity.

Benefits of technology

It improves the clarity of target content in visual content, solves the problem of detail distortion, and enhances the generation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897103A_ABST
    Figure CN120897103A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a visual content generation method and device, equipment, a medium and a program product, and the method comprises the steps: obtaining first visual content generated by a model based on input information, and enabling the input information to comprise second visual content, the second visual content comprises target content in the first visual content, and the definition of the target content is higher than that in the first visual content; according to the first visual content and the second visual content, third visual content is generated, the target content is contained in the third visual content, and the definition of the target content is higher than that of the first visual content; and content fusion is performed on the first visual content and the third visual content to generate fourth visual content, and the fourth visual content comprises other content except the target content in the first visual content and the target content in the third visual content. By means of the method, it is guaranteed that detailed content in the generated visual content is clearly displayed, the generation effect of the visual content is effectively improved, and then the application effect of the generated visual content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to the technical field of visual processing, and particularly relate to a visual content generation method, device, equipment, medium and program product. BACKGROUND

[0002] At present, the application of artificial intelligence generated content has been perfected, which can generate the desired target information based on some input information. The generation of visual content is one application of artificial intelligence generated content, such as the generation of visual content such as images or videos.

[0003] In the existing visual content generation implementation, the generated visual results often contain part or all of the content in the input information. However, compared with the content originally contained in the input information, the part of the content contained in the visual results often has the problem of content distortion. Many details that may exist in the original content are blurred or distorted in the generated visual content, thereby affecting the generation effect of the visual content and the application effect of the generated visual content. SUMMARY

[0004] Embodiments of the present disclosure provide a visual content generation method, device, equipment, medium and program product, which can ensure clear display of detailed content in the generated visual content and improve the generation effect of the visual content.

[0005] In a first aspect, the embodiments of the present disclosure provide a visual content generation method, which comprises:

[0006] obtaining first visual content generated by a model based on input information, the input information comprising second visual content, target content in the second visual content being contained in the first visual content, and the clarity of the target content in the first visual content being lower than that in the second visual content;

[0007] generating third visual content according to the first visual content and the second visual content, the third visual content containing the target content, and the clarity of the target content in the third visual content being higher than that in the first visual content;

[0008] performing content fusion on the first visual content and the third visual content to generate fourth visual content, the fourth visual content containing other content in the first visual content except the target content and containing the target content in the third visual content.

[0009] In a second aspect, the embodiments of the present disclosure also provide a visual content generation device, which comprises:

[0010] The first generation module is configured to obtain first visual content generated by a model based on input information, the input information comprising second visual content, target content in the second visual content being contained in the first visual content, and the target content having a lower definition in the first visual content than in the second visual content;

[0011] The second generation module is configured to generate third visual content according to the first visual content and the second visual content, the third visual content containing the target content, and the target content having a higher definition in the third visual content than in the first visual content;

[0012] The third generation module is configured to perform content fusion on the first visual content and the third visual content to generate fourth visual content, the fourth visual content containing other content in the first visual content except the target content and containing the target content in the third visual content.

[0013] In a fourth aspect, the embodiments of the present disclosure further provide a computer device, which comprises:

[0014] one or more processors;

[0015] a storage device configured to store one or more programs,

[0016] when the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating visual content provided by any of the embodiments of the present disclosure.

[0017] In a fifth aspect, the embodiments of the present disclosure further provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method for generating visual content provided by any of the embodiments of the present disclosure.

[0018] In a sixth aspect, the embodiments of the present disclosure further provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the method for generating visual content provided by any of the embodiments of the present disclosure.

[0019] The technical manner of the embodiments of the present disclosure specifically discloses a visual content generation method, device, equipment, medium and program product. The method can first acquire first visual content generated by a model based on input information. The input information includes second visual content. Target content in the second visual content is contained in the first visual content, and the clarity of the target content in the first visual content is lower than that in the second visual content. The third visual content is generated according to the first visual content and the second visual content. The third visual content contains the target content, and the clarity of the target content in the third visual content is higher than that in the first visual content. The content fusion is performed on the first visual content and the third visual content to generate the fourth visual content. The fourth visual content contains other content in the first visual content except the target content and contains the target content in the third visual content. The technical solution of the embodiment is equivalent to optimizing the visual content generated by the model. Specifically, the clarity of the target content in the visual content can be optimized, so that the target content has higher clarity in the finally optimized visual content. Therefore, the distortion problem of the target content in the generated visual content is optimized, the generation effect of the visual content is effectively improved, and the application effect of the generated visual content is further improved. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical manner of the exemplary embodiments of the present disclosure, the drawings needed in the description of the embodiments are briefly introduced. Obviously, the drawings introduced are only a part of the drawings of the present disclosure to be described, and not all the drawings. Those skilled in the art can obtain other drawings according to these drawings without creating creative labor.

[0021] Figure 1a A flowchart of a visual content generation method provided by the embodiments of the present disclosure is given;

[0022] Figure 1b An example flowchart of a visual content generation method provided by the embodiments of the present disclosure is given;

[0023] Figure 2a An effect display diagram of visual content generated by related visual content generation technology is given;

[0024] Figure 2b An effect display diagram of visual content generated by the visual content generation method provided by the embodiments of the present disclosure is given;

[0025] Figure 3 A structural diagram of a visual content generation device provided by the embodiments of the present disclosure is given;

[0026] Figure 4A structural schematic diagram of a computer device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are merely for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0028] It should be understood that each step described in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0029] The term "comprising" and variations thereof as used in the present disclosure are open-ended, that is, "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions of other terms will be given in the description below.

[0030] It should be noted that the "first", "second", and the like concepts mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units. It should be noted that the "one", "multiple" modification mentioned in the present disclosure is illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly stated in the context, it should be understood as "one or more".

[0031] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0032] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0033] It can be understood that, before using the technical means disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.

[0034] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium performing the operation of the technical solution of the present disclosure according to the prompt information.

[0035] As an optional but non-limiting implementation, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in a text manner. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0036] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation of the present disclosure. Other manners meeting relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0037] It should be noted that in the related technical implementation of visual content generation, due to the encoding and decoding operations on the input information, lossy compression of the visual content is caused, so that the generated visual content cannot maintain the detail content of some regions in the original visual content. However, the content in the region may have high application value and high detail preservation requirement. If the detail content is missing in the generated visual content, it may affect the application effect of the generated visual content.

[0038] Based on this, the embodiment provides a visual content generation method to optimize the generated visual content and realize high-fidelity presentation of the detail content carried in the input information by the generated visual content. Thus, the generation effect of the visual content is improved and the application effect of the visual content is ensured.

[0039] Figure 1a A flowchart of a visual content generation method provided by the embodiment of the present disclosure is shown. The embodiment can be applied to the case of generating visual content. The method can be executed by a visual content generation device. The device can be implemented by software and / or hardware. The device can be configured in a terminal and / or a server to implement the visual content generation method in the embodiment of the present disclosure.

[0040] As Figure 1a shown, the visual content generation method provided by the embodiment can include:

[0041] S101, acquire first visual content generated by a model based on input information, the input information including second visual content, target content in the second visual content being contained in the first visual content, and the target content having a lower definition in the first visual content than in the second visual content.

[0042] In the embodiment, the model can be considered as a network model for visual content generation, the input information can be considered as information for visual content generation, and the input information can be composed of information of different modalities, such as visual content of video type or image type, and can also include text information of text type, which can be description information for illustrative description of visual content generation, or can include information of audio type. The embodiment can record the visual content obtained after inputting the input information into the model as the first visual content, and record the visual content included in the input information as the second visual content.

[0043] In the embodiment, the first visual content can be considered as being generated based on the second visual content, and the first visual content can perform visual enhancement on some content based on the second visual content. What kind of visual enhancement is actually performed can be determined according to the application requirements of the visual content, and the embodiment does not make specific requirements. For example, when there is an application requirement of creative design, the generated visual content can be a creative picture or a creative video with a blue sky and white clouds designed based on the second visual content. The second visual content is a white background picture, and the visual content generated by creative design is a creative picture or a creative video with a blue sky and white clouds.

[0044] It should be noted that in the first visual content generated by performing visual enhancement on the second visual content, some subject content in the second visual content is often contained, but in the first visual content, it can not be guaranteed that some details of the subject content are clearly presented. In the embodiment, the target content can be considered as being contained in both the second visual content and the first visual content, and having a lower definition in the first visual content than the original definition in the second visual content, such as text content that has a distortion problem in the process of generating visual content from the second visual content.

[0045] In the embodiment, the target content can be one or more, and different target contents can contain different content objects. As long as the content exists in the second visual content and has a distortion problem in the first visual content, it can be the target content of the embodiment.

[0046] S102, generate third visual content according to the first visual content and the second visual content, the third visual content containing the target content, and the target content having a higher definition in the third visual content than in the first visual content.

[0047] It should be noted that the target content included in the first visual content is affected in clarity due to the distortion caused by the processing logic adopted for generation. In a related technical implementation for solving the distortion problem, one way is to directly paste the region containing the target content in the second visual content to the first visual content, but this implementation has a problem that if the target content region is directly pasted to the first visual content, a part of the content region of the generated content in the first visual content will be covered, and if the part of the content region involves visual enhancement processing of the original content, the covering of the part of the content region will affect the original visual enhancement effect.

[0048] In the embodiment, the influence of the target content backfill on the visual enhancement effect can be solved, and specifically, the first visual content and the second visual content can be combined and processed to generate new third visual content, which also contains the target content and the target content has higher clarity in the third visual content than in the first visual content.

[0049] In one implementation, the process of combining the first visual content and the second visual content to generate the third visual content can be described as follows: first, the first visual content and the second visual content are blurred in the frequency domain or the spatial domain to obtain respective visual low-frequency components, and then the corresponding visual low-frequency components can be filtered out from the frequency spectrum corresponding to the second visual content in the frequency domain, and the remaining frequency spectrum after filtering out can be considered as the visual high-frequency component corresponding to the second visual content.

[0050] It should be noted that the blurring of the visual content in the embodiment only reduces the clarity of the visual content, and does not reduce the amount of actual content objects contained in the visual content. Therefore, it can be considered that the visual low-frequency components of the first visual content and the second visual content respectively correspond to visual content in the time domain, which still contains content objects consistent with the original visual content. At the same time, it can be known that the low-frequency component of the visual content in the frequency domain is generally a description of the main content objects in the visual content. For the second visual content, the visual low-frequency component of the second visual content can be filtered out from the frequency spectrum of the second visual content in the frequency domain to obtain the visual high-frequency component of the second visual content, which can be considered as a description of the edges or contours of the visual content.

[0051] In the embodiment, the influence of the first visual content on the clarity of the target content can be considered as mainly the distortion of the edge or contour of the target content. Meanwhile, considering that the visual low-frequency component represents the description of the main content object, the embodiment can extract the visual high-frequency component representing the description of the edge or contour of the target content from the second visual content in the above manner; and then the visual high-frequency component of the second visual content can be combined with the visual low-frequency component of the first visual content.

[0052] The above manner of the embodiment can superimpose the visual low-frequency component of the target content of the first visual content and the visual high-frequency component of the target content in the second visual content, so as to form the target content with both the visual high-frequency component and the visual low-frequency component. The target content formed by the above manner clearly shows the main content object contained therein and the edge or contour of the main content object, which is equivalent to realizing the visual enhancement of the target content. The embodiment can record the visual content containing the target content formed by the above manner as the third visual content.

[0053] S103, content fusion is performed on the first visual content and the third visual content to generate fourth visual content, the fourth visual content containing other content in the first visual content except the target content and containing the target content in the third visual content.

[0054] In the embodiment, it can be known from the above description that the third visual content generated contains the target content with higher clarity. The embodiment can fuse the third visual content with the first visual content through the step, so as to obtain the fourth visual content containing high-definition content. The fourth visual content can be considered to contain other content in the first visual content except the original target content and contain the target content in the third visual content.

[0055] In the embodiment, one implementation of generating the fourth visual content can be to cut out the layer containing the target content from the third visual content, and obtain the other content layer after excluding the target content from the first visual content. Then, the layer containing the target content and the other content layer are fused again, and the fused visual content constitutes the fourth visual content.

[0056] The method for generating visual content provided by the embodiment of the disclosure is equivalent to optimizing the visual content generated by the model. Specifically, the clarity of the target content in the visual content can be optimized, so that the target content has higher clarity in the finally optimized visual content, thereby optimizing the distortion problem of the target content in the generated visual content, effectively improving the generation effect of the visual content, and further improving the application effect of the generated visual content.

[0057] In the embodiment, the visual content is common in image type and video type, and in the generation implementation of the visual content, the generation of the visual content in the image type based on the visual content in the image type is mainly considered, the generation of the visual content in the video type based on the visual content in the image type is also considered, and the generation of the visual content in the video type based on the visual content in the video type is also considered. Different types of visual content used in the input information have different specific implementations of distortion optimization of the generated visual content.

[0058] As a first optional embodiment of the embodiment, on the basis of the above embodiment, in the case that the first visual content and the second visual content are both image types, generating third visual content according to the first visual content and the second visual content can be specifically implemented as the following steps:

[0059] a1) Gaussian blur processing is performed on the first visual content and the second visual content respectively to obtain a first image low-frequency component of the first visual content and a second image low-frequency component of the second visual content.

[0060] In the embodiment, the second visual content in the input information and the first visual content generated by the model can be considered as image types, and the first visual content and the second visual content of the image type can be processed by the blur processing, and the Gaussian blur processing can be specifically performed.

[0061] In the embodiment, the Gaussian blur processing can be considered as being performed on the first visual content and the second visual content in the frequency domain, and the Fourier transform can be first performed on the first visual content and the second visual content to obtain the spectral data of the visual content in the frequency domain. Then, the same convolution kernel can be used to filter the spectral data of the visual content by using the Gaussian blur filter. Finally, the image low-frequency component of each visual content is obtained with respect to the first visual content and the second visual content, and the first image low-frequency component corresponds to the first visual content, and the second image low-frequency component corresponds to the second visual content.

[0062] In the embodiment, the information obtained after the Gaussian blur processing of the visual content can be considered as the visual low-frequency component of the visual content in the frequency domain. For the visual content in the image type, the image low-frequency component is obtained, which can represent the region with slow change of brightness or grayscale value in the image, and is equivalent to the content region with large proportion for describing the main content of the image.

[0063] b1) The second image low-frequency component is screened out from the spectrum of the second visual content to obtain a first image high-frequency component, and the target spectrum corresponding to the target content is reserved in the first image high-frequency component.

[0064] In the embodiment, the second visual content is the visual content in the input information, and the second image low-frequency component of the second visual content can be considered as a main content region with slow luminance or grayscale value change in the second visual content. The embodiment can use the frequency spectrum of the second visual content in the frequency domain to filter the low-frequency component of the second image low-frequency component, and the frequency spectrum after filtering the low-frequency component can be considered as the visual high-frequency component of the second visual content. The embodiment can record the visual high-frequency component corresponding to the image type visual content as an image high-frequency component, and can record the image high-frequency component of the second visual content as a first image high-frequency component.

[0065] In the embodiment, the image high-frequency component can represent a content region with a distance of luminance or grayscale value change in the image, and the region is often an edge (outline) and a detail part of the image. Therefore, the image high-frequency component can be considered as a measure of the edge or detail content of the image.

[0066] It can be known that the second visual content contains target content, and the target content can also be composed of a main content with slow luminance or grayscale value change and an edge (outline) or detail content. After filtering the image low-frequency component of the frequency spectrum of the second visual content by using the step, the first image high-frequency component obtained contains the frequency spectrum representing the edge or detail content of the target content. The embodiment records the frequency spectrum of this kind as a target frequency spectrum.

[0067] c1) generating the third visual content of the image type according to the first image high-frequency component and the first image low-frequency component.

[0068] In the embodiment, the embodiment also performs Gaussian blur processing on the first visual content in the above manner, and the obtained image low-frequency component represents a main content region in the first visual content, and the main content region also contains a main content region of the target content. In the embodiment, based on the first image high-frequency component containing the edge and detail of the target content and the first image low-frequency component containing the main content of the target content, the fusion of the first image high-frequency component and the first image low-frequency component can obtain frequency spectrum data containing the complete frequency spectrum of the target content. After time domain conversion, the frequency spectrum data can form an image in the time domain, and the image can be used as the third visual content of the image type.

[0069] It can be known that the third visual content contains the target content, and specifically contains the main content and the edge and detail content of the target content. Based on the main content and the detail content of the target content, the clear presentation of the target content can be realized. Therefore, compared with the loss of details of the target content in the first visual content, the clarity of the target content contained in the third visual content is higher than that of the first visual content.

[0070] The technical solution of the embodiment provides implementation of the generation of the third visual content when the first visual content and the second visual content are both image types. The method provided by the embodiment is equivalent to supplementing and enhancing the target content in the first visual content by using the detailed content of the target content in the second visual content, thereby obtaining the third visual content with clear target content. The obtained third visual content provides basic support for content optimization of the first visual content and improves the generation effect of the visual content.

[0071] On the basis of the first optional embodiment, as an implementation manner, the first visual content and the third visual content can be fused to generate a fourth visual content by the following steps:

[0072] a2) determining a first segmentation map of the first visual content, the first segmentation map containing image content other than the target content.

[0073] The embodiment can be considered as further optimization when the first visual content and the second visual content are both image types. Specifically, the embodiment can first perform image segmentation on the first visual content by the step. For example, the step can perform segmentation on the first visual content by using segmentation logic of image segmentation, for example, foreground-background segmentation of the first visual content or specific segmentation of a certain content object contained in the first visual content. The foreground-background segmentation can obtain a foreground segmentation map and a background segmentation map of the first visual content, and the specific segmentation can obtain a segmentation map containing a specific content object. After the image segmentation of the first visual content, the embodiment can obtain the first segmentation map containing other content except the target content.

[0074] b2) determining a second segmentation map of the third visual content, the second segmentation map containing the target content.

[0075] In the embodiment, the third visual content can also be segmented by using the same segmentation logic as described above, and the second segmentation map containing the target content can be obtained after the segmentation.

[0076] c2) performing image fusion on the first segmentation map and the second segmentation map to generate the fourth visual content in an image type.

[0077] In the embodiment, the first segmentation map and the second segmentation map can be considered as layers with the same size and can be mask maps. The step can perform content fusion on the two layers, and specifically, the content other than the target content in the first visual content can be fused with the target content in the third visual content by fusing pixel information. The fused visual content can be recorded as the fourth visual content, and the fourth visual content can be in the same image type as the first visual content.

[0078] The technical solution of the embodiment above gives the implementation of generating the fourth visual content when the first visual content and the second visual content are both image types. The embodiment is equivalent to replacing the target content in the first visual content with the optimized target content. The optimized target content is from the third visual content, and the third visual content includes the clear target content and the main content of other content objects in the first visual content. Therefore, replacing the target content in the first visual content with the target content in the third visual content can well avoid the occlusion and influence of the replaced target content on other regions of the first visual content, and better solve the detail distortion in the first visual content.

[0079] As a second optional embodiment of the embodiment, on the basis of the embodiment above, in the case where the first visual content is of a video type and the second visual content is of an image type, generating the third visual content according to the first visual content and the second visual content can be further specified as:

[0080] a3) performing Gaussian blur processing on each video frame content included in the first visual content to obtain a video frame low-frequency component of each video frame content.

[0081] In the embodiment, the optional embodiment can be considered to mainly generate the third visual content in the case where the first visual content is of a video type and the second visual content is of an image type. The type of the third visual content can be considered to be consistent with the type of the first visual content, and is also of a video type.

[0082] Specifically, Gaussian blur processing can be performed on the first visual content through the step. Considering that the first visual content is of a video type, the first visual content is equivalent to a video, and Gaussian blur processing can be performed on each video frame content, so that a video frame low-frequency component can be obtained for each video frame content.

[0083] Similarly, the video frame low-frequency component of each video frame content can also be considered as a representation of the main content object in the corresponding video frame content, and it can be considered that each video frame content includes target content, and the main content region of the target content can also be represented in the video frame low-frequency component of each video frame content.

[0084] b3) performing Gaussian blur processing on the second visual content to obtain an image low-frequency component of the second visual content, and screening out the image low-frequency component from the frequency spectrum of the second visual content to obtain a second image high-frequency component, and the second image high-frequency component retains a target spectrum corresponding to the target content.

[0085] In the embodiment, the second visual content is of image type, and the step can perform Gaussian blur processing on the second visual content, and the image low frequency component of the second visual content can be obtained, and the second image low frequency component is recorded.

[0086] In the embodiment, after the second image low frequency component is filtered from the spectrum of the second visual content, the image high frequency component of the second visual content is reserved, which is recorded as the second image high frequency component, and the second image high frequency component is equivalent to the edge and details of the target content.

[0087] c3) generating the first frame superimposed content of each video frame content according to the image high frequency component and each frame low frequency component, and generating the third visual content of video type according to each first frame superimposed content.

[0088] In the embodiment, the video frame low frequency component of each video frame content can be superimposed with the second image high frequency component of the second visual content, and the superimposed content can be converted in time domain to obtain the frame superimposed content corresponding to each video frame content, which is recorded as the first frame superimposed content in the embodiment. Each first frame superimposed content is equivalent to containing the edge and details of the target content, and also contains the main content of the target content, thereby ensuring the high definition of the target content.

[0089] In the embodiment, the first frame superimposed content can be combined according to the frame number of the video frame, thereby forming the third video content of video type.

[0090] The above technical solution of the embodiment provides the generation and implementation of the third visual content when the first visual content is of video type and the second visual content is of image type. The method provided in the embodiment is equivalent to supplementing and enhancing the target content in the first visual content by using the detail content of the target content in the second visual content, thereby obtaining the third visual content with clear target content. The third visual content provides a basis for content optimization of the first visual content, and improves the generation effect of the visual content.

[0091] As a third optional embodiment of the embodiment, on the basis of the above-mentioned embodiment, when the first visual content and the second visual content are both of video type, the third visual content can be generated according to the first visual content and the second visual content, and the generation of the third visual content can be further optimized as follows:

[0092] a4) performing Gaussian blur processing on each first video frame content included in the first visual content respectively to obtain a first frame low frequency component of each first video frame content.

[0093] In the embodiment, it can be considered that the optional embodiment mainly performs generation processing of the third visual content in the case that the first visual content and the second visual content are both video types. The type of the third visual content can be considered to be consistent with the type of the first visual content, and is also a video type.

[0094] Specifically, Gaussian blur processing can also be performed on the first visual content through the step first. Considering that the first visual content is a video type, the first visual content is equivalent to a video. The step can perform Gaussian blur processing on each video frame content, so that a video frame low frequency component can be obtained for each video frame content. In the embodiment, the video frame low frequency component can be referred to as a first frame low frequency component.

[0095] Similarly, the first frame low frequency component of each video frame content can also be considered as a representation of the main content object in the corresponding video frame content, and it can be considered that each video frame content contains target content. The main content area of the target content can also be represented in the first frame low frequency component of each video frame content.

[0096] b4) performing Gaussian blur processing on each second video frame content included in the second visual content respectively to obtain a second frame low frequency component of each second video frame content, and screening out the corresponding second frame low frequency component from the spectrum of each second video frame content to obtain a frame high frequency component, and each frame high frequency component retains a target spectrum of the target content.

[0097] In the embodiment, Gaussian blur processing can also be performed on the second visual content which is also a video type through the step. Gaussian blur processing can be performed on each video frame content of the second visual content respectively, so that a video frame low frequency component can also be obtained for each video frame content. In the embodiment, the video frame low frequency component can be referred to as a second frame low frequency component.

[0098] It can be known that the second frame low frequency component can also be considered as a representation of the main content object in the corresponding video frame content, and it can be considered that each video frame content contains target content. The main content area of the target content can also be represented in the second frame low frequency component of each video frame content.

[0099] In the embodiment, after screening out the second frame low frequency component from each video frame content of the second visual content, the frame high frequency component of each video frame content of the second visual content is retained, and each frame high frequency component is equivalent to representing the edge and details of the target content.

[0100] c4) generating a second frame superimposition content of each of the first video frame contents according to each of the first frame low frequency components and the frame high frequency component corresponding to the frame number.

[0101] In the embodiment, for each of the first frame low frequency components and each of the frame high frequency components, the superimposition processing of the frequency domain data can be respectively performed, and after the time domain conversion of the superimposed content, the frame superimposition content corresponding to each of the video frame contents can be obtained. The frame superimposition content is recorded as the second frame superimposition content in the embodiment. The second frame superimposition content contains the edge and details of the target content, and also contains the main content of the target content, thereby ensuring the high definition of the target content.

[0102] d4) generating the third video content of the video type based on each of the second frame superimposition contents.

[0103] The embodiment can also combine the respective second frame superimposition contents according to the frame numbers of the video frames, thereby forming the third video content of the video type.

[0104] The above technical solution of the embodiment provides the generation and implementation of the third visual content when the first visual content and the second visual content are both of the video type. The provided mode is equivalent to supplementing and enhancing the target content in the first visual content by using the detail content of the target content in the second visual content, thereby obtaining the third visual content with clear target content. The obtained third visual content provides basic support for the content optimization of the first visual content, and also improves the generation effect of the visual content.

[0105] On the basis of the second optional embodiment or the third embodiment, as an implementation mode, the content fusion of the first visual content and the third visual content can be performed to generate the fourth visual content further optimized as follows:

[0106] a5) determining a third segmentation map of each of the first video frame contents in the first visual content, and each of the third segmentation maps containing image content other than the target content.

[0107] In the embodiment, it can be considered as a further refinement on the second optional embodiment, and specifically provides the optimization of the first visual content based on the third visual content. First, the first video frame contents in the first visual content can be segmented by the step. The image segmentation logic can also be used to segment each of the first video frame contents, thereby obtaining the corresponding segmentation map relative to each of the video frame contents. The segmentation map is recorded as the third segmentation map, and each of the third segmentation maps can be considered to include image content other than the target content in the video frame content.

[0108] b5) determining fourth segmentation maps of each third video frame content in the third visual content, each of the fourth segmentation maps comprising the target content.

[0109] In the embodiment, each third video frame content in the third visual content can be segmented by using the same segmentation logic as the above steps, thereby determining a segmentation map of each video frame content, which is denoted as a fourth segmentation map, and each fourth segmentation map can be considered to comprise the target content.

[0110] c5) fusing each third segmentation map with a fourth segmentation map corresponding to the frame number to generate frame fusion content.

[0111] In the embodiment, the frame image fusion of each third segmentation map and the fourth segmentation map with the same frame number can be implemented by the step, and the fused image content can be denoted as frame fusion content. Each frame fusion content is equivalent to comprising other content in the first video frame content involved in the first visual content except the target content, and also comprising the target content in the third video frame content involved in the third visual content.

[0112] d5) constructing the fourth visual content of the video type based on each frame fusion content.

[0113] In the embodiment, each fused image can be integrated in the order of the frame number to obtain the fourth visual content of the video type, and the fourth visual content is equivalent to comprising the target content with higher definition in each frame.

[0114] The above technical solution of the embodiment gives the generation and implementation of the fourth visual content when the first visual content is of the video type and the second visual content is of the image type. The embodiment is equivalent to replacing the target content in the first visual content with the optimized target content. The optimized target content comes from the third visual content, and the third visual content comprises the clear target content and the main content of other content objects in the first visual content. Therefore, the target content in the third visual content can be used to replace the target content in the first visual content, which can well avoid the shielding and influence of the replaced target content on other regional content in the first visual content, and better solve the detail distortion in the first visual content.

[0115] As a fourth optional embodiment of the embodiment, on the basis of the above embodiment, the fourth visual content can be further optimized and added as business application data to the business application platform in response to the triggering of the business application of the generated content.

[0116] In the embodiment, the generated visual content is mainly used to provide application support for a business application, which can take the fourth visual content as the finally generated visual content after optimization. The fourth visual content can be pre-stored or directly pushed to the business application platform in need. In one implementation, the fourth visual content meeting the demand of the business application can be obtained and provided to the business application platform after receiving a business application trigger for the generated visual content. In the embodiment, the business application platform can be different according to different supported business applications. When the business application involves e-commerce interaction, the business application platform can be determined as an e-commerce platform. When the business application involves entertainment interaction, the business application platform can be determined as a multimedia resource interaction platform.

[0117] The technical solution described above provides business application support for the generated visual content. Compared with the implementation of additional content production for the required visual content in the business application, the embodiment can provide the required business application data more simply and conveniently for the business application platform, greatly increasing the application value of the generated visual content in the business application.

[0118] To better understand the implementation of the visual content generation method provided in the embodiment, Figure 1b An example flowchart of the visual content generation method provided in the embodiment is given. As shown in Figure 1b The generation process of the visual content can be described as the following example steps:

[0119] S11, obtaining first visual content generated by a model based on input information, the input information containing second visual content, the first visual content and the second visual content both including target content.

[0120] S12, performing Gaussian blur processing on the first visual content and the second visual content respectively to obtain a first low-frequency component of the first visual content and a second low-frequency component of the second visual content.

[0121] S13, filtering the second low-frequency component from the frequency spectrum of the second visual content to obtain a visual high-frequency component of the second visual content.

[0122] S14, superimposing the visual high-frequency component and the first low-frequency component to obtain full-band visual content, and converting the full-band visual content into third visual content in the time domain, the third visual content including target content with higher definition than the first visual content.

[0123] S15, segmenting the first visual content to obtain first visual segmentation content containing other visual content except the target content.

[0124] S16, segmenting the third visual content to obtain second visual segmentation content containing the target content.

[0125] S17, performing content fusion on the first visual segmentation content and the second visual segmentation content, and determining the fused content as fourth visual content.

[0126] Exemplarily, Figure 2a An effect display diagram of visual content generated by using the related visual content generation technology is given. As Figure 2a shown, there is a problem of unclear presentation of the detailed content of the first target content 211 in the generated visual content 21.

[0127] Figure 2b An effect display diagram of visual content generated by using the visual content generation method provided in the embodiment is given. As Figure 2b shown, when the visual content is generated by using the same input information as Figure 2a , the detailed content of the first target content 221 in the formed visual content 22 can be clearly presented. It can be seen that the visual content generation method provided in the embodiment can effectively solve the problem of distortion of the generated detailed content.

[0128] Figure 3 A structural schematic diagram of a visual content generation apparatus provided in the embodiment of the disclosure is given. The embodiment can be applicable to the case of visual content generation. The apparatus can be realized by software and / or hardware, and can be configured in a terminal and / or a server to realize the visual content generation method in the embodiment of the disclosure, and preferably, can be configured in the content generation end in the embodiment of the disclosure. The apparatus can specifically include: a first generation module 31, a second generation module 32, and a third generation module 33.

[0129] The first generation module 31 is configured to acquire first visual content generated by a model based on input information, the input information including second visual content, target content in the second visual content being contained in the first visual content, and a clarity of the target content in the first visual content being lower than a clarity of the target content in the second visual content.

[0130] The second generation module 32 is configured to generate third visual content according to the first visual content and the second visual content, the third visual content containing the target content, and a clarity of the target content in the third visual content being higher than the clarity of the target content in the first visual content.

[0131] The third generation module 33 is configured to perform content fusion on the first visual content and the third visual content to generate fourth visual content, where the fourth visual content includes other content in the first visual content except the target content and includes the target content in the third visual content.

[0132] The visual content generation apparatus provided by the embodiments of the present disclosure is equivalent to optimizing the visual content generated by the model, and specifically, the clarity of the target content in the visual content can be optimized, so that the target content has higher clarity in the final optimized visual content, thereby optimizing the distortion problem of the target content in the generated visual content, effectively improving the generation effect of the visual content, and further improving the application effect of the generated visual content.

[0133] Further, the second generation module 32 can be specifically configured to:

[0134] In the case that the first visual content and the second visual content are both images, the first visual content and the second visual content are respectively subjected to Gaussian blur processing to obtain a first image low-frequency component of the first visual content and a second image low-frequency component of the second visual content; the second image low-frequency component is filtered out from a frequency spectrum of the second visual content to obtain a first image high-frequency component, and the target spectrum corresponding to the target content is reserved in the first image high-frequency component; and the first image high-frequency component and the first image low-frequency component are used to generate the third visual content in the image type.

[0135] Further, the third generation module 33 can be specifically configured to:

[0136] A first segmentation map of the first visual content is determined, and the first segmentation map includes other image content except the target content; a second segmentation map of the third visual content is determined, and the second segmentation map includes the target content; and the first segmentation map and the second segmentation map are subjected to image fusion to generate the fourth visual content in the image type.

[0137] Further, the second generation module 32 can be specifically configured to:

[0138] In a case that the first visual content is of a video type and the second visual content is of an image type, each video frame content included in the first visual content is respectively subjected to Gaussian blur processing to obtain a video frame low-frequency component of each video frame content; the second visual content is subjected to Gaussian blur processing to obtain an image low-frequency component of the second visual content, and the image low-frequency component is screened out from a spectrum of the second visual content to obtain a second image high-frequency component, the second image high-frequency component retaining a target spectrum corresponding to the target content; first frame superposition contents of each video frame content are generated according to the image high-frequency component and each frame low-frequency component, and the third visual content of the video type is generated according to each first frame superposition content.

[0139] Further, the second generation module 32 can also be specifically used for:

[0140] In a case that the first visual content and the second visual content are both of a video type, each first video frame content included in the first visual content is respectively subjected to Gaussian blur processing to obtain a first frame low-frequency component of each first video frame content; each second video frame content included in the second visual content is respectively subjected to Gaussian blur processing to obtain a second frame low-frequency component of each second video frame content, and the corresponding second frame low-frequency component is screened out from a spectrum of each second video frame content to obtain a corresponding frame high-frequency component, each frame high-frequency component retaining a target spectrum of the target content; second frame superposition contents of each first video frame content are generated according to each first frame low-frequency component and the frame high-frequency component of the corresponding frame number; and the third visual content of the video type is constituted based on each second frame superposition content.

[0141] Further, the third generation module 33 can also be specifically used for:

[0142] determining third segmentation maps of each first video frame content in the first visual content, each third segmentation map containing image content other than the target content;

[0143] determining fourth segmentation maps of each third video frame content in the third visual content, each fourth segmentation map containing the target content;

[0144] performing content fusion on each third segmentation map and the fourth segmentation map of the corresponding frame number to generate frame fusion content;

[0145] constituting the fourth visual content of the video type based on each frame fusion content.

[0146] Further, the apparatus can also include:

[0147] The response module is used to respond to the business application triggering the generated content and feed the fourth visual content back to the business application platform as business application data.

[0148] The above-described apparatus can execute the visual content generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0149] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0150] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Reference is made below. Figure 4 It illustrates a computer device suitable for implementing embodiments of the present disclosure (e.g., Figure 4 The diagram below shows the structure of the terminal device or server 40. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0151] like Figure 4 As shown, the computer device 40 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 41, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 42 or a program loaded from a storage device 48 into a random access memory (RAM) 43. The RAM 43 also stores various programs and data required for the operation of the computer device 40. The processing unit 41, ROM 42, and RAM 43 are interconnected via a bus 45. An edit / output (I / O) interface 44 is also connected to the bus 45.

[0152] Typically, the following devices can be connected to I / O interface 44: input devices 46 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 47 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 48 including, for example, magnetic tapes, hard disks, etc.; and communication devices 49. Communication device 49 allows computer device 40 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4The computer device 40 is shown with various devices, but it is understood that not all of the shown devices are required to be implemented or present. More or fewer devices can alternatively be implemented or present.

[0153] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication device 49, or installed from the storage device 48, or installed from the ROM 42. When the computer program is executed by the processing device 41, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0154] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0155] The computer device provided by the embodiments of the present disclosure and the method for generating visual content provided by the above-mentioned embodiments belong to the same concept, and the technical details not described in detail in the present embodiment can be referred to the above-mentioned embodiments, and the present embodiment has the same beneficial effects as the above-mentioned embodiments.

[0156] The embodiments of the present disclosure provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating visual content provided by the above-mentioned embodiments.

[0157] It should be noted that the computer readable medium of the present disclosure described above can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0158] In this disclosure, a computer readable storage medium can be any tangible medium that can contain, or store computer readable program codes. In this disclosure, a computer readable signal medium can include a data signal traveling in baseband or traveling as part of a carrier wave traveling in baseband upon a data channel, wherein the data channel is time, frequency, or code phase shifted by the computer readable program codes. Such data signals can take a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer readable medium can be any computer readable medium except for a transitory computer readable medium. The computer readable medium can be a computer readable storage medium or a computer readable signal medium. The computer readable storage medium can be any available medium or a combination thereof that is accessible by a computer. The computer readable storage medium can be a computer readable storage medium that is non-transitory. The computer readable storage medium can be a computer readable storage medium that is tangible or a computer readable signal medium.

[0159] In some embodiments, the data requestor, the server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications of any form or medium (e.g., a communications network). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0160] The computer readable medium described above can be contained in the computer device described above; or can exist separately, without being assembled into the computer device.

[0161] The computer readable medium described above carries one or more programs, which, when executed by the computer device, cause the computer device to:

[0162] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0163] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0164] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.

[0165] The functions described in this specification can be performed at least in part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0166] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a lined- up electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0167] The foregoing description merely exemplifies the preferred embodiments of the disclosure and the principles of the technology applied. It is understood by those skilled in the art that the disclosed scope of the disclosure is not limited to the technical manner formed by the specific combination of the technical features described above, and also covers other technical manners formed by any combination of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical manner formed by the mutual replacement of the above-described features and the technical features disclosed in the disclosure (but not limited to) having similar functions.

[0168] Furthermore, although each operation is depicted in a particular order, this should not be understood as requiring these operations to be performed in the particular order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination.

[0169] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A method for generating visual content, characterized in that, include: Obtain first visual content generated by the model based on input information, wherein the input information includes second visual content, the target content in the second visual content is contained in the first visual content, and the clarity of the target content in the first visual content is lower than the clarity in the second visual content; Based on the first visual content and the second visual content, a third visual content is generated, wherein the third visual content includes the target content, and the clarity of the target content in the third visual content is higher than the clarity in the first visual content; The first visual content and the third visual content are fused to generate a fourth visual content, which includes other content in the first visual content except for the target content, and includes the target content in the third visual content.

2. The method according to claim 1, characterized in that, When both the first visual content and the second visual content are image types. The step of generating third visual content based on the first visual content and the second visual content includes: Gaussian blurring is applied to the first visual content and the second visual content respectively to obtain the first low-frequency component of the first visual content and the second low-frequency component of the second visual content. The low-frequency component of the second image is filtered out from the spectrum of the second visual content to obtain the high-frequency component of the first image, wherein the target spectrum corresponding to the target content is retained in the high-frequency component of the first image. The third visual content of the image type is generated based on the high-frequency component and the low-frequency component of the first image.

3. The method according to claim 2, characterized in that, The step of fusing the first visual content and the third visual content to generate the fourth visual content includes: A first segmentation map of the first visual content is determined, wherein the first segmentation map contains other image content besides the target content; A second segmentation map of the third visual content is determined, wherein the second segmentation map contains the target content; The first segmentation image and the second segmentation image are fused to generate the fourth visual content of image type.

4. The method according to claim 1, characterized in that, When the first visual content is a video and the second visual content is an image. The step of generating third visual content based on the first visual content and the second visual content includes: Gaussian blurring is applied to each video frame content included in the first visual content to obtain the low-frequency components of each video frame content. The second visual content is subjected to Gaussian blur processing to obtain the low-frequency component of the image of the second visual content, and the low-frequency component of the image is filtered out from the spectrum of the second visual content to obtain the high-frequency component of the second image, wherein the target spectrum corresponding to the target content is retained in the high-frequency component of the second image. Based on the high-frequency components of the second image and the low-frequency components of each frame, a first frame overlay content of each video frame is generated, and the third visual content of the video type is generated based on the first frame overlay content.

5. The method according to claim 1, characterized in that, When both the first visual content and the second visual content are video types. The step of generating third visual content based on the first visual content and the second visual content includes: Gaussian blurring is applied to each of the first video frames included in the first visual content to obtain the first low-frequency component of each first video frame content. Gaussian blurring is performed on each of the second video frames included in the second visual content to obtain the second low-frequency component of each second video frame content, and the corresponding second low-frequency component is filtered out from the spectrum of each second video frame content to obtain the corresponding high-frequency component of the frame, wherein the target spectrum of the target content is retained in each of the high-frequency components of the frame. Based on the low-frequency components of each first frame and the high-frequency components of the corresponding frame number, the second frame overlay content of each first video frame is generated. The third video content is composed of the superimposed content of each of the second frames, forming a video type.

6. The method according to claim 4 or 5, characterized in that, The step of fusing the first visual content and the third visual content to generate the fourth visual content includes: A third segmentation map of the content of each first video frame in the first visual content is determined, and each third segmentation map contains other image content besides the target content; A fourth segmentation map is determined for each third video frame content in the third visual content, and each of the fourth segmentation maps contains the target content; Each of the third segmentation images is fused with the fourth segmentation image of the corresponding frame number to generate frame fused content; The fourth visual content of the video type is constituted based on the fused content of each frame.

7. The method according to any one of claims 1-6, characterized in that, Also includes: In response to the business application triggering the generated content, the fourth visual content is fed back to the business application platform as business application data.

8. A device for generating visual content, characterized in that, include: The first generation module is used to obtain first visual content generated by the model based on input information, wherein the input information includes second visual content, the target content in the second visual content is contained in the first visual content, and the clarity of the target content in the first visual content is lower than the clarity in the second visual content. The second generation module is used to generate a third visual content based on the first visual content and the second visual content, wherein the third visual content includes the target content, and the clarity of the target content in the third visual content is higher than the clarity in the first visual content. The third generation module is used to perform content fusion on the first visual content and the third visual content to generate a fourth visual content. The fourth visual content includes other content in the first visual content except for the target content, and includes the target content in the third visual content.

9. A computer device, characterized in that, The computer device includes: One or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method for generating visual content as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the method for generating visual content as described in any one of claims 1-7.

11. A computer program product comprising a computer program that, when executed by a processor, implements a method for generating visual content according to any one of claims 1-7.