Media processing method and device, equipment and storage medium
By generating descriptive text for media content and using machine learning models to select the best image enhancement scheme, the problem of poor image enhancement effect in existing technologies has been solved, and better visual quality has been achieved.
Patent Information
- Application Number
- CN202511013436.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies for image enhancement suffer from the problem that it is difficult to determine the combined calling strategy and the image enhancement effect cannot be guaranteed.
It generates descriptive text for media content, uses preset image processing tools and machine learning models to generate multiple image enhancement schemes, and selects the best scheme for processing based on evaluation information.
It enables the determination of the target image enhancement scheme based on the evaluation information of multiple schemes, ensuring that the generated media content has better visual quality.
Smart Images

Figure CN120852218A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to media processing methods, apparatus, devices, and computer-readable storage media. Background Technology
[0002] With the advancement of computer technology, various forms of electronic devices have greatly enriched people's daily lives. For example, electronic devices can be used to enhance the image quality of media content. Image enhancement technology has been widely applied in various fields, such as medical imaging, autonomous driving, the internet, and social media. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for media processing is provided. The method includes: generating descriptive text for first media content, the descriptive text describing the image information of the first media content; providing the descriptive text to a first model to generate multiple image enhancement schemes corresponding to the first media content, wherein each image enhancement scheme includes at least one image processing step performed using a preset image processing tool; determining a target image enhancement scheme from the multiple image enhancement schemes based on first evaluation information for the multiple image enhancement schemes; and outputting second media content generated by processing the first media content using the target image enhancement scheme.
[0004] In a second aspect of this disclosure, an apparatus for media processing is provided. The apparatus includes: a first generation module configured to generate descriptive text for first media content, the descriptive text describing image information of the first media content; a providing module configured to provide the descriptive text to a first model to generate multiple image enhancement schemes corresponding to the first media content, wherein each image enhancement scheme includes at least one image processing step performed using a preset image processing tool; a first determining module configured to determine a target image enhancement scheme from the multiple image enhancement schemes based on first evaluation information for the multiple image enhancement schemes; and an output module configured to output second media content generated by processing the first media content using the target image enhancement scheme.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method according to a first aspect of this disclosure.
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0011] Figure 2 A flowchart of a media processing procedure according to some embodiments of the present disclosure is shown;
[0012] Figure 3 An example filtering flowchart of an image enhancement scheme according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A flowchart illustrating the process for determining the success rate of an image enhancement scheme according to some embodiments of the present disclosure is shown.
[0014] Figure 5 An example flowchart of media processing according to some embodiments of the present disclosure is shown;
[0015] Figure 6 A schematic structural block diagram of an apparatus for media processing according to certain embodiments of the present disclosure is shown;
[0016] Figure 7 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0020] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0021] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0022] Traditionally, current image enhancement technologies have formed a processing system centered on modular atomic capabilities. Mainstream solutions encapsulate basic algorithms such as denoising, deblurring, super-resolution, and color correction into independent modules, and then combine and call these independent modules based on the image characteristics of the media content to perform image enhancement processing on the media content.
[0023] In related technologies, these independent modules can be combined and invoked based on prior knowledge. However, due to the increase in the number of atomic capabilities, and the excessive reliance on prior knowledge for the combination and invocation of atomic capabilities, determining the combination and invocation strategy becomes somewhat difficult, and this approach is unlikely to cover all image enhancement scenarios.
[0024] Of course, there are also ways to use models to predict image enhancement schemes in related technologies. If the image enhancement effect is not good when the image enhancement scheme is executed, the image enhancement scheme can be adjusted. However, the related technologies can only ensure that the image enhancement effect of the adjusted image enhancement scheme is relatively stronger than that of the original image enhancement scheme after processing the media content. However, the actual effect of image enhancement cannot be guaranteed.
[0025] Embodiments of this disclosure propose a media processing scheme. According to this scheme, descriptive text for first media content can be generated, the descriptive text describing the image information of the first media content. Further, the descriptive text can be provided to a first model to generate multiple image quality enhancement schemes corresponding to the first media content, wherein each image quality enhancement scheme includes at least one image processing step performed using a preset image processing tool. Further, a target image quality enhancement scheme can be determined from the multiple image quality enhancement schemes based on first evaluation information for the multiple image quality enhancement schemes. Further, second media content generated by processing the first media content using the target image quality enhancement scheme can be output.
[0026] Based on this approach, embodiments of this disclosure can determine a target image enhancement scheme with better image enhancement effect based on the evaluation information of multiple image enhancement schemes, thereby enabling the processing of the first media content based on the better image enhancement scheme and effectively ensuring the visual quality of the generated second media content.
[0027] Example Environment
[0028] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include electronic device 110.
[0029] In this example environment 100, the electronic device 110 can generate descriptive text for first media content, the descriptive text being used to describe the screen information of the first media content; provide the descriptive text to the first model 120 to generate multiple image quality enhancement schemes corresponding to the first media content, wherein the image quality enhancement scheme includes at least one image processing step performed using a preset image processing tool; determine a target image quality enhancement scheme from the multiple image quality enhancement schemes based on first evaluation information for the multiple image quality enhancement schemes; and output second media content generated by processing the first media content using the target image quality enhancement scheme.
[0030] In some embodiments, the first model 120 can be a model deployed on an electronic device 110, or it can be a model deployed on other devices. The first model 120 can be any suitable machine learning model, such as a language model.
[0031] In some embodiments, the electronic device 110 may be any type of mobile terminal, fixed terminal, or portable terminal equipped with a display device, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 may also support any type of interface for the target user (such as "wearable" circuitry).
[0032] Electronic device 110 can also be a server, which can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Electronic device 110 may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, etc.
[0033] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0034] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0035] Example process
[0036] Figure 2 A flowchart of a media processing procedure 200 according to some embodiments of the present disclosure is shown. Procedure 200 can be implemented at electronic device 110. Reference is made below. Figure 1 Describe the process 200.
[0037] In box 210, electronic device 110 generates descriptive text for the first media content, which is used to describe the screen information of the first media content.
[0038] In some embodiments, the first media content can be user-generated content (UGC), such as content that users create, share, or publish. Of course, the first media content can also be any suitable media content obtained through other means, which will not be elaborated here.
[0039] In some embodiments, the first media content can be any type of content with visuals, such as images, videos, etc.
[0040] In some embodiments, the descriptive text can be any appropriate text used to describe the visual information of the first media content, such as, but not limited to, the low-level visual information and high-level visual information of the first media content.
[0041] In some embodiments, underlying visual information can indicate basic, underlying features of the first media content that reflect its physical properties or local structure. For example, underlying visual information may include, but is not limited to, whether the media content is distorted, and the type of distortion. Distortion types may include, but are not limited to, compression distortion, noise distortion, blur distortion, color and exposure distortion, etc.
[0042] In some embodiments, high-level visual information can indicate a semantic-level understanding of the first media content, involving cognitive information such as object recognition, scene classification, and logical reasoning. For example, high-level visual information may include whether the first media content includes a preset object. The preset object can be any suitable person or thing, such as a face, text, etc.
[0043] In box 220, electronic device 110 provides descriptive text to the first model to generate multiple image enhancement schemes corresponding to the first media content, wherein the image enhancement schemes include at least one image processing step performed using a preset image processing tool.
[0044] In some embodiments, the first model can be any suitable machine learning model, such as a language model.
[0045] In some embodiments, these multiple image enhancement schemes can be different schemes for improving the visual quality of the first media content, and each image enhancement scheme can indicate at least one image processing step. In some embodiments, each image enhancement scheme can also indicate the execution order corresponding to the at least one image processing step. For example, image enhancement scheme A can indicate three image processing steps: a denoising step, a decompression step, and a super-resolution step, and the execution order of these three image processing steps is, in order, denoising step, super-resolution step, and decompression step.
[0046] In some embodiments, the at least one image processing step is performed using a preset image tool. The preset image processing tool can correspond to any suitable tool, such as including, but not limited to, a denoising tool, a deblurring tool, a color correction tool, etc. For example, the denoising step can be performed using a denoising tool.
[0047] In some embodiments, the preset image tools may include, but are not limited to, generative image enhancement atoms and non-generative image enhancement atoms. A generative image enhancement atom refers to a tool that uses a predetermined generative model to enhance the image quality of media content. Generative image enhancement atoms can be associated with this generative model to improve the visual quality of media content in a generative or reconstructive manner.
[0048] As an example, generative image enhancement atoms may include, but are not limited to, generative video super-resolution and generative portrait enhancement.
[0049] In some embodiments, non-generative enhancement atoms refer to image enhancement tools that do not rely on generative models. These tools typically process media content directly to improve its visual quality.
[0050] As an example, non-generative enhancement atoms can include, but are not limited to, video denoising, video super-resolution, video deblurring, decompression distortion, color enhancement, text enhancement, and so on.
[0051] In order to improve the accuracy of the processing decisions output by the first model and improve the processing efficiency of the first model, in some embodiments, the electronic device 110 may provide descriptive text to the first model so that the first model can generate a variety of image enhancement schemes based on the first memory information.
[0052] In some embodiments, the first memory information can be learned experience and knowledge to assist the first model in quickly and accurately generating information for image enhancement schemes based on descriptive text.
[0053] In some embodiments, the first memory information may at least indicate a first image enhancement scheme that matches the first sample media content and a first sample description text corresponding to the first sample media content.
[0054] In some embodiments, the number of first image enhancement schemes matched with the first sample media content can be set as needed, and can be one or more. The first image enhancement scheme matched with the first sample media content can be a predetermined image enhancement scheme that can effectively improve the visual quality of the first sample media content.
[0055] The following explains the process of determining the first image enhancement scheme that matches the first sample media content.
[0056] In some embodiments, the electronic device 110 may process the first sample media content using a first set of image enhancement schemes to obtain a first set of media processing results. The first set of image enhancement schemes may be all implementable image enhancement schemes generated by iteratively combining all existing image enhancement atomic capabilities, and each image enhancement scheme may include at least one image enhancement atomic capability.
[0057] Furthermore, the electronic device 110 can generate second evaluation information for the first set of image enhancement schemes based on the first set of media processing results. The second evaluation information can indicate the image enhancement effect corresponding to the first set of media processing results, or indicate the image quality corresponding to the first set of media processing results, or indicate whether the semantic information corresponding to the first set of media processing results is consistent with the semantic information corresponding to the first sample media content, etc.
[0058] As an example, electronic device 110 can provide this first set of media processing results to a pre-trained evaluation model to obtain second evaluation information for this first set of image enhancement schemes.
[0059] In some embodiments, the pre-trained evaluation model can be a single evaluation model or multiple evaluation models. To improve the accuracy of the evaluation, in some embodiments, the electronic device 110 can generate a set of evaluation results for each image enhancement scheme in the first set of image enhancement schemes using multiple evaluation models. Further, the electronic device 110 can determine the second evaluation information of the image enhancement scheme based on this set of evaluation results. Specifically, for each image enhancement scheme in the first set of image enhancement schemes, the electronic device 110 can determine the sum of a set of evaluation results corresponding to that image enhancement scheme as the second evaluation information of that image enhancement scheme.
[0060] Furthermore, the electronic device 110 can determine a first image quality enhancement scheme from the first group of image quality enhancement schemes based on the second evaluation information. As an example, the electronic device 110 can sort these first group of image quality enhancement schemes according to the evaluation information from high to low to obtain a sorting result. Further, the electronic device 110 can determine a preset number of image quality enhancement schemes that rank highest in the sorting result as the first image quality enhancement scheme matching the content of the first sample media. The preset number can be set according to needs, such as 5, 10, etc.
[0061] In some embodiments, the first sample description text can be any appropriate text used to describe the image information of the first sample media content, such as including, but not limited to, the low-level visual information and high-level visual information of the first sample media content. For example, similar to the description text of the first media content, the first sample description text can also indicate whether the first sample media content is distorted, the type of distortion corresponding to the distortion, whether the first sample media content includes a preset object (such as a face or text), etc.
[0062] For example, if the sample description texts corresponding to both Video 1 and Video 2 indicate that the videos contain high noise and have low clarity, and the noise reduction step in the image enhancement scheme A matching Video 1 is before the clarity improvement step, and the noise reduction step in the image enhancement scheme B matching Video 2 is also before the clarity improvement step, then the first memory information can be "When a video contains both high noise and low clarity, it is best to perform noise reduction first and then improve clarity."
[0063] In some embodiments, the electronic device 110 may provide a first image enhancement scheme matching the first sample media content and the first sample descriptive text to a pre-trained model to obtain first memory information. This first memory information is the experience summarized by the pre-trained model based on the first image enhancement scheme matching the first sample media content and the first sample descriptive text. This pre-trained model can be any suitable machine learning model, such as a language model. This pre-trained model can be any suitable machine learning model with a large number of parameters, which can be configured to handle any suitable complex task.
[0064] Taking the server's training of the first model as an example, the following explains the process of training the first model.
[0065] The server can use the second set of image enhancement schemes to process the second sample media content to obtain the second set of media processing results. The second set of image enhancement schemes can be all implementable image enhancement schemes generated by iterating and combining all existing image enhancement atomic capabilities, and each image enhancement scheme can include at least one image enhancement atomic capability.
[0066] Furthermore, the server can generate third evaluation information based on the second set of media processing results for the second set of image enhancement schemes. The second set of media processing results can be a new set of media content generated by processing the second sample media content using the second set of image enhancement schemes. The third evaluation information can indicate the image enhancement effect corresponding to the second set of media processing results, or indicate the image quality corresponding to the second set of media processing results, or indicate whether the semantic information corresponding to the second set of media processing results is consistent with the semantic information corresponding to the second sample media content, etc.
[0067] Furthermore, the server can determine the second image enhancement scheme from the second group of image enhancement schemes based on the third evaluation information. As an example, the electronic device 110 can sort these second group of image enhancement schemes according to the third evaluation information from highest to lowest to obtain a sorting result. Further, the electronic device 110 can determine a preset number of image enhancement schemes that rank highest in the sorting result as the second image enhancement scheme. The preset number can be set according to needs, such as 5, 8, etc.
[0068] Figure 3 An example filtering flowchart of an image enhancement scheme according to some embodiments of the present disclosure is shown, now targeting Figure 3 Please provide an explanation.
[0069] like Figure 3 As an example, electronic device 110 can acquire video 301 (second sample media content). Furthermore, electronic device 110 can utilize, for example... Figure 3 The n image enhancement schemes (the second group of image enhancement schemes) shown in box 302 process video 301 to generate video 303-1, video 303-2, ..., video 303-n, with one image enhancement scheme corresponding to one video. These n image enhancement schemes can be generated by iterating through a predetermined set of image enhancement operators, which may include artifact removal, deblurring, noise reduction, etc. Each image enhancement scheme may include at least one image enhancement operator.
[0070] Furthermore, the electronic device 110 can use the first evaluation model 304-1 to determine the first scores corresponding to videos 303-1, 303-2, ..., 303-2 respectively. The electronic device 110 can also use the second evaluation model 304-2 to determine the second scores corresponding to videos 303-1, 303-2, ..., 303-2 respectively. Furthermore, the electronic device 110 can determine the final score corresponding to video 303-1 based on the sum of the first score and the second score corresponding to video 303-1. The method for determining the final scores corresponding to videos 303-2..., 303-n is the same as the method for determining the final score corresponding to video 303-1, and will not be elaborated here.
[0071] Furthermore, the electronic device 110 can determine the m videos with the highest scores based on the final scores corresponding to videos 303-1, 303-2, ..., 303-2. Furthermore, the electronic device 110 can determine the m image enhancement schemes that generated these m videos as the second image enhancement scheme.
[0072] It should be noted that the first image enhancement solution can also be based on Figure 3 The example filtering process described above will not be repeated here.
[0073] In some embodiments, the first set of image enhancement schemes may be the same as the second set of image enhancement schemes, and the corresponding first set of media processing results may be the same as the second set of media processing results. The second evaluation information and the third evaluation information may also be the same, and the first image enhancement scheme and the second image enhancement scheme may also be the same.
[0074] Furthermore, the server can train the first model based at least on the second image enhancement scheme and the second sample media content.
[0075] As an example, the server can provide the second sample media content to the second model to generate second sample descriptive text. The second sample descriptive text can be any appropriate text used to describe the visual information of the second sample media content, such as, but not limited to, the underlying visual information and high-level visual information of the second sample media content. For example, similar to the descriptive text of the first media content, the second sample descriptive text can also indicate whether the second sample media content is distorted, the type of distortion, whether the second sample media content includes a preset object (such as a face or text), etc.
[0076] Furthermore, the server can provide the first model with a second image enhancement scheme and a second sample description text to determine a third image enhancement scheme. As an example, the server can provide the first model with a second image enhancement scheme and a second sample description text so that the first model can utilize the first memory information to determine the third image enhancement scheme.
[0077] Furthermore, the server can train the first model based on the second and third image enhancement schemes. As an example, the server can determine the loss value based on the second and third image enhancement schemes and a predetermined loss function. Further, the server can train the first model based on the loss value until predetermined training completion conditions are met. These predetermined training completion conditions could be the loss value reaching a threshold, the training time reaching a threshold, etc., which will not be elaborated upon here.
[0078] In frame 230, electronic device 110 determines a target image enhancement scheme from among multiple image enhancement schemes based on first evaluation information for multiple image enhancement schemes.
[0079] In some embodiments, the first evaluation information may indicate the image enhancement effect of the multiple image enhancement schemes relative to the first media content, or indicate the image quality information of the media content obtained after processing the first media content based on the multiple image enhancement schemes, or indicate whether the semantic information of the media content obtained after processing the first media content based on the multiple image enhancement schemes is consistent with the semantic information corresponding to the first media content, etc.
[0080] In some embodiments, the electronic device 110 can determine first evaluation information based on the picture quality information of the media content generated by the corresponding picture quality enhancement scheme. Specifically, for each picture quality enhancement scheme, the higher the picture quality information of the media content generated based on that scheme, the higher the first evaluation information.
[0081] In other embodiments, the electronic device 110 may determine first evaluation information based on the semantic information of the media content generated by the corresponding image enhancement scheme. Specifically, for each image enhancement scheme, the more consistent the semantic information of the media content generated by that image enhancement scheme is with the semantic information of the first media content, the higher the first evaluation information.
[0082] In other embodiments, the electronic device 110 may determine first evaluation information based on the picture quality information of the media content generated by the corresponding picture quality enhancement scheme and the semantic information of the media content generated by the corresponding picture quality enhancement scheme.
[0083] In some embodiments, the electronic device 110 can process the first media content based on a variety of image enhancement schemes to obtain a variety of third media content.
[0084] Specifically, for each image enhancement scheme, the image enhancement scheme may include multiple image processing steps, and the execution order of these multiple image processing steps may be indicated. The electronic device 110 may, based on the execution order of these multiple image processing steps, sequentially use each preset image processing tool to perform image enhancement processing on the first media content, thereby obtaining the third media content generated based on the image enhancement scheme.
[0085] In some embodiments, the electronic device 110 can execute multiple image enhancement schemes to determine first evaluation information for the multiple image enhancement schemes. Specifically, the electronic device 110 can execute these multiple image enhancement schemes to perform image enhancement processing on the first media content to generate multiple media content. Further, the electronic device 110 can generate the first evaluation information for the multiple image enhancement schemes based on the image quality information and / or semantic information of the multiple media content.
[0086] To improve the accuracy of evaluating these multiple image enhancement schemes, in some embodiments, the electronic device 110 can utilize multiple evaluation models to generate multiple evaluation results corresponding to the third media content. These multiple evaluation models can correspond to different evaluation perspectives. For each item of third media content, the electronic device 110 can provide that third media content to each of the multiple evaluation models to obtain multiple evaluation results for that third media content, with one evaluation model outputting one evaluation result. These multiple evaluation models can be any suitable machine learning model, corresponding to different model structures and different model parameters.
[0087] Furthermore, the electronic device 110 can determine first evaluation information of the image enhancement scheme corresponding to the third media content based on multiple evaluation results. As an example, for each of the multiple third media contents, the electronic device 110 can determine the first evaluation information of the image enhancement scheme corresponding to that third media content based on the sum of the multiple evaluation results for that third media content. The image enhancement scheme corresponding to the third media content is the image enhancement scheme used to generate that third media content.
[0088] In some embodiments, the plurality of image enhancement schemes includes a fourth image enhancement scheme. This fourth image enhancement scheme can be any one of the plurality of image enhancement schemes that includes multiple processing steps.
[0089] In some embodiments, the electronic device 110 may execute a first processing step among a plurality of processing steps to determine intermediate media content. The first processing step can be any one of these plurality of processing steps. Specifically, if the first processing step is the first processing step among the plurality of processing steps, then the electronic device 110 executes this first processing step based on the first media content to obtain the intermediate media content. If the first processing step is not the first processing step among the plurality of processing steps, then the electronic device 110 executes this first processing step based on the target media content to obtain the intermediate media content. The target media content is the media content obtained by executing at least one processing step based on the first media content.
[0090] Furthermore, the electronic device 110 can adjust the fourth image enhancement scheme in response to the intermediate media content not meeting the preset conditions. The preset conditions can be any appropriate conditions, such as the image quality information meeting preset quality requirements and / or the semantic information of the intermediate media content having a similarity to the semantic information of the first media content greater than a threshold, etc.
[0091] As an example, electronic device 110 can remove the first processing step from the fourth image enhancement scheme.
[0092] For example, the fourth image enhancement scheme includes steps 1, 2, and 3. If the intermediate media content obtained after performing step 2 does not meet the preset conditions, the electronic device 110 can adjust the fourth image enhancement scheme, and the adjusted fourth image enhancement scheme does not include step 2. For example, the adjusted fourth image enhancement scheme may include steps 1, 4, and 3. Furthermore, the adjusted fourth image enhancement scheme may also include steps 7, 5, 6, etc.
[0093] Since the second processing step may have already been executed before the first step, to improve processing efficiency, the adjusted fourth image enhancement scheme can also retain the second processing step executed before the first processing step. For example, the fourth image enhancement scheme includes steps 1, 2, and 3. If the intermediate media content obtained after executing step 2 does not meet the preset conditions, and step 1 has already been executed before executing step 2, then the adjusted fourth image enhancement scheme can include steps 1, 4, and 5, where step 1 is retained.
[0094] Furthermore, the electronic device 110 can execute the adjusted fourth image enhancement scheme. Specifically, in response to the retention of the second processing step in the fourth image enhancement scheme, the electronic device 110 can continue to execute other processing steps in the fourth image enhancement scheme based on the media content obtained after executing the second processing step. The execution order of these other processing steps is the step following the second processing step included in the adjusted fourth image enhancement scheme.
[0095] In some embodiments, to improve the accuracy of processing decisions output by the first model and increase the processing efficiency of the first model, the electronic device 110 may also update the second memory information associated with the first model based on the execution result of at least one of the processing steps in the fourth image enhancement scheme. The execution result may indicate whether the intermediate media content obtained by performing the corresponding processing step meets preset conditions. The second memory information may be experience and knowledge associated with the first model, and may be reference information used to assist the first model in generating the image enhancement scheme. The first memory information and the second memory information may be the same memory information or different memory information.
[0096] As an example, the electronic device 110 can determine the success rate based on the execution result of at least one of the processing steps in the fourth image enhancement scheme. This success rate can be the success rate corresponding to the fourth image enhancement scheme or the success rate of the at least one processing step.
[0097] As an example, electronic device 110 can provide this success rate to a trained model to obtain additional memory information. Furthermore, electronic device 110 can update second memory information associated with the first model based on the additional memory information. This pre-trained model can be any suitable machine learning model, such as a large language model.
[0098] As another example, electronic device 110 can provide the trained model with the success rate, the fourth image enhancement scheme, and the descriptive text of the first media content to obtain additional memory information. Furthermore, electronic device 110 can update the second memory information associated with the first model based on the additional memory information.
[0099] Figure 4 A flowchart illustrating the process for determining the success rate of image enhancement schemes according to some embodiments of this disclosure is shown. Now, regarding... Figure 4 Please provide an explanation.
[0100] like Figure 4As shown, the electronic device 110 can generate descriptive text 402 for video 401. The descriptive text 402 can be used to indicate whether video 401 is distorted, the type of distortion, whether it includes faces, whether it includes text, etc. Furthermore, the electronic device 110 can generate image enhancement schemes 403-1, 403-2, and 403-3 based on the descriptive text 402. For example... Figure 4 As shown, image enhancement scheme 403-1 may include two image processing steps: noise reduction and deblurring; image enhancement scheme 403-2 may include two image processing steps: artifact removal and deblurring; and image enhancement scheme 403-3 may include three image processing steps: artifact removal, deblurring, and face enhancement. Furthermore, electronic device 110 may process video 401 in parallel using image enhancement schemes 403-1, 403-2, and 403-3.
[0101] Specifically, when electronic device 110 processes video 401 using image enhancement scheme 403-1, it first denoises video 401. Further, electronic device 110 can determine whether denoising of video 401 is successful. If the resulting video 404-1 (not shown in the figure) after denoising of video 401 meets preset conditions, then denoising is successful; otherwise, denoising of video 401 is determined to have failed. Further, in response to successful denoising of video 401, electronic device 110 can perform deblurring processing on video 404-1. Further, electronic device 110 can determine whether deblurring of video 404-1 is successful. If the resulting video 404-2 (not shown in the figure) after deblurring of video 404-1 meets preset conditions, then deblurring is successful; otherwise, deblurring of video 404-2 fails. Figure 4 As an example, both image processing steps in image enhancement scheme 403-1 were executed successfully. Furthermore, electronic device 110 can determine that the success rate of image enhancement scheme 403-1 is 100%.
[0102] Specifically, when electronic device 110 processes video 401 using image enhancement scheme 403-2, it first performs artifact removal processing on video 401. Further, electronic device 110 can determine whether artifact removal on video 401 is successful. If video 404-3 (not shown in the figure) obtained after artifact removal processing meets preset conditions, it indicates successful artifact removal; otherwise, artifact removal on video 401 is considered a failure. Further, in response to successful artifact removal on video 401, electronic device 110 can perform deblurring processing on video 404-3. Further, electronic device 110 can determine whether deblurring of video 404-3 is successful. If video 404-4 (not shown in the figure) obtained after deblurring video 404-3 meets preset conditions, it indicates successful deblurring; otherwise, deblurring of video 404-3 is considered a failure. Figure 4 As an example, the artifact removal step in image enhancement scheme 403-2 was successfully executed, but the deblurring step was not. Further, the electronic device 110 can reschedule and adjust image enhancement scheme 403-2 to obtain an adjusted image enhancement scheme. This adjusted scheme, compared to the original image enhancement scheme 403-2, replaces the deblurring operation with a denoising operation. Further, the electronic device 110 can perform denoising processing on video 404-3. Further, the electronic device 110 can determine whether denoising of video 404-3 was successful. If video 404-5 (not shown in the figure) obtained after denoising video 404-3 meets preset conditions, then denoising is successful; otherwise, denoising of video 404-3 is determined to have failed. Figure 4 As an example, the image processing steps of artifact removal and noise reduction were successfully performed, but the image processing step of deblurring was not successfully performed. Furthermore, the electronic device 110 can determine that the success rate of the image enhancement scheme 403-2 is 50%.
[0103] It should be noted that the process by which electronic device 110 processes video 401 using image enhancement scheme 403-3 is the same as the process by which it processes video 401 using image enhancement scheme 403-1 and image enhancement scheme 403-2, and will not be elaborated upon here. It should also be noted that during image enhancement processing, the more times a particular image enhancement scheme's image processing steps fail, the lower its success rate.
[0104] In frame 240, electronic device 110 outputs second media content generated by processing the first media content using a target image enhancement scheme.
[0105] In some embodiments, the second media content can be any type of content with visuals, such as images, videos, etc.
[0106] In some embodiments, if the first evaluation information is determined by executing multiple image enhancement schemes, the electronic device 110 can directly determine the media content obtained by executing the target image enhancement scheme with the highest first evaluation information as the second media content and output it.
[0107] Figure 5 Example flowcharts of media processing according to some embodiments of the present disclosure are shown, now for... Figure 5 Please provide an explanation.
[0108] The electronic device 110 can perform the image quality understanding operation shown in box 510, that is, use a model to perform image quality understanding on the media content to be processed, so as to identify the image quality information and semantic information of the media content to be processed. Further, the electronic device 110 can use a first model, based on the first memory information shown in box 520, and based on the image quality information and semantic information determined in box 510, to perform the prediction of the image quality enhancement scheme shown in box 530. The first memory information shown in box 520 can indicate the knowledge or experience corresponding to each scenario, for example, it can indicate... Figure 5 The examples shown are experience 501 in scenario 1, experience 502 in scenario 2, and experience 503 in scenario 3.
[0109] For example, electronic device 110 can initially determine that the media content to be processed can be processed based on image enhancement scheme 1, image enhancement scheme 2, and image enhancement scheme 3. Further, electronic device 110 can use image enhancement scheme 1, image enhancement scheme 2, and image enhancement scheme 3 to process this media content to generate media content A, media content B, and media content C (not shown in the figure).
[0110] Furthermore, the electronic device 110 can determine the image quality information and semantic information corresponding to media content A, media content B, and media content C based on a multi-dimensional evaluation system (such as using multiple evaluation models). Further, the electronic device 110 can compare the media content to be processed with media content A, media content B, and media content C to evaluate image quality enhancement scheme 1, image quality enhancement scheme 2, and image quality enhancement scheme 3. Specifically, the electronic device 110 can compare the processed media content (media content A, media content B, and media content C) with the media content to be processed to determine whether the image quality enhancement effect is greater than a threshold and / or whether the difference in semantic information is less than a threshold, thereby evaluating image quality enhancement scheme 1, image quality enhancement scheme 2, and image quality enhancement scheme 3. The better the image quality enhancement effect, the higher the evaluation; and the smaller the difference in semantic information, the higher the evaluation.
[0111] Furthermore, the electronic device 110 can determine the optimal image enhancement scheme based on the evaluation of image enhancement scheme 1, image enhancement scheme 2, and image enhancement scheme 3. Furthermore, the electronic device 110 can process the media content to be processed using the preset image tools shown in box 540 based on the optimal image enhancement scheme to obtain processed media content. Furthermore, the electronic device 110 can output the processed media content.
[0112] In some embodiments, the preset image tools may include, but are not limited to, generative image enhancement atoms and non-generative image enhancement atoms. Generative image enhancement atoms refer to tools that use a predetermined generative model to enhance the image quality of media content. Generative image enhancement atoms can be associated with this generative model to improve the visual quality of the media content in a generative or reconstruction-based manner. As an example, generative image enhancement atoms may include, but are not limited to, generative video super-resolution and generative portrait enhancement. In some embodiments, non-generative enhancement atoms refer to image enhancement tools that do not rely on a generative model. These tools typically process the media content directly to improve its visual quality.
[0113] As an example, non-generative enhancement atoms can include, but are not limited to, video denoising, video super-resolution, video deblurring, decompression distortion, color enhancement, text enhancement, and so on.
[0114] Based on this approach, embodiments of this disclosure can determine a target image enhancement scheme with better image enhancement effect based on the evaluation information of multiple image enhancement schemes, thereby enabling the processing of the first media content based on the better image enhancement scheme and effectively ensuring the visual quality of the generated second media content.
[0115] Example devices and equipment
[0116] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 6 A schematic structural block diagram of an apparatus 600 for media processing according to certain embodiments of the present disclosure is shown. The apparatus 600 may be implemented as or included in the electronic device 110 discussed above. The various modules / components in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0117] like Figure 6 As shown, the device 600 includes a first generation module 610 configured to generate descriptive text for first media content, the descriptive text describing the image information of the first media content; a providing module 620 configured to provide the descriptive text to a first model to generate multiple image quality enhancement schemes corresponding to the first media content, wherein the image quality enhancement scheme includes at least one image processing step performed using a preset image processing tool; a first determining module 630 configured to determine a target image quality enhancement scheme from the multiple image quality enhancement schemes based on first evaluation information for the multiple image quality enhancement schemes; and an output module 640 configured to output second media content generated by processing the first media content using the target image quality enhancement scheme.
[0118] In some embodiments, the providing module 620 is further configured to: provide descriptive text to the first model so that the first model generates multiple image enhancement schemes based on the first memory information, wherein the first memory information at least indicates the first image enhancement scheme matching the first sample media content and the first sample descriptive text corresponding to the sample media content.
[0119] In some embodiments, a first image enhancement scheme is determined based on the following process: processing a first sample media content using a first set of image enhancement schemes to obtain a first set of media processing results; generating second evaluation information for the first set of image enhancement schemes based on the first set of media processing results; and determining a first image enhancement scheme from the first set of image enhancement schemes based on the second evaluation information.
[0120] In some embodiments, the first model is trained based on the following process: processing the second sample media content using a second set of image enhancement schemes to obtain the second set of media processing results; generating third evaluation information for the second set of image enhancement schemes based on the second set of media processing results; determining a second image enhancement scheme from the second set of image enhancement schemes based on the third evaluation information; and training the first model based at least on the second image enhancement scheme and the second sample media content.
[0121] In some embodiments, training a first model based at least on a second image enhancement scheme and second sample media content includes: providing the second sample media content to the second model to generate second sample descriptive text; providing the second image enhancement scheme and the second sample descriptive text to the first model to determine a third image enhancement scheme; and training the first model based on the second image enhancement scheme and the third image enhancement scheme.
[0122] In some embodiments, the apparatus 600 further includes a processing module configured to: process first media content based on multiple image enhancement schemes to obtain multiple third media content; a second generation module configured to: generate multiple evaluation results corresponding to the third media content using multiple evaluation models; and a second determination module configured to: determine first evaluation information of the image enhancement scheme corresponding to the third media content based on the multiple evaluation results.
[0123] In some embodiments, the apparatus 600 further includes an execution module configured to execute multiple image enhancement schemes to determine first evaluation information for the multiple image enhancement schemes.
[0124] In some embodiments, the multiple image enhancement schemes include a fourth image enhancement scheme, the fourth image enhancement scheme including multiple processing steps, and performing the fourth image enhancement scheme includes: performing a first processing step among the multiple processing steps to determine intermediate media content; adjusting the fourth image enhancement scheme in response to the intermediate media content not meeting preset conditions; and performing the adjusted fourth image enhancement scheme.
[0125] In some embodiments, adjusting the fourth image enhancement scheme includes removing the first processing step from the fourth image enhancement scheme.
[0126] In some embodiments, the adjusted fourth image enhancement scheme is retained in the second processing step performed before the first processing step.
[0127] In some embodiments, the apparatus 600 further includes an update module configured to update the second memory information associated with the first model based on the execution result of at least one of the plurality of processing steps.
[0128] In some embodiments, the first evaluation information is determined based on at least one of the following: picture quality information of media content generated based on the corresponding picture quality enhancement scheme; semantic information of media content generated based on the corresponding picture quality enhancement scheme.
[0129] The units included in device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 600 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0130] Figure 7 A block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The electronic device 700 shown can be used to achieve Figure 1 The electronic device 110 shown.
[0131] like Figure 7 As shown, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors 710 or processing units, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processor 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 700.
[0132] Electronic device 700 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (PROM), flash memory), or some combination thereof. Storage device 730 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 700.
[0133] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 7 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0134] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0135] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0136] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0137] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0138] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0139] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0141] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A media processing method, comprising: Generate descriptive text for the first media content, wherein the descriptive text is used to describe the image information of the first media content; The descriptive text is provided to the first model to generate multiple image quality enhancement schemes corresponding to the first media content, wherein the image quality enhancement scheme includes at least one image processing step performed using a preset image processing tool; Based on the first evaluation information for the various image enhancement schemes, a target image enhancement scheme is determined from the various image enhancement schemes; as well as The output is the second media content generated by processing the first media content using the target image quality enhancement scheme.
2. The method according to claim 1, wherein providing the descriptive text to the first model to generate multiple image enhancement schemes corresponding to the first media content includes: The first model is provided with the description text so that the first model generates the multiple image enhancement schemes based on the first memory information, wherein the first memory information at least indicates the first image enhancement scheme that matches the first sample media content and the first sample description text corresponding to the first sample media content.
3. The method of claim 2, wherein the first image enhancement scheme is determined based on the following process: The first sample media content is processed using the first set of image quality enhancement schemes to obtain the first set of media processing results; Based on the first set of media processing results, a second evaluation information is generated for the first set of image quality enhancement schemes; as well as Based on the second evaluation information, the first image quality enhancement scheme is determined from the first group of image quality enhancement schemes.
4. The method of claim 1, wherein the first model is trained based on the following process: The second set of image enhancement schemes is used to process the second sample media content to obtain the second set of media processing results. Based on the second set of media processing results, a third evaluation information is generated for the second set of image enhancement schemes; Based on the third evaluation information, a second image enhancement scheme is determined from the second group of image enhancement schemes; as well as The first model is trained based at least on the second image enhancement scheme and the second sample media content.
5. The method of claim 4, wherein training the first model, based at least on the second image enhancement scheme and the second sample media content, comprises: The second sample media content is provided to the second model to generate the second sample description text; The first model is provided with the second image enhancement scheme and the second sample description text to determine the third image enhancement scheme; as well as The first model is trained based on the second image enhancement scheme and the third image enhancement scheme.
6. The method according to claim 1, further comprising: The first media content is processed based on the aforementioned multiple image enhancement schemes to obtain multiple third media contents; Multiple evaluation models are used to generate multiple evaluation results corresponding to third-media content; Based on the multiple evaluation results, the first evaluation information of the image quality enhancement scheme corresponding to the third media content is determined.
7. The method according to claim 1, further comprising: The various image enhancement schemes are executed to determine the first evaluation information of the various image enhancement schemes.
8. The method of claim 7, wherein the plurality of image enhancement schemes includes a fourth image enhancement scheme, the fourth image enhancement scheme including a plurality of processing steps, and performing the fourth image enhancement scheme includes: Perform the first of the plurality of processing steps to determine intermediate media content; In response to the intermediate media content not meeting the preset conditions, the fourth image quality enhancement scheme is adjusted; as well as Implement the adjusted fourth image enhancement scheme.
9. The method according to claim 8, wherein adjusting the fourth image enhancement scheme comprises: Remove the first processing step from the fourth image enhancement scheme.
10. The method of claim 8, wherein the adjusted fourth image enhancement scheme is retained in the second processing step performed prior to the first processing step.
11. The method of claim 8, further comprising: Based on the execution result of at least one of the plurality of processing steps, the second memory information associated with the first model is updated.
12. The method of claim 1, wherein the first evaluation information is determined based on at least one of the following: Image quality information of media content generated based on the corresponding image enhancement scheme; Semantic information of media content generated based on corresponding image enhancement solutions.
13. An apparatus for media processing, comprising: The first generation module is configured to generate descriptive text for the first media content, wherein the descriptive text is used to describe the image information of the first media content. A module is configured to provide the descriptive text to a first model to generate multiple image quality enhancement schemes corresponding to the first media content, wherein the image quality enhancement scheme includes at least one image processing step performed using a preset image processing tool; The first determining module is configured to determine a target image quality enhancement scheme from the multiple image quality enhancement schemes based on first evaluation information for the multiple image quality enhancement schemes; as well as The output module is configured to output second media content generated by processing the first media content using the target image quality enhancement scheme.
14. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 12 when executed by the at least one processor.
15. A computer-readable storage medium having stored thereon computer-executable instructions that, when executed by a processor, implement the method according to any one of claims 1 to 12.
16. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 12.