Media content processing method, apparatus, device, medium, and program product

By acquiring media content and adjusting the second machine learning model using the trained machine learning model, the problem of poor image quality consistency in media content processing is solved, and efficient and stable image generation is achieved.

CN122492872APending Publication Date: 2026-07-31BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-04-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to consistently produce image results that meet processing requirements in media content processing, resulting in poor image quality consistency.

Method used

By acquiring media content containing text and images, a second machine learning model is adjusted using a trained machine learning model, internalizing the constraints of effect verification, and generating images that meet the processing requirements.

Benefits of technology

It improves the efficiency of media content processing and image processing effects, ensuring the stability and consistency of output images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492872A_ABST
    Figure CN122492872A_ABST
Patent Text Reader

Abstract

A media content processing method, apparatus, device, medium, and program product are disclosed, relating to the field of computer application technology. The method includes: acquiring first media content, which includes first text and a first image, the first text indicating a processing method for the first image; providing the first media content to a first machine learning model to obtain a second image, the first machine learning model adjusting the second machine learning model based on first detection results of multiple fourth images, the fourth image being obtained by the second machine learning model based on the second media content, the second media content including a third image, the first detection results being associated with second text, the second text including multiple first strategies, the first strategies being used to evaluate the effect of the third image. The provided solution effectively overcomes the defect of poor image quality consistency in related technologies, and can stably generate second images that meet processing requirements when faced with different first media content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer processing technology, and in particular to a media content processing method, apparatus, device, medium, and program product. Background Technology

[0002] In advertising design, film and television production, e-commerce display, and virtual reality, media content often needs to be edited, composited, or enhanced to obtain new images that meet subsequent usage requirements.

[0003] However, when processing media content to obtain new images, the relevant technologies are insufficient in terms of the stability of the output effect, making it difficult to consistently obtain image results that meet the processing requirements. Summary of the Invention

[0004] This invention provides a media content processing method, apparatus, device, medium, and program product that enables efficient and high-quality processing of media content.

[0005] In one scenario, this paper provides a media content processing method, which includes:

[0006] Acquire first media content, the first media content including first text and first image, the first text being used to indicate the processing method of the first image;

[0007] The first media content is provided to the first machine learning model to obtain the second image. The first machine learning model is adjusted based on the first detection results of multiple fourth images to obtain the second machine learning model. The fourth image is obtained by the second machine learning model based on the second media content. The second media content includes a third image. The first detection results are associated with second text. The second text includes multiple first strategies. The first strategies are used to test the effect of the third image.

[0008] In one instance, this document also provides a media content processing apparatus, the apparatus comprising:

[0009] The first module is used to acquire first media content, which includes first text and a first image, wherein the first text is used to indicate the processing method of the first image;

[0010] The second module is used to provide the first media content to the first machine learning model to obtain the second image. The first machine learning model adjusts the second machine learning model based on the first detection results of multiple fourth images. The fourth image is obtained by the second machine learning model based on the second media content. The second media content includes a third image. The first detection results are associated with second text. The second text includes multiple first strategies. The first strategies are used to test the effect of the third image.

[0011] In one instance, this document also provides an electronic device comprising:

[0012] One or more processors;

[0013] Storage device for storing one or more programs.

[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the media content processing method as described herein.

[0015] In one instance, this document also provides a storage medium containing computer-executable instructions that, when executed by a computer processor, are used to perform media content processing methods as described herein.

[0016] In another scenario, this document also provides a computer program product, including a computer program that, when executed by a processor, implements the media content processing methods as described herein.

[0017] In one scenario, the provided solution acquires first media content containing first text and a first image, where the first text explicitly indicates the processing method for the first image, providing clear and operable input for subsequent generation and helping to avoid output fluctuations caused by ambiguous processing requirements. In another scenario, the provided solution provides the aforementioned first media content to a first machine learning model, thereby obtaining a second image. During the training phase, the first machine learning model has been adjusted based on multiple fourth images and their corresponding first detection results. The first detection results are associated with multiple first strategies used to verify the effect of the third image. This allows the first machine learning model to internalize the constraint ability to verify the effect of the output image during the learning process. Therefore, when faced with different first media content during the inference phase, the first machine learning model can stably generate a second image that meets the processing requirements, effectively overcoming the defect of poor image quality consistency in related technologies, improving the processing efficiency of media content, and enhancing the image processing effect. Attached Figure Description

[0018] The above and other features, advantages, and aspects of the embodiments described herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0019] Figure 1 This is a schematic diagram of the structure of a media content processing system under one scenario.

[0020] Figure 2 This is a flowchart illustrating a media content processing method under one specific scenario.

[0021] Figure 3 This is a flowchart illustrating a media content processing method under another scenario.

[0022] Figure 4 This is a flowchart illustrating the positive and negative sample comparison process used in another media content processing method.

[0023] Figure 5 This is a schematic diagram illustrating the processing flow of image detection results used in another media content processing method.

[0024] Figure 6 This is a schematic diagram of a media content processing device in one scenario.

[0025] Figure 7 This is a schematic diagram of the structure of an electronic device used in one scenario. Detailed Implementation

[0026] The embodiments will now be described in more detail with reference to the accompanying drawings. While some embodiments are shown in the drawings, it should be understood that the technical solutions can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the technical solutions herein. It should be understood that the illustrated drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the technical solutions.

[0027] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.

[0028] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one situation" means "at least one situation"; the term "another situation" means "at least one additional situation"; the term "some situations" means "at least some situations". Definitions of other terms will be given in the following description.

[0029] It should be noted that the concepts of "first" and "second" mentioned are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.

[0030] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0031] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0032] It is understood that before using the technical solutions disclosed in the various embodiments of this document, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this document in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0033] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations described herein, based on the prompt message.

[0034] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0035] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.

[0036] It is understood that the data involved in the technical solutions in this article (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0037] In some cases, the provided solution can be applied to Figure 1 The media content processing system shown may include a client 101 and a server 102. The client 101 may include, but is not limited to, browsers, applications (Apps), HyperText Markup Language (HTML) applications, lightweight applications (also known as mini-programs), or cloud applications. The client 101 may be deployed on an electronic device and relies on the operation of that device or certain applications on the device to implement its functions. The electronic device may be, for example, a device with a display screen that supports information browsing, such as a smartphone, tablet, personal computer, or other client terminal. For ease of understanding, Figure 1 The client is primarily represented in the form of a device. Other types of applications can also be configured on the electronic device, such as media content publishing applications, session applications, etc. Server 102 can be one or more servers providing various services. That is, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; furthermore, it can be a server for a distributed system, a server integrating blockchain technology, a cloud server, or an intelligent cloud computing server or intelligent cloud host deployed with machine learning models, etc.

[0038] The media content processing method described herein allows interaction between client 101 and server 102, such as receiving or sending messages. For example, in this paper, server 102 can receive a media content processing request sent by client 101 based on an information carrier, obtain first media content, then provide the first media content to a first machine learning model to obtain a second image, and send the second image to client 101 for display on the display interface.

[0039] The first media content includes first text and first image, wherein the first text is used to indicate the processing method of the first image; the first machine learning model is adjusted to the second machine learning model based on the first detection results of multiple fourth images, wherein the fourth images are obtained by the second machine learning model based on the second media content, the second media content includes a third image, the first detection results are associated with the second text, the second text includes multiple first strategies, and the first strategies are used to test the effect of the third image.

[0040] It should be noted that the media content processing method can be executed by client 101, or by client 101 and server 102, with different functional parts of the corresponding media content processing device deployed on client 101 and server 102 respectively; wherein, the media content processing request receiving module and media content display module of the media content processing device can be deployed on client 101, and the first module and the second module can be deployed on server 102. Client 101 and server 102 achieve data interaction and functional collaboration through network communication. It should be understood that... Figure 1 The number of clients and servers shown is for illustrative purposes only. Any number of clients and servers can be configured to meet specific implementation requirements.

[0041] In some cases, client 101 and server 102 can implement the provided information interaction methods and payment information processing methods by running computer programs. For example, computer executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as video APPs or instant messaging APPs; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In short, the aforementioned computer executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0042] Figure 2 This is a flowchart illustrating a media content processing method for one scenario. This method is applicable to image editing and content generation scenarios, particularly those requiring precise image processing based on text descriptions and demanding stable and reliable output results, such as automated advertising material production, batch e-commerce main image adjustment, and film post-production special effects compositing. This media content processing method can be executed by a media content processing device, which can be implemented through software and / or hardware, optionally through a computer program. Figure 2 As shown, the media content processing method may include:

[0043] S210. Obtain first media content, which includes first text and first image, wherein the first text is used to indicate the processing method of the first image.

[0044] Media content refers to digital data that includes one or more information formats such as text, images, audio, and video. The first media content specifically refers to a set of data used as input during the inference phase, comprising two parts: the first text and the first image. The first text describes the desired operation to be performed on the first image. For example, the first text might be "Change the red car in the picture to blue, while keeping the wheels unchanged," and the first image might be a photograph of a red car on a road.

[0045] The first text can be understood as textual information composed of natural or structured language. In one scenario, the first text can be used to explicitly indicate the processing to be performed on the first image, which may include one or more operations such as color adjustment, object replacement, style transfer, or local modification. The first text is the carrier of processing instructions, and its semantic content directly determines the expected effect of the second image. For example, the first text could be "blur the background of the person," "increase the image brightness to a medium level," "remove the vehicle in the lower left corner of the image," or "change the background of the trash can to a city street."

[0046] The first image can be understood as digitized static visual information. The first image is the object that needs to be processed based on the first text. The content, size, color, and other attributes of the first image will affect the final output. The first image can be one or more of the following: a photograph, a rendering, a hand-drawn sketch, or a scanned copy. For example, the first image could be an indoor photograph containing multiple people, a product image on a white background, or a scanned landscape postcard.

[0047] In one scenario, acquiring first media content includes: receiving the first media content via an interactive interface associated with a first machine learning model.

[0048] In one scenario, the interactive interface provides text input boxes and image upload areas. Users type text in the text boxes and select images by dragging or browsing. After the user clicks the "Submit" button, the media content processing system can use the text and images obtained from the interactive interface as the primary media content.

[0049] In one scenario, a voice input control is provided in the interactive interface. The user clicks the voice input control and enters voice information. Further, the media content processing system can convert this voice information into text, using the converted text as the first text.

[0050] In one scenario, a camera control is provided in the user interface. The user can trigger the camera control to capture an image. This captured image can then be used as the first image.

[0051] In some cases, acquiring the first media content may include, but is not limited to, at least one of the following: acquiring the first media content acquired through a preset media content acquisition control; acquiring the first media content uploaded through a preset media content upload control; receiving the first media content transmitted through a preset data interface; acquiring the first media content generated by a preset model; extracting the first media content from the target content; acquiring preset media content as the first media content, etc.

[0052] S220. Provide the first media content to the first machine learning model to obtain the second image. The first machine learning model adjusts the second machine learning model based on the first detection results of multiple fourth images. The fourth image is obtained by the second machine learning model based on the second media content. The second media content includes the third image. The first detection results are associated with the second text. The second text includes multiple first strategies. The first strategies are used to test the effect of the third image.

[0053] The first machine learning model is a trained model that accepts multimodal input and outputs images. The first machine learning model is obtained by adjusting a second machine learning model. The second machine learning model can be a pre-trained model capable of processing multimodal data. The first machine learning model can be a generative model, such as a model that generates images based on text and images. The parameter adjustment of the first machine learning model is based on at least the first detection results of multiple fourth images.

[0054] The second image is the output image of the first machine learning model based on the input first media content. At least some of the display information in the second image differs from that of the first image. This display information may include, but is not limited to, at least one of the following: display color, display size, display position of at least some content in the image, and display style of at least some content in the image. The second image may be obtained by editing at least some content in the first image based on the first text.

[0055] The fourth image can be understood as the output image generated by the second machine learning model based on the second media content. Alternatively, the fourth image is the output image of the first machine learning model during the training phase. Multiple fourth images can be the processing results of multiple images corresponding to the same or different second media content. These fourth images cover various input conditions and model states in the training samples and can be used for subsequent detection and adjustment of the second machine learning model. For example, the second media content can be input into the first machine learning model to obtain one or more fourth images corresponding to the second media content.

[0056] The first detection result can be understood as a quantitative or qualitative conclusion obtained after examining the fourth image using multiple first strategies from the second text. The first detection result reflects the degree to which the fourth image conforms to the first strategies. The first detection result can be, for example, a set of Boolean values ​​(to characterize whether each rule passes or fails), a comprehensive score, or a report containing specific deviation values. For example, for a fourth image, the detection result is "0.85".

[0057] The second media content can be understood as the input data used during the training phase, and may include third text and a third image. The structure of the second media content is similar to that of the first media content, but it is used in the model training and tuning process. For example, the third text could be "Move the cup on the desktop in the picture to the left," and the third image could be a photo of an office desk.

[0058] The third text can be understood as text used during the training phase to instruct on how to process the third image. It serves the same purpose as the first text but is part of the training samples. For example, the third text could be something like "Add some accessories to the task in the image."

[0059] The third image can be understood as the image to be processed during the training phase. It serves the same purpose as the first image but is a training sample. For example, the third image could be a low-contrast landscape photograph.

[0060] The second text can be understood as a set of texts containing multiple first strategies, used to define the criteria for evaluating the effect of the fourth image. The first strategies in the second text can be quantitative (e.g., pixel error less than 5%) or qualitative (e.g., "facial expression unchanged"). Different first strategies can be used to evaluate the visual effect of the fourth image in different dimensions. For example, the second text might read: "Rule A: The sky color in the generated image should match the color specified in the third text; Rule B: The edges of foreground objects should be smooth; Rule C: The overall brightness change should not exceed 10% of the original image."

[0061] In one scenario, the second media content further includes a third text, which is obtained by providing the third text, the third image, and a fifth text to a sixth machine learning model to obtain the second text, wherein the fifth text is used to indicate the content to be included in the second text.

[0062] The sixth machine learning model can be understood as a model capable of automatically generating structured or semi-structured text based on input information. For example, the sixth machine learning model could be a text generation model (e.g., a visual language model). The task of the sixth machine learning model is to generate a second text that meets the requirements, based on the instructions of the fifth text and combining the third text and third image information from the second media content.

[0063] The fifth text can be a natural language instruction that tells the sixth machine learning model what types or aspects of content should be included when generating the second text.

[0064] In one scenario, the proposed solution can automatically generate the second text based on the fifth text and the second media content using a sixth machine learning model. This avoids the tediousness and instability of manually writing rules one by one. Furthermore, it can dynamically generate targeted verification rules based on the specific input content (the third text and the third image). This allows the first strategy to better align with the characteristics of the current processing task, improving the adaptability and accuracy of subsequent detection steps, and thus providing more precise feedback for adjusting the first machine learning model.

[0065] In one scenario, the fifth text includes at least one of the following: fifth media content and its corresponding multiple third strategies; and descriptive information about the strategy generation logic, wherein the strategy generation logic is a logic for generating multiple strategies based on the media content.

[0066] The fifth media content can be understood as multimodal data containing text and / or images. In some cases, the fifth media content can be provided as an example to the sixth machine learning model as a text-image pair, demonstrating how to inductively derive testing rules from the media content. For example, the fifth media content might include the text "Dye the hair of the person in the picture brown" and a before-and-after comparison image of the hair dyeing process.

[0067] Multiple third strategies can be understood as a set of pre-existing validation rules associated with the fifth media content. These third strategies can serve as example rules, guiding the sixth machine learning model to understand the format, granularity, or focus of the second text to be generated. For example, the third strategies corresponding to the above hair dyeing example are: "Rule A: The hair area color should be brown; Rule B: The facial features position offset is less than 3 pixels; Rule C: The background brightness change does not exceed 5%."

[0068] In one scenario, the text and images (or text descriptions of the images) in the fifth media content can be concatenated with multiple third strategies to form a structured example, which is then placed in the prompt words as a few-sample example. The sixth machine learning model, upon receiving the second media content for the current task, generates the corresponding second text by referencing the rule style and content dimensions of that example.

[0069] In one scenario, an example library can be constructed, from which a fifth piece of media content and its third strategy can be randomly selected. This third strategy is then used as a context-injected prompt word through retrieval enhancement. The sixth machine learning model first learns the strategy generation pattern from the examples, and then generates second text that conforms to a similar pattern for new second media content.

[0070] In one scenario, the sixth machine learning model can be provided with the fifth media content and the third strategy, and asked "How are these rules derived from the media content?" After the sixth machine learning model answers that it understands, the second media content and the phrase "Please generate the strategy according to the same logical strategy" from the fifth text can be provided to obtain the second text.

[0071] In one scenario, the proposed solution provides a sixth machine learning model with example examples containing media content and corresponding rules. This enables the sixth machine learning model to mimic the rule structure, granularity, and evaluation criteria in the examples, thereby generating high-quality second text that is homogeneous with the examples and adapted to the current task.

[0072] In this context, the strategy generation logic can be understood as a rule-construction method, rather than a specific rule example. The strategy generation logic can be, for example, a heuristic process, a conditional decision tree, or a derivation path based on first principles. The descriptive information of the strategy generation logic can describe the methodology or steps of "how to extract verification rules from given media content." This descriptive information can also serve as instructions to guide the sixth machine learning model to generate second text according to a specific strategy or framework. For example, the descriptive information of the strategy generation logic could be: "Please generate a strategy according to the following logical strategy: First, identify the object to be modified mentioned in the third text; second, identify the target attribute or operation specified in the third text; third, generate attribute change rules for the object to be modified; fourth, generate invariant rules for objects not mentioned; fifth, add basic visual quality rules (such as sharpness and integrity) to the overall image."

[0073] In one scenario, the proposed solution, by explicitly providing descriptive information about the strategy generation logic, makes the output of the sixth machine learning model interpretable and controllable. This ensures that the second text generated under different inputs follows a unified methodology, avoiding arbitrariness in rule content. Simultaneously, this approach allows users to flexibly customize rule types and focuses by adjusting the logic description, enhancing the generalization and adaptability of the technical solution.

[0074] In one scenario, the second text may be user-edited. For example, user-inputted second text may be received via an interactive interface.

[0075] The first strategy can be understood as the basic entries that constitute the second text, with each rule describing a testable condition. The first strategy can be used to detect the quality of the generated image, either as an image quality metric or as a constraint designed for the processing of the fourth image.

[0076] In one scenario, the first strategy includes at least one of the following: content processing principles and quality inspection principles.

[0077] Content processing principles can be understood as rules used to verify whether specific objects or regions in an image are correctly processed or preserved as expected. For example, these principles can be used to verify at least one of the following: content in the image that needs to remain unchanged, and content in the image that needs to be modified. When editing or generating a first image based on first text, content processing principles are used to automatically determine whether the generated image accurately executes the instructions on "which content should be changed and which content should remain unchanged." For example, verifying content in the image that needs to remain unchanged: if the first text requires "keeping the background unchanged," the content processing principles can verify whether the pixel difference between the background area in the generated image and the background of the original first image is less than a threshold. As another example, verifying content in the image that needs to be modified: if the first text requires "changing the red car to blue," the content processing principles can verify whether the main color of the car area in the generated image becomes blue, and whether the modification does not spread to non-car areas. In one scenario, the provided solution, by verifying the generated image through content processing principles, can accurately distinguish between areas that "should be changed" and areas that "should remain unchanged," avoiding the common model problem of all changes being made or incomplete changes, thereby improving the accuracy of the output image in executing text instructions.

[0078] In this context, quality inspection principles can be understood as rules used to inspect the overall visual performance and structural rationality of the generated image. In image editing or generation tasks, quality inspection principles can be used to detect whether the output image exhibits visual distortion, damage, or unreasonable phenomena. For example, the quality inspection principles are used to inspect at least one of the following: visual fidelity and visual integrity. Visual fidelity indicates the consistency of visual performance before and after image processing. Visual performance may include at least one of the following: color distribution, texture detail, and brightness contrast. Visual integrity indicates the structural and logical rationality of the image. Structural and logical rationality may include, for example, at least one of the following: whether object edges are continuous, whether spatial occlusion relationships are correct, and whether unrealistic breaks or redundancies occur. For example, it is required that the relative positions of the eyes, nose, and mouth in the generated facial image conform to anatomical proportions and have no missing or repetitive parts.

[0079] In some cases, the visual fidelity of an image can be obtained based on the peak signal-to-noise ratio between the fourth and third images or by learning the similarity of image patches.

[0080] In some cases, visual integrity detection methods may include at least one of the following: using discriminator feature distance, which is commonly used in adversarial generative networks, to detect whether there are broken edges or abnormal pixel blocks in the fourth image; or using a pre-trained scene graph model to check whether the spatial relationships between objects violate common sense, such as a cup floating above a table.

[0081] In some cases, the fourth image can be input into a trained quality regression network to obtain quantitative indicators of video fidelity and visual integrity. Then, visual fidelity and visual integrity can be obtained based on the quantitative indicators of visual fidelity and visual integrity, respectively.

[0082] The effect of the third image can be understood as the visual result of the fourth image generated after the third image has been processed by the first machine learning model. The first strategy can be used to verify whether the effect meets the expected standard. For example, if the third image is a red rose and the third text is "change the petal color to pink", the effect of the third image can be that the petals in the fourth image are pink and the shape is preserved.

[0083] In one scenario, the training process of the first machine learning model includes: acquiring multiple second media contents, which include third text and third images; using the second machine learning model to generate multiple fourth images based on each second media content; for each fourth image, using a trained "third machine learning model" to obtain its "first detection result", where the first detection result can be a numerical reward score; and based on these fourth images and their corresponding first detection results, adjusting the parameters of the second machine learning model using a reinforcement learning algorithm to obtain a more powerful first machine learning model.

[0084] In one scenario, the training process of the first machine learning model may include: providing the second media content to the second machine learning model to obtain multiple fourth images corresponding to the second media content; determining the difference information corresponding to each of the multiple fourth images based on the first detection results of the multiple fourth images; and adjusting the second machine learning model based on the difference information corresponding to the multiple images to obtain the first machine learning model.

[0085] In one scenario, the proposed solution obtains the first media content indicating the processing method and inputs it into the first machine learning model obtained through the training phase (generating the fourth image, determining the difference information, and adjusting the model accordingly), so that the final output second image can stably follow the preset rules, effectively solving the problem of unstable image generation effect in related technologies and improving the reliability and consistency of media content processing results.

[0086] The difference information can be understood as the comparative conclusions obtained by further processing based on the first detection result.

[0087] In one scenario, determining the difference information corresponding to each of the multiple fourth images based on the first detection results of the fourth images includes: comparing each fourth image with other randomly selected fourth images, and recording the difference information as a first value or a second value based on the quality of the first detection results of the current fourth image relative to the compared fourth images.

[0088] In one scenario, determining the difference information corresponding to each of the multiple fourth images based on the first detection results of the multiple fourth images may include: obtaining the first detection results of the multiple fourth images, averaging the multiple first detection results, and then subtracting the multiple first detection results from the average detection result to obtain the difference information.

[0089] In one scenario, the proposed solution, by transforming the absolute first detection result into a relative comparison result, can amplify the differences in performance between different generated images, providing a clearer direction for model tuning.

[0090] In one scenario, adjusting the second machine learning model based on the difference information corresponding to multiple images to obtain the first machine learning model includes: using the difference information as a reward signal (positive values ​​reward good results, negative values ​​penalize poor results), updating the parameters of the second machine learning model using a policy gradient method, and obtaining the first machine learning model after multiple iterations.

[0091] In one scenario, the second machine learning model is adjusted based on the difference information corresponding to multiple images to obtain the first machine learning model, including: taking the fourth image with positive difference information as a positive sample and the one with negative difference information as a negative sample, constructing a contrast loss function, and fine-tuning the model through gradient descent so that the model is more inclined to generate outputs of the positive sample type.

[0092] In one scenario, the second machine learning model is adjusted based on the difference information corresponding to multiple images to obtain the first machine learning model. This includes: assigning training weights to the second media content corresponding to each fourth image based on the absolute value of the difference information. For example, the weights of good samples are greater than 0 but less than 1, and the weights of poor samples are greater than 1. The second machine learning model is then retrained to make it more difficult for it to remember how the poor samples were generated.

[0093] In one scenario, the proposed solution utilizes discrepancy information to adjust the model, enabling it to distinguish between good and bad generated results and optimize towards a better outcome. The adjusted first machine learning model significantly improves consistency with the first strategy when generating images, thereby consistently outputting high-quality second images during the inference phase.

[0094] In one scenario, providing the first media content to a first machine learning model to obtain a second image may include: preprocessing the acquired first text and first image according to the input format required by the first machine learning model. This preprocessing may include, but is not limited to, at least one of the following: text segmentation and conversion into embedding vectors, image scaling and normalization of pixel values. Then, the forward inference function of the first machine learning model is called, passing the first text and first image into the first machine learning model. After processing the first text and first image internally, the first machine learning model outputs a generated image tensor, which is then post-processed (e.g., denormalization, resizing) to obtain the second image.

[0095] In one scenario, providing the first media content to a first machine learning model to obtain a second image may include: sending the acquired first text and the first image to a server, the server adding the first machine learning model, performing inference, generating the second image, and returning the second image to the client for display on the client.

[0096] In one scenario, the provided solution acquires first media content containing first text and a first image, where the first text explicitly indicates the processing method for the first image, providing clear and operable input for subsequent generation and helping to avoid output fluctuations caused by ambiguous processing requirements. In another scenario, the provided solution provides the aforementioned first media content to a first machine learning model, thereby obtaining a second image. During the training phase, the first machine learning model has been adjusted based on multiple fourth images and their corresponding first detection results. The first detection results are associated with multiple first strategies used to verify the effect of the third image. This allows the first machine learning model to internalize the constraint ability to verify the effect of the output image during the learning process. Therefore, when faced with different first media content during the inference phase, the first machine learning model can stably generate a second image that meets the processing requirements, effectively overcoming the defect of poor image quality consistency in related technologies, improving the processing efficiency of media content, and enhancing the image processing effect.

[0097] Figure 3 This is a flowchart illustrating a media content processing method under one scenario. The technical solution in this scenario can be combined with implementation methods in other scenarios. For identical or related parts, descriptions of other scenarios can be used, and will not be repeated here. Figure 3 As shown, the method in this case may specifically include:

[0098] S310. Obtain the second media content, provide the second media content to the second machine learning model, and obtain the fourth image.

[0099] The second media content can be multimodal data containing text and images. It serves as input data for the second machine learning model. Alternatively, it can be understood as input samples used during the training phase of the first machine learning model.

[0100] The second machine learning model can be understood as a model that can process text and images simultaneously and output an image. The second machine learning model can be a model that already possesses image editing capabilities.

[0101] The fourth image can be understood as the image generated by the second machine learning model in response to the second media content.

[0102] In some cases, acquiring second media content may include, but is not limited to, at least one of the following: acquiring second media content acquired through a preset media content acquisition control; acquiring second media content uploaded through a preset media content upload control; receiving second media content transmitted through a preset data interface; acquiring second media content generated by a preset model; extracting second media content from target content; acquiring preset media content as second media content, etc.

[0103] In one scenario, the second media content may include third text and a third image. Furthermore, the acquired third text and third image can be input into a second machine learning model to obtain a fourth image.

[0104] In one scenario, the second machine learning model can be obtained through training. For example, the second machine learning model can be obtained by training a seventh machine learning model based on a sixth media content and an eighth image, where the eighth image is the labeled image corresponding to the sixth media content. For instance, the sixth media content can be input into the seventh machine learning model to obtain a ninth image, and a first loss can be determined based on the ninth and eighth images. The parameters of the seventh machine learning model are then adjusted based on this first loss to obtain the second machine learning model. In another scenario, the seventh machine learning model can be a pre-trained model with dialogue capabilities and multimodal data processing capabilities.

[0105] S320, Provide the third machine learning model with the second media content, the second text corresponding to the second media content, and the fourth image corresponding to the second media content, to obtain the first detection result.

[0106] The third machine learning model can be understood as a model that evaluates the effectiveness of the fourth image based on the first strategy in the second text. For example, it could be a rule engine integrating multiple image quality metrics or a finely tuned visual language model.

[0107] In one scenario, the third machine learning model is trained as follows: a third media content, a fifth image, and a sixth image are acquired, wherein the fifth image and the sixth image are positive and negative sample images generated based on the third media content, respectively; a fourth text corresponding to the third media content is acquired, wherein the fourth text includes multiple second strategies, and the first strategies are used to test the effectiveness of the fifth image; based on the third media content, the fourth text, the fifth image, the sixth image, and the fourth machine learning model, a second detection result corresponding to the fifth image and a third detection result corresponding to the sixth image are obtained, respectively; and the parameters of the fourth machine learning model are adjusted based on the second and third detection results to obtain the third machine learning model.

[0108] In one scenario, the proposed solution involves acquiring positive and negative sample images, obtaining corresponding rule text, using a fourth machine learning model to output comparative detection results, and adjusting parameters to train a third machine learning model capable of accurately verifying the quality of generated images based on multiple rules. This third machine learning model possesses high discrimination accuracy and generalization ability, providing reliable quality feedback for the closed-loop optimization of the first machine learning model, ultimately improving the output stability of the entire media content processing system.

[0109] The third media content can be multimodal data containing text and / or images. This third media data serves as input samples for training the third machine learning model, simulating the media content provided to the model during the inference phase.

[0110] The fifth image can be understood as an image whose at least one visual feature satisfies the first condition. This first condition is used to filter images that meet the processing expectations, or in other words, images that present the expected effect. For example, the fifth image can be understood as a high-quality image generated based on third-media content. That is, the fifth image can serve as a positive sample, representing a generated result that meets the expected effect. Correspondingly, the sixth image can be understood as a low-quality image generated based on third-media content.

[0111] Similarly, the sixth image can serve as a negative sample, representing a generated result that does not meet expectations. It should be noted that the quality of the fifth and sixth images can be a relative result derived from a comparison. For example, the fifth image may be of better quality than the sixth image.

[0112] The fourth text can be understood as a collection of texts containing multiple validation rules. Corresponding to the third media content, the fourth text defines the desired effect on the generated image. The second strategy functions similarly to the first strategy, but serves as a supervisory signal when training the third machine learning model.

[0113] In one scenario, the second detection result can be understood as the conclusion output by the fourth machine learning model after detecting the fifth image (positive sample). For example, ideally, the second detection result should be "pass" or a high score.

[0114] In one scenario, the third detection result can be understood as the conclusion output by the fourth machine learning model after detecting the sixth image (the negative sample). For example, ideally, the third detection result should be "fail" or a low score.

[0115] In one scenario, based on the third media content, the fourth text, the fifth image, the sixth image, and the fourth machine learning model, obtaining a second detection result corresponding to the fifth image and a third detection result corresponding to the sixth image respectively includes: providing the third media content, the fourth text, and the fifth image to the fourth machine learning model to obtain at least one second detection result; and providing the third media content, the fourth text, and the sixth image to the fourth machine learning model to obtain at least one third detection result.

[0116] In one scenario, obtaining a second detection result corresponding to the fifth image and a third detection result corresponding to the sixth image based on the third media content, the fourth text, the fifth image, the sixth image, and the fourth machine learning model includes: providing the third media content, the fourth text, and the fifth image to the fourth machine learning model to obtain at least one fourth detection result; and providing the third media content, the fourth text, and the sixth image to the fourth machine learning model to obtain at least one fifth detection result; and obtaining a second detection result corresponding to the fifth image and a third detection result corresponding to the sixth image based on the difference information between the fourth detection result and the fifth detection result.

[0117] In one scenario, the proposed solution, by acquiring multiple detection results and determining the final positive / negative sample classification based on their relative comparisons, effectively eliminates the random errors or model biases that may arise from a single detection. The difference information amplifies the discrepancies between positive and negative samples, enabling more robust supervisory signals to be obtained through multiple result comparisons even if the model's initial discriminative ability is limited. This improves the training efficiency when subsequently adjusting the fourth machine learning model and the detection accuracy of the final third machine learning model.

[0118] The fourth detection result can be understood as the detection result output by the fourth machine learning model for the input (third media content, fourth text, and fifth image). Multiple fourth detection results can be obtained by inputting the third media content, the fourth text, and the fifth image into the fourth machine learning model multiple times.

[0119] The fifth detection result can be understood as the detection result output by the fourth machine learning model for the input (third media content, fourth text, and sixth image). Multiple fifth detection results can be obtained by inputting the third media content, the fourth text, and the sixth image multiple times into the fourth machine learning model.

[0120] In one scenario, the sum of the number of the fourth detection result and the number of the fifth detection result can be more than two. The number of the fourth detection result and the number of the fifth detection result can be the same or different.

[0121] Difference information can be understood as a comparative conclusion obtained by comparing the differences between the fourth and fifth detection results. For example, difference information can indicate whether the detection results of positive and negative samples meet expectations, that is, whether the difference information can indicate that the fifth image is a positive sample and the sixth image is a negative sample.

[0122] In one scenario, obtaining a second detection result corresponding to the fifth image based on the difference information between the fourth detection result and the fifth detection result includes: comparing each of the fourth detection results with each of the fifth detection results to obtain a first comparison result; using the total number of first results in the first comparison result as a first quantity, where the first result indicates that the fifth image is a positive sample image and the sixth image is a negative sample image; and using the ratio of the first quantity to the second quantity as the second detection result corresponding to the fourth detection result, where the second quantity is the total number of the fifth detection results.

[0123] The first comparison result can be understood as the conclusion drawn by comparing a single fourth detection result with a certain fifth detection result. The comparison method can be numerical comparison (e.g., whether the fourth detection result is greater than the fifth detection result), threshold judgment, or a custom rule. When the comparison conclusion indicates that the fifth image conforms to the rule more than the sixth image (i.e., the fourth detection result is better than the fifth detection result), this comparison result is called the "first result." For example, if the fourth detection result = 0.92 and the fifth detection result is 0.91, then 0.92 > 0.91, and this first comparison result is "positive sample is better than negative sample."

[0124] In one scenario, the first quantity can be understood as the total number of "first results" among all the first comparison results after comparing a given fourth detection result with each of the fifth detection results. For example, for a fourth detection result of 0.68, compared one by one with three fifth detection results [0.72, 0.45, 0.37], if it is greater than two of them, then the first quantity is 2. The second quantity is the total number of fifth detection results. For example, if the size of the fifth detection result set is 3, then the second quantity is 3. Further, the ratio obtained by dividing the first quantity by the second quantity can be used as the second detection result corresponding to the current fourth detection result. This ratio reflects the relative advantage of the fourth detection result compared to all negative sample detection results. For example, if the first quantity is 2 and the second quantity is 3, then the second detection result is approximately equal to 0.66.

[0125] In one scenario, the proposed solution transforms absolute score differences into relative ranking proportions by comparing each fourth detection result with all fifth detection results one by one and calculating their proportions. This effectively eliminates the impact of inconsistent model output scales or random fluctuations. This approach allows the second detection result to directly reflect the advantage of positive samples relative to the set of negative samples, thereby providing a more robust and smoother supervisory signal for subsequent parameter adjustments and improving the stability and convergence efficiency of the third machine learning model training.

[0126] In one scenario, obtaining the third detection result corresponding to the sixth image based on the difference information between the fourth detection result and the fifth detection result includes: comparing each fifth detection result with each of the fourth detection results to obtain a second comparison result; using the total number of second results in the second comparison result as a third quantity, where the second result indicates that the fifth image is a positive sample image and the sixth image is a negative sample image; and using the ratio of the third quantity to the fourth quantity as the third detection result corresponding to the fifth detection result, where the second quantity is the total number of the fourth detection results.

[0127] The second comparison result can be understood as the conclusion drawn by comparing a single fifth detection result with a certain fourth detection result. The comparison method can be a numerical comparison (e.g., whether the fifth detection result is less than the fourth detection result) or a custom rule. If the comparison conclusion indicates that the fifth image is superior to the sixth image (i.e., positive samples are superior to negative samples), a "second result" is output. Specifically, the meaning of the second result is the same as the aforementioned first result: indicating that the fifth image is a positive sample and the sixth image is a negative sample. Therefore, when the fifth detection result < the fourth detection result, it conforms to the relationship of positive being superior to negative, generating a second result.

[0128] The third quantity can be understood as the total number of "second results" generated from all the first comparison results after comparing a given fifth detection result with each of the four fourth detection results. For example, for the fifth detection result 0.45, it is compared one by one with the three fourth detection results [0.92, 0.88, 0.95]. Since 0.45 is less than all three numbers, each comparison generates a second result, and the third quantity is 3.

[0129] The fourth quantity refers to the total number of fourth test results. For example, if the total number of fourth test results is 3, then the fourth quantity is 3.

[0130] The ratio obtained by dividing the third quantity by the fourth quantity is taken as the third detection result corresponding to the current fifth detection result. This ratio reflects the relative inferiority of the fifth detection result compared to all positive sample detection results (or the degree to which negative samples are correctly identified as negative). For example, if the third quantity is 3 and the fourth quantity is 3, then the third detection result is 1.0; if the third quantity is 1 and the fourth quantity is 3, then the third detection result is approximately equal to 0.33.

[0131] In one scenario, the proposed solution generates a quantified third detection result for the negative sample (sixth image) by comparing each fifth detection result with all fourth detection results one by one and calculating their proportions. This result directly reflects the degree of inferiority of the negative sample relative to all positive samples. This symmetrical comparison method ensures that the evaluation scale for positive and negative samples is consistent and is insensitive to the absolute value of the model output, thus providing a balanced and robust training signal for subsequent adjustments to the fourth machine learning model. Combined with the aforementioned second detection result, a complete contrastive learning supervision objective can be formed, further improving the discriminative stability of the third machine learning model.

[0132] In one scenario, multiple image pairs can be constructed based on multiple images output from the same input media content. Each image pair includes one positive sample image and one negative sample image. For each image pair, a third machine learning model generates multiple independent inference trajectories for both images, and then a score for each image is obtained through rule parsing. A positive sample group and a negative sample group are constructed based on the multiple positive sample images. The scores of all images in both the positive and negative sample groups are compared pairwise: the win rate reward for the positive sample group is the percentage of times its score is higher than all samples in the negative sample group; the negative rate reward for the negative sample group is the percentage of times its score is lower than all samples in the positive sample group. These two ratios serve as the final reward values ​​for the two groups, measuring their relative performance across groups.

[0133] like Figure 4As shown, the process involves obtaining image A, text A indicating the processing method of image A, text B including multiple first strategies, image B1, and image B2. Image B1 and image B2 are both obtained by a first machine learning model (machine learning model A) based on text A, and are respectively negative and positive samples based on that input. By processing image A, text A, text B, and image B1 through the inference trajectory of machine learning model B, multiple scores corresponding to image B1 are obtained: 0.63, 0.42, and 0.28. Similarly, by processing image A, text A, text B, and image B2 through the inference trajectory of machine learning model B, multiple scores corresponding to image B1 are obtained: 0.77, 0.58, and 0.35. Then, the two sets of images corresponding to images B1 and B2 are compared to determine the reward corresponding to each score. The reward for the score of the image B1 set is the percentage of times its score is lower than all scores of the image B2 set. The reward for the score of image B2 is the percentage of times its score is higher than the total score of image B1; the reward for the negative sample group is the percentage of times its score is lower than the total score of the positive sample group. For example, the score corresponding to thought chain 1 for image B1 is 0.63. Since image B1 is a selected negative sample image, theoretically, the score of image B1 output by machine learning model B should be lower than the score corresponding to image B2. However, upon comparison, this score is only lower than the score of 0.72 corresponding to thought chain 1 for image B2, but higher than the scores of image B2 under the other two thought chains. This means that in three reasoning attempts, machine learning model B was only reasonable or the result met expectations once, so the reward corresponding to the 0.62 score is 1 / 3. Other rewards are obtained in a similar way and will not be elaborated here.

[0134] In one scenario, the average reward can be determined based on the relative merits of multiple images within each positive and negative sample group. Then, the average reward of the corresponding group is subtracted from the relative merit of each sample from its relative merits to obtain the relative advantage value within the group, which measures whether it is better or worse than the average within the group. The model loss is determined based on the relative advantage value within each image in the positive and negative sample groups, and the parameters of the fourth machine learning model are adjusted to obtain the third machine learning model.

[0135] In one scenario, a fifth machine learning model can be trained based on fourth media content and a seventh image to obtain the fourth machine learning model, where the seventh image is the labeled image corresponding to the fourth media content. For example, the fourth media content can be input into the fifth machine learning model to obtain a tenth image, and a second loss can be determined based on the tenth and seventh images. The parameters of the fifth machine learning model are then adjusted based on this second loss to obtain the fourth machine learning model. In another scenario, the fifth machine learning model can be a pre-trained model with dialogue capabilities and multimodal data processing capabilities. For example, the fifth machine learning model can accept media content input and output image-related predictions, such as generated images.

[0136] In one scenario, the proposed solution trains the fifth machine learning model using fourth media content and the corresponding seventh image, enabling the model to establish a fundamental ability to correctly associate media content with images from the early stages of learning. This pre-training process provides good initial parameter values ​​for subsequent fine-tuning through positive and negative sample comparisons, thereby accelerating the convergence of the third machine learning model and improving the accuracy and generalization of the final detection results.

[0137] like Figure 5 As shown, the training process of machine learning model A includes: acquiring image A and text A to indicate the processing method for image A. On one hand, image A and text A are input into a sixth machine learning model (not shown in the figure, such as a visual language model) to obtain text B. Text B is used to indicate multiple rules that need to be followed in the processing results of image A. These rules are obtained through analysis of text A and image A. On the other hand, image A and text A are input into machine learning model A to obtain multiple processed images, such as image B1, image B2, image B3, etc. Inputting image A, text A, text B, and image B1 into machine learning model B yields the first detection result for image B1, such as a score of 8.5. Inputting image A, text A, text B, and image B2 into machine learning model B yields the first detection result for image B2. Inputting image A, text A, text B, and image B3 into machine learning model B yields the first detection result for image B3.

[0138] S330. Adjust the second machine learning model based on the first detection results of multiple fourth images to obtain the first machine learning model.

[0139] In one scenario, the first detection result of each fourth image is converted into the loss weight of that sample (where samples that fail can be assigned high weights), and the parameters of the second machine learning model are fine-tuned using weighted gradient descent. This process is repeated multiple times until the pass rate on the validation set stabilizes above the threshold.

[0140] In one scenario, a reinforcement learning framework can be used, where the first detection result (such as the overall score) is used as the reward signal, and the second machine learning model is used as the policy network. The model is updated through the policy gradient algorithm, and after each update, the newly generated image is used to re-detect to obtain a new reward.

[0141] In one scenario, contrastive learning can be used to construct a contrastive loss function, taking the fourth image with the best first detection result as a positive example and the one with the worst result as a negative example, and adjusting the model encoder to make the separation between positive and negative examples in its output feature space higher.

[0142] In one scenario, multiple fourth images obtained from the same second media content are treated as a sample group. Each fourth image in the sample group is detected, and its score is obtained. The average score and fluctuation range of the entire sample group are calculated based on the scores of the multiple fourth images. The score of each fourth image is then compared with the average score of the sample group, and normalization is performed according to the fluctuation within the group to obtain the relative advantage value of each fourth image. The current second machine learning model is compared with the previous second machine learning model to obtain the change ratio of the output probabilities of multiple feature processing units of the fourth image. The probability change ratio is clipped to limit the magnitude of a single parameter update of the second machine learning model and avoid training runaway. The basic policy loss is obtained based on the clipped probability ratio and the relative advantage value of the fourth image. An additional KL divergence constraint loss is added to control the new model from deviating too far from the original fine-tuned model. The policy loss and KL constraint loss are integrated, averaged, and used as the final loss for model back-up updates.

[0143] S340. Obtain first media content, which includes first text and first image, wherein the first text is used to indicate the processing method of the first image.

[0144] S350: Provide the first media content to the first machine learning model to obtain the second image.

[0145] In one scenario, the proposed solution establishes a data link between the model's current output and subsequent detection by providing explicit second media content to the second machine learning model and obtaining the fourth image it generates. This provides the necessary test images for subsequent rule-based performance evaluation, thereby supporting a quantitative assessment of the model's generation quality.

[0146] In one scenario, the proposed solution, by introducing a third machine learning model and using multiple first strategies from the second text to explicitly detect the fourth image, can transform abstract performance requirements into concrete and quantifiable first detection results. This provides an objective basis for evaluating the quality of the model's generation, avoiding inconsistencies caused by relying solely on subjective judgment, and thus providing stable and reliable feedback signals for subsequent model adjustments.

[0147] In one scenario, the proposed solution adjusts the second machine learning model based on multiple fourth images and their corresponding first detection results. This allows the final first machine learning model to internalize the constraints of the first strategy during the learning process, thereby enabling it to actively approach the effect that meets the rule requirements when generating images and significantly reducing the possibility of the generated images deviating from the expected quality.

[0148] Figure 6 This is a schematic diagram of the structure of a media content processing device in one scenario, such as... Figure 6 As shown, the media content processing device includes a first module 610 and a second module 620. The first module 610 is used to acquire first media content, which includes first text and a first image. The first text indicates the processing method for the first image. The second module 620 is used to provide the first media content to a first machine learning model to obtain a second image. The first machine learning model adjusts itself based on first detection results from multiple fourth images. The fourth images are obtained by the second machine learning model based on the second media content. The second media content includes a third image. The first detection results are associated with the second text. The second text includes multiple first strategies, which are used to evaluate the effectiveness of the third image.

[0149] In one scenario, the provided solution acquires first media content containing first text and a first image, where the first text explicitly indicates the processing method for the first image, providing clear and operable input for subsequent generation and helping to avoid output fluctuations caused by ambiguous processing requirements. In another scenario, the provided solution provides the aforementioned first media content to a first machine learning model, thereby obtaining a second image. During the training phase, the first machine learning model has been adjusted based on multiple fourth images and their corresponding first detection results. The first detection results are associated with multiple first strategies used to verify the effect of the third image. This allows the first machine learning model to internalize the constraint ability to verify the effect of the output image during the learning process. Therefore, when faced with different first media content during the inference phase, the first machine learning model can stably generate a second image that meets the processing requirements, effectively overcoming the defect of poor image quality consistency in related technologies, improving the processing efficiency of media content, and enhancing the image processing effect.

[0150] In one scenario, the media content processing apparatus may include a third module. This third module is configured to provide a third machine learning model with the second media content, the second text corresponding to the second media content, and the fourth image corresponding to the second media content, to obtain a first detection result.

[0151] In one scenario, the media content processing device may include a fourth module, a fifth module, a sixth module, and a seventh module. The fourth module is used to acquire third media content, a fifth image, and a sixth image, wherein the fifth image and the sixth image are positive and negative sample images generated based on the third media content, respectively. The fifth module is used to acquire fourth text corresponding to the third media content, wherein the fourth text includes multiple second strategies, and the first strategy is used to test the effect of the fifth image. The sixth module is used to obtain a second detection result corresponding to the fifth image and a third detection result corresponding to the sixth image, respectively, based on the third media content, the fourth text, the fifth image, the sixth image, and a fourth machine learning model. The seventh module is used to adjust the parameters of the fourth machine learning model based on the second and third detection results to obtain a third machine learning model.

[0152] In some cases, the sixth module may include a first unit, a second unit, and a third unit. The first unit is used to provide the third media content, the fourth text, and the fifth image to the fourth machine learning model to obtain at least one fourth detection result; the second unit is used to provide the third media content, the fourth text, and the sixth image to the fourth machine learning model to obtain at least one fifth detection result; and the third unit is used to obtain a second detection result corresponding to the fifth image and a third detection result corresponding to the sixth image, respectively, based on the difference information between the fourth and fifth detection results.

[0153] In some cases, the third unit can be further used to: compare each of the fourth detection results with each of the fifth detection results to obtain a first comparison result; take the total number of first results in the first comparison result as a first quantity, the first result being used to indicate that the fifth image is a positive sample image and the sixth image is a negative sample image; and take the ratio of the first quantity to the second quantity as a second detection result corresponding to the fourth detection result, the second quantity being the total number of the fifth detection results.

[0154] In some cases, the media content processing apparatus further includes an eighth module. This eighth module is used to train a fifth machine learning model based on the fourth media content and a seventh image to obtain the fourth machine learning model, wherein the seventh image is a tag image corresponding to the fourth media content.

[0155] In some cases, the second media content may further include a third text; the media content processing device also includes a ninth module. The ninth module is used to provide the third text, the third image, and the fifth text to a sixth machine learning model to obtain a second text, wherein the fifth text is used to indicate the content to be included in the second text.

[0156] In some cases, the fifth text may include at least one of the following:

[0157] Fifth media content and its corresponding multiple third strategies;

[0158] The description information of the strategy generation logic, which is the logic of generating multiple strategies based on media content.

[0159] In some cases, the first strategy may include at least one of the following:

[0160] Content processing principles, which are used to examine at least one of the following: content in the image that needs to remain unchanged, and content in the image that needs to be modified;

[0161] The quality inspection principle is used to inspect at least one of the following for an image: visual fidelity and visual integrity, wherein visual fidelity is used to indicate the consistency of visual representation before and after image processing, and visual integrity is used to indicate the structural and logical rationality of the image.

[0162] The aforementioned media content processing apparatus can execute the media content processing method provided in any of the embodiments described herein, and has the corresponding functional modules and beneficial effects for executing the media content processing method.

[0163] It is worth noting that the various units and modules included in the above-mentioned media content processing device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments described herein.

[0164] The following is for reference. Figure 7 This document illustrates a schematic diagram of an electronic device (e.g., a terminal device or server) 700 suitable for implementing the above-described methods. The terminal device referred to herein may include, but is not limited to, mobile phones, laptops, digital broadcast receivers, and personal digital assistants (PDAs). ), tablet computer ), portable multimedia player ( Mobile terminals such as vehicle-mounted terminals (e.g., vehicle navigation terminals) and fixed terminals such as digital televisions and desktop computers. Figure 7The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments described herein.

[0165] like Figure 7 As shown, electronic device 700 may include processing device (e.g., central processing unit, graphics processor, etc.) 701, which can be based on data stored in read-only memory (ROM). The program in 702 or loaded from storage device 708 into random access memory (RAM) The RAM 703 stores various appropriate actions and processes through the programs in RAM 703. RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.

[0166] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; including, for example, liquid crystal displays (LCDs). The device 707 includes an output device 707 such as a speaker or vibrator; a storage device 708 including, for example, magnetic tape or hard disk; and a communication device 709. The communication device 709 allows the electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0167] In particular, according to embodiments of this document, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, the technical solutions of this document include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the methods of the embodiments of this document.

[0168] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0169] The electronic device provided in this embodiment and the media content processing method provided in the above technical solutions belong to the same inventive concept. Technical details not described in detail in this document can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0170] This article provides a computer storage medium on which a computer program is stored, which, when executed by a processor, implements the media content processing method provided in the above embodiments.

[0171] It should be noted that the computer-readable medium mentioned above can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, or an erasable programmable read-only memory (ROM). Also known as flash memory), fiber optic, portable compact disk read-only memory (FROM) Computer-readable storage media can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0172] Based on one or more scenarios described herein, Example 1 provides a media content processing method, comprising: acquiring first media content, the first media content including first text and a first image, the first text being used to indicate the processing method for the first image; providing the first media content to a first machine learning model to obtain a second image, the first machine learning model adjusting the second machine learning model based on first detection results of multiple fourth images, the fourth image being obtained by the second machine learning model based on the second media content, the second media content including a third image, the first detection results being associated with the second text, the second text including multiple first strategies, the first strategies being used to test the effect of the third image.

[0173] According to one or more scenarios in this paper, Example 2 provides the method of Example 1. Optionally, the first detection result is obtained by providing the second media content, the second text corresponding to the second media content, and the fourth image corresponding to the second media content to a third machine learning model to obtain the first detection result.

[0174] Based on one or more scenarios described herein, Example 3 provides the method of Example 2. Optionally, the third machine learning model is trained as follows: acquiring third media content, a fifth image, and a sixth image, wherein the fifth image and the sixth image are positive sample images and negative sample images generated based on the third media content, respectively; acquiring fourth text corresponding to the third media content, wherein the fourth text includes multiple second strategies, and the first strategies are used to test the effect of the fifth image; obtaining a second detection result corresponding to the fifth image and a third detection result corresponding to the sixth image based on the third media content, the fourth text, the fifth image, the sixth image, and the fourth machine learning model, respectively; and adjusting the parameters of the fourth machine learning model based on the second detection result and the third detection result to obtain the third machine learning model.

[0175] According to one or more scenarios described herein, Example 4 provides the method of Example 3. Optionally, obtaining the second detection result corresponding to the fifth image and the third detection result corresponding to the sixth image based on the third media content, the fourth text, the fifth image, the sixth image, and the fourth machine learning model includes: providing the third media content, the fourth text, and the fifth image to the fourth machine learning model to obtain at least one fourth detection result; and providing the third media content, the fourth text, and the sixth image to the fourth machine learning model to obtain at least one fifth detection result; and obtaining the second detection result corresponding to the fifth image and the third detection result corresponding to the sixth image based on the difference information between the fourth detection result and the fifth detection result.

[0176] According to one or more scenarios in this document, Example 5 provides the method of Example 4. Optionally, obtaining the second detection result corresponding to the fifth image based on the difference information between the fourth detection result and the fifth detection result includes: comparing the fourth detection result with each of the fifth detection results for a single fourth detection result to obtain a first comparison result; taking the total number of first results in the first comparison result as a first quantity, where the first result is used to indicate that the fifth image is a positive sample image and the sixth image is a negative sample image; and taking the ratio of the first quantity to the second quantity as the second detection result corresponding to the fourth detection result, where the second quantity is the total number of the fifth detection results.

[0177] According to one or more scenarios in this article, Example 6 provides the method of Example 3, which further includes: training a fifth machine learning model based on the fourth media content and a seventh image to obtain the fourth machine learning model, wherein the seventh image is a tag image corresponding to the fourth media content.

[0178] According to one or more scenarios in this document, Example 7 provides the method of Example 1. Optionally, the second media content further includes a third text, which is obtained by providing the third text, the third image, and a fifth text to a sixth machine learning model to obtain the second text, wherein the fifth text is used to indicate the content to be included in the second text.

[0179] According to one or more scenarios in this document, Example 8 provides the method of Example 7, which further includes: optionally, the fifth text includes at least one of the following: fifth media content and its corresponding multiple third strategies; descriptive information of strategy generation logic, wherein the strategy generation logic is logic for generating multiple strategies based on media content.

[0180] Based on one or more scenarios described herein, Example 9 provides the method of Example 7. Optionally, the first strategy includes at least one of the following: a content processing principle, which is used to examine at least one of the following: content in the image that needs to remain unchanged, and content in the image that needs to be modified; a quality inspection principle, which is used to examine at least one of the following of the image: visual fidelity and visual integrity, wherein visual fidelity is used to indicate the consistency of visual representation before and after image processing, and visual integrity is used to indicate the structural and logical rationality of the image.

[0181] According to one or more scenarios described herein, Example 10 provides a media content processing apparatus, comprising: a first module for acquiring first media content, the first media content including first text and a first image, the first text being used to indicate a processing method for the first image; and a second module for providing the first media content to a first machine learning model to obtain a second image, wherein the first machine learning model adjusts the second machine learning model based on first detection results of multiple fourth images, the fourth images being obtained by the second machine learning model based on the second media content, the second media content including third text and a third image, the first detection results being associated with the second text, the second text including multiple first strategies, the first strategies being used to verify the effect of the third image.

[0182] In some implementations, the client and server can utilize, for example... It can communicate with any currently known or future-developed network protocol, such as Hypertext Transfer Protocol (HTTP), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs). The Internet (e.g., the Internet) and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any network currently known or under future development.

[0183] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0184] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire first media content, the first media content including first text and a first image, the first text being used to indicate a processing method for the first image; provide the first media content to a first machine learning model to obtain a second image, the first machine learning model being adjusted based on first detection results of multiple fourth images, the fourth image being obtained by the second machine learning model based on the second media content, the second media content including third text and a third image, the first detection results being associated with the second text, the second text including multiple first strategies, the first strategies being used to verify the effect of the third image.

[0185] Computer program code for performing the operations described herein can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0186] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this document. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0187] The modules or units described herein can be implemented in software or hardware. The names of modules or units do not necessarily limit the module or unit itself; for example, the first unit can also be described as "a unit that obtains at least one fourth detection result".

[0188] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include at least one of the following: field-programmable gate arrays (FPGAs). Application-Specific Integrated Circuits Specialized standard products System-on-a-Chip Complex programmable logic devices etc.

[0189] In the context of this document, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (flash memory), optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0190] The above description is merely a preferred embodiment and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.

[0191] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this document. Certain features described in the context of individual implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.

[0192] Although the subject matter has been described using a programming language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims.

Claims

1. A media content processing method, comprising: Acquire first media content, the first media content including first text and first image, the first text being used to indicate the processing method of the first image; The first media content is provided to the first machine learning model to obtain the second image. The first machine learning model is adjusted based on the first detection results of multiple fourth images to obtain the second machine learning model. The fourth image is obtained by the second machine learning model based on the second media content. The second media content includes a third image. The first detection results are associated with second text. The second text includes multiple first strategies. The first strategies are used to test the effect of the third image.

2. The media content processing method according to claim 1, wherein the first detection result is obtained in the following manner: The third machine learning model is provided with the second media content, the second text corresponding to the second media content, and the fourth image corresponding to the second media content to obtain a first detection result.

3. The media content processing method according to claim 2, wherein the third machine learning model is trained in the following manner: Acquire third media content, a fifth image, and a sixth image, wherein the fifth image and the sixth image are positive sample images and negative sample images generated based on the third media content, respectively; Obtain a fourth text corresponding to the third media content, the fourth text including multiple second strategies, the second strategies being used to verify the effect of the fifth image; Based on the third media content, the fourth text, the fifth image, the sixth image, and the fourth machine learning model, the second detection result corresponding to the fifth image and the third detection result corresponding to the sixth image are obtained respectively. The parameters of the fourth machine learning model are adjusted based on the second and third detection results to obtain the third machine learning model.

4. The media content processing method according to claim 3, wherein obtaining the second detection result corresponding to the fifth image and the third detection result corresponding to the sixth image based on the third media content, the fourth text, the fifth image, the sixth image, and the fourth machine learning model, respectively, includes: The third media content, the fourth text, and the fifth image are provided to the fourth machine learning model to obtain at least one fourth detection result; as well as, The third media content, the fourth text, and the sixth image are provided to the fourth machine learning model to obtain at least one fifth detection result; Based on the difference information between the fourth detection result and the fifth detection result, the second detection result corresponding to the fifth image and the third detection result corresponding to the sixth image are obtained respectively.

5. The media content processing method according to claim 4, wherein obtaining the second detection result corresponding to the fifth image based on the difference information between the fourth detection result and the fifth detection result includes: For each of the fourth detection results, the fourth detection result is compared with each of the fifth detection results to obtain a first comparison result; The total number of first results in the first comparison results is taken as the first number, and the first results are used to indicate that the fifth image is a positive sample image and the sixth image is a negative sample image; The ratio of the first quantity to the second quantity is taken as the second detection result corresponding to the fourth detection result, and the second quantity is the total number of the fifth detection results.

6. The media content processing method according to claim 3 further includes: The fifth machine learning model is trained based on the fourth media content and the seventh image to obtain the fourth machine learning model, wherein the seventh image is the tag image corresponding to the fourth media content.

7. The media content processing method according to claim 1, wherein the second media content further includes a third text; the second text is obtained in the following manner: The third text, the third image, and the fifth text are provided to the sixth machine learning model to obtain the second text, wherein the fifth text is used to indicate the content that should be included in the second text.

8. The media content processing method according to claim 7, wherein the fifth text includes at least one of the following: Fifth media content and its corresponding multiple third strategies; The description information of the strategy generation logic, which is the logic of generating multiple strategies based on media content.

9. The media content processing method according to claim 1, wherein the first strategy includes at least one of the following: Content processing principles, which are used to examine at least one of the following: content in the image that needs to remain unchanged, and content in the image that needs to be modified; quality inspection principles for inspecting at least one of the following of an image: visual fidelity and visual integrity, wherein, The visual fidelity is used to indicate the consistency of visual representation before and after image processing, and the visual integrity is used to indicate the rationality of the image in terms of structure and logic.

10. A media content processing apparatus, comprising: The first module is used to acquire first media content, which includes first text and a first image, wherein the first text is used to indicate the processing method of the first image; The second module is used to provide the first media content to the first machine learning model to obtain the second image. The first machine learning model adjusts the second machine learning model based on the first detection results of multiple fourth images. The fourth image is obtained by the second machine learning model based on the second media content. The second media content includes third text and the third image. The first detection results are associated with the second text. The second text includes multiple first strategies. The first strategies are used to test the effect of the third image.

11. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the media content processing method as described in any one of claims 1-9.

12. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the media content processing method as described in any one of claims 1-9.

13. A computer program product comprising a computer program that, when executed by a processor, implements the media content processing method as described in any one of claims 1-9.