A multimedia information processing method and device, electronic equipment and storage medium

CN122547772APending Publication Date: 2026-08-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

由于第一权重系数表征加噪过程中原始多媒体信息在已加噪多媒体信息中所占的比例,通过对原始处理因子进行重构,使得重构后的第一目标处理因子的分母不包含第一权重系数,可以避免在去噪初始阶段,由于第一权重系数趋于0而导致去噪常微分方程数值解的步长较大的问题,从而减小去噪初始阶段的误差,提升通过去噪过程所生成的编辑后的目标多媒体信息的质量

Benefits of technology

[0017]本申请公开的多媒体信息处理方法,通过对与去噪过程相关的去噪常微分方程中包括的原始处理因子进行重构,得到分母不包含原始多媒体信息对应的第一权重系数的第一目标处理因子,基于第一目标处理因子建立目标多媒体信息处理模型,将原始多媒体信息以及编辑指令输入目标多媒体信息处理模型,可以得到编辑后的目标多媒体信息。由于第一权重系数表征加噪过程中原始多媒体信息在已加噪多媒体信息中所占的比例,通过对原始处理因子进行重构,使得重构后的第一目标处理因子的分母不包含第一权重系数,可以避免在去噪初始阶段,由于第一权重系数趋于0而导致去噪常微分方程数值解的步长较大的问题,从而减小去噪初始阶段的误差,提升通过去噪过程所生成的编辑后的目标多媒体信息的质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547772A_ABST
    Figure CN122547772A_ABST
Patent Text Reader

Abstract

This application relates to a multimedia information processing method, apparatus, electronic device, and storage medium, comprising: acquiring original multimedia information, editing instructions, and a target multimedia information processing model; inputting the original multimedia information into a denoising sub-model to obtain denoised multimedia information corresponding to the original multimedia information; and inputting the denoised multimedia information and editing instructions into a denoising sub-model to obtain target multimedia information corresponding to the original multimedia information. This application reconstructs the original processing factors involved in the denoising process to obtain a first target processing factor whose denominator does not contain the first weight coefficient corresponding to the original multimedia information. Based on the first target processing factor, a target multimedia information processing model is established. Inputting the original multimedia information and editing instructions into the target multimedia information processing model yields edited target multimedia information. Through reconstruction processing, the error in the initial stage of denoising can be reduced, improving the quality of the target multimedia information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimedia information processing technology, and in particular to a multimedia information processing method, apparatus, electronic device and storage medium. Background Technology

[0002] Since the advent of Generative Adversarial Networks (GANs), multimedia information generation has become an important frontier topic in the field of artificial intelligence. Multimedia information generation can include image generation, text generation, and audio generation. In addition to the well-known GANs, mainstream methods also include Variational Autoencoders (VAEs), flow-based generative models, and the recently popular diffusion models.

[0003] The diffusion model includes two main processes: forward diffusion and reverse generation. In the forward diffusion process, noise is gradually added to the original multimedia information, turning it into noise. In the reverse generation process, the original multimedia information is gradually recovered from the noise through denoising.

[0004] In existing techniques for generating multimedia information using diffusion models, the numerical solution step size of the ordinary differential equations related to the denoising process is relatively large in the initial stage of the reverse generation process, i.e., the initial stage of denoising. This leads to a large error in the initial stage of denoising, which affects the quality of the finally generated multimedia information. Summary of the Invention

[0005] To address the aforementioned technical problems, this application discloses a multimedia information processing method, apparatus, electronic device, and storage medium. By reconstructing the original processing factors included in the denoising ordinary differential equation related to the denoising process, a first target processing factor is obtained whose denominator does not contain the first weight coefficient corresponding to the original multimedia information. A target multimedia information processing model is established based on the first target processing factor. Inputting the original multimedia information and editing instructions into the target multimedia information processing model yields the edited target multimedia information. Since the first weight coefficient represents the proportion of the original multimedia information in the already denoised multimedia information during the denoising process, reconstructing the original processing factors ensures that the denominator of the reconstructed first target processing factor does not contain the first weight coefficient. This avoids the problem of a large step size in the numerical solution of the denoising ordinary differential equation due to the first weight coefficient approaching 0 in the initial stage of denoising, thereby reducing the error in the initial stage of denoising and improving the quality of the edited target multimedia information generated through the denoising process.

[0006] On the one hand, this application provides a multimedia information processing method, the method comprising:

[0007] The system acquires original multimedia information, editing instructions for the original multimedia information, and a target multimedia information processing model. The target multimedia information processing model includes a noise-adding sub-model and a noise-reducing sub-model. The noise-reducing sub-model is established based on a first target processing factor, which is obtained by reconstructing the original processing factor. The denominator of the original processing factor includes a first weight coefficient corresponding to the original multimedia information, while the denominator of the first target processing factor does not include the first weight coefficient. The first weight coefficient represents the proportion of the original multimedia information in the noisy multimedia information during the noise-adding process.

[0008] The original multimedia information is input into the noise-adding sub-model, and the original multimedia information is subjected to step-by-step noise-adding processing to obtain the noise-adding multimedia information corresponding to the original multimedia information.

[0009] The noisy multimedia information and the editing instructions are input into the denoising sub-model to perform step-by-step denoising processing on the noisy multimedia information, thereby obtaining the target multimedia information corresponding to the original multimedia information.

[0010] On the other hand, this application also provides a multimedia information processing apparatus, the apparatus comprising:

[0011] The acquisition module is used to acquire original multimedia information, editing instructions for the original multimedia information, and a target multimedia information processing model. The target multimedia information processing model includes a noise-adding sub-model and a noise-reducing sub-model. The noise-reducing sub-model is established based on a first target processing factor. The first target processing factor is obtained by reconstructing the original processing factor. The denominator of the original processing factor includes a first weight coefficient corresponding to the original multimedia information. The denominator of the first target processing factor does not include the first weight coefficient. The first weight coefficient represents the proportion of the original multimedia information in the noisy multimedia information during the noise-adding process.

[0012] The noise-adding module is used to input the original multimedia information into the noise-adding sub-model, perform step-by-step noise-adding processing on the original multimedia information, and obtain the noise-adding multimedia information corresponding to the original multimedia information.

[0013] The denoising module is used to input the noisy multimedia information and the editing instructions into the denoising sub-model, and to perform step-by-step denoising processing on the noisy multimedia information to obtain the target multimedia information corresponding to the original multimedia information.

[0014] On the other hand, this application also provides an electronic device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the multimedia information processing method as described above.

[0015] On the other hand, this application also provides a computer-readable storage medium storing at least one instruction or at least one program, which is loaded and executed by a processor to implement the multimedia information processing method as described above.

[0016] Implementing the embodiments of this application has the following beneficial effects:

[0017] The multimedia information processing method disclosed in this application reconstructs the original processing factors included in the denoising ordinary differential equation related to the denoising process, obtaining a first target processing factor whose denoiser does not contain the first weight coefficient corresponding to the original multimedia information. A target multimedia information processing model is established based on the first target processing factor. The original multimedia information and editing instructions are input into the target multimedia information processing model to obtain the edited target multimedia information. Since the first weight coefficient represents the proportion of the original multimedia information in the denoised multimedia information during the denoising process, reconstructing the original processing factors ensures that the denoiser of the reconstructed first target processing factor does not contain the first weight coefficient. This avoids the problem of a large step size in the numerical solution of the denoising ordinary differential equation due to the first weight coefficient tending to 0 in the initial stage of denoising, thereby reducing the error in the initial stage of denoising and improving the quality of the edited target multimedia information generated through the denoising process. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram illustrating an application scenario of a multimedia information processing method provided in an embodiment of this application.

[0020] Figure 2 A flowchart illustrating a multimedia information processing method provided in an embodiment of this application;

[0021] Figure 3 A flowchart illustrating a noise-adding process provided in an embodiment of this application;

[0022] Figure 4 A flowchart illustrating a noise reduction process provided in an embodiment of this application;

[0023] Figure 5 A flowchart illustrating a method for determining a denoising sub-model provided in an embodiment of this application;

[0024] Figure 6 This is a schematic diagram of an interpolation process provided in an embodiment of this application;

[0025] Figure 7 A flowchart illustrating a method for determining a noise-generating sub-model provided in an embodiment of this application;

[0026] Figure 8 A flowchart illustrating a method for generating target multimedia information provided in an embodiment of this application;

[0027] Figure 9 A flowchart illustrating a second noise determination method provided in an embodiment of this application;

[0028] Figure 10 A flowchart illustrating a method for generating denoised multimedia information provided in an embodiment of this application;

[0029] Figure 11 This is a schematic diagram of the structure of a multimedia information processing device provided in an embodiment of this application;

[0030] Figure 12 This is a schematic diagram of the hardware structure of a device for implementing the method provided in the embodiments of this application. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0032] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. Furthermore, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such information can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that illustrated or described herein.

[0033] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0034] 1) Diffusion Models: Diffusion models include two main processes: forward diffusion and reverse generation. In the forward diffusion process, noise is gradually added to the original multimedia information, turning it into noise. In the reverse generation process, the original multimedia information is gradually recovered from the noise through denoising.

[0035] 2) Ordinary Differential Equation (ODE): An equation that represents the relationship between an unknown function, its derivative, and the independent variable is called a differential equation. The unknown function is a univariate function, that is, a differential equation with only one independent variable is called an ordinary differential equation.

[0036] 3) Denoising Diffusion Implicit Models (DDIM): A type of diffusion model that can achieve skip-step denoising and speed up the denoising process.

[0037] See Figure 1 , Figure 1 This is a schematic diagram of an application scenario for a multimedia information processing method provided in an embodiment of this application. The application scenario includes a server 100, a network 200, and a terminal 300. The terminal 300 is equipped with an information processing client, which is displayed on the display interface. The terminal 300 is connected to the server 100 through the network 200.

[0038] The terminal 300 is used to send an information processing request to the server 100. The information processing request may include original multimedia information and editing instructions for the original multimedia information.

[0039] Server 100 is configured to receive information processing requests sent by terminal 300, and based on the information processing requests, acquire original multimedia information, editing instructions for the original multimedia information, and a target multimedia information processing model. The target multimedia information processing model includes a noise-adding sub-model and a noise-reducing sub-model. The noise-reducing sub-model is established based on a first target processing factor, which is obtained by reconstructing an original processing factor. The denominator of the original processing factor includes a first weight coefficient corresponding to the original multimedia information, while the denominator of the first target processing factor does not include the first weight coefficient. The first weight coefficient represents the proportion of the original multimedia information in the noisy multimedia information during the noise-adding process. The original multimedia information is input into the noise-adding sub-model, and the original multimedia information is subjected to progressive noise-adding processing to obtain noisy multimedia information corresponding to the original multimedia information. The noisy multimedia information and the editing instructions are input into the noise-reducing sub-model, and the noisy multimedia information is subjected to progressive noise-reducing processing to obtain the target multimedia information corresponding to the original multimedia information. The target multimedia information is then sent to terminal 300.

[0040] Terminal 300 is also used to receive target multimedia information sent by server 100 and display the target multimedia information on the display interface.

[0041] In some embodiments, server 100 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Network 200 may be a wide area network (WAN), a local area network (LAN), or a combination of both, using wireless or wired links to achieve data transmission. Terminal 300 may be a smartphone, tablet computer, laptop computer, desktop computer, set-top box, smart voice interaction device, smart home appliance, vehicle terminal, aircraft, portable music player, personal digital assistant, dedicated messaging device, portable gaming device, smart speaker, and smartwatch, etc., but is not limited to these.

[0042] The multimedia information processing method provided in the embodiments of this application will be described below. In actual implementation, the multimedia information processing method provided in the embodiments of this application can be implemented by a terminal or a server alone, or by a terminal and a server working together, so that... Figure 1 The following description uses the example of server 100 executing the multimedia information processing method provided in this application embodiment alone. See also... Figure 2 , Figure 2This is a flowchart illustrating a multimedia information processing method provided in an embodiment of this application. The method includes:

[0043] S201, acquire the original multimedia information, editing instructions for the original multimedia information, and a target multimedia information processing model; the target multimedia information processing model includes a noise-adding sub-model and a noise-reducing sub-model, the noise-reducing sub-model is established based on a first target processing factor, the first target processing factor is obtained by reconstructing the original processing factor, the denominator of the original processing factor includes a first weight coefficient corresponding to the original multimedia information, the denominator of the first target processing factor does not include the first weight coefficient, the first weight coefficient represents the proportion of the original multimedia information in the noise-added multimedia information during the noise-adding process;

[0044] In some embodiments, the original multimedia information can be various multimedia information such as images, videos, audio, and text. The original multimedia information may include original multimedia data and descriptive information about the original multimedia data. Editing instructions for the original multimedia information may be descriptive information about the target multimedia information to be generated. The target multimedia information processing model processes the input original multimedia information and the editing instructions for the original multimedia information to obtain the target multimedia information corresponding to the original multimedia information. For example, when the original multimedia data is an image, the descriptive information for the original image is "a woman," and the editing instruction for the original image is "an angry woman." This means that by inputting the original image, "a woman," and "an angry woman" into the target multimedia information processing model, the original image representing "a woman" can be edited to obtain the target image representing "an angry woman."

[0045] In some embodiments, the target multimedia information processing model includes a noise-adding sub-model and a noise-reducing sub-model. The noise-adding sub-model corresponds to the noise-adding process, and the noise-reducing sub-model corresponds to the noise-reducing process. The noise-reducing sub-model is established based on a first target processing factor. The first target processing factor is obtained by reconstructing the original processing factor. The original processing factor is a factor contained in a preset noise-reducing equation related to the noise-reducing process. The denominator of the original processing factor includes a first weight coefficient corresponding to the original multimedia information. The denominator of the first target processing factor does not include the first weight coefficient corresponding to the original multimedia information. The first weight coefficient corresponding to the original multimedia information represents the proportion of the original multimedia information in the noise-added multimedia information during the noise-adding process, that is, the retention ratio of the original multimedia information during the noise-adding process. During the noise addition process, noise is gradually added to the original multimedia information to obtain the noisy multimedia information corresponding to the original multimedia information. That is, the original multimedia information is gradually transformed into noise. Therefore, in the noisy multimedia information, the first weight coefficient corresponding to the original multimedia information tends to 0. That is, the denominator of the original processing factor tends to 0 in the initial stage of denoising. This results in a large step size of the numerical solution of the preset denoising equation, which leads to a large error in the initial stage of denoising and affects the quality of the final generated target multimedia information. The embodiments of this application reconstruct the original processing factor so that the denominator of the reconstructed target processing factor does not contain the first weight coefficient corresponding to the original multimedia information. This can reduce the error in the initial stage of denoising and improve the quality of the target multimedia information generated by the denoising process.

[0046] For example, the preset denoising equation includes an original processing factor that is the ratio of a second weighting coefficient to a first weighting coefficient. The first weighting coefficient represents the proportion of the original multimedia information in the noisy multimedia information during the noise addition process, and the second weighting coefficient represents the cumulative intensity of the first noise added during the noise addition process. In the initial stage of denoising, the proportion of the original multimedia information in the noisy multimedia information tends to 0, that is, in the initial stage of denoising, the denominator of the original processing factor tends to 0. The first target processing factor obtained by reconstructing the original processing factor is the ratio of the first weighting coefficient to the second weighting coefficient. In the initial stage of denoising, although the first weighting coefficient tends to 0, the denominator of the first target processing factor does not include the first weighting coefficient, that is, in the initial stage of denoising, the denominator of the first target processing factor will not tend to 0.

[0047] S203, input the original multimedia information into the noise-adding sub-model, perform step-by-step noise-adding processing on the original multimedia information, and obtain the noise-adding multimedia information corresponding to the original multimedia information;

[0048] In some embodiments, during the noise-adding stage, the original multimedia information is input into a noise-adding sub-model. The noise-adding sub-model is used to progressively add noise to the original multimedia information, resulting in the noise-added multimedia information corresponding to each noise-adding time step in a preset noise-adding time sequence. The noise-added multimedia information corresponding to the last noise-adding time step in the preset noise-adding time sequence is determined as the noise-added multimedia information corresponding to the original multimedia information. The preset noise-adding time sequence includes multiple noise-adding time steps. For each noise-adding time step, the noise-adding sub-model is used to add noise to the noise-added multimedia information corresponding to that time step, resulting in the noise-added multimedia information corresponding to the next noise-adding time step.

[0049] See Figure 3 , Figure 3 This is a flowchart illustrating a noise-adding process provided in an embodiment of this application. The preset noise-adding time sequence is {t0,...,t...} i-1 ,t i ,...,t N}, where i = 1, ..., N, t0 ~ t N All are noisy time steps. The original multimedia information is x0, and the noisy multimedia information corresponding to the original multimedia information is x. N During the noise addition process, x0~x N Both can represent noisy multimedia information. Here, x0 represents the original multimedia information, which does not contain noise. Therefore, the noisy multimedia information corresponding to the noisy time step t0 can be considered the original multimedia information; that is, for the noisy time step t0, the proportion of the original multimedia information in the noisy multimedia information is 1. For the noisy time step t... i-1 The corresponding noisy multimedia information is x. i-1 Using the noise-added sub-model to analyze x i-1 By performing noise addition processing, the next noise addition time step t can be obtained. i Corresponding noisy multimedia information x i .

[0050] It should be noted that the number of noise additions, N, can be set according to actual needs. To ensure the quality of the final generated target multimedia information, the number of noise additions can be set to N≥10. Furthermore, any two adjacent noise addition time steps in the preset noise addition time sequence can be either adjacent or non-adjacent. For example, the preset noise addition time sequence could be L1={0,1,2,...,99,100}, in which case N=100, t 100 The preset noise-adding time sequence is 100, meaning any two adjacent noise-adding time steps in the preset noise-adding time sequence are considered adjacent time steps, requiring 100 noise additions. Alternatively, the preset noise-adding time sequence can be L2 = {0, 10, 20, ..., 90, 100}, in which case N = 10, t 10With a value of 100, any two adjacent noise-adding time steps in the preset noise-adding time sequence are non-adjacent time steps. L2 only needs to add noise 10 times to achieve a noise-adding effect similar to L1 adding noise 100 times. That is, L2 can achieve skip-step noise addition and speed up the noise addition speed.

[0051] S205, the noisy multimedia information and the editing instructions are input into the denoising sub-model, and the noisy multimedia information is processed step by step to obtain the target multimedia information corresponding to the original multimedia information.

[0052] In some embodiments, during the denoising stage, the noisy multimedia information and editing instructions are input into a denoising sub-model. The denoising sub-model is used to progressively denoise the noisy multimedia information to obtain the denoised multimedia information corresponding to each denoising time step in a preset denoising time sequence. The denoised multimedia information corresponding to the last denoising time step in the preset denoising time sequence is determined as the target multimedia information corresponding to the original multimedia information. The preset denoising time sequence includes multiple denoising time steps. For each denoising time step, the denoising sub-model is used to denoise the denoised multimedia information corresponding to that denoising time step to obtain the denoised multimedia information corresponding to the next denoising time step.

[0053] See Figure 4 , Figure 4 This is a flowchart illustrating a denoising process provided in an embodiment of this application. The preset denoising time series is {t}. N , ..., t i , t i-1 , ..., t0}, where i = 1, ..., N, t N ~t0 represent denoising time steps. The preset denoising time sequence includes the same denoising time steps as the preset denoising time sequence, but in reverse order. The denoised multimedia information corresponding to the original multimedia information is x. N The target multimedia information corresponding to the original multimedia information is x'0. During the denoising process, x... N ~x'0 can all represent denoised multimedia information, where x N As noisy multimedia information, it does not contain the original multimedia information, and the denoising time step t can be considered as... N The corresponding denoised multimedia information is the same as the noisy multimedia information corresponding to the original multimedia information. For the denoising time step t... i The corresponding denoised multimedia information is x' i Using a denoising sub-model to analyze x' i After denoising, the next denoising time step t can be obtained. i-1 Corresponding denoised multimedia information x' i-1 .

[0054] It should be noted that the number of denoising iterations N can be set according to actual needs. To ensure the quality of the final generated target multimedia information, the number of denoising iterations can be set to N≥10. Furthermore, any two adjacent denoising time steps in the preset denoising time sequence can be either adjacent or non-adjacent. For example, the preset denoising time sequence could be L'1={100, 99, ..., 2, 1, 0}, in which case N=100, t 100 The preset denoising time series is 100, meaning any two adjacent denoising time steps in the preset denoising time series are considered adjacent time steps, requiring 100 denoising iterations. Alternatively, the preset denoising time series can be L'2 = {100, 90, ..., 20, 10, 0}, in which case N = 10, t 10 With a value of 100, any two adjacent denoising time steps in the preset denoising time sequence are non-adjacent time steps. L'2 only needs to denoise 10 times to achieve a denoising effect similar to that of L'1 which needs to denoise 100 times. In other words, L'2 can achieve skip-step denoising and speed up the denoising speed.

[0055] In some embodiments, see Figure 5 , Figure 5 This is a flowchart illustrating a method for determining a denoising sub-model provided in an embodiment of this application. Before obtaining the original multimedia information, the editing instructions for the original multimedia information, and the target multimedia information processing model, the method further includes:

[0056] S501, Obtain the preset denoising equation corresponding to the denoising sub-model; The preset denoising equation includes the original processing factor;

[0057] In some embodiments, the preset denoising equation corresponding to the denoising sub-model is the deterministic sampling ordinary differential equation of the diffusion model, that is:

[0058]

[0059] in, x(t) represents the denoised multimedia information at time step t, q(x(t), t) represents the distribution of x(t), and sampling q(x(t), t) yields x(t). Then f(t) and g 2 Substituting (t) into the above equation, we get:

[0060]

[0061] in, α is the original treatment factor. t σ represents the first weighting coefficient corresponding to the original multimedia information, characterizing the proportion of the original multimedia information in the noisy multimedia information during the noise addition process. tThe second weighting coefficient represents the cumulative intensity of the first noise added during the noise addition process, and ∈ θ A preset neural network is used to predict the second noise corresponding to each denoising time step in the preset denoising time series. The second noise corresponding to each denoising time step is the predicted value of the first noise added at each denoising time step in the preset denoising time series.

[0062] S503, the preset denoising equation is reconstructed, and the original processing factor is converted into the first target processing factor to obtain a target denoising equation containing the first target processing factor; the original processing factor is the ratio of the second weighting coefficient to the first weighting coefficient, the first target processing factor is the ratio of the first weighting coefficient to the second weighting coefficient, and the second weighting coefficient represents the cumulative intensity of the first noise added during the noise addition process.

[0063] In some embodiments, the chain rule can be used to reconstruct the preset denoising equation, thereby... Convert to As the first objective processing factor, the reconstructed objective denoising equation is:

[0064]

[0065] in, x t This is the abbreviated form of x(t). α t and σ t It is a function related to the noise addition time step t, and the specific functional relationship can be set according to actual needs.

[0066] S505, interpolate the target denoising equation to obtain the denoising sub-model.

[0067] In some embodiments, after obtaining the reconstructed target denoising equation, a denoising sub-model representing the relationship between adjacent denoising time steps can be obtained by interpolating the target denoising equation. For example, the Lagrange interpolation method can be used to interpolate the target denoising equation to obtain the denoising sub-model.

[0068] This application embodiment reconstructs the original processing factors by swapping the denominator and numerator, so that the reconstructed processing factors do not have the problem of the denominator tending to 0 in the initial stage of denoising. That is, through simple factor reconstruction processing, the error in the initial stage of denoising is reduced and the quality of the target multimedia information generated by the denoising process is improved without increasing the computational complexity.

[0069] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of an interpolation process provided in an embodiment of this application. The target denoising equation includes multimedia information parameters, which represent the denoised multimedia information corresponding to different denoising time steps. The interpolation process on the target denoising equation to obtain the denoising sub-model includes:

[0070] S601, Based on the first target processing factor and the second target processing factor contained in the target denoising equation, a target interpolation polynomial is established; the second target processing factor is the ratio of the multimedia information parameter to the second weight coefficient;

[0071] In some embodiments, the target denoising equation includes multimedia information parameters x. t Multimedia information parameter x t The second target processing factor in the target denoising equation represents the denoised multimedia information corresponding to the denoising time step t. Based on the first target processing factor and the second target treatment factor A target interpolation polynomial can be established. Specifically, let... The Lagrange interpolation method is used to interpolate the target denoising equation. Let p(λ) be the value of the target denoising equation in z. t-1 z t z t+1 The Lagrange interpolation is then:

[0072]

[0073] p(λ) is the target interpolation polynomial.

[0074] S603, based on the first derivative of the target interpolation polynomial and the target denoising equation, the denoising sub-model is obtained; the denoising sub-model represents the relationship between the denoised multimedia information corresponding to the next denoising time step, the denoised multimedia information corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step during the denoising process.

[0075] In some embodiments, the target denoising equation can be rearranged to obtain:

[0076]

[0077] It can be observed that, It can be equivalent to the first derivative p′(λ) of the target interpolation polynomial at λ=λ i The expression for time is given, therefore we can obtain:

[0078]

[0079] Let δ1 = λi -λ i-1 δ2=λ i+1 -λ i δ3=λ i+1 -λ i-1 ,but:

[0080]

[0081] x is obtained by rearranging terms. i-1 With x i x i+1 The relationship between them:

[0082]

[0083] By unifying the denoising time step in the above formula to i, we can obtain the denoising sub-model:

[0084]

[0085] Where, x i x represents the denoised multimedia information corresponding to the current denoising time step. i+1 x represents the denoised multimedia information corresponding to the previous denoising time step. i-1 This is the denoised multimedia information corresponding to the next denoising time step.

[0086] This application embodiment performs interpolation processing on the reconstructed target denoising equation. Based on the mathematical relationship between the target denoising equation and the interpolation polynomial, a denoising sub-model representing the relationship between adjacent terms of denoised multimedia information can be obtained, thereby improving the generation efficiency of the denoising sub-model. Moreover, since the target processing factor in the reconstructed target denoising equation will not have the problem of the denominator tending to 0 in the initial stage of denoising, the processing error in the initial stage of denoising can be reduced, thereby improving the quality of the target multimedia information generated through the denoising process.

[0087] In some embodiments, see Figure 7 , Figure 7 This is a flowchart illustrating a method for determining a denoising sub-model according to an embodiment of this application. After obtaining the denoising sub-model based on the first derivative of the target interpolation polynomial and the target denoising equation, the method further includes:

[0088] S701, the denoised multimedia information corresponding to the previous denoising time step in the denoising sub-model is transposed to obtain the denoising sub-model; the denoising sub-model represents the relationship between the denoised multimedia information corresponding to the next denoising time step, the denoised multimedia information corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step during the denoising process.

[0089] In some embodiments, by analyzing the denoised multimedia information x corresponding to the previous denoising time step in the denoising sub-model i+1 By performing a rearrangement process, we can obtain the noise-added sub-model:

[0090]

[0091] During the noise addition process, the order of the noise addition time steps is the reverse of the denoising time steps, therefore x i x represents the noisy multimedia information corresponding to the current noisy time step. i-1 x represents the noisy multimedia information corresponding to the previous noisy time step. i+1 This is the noisy multimedia information corresponding to the next noisy time step.

[0092] It should be noted that the noise-adding sub-model involves x i x i-1 and x i+1 That is, we need to substitute x. i With x i-1 Only then can we obtain x i+1 In the noise addition process, for the first noise addition, the only known quantity is the original multimedia information x0, and there is no noised multimedia information corresponding to the previous noise addition time step. In this case, the Denoising Diffusion Implicit Models (DDIM) can be used to first obtain the noised multimedia information x1 corresponding to the next noise addition time step based on the original multimedia information x0. The noise addition process of the DDIM model can be expressed as:

[0093]

[0094] Then, given x0 and x1, the noise-adding sub-model in the embodiments of this application is used to obtain the noise-adding multimedia information corresponding to the subsequent noise-adding time steps.

[0095] S703, Based on the noise-adding sub-model and the noise-reducing sub-model, the target multimedia information processing model is obtained.

[0096] In some embodiments, the target multimedia information processing model includes a noise-adding sub-model and a noise-reducing sub-model. After obtaining the noise-adding sub-model and the noise-reducing sub-model, a target multimedia information processing model for processing multimedia information can be established.

[0097] In this embodiment, after obtaining the denoising sub-model, a denoising sub-model can be further obtained through a term-shifting process. This means the denoising sub-model and the denoising sub-model are precisely reversible, ensuring consistency between the unedited portions before and after information processing and improving the quality of the final generated target multimedia information. For example, in image stylization editing scenarios, the target multimedia information processing model ensures that only the style is edited, while maintaining the rest of the image structure unchanged; in image entity editing scenarios, the target multimedia information processing model ensures that only target entities are edited, and non-editable targets remain unmodified.

[0098] In some embodiments, see Figure 8 , Figure 8 This is a flowchart illustrating a method for generating target multimedia information according to an embodiment of this application. The method involves inputting the noisy multimedia information and the editing instructions into the denoising sub-model, and performing progressive denoising processing on the noisy multimedia information to obtain the target multimedia information corresponding to the original multimedia information. The steps include:

[0099] S801, the editing instruction is input into a preset neural network, and the second noise corresponding to each denoising time step in the preset denoising time series is predicted based on the preset neural network; the second noise corresponding to each denoising time step is the predicted value of the first noise added at each denoising time step in the preset denoising time series;

[0100] In some embodiments, during the denoising process, it is necessary to predict the second noise corresponding to each denoising time step in the preset denoising time series based on a preset neural network, i.e., ∈ θ (x i In each denoising time step, the second noise is the predicted value of the first noise added at each denoising time step in the preset denoising time series. For example, the preset denoising time series is {t... N ,...,t i ,t i-1 The preset noisy time series is {t0,...,t}. i-1 ,t i ,...,t N In the denoising process, the denoising time step t in the preset denoising time series is predicted based on the preset neural network. i The corresponding second noise is the noise addition time step t in the preset noise addition time sequence during the noise addition process. i The predicted value of the first noise added.

[0101] S803, based on the second noise corresponding to each time step, the noisy multimedia information, and the denoising sub-model, the denoised multimedia information corresponding to each denoising time step is obtained;

[0102] In some embodiments, in order to obtain the denoised multimedia information x corresponding to the next denoising time step i-1 It is necessary to obtain the denoised multimedia information x corresponding to the current denoising time step. i The denoised multimedia information x corresponding to the previous denoising time step i+1 and the second noise corresponding to the current denoising time step ∈ θ (x i i) Input the denoising sub-model.

[0103] S805, the denoised multimedia information corresponding to the last denoised time step in the preset denoising time sequence is determined as the target multimedia information.

[0104] In some embodiments, the preset denoising time series is {t} N , ..., t i , t i-1 Therefore, the denoised multimedia information corresponding to the denoising time step t0 is the target multimedia information corresponding to the original multimedia information, that is, the target multimedia information obtained after editing the original multimedia information according to the editing instructions.

[0105] This embodiment of the application reconstructs the ordinary differential equations involved in the denoising process to obtain a denoising sub-model. The noisy multimedia information corresponding to the original multimedia information and editing instructions are input into the denoising sub-model. Combined with noise predicted by a preset neural network, the noisy multimedia information is progressively denoised, ultimately obtaining the edited target multimedia information. Since the target processing factor in the reconstructed ordinary differential equation does not exhibit the problem of the denominator approaching 0 in the initial stage of denoising, the processing error in the initial stage of denoising can be reduced, improving the quality of the target multimedia information generated through the denoising process.

[0106] In some embodiments, see Figure 9 , Figure 9 This is a flowchart illustrating a second noise determination method provided in an embodiment of this application. The step of inputting the editing instruction into a preset neural network and predicting the second noise corresponding to each denoising time step in a preset denoising time series based on the preset neural network includes:

[0107] S901, Obtain the denoised multimedia information corresponding to the current denoising time step;

[0108] In some embodiments, for the current denoising time step t i The corresponding denoised multimedia information is x. i .

[0109] S903, the current denoising time step, the denoised multimedia information corresponding to the current denoising time step, and the editing instruction are input into the preset neural network to obtain the second noise corresponding to the current denoising time step.

[0110] In some embodiments, the current denoising time step t i The denoised multimedia information x corresponding to the current denoising time step i and editing instructions input preset neural network ∈ θ Preset neural network ∈ θ The output is the current denoising time step t. i The corresponding second noise.

[0111] It should be noted that the second noise refers to the gradient of the logarithmic probability distribution of the denoised multimedia information, also known as the score, i.e., ∈ θ (x i i) The prediction is

[0112] In some embodiments, the preset neural network can be obtained through training, and the training process includes:

[0113] 1) Sample the original multimedia information x0~q(x0);

[0114] 2) Select a specific step t ~ U({1, 2, ..., T}) during the diffusion process;

[0115] 3) Add Gaussian noise ∈ ~N(0, I);

[0116] 4) Estimating noise

[0117] 5) Learn the loss on the network using gradient descent.

[0118] For example, the default neural network could be a U-Net network.

[0119] This embodiment of the application reconstructs the ordinary differential equations involved in the denoising process to obtain a denoising sub-model. The noisy multimedia information corresponding to the original multimedia information and editing instructions are input into the denoising sub-model. Combined with noise predicted by a preset neural network, the noisy multimedia information is progressively denoised, ultimately obtaining the edited target multimedia information. Since the target processing factor in the reconstructed ordinary differential equation does not exhibit the problem of the denominator approaching 0 in the initial stage of denoising, the processing error in the initial stage of denoising can be reduced, improving the quality of the target multimedia information generated through the denoising process.

[0120] In some embodiments, see Figure 10 , Figure 10This is a flowchart illustrating a method for generating denoised multimedia information according to an embodiment of this application. The step of obtaining the denoised multimedia information corresponding to each denoised time step based on the second noise corresponding to each time step, the noisy multimedia information, and the denoising sub-model includes:

[0121] S1001, Determine the current denoising time step based on the preset denoising time sequence;

[0122] In some embodiments, the preset denoising time series is {t} N ,...,t i ,t i-1 ,...,t0}, the noisy multimedia information x corresponding to the original multimedia information obtained during the noisy process. N This can be viewed as the denoised multimedia information corresponding to the first denoising time step in the denoising process, at which point the current denoising time step is t. N .

[0123] S1003, based on the denoised multimedia information corresponding to the current denoising time step, the second noise corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step, perform denoising processing on the denoised multimedia information corresponding to the current denoising time step to obtain the denoised multimedia information corresponding to the next denoising time step.

[0124] In some embodiments, the current denoising time step t N Corresponding denoised multimedia information x N The denoised multimedia information corresponding to the previous denoising time step and the current denoising time step t N The corresponding second noise ∈ θ (x N By inputting the denoising sub-model (N), the next denoising time step t can be obtained. N-1 Corresponding denoised multimedia information x N-1 .

[0125] It should be noted that, due to t N As the first denoising time step in the preset denoising time series, x does not actually exist. N+1 Therefore, in this case, the noisy multimedia information x corresponding to the original multimedia information obtained in the noisy process can be used as the final noisy multimedia information. N The previous noise-adding time step t N-1 The corresponding noisy multimedia information is considered as x N+1 .

[0126] S1005, the next denoising time step is determined as the current denoising time step;

[0127] In some embodiments, the next denoising time step t is obtained. N-1 Corresponding denoised multimedia information x N-1 Next, the current denoising time step needs to be updated, that is, the next denoising time step t needs to be updated. N-1 The current denoising time step is determined, and the denoised multimedia information corresponding to each denoising time step in the preset denoising time sequence can be obtained by looping.

[0128] S1007, Repeated execution: Based on the denoised multimedia information corresponding to the current denoising time step, the second noise corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step, perform denoising processing on the denoised multimedia information corresponding to the current denoising time step to obtain the denoised multimedia information corresponding to the next denoising time step, until the next denoising time step is determined as the current denoising time step, until the current denoising time step is the last denoising time step in the preset denoising time sequence, to obtain the denoised multimedia information corresponding to each denoising time step.

[0129] In some embodiments, steps S1003 to S1005 are repeated until the updated current denoising time step is the last denoising time step t0 in the preset denoising time sequence. Since x0 can be obtained based on x2 and x1, when the current denoising time step is the last denoising time step in the preset denoising time sequence, the denoised multimedia information corresponding to each denoising time step in the preset denoising time sequence can be obtained. The denoised multimedia information corresponding to the last denoising time step in the preset denoising time sequence is the target multimedia information corresponding to the original multimedia information.

[0130] This embodiment of the application reconstructs the ordinary differential equations involved in the denoising process to obtain a denoising sub-model. The noisy multimedia information corresponding to the original multimedia information and editing instructions are input into the denoising sub-model. Combined with noise predicted by a preset neural network, the noisy multimedia information is progressively denoised, ultimately obtaining the edited target multimedia information. Since the target processing factor in the reconstructed ordinary differential equation does not exhibit the problem of the denominator approaching 0 in the initial stage of denoising, the processing error in the initial stage of denoising can be reduced, improving the quality of the target multimedia information generated through the denoising process.

[0131] This application provides a multimedia information processing method, which includes: acquiring original multimedia information, editing instructions for the original multimedia information, and a target multimedia information processing model; the target multimedia information processing model includes a noise-adding sub-model and a noise-reducing sub-model, the noise-reducing sub-model being established based on a first target processing factor, the first target processing factor being obtained by reconstructing an original processing factor, the denominator of the original processing factor including a first weight coefficient corresponding to the original multimedia information, and the denominator of the first target processing factor not including the first weight coefficient, the first weight coefficient representing the proportion of the original multimedia information in the noise-adding multimedia information during the noise-adding process; inputting the original multimedia information into the noise-adding sub-model to perform stepwise noise-adding processing on the original multimedia information to obtain noise-added multimedia information corresponding to the original multimedia information; inputting the noise-added multimedia information and the editing instructions into the noise-reducing sub-model to perform stepwise noise-reducing processing on the noise-added multimedia information to obtain the target multimedia information corresponding to the original multimedia information. This method reconstructs the original processing factors included in the denoising ordinary differential equation related to the denoising process, obtaining a first target processing factor whose denominator does not contain the first weight coefficient corresponding to the original multimedia information. A target multimedia information processing model is established based on this first target processing factor. Inputting the original multimedia information and editing instructions into the target multimedia information processing model yields the edited target multimedia information. Since the first weight coefficient represents the proportion of the original multimedia information in the already denoised multimedia information during the denoising process, reconstructing the original processing factors ensures that the denominator of the reconstructed first target processing factor does not contain the first weight coefficient. This avoids the problem of a large step size in the numerical solution of the denoising ordinary differential equation due to the first weight coefficient approaching 0 in the initial stage of denoising, thereby reducing the error in the initial stage of denoising and improving the quality of the edited target multimedia information generated through the denoising process.

[0132] This application also provides a multimedia information processing device, see [link to relevant documentation] Figure 11 The device includes:

[0133] The acquisition module 1110 is used to acquire original multimedia information, editing instructions for the original multimedia information, and a target multimedia information processing model. The target multimedia information processing model includes a noise-adding sub-model and a noise-reducing sub-model. The noise-reducing sub-model is established based on a first target processing factor. The first target processing factor is obtained by reconstructing the original processing factor. The denominator of the original processing factor includes a first weight coefficient corresponding to the original multimedia information. The denominator of the first target processing factor does not include the first weight coefficient. The first weight coefficient represents the proportion of the original multimedia information in the noisy multimedia information during the noise-adding process.

[0134] The noise-adding module 1120 is used to input the original multimedia information into the noise-adding sub-model, perform step-by-step noise-adding processing on the original multimedia information, and obtain the noise-adding multimedia information corresponding to the original multimedia information.

[0135] The denoising module 1130 is used to input the noisy multimedia information and the editing instructions into the denoising sub-model, and to perform step-by-step denoising processing on the noisy multimedia information to obtain the target multimedia information corresponding to the original multimedia information.

[0136] In some embodiments, the apparatus further includes:

[0137] A preset denoising equation acquisition module is used to acquire the preset denoising equation corresponding to the denoising sub-model; the preset denoising equation includes the original processing factor.

[0138] The reconstruction module is used to reconstruct the preset denoising equation, convert the original processing factor into the first target processing factor, and obtain a target denoising equation containing the first target processing factor; the original processing factor is the ratio of the second weighting coefficient to the first weighting coefficient, the first target processing factor is the ratio of the first weighting coefficient to the second weighting coefficient, and the second weighting coefficient represents the cumulative intensity of the first noise added during the noise addition process;

[0139] The interpolation module is used to interpolate the target denoising equation to obtain the denoising sub-model.

[0140] In some embodiments, the target denoising equation includes multimedia information parameters, which characterize the denoised multimedia information corresponding to different denoising time steps, and the interpolation module includes:

[0141] An interpolation polynomial establishment unit is used to establish a target interpolation polynomial based on the first target processing factor and the second target processing factor contained in the target denoising equation; the second target processing factor is the ratio of the multimedia information parameter to the second weighting coefficient;

[0142] The denoising sub-model determination unit is used to obtain the denoising sub-model based on the first derivative of the target interpolation polynomial and the target denoising equation; the denoising sub-model represents the relationship between the denoised multimedia information corresponding to the next denoising time step, the denoised multimedia information corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step during the denoising process.

[0143] In some embodiments, the interpolation module further includes:

[0144] The noise-adding sub-model determination unit is used to perform a term shifting process on the denoised multimedia information corresponding to the previous denoising time step in the denoising sub-model to obtain the noise-adding sub-model; the noise-adding sub-model represents the relationship between the denoised multimedia information corresponding to the next denoising time step, the denoised multimedia information corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step during the noise-adding process.

[0145] The target multimedia information processing model determination unit is used to obtain the target multimedia information processing model based on the noise-adding sub-model and the noise-reducing sub-model.

[0146] In some embodiments, the noise reduction module 1130 includes:

[0147] The second noise prediction unit is used to input the editing instructions into a preset neural network and predict the second noise corresponding to each denoising time step in the preset denoising time series based on the preset neural network; the second noise corresponding to each denoising time step is the predicted value of the first noise added at each denoising time step in the preset denoising time series.

[0148] The denoised multimedia information determination unit is used to obtain the denoised multimedia information corresponding to each denoised time step based on the second noise corresponding to each time step, the noisy multimedia information, and the denoising sub-model.

[0149] The target multimedia information determination unit is used to determine the denoised multimedia information corresponding to the last denoised time step in the preset denoised time sequence as the target multimedia information.

[0150] In some embodiments, the second noise prediction unit includes:

[0151] The denoised multimedia information acquisition subunit is used to acquire the denoised multimedia information corresponding to the current denoising time step;

[0152] The second noise determination subunit is used to input the current denoising time step, the denoised multimedia information corresponding to the current denoising time step, and the editing instruction into the preset neural network to obtain the second noise corresponding to the current denoising time step.

[0153] In some embodiments, the noise-reduced multimedia information determination unit includes:

[0154] The current denoising time step determination subunit is used to determine the current denoising time step based on the preset denoising time sequence;

[0155] The denoising subunit is used to perform denoising processing on the denoised multimedia information corresponding to the current denoising time step based on the denoised multimedia information corresponding to the current denoising time step, the second noise corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step, to obtain the denoised multimedia information corresponding to the next denoising time step.

[0156] The current denoising time step update subunit is used to determine the next denoising time step as the current denoising time step;

[0157] The loop subunit is used to repeatedly execute: based on the denoised multimedia information corresponding to the current denoising time step, the second noise corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step, to perform denoising processing on the denoised multimedia information corresponding to the current denoising time step to obtain the denoised multimedia information corresponding to the next denoising time step, until the next denoising time step is determined as the current denoising time step, until the current denoising time step is the last denoising time step in the preset denoising time sequence, to obtain the denoised multimedia information corresponding to each denoising time step.

[0158] The apparatus provided in the above embodiments can execute the method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the above embodiments can be found in a multimedia information processing method provided in any embodiment of this application.

[0159] This embodiment also provides a computer-readable storage medium storing computer-executable instructions, which are loaded by a processor and executed by the multimedia information processing method described above in this embodiment.

[0160] This embodiment also provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program adapted to be loaded by the processor and executed by the multimedia information processing method described above in this embodiment.

[0161] The electronic device may be a computer terminal, a mobile terminal, or a server, and may also participate in constituting the apparatus or system provided in the embodiments of this application. For example... Figure 12As shown, the electronic device 12 may include one or more processors 1202 (shown as 1202a, 1202b, ..., 1202n in the figure) (processor 1202 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPLD, etc.), a memory 1204 for storing information, and a transmission device 1206 for communication functions. In addition, it may also include: input / output interfaces (I / O interfaces) and network interfaces. Those skilled in the art will understand that... Figure 12 The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, electronic device 12 may also include components that are more... Figure 12 The more or fewer components shown, or having the same Figure 12 The different configurations shown.

[0162] It should be noted that the aforementioned one or more processors 1202 and / or other information processing circuits are generally referred to herein as "information processing circuits". These information processing circuits may be wholly or partially embodied in software, hardware, firmware, or any other combination thereof. Furthermore, the information processing circuits may be a single, independent processing module, or may be wholly or partially integrated into any other element within the electronic device 12.

[0163] The memory 1204 can be used to store software programs and modules of application software, such as the program instruction / information storage device corresponding to the method described in the embodiments of this application. The processor 1202 executes various functional applications and information processing by running the software programs and modules stored in the memory 1204, thereby realizing the above-mentioned multimedia information processing method. The memory 1204 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1204 may further include memory remotely located relative to the processor 1202, and these remote memories can be connected to the electronic device 12 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0164] The transmission device 1206 is used to receive or send information via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 12. In one example, the transmission device 1206 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 1206 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0165] This specification provides the operational steps of the methods described in the embodiments or flowcharts, but more or fewer operational steps may be included based on conventional or non-inventive labor. The steps and order listed in the embodiments are merely one possible execution order among many steps and do not represent the only execution order. In actual system or interrupt product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0166] The structure shown in this embodiment is only a partial structure related to the solution of this application and does not constitute a limitation on the device to which the solution of this application is applied. Specific devices may include more or fewer components than shown, or combinations of certain components, or arrangements of different components. It should be understood that the methods, apparatuses, etc., disclosed in this embodiment can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or unit modules through some interfaces.

[0167] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0168] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0169] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A multimedia information processing method, characterized in that, The method includes: The system acquires original multimedia information, editing instructions for the original multimedia information, and a target multimedia information processing model. The target multimedia information processing model includes a noise-adding sub-model and a noise-reducing sub-model. The noise-reducing sub-model is established based on a first target processing factor, which is obtained by reconstructing the original processing factor. The denominator of the original processing factor includes a first weight coefficient corresponding to the original multimedia information, while the denominator of the first target processing factor does not include the first weight coefficient. The first weight coefficient represents the proportion of the original multimedia information in the noisy multimedia information during the noise-adding process. The original multimedia information is input into the noise-adding sub-model, and the original multimedia information is subjected to step-by-step noise-adding processing to obtain the noise-adding multimedia information corresponding to the original multimedia information. The noisy multimedia information and the editing instructions are input into the denoising sub-model to perform step-by-step denoising processing on the noisy multimedia information, thereby obtaining the target multimedia information corresponding to the original multimedia information.

2. The multimedia information processing method according to claim 1, characterized in that, Before acquiring the original multimedia information, the editing instructions for the original multimedia information, and the target multimedia information processing model, the method further includes: Obtain the preset denoising equation corresponding to the denoising sub-model; the preset denoising equation includes the original processing factor; The preset denoising equation is reconstructed to convert the original processing factor into the first target processing factor, resulting in a target denoising equation containing the first target processing factor. The original processing factor is the ratio of the second weighting coefficient to the first weighting coefficient, and the first target processing factor is the ratio of the first weighting coefficient to the second weighting coefficient. The second weighting coefficient represents the cumulative intensity of the first noise added during the noise addition process. The target denoising equation is interpolated to obtain the denoising sub-model.

3. The multimedia information processing method according to claim 2, characterized in that, The target denoising equation includes multimedia information parameters, which characterize the denoised multimedia information at different denoising time steps. The interpolation process performed on the target denoising equation to obtain the denoising sub-model includes: Based on the first target processing factor and the second target processing factor contained in the target denoising equation, a target interpolation polynomial is established; the second target processing factor is the ratio of the multimedia information parameter to the second weight coefficient. Based on the first derivative of the target interpolation polynomial and the target denoising equation, the denoising sub-model is obtained; the denoising sub-model represents the relationship between the denoised multimedia information corresponding to the next denoising time step, the denoised multimedia information corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step during the denoising process.

4. The multimedia information processing method according to claim 3, characterized in that, After obtaining the denoising sub-model based on the first derivative of the target interpolation polynomial and the target denoising equation, the method further includes: The denoised multimedia information corresponding to the previous denoising time step in the denoising sub-model is rearranged to obtain the denoising sub-model; the denoising sub-model represents the relationship between the denoised multimedia information corresponding to the next denoising time step, the denoised multimedia information corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step during the denoising process. Based on the noise-adding sub-model and the noise-reducing sub-model, the target multimedia information processing model is obtained.

5. The multimedia information processing method according to claim 1, characterized in that, The step of inputting the noisy multimedia information and the editing instructions into the denoising sub-model, and performing stepwise denoising processing on the noisy multimedia information to obtain the target multimedia information corresponding to the original multimedia information includes: The editing instructions are input into a preset neural network, and the preset neural network is used to predict the second noise corresponding to each denoising time step in the preset denoising time series; the second noise corresponding to each denoising time step is the predicted value of the first noise added at each denoising time step in the preset denoising time series; Based on the second noise corresponding to each time step, the noisy multimedia information, and the denoising sub-model, the denoised multimedia information corresponding to each denoising time step is obtained; The denoised multimedia information corresponding to the last denoised time step in the preset denoising time sequence is determined as the target multimedia information.

6. The multimedia information processing method according to claim 5, characterized in that, The step of inputting the editing instructions into a preset neural network and predicting the second noise corresponding to each denoising time step in a preset denoising time series based on the preset neural network includes: Obtain the denoised multimedia information corresponding to the current denoising time step; The current denoising time step, the denoised multimedia information corresponding to the current denoising time step, and the editing instruction are input into the preset neural network to obtain the second noise corresponding to the current denoising time step.

7. The multimedia information processing method according to claim 5, characterized in that, The step of obtaining the denoised multimedia information corresponding to each denoised time step based on the second noise corresponding to each time step, the noisy multimedia information, and the denoising sub-model includes: The current denoising time step is determined based on the preset denoising time sequence; Based on the denoised multimedia information corresponding to the current denoising time step, the second noise corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step, the denoised multimedia information corresponding to the current denoising time step is denoised to obtain the denoised multimedia information corresponding to the next denoising time step. The next denoising time step is determined as the current denoising time step; Repeated execution: Based on the denoised multimedia information corresponding to the current denoising time step, the second noise corresponding to the current denoising time step, and the denoised multimedia information corresponding to the previous denoising time step, the denoised multimedia information corresponding to the current denoising time step is denoised to obtain the denoised multimedia information corresponding to the next denoising time step, until the next denoising time step is determined as the current denoising time step, until the current denoising time step is the last denoising time step in the preset denoising time sequence, to obtain the denoised multimedia information corresponding to each denoising time step.

8. A multimedia information processing device, characterized in that, The device includes: The acquisition module is used to acquire original multimedia information, editing instructions for the original multimedia information, and a target multimedia information processing model. The target multimedia information processing model includes a noise-adding sub-model and a noise-reducing sub-model. The noise-reducing sub-model is established based on a first target processing factor. The first target processing factor is obtained by reconstructing the original processing factor. The denominator of the original processing factor includes a first weight coefficient corresponding to the original multimedia information. The denominator of the first target processing factor does not include the first weight coefficient. The first weight coefficient represents the proportion of the original multimedia information in the noisy multimedia information during the noise-adding process. The noise-adding module is used to input the original multimedia information into the noise-adding sub-model, perform step-by-step noise-adding processing on the original multimedia information, and obtain the noise-adding multimedia information corresponding to the original multimedia information. The denoising module is used to input the noisy multimedia information and the editing instructions into the denoising sub-model, and to perform step-by-step denoising processing on the noisy multimedia information to obtain the target multimedia information corresponding to the original multimedia information.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the multimedia information processing method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the multimedia information processing method as described in any one of claims 1-7.