Multimedia file synthesis method and device, electronic equipment and storage medium

By decoding multimedia files to obtain audio and video frame data, analyzing audio waveform parameters to control video special effects, solving the problem of insufficient audio and video fusion in the prior art, realizing dynamic visual effects and efficient creation of multimedia files.

CN120264071APending Publication Date: 2025-07-04BEIJING TRICOLOR TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510408656.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing video processing technology is difficult to achieve deep integration of audio and video of multiple materials, cannot create unique visual effects, and cannot meet the market's demand for video picture quality and visual impact.

Method used

By decoding the multimedia file, obtaining video and audio frame data, constructing audio waveform data and analyzing their statistical parameters, controlling the video special effect output based on these parameters, synthesizing the target video frame data and combining it with the audio frame, and generating a multimedia file to dynamically display the special effects.

Benefits of technology

It enhances the interactivity and viewing of multimedia files, improves the efficiency and quality of video special effects generation, supports a variety of multimedia file formats, and promotes creation and personalized customization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120264071A_ABST
    Figure CN120264071A_ABST
Patent Text Reader

Abstract

The invention provides a multimedia file synthesis method and device, electronic equipment and a storage medium, and the method comprises the steps: decoding a first multimedia file to obtain video and audio frame data, constructing audio waveform data, and analyzing statistical parameters and a specific time interval average value of the audio waveform data; determining a control parameter to regulate and control the special effect output of the video frame, synthesizing target video data, combining the target frame and the audio frame to generate a second multimedia file, and dynamically displaying the special effect based on the audio content when the file is played; the video special effect can be controlled through audio data analysis, and the function of dynamically displaying rich visual effects according to the audio content when the multimedia file is played is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimedia technologies, and in particular, to a method, apparatus, electronic device, and storage medium for synthesizing multimedia files. Background Art

[0002] In today's booming video content market, the visual effects of video images directly determine their market popularity. Cool video images with strong visual impact can quickly capture the audience's attention in scenarios such as advertisements, movies, and game promotions, helping the content stand out in the fierce competition. Users' strong interest in exciting video clips further highlights the importance of high-quality video images.

[0003] However, there are obvious limitations in the current video display effects. Existing video processing technologies are difficult to meet the growing market demand in terms of improving image quality and enhancing visual impact. Especially in synthesizing various materials for audio-visual integration to create unique visual effects, there are significant technological gaps. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus, electronic device, and storage medium for synthesizing multimedia files, which can control video special effects through audio data analysis and realize the function of dynamically displaying rich visual effects according to audio content when playing multimedia files.

[0005] The technical solution of the embodiments of this application is implemented as follows:

[0006] In a first aspect, an embodiment of this application provides a method for synthesizing multimedia files, the method including:

[0007] Performing decoding processing on a first multimedia file to obtain video frame data and audio frame data;

[0008] Constructing audio waveform data through the audio frame data, and determining a statistical parameter set of the audio waveform data and the waveform average value in a specific time interval based on the audio waveform data; wherein, the statistical parameter set at least includes the mathematical mean, mathematical variance, and mathematical mean square deviation of the audio waveform data;

[0009] Determining a control parameter based on the audio frame data and the statistical parameter set; wherein, the control parameter is used to control the special effect output of the video frame data;

[0010] Constructing target frame data for video frame synthesis based on the statistical parameter set, the control parameter, and the waveform average value, and performing synthesis processing on the target frame data and the video frame data to obtain target video data; wherein, each frame in the target frame data corresponds one-to-one with each frame in the video frame data;

[0011] Merge the target frame data with the audio frame data to obtain a second multimedia file; wherein, special effects are displayed based on the audio frame data when the second multimedia file is played.

[0012] In a second aspect, an embodiment of the present application further provides a multimedia file synthesis device, and the device includes:

[0013] A decoding module, configured to perform decoding processing on a first multimedia file to obtain video frame data and audio frame data;

[0014] A construction module, configured to construct audio waveform data through the audio frame data, and determine a statistical parameter set of the audio waveform data and a waveform average value of a specific time interval based on the audio waveform data; wherein, the statistical parameter set at least includes the mathematical mean, mathematical variance, and mathematical mean square deviation of the audio waveform data;

[0015] A determination module, configured to determine a control parameter based on the audio frame data and the statistical parameter set; wherein, the control parameter is used to control the special effect output of the video frame data;

[0016] A synthesis module, configured to construct target frame data for video frame synthesis based on the statistical parameter set, the control parameter, and the waveform average value, and perform synthesis processing on the target frame data and the video frame data to obtain target video data; wherein, each frame in the target frame data corresponds one-to-one with each frame in the video frame data;

[0017] A merging module, configured to merge the target frame data with the audio frame data to obtain a second multimedia file; wherein, special effects are displayed based on the audio frame data when the second multimedia file is played.

[0018] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor, a storage medium, and a bus, where the storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to execute the multimedia file synthesis method according to any one of the first aspects.

[0019] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the multimedia file synthesis method according to any one of the first aspects.

[0020] The embodiments of the present application have the following beneficial effects:

[0021] The video and audio frame data are obtained by decoding the first multimedia file, and a waveform is constructed using the audio frame data and its statistical parameters and average values ​​of a specific time interval are analyzed. Based on this, control parameters are determined to adjust the special effects output of the video frame, and then the target video data containing these special effects are synthesized, and it is ensured that the target frame corresponds to the original video frame one by one. Finally, the target frame and the audio frame are merged to generate a second multimedia file, so that special effects can be dynamically displayed according to the audio content during playback, thereby enhancing the interactivity and viewing experience of the multimedia file. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 is a flowchart of steps S101-S105 provided in an embodiment of the present application;

[0024] Figure 2 It is a flowchart of steps S201-S202 provided in an embodiment of the present application;

[0025] Figure 3 It is a flowchart of steps S301-S302 provided in an embodiment of the present application;

[0026] Figure 4 It is a flowchart of steps S401-S402 provided in an embodiment of the present application;

[0027] Figure 5 is a schematic diagram of the structure of a multimedia file synthesis device provided in an embodiment of the present application;

[0028] Figure 6 It is a schematic diagram of the composition structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. It should be understood that the accompanying drawings in this application are only for the purposes of illustration and description, and are not used to limit the protection scope of this application. Additionally, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of this application. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application.

[0030] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0031] Furthermore, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. The components of the embodiments of this application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents the selected embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of this application.

[0032] In the following description, the terms "first", "second", and "third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first", "second", and "third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.

[0033] It should be noted that the term "including" will be used in the embodiments of this application to indicate the existence of the features stated thereafter, but does not exclude the addition of other features.

[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application and do not limit this application.

[0035] See Figure 1 , Figure 1It is a schematic flowchart of steps S101 - S105 of the multimedia file synthesis method provided by an embodiment of the present application, which will be described in conjunction with Figure 1 the steps S101 - S105 shown.

[0036] In step S101, the first multimedia file is decoded to obtain video frame data and audio frame data.

[0037] Here, decoding the first multimedia file is a basic step to obtain video frame data and audio frame data. The decoding process usually involves demultiplexing (Demux), that is, separating the bundled audio and video streams.

[0038] In some embodiments, the first multimedia file is respectively decoded for video and audio based on the FFmpeg codec library. The video frame data is in YUV format, and the audio frame data is in PCM format.

[0039] Here, using the FFmpeg library to decode the first multimedia file is a key step to obtain video frame data and audio frame data. FFmpeg supports a wide range of multimedia formats, including common video and audio container formats such as MP4, MKV, AVI, etc. The video part is decoded into frame data in YUV format. YUV is a color encoding format, where Y represents luminance, and U and V represent chrominance information. The YUV format is very common in video processing because it allows independent processing of luminance and chrominance, which is very useful in compression and special effects processing. The audio part is decoded into frame data in PCM (Pulse Code Modulation) format. PCM is an uncompressed digital audio representation method, where the audio signal is sampled and quantized into a series of values. The PCM format provides high-quality audio data, suitable for further analysis and processing.

[0040] In step S102, audio waveform data is constructed from the audio frame data, and a set of statistical parameters of the audio waveform data and the waveform average value in a specific time interval are determined; wherein, the set of statistical parameters at least includes the mathematical mean, mathematical variance, and mathematical mean square deviation of the audio waveform data.

[0041] Here, constructing audio waveform data from audio frame data is the basis for understanding audio characteristics; determining a set of statistical parameters of the audio waveform data, including the mathematical mean, mathematical variance, and mathematical mean square deviation, etc., these parameters help to quantify the characteristics of the audio; calculating the waveform average value in a specific time interval, which can reflect the average intensity or level of the audio in this time period.

[0042] In some embodiments, refer toFigure 2 , Figure 2 is a schematic flowchart of steps S201 - S202 provided by an embodiment of the present application. The construction of the audio waveform data from the audio frame data can be achieved through steps S201 - S202, and will be described in conjunction with each step.

[0043] In step S201, the number of channels, sampling rate, bit depth, and number of frames are extracted from the audio frame data.

[0044] In step S202, the number of channels, the sampling rate, the bit depth, and the number of frames are substituted into a waveform generation function to obtain the audio waveform data.

[0045] Here, the audio channel number represents the number of independent sound channels contained in the audio signal. Common channel numbers include mono (1 channel), stereo (2 channels), and multi - channel (such as 5.1, 7.1, etc.).

[0046] The sampling rate refers to the number of times the audio signal is sampled per second, usually measured in Hertz (Hz). The higher the sampling rate, the higher the restoration degree of the audio signal and the better the sound quality. The bit depth represents the number of data bits for each sampling point, which determines the dynamic range and accuracy of the audio signal. Common bit depths include 8 - bit, 16 - bit, 24 - bit, and 32 - bit, etc. The number of frames represents the total number of sampling points in the audio data, which is related to the audio duration and sampling rate.

[0047] After extracting the above parameters of the audio frame data, these parameters will be substituted into the waveform generation function to generate the audio waveform data.

[0048] In some embodiments, the mathematical mean mathematical variance and mathematical mean square deviation where Wi represents any sample value in the audio waveform data, and n represents the number of samples in the audio waveform data;

[0049] The value range of a specific time interval is greater than or equal to 0.01 seconds and less than 0.1 seconds. The waveform average value is calculated with the specific time interval as the step size, and the waveform average value Agv(t) = (W(t + Δt) - W(t)) / Δt; where W(t) is the audio waveform data at time t, Δt is the step size, and [t, t + Δt] represents the specific time interval.

[0050] Here, the mathematical mean is the sum of all sample values divided by the number of samples. In the context of audio waveform data, the mathematical mean can reflect the average volume or intensity of the audio signal. The mathematical variance is a statistic that measures the degree of dispersion of the audio waveform data, that is, the average of the squares of the differences between the sample values and the mathematical mean. The larger the variance, the higher the degree of dispersion of the audio waveform data, that is, the greater the variation in volume or intensity. The mathematical mean square deviation is the square root of the variance, which provides another measure of the degree of dispersion of the audio waveform data. The mean square deviation can be understood physically as the standard deviation of the differences between the sample values and the mathematical mean.

[0051] The average value of the waveform in a specific time interval is used to measure the average rate of change of the audio waveform data within a given time interval. It is particularly useful for capturing rapid changes in the audio signal. The calculation formula is: Agv(t) = (W(t + Δt) - W(t)) / Δt; where W(t) represents the audio waveform data at time t, Δt represents the step size of the specific time interval, and [t, t + Δt] represents the time interval. In the embodiments of the present application, the value range of the specific time interval is limited between 0.01 seconds and 0.1 seconds. This range is selected based on considerations of the actual application scenario to ensure that both rapid changes in the audio signal can be captured and the calculation results are not overly sensitive or unstable due to too short a time interval.

[0052] In step S103, control parameters are determined based on the audio frame data and the set of statistical parameters; wherein, the control parameters are used to control the special effect output of the video frame data.

[0053] Here, based on the audio frame data and the set of statistical parameters, control parameters for controlling the special effect output of the video frame data are determined. These parameters relate to the intensity, speed, direction, etc. of the video special effects.

[0054] In some embodiments, refer to Figure 3 , Figure 3 is a schematic flowchart of steps S301 - S302 provided by the embodiments of the present application. The determination of the control parameters based on the audio frame data and the set of statistical parameters can be implemented through steps S301 - S302, and will be described in combination with each step.

[0055] In step S301, the weight of the mathematical mean square deviation is determined based on the set of statistical parameters, and the number of audio tracks, the center point of the video frame, and the RGB color are determined based on the audio frame data; wherein, the weight characterizes the smoothness of the audio waveform data, and the value range of the weight is [0, 1].

[0056] In step S302, the control parameters are determined based on the number of audio tracks, the weight of the mathematical mean square deviation, the center point of the video frame, and the RGB color.

[0057] Here, a parameter U(n, k, o, c) is defined to control the special effect output.

[0058] n: The number of audio tracks;

[0059] k: This is the weight corresponding to the mathematical mean square error, used to reflect the smoothness of the waveform, and the range is [0, 1];

[0060] o: The center point of the video frame;

[0061] c: RGB color;

[0062] After determining the parameters such as the weight of the mathematical mean square error, the number of audio tracks, the center point of the video frame, and the RGB color, we can determine the control parameters based on these parameters. The control parameters are the key information guiding the generation and output of video special effects, and it may include the type, intensity, action area, duration, etc. of the special effects. The specific way of determining the control parameters varies due to different application scenarios and algorithm designs. For example, in a possible implementation, the intensity of the video special effect can be adjusted according to the value of the weight. The greater the weight, the higher the special effect intensity; at the same time, the action area and direction of the special effect are determined according to the number of audio tracks and the center point of the video frame; finally, the color attribute of the special effect is determined according to the RGB color.

[0063] In step S104, based on the statistical parameter set, the control parameters, and the waveform average value, target frame data for video frame synthesis is constructed, and the target frame data is synthesized with the video frame data to obtain target video data; wherein, each frame in the target frame data corresponds one-to-one with each frame in the video frame data.

[0064] Here, according to the statistical parameter set, the control parameters, and the waveform average value, target frame data for video frame synthesis is constructed. Each frame in the target frame data corresponds one-to-one with each frame in the original video frame data. The target frame data is synthesized with the original video frame data to obtain target video data. This process involves special effect processing such as image overlay, color adjustment, and filter application.

[0065] In some embodiments, refer to Figure 4 , Figure 4 is the flowchart of steps S401 - S402 provided by the embodiments of the present application. The construction of the target frame data for video frame synthesis based on the statistical parameter set, the control parameters, and the waveform average value can be realized through steps S401 - S402, and will be described in combination with each step.

[0066] In step S401, the i-th frame image Fr(i) is obtained from the video frame data.

[0067] In step S402, a two-dimensional plane Fa(t) of the same size as Fr(i) is constructed, and the initial value of Fa(t) is set to 1 to obtain the target frame data; where Q represents the audio waveform data, U represents the control parameter, and Agv(t) represents the waveform average value; Fa(t) represents a fluctuation centered on the center point of the video frame, which fluctuates around with time t, and the affected coordinate positions are changed within [0, 1], while the unaffected coordinate positions remain 1.

[0068] Here, the system obtains the i-th frame image from the video frame data, denoted as Fr(i). This step is a common operation in the video processing flow, aiming to select a specific video frame as the starting point for subsequent processing.

[0069] The system performs the following operations to construct the target frame data:

[0070] Initializing the two-dimensional plane: First, the system creates a two-dimensional plane of the same size as Fr(i), denoted as Fa(t). The initial value of each pixel point in this plane is set to 1. This two-dimensional plane will serve as the basis for the target frame data.

[0071] Calculating the target frame data: Next, the system calculates the target frame data according to the formula Here, the formula involves several key elements:

[0072] Q: Represents the audio waveform data. In practical applications, it can refer to the sample value at a specific moment in the audio frame or the processed audio feature value.

[0073] U: Represents the control parameter. This parameter is determined based on the statistical parameter set and possible other factors (such as the number of audio tracks, the center point of the video frame, etc.), and is used to adjust the intensity or behavior of the video special effects.

[0074] Agv(t): Represents the waveform average value, which reflects the average change rate of the audio waveform data within a specific time interval.

[0075] Fluctuation effect: Fa(t) in the formula represents a fluctuation effect centered on the center point of the video frame. As time t progresses, this fluctuation effect spreads around. The values of the affected coordinate positions change within [0, 1], while the unaffected coordinate positions remain 1. This fluctuation effect is related to the intensity or change rate of the audio signal, thus realizing the interaction between audio and video.

[0076] Finally, the obtained Fa(t) after the above calculations is the target frame data. This data can be used in the subsequent video frame synthesis process to achieve audio-driven video effects.

[0077] Application scenarios and significance.

[0078] In some embodiments, the synthesizing the target frame data with the video frame data to obtain target video data includes:

[0079] Synthesizing the target frame data with the video frame data through the following formula:

[0080] F(i) = Fr(i) × Fa(t), where the value of F(i) is constrained to [0, 255];

[0081]

[0082] Save the video frame sequence V as the target video data.

[0083] Here, for each frame Fr(i) in the video (where i represents the frame index), the corresponding target frame data Fa(t) is used for synthesis.

[0084] The synthesis formula is: F(i) = Fr(i) × Fa(t).

[0085] Here, F(i) represents the i-th synthesized image, Fr(i) represents the original i-th image, and Fa(t) represents the target frame data corresponding to the i-th frame (where t is related to i, representing a relative time or frame offset, depending on the calculation method of Fa(t)).

[0086] Value constraint: The value of the synthesized F(i) is constrained between [0, 255]. This is because in most video coding standards, the pixel value range is usually 0 to 255 (for 8-bit depth). If the calculated value of F(i) exceeds this range, appropriate scaling or clipping is required to ensure its validity.

[0087] Sequence synthesis: After frame-by-frame synthesis of all frames, the obtained frame sequence F1, F2,..., FN (where N is the total number of frames in the video) is combined to form the final video sequence V.

[0088] The synthesis formula can be expressed as:

[0089] Save the target video data: Finally, save the synthesized video sequence V as the target video data. This usually involves encoding the video sequence into a specific video format (such as MP4, AVI, etc.) and saving it to disk.

[0090] In step S105, the target frame data and the audio frame data are merged to obtain a second multimedia file; wherein, when the second multimedia file is played, special effects are displayed based on the audio frame data.

[0091] Finally, the processed target frame data and the audio frame data are merged to generate a second multimedia file. When playing this multimedia file, the audio frame data will trigger the display of corresponding video special effects.

[0092] In summary, the embodiments of the present application have the following beneficial effects:

[0093] (1) Enhancing the interactivity and expressiveness of multimedia files: Through the embodiments of the present application, the audio frame data is effectively utilized to construct audio waveform data, and control parameters are determined based on these data, thereby controlling the special effect output of the video frame data. This interactive manner between audio and video greatly enhances the dynamic effect and expressiveness of multimedia files, enabling users to obtain a richer and more immersive experience when watching videos.

[0094] (2) Improving the flexibility and accuracy of multimedia file synthesis: The embodiments of the present application calculate a set of statistical parameters of the audio waveform data (including mathematical mean, mathematical variance, and mathematical mean square deviation, etc.), as well as the waveform average value in a specific time interval, providing a reliable data basis for the precise control of video frame special effects. At the same time, the target frame data constructed based on these parameters and control parameters can ensure that the video special effects are closely synchronized with the changes in the audio signal, improving the flexibility and accuracy of synthesis.

[0095] (3) Optimizing the generation efficiency and quality of video special effects: When constructing the target frame data, the efficient algorithm of the embodiments of the present application calculates the weighted sum of the audio waveform data, control parameters, and waveform average value to generate special effect frames corresponding one by one to the video frame data. This method not only simplifies the generation process of video special effects but also improves the quality of special effect frames, making the final synthesized video more smooth and natural in visual effects.

[0096] (4) Supporting multiple multimedia file formats: The embodiments of the present application can process multimedia files in multiple formats, especially through the FFmpeg codec library for video and audio decoding processing, making this method widely applicable.

[0097] (5) Promoting multimedia creation and personalized customization: The multimedia file synthesis method of the embodiments of the present application provides more creative space and possibilities for personalization for creators. By adjusting control parameters, the processing method of audio waveform data, and the type of video special effects, etc., creators can easily achieve various creative effects to meet the diverse and personalized needs of users for multimedia content.

[0098] Based on the same inventive concept, an embodiment of the present application further provides a multimedia file synthesis device corresponding to the multimedia file synthesis method in the first embodiment. Since the principle of problem-solving of the device in the embodiment of the present application is similar to that of the above multimedia file synthesis method, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0099] As Figure 5 shown, Figure 5 FIG. is a schematic structural diagram of a multimedia file synthesis device 500 provided by an embodiment of the present application. The multimedia file synthesis device 500 includes:

[0100] A decoding module 501, configured to perform decoding processing on the first multimedia file to obtain video frame data and audio frame data;

[0101] A construction module 502, configured to construct audio waveform data through the audio frame data, and determine a statistical parameter set of the audio waveform data and a waveform average value of a specific time interval based on the audio waveform data; wherein, the statistical parameter set at least includes a mathematical mean, a mathematical variance, and a mathematical mean square deviation of the audio waveform data;

[0102] A determination module 503, configured to determine a control parameter based on the audio frame data and the statistical parameter set; wherein, the control parameter is used to control the special effect output of the video frame data;

[0103] A synthesis module 504, configured to construct target frame data for video frame synthesis based on the statistical parameter set, the control parameter, and the waveform average value, and perform synthesis processing on the target frame data and the video frame data to obtain target video data; wherein, each frame in the target frame data corresponds to each frame in the video frame data one by one;

[0104] A merging module 505, configured to merge the target frame data and the audio frame data to obtain a second multimedia file; wherein, the second multimedia file displays special effects based on the audio frame data during playback.

[0105] Those skilled in the art should understand that Figure 5 the implementation functions of the various units in the multimedia file synthesis device 500 shown can be understood with reference to the relevant descriptions of the foregoing multimedia file synthesis method. Figure 5 The functions of the various units in the multimedia file synthesis device 500 shown can be implemented by a program running on a processor, or can be implemented by specific logic circuits.

[0106] In a possible implementation, the first multimedia file is subjected to video decoding and audio decoding respectively based on the FFmpeg codec library. The video frame data is in YUV format, and the audio frame data is in PCM format.

[0107] In a possible implementation, the construction module 502 constructs audio waveform data from the audio frame data, including:

[0108] Extracting the number of channels, sampling rate, bit depth, and number of frames from the audio frame data;

[0109] Substituting the number of channels, the sampling rate, the bit depth, and the number of frames into a waveform generation function to obtain the audio waveform data.

[0110] In a possible implementation, the mathematical mean mathematical variance and mathematical mean square deviation where Wi represents any sample value in the audio waveform data, and n represents the number of samples in the audio waveform data;

[0111] The value range of the specific time interval is greater than or equal to 0.01 second and less than 0.1 second. The waveform average value is calculated with the specific time interval as the step size. The waveform average value Agv(t) = (W(t + Δt) - W(t)) / Δt; where W(t) is the audio waveform data at time t, Δt is the step size, and [t, t + Δt] represents the specific time interval.

[0112] In a possible implementation, the determination module 503 determines control parameters based on the audio frame data and the statistical parameter set, including:

[0113] Determining the weight of the mathematical mean square deviation based on the statistical parameter set, and determining the number of audio tracks, the center point of the video frame, and the RGB color based on the audio frame data; where the weight characterizes the smoothness of the audio waveform data, and the value range of the weight is [0, 1];

[0114] Determining the control parameters based on the number of audio tracks, the weight of the mathematical mean square deviation, the center point of the video frame, and the RGB color.

[0115] In a possible implementation, the synthesis module 504 constructs target frame data for video frame synthesis based on the statistical parameter set, the control parameters, and the waveform average value, including:

[0116] Obtaining the i-th frame image Fr(i) from the video frame data;

[0117] Construct a two-dimensional plane Fa(t) with the same size as Fr(i), and set the initial value of Fa(t) to 1 to obtain the target frame data; wherein, Q represents the audio waveform data, U represents the control parameter, and Agv(t) represents the waveform average value; Fa(t) represents that with the center point of the video frame as the center, it fluctuates around continuously over time t, and the affected coordinate positions are changed within [0, 1], while the unaffected coordinate positions remain 1.

[0118] In a possible implementation manner, the synthesis module 504 synthesizes the target frame data and the video frame data to obtain target video data, including:

[0119] Synthesize the target frame data and the video frame data through the following formula:

[0120] F(i) = Fr(i) × Fa(t), and the value of F(i) is constrained to [0, 255];

[0121]

[0122] Save the video frame sequence V as the target video data.

[0123] The above multimedia file synthesis device has the following beneficial effects:

[0124] (1) Enhance the interactivity and expressiveness of multimedia files: Through the embodiments of the present application, the audio frame data is effectively utilized to construct the audio waveform data, and based on these data, the control parameters are determined to control the special effect output of the video frame data. This interactive manner between audio and video greatly enhances the dynamic effect and expressiveness of multimedia files, enabling users to obtain a more rich and immersive experience when watching videos.

[0125] (2) Improve the flexibility and accuracy of multimedia file synthesis: The embodiments of the present application calculate the statistical parameter set of the audio waveform data (including mathematical mean, mathematical variance, mathematical mean square deviation, etc.), as well as the waveform average value in a specific time interval, providing a reliable data basis for the precise control of video frame special effects. At the same time, the target frame data constructed based on these parameters and control parameters can ensure that the video special effects are closely synchronized with the changes in the audio signal, improving the flexibility and accuracy of synthesis.

[0126] (3) Optimize the generation efficiency and quality of video special effects: When constructing the target frame data, the efficient algorithm of the embodiments of the present application generates special effect frames corresponding one by one to the video frame data by calculating the weighted sum of the audio waveform data, control parameters, and waveform average values. This method not only simplifies the process of generating video special effects but also improves the quality of the special effect frames, making the final synthesized video more smooth and natural in visual effects.

[0127] (4) Support multiple multimedia file formats: The embodiments of the present application can process multimedia files in multiple formats. In particular, through the FFmpeg codec library for video and audio decoding processing, this method has wide applicability.

[0128] (5) Facilitate multimedia creation and personalized customization: The multimedia file synthesis method of the embodiments of the present application provides more creative space and possibilities for personalization for creators. By adjusting control parameters, the processing method of audio waveform data, and the type of video special effects, etc., creators can easily achieve various creative effects and meet the diverse and personalized needs of users for multimedia content.

[0129] As Figure 6 shown, Figure 6 is a schematic structural diagram of an electronic device 600 provided by the embodiments of the present application. The electronic device 600 includes:

[0130] A processor 601, a storage medium 602, and a bus 603. The storage medium 602 stores machine-readable instructions executable by the processor 601. When the electronic device 600 runs, the processor 601 communicates with the storage medium 602 through the bus 603, and the processor 601 executes the machine-readable instructions to perform the steps of the multimedia file synthesis method described in the embodiments of the present application.

[0131] In practical applications, the various components in the electronic device 600 are coupled together through the bus 603. It can be understood that the bus 603 is used to realize the connection and communication between these components. In addition to the data bus, the bus 603 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 6 all kinds of buses are labeled as bus 603.

[0132] The above-mentioned electronic device has the following beneficial effects:

[0133] (1)Enhance the interactivity and expressiveness of multimedia files: Through the embodiments of the present application, audio frame data is effectively utilized to construct audio waveform data, and control parameters are determined based on this data to control the special effect output of video frame data. This interactive manner between audio and video greatly enhances the dynamic effect and expressiveness of multimedia files, enabling users to obtain a richer and more immersive experience when watching videos.

[0134] (2)Improve the flexibility and accuracy of multimedia file synthesis: The embodiments of the present application calculate a set of statistical parameters of audio waveform data (including mathematical mean, mathematical variance, and mathematical mean square deviation, etc.), as well as the waveform average value in a specific time interval, providing a reliable data basis for the precise control of video frame special effects. At the same time, the target frame data constructed based on these parameters and control parameters can ensure that the video special effects are closely synchronized with the changes in the audio signal, improving the flexibility and accuracy of synthesis.

[0135] (3)Optimize the generation efficiency and quality of video special effects: When constructing the target frame data, the efficient algorithm of the embodiments of the present application generates special effect frames corresponding one by one to the video frame data by calculating the weighted sum of the audio waveform data, control parameters, and waveform average value. This method not only simplifies the process of generating video special effects but also improves the quality of the special effect frames, making the final synthesized video more smooth and natural in visual effects.

[0136] (4)Support multiple multimedia file formats: The embodiments of the present application can process multimedia files in multiple formats, especially by decoding video and audio through the FFmpeg codec library, making this method widely applicable.

[0137] (5)Promote multimedia creation and personalized customization: The multimedia file synthesis method of the embodiments of the present application provides more creative space and possibilities for personalization for creators. By adjusting control parameters, the processing method of audio waveform data, and the type of video special effects, etc., creators can easily achieve various creative effects to meet the diverse and personalized needs of users for multimedia content.

[0138] The embodiments of the present application also provide a computer-readable storage medium, and the storage medium stores executable instructions. When the executable instructions are executed by at least one processor 601, the multimedia file synthesis method described in the embodiments of the present application is implemented.

[0139] In some embodiments, the storage medium may be a ferromagnetic random access memory (FRAM), read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), flash memory, magnetic surface memory, optical disc, or compact disc read only memory (CDROM), etc.; or it may be various devices including one or any combination of the above memories.

[0140] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0141] As an example, the executable instructions may or may not correspond to a file in the file system, and may be stored as part of a file that holds other programs or data. For example, they may be stored in one or more scripts in a hypertext markup language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code).

[0142] As an example, the executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or alternatively, on multiple computing devices distributed across multiple locations and interconnected by a communication network.

[0143] The above computer-readable storage medium has the following beneficial effects:

[0144] (1)Enhance the interactivity and expressiveness of multimedia files: Through the embodiments of the present application, audio frame data is effectively utilized to construct audio waveform data, and control parameters are determined based on these data to control the special effect output of video frame data. This interactive manner between audio and video greatly enhances the dynamic effect and expressiveness of multimedia files, enabling users to obtain a richer and more immersive experience when watching videos.

[0145] (2)Improve the flexibility and accuracy of multimedia file synthesis: The embodiments of the present application provide a reliable data basis for the precise control of video frame special effects by calculating the statistical parameter set of audio waveform data (including mathematical mean, mathematical variance, and mathematical mean square deviation, etc.) and the waveform average value in a specific time interval. At the same time, the target frame data constructed based on these parameters and control parameters can ensure that the video special effects are closely synchronized with the changes in the audio signal, improving the flexibility and accuracy of synthesis.

[0146] (3)Optimize the generation efficiency and quality of video special effects: When constructing the target frame data, the efficient algorithm of the embodiments of the present application generates special effect frames corresponding one by one to the video frame data by calculating the weighted sum of audio waveform data, control parameters, and waveform average value. This method not only simplifies the generation process of video special effects but also improves the quality of special effect frames, making the final synthesized video more smooth and natural in visual effects.

[0147] (4)Support multiple multimedia file formats: The embodiments of the present application can process multimedia files in multiple formats, especially by decoding and encoding videos and audios through the FFmpeg codec library, making this method widely applicable.

[0148] (5)Promote multimedia creation and personalized customization: The multimedia file synthesis method of the embodiments of the present application provides more creative space and possibilities for personalization for creators. By adjusting control parameters, the processing method of audio waveform data, and the type of video special effects, etc., creators can easily achieve various creative effects to meet the diverse and personalized needs of users for multimedia content.

[0149] In several embodiments provided by the present application, it should be understood that the disclosed methods and electronic devices can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0150] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0151] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit.

[0152] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a platform server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0153] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for synthesizing multimedia files, characterized in that, The method includes: Performing decoding processing on a first multimedia file to obtain video frame data and audio frame data; Constructing audio waveform data from the audio frame data, and determining a statistical parameter set of the audio waveform data and a waveform average value in a specific time interval based on the audio waveform data; wherein, the statistical parameter set at least includes the mathematical mean, mathematical variance, and mathematical mean square deviation of the audio waveform data; Determining a control parameter based on the audio frame data and the statistical parameter set; wherein, the control parameter is used to control the special effect output of the video frame data; Constructing target frame data for video frame synthesis based on the statistical parameter set, the control parameter, and the waveform average value, and performing synthesis processing on the target frame data and the video frame data to obtain target video data; wherein, each frame in the target frame data corresponds one-to-one with each frame in the video frame data; Merging the target frame data and the audio frame data to obtain a second multimedia file; wherein, the second multimedia file displays special effects based on the audio frame data during playback.

2. The method according to claim 1, wherein The first multimedia file performs video decoding and audio decoding respectively based on the FFmpeg codec library, the video frame data is in YUV format, and the audio frame data is in PCM format.

3. The method according to claim 1, characterized in that The constructing audio waveform data from the audio frame data includes: Extracting the number of channels, sampling rate, bit depth, and number of frames from the audio frame data; Substituting the number of channels, the sampling rate, the bit depth, and the number of frames into a waveform generation function to obtain the audio waveform data.

4. The method according to claim 1, wherein The mathematical mean Mathematical variance And mathematical mean square deviation Wherein, Wi represents any sample value in the audio waveform data, and n represents the number of samples in the audio waveform data; The value range of the specific time interval is greater than or equal to 0.01 second and less than 0.1 second, the waveform average value is calculated with the specific time interval as the step size, and the waveform average value Agv(t) = (W(t + Δt) - W(t)) / Δt; where W(t) is the audio waveform data at time t, Δt is the step size, and [t, t + Δt] represents the specific time interval.

5. The method according to claim 1, characterized in that The determining a control parameter based on the audio frame data and the statistical parameter set includes: Determining the weight of the mathematical mean square deviation based on the statistical parameter set, and determining the number of audio tracks, the center point of the video frame, and the RGB color based on the audio frame data; wherein, the weight characterizes the smoothness of the audio waveform data, and the value range of the weight is [0, 1]; Determining the control parameter based on the number of audio tracks, the weight of the mathematical mean square deviation, the center point of the video frame, and the RGB color.

6. The method according to claim 5, wherein The constructing target frame data for video frame synthesis based on the statistical parameter set, the control parameter, and the waveform average value includes: Obtaining the i-th frame image Fr(i) from the video frame data; Construct a two-dimensional plane Fa(t) with the same size as Fr(i), and set the initial value of Fa(t) to 1 to obtain the target frame data; where Q represents the audio waveform data, U represents the control parameter, and Agv(t) represents the waveform average value; Fa(t) represents that with the center point of the video frame as the center, it fluctuates continuously around with time t, and the affected coordinate positions are changed between [0, 1], and the unaffected coordinate positions remain 1.

7. The method according to claim 6, wherein The performing synthesis processing on the target frame data and the video frame data to obtain target video data includes: Performing synthesis processing on the target frame data and the video frame data through the following formula: F(i) = Fr(i) × Fa(t), and the value of F(i) is constrained to [0, 255]; Saving the video frame sequence V as target video data.

8. A multimedia file synthesis device, characterized in that The device includes: A decoding module for decoding a first multimedia file to obtain video frame data and audio frame data; A construction module for constructing audio waveform data from the audio frame data, and determining a statistical parameter set of the audio waveform data and a waveform average value in a specific time interval based on the audio waveform data; wherein the statistical parameter set at least includes the mathematical mean, mathematical variance, and mathematical mean square deviation of the audio waveform data; A determination module for determining a control parameter based on the audio frame data and the statistical parameter set; wherein the control parameter is used to control the special effect output of the video frame data; A synthesis module for constructing target frame data for video frame synthesis based on the statistical parameter set, the control parameter, and the waveform average value, and performing a synthesis process on the target frame data and the video frame data to obtain target video data; wherein each frame in the target frame data corresponds one-to-one with each frame in the video frame data; A merging module for merging the target frame data and the audio frame data to obtain a second multimedia file; wherein the second multimedia file displays special effects based on the audio frame data during playback.

9. An electronic device, characterized in that, Comprising: A processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the multimedia file synthesis method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by the processor, it executes the multimedia file synthesis method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Satellite channel high fault tolerance audio and video joint coding and decoding method

    CN120935358A

  • Satellite channel high fault-tolerant audio and video joint coding method

    CN120935358B