WebM video synthesis method and device, equipment and medium
By extracting and segmenting audio files and using a deep learning framework to process audio feature segments in parallel, the problem of time-consuming WebM video synthesis is solved, and the audio and video synthesis efficiency is achieved quickly responding to customer interaction needs.
Patent Information
- Application Number
- CN202510193060.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
AI Technical Summary
The existing WebM video synthesis method has a time-consuming problem, which makes it impossible to respond to customer interaction needs in a timely manner during the digital human synthesis process.
By loading the pre-stored audio files, extracting audio feature data, dividing them into multiple audio feature fragments, and processing these fragments in parallel based on the deep learning framework to generate corresponding video clips, and finally merge them into audio and video.
This method effectively reduces the time for audio and video synthesis, improves synthesis efficiency, and can quickly respond to customers' interactive needs.
Smart Images

Figure CN120050448A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a WebM video synthesis method, device, equipment and medium. Background Art
[0002] Digital people are digital characters that are close to human images and created using digital technology. As artificial intelligence becomes more and more popular, more and more digital people generated by artificial intelligence are taking on the role of customer service. For example, in the banking industry, digital people who can communicate with customers also take on some of the work of lobby managers in financial equipment.
[0003] In the prior art, data browsing is usually achieved through the Web, and WebM is an encoding method used by mainstream browsers to display digital human videos. However, video WebM encoding uses CPU encoding, and there is a problem of long transcoding time. During the digital human synthesis process, if the transcoding time is too long, the digital human cannot be synthesized in time, which in turn leads to the problem of being unable to meet the customer's interactive needs in time.
[0004] Therefore, the existing WebM video synthesis method has the problem of being time-consuming. Summary of the invention
[0005] The embodiments of the present invention provide a WebM video synthesis method, device, equipment and medium, aiming to solve the problem that the existing WebM video synthesis method is time-consuming.
[0006] In a first aspect, an embodiment of the present invention provides a WebM video synthesis method, the method comprising:
[0007] Load the pre-stored audio file to obtain the corresponding audio data;
[0008] Extracting features from the audio data using a preset audio feature extraction model to obtain corresponding audio feature data;
[0009] Segmenting the audio feature data according to a preset segmentation coefficient to obtain a plurality of audio feature segments;
[0010] Processing the plurality of audio feature segments in parallel based on a deep learning framework to obtain a video segment corresponding to each audio feature segment;
[0011] The audio data and the plurality of video clips are combined to obtain corresponding audio and video. In a second aspect, an embodiment of the present invention further provides a WebM video synthesis device, the device comprising:
[0012] A loading unit, used to load a pre-stored audio file to obtain corresponding audio data;
[0013] An extraction unit, configured to extract features from the audio data using a preset audio feature extraction model to obtain corresponding audio feature data;
[0014] A segmentation unit, used to segment the audio feature data according to a preset segmentation coefficient to obtain a plurality of audio feature segments;
[0015] A processing unit, configured to perform parallel processing on the plurality of audio feature segments based on a deep learning framework to obtain a video segment corresponding to each of the audio feature segments;
[0016] The merging unit is used to merge the audio data and the multiple video clips to obtain corresponding audio and video.
[0017] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0018] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program includes program instructions, and wherein the program instructions, when executed by a processor, can implement the method described in the first aspect above.
[0019] The present invention provides a WebM video synthesis method, device, equipment and medium, the method comprising: loading a pre-stored audio file to obtain corresponding audio data; extracting features from the audio data using a preset audio feature extraction model to obtain corresponding audio feature data; segmenting the audio feature data according to a preset segmentation coefficient to obtain multiple audio feature segments; processing multiple audio feature segments in parallel based on a deep learning framework to obtain a video segment corresponding to each audio feature segment; merging the audio data and multiple video segments to obtain corresponding audio and video. The embodiment of the present invention can process multiple audio feature segments in parallel based on a deep learning framework at the same time, which can effectively reduce the time for synthesizing audio and video, improve the efficiency of audio and video synthesis, and achieve rapid response to customer interaction needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0021] Figure 1A schematic diagram of a flow chart of a WebM video synthesis method provided by an embodiment of the present invention;
[0022] Figure 2 A schematic block diagram of a WebM video synthesis device provided by an embodiment of the present invention;
[0023] Figure 3 A schematic block diagram of an electronic device provided by an embodiment of the present invention;
[0024] Figure 4 A schematic diagram of an application environment of a WebM video synthesis method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0027] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0028] It should be further understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. The embodiment of the present invention provides a WebM video synthesis method, apparatus, device and medium. The WebM video synthesis method can be found in Figure 4 , Figure 4 Schematic diagram of the application environment of the WebM video synthesis method provided by the embodiment of the present invention. The WebM video synthesis method is applied in Figure 4In an application environment, the client communicates with the server through a network. The server can receive customer questions through the client, and perform question recall processing from a preset standard question library according to the customer questions to obtain the corresponding audio file; load the audio file to obtain the corresponding audio data; extract features from the audio data by a preset audio feature extraction model to obtain corresponding audio feature data; segment the audio feature data according to a preset segmentation coefficient to obtain multiple audio feature segments; perform parallel processing on multiple audio feature segments based on a deep learning framework to obtain a video segment corresponding to each audio feature segment; merge the audio data and multiple video segments to obtain corresponding audio and video; the server sends the audio and video to the client to play the audio and video on the target display interface of the client to respond to the customer's questions. The embodiment of the present invention can simultaneously process multiple audio feature segments in parallel based on a deep learning framework, which can effectively reduce the time for synthesizing audio and video, improve the efficiency of audio and video synthesis, and achieve rapid response to customer interaction needs. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0029] Figure 1 Schematic diagram of the flow of the WebM video synthesis method provided by the embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110-S150.
[0030] S110: Load a pre-stored audio file to obtain corresponding audio data.
[0031] In this embodiment, a pre-stored audio file is loaded to obtain corresponding audio data, and the audio file is a Wav audio file, which can record various mono or stereo sound information and ensure that the sound is not distorted; the audio file is used to answer questions raised by customers to achieve human-computer interaction; specifically, the server performs question recall processing from a preset standard question library based on customer questions from the client to obtain corresponding audio files.
[0032] In one embodiment, before step S110, the method further includes: determining whether the sampling rate of the audio file is a target sampling rate to obtain a determination result; if the determination result is no, converting the sampling rate of the audio file according to the target sampling rate to obtain a converted audio file.
[0033] In this embodiment, it is determined whether the sampling rate of the audio file is the target sampling rate to obtain a determination result; the target sampling rate can be set to 16K, and the present invention can improve data processing efficiency by setting a unified target sampling rate; if the determination result is no, the sampling rate of the audio file is converted according to the target sampling rate to obtain a converted audio file; if the determination result is yes, step S120 is entered without converting the sampling rate of the audio file.
[0034] S120: extract features from the audio data using a preset audio feature extraction model to obtain corresponding audio feature data.
[0035] In this embodiment, a preset audio feature extraction model is used to perform feature extraction on the audio data to obtain corresponding audio feature data; wherein, the audio feature data is a Hubert feature; the audio feature model may be a Hubert (speech recognition pre-training) model, which extracts latent variables from the audio data by self-encoding, and these latent variables are quantified and used for subsequent speech processing tasks.
[0036] S130: Segment the audio feature data according to a preset segmentation coefficient to obtain a plurality of audio feature segments.
[0037] In this embodiment, the audio feature data is segmented according to a preset segmentation coefficient to obtain a plurality of audio feature segments; if the preset segmentation coefficient is 3, the audio feature data is evenly segmented into 3 equal parts to obtain 3 audio feature segments. In this embodiment of the present invention, the audio feature data is evenly segmented into a plurality of segments so that a plurality of audio feature segments can be processed in parallel in a subsequent processing process.
[0038] S140. Process the multiple audio feature segments in parallel based on a deep learning framework to obtain a video segment corresponding to each of the audio feature segments.
[0039] In this embodiment, multiple audio feature segments are processed in parallel based on a deep learning framework (pytorch) to obtain video segments corresponding to each audio feature segment; wherein the audio feature segments each include multiple audio features, and the video segments each include multiple image frames, and the format of the image frames is VP8 or VP9 format for WebM files, and the audio features and the image frames are in a one-to-one correspondence.
[0040] In one embodiment, step S140 includes: inferring the audio features in the audio feature segment based on a deep learning framework to obtain an image frame corresponding to each audio feature in the audio feature segment; wherein the audio feature segment includes multiple audio features; and merging multiple image frames to obtain a video segment corresponding to the audio feature segment.
[0041] In this embodiment, multiple audio feature segments are processed in parallel based on a deep learning framework (pytorch). Here, taking the processing of a single audio feature segment as an example, the audio features in the audio feature segment are inferred based on the deep learning framework to obtain an image frame corresponding to each audio feature in the audio feature segment; wherein the audio feature segment includes multiple audio features; the multiple image frames are merged to obtain a video segment corresponding to the audio feature segment; specifically, each image frame is associated with a unique first timestamp, and the multiple image frames can be merged according to the monotonically increasing characteristic of the first timestamp.
[0042] In one embodiment, before the deep learning framework is used to infer the audio features in the audio feature segment to obtain an image frame corresponding to each audio feature in the audio feature segment, it also includes: acquiring template image data; converting the size of the template image data according to a preset pixel value to obtain a target template image.
[0043] In this embodiment, template image data is obtained; the size of the template image data is converted according to a preset pixel value to obtain a target template image; wherein the preset pixel value can be set to 256×256, and the present invention improves data processing efficiency by setting a uniform pixel value; the target template image is a digital human image, and the mouth shape of the digital human in the target template image is a default mouth shape, such as not open.
[0044] In one embodiment, after converting the size of the template image data according to the preset pixel value to obtain the target template image, it also includes: inputting the target template image and the audio feature into the deep learning framework to obtain the initial image frame corresponding to the audio feature.
[0045] In this embodiment, the target template image and the audio feature are input into the deep learning framework to obtain an initial image frame corresponding to the audio feature; wherein, the format of the initial image frame is BGR format; specifically, the target template image and the audio feature T1 are input into the deep learning framework, and the digital population type T2 corresponding to the audio feature T1 is determined based on the deep learning framework, and the initial image frame corresponding to the audio feature T1 is obtained according to the digital population type T2 and the target template image.
[0046] In one embodiment, after obtaining the initial image frame corresponding to the audio feature, the method further includes: performing encoding conversion on the initial image frame according to a preset encoding format to obtain the image frame corresponding to the audio feature.
[0047] In this embodiment, the initial image frame is encoded and converted according to a preset encoding format to obtain an image frame corresponding to the audio feature; wherein the preset encoding format can be VP8 format or VP9 format; VP8 and VP9 are video encoding formats and are widely used in WebM video files.
[0048] S150: Merge the audio data and the multiple video clips to obtain corresponding audio and video.
[0049] In this embodiment, each of the video clips is associated with a unique second timestamp, and multiple video clips can be merged according to the monotonically increasing characteristic of the second timestamp to obtain a corresponding merged video; the audio data and the merged video are merged to obtain corresponding audio and video; wherein the file format of the audio and video is WebM format.
[0050] In one embodiment, after step S150, the method further includes: playing the audio and video at a preset playing speed on the target display interface.
[0051] In this embodiment, after the corresponding audio and video are obtained, the audio and video can be played at a preset playback speed on the target display interface to realize real-time question-and-answer interaction of the digital human. For ease of description, a medical business application scenario is taken as an example: the user asks a question through the interactive interface of the client, such as "Which department should I go to for a cold or fever?" After receiving the question raised by the user, the server will perform question recall processing from a preset standard question library to obtain the corresponding audio file; load the audio file to obtain the corresponding audio data; extract features from the audio data using a preset audio feature extraction model to obtain corresponding audio feature data; segment the audio feature data according to a preset segmentation coefficient to obtain multiple audio feature segments; based on depth The learning framework processes multiple audio feature segments in parallel to obtain a video segment corresponding to each audio feature segment; merges the audio data and multiple video segments to obtain corresponding audio and video; wherein the audio and video are digital human audio and video; the server sends the audio and video to the client to play the audio and video on the target display interface of the client. During the audio and video playback, the digital human in the audio and video will make lip shapes that match the audio content (i.e., the content of the audio file) to respond to customer questions; wherein the audio content is the content of the answer to the question raised by the user (such as "Which department should I go to for a cold or fever?") (such as "When you have a cold or fever, you can go to the respiratory department for treatment").
[0052] In summary, the embodiments of the present invention can simultaneously process multiple audio feature segments in parallel based on a deep learning framework, which can effectively reduce the time for synthesizing audio and video, improve the efficiency of audio and video synthesis, and achieve rapid response to customer interaction needs.
[0053] Figure 2 Schematic block diagram of a WebM video synthesis device provided by an embodiment of the present invention. Figure 2 As shown, corresponding to the above WebM video synthesis method, the present invention also provides a WebM video synthesis device, the device is configured as follows Figure 4In an application environment, the client communicates with the server through a network. The server can receive customer questions through the client, and perform question recall processing from a preset standard question library based on the customer questions to obtain the corresponding audio file; load the audio file to obtain the corresponding audio data; perform feature extraction on the audio data by a preset audio feature extraction model to obtain the corresponding audio feature data; segment the audio feature data according to a preset segmentation coefficient to obtain multiple audio feature segments; perform parallel processing on multiple audio feature segments based on a deep learning framework to obtain a video segment corresponding to each audio feature segment; merge the audio data and multiple video segments to obtain the corresponding audio and video; the server sends the audio and video to the client to play the audio and video on the target display interface of the client to respond to the customer's questions; wherein the client can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. Specifically, please refer to Figure 2 , the WebM video synthesis device 700 comprises:
[0054] The loading unit 701 is used to load a pre-stored audio file to obtain corresponding audio data;
[0055] An extraction unit 702 is used to extract features from the audio data using a preset audio feature extraction model to obtain corresponding audio feature data;
[0056] A segmentation unit 703, configured to segment the audio feature data according to a preset segmentation coefficient to obtain a plurality of audio feature segments;
[0057] A processing unit 704 is used to process the multiple audio feature segments in parallel based on a deep learning framework to obtain a video segment corresponding to each audio feature segment;
[0058] The merging unit 705 is used to merge the audio data and the multiple video clips to obtain corresponding audio and video.
[0059] In some embodiments, when the processing unit 704 performs the step of processing the plurality of audio feature segments in parallel based on the deep learning framework to obtain a video segment corresponding to each audio feature segment, the processing unit 704 is specifically configured to:
[0060] Inferring the audio features in the audio feature segment based on a deep learning framework to obtain an image frame corresponding to each audio feature in the audio feature segment; wherein each audio feature segment includes multiple audio features;
[0061] The plurality of image frames are merged to obtain a video segment corresponding to the audio feature segment.
[0062] In some embodiments, before executing the step of inferring the audio features in the audio feature segment based on the deep learning framework to obtain the image frame corresponding to each audio feature in the audio feature segment, the processing unit 704 is further configured to:
[0063] Get template image data;
[0064] The size of the template image data is converted according to a preset pixel value to obtain a target template image.
[0065] In some embodiments, after executing the step of converting the size of the template image data according to the preset pixel value to obtain the target template image, the processing unit 704 is further configured to:
[0066] The target template image and the audio features are input into the deep learning framework to obtain an initial image frame corresponding to the audio features.
[0067] In some embodiments, after executing the step of obtaining the initial image frame corresponding to the audio feature, the processing unit 704 is further configured to:
[0068] The initial image frame is converted into code according to a preset coding format to obtain an image frame corresponding to the audio feature.
[0069] In some embodiments, before executing the step of loading a pre-stored audio file to obtain corresponding audio data, the loading unit 701 is further used to:
[0070] Determine whether the sampling rate of the audio file is a target sampling rate, and obtain a determination result;
[0071] If the judgment result is no, the sampling rate of the audio file is converted according to the target sampling rate to obtain a converted audio file.
[0072] In some embodiments, after executing the step of obtaining the corresponding audio and video, the merging unit 705 is further used to:
[0073] The audio and video are played at a preset playback speed on the target display interface.
[0074] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned WebM video synthesis device and each unit can refer to the corresponding description in the aforementioned method embodiment, and for the convenience and brevity of description, it will not be repeated here.
[0075] The WebM video synthesis device can be implemented in the form of a computer program. Figure 3 The electronic device shown is running.
[0076] See also Figure 3 , Figure 3 800 is a schematic block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 800 may be a terminal or a server, wherein the terminal may be an electronic device with a communication function. The server may be an independent server or a server cluster composed of multiple servers.
[0077] See also Figure 3 The electronic device 800 includes a processor 802 , a memory and a network interface 805 connected via a system bus 801 , wherein the memory may include a non-volatile storage medium 803 and an internal memory 804 .
[0078] The non-volatile storage medium 803 can store an operating system 8031 and a computer program 8032. The computer program 8032 includes program instructions, and when the program instructions are executed, the processor 802 can execute a WebM video synthesis method.
[0079] The processor 802 is used to provide computing and control capabilities to support the operation of the entire electronic device 800 .
[0080] The internal memory 804 provides an environment for the operation of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, the processor 802 can execute a WebM video synthesis method.
[0081] The network interface 805 is used to communicate with other devices over the network. Figure 3 The structure shown in the figure is merely a block diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the electronic device 800 to which the solution of the present invention is applied. The specific electronic device 800 may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0082] The processor 802 is used to run the computer program 8032 stored in the memory to implement the following steps:
[0083] Load the pre-stored audio file to obtain the corresponding audio data;
[0084] Extracting features from the audio data using a preset audio feature extraction model to obtain corresponding audio feature data;
[0085] Segmenting the audio feature data according to a preset segmentation coefficient to obtain a plurality of audio feature segments;
[0086] Processing the plurality of audio feature segments in parallel based on a deep learning framework to obtain a video segment corresponding to each audio feature segment;
[0087] The audio data and the plurality of video clips are combined to obtain corresponding audio and video.
[0088] In some embodiments, when the processor 802 implements the step of processing the plurality of audio feature segments in parallel based on the deep learning framework to obtain a video segment corresponding to each of the audio feature segments, the processor 802 is specifically configured to:
[0089] Inferring the audio features in the audio feature segment based on a deep learning framework to obtain an image frame corresponding to each audio feature in the audio feature segment; wherein each audio feature segment includes multiple audio features;
[0090] The plurality of image frames are merged to obtain a video segment corresponding to the audio feature segment.
[0091] In some embodiments, before implementing the step of inferring the audio features in the audio feature segment based on the deep learning framework to obtain the image frame corresponding to each audio feature in the audio feature segment, the processor 802 is further configured to:
[0092] Get template image data;
[0093] The size of the template image data is converted according to a preset pixel value to obtain a target template image.
[0094] In some embodiments, after implementing the step of converting the size of the template image data according to the preset pixel value to obtain the target template image, the processor 802 is further configured to:
[0095] The target template image and the audio features are input into the deep learning framework to obtain an initial image frame corresponding to the audio features.
[0096] In some embodiments, after implementing the step of obtaining the initial image frame corresponding to the audio feature, the processor 802 is further configured to:
[0097] The initial image frame is converted into code according to a preset coding format to obtain an image frame corresponding to the audio feature.
[0098] In some embodiments, before implementing the step of loading a pre-stored audio file to obtain corresponding audio data, the processor 802 is further configured to:
[0099] Determine whether the sampling rate of the audio file is a target sampling rate, and obtain a determination result;
[0100] If the judgment result is no, the sampling rate of the audio file is converted according to the target sampling rate to obtain a converted audio file.
[0101] In some embodiments, after implementing the step of obtaining the corresponding audio and video, the processor 802 is further configured to:
[0102] The audio and video are played at a preset playback speed on the target display interface.
[0103] It should be understood that in the embodiment of the present invention, the processor 802 may be a central processing unit (CPU), and the processor 802 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0104] It can be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment can be completed by instructing the relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiment of the above method.
[0105] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor executes the following steps:
[0106] Load the pre-stored audio file to obtain the corresponding audio data;
[0107] Extracting features from the audio data using a preset audio feature extraction model to obtain corresponding audio feature data;
[0108] Segmenting the audio feature data according to a preset segmentation coefficient to obtain a plurality of audio feature segments;
[0109] Processing the plurality of audio feature segments in parallel based on a deep learning framework to obtain a video segment corresponding to each audio feature segment;
[0110] The audio data and the plurality of video clips are combined to obtain corresponding audio and video.
[0111] In one embodiment, when the processor executes the program instructions to implement the step of performing parallel processing on the plurality of audio feature segments based on a deep learning framework to obtain a video segment corresponding to each audio feature segment, the processor is specifically configured to:
[0112] Inferring the audio features in the audio feature segment based on a deep learning framework to obtain an image frame corresponding to each audio feature in the audio feature segment; wherein each audio feature segment includes multiple audio features;
[0113] The plurality of image frames are merged to obtain a video segment corresponding to the audio feature segment.
[0114] In one embodiment, before the processor executes the program instructions to implement the step of inferring the audio features in the audio feature segment based on the deep learning framework to obtain the image frame corresponding to each audio feature in the audio feature segment, it is further used to:
[0115] Get template image data;
[0116] The size of the template image data is converted according to a preset pixel value to obtain a target template image.
[0117] In one embodiment, after the processor executes the program instructions to convert the size of the template image data according to the preset pixel value to obtain the target template image, it is further configured to:
[0118] The target template image and the audio features are input into the deep learning framework to obtain an initial image frame corresponding to the audio features.
[0119] In one embodiment, after executing the program instructions to implement the step of obtaining the initial image frame corresponding to the audio feature, the processor is further configured to:
[0120] The initial image frame is converted into code according to a preset coding format to obtain an image frame corresponding to the audio feature.
[0121] In one embodiment, before the processor executes the program instructions to load the pre-stored audio file and obtain the corresponding audio data step, it is further used to:
[0122] Determine whether the sampling rate of the audio file is a target sampling rate, and obtain a determination result;
[0123] If the judgment result is no, the sampling rate of the audio file is converted according to the target sampling rate to obtain a converted audio file.
[0124] In one embodiment, after executing the program instructions to obtain the corresponding audio and video steps, the processor is further configured to:
[0125] The audio and video are played at a preset playback speed on the target display interface.
[0126] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, etc., which are computer-readable storage media that can store program codes.
[0127] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0128] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0129] The steps in the method of the embodiment of the present invention can be adjusted in order, combined and deleted according to actual needs. The units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0130] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling an electronic device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.
[0131] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A WebM video synthesis method, characterized in that: The method comprises: Load the pre-stored audio file to obtain the corresponding audio data; Extracting features from the audio data using a preset audio feature extraction model to obtain corresponding audio feature data; Segmenting the audio feature data according to a preset segmentation coefficient to obtain a plurality of audio feature segments; Processing the plurality of audio feature segments in parallel based on a deep learning framework to obtain a video segment corresponding to each audio feature segment; The audio data and the plurality of video clips are combined to obtain corresponding audio and video.
2. The WebM video synthesis method according to claim 1, characterized in that: The method of performing parallel processing on the plurality of audio feature segments based on a deep learning framework to obtain a video segment corresponding to each audio feature segment includes: Inferring the audio features in the audio feature segment based on a deep learning framework to obtain an image frame corresponding to each audio feature in the audio feature segment; wherein each audio feature segment includes multiple audio features; The plurality of image frames are merged to obtain a video segment corresponding to the audio feature segment.
3. The WebM video synthesis method according to claim 2, characterized in that: Before the audio features in the audio feature segment are inferred based on the deep learning framework to obtain the image frame corresponding to each audio feature in the audio feature segment, the method further includes: Get template image data; The size of the template image data is converted according to a preset pixel value to obtain a target template image.
4. The WebM video synthesis method according to claim 3, characterized in that: After converting the size of the template image data according to the preset pixel value to obtain the target template image, the method further includes: The target template image and the audio features are input into the deep learning framework to obtain an initial image frame corresponding to the audio features.
5. The WebM video synthesis method according to claim 4, characterized in that: After obtaining the initial image frame corresponding to the audio feature, the method further includes: The initial image frame is converted into code according to a preset coding format to obtain an image frame corresponding to the audio feature.
6. The WebM video synthesis method according to claim 1, characterized in that: Before the pre-stored audio file is loaded to obtain the corresponding audio data, the method further includes: Determine whether the sampling rate of the audio file is a target sampling rate, and obtain a determination result; If the judgment result is no, the sampling rate of the audio file is converted according to the target sampling rate to obtain a converted audio file.
7. The WebM video synthesis method according to claim 1, characterized in that: After obtaining the corresponding audio and video, the method further includes: The audio and video are played at a preset playback speed on the target display interface.
8. A WebM video synthesis device, characterized in that: The device comprises: A loading unit, used to load a pre-stored audio file to obtain corresponding audio data; An extraction unit, configured to extract features from the audio data using a preset audio feature extraction model to obtain corresponding audio feature data; A segmentation unit, used to segment the audio feature data according to a preset segmentation coefficient to obtain a plurality of audio feature segments; A processing unit, configured to perform parallel processing on the plurality of audio feature segments based on a deep learning framework to obtain a video segment corresponding to each of the audio feature segments; The merging unit is used to merge the audio data and the multiple video clips to obtain corresponding audio and video.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the WebM video synthesis method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the WebM video synthesis method according to any one of claims 1 to 7.