Method for producing an audio file, server and storage medium

CN116386586BActive Publication Date: 2026-09-22TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310217680.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2026-09-22
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

然而,目前的由人工听录的歌声曲谱的质量容易受到听录环境和人工操作的影响,从而从歌声曲谱中提取出的演唱内容和音高旋律往往表现的不够准确,导致合成的歌声的自然度不高,以及缺乏表现力和感染力

Benefits of technology

[0052]该方法先通过获取输入音频,以及输入音频对应的歌词序列;其中,输入音频包括用于表达歌词序列的干声音频,干声音频为由多个连续的音频帧组成;从输入音频中提取出关于干声音频的音频帧序列,并对音频帧序列进行基频检测,得到每一音频帧的基频值;以及对音频帧序列和歌词序列进行对齐处理,得到与音频帧序列单调对齐的音节信息,其中,音节信息包括音素时间戳;基于每一音频帧的基频值相对于对应参考值的偏离程度,对音频帧序列中对应偏离程度异常的音频帧的基频值进行修复处理,得到修复后的音频帧序列;根据音节信息和修复后的音频帧序列,生成匹配于音素时间戳的目标音频文件。这样,一方面,先对音频帧序列和歌词序列进行对齐和对音频帧进行修复,得到对应对齐的音节信息和修复的音频帧序列,再根据对齐的音节信息和修复的音频帧序列来生成音频文件,从而优化了音频文件制作的流程,降低了人力和时间成本的消耗;另一方面,利用每一音频帧的基频值相对于对应参考值的偏离程度对异常的音频帧的基频值进行修复,以生成针对于输入音频的音频文件,能够提升制作的音频文件的自然度和表现力,从而利于基于制作的音频文件进行后续的音频处理。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386586B_ABST
    Figure CN116386586B_ABST
Patent Text Reader

Abstract

The application relates to a method for making an audio file, a server and a storage medium. The method comprises the following steps: acquiring input audio and a corresponding lyric sequence; extracting an audio frame sequence of dry audio from the input audio, performing fundamental frequency detection on the audio frame sequence to obtain a fundamental frequency value of each audio frame; performing alignment processing on the audio frame sequence and the lyric sequence to obtain syllable information monotonously aligned with the audio frame sequence, wherein the syllable information comprises a phoneme timestamp; performing repair processing on the fundamental frequency value of an audio frame with an abnormal deviation degree in the audio frame sequence based on the deviation degree of the fundamental frequency value of each audio frame relative to a corresponding reference value to obtain a repaired audio frame sequence; and generating a target audio file matched with the phoneme timestamp according to the syllable information and the repaired audio frame sequence. The method can optimize the process of making an audio file and improve the naturalness and expressiveness of the made audio file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method for creating audio files, a server, and a storage medium. Background Technology

[0002] With the development of internet technology, voice synthesis has emerged as a new application area of ​​speech synthesis technology. It utilizes related technologies to enable computers to produce beautiful and melodious singing voices, much like humans. Therefore, voice synthesis has considerable application value and promising prospects in fields such as virtual singers, record production, and digital music creation.

[0003] Traditional methods of vocal synthesis typically involve synthesizing vocal content (text and duration) and pitch and melody (notes and duration) from a manually recorded musical score. However, the quality of currently recorded musical scores is easily affected by the recording environment and human intervention, resulting in inaccurate extraction of vocal content and pitch / melody. This leads to a lack of naturalness, expressiveness, and emotional impact in the synthesized vocals. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, server, and storage medium for creating audio files that can improve the naturalness and expressiveness of synthesized audio, in response to the aforementioned technical problems.

[0005] According to a first aspect of the present disclosure, a method for creating an audio file is provided, comprising:

[0006] The input audio and the corresponding lyrics sequence are obtained; the input audio includes dry audio for expressing the lyrics sequence, and the dry audio consists of multiple consecutive audio frames;

[0007] Extract an audio frame sequence of the dry audio from the input audio, and perform fundamental frequency detection on the audio frame sequence to obtain the fundamental frequency value of each audio frame; and

[0008] The audio frame sequence and the lyrics sequence are aligned to obtain syllable information that is monotonically aligned with the audio frame sequence. The syllable information includes phoneme timestamps.

[0009] Based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, the fundamental frequency values ​​of audio frames with abnormal deviation in the audio frame sequence are repaired to obtain the repaired audio frame sequence.

[0010] Based on the syllable information and the repaired audio frame sequence, a target audio file matching the phoneme timestamp is generated.

[0011] In one exemplary embodiment, the step of repairing the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence based on the degree of deviation of the fundamental frequency value of each audio frame relative to the corresponding reference value, to obtain a repaired audio frame sequence, includes:

[0012] Based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, abnormal audio frames and empty audio frames are identified in each audio frame; the fundamental frequency value of the empty audio frame does not exist;

[0013] The fundamental frequency value of the abnormal audio frame is corrected, and the fundamental frequency value of the empty audio frame is interpolated to obtain the repaired audio frame sequence.

[0014] In one exemplary embodiment, determining empty audio frames in each audio frame based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value includes:

[0015] Based on the fundamental frequency value of each audio frame, fundamental frequency data corresponding to each audio frame is generated;

[0016] In the baseband data, the audio frames corresponding to the parts where the baseband value does not exist are determined as the empty audio frames;

[0017] The interpolation process for the fundamental frequency value of the empty audio frame includes:

[0018] In the audio frame sequence, the corresponding logarithmic function value is generated using the fundamental frequency value of the preceding adjacent audio frame and the fundamental frequency value of the following adjacent audio frame relative to the empty audio frame.

[0019] The logarithmic function value is interpolated into the empty audio frame to interpolate the fundamental frequency value of the empty audio frame.

[0020] In one exemplary embodiment, identifying abnormal audio frames in each audio frame based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value includes:

[0021] Based on the fundamental frequency value of each audio frame, differential data corresponding to each audio frame is generated; the horizontal axis of the differential data represents the order of each audio frame in the audio frame sequence, and the vertical axis of the differential data represents the differential value of the audio frame corresponding to the order.

[0022] In the differential data, based on the relationship between the distance of the differential value of each position from the corresponding reference value and the preset distance, abnormal audio frames are determined in each audio frame.

[0023] In an exemplary embodiment, determining abnormal audio frames in each audio frame based on the relationship between the distance of the difference value of each position deviating from the corresponding reference value and a preset distance in the differential data includes:

[0024] The difference value of the previous adjacent position or the difference value of the next adjacent position relative to the difference value of each position is determined as the reference value of the difference value of each position.

[0025] Among the difference values ​​of each arrangement position, abnormal difference values ​​are identified where the distance between the difference value and the corresponding reference value is greater than the preset distance, and the audio frame corresponding to the abnormal difference value is identified as the abnormal audio frame.

[0026] In an exemplary embodiment, the step of correcting the fundamental frequency value of the abnormal audio frame includes...

[0027] In the audio frame sequence, the fundamental frequency value of the preceding or following adjacent audio frame in the fundamental frequency image of the audio frame sequence corresponding to the abnormal audio frame is used as the correction fundamental frequency value of the abnormal audio frame.

[0028] The fundamental frequency value of the abnormal audio frame is replaced with the correction fundamental frequency value to perform correction processing on the fundamental frequency value of the abnormal audio frame.

[0029] In one exemplary embodiment, the audio frames in the audio frame sequence include unvoiced frames and voiced frames;

[0030] After repairing the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence to obtain the repaired audio frame sequence, the method further includes:

[0031] Based on the syllable information, the unvoiced and voiced frames in the repaired audio frame sequence are classified.

[0032] Based on a preset marker mask, the unvoiced frames and the voiced frames after classification are restored to obtain a restored audio frame sequence; the restoration process is used to restore the fundamental frequency value of the restored audio frame to the fundamental frequency value obtained by fundamental frequency detection.

[0033] The step of generating a target audio file matching the phoneme timestamp based on the syllable information and the repaired audio frame sequence includes:

[0034] The target audio file is generated by performing data fusion processing on the syllable information and the restored audio frame sequence according to the phoneme timestamp.

[0035] In an exemplary embodiment, the syllable information represents multiple phoneme information and the audio frames occupied by the multiple phoneme information;

[0036] The process of classifying unvoiced and voiced frames in the repaired audio frame sequence based on the syllable information includes:

[0037] Based on the syllable information, the phoneme information corresponding to each audio frame is determined in the repaired audio frame sequence;

[0038] Based on the type of phoneme information corresponding to each audio frame, each audio frame in the repaired audio frame sequence is classified as either an unvoiced frame or a voiced frame.

[0039] According to a second aspect of the present disclosure, an apparatus for creating audio files is provided, comprising:

[0040] The information acquisition unit is configured to acquire input audio and the lyrics sequence corresponding to the input audio; the input audio includes dry audio for expressing the lyrics sequence, and the dry audio consists of multiple consecutive audio frames;

[0041] The audio processing unit is configured to extract an audio frame sequence about the dry audio from the input audio, and perform fundamental frequency detection on the audio frame sequence to obtain the fundamental frequency value of each audio frame;

[0042] A monotonic alignment unit is configured to perform alignment processing on the audio frame sequence and the lyrics sequence to obtain syllable information monotonically aligned with the audio frame sequence, the syllable information including phoneme timestamps;

[0043] The audio repair unit is configured to perform repair processing on the fundamental frequency values ​​of audio frames in the audio frame sequence that have abnormal deviations based on the degree of deviation of the fundamental frequency value of each audio frame relative to the corresponding reference value, so as to obtain a repaired audio frame sequence.

[0044] The template generation unit is configured to generate a target audio file matching the phoneme timestamp based on the syllable information and the repaired audio frame sequence.

[0045] According to a third aspect of the present disclosure, a server is provided, comprising:

[0046] processor;

[0047] Memory for storing the executable instructions of the processor;

[0048] The processor is configured to execute the executable instructions to implement the method for creating an audio file as described in any of the preceding claims.

[0049] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, the computer-readable storage medium including program data, which, when executed by a processor of a server, enables the server to perform the method for creating an audio file as described in any of the preceding claims.

[0050] According to a fifth aspect of the present disclosure, a computer program product is also provided, the computer program product including program instructions that, when executed by a processor of a server, enable the server to perform the method for creating an audio file as described in any of the preceding claims.

[0051] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0052] The method first acquires input audio and the corresponding lyrics sequence. The input audio includes dry audio representing the lyrics sequence, which consists of multiple consecutive audio frames. An audio frame sequence related to the dry audio is extracted from the input audio, and fundamental frequency detection is performed on the audio frame sequence to obtain the fundamental frequency value of each audio frame. The audio frame sequence and the lyrics sequence are then aligned to obtain syllable information monotonically aligned with the audio frame sequence, where the syllable information includes phoneme timestamps. Based on the deviation of the fundamental frequency value of each audio frame from its corresponding reference value, the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence are repaired to obtain a repaired audio frame sequence. Finally, a target audio file matching the phoneme timestamps is generated based on the syllable information and the repaired audio frame sequence. In this way, on the one hand, the audio frame sequence and lyrics sequence are aligned and the audio frames are repaired to obtain the corresponding aligned syllable information and the repaired audio frame sequence. Then, the audio file is generated based on the aligned syllable information and the repaired audio frame sequence, thereby optimizing the audio file production process and reducing the consumption of manpower and time costs. On the other hand, the deviation of the fundamental frequency value of each audio frame from the corresponding reference value is used to repair the fundamental frequency value of abnormal audio frames in order to generate an audio file for the input audio. This can improve the naturalness and expressiveness of the produced audio file, thus facilitating subsequent audio processing based on the produced audio file.

[0053] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0054] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0055] Figure 1 This is an application environment diagram illustrating a method for creating an audio file according to an exemplary embodiment.

[0056] Figure 2 This is a flowchart illustrating a method for creating an audio file according to an exemplary embodiment.

[0057] Figure 3 This is a flowchart illustrating a step for repairing the fundamental frequency value of an audio frame according to an exemplary embodiment.

[0058] Figure 4 This is a flowchart illustrating a step of determining empty audio frames in each audio frame according to an exemplary embodiment.

[0059] Figure 5 This is a schematic diagram illustrating the fundamental frequency image of an audio frame sequence according to an exemplary embodiment.

[0060] Figure 6 This is a flowchart illustrating a step of interpolating the fundamental frequency value of an empty audio frame according to an exemplary embodiment.

[0061] Figure 7 This is a schematic diagram illustrating a logarithmic curve image according to an exemplary embodiment.

[0062] Figure 8 This is a flowchart illustrating a step of identifying abnormal audio frames in each audio frame according to an exemplary embodiment.

[0063] Figure 9 This is a schematic diagram illustrating a differential image of an audio frame sequence according to an exemplary embodiment.

[0064] Figure 10 This is a flowchart illustrating a step for correcting the fundamental frequency value of an abnormal audio frame according to an exemplary embodiment.

[0065] Figure 11 This is a flowchart illustrating a step of restoring an audio frame according to an exemplary embodiment.

[0066] Figure 12 This is a flowchart illustrating a step of restoring classified audio frames according to another exemplary embodiment.

[0067] Figure 13 This is a flowchart illustrating another method for creating an audio file according to an exemplary embodiment.

[0068] Figure 14 This is a block diagram illustrating a method for creating an audio file according to another exemplary embodiment.

[0069] Figure 15 This is a block diagram of an audio file creation apparatus according to an exemplary embodiment.

[0070] Figure 16 This is a block diagram illustrating a server for creating audio files according to an exemplary embodiment.

[0071] Figure 17 This is a block diagram illustrating a computer-readable storage medium for creating audio files according to an exemplary embodiment.

[0072] Figure 18 This is a block diagram illustrating a computer program product for creating audio files according to an exemplary embodiment. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0074] The term "and / or" in the embodiments of this application refers to any and all possible combinations including one or more of the associated listed items. It should also be noted that, when used in this specification, "including / comprising" specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or components and / or groups thereof.

[0075] The terms "first," "second," etc., used in this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0076] Furthermore, although the terms "first," "second," etc., are used repeatedly in this application to describe various operations (or various elements, or various applications, or various instructions, or various data), these operations (or elements, or applications, or instructions, or data) should not be limited by these terms. These terms are only used to distinguish one operation (or element, or application, or instruction, or data) from another operation (or element, or application, or instruction, or data). For example, a first marker mask can be called a second marker mask, and a second marker mask can be called a first marker mask; the only difference is the scope they encompass, but it does not depart from the scope of this application. Both the first marker mask and the second marker mask are sets of marker masks corresponding to various categories of audio frames, but they are not sets of marker masks corresponding to the same category of audio frames.

[0077] The audio file creation method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a communication network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on other network servers.

[0078] In some embodiments, reference Figure 1 First, server 104 acquires the input audio and the corresponding lyrics sequence. The input audio includes dry audio used to express the lyrics sequence, which consists of multiple consecutive audio frames. Then, server 104 extracts the audio frame sequence related to the dry audio from the input audio, performs fundamental frequency detection on the audio frame sequence to obtain the fundamental frequency value of each audio frame, and performs alignment processing on the audio frame sequence and the lyrics sequence to obtain syllable information monotonically aligned with the audio frame sequence. The syllable information includes phoneme timestamps. Then, based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, server 104 repairs the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence to obtain a repaired audio frame sequence. Finally, server 104 generates a target audio file matching the phoneme timestamps based on the syllable information and the repaired audio frame sequence.

[0079] In some embodiments, terminal 102 (such as a mobile terminal or a fixed terminal) can be implemented in various forms. Terminal 102 can be a mobile terminal, including mobile phones, smartphones, laptops, portable handheld devices, personal digital assistants (PDAs), tablet computers (PADs), etc., capable of generating target audio files matching phoneme timestamps based on syllable information and a repaired audio frame sequence. Terminal 102 can also be a fixed terminal, including Automated Teller Machines (ATMs), automated multi-function machines, digital TVs, desktop computers, fixed-line computers, etc., capable of generating target audio files matching phoneme timestamps based on syllable information and a repaired audio frame sequence.

[0080] Hereinafter, it is assumed that terminal 102 is a fixed terminal. However, those skilled in the art will understand that, if there are operations or elements specifically designed for mobile purposes, the construction according to the embodiments disclosed in this application can also be applied to mobile type terminal 102.

[0081] In some embodiments, the data processing component running on server 104 may load any of the various additional server applications and / or middleware applications being executed, such as HTTP (Hypertext Transfer Protocol), FTP (File Transfer Protocol), CGI (Common Gateway Interface), RDBMS (Relational Database Management System), etc.

[0082] In some embodiments, server 104 may be implemented using a standalone server or a server cluster consisting of multiple servers. Server 104 may be adapted to run one or more application services or software components that provide the terminal 102 described in the foregoing disclosure.

[0083] In some embodiments, the application service may include a service interface that provides users with audio file selection and generation, as well as corresponding program services, etc. The software component may include, for example, an application (SDK) or client (APP) that has the function of extracting lyrics and fundamental frequency of a song based on the user's input.

[0084] In some embodiments, the application or client provided by server 104, which has the function of extracting lyrics and baseband of a song based on the user's input song, includes a portal port that provides one-to-one application services to the user in the foreground and multiple business systems that perform data processing in the background. This extends the relevant application functions for the audio file obtained by subsequently repairing and synthesizing the audio frame sequence to the APP or client, so that the user can use and access the functions associated with creating audio files anytime and anywhere.

[0085] In some embodiments, the resource transfer function of an APP or client can be a computer program running in user mode to complete one or more specific tasks, which can interact with the user and has a visual user interface. The APP or client can include two parts: a graphical user interface (GUI) and an engine, which together provide users with a variety of application services in the form of a user interface in a digital client system.

[0086] In some embodiments, users can input corresponding code data or control parameters into the APP or client through a preset input device or automatic control program to execute application services of the computer program in the server 104 and display application services in the user interface.

[0087] As an example, when a user needs to synthesize a song into a vocal file on terminal 102, the user can input the input audio and the corresponding lyrics sequence control parameters to terminal 102 through the input device. Then, server 104 executes the audio file creation method on the input audio and lyrics sequence, thereby generating an audio file for the input audio based on the input audio and lyrics sequence. Finally, server 104 sends information data about the audio file to terminal 102 so that the created audio file can be run in the APP or client of terminal 102.

[0088] In some embodiments, the operating system running the app or client may include various versions of Microsoft... Apple and / or Linux operating system, various commercial or similar Operating systems (including but not limited to various GNU / Linux operating systems, Google) OS and / or mobile operating systems, such as Phone OS OS OS operating systems, as well as other online or offline operating systems, are not specifically limited here.

[0089] In some embodiments, such as Figure 2 As shown, a method for creating audio files is provided, which can be applied to... Figure 1 Taking server 104 as an example, the method includes the following steps:

[0090] Step S11: Obtain the input audio and the corresponding lyrics sequence.

[0091] In some embodiments, the server obtains the input audio transmitted by the user account and the corresponding lyrics sequence from the terminal application (such as a mobile phone, tablet, etc.).

[0092] The input audio can be a released, officially released version of a music song, or a local song recorded by the terminal application (e.g., a live song recorded offline by the terminal application and a web song recorded online).

[0093] In some embodiments, the input audio includes dry audio for expressing a sequence of lyrics and accompaniment audio for expressing a musical melody. The dry audio consists of multiple consecutive audio frames.

[0094] In some embodiments, the server may first acquire the input audio transmitted by the user account (e.g., a live song recorded offline by the terminal application), and then pass the input audio into a preset audio-accompaniment separation model (e.g., the Spleeter algorithm) to separate the dry audio from the accompaniment, thereby extracting the dry audio of the input audio, which consists of multiple consecutive audio frames. Then, a preset lyrics parsing model is used to extract lyrics from the dry audio of each audio frame to obtain the lyrics sequence of the input audio.

[0095] As an example, the Spleeter algorithm first re-segments the song fragments and projects the segmented song into a low-dimensional space to reduce dimensionality and compress the information data, thereby extracting audio depth features related to the song. Then, a multilayer perceptron based on MLP is used to classify the audio depth features, obtaining audio depth features for the dry vocals and the accompaniment. Finally, the compressed low-dimensional features are restored to the original dimensions of the dry vocals and accompaniment audio.

[0096] As an example, the lyrics parsing model first extracts the text information of the dry audio from each audio frame using a pre-defined speech recognition algorithm. This speech recognition algorithm can be based on Dynamic Time Warping, Hidden Markov Models (HMMs) based on parametric models, artificial neural networks, or a hybrid algorithm, etc. Then, the lyrics parsing model performs word segmentation on the text information of the dry audio to obtain the words to be recognized. Finally, the words to be recognized in the dry audio are matched with the lyrics text information corresponding to multiple songs to obtain the dry audio lyrics sequence.

[0097] Step S12: Extract the audio frame sequence of the dry audio from the input audio, and perform fundamental frequency detection on the audio frame sequence to obtain the fundamental frequency value of each audio frame.

[0098] In some embodiments, the server first obtains the separated vocal audio from the vocal accompaniment separation model, and then decomposes the vocal audio according to a preset frame rate (e.g., 0.10 ms, 0.20 ms) to obtain a plurality of consecutive audio frames, so as to form an audio frame sequence for the vocal audio. Then, the server performs fundamental frequency detection on each audio frame in the audio frame sequence through a preset fundamental frequency detection algorithm (e.g., yin algorithm, pyin algorithm, crepe algorithm, etc.) to obtain the fundamental frequency value of each audio frame.

[0099] As an example, the preset fundamental frequency detection algorithm first preprocesses each audio frame (including calculating the short-time energy and zero-crossing rate of the audio frame signal s(t), adopting a double-threshold method for prosodic segmentation, and filtering the segmented signal through a band-pass filter of 50 to 1500 Hz) to obtain the preprocessed speech signal. Then, Fourier transform is performed on the preprocessed speech signal to obtain the spectrum of the speech signal; Top-hat transform is performed on the speech spectrum to detect the spectrum envelope; then a local minimum-maximum method is used for peak detection on the spectrum envelope to determine each peak region in the spectrum envelope, and Hilbert transform is used to solve the peak regions whose peak energy exceeds half of the maximum peak energy to obtain the instantaneous fundamental frequency value of the audio frame corresponding to a certain frame. Finally, the rectangular window function is used to smooth the instantaneous fundamental frequency value to complete the fundamental frequency detection of the audio frame of this frame, and the corresponding fundamental frequency value is obtained.

[0100] Step S13: Align the audio frame sequence and the lyric sequence to obtain syllable information that is monotonically aligned with the audio frame sequence.

[0101] In some embodiments, the server uses a forced alignment algorithm (e.g., monotonic alignment search (MAS), Needleman–Wunsch algorithm, etc.) to perform phoneme / syllable-wise monotonic alignment on each audio frame in the audio frame sequence and the lyric sequence, so as to obtain phoneme information and / or syllable information that is monotonically aligned with the audio frame sequence.

[0102] In some embodiments, a phoneme (phoneme) is the smallest unit of Chinese pronunciation (the smallest unit of writing is grapheme), for example, for the Chinese character "好", the pronunciation units are h and ao.

[0103] In some embodiments, a syllable is the smallest speech unit of combined pronunciation of a single vowel phoneme and consonant phonemes in language, and a single vowel phoneme can also form a syllable by itself. A Chinese syllable is a speech unit formed by combining an initial consonant and a final vowel, and a single vowel can also form a syllable by itself.

[0104] In some embodiments, each phoneme / syllable in the phoneme information and / or syllable information that is monotonically aligned with the audio frame sequence is aligned with at least one audio frame in the audio frame sequence. Wherein, the phoneme information and / or syllable information comprises a phoneme timestamp for each phoneme.

[0105] As an example, the server uses a forced alignment algorithm to perform phoneme-by-phoneme monotonic alignment on each audio frame in an audio frame sequence and the corresponding lyric sequence according to the phoneme timestamp of each phoneme, so as to obtain lyric information monotonically aligned with the audio frame sequence. Wherein, the lyric "hao de" in the lyric sequence is sequentially decomposed into four phonemes as pronunciation units: "h", "ao", "d", and "e", the phoneme "h" is aligned with the first to third audio frames in the audio frame sequence, the phoneme "ao" is aligned with the fourth to tenth audio frames in the audio frame sequence, the phoneme "d" is aligned with the eleventh to fifteenth audio frames in the audio frame sequence, and the phoneme "e" is aligned with the sixteenth to twentieth audio frames in the audio frame sequence.

[0106] As another example, the server uses a forced alignment algorithm to perform syllable-by-syllable monotonic alignment on each audio frame in an audio frame sequence and the corresponding lyric sequence according to the phoneme timestamp of each phoneme, so as to obtain lyric information monotonically aligned with the audio frame sequence. Wherein, the lyric "hao de" in the lyric sequence is sequentially decomposed into two syllables as pronunciation units: "hao" and "de", the syllable "hao" is aligned with the first to tenth audio frames in the audio frame sequence, and the syllable "de" is aligned with the eleventh to twentieth audio frames in the audio frame sequence.

[0107] Step S14: Based on the deviation degree of the fundamental frequency value of each audio frame relative to the corresponding reference value, repair processing is performed on the fundamental frequency value of the audio frame with abnormal deviation degree in the audio frame sequence, so as to obtain a repaired audio frame sequence.

[0108] In some embodiments, the server first determines the reference value corresponding to the fundamental frequency value of each audio frame, wherein the reference value can be manually set by design engineers or calculated by the server according to a preset calculation rule; then, the server further determines the deviation degree between the fundamental frequency value of each audio frame and its corresponding reference value, and determines abnormal audio frames in the audio frame sequence according to the deviation degree corresponding to each audio frame; finally, the server performs repair processing on the abnormal audio frames according to a preset repair method, so as to obtain a repaired audio frame sequence.

[0109] Wherein, the preset repair method can be a method of interpolating or correcting the fundamental frequency value of abnormal audio frames, which is not specifically limited herein.

[0110] Step S15: Generate a target audio file that matches the phoneme timestamp based on the syllable information and the repaired audio frame sequence.

[0111] In some embodiments, the server can adjust the syllable information (or phoneme information) and the repaired audio frame sequence into aligned syllable information (or phoneme information) and audio frame sequences with the same vector length, according to the duration parameter and fundamental frequency value of each audio frame, the audio frame position occupied by the syllable / phoneme, and the phoneme timestamp of each phoneme. Then, a preset synthesizer is used to generate a target audio file matching the phoneme timestamps.

[0112] Among them, audio synthesis technology has been widely used in speech synthesis due to its advantages such as large adjustment capability and strong speech plasticity. In practice, LPC (linear predictive coding) filters can be used as synthesizers, and this application does not limit the specific synthesizer.

[0113] Since the syllable information (or phoneme information) and audio frame sequence with the same vector length and alignment are added, the synthesized audio file has the same melody and rhythm as the input audio.

[0114] In the aforementioned audio file creation process, the server first acquires the input audio and the corresponding lyrics sequence. The input audio includes dry audio used to express the lyrics sequence, which consists of multiple consecutive audio frames. Then, the server extracts the audio frame sequence related to the dry audio from the input audio, performs fundamental frequency detection on the audio frame sequence to obtain the fundamental frequency value of each audio frame, and performs monotonic alignment processing on the audio frame sequence and the lyrics sequence to obtain syllable information monotonically aligned with the audio frame sequence. This syllable information includes phoneme timestamps. Next, based on the deviation of the fundamental frequency value of each audio frame from its corresponding reference value, the server repairs the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence to obtain a repaired audio frame sequence. Finally, the server generates a target audio file matching the phoneme timestamps based on the syllable information and the repaired audio frame sequence. In this way, on the one hand, the audio frame sequence and lyrics sequence are first monotonically aligned and the audio frames are repaired, and then the aligned lyrics information and the repaired audio frame sequence are merged to generate an audio template, thereby optimizing the audio template production process and reducing the consumption of manpower and time costs; on the other hand, the deviation of the fundamental frequency value of each audio frame from the corresponding reference value is used to repair the fundamental frequency value of abnormal audio frames in order to generate an audio template for the input audio, which can improve the naturalness and expressiveness of the produced audio template, thus facilitating subsequent audio processing based on the produced audio template.

[0115] Those skilled in the art will understand that the methods disclosed in the above-described specific embodiments can be implemented in more specific ways. For example, the implementation described above, in which the server repairs the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, to obtain the repaired audio frame sequence, is merely illustrative.

[0116] For example, the server extracts an audio frame sequence of dry audio from the input audio and performs fundamental frequency detection on the audio frame sequence to obtain the fundamental frequency value of each audio frame; or the server aligns the audio frame sequence and the lyrics sequence to obtain syllable information monotonically aligned with the audio frame sequence, etc. This is just one way of setting things together. In actual implementation, there may be other ways of dividing them. For example, the audio frame sequence of dry audio and the monotonically aligned syllable information can be combined or set into another system, or some features can be ignored or not performed.

[0117] In one exemplary embodiment, see Figure 3 , Figure 3 This is a flowchart illustrating an embodiment of the fundamental frequency value repair processing for audio frames in this application. In step S14, the server repairs the fundamental frequency values ​​of audio frames with abnormal deviations based on the degree of deviation of each audio frame's fundamental frequency value relative to the corresponding reference value, obtaining a repaired audio frame sequence. This process can be implemented in the following way:

[0118] Step S141: Based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, abnormal audio frames and empty audio frames are identified in each audio frame.

[0119] Step S142: Correct the fundamental frequency value of the abnormal audio frame and interpolate the fundamental frequency value of the empty audio frame to obtain the repaired audio frame sequence.

[0120] Specifically, the server can use the following steps a1-a4 to identify empty audio frames in each audio frame and perform interpolation processing on the fundamental frequency value of the empty audio frames; and use steps b1-b4 to identify abnormal audio frames in each audio frame and perform bias correction processing on the fundamental frequency value of the abnormal audio frames.

[0121] In some embodiments, the order in which the server repairs audio frames can be either to first correct the fundamental frequency value of abnormal audio frames and then interpolate the fundamental frequency value of empty audio frames; or to first interpolate the fundamental frequency value of empty audio frames and then correct the fundamental frequency value of abnormal audio frames.

[0122] Since interpolation can affect the accuracy of abnormal audio frame identification, it is advisable to first identify and correct the fundamental frequency value of the abnormal audio frame, and then identify and interpolate the fundamental frequency value of the empty audio frame.

[0123] In one exemplary embodiment, see Figure 4 , Figure 4 This is a schematic flowchart illustrating an embodiment of determining empty audio frames in each audio frame according to this application. In step S141, the server determines empty audio frames in each audio frame based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value. This process can be implemented in the following way:

[0124] Step a1: Generate the fundamental frequency data corresponding to each audio frame based on the fundamental frequency value of each audio frame.

[0125] See Figure 5 , Figure 5 This is a schematic diagram of an embodiment of the fundamental frequency image of an audio frame sequence in this application. The fundamental frequency data of the audio frames is represented based on the corresponding fundamental frequency image. The horizontal axis of the fundamental frequency image represents the position of each audio frame in the audio frame sequence, and the vertical axis represents the fundamental frequency value of each audio frame in the audio frame sequence. That is, for the coordinate (X, Y) in the fundamental frequency image, X represents the Xth audio frame in the audio frame sequence, and Y represents the fundamental frequency value of the corresponding Xth audio frame in the audio frame sequence.

[0126] Step a2: In the baseband data, the audio frames corresponding to the parts where the baseband value does not exist are identified as empty audio frames.

[0127] In some embodiments, the fundamental frequency value of an empty audio frame does not exist, which can be understood as the fundamental frequency value of an empty audio frame being 0.

[0128] Continue as Figure 5 As shown, if there are multiple image segments near the parallel line with a vertical coordinate of 0 in the fundamental frequency image, the server will identify the audio frame corresponding to each of these multiple image segments as an empty audio frame. For example... Figure 5 If the vertical coordinate of the image segment containing the server audio frame A1 is 0, then the server will identify audio frame A1 as an empty audio frame.

[0129] In one exemplary embodiment, see Figure 6 , Figure 6 This is a flowchart illustrating an embodiment of the interpolation processing of the fundamental frequency value of the empty audio frame in this application. In step S142, the process of the server interpolating the fundamental frequency value of the empty audio frame can be implemented in the following way:

[0130] Step a3: In the audio frame sequence, the corresponding logarithmic function value is generated using the fundamental frequency value of the preceding and following adjacent audio frames relative to the empty audio frame.

[0131] See Figure 7 , Figure 7 This is a schematic diagram of an embodiment of the logarithmic curve in this application. The horizontal axis of the first endpoint A of the logarithmic curve represents the rank of the preceding adjacent audio frame of the empty audio frame, and the horizontal axis of the second endpoint B represents the rank of the following adjacent audio frame of the empty audio frame. The vertical axis of the first endpoint A of the logarithmic curve represents the fundamental frequency value of the preceding adjacent audio frame of the empty audio frame, and the vertical axis of the second endpoint B represents the fundamental frequency value of the following adjacent audio frame of the empty audio frame. That is, the server generates the corresponding logarithmic function value and the corresponding function curve based on the rank and fundamental frequency value of the preceding and following adjacent audio frames of the empty audio frame. The vertical axis of the midpoint C of the logarithmic curve is the logarithmic function value.

[0132] Step a4 involves interpolating the logarithmic function value into the empty audio frame to interpolate the fundamental frequency value of the empty audio frame.

[0133] In some embodiments, the server generates a corresponding logarithmic function value and inserts it into the empty audio frame to fill the fundamental frequency value of the empty audio frame. Continuing as... Figure 7 As shown, the first endpoint A of the logarithmic curve is connected to the previous adjacent audio frame of the empty audio frame, and the second endpoint B of the logarithmic curve is connected to the next adjacent audio frame of the empty audio frame, thereby interpolating the fundamental frequency value of the empty audio frame.

[0134] In some embodiments, the logarithmic curve corresponding to the audio frame sequence can be a sigmoid function. The sigmoid function not only has a smooth curve shape, but its derivative at the connection points is close to 0, allowing it to connect well with the non-zero fundamental frequency values ​​before and after empty audio frames.

[0135] In other embodiments, the logarithmic curve corresponding to the audio frame sequence can be a linear function or a spline function.

[0136] The linear function is a straight line connecting the non-zero fundamental frequency values ​​at the beginning and end of an empty audio frame, and the inserted value lies on this straight line.

[0137] The spline function is a quadratic spline interpolation function. It obtains a set of valid discrete data and then fits this data with a quadratic function curve. Assuming there are 4 points, x0, x1, x2, and x3, and 3 intervals, 3 quadratic splines are needed. Each quadratic spline is ax^2 + bx + c, resulting in a total of 9 unknowns (3 * 3 = 9). The fitted curve is then used to interpolate the missing fundamental frequency curve.

[0138] In one exemplary embodiment, see Figure 8 , Figure 8 This is a schematic flowchart illustrating an embodiment of identifying abnormal audio frames in each audio frame according to this application. In step S141, the server determines the abnormal audio frame in each audio frame based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value. This process can be implemented in the following way:

[0139] Step b1: Generate differential data corresponding to each audio frame based on the fundamental frequency value of each audio frame.

[0140] See Figure 9 , Figure 9 This is a schematic diagram of an embodiment of the differential image of an audio frame sequence in this application. The differential data of the audio frames is represented based on the corresponding differential image. The horizontal axis of the differential image represents the position of each audio frame in the audio frame sequence, and the vertical axis represents the difference value of each audio frame at the corresponding position in the audio frame sequence. That is, for the coordinates (X', Y') in the differential image, X' represents the Xth audio frame in the audio frame sequence, and Y' represents the difference value of the corresponding Xth audio frame in the audio frame sequence.

[0141] Step b2: In the differential data, based on the relationship between the distance of the differential value of each position from the corresponding reference value and the preset distance, abnormal audio frames are identified in each audio frame.

[0142] In some embodiments, the process by which the server identifies abnormal audio frames in each audio frame based on the relationship between the distance of the difference value of each position deviating from the corresponding reference value and a preset distance in the difference image can be implemented in the following ways:

[0143] Step b21: Determine the difference value of the previous adjacent position or the difference value of the next adjacent position relative to the difference value of each arrangement position as the reference value of the difference value of each arrangement position.

[0144] Step b22: Among the difference values ​​of each arrangement position, identify the abnormal difference values ​​where the difference value deviates from the corresponding reference value by a distance greater than a preset distance, and identify the audio frame corresponding to the abnormal difference value as an abnormal audio frame.

[0145] In some embodiments, the server uses the difference value of the audio frame in the preceding position of the difference value of each permutation position as a reference value for the difference value of each permutation position. Then, the server determines a first distance (e.g., a straight-line distance or a perpendicular distance) between the difference value of each permutation position and its corresponding reference value. Next, the server identifies abnormal difference values ​​among the difference values ​​of each permutation position where the first distance is greater than a preset second distance, and determines the audio frame corresponding to the abnormal difference value as an abnormal audio frame.

[0146] In one exemplary embodiment, see Figure 10 , Figure 10 This is a flowchart illustrating an embodiment of the fundamental frequency correction processing for abnormal audio frames in this application. In step S142, the server performs fundamental frequency correction processing on the abnormal audio frame, which can be implemented in the following way:

[0147] Step b3: In the audio frame sequence, the fundamental frequency value of the preceding or following adjacent audio frame in the fundamental frequency image of the audio frame sequence corresponding to the abnormal audio frame is used as the correction fundamental frequency value of the abnormal audio frame.

[0148] Step b4: Replace the fundamental frequency value of the abnormal audio frame with the correction fundamental frequency value to perform correction processing on the fundamental frequency value of the abnormal audio frame.

[0149] As an example, the fundamental frequency value of the abnormal audio frame S1 is P1, and the fundamental frequency value of the preceding adjacent audio frame S2 in the differential image of the audio frame sequence is P2. Then the server uses the fundamental frequency value of the preceding adjacent audio frame S2, P2, as the correction fundamental frequency value of the abnormal audio frame S1, and replaces the fundamental frequency value P1 of the abnormal audio frame with the correction fundamental frequency value P2 to perform correction processing on the fundamental frequency value P1 of the abnormal audio frame S1.

[0150] In some embodiments, syllable information (or phoneme information) represents multiple phoneme information and the audio frames occupied by the multiple phoneme information. That is, each phoneme in the syllable information (or phoneme information) monotonically aligned with the audio frame sequence is aligned with at least one audio frame in the audio frame sequence.

[0151] In some embodiments, the audio frames in the audio frame sequence include voiced frames and unvoiced frames.

[0152] In one exemplary embodiment, see Figure 11 , Figure 11 This is a schematic flowchart illustrating an embodiment of the audio frame restoration process in this application. After step S15, the server can also implement it in the following ways:

[0153] Step c1: Based on syllable information, classify and process unvoiced and voiced frames in the repaired audio frame sequence.

[0154] In one embodiment, the process of classifying unvoiced and voiced frames in the repaired audio frame sequence based on syllable information by the server can be implemented in the following way:

[0155] Step c11: Based on syllable information, determine the phoneme information corresponding to each audio frame in the repaired audio frame sequence.

[0156] Step c12: Based on the type of phoneme information corresponding to each audio frame, classify each audio frame in the repaired audio frame sequence into unvoiced frames or voiced frames.

[0157] In some embodiments, the server inputs the repaired audio frame sequence and syllable information into a preset convolutional network to identify the phoneme information corresponding to each audio frame, and classifies each audio frame in the repaired audio frame sequence into unvoiced frames or voiced frames according to the type of phoneme information corresponding to each audio frame.

[0158] Specifically, the convolutional network normalizes the audio frames in the repaired audio frame sequence that correspond to each phoneme based on the aligned syllable information. Then, it uses a preset convolutional layer to perform data augmentation on the normalized audio frames. Next, it uses a preset linear layer to perform data channel descent on the data-augmented audio frames to obtain one-dimensional audio frames. Finally, it uses the sigmoid function to identify and classify the one-dimensional audio frames to determine whether each audio frame in the repaired audio frame sequence is a voiceless frame or a voiced frame.

[0159] Step c2: Based on the preset marker mask, the unvoiced frames and voiced frames after classification are restored to obtain the restored audio frame sequence.

[0160] In some embodiments, the restoration process is used to restore the fundamental frequency value of the repaired audio frame to the fundamental frequency value obtained from fundamental frequency detection.

[0161] In one exemplary embodiment, see Figure 12 , Figure 12 This is a flowchart illustrating an embodiment of the audio frame restoration process according to this application. In step c2, the server performs restoration processing on the unvoiced audio frames and the voiced audio frames after classification, based on a preset marker mask. This process can be implemented in the following ways:

[0162] Step d1: Generate a first marker mask for the classified unvoiced frames and a second marker mask for the classified voiced frames.

[0163] Step d2: Assign a first restored value to the first marker mask and a second restored value to the second marker mask.

[0164] Step d3 involves fusing the first restored value of the first marker mask with the fundamental frequency value of the corresponding classified unvoiced frame to restore the classified unvoiced frames in the repaired audio frame sequence.

[0165] Step d4: The second restored value of the second marker mask is fused with the fundamental frequency value of the corresponding classified voiced frame to restore the classified voiced frame in the repaired audio frame sequence.

[0166] As an example, the server applies a first marker mask to the classified unvoiced frames and a second marker mask to the classified voiced frames. Then, the server assigns a first restored value of "0" to the first marker mask of each unvoiced frame and a second restored value of "1" to the second marker mask of each voiced frame. Finally, the server fuses each unvoiced frame with its corresponding first marker mask to multiply the fundamental frequency value of each unvoiced frame by the first restored value "0" to obtain the fundamental frequency value of the unvoiced frame restored to the value of "0". The server also fuses each voiced frame with its corresponding second marker mask to multiply the fundamental frequency value of each voiced frame by the second restored value "1" to obtain the fundamental frequency value of the voiced frame restored to its original value. That is, the fundamental frequency value of the voiced frame remains unchanged after restoration.

[0167] In some embodiments, after the server restores the classified unvoiced frames in the repaired audio frame sequence, the server generates a target audio file matching the phoneme timestamp based on the syllable information and the repaired audio frame sequence. Specifically, this process includes: performing data fusion processing on the syllable information and the restored audio frame sequence according to the phoneme timestamp to generate the target audio file.

[0168] In some embodiments, the server integrates the syllable information and the restored audio frame sequence into an aligned syllable information and audio frame sequence with the same vector length, according to the duration parameter and fundamental frequency value of each audio frame, the audio frame position occupied by the syllable / phoneme, and the phoneme timestamp of each phoneme. Then, a preset synthesizer is used to perform data fusion processing on the lyrics information and audio frame sequence to generate a target audio file for the input audio.

[0169] Among them, audio synthesis technology has been widely used in speech synthesis due to its advantages such as large adjustment capability and strong speech plasticity. In practice, LPC (linear predictive coding) filters can be used as synthesizers, and this application does not limit the specific synthesizer.

[0170] Since the lyrics information and audio frame sequence with the same vector length and alignment are added, the synthesized audio template has the same melody and rhythm as the input audio.

[0171] To more clearly illustrate the method for creating audio files provided in this disclosure, a specific embodiment is given below for detailed description. In an exemplary embodiment, reference is made to... Figure 13 and Figure 14 , Figure 13 This is a flowchart illustrating a method for creating an audio file according to another exemplary embodiment. Figure 14 This is a block diagram illustrating a method for creating an audio file according to another exemplary embodiment. The method is used in server 104 and specifically includes the following:

[0172] Step S21: Obtain the online song input by the user.

[0173] The online song can be a formally released music song or a live song recorded offline through a terminal (such as a mobile phone, tablet, etc.). This application does not specifically limit the source and form of the input online song.

[0174] Step S22: The server uses a preset speech recognition network to recognize the lyrics of the online song and obtains the corresponding lyric sequence.

[0175] The pre-trained speech recognition network can be a Transformer model or other speech recognition models (such as attention-based RNNs, LSTMs, etc.).

[0176] Step S23: The server separates the audio and accompaniment of the online song using an audio-accompaniment separation algorithm to obtain the dry audio of the online song.

[0177] The dry audio is composed of multiple audio frames spliced ​​together, and these audio frames include unvoiced frames and voiced frames.

[0178] Among them, the sound companion separation algorithm includes the Spleeter algorithm.

[0179] Step S24: The server extracts the fundamental frequency of the dry audio using a fundamental frequency extraction algorithm to obtain the fundamental frequency information of each frame of dry audio; and, the server performs word-by-word alignment of the lyrics and dry audio using a forced alignment algorithm to obtain the aligned dry audio.

[0180] The fundamental frequency information includes the fundamental frequency value of the dry audio in each audio frame.

[0181] Wherein, the fundamental frequency extraction algorithms include the yin algorithm, the pyin algorithm, the crepe algorithm, etc.

[0182] Wherein, the forced alignment algorithm includes a Monotonic Alignment Search (MAS) algorithm.

[0183] Wherein, the aligned dry voice audio includes: a plurality of phonemes and audio frames occupied by the plurality of phonemes correspondingly.

[0184] Wherein, phoneme (phoneme): the smallest unit of Chinese pronunciation (the smallest unit of writing is grapheme), for example, for the Chinese character "好", the pronunciation units are h and ao.

[0185] Wherein, the aligned dry voice audio may also include: a plurality of syllables and audio frames occupied by the plurality of syllables correspondingly.

[0186] Wherein, a syllable is the smallest speech unit formed by combined pronunciation of a single vowel phoneme and consonant phoneme, and a single vowel phoneme can also form a syllable independently. A Chinese syllable is a speech unit formed by combining an initial consonant and a final vowel, and a single final vowel can also form a syllable independently.

[0187] Wherein, there is a mapping relationship between each audio frame of the aligned dry voice audio and a corresponding phoneme, that is, each audio frame of the aligned dry voice audio follows the Gaussian distribution (log probability value of normal distribution) of the corresponding phoneme.

[0188] In step S25, the server draws a corresponding original fundamental frequency curve according to the fundamental frequency information of the dry voice audio, and draws a corresponding first-order difference fundamental frequency curve according to the original fundamental frequency curve.

[0189] Wherein, the horizontal axis value of the original fundamental frequency curve represents an audio frame position, and the vertical axis value of the original fundamental frequency curve represents a fundamental frequency value of the dry voice audio in the corresponding audio frame.

[0190] Wherein, a pitch value is the fundamental frequency value. When a sounding body vibrates to produce sound, the sound can generally be decomposed into many simple sine waves. That is, all natural sounds are basically composed of sine waves with different frequencies. The sine wave with the lowest frequency is the fundamental tone, while other sine waves with higher frequencies are overtones.

[0191] Wherein, the horizontal axis value of the first-order difference fundamental frequency curve represents an audio frame position, and the vertical axis value of the first-order difference fundamental frequency curve represents a difference value of the fundamental frequency value of the dry voice audio in the corresponding audio frame.

[0192] Step S26: The server identifies transition points in the first-order differential fundamental frequency curve based on the relationship between the deviation of the difference values ​​at each position point in the first-order differential fundamental frequency curve and the preset threshold; and identifies position points in the original fundamental frequency curve where the fundamental frequency value is 0.

[0193] The deviation of the difference values ​​in the first-order differential fundamental frequency curve represents the degree to which the difference values ​​at the corresponding positions deviate from the first-order differential fundamental frequency curve.

[0194] Among them, the jump point in the first-order differential fundamental frequency curve corresponds to the error detection point in the original fundamental frequency curve.

[0195] Among them, the points where the fundamental frequency value is 0 in the original fundamental frequency curve correspond to the missed detection points in the original fundamental frequency curve.

[0196] Step S27: The server repairs the fundamental frequency value of the corresponding error detection point in the original fundamental frequency curve based on the position of the jump point in the first-order differential fundamental frequency curve, and obtains the fundamental frequency curve after error detection repair.

[0197] Specifically, the fundamental frequency value adjacent to the error detection point at the corresponding position in the original fundamental frequency curve, or the fundamental frequency value adjacent to the error detection point at the corresponding position in the original fundamental frequency curve, is taken as the fundamental frequency value to be repaired.

[0198] In step S28, the server determines the position of the audio frame occupied by each phoneme in the fundamental frequency curve after error detection and repair, based on the aligned dry audio.

[0199] In step S29, the server interpolates the fundamental frequency values ​​of the missed detection points in the audio frame positions occupied by each phoneme in the fundamental frequency curve after error detection and repair according to the preset interpolation function, so as to obtain the fundamental frequency curve after error detection and interpolation.

[0200] The server converts the missing detection interpolated fundamental frequency curve into the corresponding fundamental frequency data.

[0201] The interpolation function is the preset sigmoid function:

[0202] In the sigmoid function, A is a constant, the horizontal axis value of S(x) represents the audio frame position where the missed detection point is located, and the vertical axis value of S(x) after inverse transformation represents the fundamental frequency value to be interpolated.

[0203] If there are multiple consecutive and adjacent missed detection points in the audio frame position occupied by a phoneme, the sigmoid function performs a one-time interpolation of the fundamental frequency value for these multiple missed detection points.

[0204] In this process, the fundamental frequency values ​​corresponding to both unvoiced frames and voiced frames in the fundamental frequency curve after interpolation are interpolated.

[0205] Step S30: The server inputs the aligned dry audio and fundamental frequency data into a preset convolutional network to generate mask marks for the fundamental frequency data aligned with the unvoiced frame position and mask marks for the fundamental frequency data aligned with the voiced frame position.

[0206] The convolutional network normalizes the fundamental frequency data aligned to the audio frame positions occupied by each phoneme based on the aligned dry audio. Then, it uses a preset convolutional layer to perform data augmentation on the normalized fundamental frequency data. Next, it uses a preset linear layer to perform data channel descent on the data augmented fundamental frequency data to obtain one-dimensional fundamental frequency data. Finally, it uses the sigmoid function to perform mask marking on the one-dimensional fundamental frequency data to output a first mask mark aligned to the unvoiced frame position and a second mask mark aligned to the voiced frame position.

[0207] Step S31: The server assigns a value of 0 to the first mask marker and a value of 1 to the second mask marker.

[0208] Step S32: The server multiplies the fundamental frequency data of the audio frame position occupied by each phoneme with the mask mark of the corresponding audio frame position to perform data restoration processing on the fundamental frequency data and obtain the restored fundamental frequency data.

[0209] Specifically, the fundamental frequency value of the fundamental frequency data aligned with the unvoiced frame position is restored to 0, while the fundamental frequency value of the fundamental frequency data aligned with the voiced frame position remains unchanged.

[0210] Step S33: The server merges aligned dry audio with the same vector length and the fundamental frequency data after data restoration to generate a vocal synthesis template.

[0211] In this way, on the one hand, the audio frame sequence and lyrics sequence are monotonically aligned and the audio frames are repaired to obtain the corresponding aligned syllable information and the repaired audio frame sequence. Then, the audio file is generated based on the aligned syllable information and the repaired audio frame sequence, thereby optimizing the audio file production process and reducing the consumption of manpower and time costs. On the other hand, the deviation of the fundamental frequency value of each audio frame from the corresponding reference value is used to repair the fundamental frequency value of abnormal audio frames in order to generate an audio file for the input audio. This can improve the naturalness and expressiveness of the produced audio file, thus facilitating subsequent audio processing based on the produced audio file.

[0212] It should be understood that, although Figures 2-14The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figures 2-14 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.

[0213] It is understood that the same / similar parts between the various embodiments of the methods described above in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments, and relevant parts can be referred to the description of other method embodiments.

[0214] Figure 15 This is a block diagram of an audio file creation apparatus provided in an embodiment of this application. (Refer to...) Figure 15 The audio file production device 10 includes: an information acquisition unit 11, an audio processing unit 12, a monotonic alignment unit 13, an audio repair unit 14, and a template generation unit 15.

[0215] The information acquisition unit 11 is configured to acquire input audio and the lyrics sequence corresponding to the input audio; the input audio includes dry audio for expressing the lyrics sequence, and the dry audio is composed of multiple consecutive audio frames.

[0216] The audio processing unit 12 is configured to extract an audio frame sequence about the dry audio from the input audio, and perform fundamental frequency detection on the audio frame sequence to obtain the fundamental frequency value of each audio frame.

[0217] The monotonic alignment unit 13 is configured to perform alignment processing on the audio frame sequence and the lyrics sequence to obtain syllable information monotonically aligned with the audio frame sequence, wherein the syllable information includes phoneme timestamps.

[0218] The audio repair unit 14 is configured to perform repair processing on the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, so as to obtain a repaired audio frame sequence.

[0219] The template generation unit 15 is configured to generate a target audio file matching the phoneme timestamp based on the syllable information and the repaired audio frame sequence.

[0220] In some embodiments, in order to repair the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence based on the degree of deviation of the fundamental frequency value of each audio frame relative to the corresponding reference value, and to obtain a repaired audio frame sequence, the audio repair unit 14 is further configured to:

[0221] Based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, abnormal audio frames and empty audio frames are identified in each audio frame; the fundamental frequency value of the empty audio frame does not exist;

[0222] The fundamental frequency value of the abnormal audio frame is corrected, and the fundamental frequency value of the empty audio frame is interpolated to obtain the repaired audio frame sequence.

[0223] In some embodiments, in determining empty audio frames in each audio frame based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, the audio repair unit 14 is further configured to:

[0224] Based on the fundamental frequency value of each audio frame, fundamental frequency data corresponding to each audio frame is generated;

[0225] In the baseband data, the audio frames corresponding to the parts where the baseband value does not exist are determined as the empty audio frames.

[0226] In some embodiments, in identifying the target video as the indoor scene or the outdoor scene based on the second number of target image scene categories, the audio restoration unit 14 is further configured to:

[0227] In the audio frame sequence, the corresponding logarithmic function value is generated using the fundamental frequency value of the preceding adjacent audio frame and the fundamental frequency value of the following adjacent audio frame relative to the empty audio frame.

[0228] The logarithmic function value is interpolated into the empty audio frame to interpolate the fundamental frequency value of the empty audio frame.

[0229] In some embodiments, in determining abnormal audio frames among the audio frames based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, the audio repair unit 14 is further configured to:

[0230] Based on the fundamental frequency value of each audio frame, differential data corresponding to each audio frame is generated; the horizontal axis of the differential data represents the order of each audio frame in the audio frame sequence, and the vertical axis of the differential data represents the differential value of the audio frame corresponding to the order.

[0231] In the differential data, based on the relationship between the distance of the differential value of each position from the corresponding reference value and the preset distance, abnormal audio frames are determined in each audio frame.

[0232] In some embodiments, the audio repair unit 14 is further configured to: determine abnormal audio frames in each audio frame based on the relationship between the distance of the difference value of each position deviating from the corresponding reference value and a preset distance in the differential data;

[0233] The difference value of the previous adjacent position or the difference value of the next adjacent position relative to the difference value of each position is determined as the reference value of the difference value of each position.

[0234] Among the difference values ​​of each arrangement position, abnormal difference values ​​are identified where the distance between the difference value and the corresponding reference value is greater than the preset distance, and the audio frame corresponding to the abnormal difference value is identified as the abnormal audio frame.

[0235] In some embodiments, in terms of correcting the fundamental frequency value of the abnormal audio frame, the audio repair unit 14 is further configured to:

[0236] In the audio frame sequence, the fundamental frequency value of the preceding or following adjacent audio frame in the fundamental frequency image of the audio frame sequence corresponding to the abnormal audio frame is used as the correction fundamental frequency value of the abnormal audio frame.

[0237] The fundamental frequency value of the abnormal audio frame is replaced with the correction fundamental frequency value to perform correction processing on the fundamental frequency value of the abnormal audio frame.

[0238] In some embodiments, the audio frames in the audio frame sequence include unvoiced frames and voiced frames; after repairing the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence to obtain a repaired audio frame sequence, the audio template creation apparatus 10 is further configured to:

[0239] Based on the syllable information, the unvoiced and voiced frames in the repaired audio frame sequence are classified.

[0240] Based on a preset marker mask, the unvoiced audio frames and the voiced audio frames after classification are restored to obtain a restored audio frame sequence; the restoration process is used to restore the fundamental frequency value of the repaired audio frames to the fundamental frequency value obtained by fundamental frequency detection.

[0241] In some embodiments, in generating a target audio file matching the phoneme timestamp based on the syllable information and the repaired audio frame sequence, the template generation unit 15 is further configured to:

[0242] The target audio file is generated by performing data fusion processing on the syllable information and the restored audio frame sequence according to the phoneme timestamp.

[0243] Figure 16 This is a block diagram of a server 20 provided in an embodiment of this application. For example, server 20 can be an electronic device, an electronic component, or a server array, etc. (Refer to...) Figure 16 Server 20 includes processor 21, which may be a collection of processors, including one or more processors. Server 20 also includes memory resources represented by memory 22, on which computer programs, such as application programs, are stored. The computer programs stored in memory 22 may include one or more modules, each corresponding to a set of executable instructions. Furthermore, processor 21 is configured to implement the method for creating audio files as described above when executing the computer program.

[0244] In some embodiments, server 20 is an electronic device whose computing system can run one or more operating systems, including any of the operating systems discussed above and any commercially available server operating system. Server 20 can also run any of a variety of additional server applications and / or middleware applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, super servers, database servers, etc. Exemplary database servers include, but are not limited to, commercially available database servers from companies such as IBM.

[0245] In some embodiments, processor 21 typically controls the overall operation of server 20, such as operations associated with display, data processing, data communication, and recording operations. Processor 21 may include one or more processor components to execute computer programs to perform all or part of the steps of the methods described above. Furthermore, processor components may include one or more modules to facilitate interaction between processor components and other components. For example, processor components may include a multimedia module to facilitate control of the interaction between user server 20 and processor 21 using multimedia components.

[0246] In some embodiments, the processor component in processor 21 may also be referred to as a CPU (Central Processing Unit). The processor component may be an electronic chip with signal processing capabilities. The processor may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor component. Furthermore, the processor component may be implemented using integrated circuit chips.

[0247] In some embodiments, memory 22 is configured to store various types of data to support operation on server 20. Examples of such data include instructions for any application or method operating on server 20, acquired data, messages, images, videos, etc. Memory 22 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, optical disk, or graphene storage.

[0248] In some embodiments, the memory 22 can be a memory module, TF card, etc., and can store all information in the server 20, including the input raw data, computer programs, intermediate running results, and final running results. In some embodiments, it stores and retrieves information according to the location specified by the processor. In some embodiments, the server 20 has a memory function and can ensure normal operation because of the memory 22. In some embodiments, the memory 22 of the server 20 can be classified into main memory (RAM) and auxiliary memory (external memory) according to its purpose, or it can be classified into external memory and internal memory. External memory is usually magnetic media or optical discs, which can store information for a long time. RAM refers to the storage component on the motherboard, which is used to store the currently executing data and programs, but it is only used to temporarily store programs and data. The data will be lost when the power is turned off or the power is cut off.

[0249] In some embodiments, server 20 may further include: a power supply component 23 configured to perform power management of server 20, a wired or wireless network interface 24 configured to connect server 20 to a network, and an input / output (I / O) interface 25. Server 20 may operate on an operating system stored in memory 22, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, or similar.

[0250] In some embodiments, power supply component 23 provides power to various components of server 20. Power supply component 23 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to server 20.

[0251] In some embodiments, the wired or wireless network interface 24 is configured to facilitate wired or wireless communication between the server 20 and other devices. The server 20 may access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof.

[0252] In some embodiments, the wired or wireless network interface 24 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the wired or wireless network interface 24 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0253] In some embodiments, the input / output (I / O) interface 25 provides an interface between the processor 21 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, volume buttons, a power button, and a lock button.

[0254] Figure 17 This is a block diagram of a computer-readable storage medium 30 provided in an embodiment of this application. The computer-readable storage medium 30 stores a computer program 31, which, when executed by a processor, implements the audio file creation method described above.

[0255] If the integrated units of the various functional units in the various embodiments of this application are implemented as software functional units and sold or used as independent products, they can be stored in the computer-readable storage medium 30. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer-readable storage medium 30 includes a computer program 31, which includes several instructions to cause a computer device (which may be a personal computer, system server, or network device, etc.), an electronic device (e.g., MP3, MP4, etc., or a mobile phone, tablet computer, wearable device, etc., or a desktop computer, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of this application.

[0256] Figure 18 This is a block diagram of a computer program product 40 provided in an embodiment of this application. The computer program product 40 includes program instructions 41, which can be executed by the processor of the server 20 to implement the audio file creation method described above.

[0257] Those skilled in the art will understand that embodiments of this application may provide a method for creating audio files, an apparatus 10 for creating audio files, a server 20, a computer-readable storage medium 30, or a computer program product 40. Therefore, this application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application may take the form of a computer program product 40 embodied on one or more computer program instructions 41 (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0258] This application is described with reference to flowchart illustrations and / or block diagrams of an audio file creation method, an audio file creation apparatus 10, a server 20, a computer-readable storage medium 30, or a computer program product 40 according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by the computer program product 40. These computer program products 40 can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, such that program instructions 41, executable by the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0259] These computer program products 40 may also be stored in a computer-readable storage medium capable of directing a computer or other programmable data processing device to function in a particular manner, such that program instructions 41 stored in the computer program product 40 produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0260] These program instructions 41 may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing the program instructions 41 that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0261] It should be noted that the various methods, apparatuses, electronic devices, computer-readable storage media, computer program products, etc. described above may also include other implementation methods according to the description of the method embodiments. For specific implementation methods, please refer to the description of the relevant method embodiments, which will not be elaborated here.

[0262] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0263] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for creating an audio file, characterized in that, The method includes: The input audio and the corresponding lyrics sequence are obtained; the input audio includes dry audio for expressing the lyrics sequence, and the dry audio consists of multiple consecutive audio frames; Extract an audio frame sequence of the dry audio from the input audio, and perform fundamental frequency detection on the audio frame sequence to obtain the fundamental frequency value of each audio frame; and The audio frame sequence and the lyrics sequence are aligned to obtain syllable information that is monotonically aligned with the audio frame sequence. The syllable information includes phoneme timestamps. Based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, the fundamental frequency values ​​of audio frames with abnormal deviation in the audio frame sequence are repaired to obtain a repaired audio frame sequence; the abnormal audio frames include abnormal audio frames; the method for determining abnormal audio frames includes: taking the difference value of the audio frame in the position preceding the difference value of each arranged position in the audio frame sequence as the reference value of the difference value of each arranged position; determining a first distance between the difference value of each arranged position and its corresponding reference value; identifying abnormal difference values ​​in the difference values ​​of each arranged position where the first distance is greater than a preset second distance, and determining the audio frame corresponding to the abnormal difference value as the abnormal audio frame; Based on the syllable information and the repaired audio frame sequence, a target audio file matching the phoneme timestamp is generated.

2. The method according to claim 1, characterized in that, The step of correcting the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence based on the degree of deviation of the fundamental frequency value of each audio frame relative to the corresponding reference value, to obtain a corrected audio frame sequence, includes: Based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, abnormal audio frames and empty audio frames are identified in each audio frame; the fundamental frequency value of the empty audio frame does not exist; The fundamental frequency value of the abnormal audio frame is corrected, and the fundamental frequency value of the empty audio frame is interpolated to obtain the repaired audio frame sequence.

3. The method according to claim 2, characterized in that, in, Based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, empty audio frames are determined in each audio frame, including: Based on the fundamental frequency value of each audio frame, fundamental frequency data corresponding to each audio frame is generated; In the baseband data, the audio frames corresponding to the parts where the baseband value does not exist are determined as the empty audio frames; The interpolation process for the fundamental frequency value of the empty audio frame includes: In the audio frame sequence, the corresponding logarithmic function value is generated using the fundamental frequency value of the preceding adjacent audio frame and the fundamental frequency value of the following adjacent audio frame relative to the empty audio frame. The logarithmic function value is interpolated into the empty audio frame to interpolate the fundamental frequency value of the empty audio frame.

4. The method according to claim 2, characterized in that, in, Based on the degree of deviation of the fundamental frequency value of each audio frame from the corresponding reference value, abnormal audio frames are identified in each audio frame, including: Based on the fundamental frequency value of each audio frame, differential data corresponding to each audio frame is generated; the horizontal axis of the differential data represents the order of each audio frame in the audio frame sequence, and the vertical axis of the differential data represents the differential value of the audio frame corresponding to the order. In the differential data, based on the relationship between the distance of the differential value of each position from the corresponding reference value and the preset distance, abnormal audio frames are determined in each audio frame.

5. The method according to claim 4, characterized in that, In the differential data, based on the relationship between the distance of the differential value of each position deviating from the corresponding reference value and a preset distance, abnormal audio frames are identified in each audio frame, including: The difference value of the previous adjacent position or the difference value of the next adjacent position relative to the difference value of each position is determined as the reference value of the difference value of each position. Among the difference values ​​of each arrangement position, abnormal difference values ​​are identified where the distance between the difference value and the corresponding reference value is greater than the preset distance, and the audio frame corresponding to the abnormal difference value is identified as the abnormal audio frame.

6. The method according to claim 4, characterized in that, The correction process for the fundamental frequency value of the abnormal audio frame includes... In the audio frame sequence, the fundamental frequency value of the preceding or following adjacent audio frame relative to the abnormal audio frame is used as the correction fundamental frequency value of the abnormal audio frame. The fundamental frequency value of the abnormal audio frame is replaced with the correction fundamental frequency value to perform correction processing on the fundamental frequency value of the abnormal audio frame.

7. The method according to claim 1, characterized in that, The audio frames in the audio frame sequence include unvoiced frames and voiced frames; After repairing the fundamental frequency values ​​of audio frames with abnormal deviations in the audio frame sequence to obtain the repaired audio frame sequence, the method further includes: Based on the syllable information, the unvoiced and voiced frames in the repaired audio frame sequence are classified. Based on a preset marker mask, the unvoiced frames and the voiced frames after classification are restored to obtain a restored audio frame sequence; the restoration process is used to restore the fundamental frequency value of the restored audio frame to the fundamental frequency value obtained by fundamental frequency detection. The step of generating a target audio file matching the phoneme timestamp based on the syllable information and the repaired audio frame sequence includes: The target audio file is generated by performing data fusion processing on the syllable information and the restored audio frame sequence according to the phoneme timestamp.

8. The method according to claim 7, characterized in that, The syllable information represents multiple phoneme information and the audio frames occupied by the multiple phoneme information; The process of classifying unvoiced and voiced frames in the repaired audio frame sequence based on the syllable information includes: Based on the syllable information, the phoneme information corresponding to each audio frame is determined in the repaired audio frame sequence; Based on the type of phoneme information corresponding to each audio frame, each audio frame in the repaired audio frame sequence is classified as either an unvoiced frame or a voiced frame.

9. A server, characterized in that, include: processor; Memory for storing the executable instructions of the processor; The processor is configured to execute the executable instructions to implement the method for creating an audio file as described in any one of claims 1 to 8.

10. A computer-readable storage medium comprising program data, characterized in that, When the program data is executed by the server's processor, the server is able to perform the method for creating an audio file as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Sound correction method, computer equipment and computer readable storage medium

    CN115101080A