Audio processing method, device and storage medium
By extracting the audio feature, obtaining the speech speed, fundamental frequency fluctuations and energy proportion characteristics, the problem of low accuracy in the discrimination of audio types in the prior art is solved, and a more accurate distinction between singing and speaking voices is achieved.
Patent Information
- Application Number
- CN202211250389.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-12
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-10-12
AI Technical Summary
The prior art has low accuracy when distinguishing audio types, making it difficult to distinguish between singing and speaking.
By extracting the target audio feature, the speech speed feature, fundamental frequency fluctuation feature and energy proportion feature are obtained, and the audio type is determined based on the comparison results of these features with the preset characteristic value range.
Improve the accuracy of audio type discrimination, allowing for more accurate distinction between singing and speaking voices.
Smart Images

Figure CN115641874B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an audio processing method, device and storage medium. Background Art
[0002] In the relevant scenarios of audio processing, since there are different processing requirements for singing type audio and speaking type audio, for example, it is necessary to enhance the speaking type audio to make the speaking voice more prominent, and it is necessary to correct the sound of the singing type audio to make the singing voice more beautiful and pleasant; based on this, it is very important to accurately distinguish whether the audio is of singing type or speaking type, but in the existing distinction schemes, the distinction is usually made based on the bandwidth of the audio, that is, since the frequency of speaking is generally between 300 and 3400 Hz, and the frequency of singing covers the 20 to 20 kHz that can be heard by the human ear, the audio can be distinguished as a speaking type when the frequency of the audio is between 300 and 3400 Hz, and the audio can be distinguished as a singing type when the frequency of the audio is between 20 and 20 kHz, and the audio can be distinguished as a singing type when the frequency of the audio is between 20 and 20 kHz, and the distinction accuracy is low. Summary of the invention
[0003] The embodiments of the present application provide an audio processing method, apparatus, device, storage medium and computer program product, which can accurately determine the audio type of audio.
[0004] On the one hand, an embodiment of the present application provides an audio processing method, including:
[0005] Performing feature extraction processing on the target audio to obtain a type discrimination feature of the target audio; the type discrimination feature includes at least any one of the following: a speech rate feature, a fundamental frequency fluctuation feature, and an energy proportion feature, the fundamental frequency fluctuation feature is used to indicate the fluctuation of the fundamental frequency of the target audio within a continuous target time period, and the energy proportion feature is used to indicate the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the full frequency band;
[0006] The audio type of the target audio is determined based on a comparison result between the type discrimination feature of the target audio and a preset feature value range corresponding to the type discrimination feature; the audio type is a singing type or a speaking type.
[0007] On the one hand, an embodiment of the present application provides an audio processing device, including:
[0008] An extraction unit is used to perform feature extraction processing on the target audio to obtain a type discrimination feature of the target audio; the type discrimination feature includes at least any one of the following: a speech rate feature, a fundamental frequency fluctuation feature, and an energy proportion feature, the fundamental frequency fluctuation feature is used to indicate the fluctuation of the fundamental frequency of the target audio within a continuous target time period, and the energy proportion feature is used to indicate the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the full frequency band;
[0009] A discrimination unit is used to determine the audio type of the target audio according to a comparison result of the type discrimination feature of the target audio and a preset feature value range corresponding to the type discrimination feature; the audio type is a singing type or a speaking type.
[0010] On the one hand, an embodiment of the present application provides an audio processing device, characterized in that the audio processing device includes an input interface and an output interface, and further includes:
[0011] a processor adapted to implement one or more instructions; and,
[0012] A computer storage medium storing one or more instructions, wherein the one or more instructions are suitable for being loaded by the processor and executing the above-mentioned audio processing method.
[0013] On the one hand, an embodiment of the present application provides a computer storage medium, characterized in that computer program instructions are stored in the computer storage medium, and when the computer program instructions are executed by a processor, they are used to execute the above-mentioned audio processing method.
[0014] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program stored in a computer storage medium; a processor of an audio processing device reads the computer program from the computer storage medium, and the processor executes the computer program, so that the audio processing device performs the above-mentioned audio processing method.
[0015] In an embodiment of the present application, the audio type of the target audio can be discriminated based on one or more type discrimination features of the target audio's speech rate feature, the target audio's fundamental frequency fluctuation feature, and the target audio's energy ratio feature. The target audio's speech rate feature can be used to indicate the target audio's speech rate. When the target audio's audio type is discriminated based on the target audio's speech rate feature, the audio type can be discriminated based on the speed of speech caused by the physical vocalization difference between speaking and singing, and the audio type of the audio can be accurately discriminated based on the difference in speech rate. The fundamental frequency fluctuation feature of the target audio can be used to indicate the fluctuation of the fundamental frequency of the target audio within a continuous target time period. Since the fundamental frequency determines the tone of the sound, when the target audio's audio type is discriminated based on the fundamental frequency fluctuation feature of the target audio, the audio type can be discriminated based on the fluctuation of the tone caused by the physical vocalization difference between speaking and singing, and the audio type of the audio can be accurately discriminated based on the difference in tone fluctuation size. The energy proportion feature of the target audio can be used to indicate the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the entire frequency band. When the preset frequency band is set to a speaking frequency band used to indicate the frequency of the speaking voice, when the audio type of the target audio is determined based on the energy proportion feature of the target audio, the audio type can be determined based on the significant energy of the speaking frequency band caused by the physical vocalization difference between speaking and singing. The audio type of the audio can be accurately determined based on the significant energy of the frequency band. Moreover, when the audio type of the target audio is determined based on multiple types of distinguishing features such as the speech rate feature of the target audio, the fundamental frequency fluctuation feature of the target audio, and the energy proportion feature of the target audio, multiple types of distinguishing features can be fully utilized to further improve the accuracy of audio type determination. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0017] Figure 1 It is a flow chart for distinguishing the audio type of audio provided by an embodiment of the present application;
[0018] Figure 2 is another flow chart for determining the audio type of audio provided by an embodiment of the present application;
[0019] Figure 3It is a flowchart of an audio processing method provided in an embodiment of the present application;
[0020] Figure 4 is a flowchart of another audio processing method provided in an embodiment of the present application;
[0021] Figure 5a is a schematic diagram of the result of performing speech recognition processing on target audio provided by an embodiment of the present application;
[0022] Figure 5b is a schematic diagram of the result of performing speech recognition processing on target audio provided by an embodiment of the present application;
[0023] Figure 6 is a flowchart of another audio processing method provided in an embodiment of the present application;
[0024] Figure 7a is a schematic diagram of a fundamental frequency of a target audio obtained by extraction provided in an embodiment of the present application;
[0025] Figure 7b is another schematic diagram of the fundamental frequency of the target audio extracted according to an embodiment of the present application;
[0026] Figure 8a is a schematic diagram of a note mapping result corresponding to a valid sampling point provided in an embodiment of the present application;
[0027] Figure 8b is a schematic diagram of another note mapping result corresponding to a valid sampling point provided in an embodiment of the present application;
[0028] Fig. 9 is a flowchart of another audio processing method provided in an embodiment of the present application;
[0029] Fig.10 is a comparative schematic diagram of an average power spectrum provided in an embodiment of the present application;
[0030] Fig.11 is a flowchart of another audio processing method provided in an embodiment of the present application;
[0031] Fig.12 is another flow chart for determining the audio type of audio provided by an embodiment of the present application;
[0032] Fig.13 is a structural diagram of an audio processing device provided in an embodiment of the present application;
[0033] Fig.14 It is a structural diagram of the audio processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0034] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0035] An embodiment of the present application provides an audio processing solution, which can perform feature extraction processing on a target audio to obtain a type discrimination feature of the target audio; and then determine whether the audio type of the target audio is a singing type or a speaking type based on a comparison result of the type discrimination feature of the target audio and a preset feature value range corresponding to the type discrimination feature; wherein the type discrimination feature of the target audio is used to discriminate whether the audio type of the target audio is a singing type or a speaking type, and the type discrimination feature includes at least any one of the following: a speech rate feature, a fundamental frequency fluctuation feature, and an energy proportion feature, the fundamental frequency fluctuation feature of the target audio is used to indicate: the fluctuation of the fundamental frequency of the target audio within a continuous target time period, and the energy proportion feature of the target audio is used to indicate: the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the entire frequency band.
[0036] In a specific implementation, the audio processing scheme proposed in the embodiment of the present application can be executed by an audio processing device, which can be a terminal device or a server; the terminal device here can include but is not limited to: computers, smart phones, tablet computers, laptops, smart home appliances, vehicle terminals, smart wearable devices, etc.; the server here can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. Further optionally, the audio processing scheme proposed in the embodiment of the present application can also be executed individually or collaboratively by other electronic devices with computing power, which is not limited by the embodiment of the present application.
[0037] In one embodiment, the audio processing device determines the audio type of the target audio based on the comparison result of the type discrimination feature of the target audio and the preset feature value range corresponding to the type discrimination feature, which may include: if the type discrimination feature of the target audio is one, and the type discrimination feature of the target audio meets the preset feature value range corresponding to the corresponding type discrimination feature, the audio type of the target audio is judged as a singing type; if the type discrimination feature of the target audio is one, and the type discrimination feature of the target audio does not meet the preset feature value range corresponding to the corresponding type discrimination feature, the audio type of the target audio is judged as a speaking type; if the type discrimination feature of the target audio is multiple, and among the multiple type discrimination features of the target audio, there is at least one type discrimination feature that meets the preset feature value range corresponding to the at least one type discrimination feature, the audio type of the target audio is judged as a singing type; if the type discrimination feature of the target audio is multiple, and among the multiple type discrimination features of the target audio, each type discrimination feature does not meet the preset feature value range corresponding to each type discrimination feature, the audio type of the target audio is judged as a speaking type.
[0038] In one embodiment, when the type discrimination feature of the target audio is one, see Figure 1 , which is a discrimination flow chart for discriminating the audio type of audio provided in an embodiment of the present application; its steps may include: ① performing feature extraction processing on the target audio to obtain the type discrimination feature of the target audio; ② judging whether the type discrimination feature of the target audio does not conform to the preset feature value range corresponding to the corresponding type discrimination feature, that is, judging whether the type discrimination feature of the target audio is outside the preset feature value range corresponding to the corresponding type discrimination feature; ③ if the type discrimination feature of the target audio does not conform to the preset feature value range corresponding to the corresponding type discrimination feature, that is, the type discrimination feature of the target audio is outside the preset feature value range corresponding to the corresponding type discrimination feature, then the audio type of the target audio is discriminated as a speaking voice type; ④ otherwise, the audio type of the target audio is discriminated as a singing voice type.
[0039] In one embodiment, when there are multiple types of target audio discrimination features, see Figure 2, which is another discrimination flow chart for discriminating the audio type of audio provided in an embodiment of the present application; its steps may include: ① performing feature extraction processing on the target audio to obtain multiple type discrimination features of the target audio; ② judging whether each type discrimination feature among the multiple type discrimination features of the target audio does not conform to the preset feature value range corresponding to each type discrimination feature, that is, judging whether each type discrimination feature among the multiple type discrimination features of the target audio is outside the preset feature value range corresponding to each type discrimination feature; ③ if each type discrimination feature among the multiple type discrimination features of the target audio does not conform to the preset feature value range corresponding to each type discrimination feature, that is, each type discrimination feature among the multiple type discrimination features of the target audio is outside the preset feature value range corresponding to each type discrimination feature, then the audio type of the target audio is discriminated as a speaking voice type; ④ otherwise, the audio type of the target audio is discriminated as a singing voice type.
[0040] It is particularly important to note that in the specific implementation of this application, object-related data is involved, such as the target audio is the audio recorded by the object, etc. When the embodiment of this application is applied to a specific product or technology, it is necessary to obtain the object's permission or consent, and the collection, use and processing of relevant data must comply with local laws, regulations and standards.
[0041] Based on the above audio processing solution, the present application embodiment provides an audio processing method. Figure 3 A flowchart of an audio processing method provided in an embodiment of the present application. Figure 3 The audio processing method shown can be executed by an audio processing device, and can also be executed individually or collaboratively by other electronic devices with computing power. The embodiment of the present application takes an audio processing device as an example. Figure 3 The audio processing method shown may include the following steps:
[0042] S301, performing feature extraction processing on the target audio to obtain type discrimination features of the target audio.
[0043] In one embodiment, the target audio may be any original audio to be subjected to audio type discrimination, or may be audio after the original audio has been subjected to processes such as speech enhancement and noise reduction. When the original audio is subjected to speech enhancement and noise reduction, it may be implemented through existing tools, for example, signal processing modules in tools such as webrtc (Web Real-Time Communication) and speex, and lightweight NN noise reduction tools such as rnnnoise may be used. After the original audio has been subjected to speech enhancement and noise reduction, it is beneficial to improve the robustness of the discrimination effect.
[0044] The type discrimination feature of the target audio is used to discriminate whether the audio type of the target audio is a singing type or a speaking type, and the type discrimination feature includes at least any one of the following: a speech rate feature, a fundamental frequency fluctuation feature, and an energy proportion feature. Among them, the speech rate feature of the target audio can be used to indicate: the speaking speed of the target audio. The fundamental frequency fluctuation feature of the target audio can be used to indicate: the fluctuation of the fundamental frequency of the target audio in the continuous target time period; since the fundamental frequency is the lowest frequency in the audio, the fundamental frequency determines the pitch of the sound, so the fundamental frequency fluctuation feature of the target audio can also be used to indicate: the fluctuation of the pitch of the target audio. The energy proportion feature of the target audio is used to indicate: the energy of the target audio in the preset frequency band, and the difference between the energy of the target audio in the full frequency band; wherein, the preset frequency band can be set according to specific needs, for example, the preset frequency band can be set as: the speaking frequency band used to indicate the frequency of the speaking voice, since the frequency of the speaking voice is generally between 300 and 3400 Hz, the speaking frequency band can be a frequency band between 300 and 3400 Hz.
[0045] S302, determining the audio type of the target audio according to the comparison result between the type discrimination feature of the target audio and the preset feature value range corresponding to the type discrimination feature; the audio type is a singing type or a speaking type.
[0046] In one embodiment, the audio processing device determines the audio type of the target audio based on the comparison result of the type discrimination feature of the target audio and the preset feature value range corresponding to the type discrimination feature, which may include: if the type discrimination feature of the target audio is one, and the type discrimination feature of the target audio meets the preset feature value range corresponding to the corresponding type discrimination feature, the audio type of the target audio is judged as a singing type; if the type discrimination feature of the target audio is one, and the type discrimination feature of the target audio does not meet the preset feature value range corresponding to the corresponding type discrimination feature, the audio type of the target audio is judged as a speaking type; if the type discrimination feature of the target audio is multiple, and among the multiple type discrimination features of the target audio, there is at least one type discrimination feature that meets the preset feature value range corresponding to the at least one type discrimination feature, the audio type of the target audio is judged as a singing type; if the type discrimination feature of the target audio is multiple, and among the multiple type discrimination features of the target audio, each type discrimination feature does not meet the preset feature value range corresponding to each type discrimination feature, the audio type of the target audio is judged as a speaking type.
[0047] In an embodiment of the present application, the audio type of the target audio can be discriminated based on one or more type discrimination features of the target audio's speech rate feature, the target audio's fundamental frequency fluctuation feature, and the target audio's energy ratio feature. The target audio's speech rate feature can be used to indicate the target audio's speech rate. When the target audio's audio type is discriminated based on the target audio's speech rate feature, the audio type can be discriminated based on the speed of speech caused by the physical vocalization difference between speaking and singing, and the audio type of the audio can be accurately discriminated based on the difference in speech rate. The fundamental frequency fluctuation feature of the target audio can be used to indicate the fluctuation of the fundamental frequency of the target audio within a continuous target time period. Since the fundamental frequency determines the tone of the sound, when the target audio's audio type is discriminated based on the fundamental frequency fluctuation feature of the target audio, the audio type can be discriminated based on the fluctuation of the tone caused by the physical vocalization difference between speaking and singing, and the audio type of the audio can be accurately discriminated based on the difference in tone fluctuation size. The energy proportion feature of the target audio can be used to indicate the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the entire frequency band. When the preset frequency band is set to a speaking frequency band used to indicate the frequency of the speaking voice, when the audio type of the target audio is determined based on the energy proportion feature of the target audio, the audio type can be determined based on the significant energy of the speaking frequency band caused by the physical vocalization difference between speaking and singing. The audio type of the audio can be accurately determined based on the significant energy of the frequency band. Moreover, when the audio type of the target audio is determined based on multiple types of distinguishing features such as the speech rate feature of the target audio, the fundamental frequency fluctuation feature of the target audio, and the energy proportion feature of the target audio, multiple types of distinguishing features can be fully utilized to further improve the accuracy of audio type determination.
[0048] Based on the above audio processing solution, the embodiment of the present application provides another audio processing method, which is described by taking the type discrimination feature of the target audio as one type, and the type discrimination feature of the target audio includes the speech speed feature of the target audio. Figure 4 A flowchart of another audio processing method provided in an embodiment of the present application. Figure 4 The audio processing method shown can be executed by an audio processing device, and can also be executed individually or collaboratively by other electronic devices with computing power. The embodiment of the present application takes an audio processing device as an example. Figure 4 The audio processing method shown may include the following steps:
[0049] S401, performing feature extraction processing on the target audio to obtain type discrimination features of the target audio; the type discrimination features include speech speed features.
[0050] In one embodiment, the audio processing device performs feature extraction processing on the target audio to obtain the speech speed characteristics of the target audio, which may include: performing speech recognition processing on the target audio to obtain the text content corresponding to the target audio and the utterance start time and utterance end time of each text content in the target audio; determining the average utterance duration of each text content based on the utterance start time and utterance end time of each text content; determining the speech speed characteristics of the target audio based on the average utterance duration, wherein the speech speed characteristics of the target audio are negatively correlated with the average utterance duration.
[0051] In a specific implementation, the audio processing device performs speech recognition processing on the target audio to obtain the text content corresponding to the target audio and the start time and end time of each text content in the target audio. This can be achieved by calling an existing open source speech recognition (Automatic Speech Recognition, ASR) tool, which is not limited in the embodiment of the present application. Further, the speech rate feature of the target audio can be characterized by the number of text contents included in a unit time. At this time, the speech rate feature of the target audio can be shown by the following formula 1:
[0052]
[0053] Among them, wpm can represent the amount of text content included in a unit of time, that is, the amount of text content included in 1 minute, that is, the speaking speed feature of the target audio; Tw can represent the average utterance time.
[0054] For example, see Figure 5a , is a schematic diagram of the result of speech recognition processing on a target audio provided by an embodiment of the present application, wherein the target audio is an audio 1 of a speaking voice type, and the text content corresponding to the obtained target audio is: five, star, red, flag, welcome, wind, piao, and yang, which are expressed in pinyin as: wu, xing, hong, qi, ying, feng, piao, and yang. For example, the utterance start time of the text content "five" can be shown as marked by 501, and the utterance end time can be shown as marked by 502; see Table 1, which shows the text contents corresponding to the target audio. The utterance start time and utterance end time, as well as the utterance duration and average utterance duration corresponding to each text content determined based on the utterance start time and utterance end time of each text content; wherein, the utterance duration corresponding to any text content is: the utterance end time of the text content, minus the utterance start time of the text content; the average utterance duration of each text content is obtained by averaging the utterance duration corresponding to each text content; further, based on the average utterance duration, the speech rate feature of the target audio is determined to be approximately 258.34 (60 ÷ 0.23225).
[0055] Table 1
[0056]
[0057] For example, see Figure 5b , which is another schematic diagram of the result of speech recognition processing on the target audio provided by an embodiment of the present application, wherein the target audio is a singing type audio 2, and the text content corresponding to the obtained target audio is: five, star, red, flag, welcome, wind, piao, and yang, which are expressed in pinyin as: wu, xing, hong, qi, ying, feng, piao, and yang. For example, the utterance start time of the text content "five" can be shown as the 511 mark, and the utterance end time can be shown as the 512 mark; see Table 2, which shows the text contents corresponding to the target audio The utterance start time and utterance end time of each text content, as well as the utterance duration and the average utterance duration corresponding to each text content are determined based on the utterance start time and utterance end time of each text content; wherein, the utterance duration corresponding to any text content is: the utterance end time of the text content, minus the utterance start time of the text content; the average utterance duration of each text content is obtained by averaging the utterance duration corresponding to each text content; further, based on the average utterance duration, the speech speed feature of the target audio is determined to be approximately 93.60 (60÷0.641).
[0058] Table 2
[0059]
[0060] S402: If the speech rate feature of the target audio meets the preset feature value range corresponding to the speech rate feature, the audio type of the target audio is determined to be a singing type.
[0061] S403: If the speech rate feature of the target audio does not conform to the preset feature value range corresponding to the speech rate feature, the audio type of the target audio is determined to be a speaking voice type.
[0062] In step S402 to step S403, the preset characteristic value range corresponding to the speech rate feature can be a range less than the first speech rate threshold or greater than the second speech rate threshold, wherein the first speech rate threshold is less than the second speech rate threshold; the first speech rate threshold and the second speech rate threshold can be set according to specific needs. In a specific implementation, if the speech rate feature of the target audio meets the preset characteristic value range corresponding to the speech rate feature, the audio type of the target audio is judged as a singing type, which may include: if the speech rate feature of the target audio is within the preset characteristic value range corresponding to the speech rate feature, the audio type of the target audio is judged as a singing type; that is, if the speech rate feature of the target audio is less than the first speech rate threshold or greater than the second speech rate threshold, the audio type of the target audio is judged as a singing type.
[0063] Furthermore, if the speech rate feature of the target audio does not conform to the preset feature value range corresponding to the speech rate feature, the audio type of the target audio is judged as a speaking voice type, which may include: if the speech rate feature of the target audio is outside the preset feature value range corresponding to the speech rate feature, the audio type of the target audio is judged as a speaking voice type; that is, if the speech rate feature of the target audio is greater than or equal to the first speech rate threshold, and less than or equal to the second speech rate threshold, the audio type of the target audio is judged as a speaking voice type.
[0064] For example, the number of text contents spoken within 1 minute is usually between 150 and 260 words, that is, [150, 260]. Therefore, the first speaking rate threshold can be set to 150 and the second speaking rate threshold can be set to 260; then, the audio type of audio 1 corresponding to Table 1 is judged as a speaking type, and the audio type of audio 2 corresponding to Table 2 is judged as a singing type.
[0065] In another feasible implementation, the speaking speed feature of the target audio can also be represented by the average utterance duration; then at this time, the preset characteristic value range corresponding to the speaking speed feature can be a range greater than the first utterance duration threshold or less than the second utterance duration threshold, wherein the first utterance duration threshold is greater than the second utterance duration threshold; the first utterance duration threshold and the second utterance duration threshold can be set according to specific needs; further, the first utterance duration threshold is inversely proportional to the first speaking speed threshold, and the second utterance duration threshold is inversely proportional to the second speaking speed threshold.
[0066] In a specific implementation, if the speech speed feature of the target audio meets the preset feature value range corresponding to the speech speed feature, the audio type of the target audio is judged as a singing type, which may include: if the speech speed feature of the target audio is within the preset feature value range corresponding to the speech speed feature, the audio type of the target audio is judged as a singing type; that is, if the speech speed feature of the target audio is greater than the first utterance duration threshold or less than the second utterance duration threshold, the audio type of the target audio is judged as a singing type.
[0067] Furthermore, if the speech rate feature of the target audio does not conform to the preset feature value range corresponding to the speech rate feature, the audio type of the target audio is judged as a speaking voice type, which may include: if the speech rate feature of the target audio is outside the preset feature value range corresponding to the speech rate feature, the audio type of the target audio is judged as a speaking voice type; that is, if the speech rate feature of the target audio is less than or equal to the first utterance duration threshold, and greater than or equal to the second utterance duration threshold, the audio type of the target audio is judged as a speaking voice type.
[0068] In the embodiment of the present application, the target audio can be subjected to feature extraction processing to obtain the speech speed feature of the target audio; and then, when the speech speed feature of the target audio meets the preset feature value range corresponding to the speech speed feature, the audio type of the target audio is judged as a singing type; when the speech speed feature of the target audio does not meet the preset feature value range corresponding to the speech speed feature, the audio type of the target audio is judged as a speaking type; that is, when the speech speed feature of the target audio is less than the first speech speed threshold or greater than the second speech speed threshold, the audio type of the target audio is judged as a singing type; when the speech speed feature of the target audio is greater than or equal to the first speech speed threshold and less than or equal to the second speech speed threshold, the audio type of the target audio is judged as a speaking type; the audio type of the target audio can be judged according to the speed of the target audio; when the speech speed is too fast or too slow, the audio type is judged as a singing type, otherwise the audio type is judged as a speaking type, that is, the audio type can be judged according to the speed caused by the physical vocalization difference when speaking and singing, and the audio type of the audio can be accurately judged.
[0069] Based on the above audio processing solution, the embodiment of the present application provides another audio processing method, which is described by taking the type discrimination feature of the target audio as one, and the type discrimination feature of the target audio includes the fundamental frequency fluctuation feature of the target audio. Figure 6 A flowchart of another audio processing method provided in an embodiment of the present application. Figure 6 The audio processing method shown can be executed by an audio processing device, and can also be executed individually or collaboratively by other electronic devices with computing power. The embodiment of the present application takes an audio processing device as an example. Figure 6 The audio processing method shown may include the following steps:
[0070] S601, performing feature extraction processing on the target audio to obtain type discrimination features of the target audio; the type discrimination features include fundamental frequency fluctuation features.
[0071] In one embodiment, the audio processing device performs feature extraction processing on the target audio to obtain the fundamental frequency fluctuation characteristics of the target audio, which may include: extracting the fundamental frequency of the target audio; sampling the fundamental frequency of the target audio within a target time period to obtain the fundamental frequency corresponding to each sampling point; and determining the fundamental frequency fluctuation characteristics of the target audio based on the difference between the fundamental frequencies corresponding to each sampling point.
[0072] In the specific implementation, when the audio processing device extracts the fundamental frequency of the target audio, it can call existing open source tools to implement it, for example, it can call pYin, crepe, harvest and other open source tools to implement it; see Figure 7a, is a schematic diagram of a fundamental frequency of a target audio extracted according to an embodiment of the present application, wherein the target audio is the audio 1 corresponding to Table 1, see Figure 7b , which is another schematic diagram of the fundamental frequency of the target audio extracted provided in the embodiment of the present application, and the target audio is the audio 2 corresponding to Table 2. Furthermore, the target time period is the time period during which the target audio is played, and the corresponding duration is the duration of the target audio; the sampling interval used by the audio processing device to sample the fundamental frequency of the target audio within the target time period can be set according to specific needs, for example, the sampling interval can be set to 5 milliseconds (ms), 10 milliseconds, etc., and the embodiments of the present application are described later using the sampling interval of 5 milliseconds as an example.
[0073] Further, the audio processing device determines the fundamental frequency fluctuation characteristics of the target audio based on the difference between the fundamental frequencies corresponding to the respective sampling points, which may include: performing standard deviation calculation processing on the fundamental frequencies corresponding to the respective sampling points to obtain a target standard deviation; and determining the target standard deviation as the fundamental frequency fluctuation characteristics of the target audio. At this time, the fundamental frequency fluctuation characteristics of the target audio may be shown by the following formula 2.1:
[0074]
[0075] Wherein, N represents the number of sampling points obtained by sampling the fundamental frequency of the target audio within the target time period; n is an independent variable, [voiced] represents a set of sampling points composed of the obtained sampling points, the number of which is N; f(n) represents the fundamental frequency corresponding to the nth sampling point; It represents the mean value of the fundamental frequency corresponding to each sampling point, which can be specifically shown by the following formula 2.2:
[0076]
[0077] In another feasible implementation, the audio processing device performs standard deviation calculation processing on the fundamental frequency corresponding to each sampling point to obtain a target standard deviation, which may include: performing note mapping processing on the fundamental frequency corresponding to each sampling point to obtain a note mapping result corresponding to each sampling point; performing standard deviation calculation processing on the note mapping result corresponding to each sampling point to obtain a target standard deviation; and then the target standard deviation can be determined as the fundamental frequency fluctuation characteristic of the target audio.
[0078] In a specific implementation, the audio processing device performs note mapping processing on the fundamental frequency corresponding to the sampling point to obtain the note mapping result corresponding to the sampling point, which can be shown by the following formula 3.1:
[0079] c(n)=12·log2(f(n) / 440)+69 (3.1)
[0080] Among them, f(n) represents the fundamental frequency corresponding to the nth sampling point, and c(n) represents the note mapping result corresponding to the nth sampling point.
[0081] Furthermore, the audio processing device may perform standard deviation calculation processing on the note mapping results corresponding to each sampling point to obtain a target standard deviation; the target standard deviation is determined as the fundamental frequency fluctuation characteristic of the target audio. At this time, the fundamental frequency fluctuation characteristic of the target audio may be shown by the following formula 3.2.1:
[0082]
[0083] Wherein, N represents the number of sampling points obtained by sampling the fundamental frequency of the target audio within the target time period; n is an independent variable, and [voiced] represents a sampling point set composed of the obtained sampling points, the number of which is N; It represents the mean value of the note mapping results corresponding to each sampling point, which can be specifically shown by the following formula 3.2.2:
[0084]
[0085] In another feasible implementation, when the audio processing device performs note mapping processing on the fundamental frequency corresponding to each sampling point, the fundamental frequency corresponding to each sampling point can be mapped to a reference note, which can be 88 piano notes and can be represented by numbers 1 to 88; based on this, after the audio processing device obtains the note mapping results corresponding to each sampling point, it can also perform standard deviation calculation processing on the note mapping results corresponding to the valid sampling points in the note mapping results corresponding to each sampling point to obtain the target standard deviation, and determine the target standard deviation as the fundamental frequency fluctuation characteristic of the target audio. Among them, the valid sampling point refers to the sampling point whose note mapping result is the reference note, that is, the sampling point whose note mapping result belongs to [1, 88], and the fundamental frequency of the target audio appears as a voiced sound at the valid sampling point; see Figure 8a , is a schematic diagram of a note mapping result corresponding to a valid sampling point provided in an embodiment of the present application, the target audio indicated by the schematic diagram of the note mapping result corresponding to the valid sampling point is the audio 1 corresponding to Table 1, wherein, as shown in the mark 801, the note mapping result corresponding to the valid sampling point in the target audio is displayed, and as shown in the mark 802, the note mapping result corresponding to the valid sampling point in the voiced sound part of the target audio is displayed. By comparing 801 and 802, it can be seen that the note mapping result corresponding to the valid sampling point in the target audio is the same as the note mapping result corresponding to the valid sampling point in the voiced sound part of the target audio. It can also be seen that the fundamental frequency of the target audio is expressed as voiced sound at the valid sampling point; see Figure 8b, which is a schematic diagram of another note mapping result corresponding to a valid sampling point provided in an embodiment of the present application, wherein the target audio indicated by the schematic diagram of the note mapping result corresponding to the valid sampling point is the audio 2 corresponding to Table 2, wherein the note mapping result corresponding to the valid sampling point in the target audio is shown as indicated by the mark 811, and the note mapping result corresponding to the valid sampling point in the voiced part of the target audio is shown as indicated by the mark 812.
[0086] Furthermore, the fundamental frequency fluctuation characteristics of the target audio can be expressed by the following formula 3.3.1:
[0087]
[0088] Wherein, N′ represents the number of valid sampling points in each sampling point, that is, the number of valid sampling points in N sampling points; n is an independent variable, [voiced′] represents the valid sampling point set composed of each valid sampling point obtained, the number is N′; It represents the mean value of the note mapping results corresponding to each valid sampling point, which can be specifically shown by the following formula 3.3.2:
[0089]
[0090] S602: If the fundamental frequency fluctuation feature of the target audio meets the preset feature value range corresponding to the fundamental frequency fluctuation feature, the audio type of the target audio is determined to be a singing type.
[0091] S603: If the fundamental frequency fluctuation feature of the target audio does not conform to the preset feature value range corresponding to the fundamental frequency fluctuation feature, the audio type of the target audio is determined to be a speech type.
[0092] In step S602 to step S603, the preset characteristic value range corresponding to the fundamental frequency fluctuation feature is a range greater than the fundamental frequency fluctuation threshold; the fundamental frequency fluctuation threshold can be set according to specific needs. In a specific implementation, if the fundamental frequency fluctuation feature of the target audio meets the preset characteristic value range corresponding to the fundamental frequency fluctuation feature, the audio type of the target audio is judged as a singing type, which may include: if the fundamental frequency fluctuation feature of the target audio is within the preset characteristic value range corresponding to the fundamental frequency fluctuation feature, the audio type of the target audio is judged as a singing type; that is, if the fundamental frequency fluctuation feature of the target audio is greater than the fundamental frequency fluctuation threshold, the audio type of the target audio is judged as a singing type.
[0093] Furthermore, if the fundamental frequency fluctuation characteristics of the target audio do not conform to the preset characteristic value range corresponding to the fundamental frequency fluctuation characteristics, the audio type of the target audio is judged as a speaking voice type, which may include: if the fundamental frequency fluctuation characteristics of the target audio are outside the preset characteristic value range corresponding to the fundamental frequency fluctuation characteristics, the audio type of the target audio is judged as a speaking voice type; that is, if the fundamental frequency fluctuation characteristics of the target audio are less than or equal to the fundamental frequency fluctuation threshold, the audio type of the target audio is judged as a speaking voice type.
[0094] In a feasible implementation, the fundamental frequency fluctuation feature experience value of the singing type audio can be extracted from the fundamental frequency fluctuation features of a large number of singing type audio, and the fundamental frequency fluctuation feature experience value of the speaking type audio can be extracted from the fundamental frequency fluctuation features of a large number of speaking type audio, and then the fundamental frequency fluctuation threshold value can be determined based on the fundamental frequency fluctuation feature experience value of the singing type audio and the fundamental frequency fluctuation feature experience value of the speaking type audio, so that the determined fundamental frequency fluctuation threshold value is greater than the fundamental frequency fluctuation feature experience value of the speaking type audio and is less than the fundamental frequency fluctuation feature experience value of the singing type audio; for example, if the extracted fundamental frequency fluctuation feature experience value of the speaking type audio is 1.38 (expressed as: δ speech =1.38), the extracted fundamental frequency fluctuation characteristic value of the singing type audio is 3.33 (expressed as: δ sing =3.33), then optionally, the determined fundamental frequency fluctuation threshold may be 3 (expressed as: δ thr =3).
[0095] In the embodiment of the present application, feature extraction processing can be performed on the target audio to obtain the fundamental frequency fluctuation characteristics of the target audio; then, when the fundamental frequency fluctuation characteristics of the target audio meet the preset characteristic value range corresponding to the fundamental frequency fluctuation characteristics, the audio type of the target audio is judged as a singing type; when the fundamental frequency fluctuation characteristics of the target audio do not meet the preset characteristic value range corresponding to the fundamental frequency fluctuation characteristics, the audio type of the target audio is judged as a speaking type; that is, when the fundamental frequency fluctuation characteristics of the target audio are greater than the fundamental frequency fluctuation threshold, the audio type of the target audio is judged as a singing type; when the fundamental frequency fluctuation characteristics of the target audio are less than or equal to the fundamental frequency fluctuation threshold, the audio type of the target audio is judged as a singing type. When the frequency fluctuation threshold is exceeded, the audio type of the target audio is judged as a speaking type; the audio type of the target audio can be judged according to the fluctuation of the fundamental frequency of the target audio within the continuous target time period. Since the fundamental frequency determines the pitch of the sound, in other words, the audio type of the target audio can be judged according to the fluctuation of the pitch of the target audio; when the fundamental frequency fluctuates violently, the audio type is judged as a singing type, otherwise the audio type is judged as a speaking type, that is, the audio type can be judged according to the fluctuation of the pitch caused by the physical vocalization difference between speaking and singing, and the audio type of the audio can be accurately judged.
[0096] Based on the above audio processing solution, the embodiment of the present application provides another audio processing method, which is described by taking the type discrimination feature of the target audio as one, and the type discrimination feature of the target audio includes the energy proportion feature of the target audio. Fig. 9 A flowchart of another audio processing method provided in an embodiment of the present application. Fig. 9 The audio processing method shown can be executed by an audio processing device, and can also be executed individually or collaboratively by other electronic devices with computing power. The embodiment of the present application takes an audio processing device as an example. Fig. 9 The audio processing method shown may include the following steps:
[0097] S901, performing feature extraction processing on the target audio to obtain type discrimination features of the target audio; the type discrimination features include energy proportion features.
[0098] In one embodiment, the energy proportion feature of the target audio is used to indicate: the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the full frequency band; wherein the preset frequency band can be set according to specific needs, for example, the preset frequency band can be set to: a speaking frequency band used to indicate the frequency of the speaking sound, since the frequency of the speaking sound is generally between 300 and 3400 Hz, the speaking frequency band can be a frequency band between 300 and 3400 Hz; for example, the preset frequency band can also be set to: a reference frequency band in the full frequency band where the target audio is located except the speaking frequency band.
[0099] In one embodiment, when the preset frequency band is a speaking frequency band used to indicate the frequency of speaking sound, the audio processing device performs feature extraction processing on the target audio to obtain the energy proportion feature of the target audio, which may include: determining the average power spectrum of the target audio in the speaking frequency band, and the average power spectrum of the target audio in the entire frequency band; summing the average power spectrum of the target audio in the speaking frequency band to obtain a first energy value, and summing the average power spectrum of the target audio in the entire frequency band to obtain a second energy value; and using the ratio of the first energy value to the second energy value as the energy proportion feature of the target audio.
[0100] In a specific implementation, the audio processing device determines the average power spectrum of the target audio within the speaking sound frequency band, which may include: performing frame processing on the target audio to obtain multiple audio frames; determining the power spectrum of each audio frame, and averaging the power spectrum of each audio frame, and using the averaged result as the average power spectrum of the target audio; in the average power spectrum of the target audio, the average power spectrum in the speaking sound frequency band is determined as the average power spectrum of the target audio in the speaking sound frequency band.
[0101] Further, if x(i) represents the i-th audio sampling signal obtained after the target audio is subjected to time domain sampling processing, wherein i represents the i-th time domain sampling point obtained after the target audio is subjected to time domain sampling processing, i∈[0,L′-1], and L′ represents the number of time domain sampling points obtained after the target audio is subjected to time domain sampling processing; then the audio processing device performs frame processing on the target audio to obtain multiple audio frames, which can be shown by the following formula 4:
[0102] x h (i) = x(h·M+i) (4)
[0103] Among them, x h (i) represents the i-th audio sampling signal in the h-th audio frame, h∈[0,H-1], H represents the number of audio frames obtained by framing the target audio, i∈[0,L-1], L represents the frame length, that is, the number of time domain sampling points in each audio frame, and M represents the frame shift, that is, the number of time domain sampling points that overlap between the current audio frame and the next audio frame. The frame length and frame shift can be set according to specific needs. For example, the frame length can be set to 0.05 seconds (expressed as: t frmhop1 =0.05s), then L = f s *t frmhop1 , where f s Indicates the sampling frequency used when performing time-domain sampling processing on the target audio. The sampling frequency can be set according to specific needs; for example, the frame shift can be set to 0.25 seconds (expressed as: t frmhop2 =0.25s), then M = f s *t frmhop2 ; For example, the frame length can be set to 0.5 seconds, the frame shift can be set to 0.25 seconds; and so on.
[0104] Furthermore, in the process of determining the power spectrum of each audio frame, the audio processing device can first perform windowing processing on each audio frame to obtain the windowing processing result of the audio frame, and then perform Fourier transform on the windowing processing result of each audio frame to obtain the Fourier transform result corresponding to each audio frame, and then obtain the power spectrum of each audio frame based on the Fourier transform result corresponding to each audio frame; wherein, the process of framing, windowing and Fourier transforming the target audio frame is the process of performing short-time Fourier transform (i.e., short-time Fourier transform, STFT) on the target audio. In a specific implementation, the audio processing device performs windowing processing on the audio frame to obtain the windowing processing result of the audio frame, which can be shown by the following formula 5.1:
[0105] x h (i) = x h(i)·w(i) (5.1)
[0106] where x h (i) represents the i-th audio sampling signal in the h-th audio frame, w(i) represents the window function, and xw h (i) represents the windowing result of the i-th audio sampling signal in the h-th audio frame. The windowing result of the h-th audio frame includes: the windowing results of each audio sampling signal in the h-th audio frame; the window function can be selected according to specific requirements. For example, a rectangular window, a Hanning window, a Hamming window, etc. can be selected. In the embodiments of the present application, the Hanning window is used for illustrative purposes. Among them, the Hanning window can be shown by the following formula 5.2:
[0107]
[0108] Further, the audio processing device performs a Fourier transform on the windowing result of the audio frame to obtain the Fourier transform result corresponding to the audio frame, which can be shown by the following formula 6:
[0109]
[0110] where xw h (i) represents the windowing result of the i-th audio sampling signal in the h-th audio frame; K represents the number of points of the Fourier transform, which can be set according to specific requirements. When the frame length L < K, it is necessary to supplement xw h (i), and when L > K, truncation processing is performed (that is, K points are intercepted for Fourier transform); k represents the number of points of the k-th Fourier transform, and also represents the frequency point corresponding to the number of points of the k-th Fourier transform; j represents the imaginary unit; X(h,k) represents the Fourier transform result of the k-th Fourier transform point corresponding to the h-th audio frame, that is, the Fourier transform result of the k-th frequency point corresponding to the h-th audio frame. The Fourier transform result corresponding to the h-th audio frame includes: the Fourier transform results of each frequency point corresponding to the h-th audio frame.
[0111] Further, the audio processing device obtains the power spectrum of the audio frame based on the Fourier transform result corresponding to the audio frame, which can be shown by the following formula 7:
[0112] P(h,k) = ‖X(h,k)‖ 2 (7)
[0113] where P(h,k) represents the power spectrum of the k-th frequency point corresponding to the h-th audio frame. The power spectrum of the h-th audio frame includes: the power spectra of each frequency point corresponding to the h-th audio frame.
[0114] Furthermore, after obtaining the power spectrum of each audio frame, the audio processing device may average the power spectrum of each audio frame, and use the averaged result as the average power spectrum of the target audio. This process can be shown by the following formula 8:
[0115]
[0116] Among them, P(h,k) represents the power spectrum of the kth frequency point corresponding to the hth audio frame, and P(k) represents the result obtained by averaging the power spectra of the kth frequency points corresponding to each audio frame in the H audio frames obtained after frame processing, which can be used as the average power spectrum of the kth frequency point corresponding to the target audio. The average power spectrum of the target audio includes: the average power spectrum of each frequency point corresponding to the target audio.
[0117] After obtaining the average power spectrum of the target audio, the average power spectrum of the target audio in the speaking frequency band is determined as the average power spectrum of the target audio in the speaking frequency band, that is, the average power spectrum of the frequency points in the speaking frequency band in the average power spectrum of each frequency point corresponding to the target audio is determined as the average power spectrum of the target audio in the speaking frequency band; the average power spectrum of the target audio in the full frequency band where the target audio is located is determined as the average power spectrum of the target audio in the full frequency band, that is, the average power spectrum of the frequency points in the full frequency band where the target audio is located in the average power spectrum of each frequency point corresponding to the target audio is determined as the average power spectrum of the target audio in the full frequency band. Further, the audio processing device sums the average power spectrum of the target audio in the speaking frequency band to obtain a first energy value, and sums the average power spectrum of the target audio in the full frequency band to obtain a second energy value, and uses the ratio of the first energy value to the second energy value as the energy proportion feature of the target audio. The process can be shown by the following formula 9:
[0118]
[0119] Among them, r p Represents the energy ratio characteristics of the target audio; K L Indicates the lowest frequency point in the speaking frequency band, K U Indicates the highest frequency point in the speaking frequency band. When the speaking frequency band is between 300 and 3400 Hz, K L It can represent the frequency point corresponding to 300 Hz, K U It can represent the frequency point corresponding to 3400 Hz; K / 2 represents the number of frequency points in the full frequency band of the target audio. When the target audio is sampled in the time domain, the sampling frequency used is f s When K / 2 can be expressed as f s / 2 corresponding frequency point; represents the first energy value, represents the second energy value.
[0120] In one embodiment, in the process of determining the average power spectrum of the target audio in the speaking frequency band, the effective sounding part in the target audio can be framed to obtain multiple audio frames, and then the average power spectrum of the target audio in the speaking frequency band can be determined based on the obtained multiple audio frames; the accuracy of the determined average power spectrum of the target audio in the speaking frequency band can be improved by retaining the effective sounding part in the target audio and discarding the silent part, non-human voice part and other ineffective sounding parts in the target audio.
[0121] In a specific implementation, the audio processing device can perform speech recognition processing on the target audio to obtain the text content corresponding to the target audio and the utterance start time and end time of each text content in the target audio; intercept the target audio based on the utterance start time and utterance end time of each text content to obtain the audio segment corresponding to each text content; splice the audio segments corresponding to each text content to obtain the spliced audio; then the spliced audio can be framed to obtain multiple audio frames; and then the power spectrum of each audio frame can be determined, and the power spectrum of each audio frame can be averaged, and the result of the average processing is used as the average power spectrum of the target audio, and in the average power spectrum of the target audio, the average power spectrum in the speaking sound frequency band is determined as the average power spectrum of the target audio in the speaking sound frequency band. Among them, the audio processing device performs speech recognition processing on the target audio to obtain the text content corresponding to the target audio and the utterance start time and utterance end time of each text content in the target audio. The relevant process has been explained in step S401 and will not be repeated here; the relevant process of performing frame processing on the spliced audio and then determining the average power spectrum of the target audio in the speaking sound frequency band is similar to the above-mentioned relevant process of performing frame processing on the target audio and then determining the average power spectrum of the target audio in the speaking sound frequency band, and will not be repeated here.
[0122] S902: If the energy proportion feature of the target audio meets the preset feature value range corresponding to the energy proportion feature, the audio type of the target audio is determined to be a singing type.
[0123] S903: If the energy proportion feature of the target audio does not conform to the preset feature value range corresponding to the energy proportion feature, the audio type of the target audio is determined to be a speech type.
[0124] In step S902 to step S903, when the preset frequency band is a speaking frequency band used to indicate the frequency of the speaking voice, the preset characteristic value range corresponding to the energy proportion feature is a range less than or equal to the energy proportion threshold; the energy proportion threshold can be set according to specific needs. In a specific implementation, if the energy proportion feature of the target audio meets the preset characteristic value range corresponding to the energy proportion feature, the audio type of the target audio is judged as a singing type, which may include: if the energy proportion feature of the target audio is within the preset characteristic value range corresponding to the energy proportion feature, the audio type of the target audio is judged as a singing type; that is, if the energy proportion feature of the target audio is less than or equal to the energy proportion threshold, the audio type of the target audio is judged as a singing type.
[0125] Furthermore, if the energy proportion feature of the target audio does not meet the preset characteristic value range corresponding to the energy proportion feature, the audio type of the target audio is judged as a speaking voice type, which may include: if the energy proportion feature of the target audio is outside the preset characteristic value range corresponding to the energy proportion feature, the audio type of the target audio is judged as a speaking voice type; that is, if the energy proportion feature of the target audio is greater than the energy proportion threshold, the audio type of the target audio is judged as a speaking voice type.
[0126] In a feasible implementation manner, an energy proportion feature experience value of the singing type audio can be extracted from the energy proportion features of a large number of singing type audios, and an energy proportion feature experience value of the speaking type audio can be extracted from the energy proportion features of a large number of speaking type audios, and then an energy proportion threshold value can be determined based on the energy proportion feature experience value of the singing type audio and the energy proportion feature experience value of the speaking type audio, so that the determined energy proportion threshold value is greater than the energy proportion feature experience value of the singing type audio and less than the energy proportion feature experience value of the speaking type audio; for example, if the extracted energy proportion feature experience value of the singing type audio is 0.238 (expressed as: r p_sing =0.238), the energy ratio feature value of the extracted speech type audio is 0.328 (expressed as: r p_speech =0.328), then optionally, the determined energy proportion threshold may be 0.3 (expressed as: r thr =0.3).
[0127] In another feasible implementation, if the preset frequency band is set to: a reference frequency band within the full frequency band where the target audio is located except for the speaking frequency band; then, the preset feature value range corresponding to the energy proportion feature can be a range greater than the energy proportion threshold; the energy proportion threshold can be set according to specific needs.
[0128] See also Fig.10, which is a comparative schematic diagram of an average power spectrum provided in an embodiment of the present application; wherein, the curve marked as 1001 represents the average power spectrum of the audio 1 of the speaking type in the whole frequency band, the curve marked as 1002 represents the average power spectrum of the audio 2 of the singing type in the whole frequency band, and the frequency band marked as 1003 is the speaking frequency band; wherein, the average power spectrum is converted into decibel (dB) units to represent, and the converted average power spectrum can be shown by the following formula 10:
[0129] P dB (k) = 10·log 10 P(K) (10)
[0130] Among them, P(k) represents the average power spectrum of the kth frequency point corresponding to the target audio, P dB (k) represents the converted average power spectrum of the kth frequency point corresponding to the target audio.
[0131] In the embodiment of the present application, feature extraction processing can be performed on the target audio to obtain the energy proportion feature of the target audio; then, when the energy proportion feature of the target audio meets the preset feature value range corresponding to the energy proportion feature, the audio type of the target audio is judged as a singing type; when the energy proportion feature of the target audio does not meet the preset feature value range corresponding to the energy proportion feature, the audio type of the target audio is judged as a speaking type; that is, when the energy proportion feature of the target audio is less than or equal to the energy proportion threshold, the audio type of the target audio is judged as a singing type, and when the energy proportion feature of the target audio is greater than the energy proportion threshold The audio type of the target audio is determined as a speaking type; the audio type of the target audio can be determined based on the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the entire frequency band. When the preset frequency band is a speaking frequency band for indicating the frequency of the speaking voice, the energy in the speaking frequency band can be significant, that is, when the energy in the speaking frequency band accounts for a high proportion, the audio type can be determined as a speaking type; otherwise, the audio type can be determined as a singing type; that is, the audio type can be determined based on the significance of the energy in the speaking frequency band caused by the physical vocalization difference between speaking and singing, and the audio type of the audio can be accurately determined.
[0132] Based on the above audio processing solution, the embodiment of the present application provides another audio processing method, which is described by using multiple types of target audio discrimination features. Fig.11 A flowchart of another audio processing method provided in an embodiment of the present application. Fig.11 The audio processing method shown can be executed by an audio processing device, and can also be executed individually or collaboratively by other electronic devices with computing power. The embodiment of the present application takes an audio processing device as an example. Fig.11The audio processing method shown may include the following steps:
[0133] S1101, perform feature extraction processing on the target audio to obtain multiple types of discriminant features of the target audio.
[0134] Among them, any type discrimination feature of the target audio is used to discriminate whether the audio type of the target audio is a singing type or a speaking type; the multiple type discrimination features may include multiple types of speech speed features, fundamental frequency fluctuation features and energy proportion features; the fundamental frequency fluctuation feature of the target audio is used to indicate: the fluctuation of the fundamental frequency of the target audio within a continuous target time period, and the energy proportion feature of the target audio is used to indicate: the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the entire frequency band.
[0135] S1102: If, among the multiple type discrimination features of the target audio, there is at least one type discrimination feature that meets the preset feature value range corresponding to the at least one type discrimination feature, the audio type of the target audio is discriminated as a singing type.
[0136] S1103: If, among the multiple type discrimination features of the target audio, each type discrimination feature does not conform to the preset feature value range corresponding to each type discrimination feature, the audio type of the target audio is discriminated as a speech type.
[0137] In step S1102 to step S1103, for example, if the multiple type discrimination features of the target audio include speech rate features, fundamental frequency fluctuation features and energy proportion features; then, when the audio type of the target audio is judged as a singing type, any one of the following three conditions needs to be met: the speech rate feature of the target audio conforms to the preset feature value range corresponding to the speech rate feature, the fundamental frequency fluctuation feature of the target audio conforms to the preset feature value range corresponding to the fundamental frequency fluctuation feature, and the energy proportion feature of the target audio conforms to the preset feature value range corresponding to the energy proportion feature; that is, any one of the following three conditions needs to be met: the speech rate feature of the target audio is less than the first speech rate threshold or greater than the second speech rate threshold; the fundamental frequency fluctuation feature of the target audio is greater than the fundamental frequency fluctuation threshold; the energy proportion feature of the target audio is less than or equal to the energy proportion threshold.
[0138] When the audio type of the target audio is judged as a speaking voice type, the following three conditions need to be met at the same time: the speaking speed feature of the target audio does not meet the preset feature value range corresponding to the speaking speed feature, the fundamental frequency fluctuation feature of the target audio does not meet the preset feature value range corresponding to the fundamental frequency fluctuation feature, and the energy proportion feature of the target audio does not meet the preset feature value range corresponding to the energy proportion feature; that is, the following three conditions need to be met at the same time: the speaking speed feature of the target audio is greater than or equal to the first speaking speed threshold and less than or equal to the second speaking speed threshold; the fundamental frequency fluctuation feature of the target audio is less than or equal to the fundamental frequency fluctuation threshold; the energy proportion feature of the target audio is greater than the energy proportion threshold.
[0139] In one embodiment, when multiple type discrimination features among speech rate features, fundamental frequency fluctuation features, and energy proportion features are used to discriminate the audio type of the target audio, the target audio can be first subjected to feature extraction processing to obtain a type discrimination feature of the target audio. If the type discrimination feature of the target audio meets the preset feature value range corresponding to the corresponding type discrimination feature, the audio type of the target audio is discriminated as a singing type; if the type discrimination feature of the target audio does not meet the preset feature value range corresponding to the corresponding type discrimination feature, the relevant process of performing feature extraction processing on the target audio is repeated to obtain another type discrimination feature of the target audio until the audio type of the target audio can be discriminated; computing resources can be saved and processing speed can be improved. For example, when three type discrimination features among speech rate features, fundamental frequency fluctuation features, and energy proportion features are used to discriminate the audio type of the target audio, see Fig.12, is another discrimination flow chart for discriminating the audio type of audio provided in an embodiment of the present application; its steps may include: an audio processing device performs feature extraction processing on the target audio to obtain a speech speed feature of the target audio; determines whether the speech speed feature of the target audio meets a preset feature value range corresponding to the speech speed feature, that is, determines whether the speech speed feature of the target audio is less than a first speech speed threshold or greater than a second speech speed threshold; if the speech speed feature of the target audio meets the preset feature value range corresponding to the speech speed feature, that is, the speech speed feature of the target audio is less than the first speech speed threshold or greater than the second speech speed threshold, the audio type of the target audio is discriminated as a singing type; otherwise, feature extraction processing is performed on the target audio to obtain a fundamental frequency fluctuation feature of the target audio; determines whether the fundamental frequency fluctuation feature of the target audio meets the preset feature value range corresponding to the fundamental frequency fluctuation feature, that is, determines whether the target audio is a singing type. Whether the fundamental frequency fluctuation feature of the audio is greater than the fundamental frequency fluctuation threshold; if the fundamental frequency fluctuation feature of the target audio meets the preset feature value range corresponding to the fundamental frequency fluctuation feature, that is, the fundamental frequency fluctuation feature of the target audio is greater than the fundamental frequency fluctuation threshold, then the audio type of the target audio is judged as a singing type; otherwise, feature extraction processing is performed on the target audio to obtain the energy proportion feature of the target audio; it is judged whether the energy proportion feature of the target audio meets the preset feature value range corresponding to the energy proportion feature, that is, it is judged whether the energy proportion feature of the target audio is less than or equal to the energy proportion threshold; if the energy proportion feature of the target audio meets the preset feature value range corresponding to the energy proportion feature, that is, the energy proportion feature of the target audio is less than or equal to the energy proportion threshold, then the audio type of the target audio is judged as a singing type; otherwise, the audio type of the target audio is judged as a speaking type.
[0140] In an embodiment of the present application, when multiple type discrimination features among speech rate features, fundamental frequency fluctuation features and energy proportion features are used to discriminate the audio type of the target audio, if among the multiple type discrimination features of the target audio, there is at least one type discrimination feature that meets the preset feature value range corresponding to the at least one type discrimination feature, then the audio type of the target audio is discriminated as a singing type; if among the multiple type discrimination features of the target audio, each type discrimination feature does not meet the preset feature value range corresponding to each type discrimination feature, then the audio type of the target audio is discriminated as a speaking type; when multiple type discrimination features are used to discriminate the audio type of the target audio, the discrimination result can be made more accurate. And, further, when multiple type discrimination features among speech rate features, fundamental frequency fluctuation features and energy proportion features are used to discriminate the audio type of the target audio, the target audio can be firstly subjected to feature extraction processing to obtain a type discrimination feature of the target audio. If a type discrimination feature of the target audio meets the preset feature value range corresponding to the corresponding type discrimination feature, the audio type of the target audio is discriminated as a singing type; if a type discrimination feature of the target audio does not meet the preset feature value range corresponding to the corresponding type discrimination feature, the relevant process of performing feature extraction processing on the target audio is repeated to obtain another type discrimination feature of the target audio until the audio type of the target audio can be discriminated; computing resources can be saved and processing speed can be improved.
[0141] Based on the above audio processing method embodiment, the present application embodiment provides an audio processing device. Fig.13 , is a structural diagram of an audio processing device provided in an embodiment of the present application, and the audio processing device may include an extraction unit 1301 and a discrimination unit 1302. Fig.13 The audio processing device shown can run the following units:
[0142] The extraction unit 1301 is used to perform feature extraction processing on the target audio to obtain a type discrimination feature of the target audio; the type discrimination feature includes at least any one of the following: a speech rate feature, a fundamental frequency fluctuation feature, and an energy proportion feature, the fundamental frequency fluctuation feature is used to indicate the fluctuation of the fundamental frequency of the target audio in a continuous target time period, and the energy proportion feature is used to indicate the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the full frequency band;
[0143] The discrimination unit 1302 is used to determine the audio type of the target audio according to the comparison result of the type discrimination feature of the target audio and the preset feature value range corresponding to the type discrimination feature; the audio type is a singing type or a speaking type.
[0144] In one embodiment, when the discrimination unit 1302 determines the audio type of the target audio according to the comparison result between the type discrimination feature of the target audio and the preset feature value range corresponding to the type discrimination feature, the following operations are specifically performed:
[0145] If there are multiple type discrimination features of the target audio, and among the multiple type discrimination features of the target audio, there is at least one type discrimination feature that meets the preset feature value range corresponding to the at least one type discrimination feature, the audio type of the target audio is discriminated as a singing type;
[0146] If there are multiple type discrimination features of the target audio, and among the multiple type discrimination features of the target audio, each type discrimination feature does not conform to the preset feature value range corresponding to each type discrimination feature, the audio type of the target audio is discriminated as a speech type.
[0147] In one embodiment, when the extraction unit 1301 performs feature extraction processing on the target audio and obtains the speech rate feature of the target audio, the following operations are specifically performed:
[0148] Perform speech recognition processing on the target audio to obtain text content corresponding to the target audio and the utterance start time and utterance end time of each text content in the target audio;
[0149] Determine the average utterance duration of each text content based on the utterance start time and the utterance end time of each text content;
[0150] Based on the average utterance duration, the speech rate feature of the target audio is determined; the speech rate feature of the target audio is negatively correlated with the average utterance duration.
[0151] In one embodiment, when the extraction unit 1301 performs feature extraction processing on the target audio and obtains the fundamental frequency fluctuation feature of the target audio, the following operations are specifically performed:
[0152] Extracting the fundamental frequency of the target audio;
[0153] In the target time period, sampling the fundamental frequency of the target audio to obtain the fundamental frequency corresponding to each sampling point;
[0154] Based on the difference between the fundamental frequencies corresponding to the respective sampling points, the fundamental frequency fluctuation characteristics of the target audio are determined.
[0155] In one embodiment, when the extraction unit 1301 determines the fundamental frequency fluctuation characteristics of the target audio based on the difference between the fundamental frequencies corresponding to the sampling points, the following operations are specifically performed:
[0156] Performing standard deviation calculation on the fundamental frequencies corresponding to the respective sampling points to obtain a target standard deviation;
[0157] The target standard deviation is determined as a fundamental frequency fluctuation feature of the target audio.
[0158] In one embodiment, the extraction unit 1301 performs standard deviation calculation processing on the fundamental frequencies corresponding to the sampling points to obtain the target standard deviation, and specifically performs the following operations:
[0159] Performing note mapping processing on the fundamental frequency corresponding to each sampling point to obtain the note mapping result corresponding to each sampling point;
[0160] The standard deviation of the note mapping results corresponding to each sampling point is calculated to obtain the target standard deviation.
[0161] In one embodiment, the preset frequency band is a speech frequency band used to indicate the frequency of the speech sound;
[0162] When the extraction unit 1301 performs feature extraction processing on the target audio and obtains the energy proportion feature of the target audio, the following operations are specifically performed:
[0163] Determine an average power spectrum of the target audio in the speaking frequency band, and an average power spectrum of the target audio in the full frequency band;
[0164] Performing a summation process on the average power spectrum of the target audio in the speaking sound frequency band to obtain a first energy value, and performing a summation process on the average power spectrum of the target audio in the full frequency band to obtain a second energy value;
[0165] The ratio of the first energy value to the second energy value is used as an energy proportion feature of the target audio.
[0166] In one embodiment, when the extraction unit 1301 determines the average power spectrum of the target audio in the speaking sound frequency band, the following operations are specifically performed:
[0167] Performing frame processing on the target audio to obtain multiple audio frames;
[0168] Determine the power spectrum of each audio frame, and average the power spectrum of each audio frame, and use the averaged result as the average power spectrum of the target audio;
[0169] Among the average power spectra of the target audio, the average power spectrum in the speaking frequency band is determined as the average power spectrum of the target audio in the speaking frequency band.
[0170] In one embodiment, when the extraction unit 1301 performs frame processing on the target audio and obtains multiple audio frames, the following operations are specifically performed:
[0171] Perform speech recognition processing on the target audio to obtain text content corresponding to the target audio and the utterance start time and utterance end time of each text content in the target audio;
[0172] Intercepting the target audio based on the utterance start time and utterance end time of each text content to obtain an audio segment corresponding to each text content;
[0173] Splicing the audio clips corresponding to the various text contents to obtain spliced audio;
[0174] The spliced audio is processed by frame division to obtain the multiple audio frames.
[0175] In one embodiment, the preset characteristic value range corresponding to the speech rate feature is a range less than a first speech rate threshold or greater than a second speech rate threshold, wherein the first speech rate threshold is less than the second speech rate threshold;
[0176] The preset characteristic value range corresponding to the fundamental frequency fluctuation feature is a range greater than the fundamental frequency fluctuation threshold;
[0177] The preset characteristic value range corresponding to the energy proportion feature is a range that is less than or equal to the energy proportion threshold.
[0178] According to one embodiment of the present application, Figure 3 , Figure 4 , Figure 6 , Fig. 9 as well as Fig.11 The various steps involved in the audio processing method shown can be performed by Fig.13 The audio processing device shown in the figure is executed by each unit. For example, Figure 3 Step S301 shown can be performed by Fig.13 The extraction unit 1301 in the audio processing device shown is executed; Figure 3 Step S302 shown can be performed by Fig.13 The determination unit 1302 in the audio processing device shown in the figure is used for execution. Figure 4 Step S401 shown can be performed by Fig.13 The extraction unit 1301 in the audio processing device shown is executed; Figure 4 Steps S402 to S403 shown in FIG. Fig.13 The determination unit 1302 in the audio processing device shown in the figure is used for execution. Figure 6 Step S601 shown can be performed by Fig.13The extraction unit 1301 in the audio processing device shown is executed; Figure 6 Steps S602 to S603 shown in FIG. Fig.13 The determination unit 1302 in the audio processing device shown in the figure is used for execution. Fig. 9 Step S901 shown can be performed by Fig.13 The extraction unit 1301 in the audio processing device shown is executed; Fig. 9 Steps S902 to S903 shown in FIG. Fig.13 The determination unit 1302 in the audio processing device shown in the figure is used for execution. Fig.11 Step S1101 shown can be performed by Fig.13 The extraction unit 1301 in the audio processing device shown is executed; Fig.11 Steps S1102 to S1103 shown in FIG. Fig.13 The determination unit 1302 in the audio processing device shown is executed.
[0179] According to another embodiment of the present application, Fig.13 The various units in the audio processing device shown can be separately or completely combined into one or several other units to form, or one (some) of the units can be further divided into multiple smaller units in function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In practical applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the audio processing device divided based on logical functions can also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0180] According to another embodiment of the present application, the program can be executed by running on a general computing device such as a computer including a central processing unit (CPU), a random access memory medium (RAM), a read-only memory medium (ROM), and other processing elements and storage elements. Figure 3 , Figure 4 , Figure 6 , Fig. 9 as well as Fig.11 A computer program (including program code) for each step involved in the corresponding method shown in FIG. Fig.13 The audio processing device shown in and the audio processing method of the embodiment of the present application are implemented. The computer program can be recorded on, for example, a computer-readable storage medium, and loaded into the above-mentioned computing device through the computer-readable storage medium and run therein.
[0181] In an embodiment of the present application, the audio type of the target audio can be discriminated based on one or more type discrimination features of the target audio's speech rate feature, the target audio's fundamental frequency fluctuation feature, and the target audio's energy ratio feature. The target audio's speech rate feature can be used to indicate the target audio's speech rate. When the target audio's audio type is discriminated based on the target audio's speech rate feature, the audio type can be discriminated based on the speed of speech caused by the physical vocalization difference between speaking and singing, and the audio type of the audio can be accurately discriminated based on the difference in speech rate. The fundamental frequency fluctuation feature of the target audio can be used to indicate the fluctuation of the fundamental frequency of the target audio within a continuous target time period. Since the fundamental frequency determines the tone of the sound, when the target audio's audio type is discriminated based on the fundamental frequency fluctuation feature of the target audio, the audio type can be discriminated based on the fluctuation of the tone caused by the physical vocalization difference between speaking and singing, and the audio type of the audio can be accurately discriminated based on the difference in tone fluctuation size. The energy proportion feature of the target audio can be used to indicate the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the entire frequency band. When the preset frequency band is set to a speaking frequency band used to indicate the frequency of the speaking voice, when the audio type of the target audio is determined based on the energy proportion feature of the target audio, the audio type can be determined based on the significant energy of the speaking frequency band caused by the physical vocalization difference between speaking and singing. The audio type of the audio can be accurately determined based on the significant energy of the frequency band. Moreover, when the audio type of the target audio is determined based on multiple types of distinguishing features such as the speech rate feature of the target audio, the fundamental frequency fluctuation feature of the target audio, and the energy proportion feature of the target audio, multiple types of distinguishing features can be fully utilized to further improve the accuracy of audio type determination.
[0182] Based on the above audio processing method embodiment and audio processing device embodiment, the present application also provides an audio processing device. Fig.14 , is a structural diagram of an audio processing device provided in an embodiment of the present application. Fig.14 The audio processing device shown may include at least a processor 1401, an input interface 1402, an output interface 1403, and a computer storage medium 1404. The processor 1401, the input interface 1402, the output interface 1403, and the computer storage medium 1404 may be connected via a bus or other means.
[0183] The computer storage medium 1404 may be stored in the memory of the audio processing device. The computer storage medium 1404 is used to store a computer program. The computer program includes program instructions. The processor 1401 is used to execute the program instructions stored in the computer storage medium 1404. The processor 1401 (or CPU (Central Processing Unit)) is the computing core and control core of the audio processing device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the above-mentioned audio processing method flow or corresponding functions.
[0184] The embodiment of the present application also provides a computer storage medium (Memory), which is a memory device in an audio processing device for storing programs and data. It is understandable that the computer storage medium here can include both the built-in storage medium in the terminal and the extended storage medium supported by the terminal. The computer storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor 1401 are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed random access memory (RAM) memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer storage medium located away from the aforementioned processor.
[0185] In one embodiment, the processor 1401 may load and execute one or more instructions stored in a computer storage medium to implement the above-mentioned Figure 3 , Figure 4 , Figure 6 , Fig. 9 as well as Fig.11 In the corresponding steps of the method in the audio processing method embodiment, in a specific implementation, the processor 1401 is used to:
[0186] Performing feature extraction processing on the target audio to obtain a type discrimination feature of the target audio; the type discrimination feature includes at least any one of the following: a speech rate feature, a fundamental frequency fluctuation feature, and an energy proportion feature, the fundamental frequency fluctuation feature is used to indicate the fluctuation of the fundamental frequency of the target audio within a continuous target time period, and the energy proportion feature is used to indicate the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the full frequency band;
[0187] The audio type of the target audio is determined based on a comparison result between the type discrimination feature of the target audio and a preset feature value range corresponding to the type discrimination feature; the audio type is a singing type or a speaking type.
[0188] In one embodiment, when the processor 1401 determines the audio type of the target audio according to the comparison result between the type discrimination feature of the target audio and the preset feature value range corresponding to the type discrimination feature, the processor 1401 specifically performs the following operations:
[0189] If there are multiple type discrimination features of the target audio, and among the multiple type discrimination features of the target audio, there is at least one type discrimination feature that meets the preset feature value range corresponding to the at least one type discrimination feature, the audio type of the target audio is discriminated as a singing type;
[0190] If there are multiple type discrimination features of the target audio, and among the multiple type discrimination features of the target audio, each type discrimination feature does not conform to the preset feature value range corresponding to each type discrimination feature, the audio type of the target audio is discriminated as a speech type.
[0191] In one embodiment, the processor 1401 performs feature extraction processing on the target audio to obtain the speech rate feature of the target audio by specifically performing the following operations:
[0192] Perform speech recognition processing on the target audio to obtain text content corresponding to the target audio and the utterance start time and utterance end time of each text content in the target audio;
[0193] Determine the average utterance duration of each text content based on the utterance start time and the utterance end time of each text content;
[0194] Based on the average utterance duration, the speech rate feature of the target audio is determined; the speech rate feature of the target audio is negatively correlated with the average utterance duration.
[0195] In one embodiment, when the processor 1401 performs feature extraction processing on the target audio and obtains the fundamental frequency fluctuation feature of the target audio, the processor 1401 specifically performs the following operations:
[0196] Extracting the fundamental frequency of the target audio;
[0197] In the target time period, sampling the fundamental frequency of the target audio to obtain the fundamental frequency corresponding to each sampling point;
[0198] Based on the difference between the fundamental frequencies corresponding to the respective sampling points, the fundamental frequency fluctuation characteristics of the target audio are determined.
[0199] In one embodiment, when the processor 1401 determines the fundamental frequency fluctuation characteristics of the target audio based on the difference between the fundamental frequencies corresponding to the respective sampling points, the processor 1401 specifically performs the following operations:
[0200] Performing standard deviation calculation on the fundamental frequencies corresponding to the respective sampling points to obtain a target standard deviation;
[0201] The target standard deviation is determined as a fundamental frequency fluctuation feature of the target audio.
[0202] In one embodiment, the processor 1401 performs standard deviation calculation processing on the fundamental frequencies corresponding to the sampling points to obtain a target standard deviation, and specifically performs the following operations:
[0203] Performing note mapping processing on the fundamental frequency corresponding to each sampling point to obtain the note mapping result corresponding to each sampling point;
[0204] The standard deviation of the note mapping results corresponding to each sampling point is calculated to obtain the target standard deviation.
[0205] In one embodiment, the preset frequency band is a speech frequency band used to indicate the frequency of the speech sound;
[0206] When the processor 1401 performs feature extraction processing on the target audio and obtains the energy proportion feature of the target audio, the processor 1401 specifically performs the following operations:
[0207] Determine an average power spectrum of the target audio in the speaking frequency band, and an average power spectrum of the target audio in the full frequency band;
[0208] Performing a summation process on the average power spectrum of the target audio in the speaking sound frequency band to obtain a first energy value, and performing a summation process on the average power spectrum of the target audio in the full frequency band to obtain a second energy value;
[0209] The ratio of the first energy value to the second energy value is used as an energy proportion feature of the target audio.
[0210] In one embodiment, when the processor 1401 determines the average power spectrum of the target audio in the speaking frequency band, it specifically performs the following operations:
[0211] Performing frame processing on the target audio to obtain multiple audio frames;
[0212] Determine the power spectrum of each audio frame, and average the power spectrum of each audio frame, and use the averaged result as the average power spectrum of the target audio;
[0213] Among the average power spectra of the target audio, the average power spectrum in the speaking frequency band is determined as the average power spectrum of the target audio in the speaking frequency band.
[0214] In one embodiment, the processor 1401 performs frame processing on the target audio to obtain multiple audio frames, and specifically performs the following operations:
[0215] Perform speech recognition processing on the target audio to obtain text content corresponding to the target audio and the utterance start time and utterance end time of each text content in the target audio;
[0216] Intercepting the target audio based on the utterance start time and utterance end time of each text content to obtain an audio segment corresponding to each text content;
[0217] Splicing the audio clips corresponding to the various text contents to obtain spliced audio;
[0218] The spliced audio is processed by frame division to obtain the multiple audio frames.
[0219] In one embodiment, the preset characteristic value range corresponding to the speech rate feature is a range less than a first speech rate threshold or greater than a second speech rate threshold, wherein the first speech rate threshold is less than the second speech rate threshold;
[0220] The preset characteristic value range corresponding to the fundamental frequency fluctuation feature is a range greater than the fundamental frequency fluctuation threshold;
[0221] The preset characteristic value range corresponding to the energy proportion feature is a range that is less than or equal to the energy proportion threshold.
[0222] The present application provides a computer program product, which includes a computer program stored in a computer storage medium; a processor of an audio processing device reads the computer program from the computer storage medium, and the processor executes the computer program, so that the audio processing device performs the above-mentioned Figure 3 , Figure 4 , Figure 6 , Fig. 9 as well as Fig.11 The computer-readable storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).
[0223] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An audio processing method, characterized in that: include: Performing feature extraction processing on the target audio to obtain a type discrimination feature of the target audio; The type discrimination feature includes an energy proportion feature, and the type discrimination feature also includes one or more of a speech rate feature and a fundamental frequency fluctuation feature, the fundamental frequency fluctuation feature is used to indicate the fluctuation of the fundamental frequency of the target audio within a continuous target time period, and the energy proportion feature is used to indicate the difference between the energy of the target audio in a preset frequency band and the energy of the target audio in the full frequency band; the energy proportion feature is determined based on a ratio of a first energy value to a second energy value, the first energy value is determined based on an average power spectrum of the target audio in the preset frequency band, and the second energy value is determined based on an average power spectrum of the target audio in the full frequency band; The audio type of the target audio is determined based on a comparison result between the type discrimination feature of the target audio and a preset feature value range corresponding to the type discrimination feature; the audio type is a singing type or a speaking type.
2. The method according to claim 1, characterized in that The step of determining the audio type of the target audio according to a comparison result between the type discrimination feature of the target audio and a preset feature value range corresponding to the type discrimination feature comprises: If, among the multiple type discrimination features of the target audio, there is at least one type discrimination feature that meets the preset feature value range corresponding to the at least one type discrimination feature, the audio type of the target audio is discriminated as a singing type; If, among the multiple type discrimination features of the target audio, each type discrimination feature does not conform to the preset feature value range corresponding to each type discrimination feature, the audio type of the target audio is discriminated as a speech type.
3. The method according to claim 1, characterized in that Performing feature extraction processing on the target audio to obtain the speech rate feature of the target audio includes: Perform speech recognition processing on the target audio to obtain text content corresponding to the target audio and the utterance start time and utterance end time of each text content in the target audio; Determine the average utterance duration of each text content based on the utterance start time and the utterance end time of each text content; Based on the average utterance duration, the speech rate feature of the target audio is determined; the speech rate feature of the target audio is negatively correlated with the average utterance duration.
4. The method according to claim 1, characterized in that Performing feature extraction processing on the target audio to obtain fundamental frequency fluctuation features of the target audio includes: Extracting the fundamental frequency of the target audio; In the target time period, sampling the fundamental frequency of the target audio to obtain the fundamental frequency corresponding to each sampling point; Based on the difference between the fundamental frequencies corresponding to the respective sampling points, the fundamental frequency fluctuation characteristics of the target audio are determined.
5. The method according to claim 4, characterized in that The determining the fundamental frequency fluctuation characteristics of the target audio based on the difference between the fundamental frequencies corresponding to the respective sampling points includes: Performing standard deviation calculation on the fundamental frequencies corresponding to the respective sampling points to obtain a target standard deviation; The target standard deviation is determined as a fundamental frequency fluctuation feature of the target audio.
6. The method according to claim 5, characterized in that The step of performing standard deviation calculation on the fundamental frequencies corresponding to the respective sampling points to obtain a target standard deviation includes: Performing note mapping processing on the fundamental frequency corresponding to each sampling point to obtain the note mapping result corresponding to each sampling point; The standard deviation of the note mapping results corresponding to each sampling point is calculated to obtain the target standard deviation.
7. The method according to claim 1, characterized in that The preset frequency band is a speaking frequency band used to indicate the frequency of the speaking sound; Performing feature extraction processing on the target audio to obtain energy proportion features of the target audio includes: Determine an average power spectrum of the target audio in the speaking frequency band, and an average power spectrum of the target audio in the full frequency band; Performing a summation process on the average power spectrum of the target audio in the speaking sound frequency band to obtain a first energy value, and performing a summation process on the average power spectrum of the target audio in the full frequency band to obtain a second energy value; The ratio of the first energy value to the second energy value is used as an energy proportion feature of the target audio.
8. The method according to claim 7, characterized in that The determining the average power spectrum of the target audio in the speaking frequency band comprises: Performing frame processing on the target audio to obtain multiple audio frames; Determine the power spectrum of each audio frame, and average the power spectrum of each audio frame, and use the averaged result as the average power spectrum of the target audio; Among the average power spectra of the target audio, the average power spectrum in the speaking frequency band is determined as the average power spectrum of the target audio in the speaking frequency band.
9. The method according to claim 8, characterized in that The frame processing of the target audio to obtain a plurality of audio frames includes: Perform speech recognition processing on the target audio to obtain text content corresponding to the target audio and the utterance start time and utterance end time of each text content in the target audio; Based on the utterance start time and utterance end time of each text content, the target audio is intercepted to obtain the audio segment corresponding to each text content; Splicing the audio clips corresponding to the various text contents to obtain spliced audio; The spliced audio is processed by frame division to obtain the multiple audio frames.
10. The method according to claim 1, characterized in that The preset characteristic value range corresponding to the speech rate feature is a range less than the first speech rate threshold or greater than the second speech rate threshold, wherein the first speech rate threshold is less than the second speech rate threshold; The preset characteristic value range corresponding to the fundamental frequency fluctuation feature is a range greater than the fundamental frequency fluctuation threshold; The preset characteristic value range corresponding to the energy proportion feature is a range that is less than or equal to the energy proportion threshold.
11. An audio processing device, characterized in that: The audio processing device includes an input interface and an output interface, and also includes: a processor adapted to implement one or more instructions; and, A computer storage medium storing one or more instructions, wherein the one or more instructions are suitable for being loaded by the processor and executing the audio processing method according to any one of claims 1 to 10.
12. A computer storage medium, characterized in that: The computer storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, they are used to execute the audio processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Environment sound control method and system for automobile intelligent cabin
CN111816199A
Audio processing method based on howling detection and electronic equipment
CN114464205A