Audio Processing Method, Apparatus, Computer Device, and Storage Medium

By obtaining the voice information density and quantity of the audio clip in the audio processing method, determining the target pause duration and inserting the pause clip, the problem of poor audio content transmission caused by excessive speech speed is solved, and the understanding and absorption effect of the audio content is improved.

CN114360501BActive Publication Date: 2025-07-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111241118.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-25
Publication Date
2025-07-01
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

In audio content, the fast and continuous voice content makes it difficult for the listener to keep up and understand, and the traditional slowdown method cannot effectively improve the effectiveness of the communication of audio content.

Method used

By obtaining the voice information density and voice information volume of the current audio clip in the audio signal, the target pause duration is determined, and the pause interval is inserted between the current audio clip and the subsequent audio clip. If the speech interval is less than the target pause duration.

Benefits of technology

It improves the effectiveness of audio content, gives listeners more time to understand and absorb audio content, and avoids content understanding and reception problems caused by too fast speech speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114360501B_ABST
    Figure CN114360501B_ABST
Patent Text Reader

Abstract

This application relates to an audio processing method, apparatus, computer device, and storage medium. The method includes: obtaining the speech information density and speech information quantity of a current audio segment in an audio signal; the speech information density is used to measure the frequency of speech information fluctuations; determining a target pause duration of the current audio segment based on the speech information density and the speech information quantity; obtaining the speech interval duration between the current audio segment and a subsequent audio segment in the audio signal; and if the speech interval duration is less than the target pause duration, inserting a pause segment between the current audio segment and the subsequent audio segment. Using this method can improve the effectiveness of audio content transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to an audio processing method, apparatus, computer device, storage medium, and computer program. Background Art

[0002] In some audio content programs, when the speaker's speech rate is fast and there is a lot of continuous speech content, it is easy for listeners to fail to keep up with the speech content, resulting in not understanding or comprehending it. In the case of not understanding the previous speech content, the subsequent speech content may be missed, affecting the reception and understanding of other subsequent speech segments.

[0003] In traditional technologies, usually, the audio with too fast speech rate is decelerated. However, the traditional audio processing method by decelerating can make the listener hear the audio content more clearly, but the listener may not be able to fully understand and absorb the audio content, and there is a problem of low effectiveness in the transmission of audio content. Summary of the Invention

[0004] Based on this, it is necessary to provide an audio processing method, apparatus, computer device, and storage medium that can improve the effectiveness of audio content transmission for the above technical problems.

[0005] An audio processing method, the method includes:

[0006] Obtain the speech information density and speech information volume of the current audio segment in the audio signal; the speech information density is used to measure the frequency of speech information fluctuations;

[0007] Based on the speech information density and the speech information volume, determine the target pause duration of the current audio segment;

[0008] Obtain the speech interval duration between the current audio segment and the subsequent audio segment in the audio signal;

[0009] If the speech interval duration is less than the target pause duration, insert a pause segment between the current audio segment and the subsequent audio segment.

[0010] A speech processing apparatus, the apparatus includes:

[0011] An obtaining module, configured to obtain the speech information density and speech information volume of the current audio segment in the audio signal; the speech information density is used to measure the frequency of speech information fluctuations;

[0012] A determining module, configured to determine the target pause duration of the current audio segment based on the speech information density and the speech information volume;

[0013] The obtaining module is further configured to obtain the speech interval duration between the current audio segment and the subsequent audio segment in the audio signal;

[0014] An insertion module, configured to insert a pause segment between the current audio segment and the subsequent audio segment if the speech interval duration is less than the target pause duration.

[0015] In one embodiment, the determining module is further configured to sequentially obtain each audio frame in the audio signal and detect the speech information of each audio frame; and determine the current audio segment based on the speech information of each audio frame.

[0016] In one embodiment, the determining module is further configured to use the first speech frame in the audio signal as the start node of the audio segment, or use the first speech frame after the end of the previous audio segment as the start node of the audio segment; wherein, the speech frame is an audio frame whose speech information meets the speech active condition; if there are consecutive non-speech frames exceeding a specified number after the start node, then based on the consecutive non-speech frames exceeding the specified number, determine the end node of the audio segment; wherein, the non-speech frame is an audio frame whose speech information does not meet the speech active condition; based on the start node and the end node, determine the current audio segment.

[0017] In one embodiment, for the currently obtained current audio frame, if the current start marker parameter is the first value and the current audio frame is a speech frame, the determining module is further configured to use the current audio frame as the start node of the audio segment and adjust the start marker parameter from the first value to the second value; wherein, when it is detected that there are consecutive non-speech frames of a specified number, the start marker parameter is adjusted from the second value to the first value.

[0018] In one embodiment, for the audio frames after the start node, if the current audio frame is a non-speech frame, the determining module is further configured to increment a non-speech count value by one; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, use the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute until the non-speech count value exceeds the specified number; use the last non-speech frame in the consecutive non-speech frames of the specified number as the end node of the audio segment.

[0019] In one embodiment, the determining module is further configured to, for the audio frames after the start node, if the current audio frame is a non-speech frame, increment a non-speech count value by one; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, use the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute until the non-speech count value is set to zero after exceeding the specified number; use the previous frame before the speech frame that triggers the non-speech count value to be set to zero as the end node of the audio segment.

[0020] In one embodiment, the determining module is further configured to, if there are consecutive non-speech frames exceeding the specified number after the start node, use the first speech frame that appears after the consecutive non-speech frames exceeding the specified number as the start node of the next audio segment.

[0021] In one embodiment, the obtaining module is further configured to obtain the current audio segment in the audio signal; perform pitch frequency detection on each audio frame in the current audio segment to obtain the number of pitch frequency fluctuations in the current audio segment; the number of pitch frequency fluctuations characterizes the speech information amount; determine the speech duration of the current audio segment; based on the comparison value between the number of pitch frequency fluctuations and the speech duration, determine the speech information density of the current audio segment.

[0022] In one embodiment, the obtaining module is further configured to perform pitch frequency detection on each audio frame in the current audio segment to obtain the pitch frequency value of each speech frame; wherein, the speech frame is an audio frame whose speech information meets the speech activity condition; based on the pitch frequency values between adjacent speech frames, determine the number of pitch frequency fluctuations in the current audio segment.

[0023] In one embodiment, the obtaining module is further configured to determine the pitch frequency state corresponding to each of the adjacent speech frames based on the pitch frequency values between the adjacent speech frames; if the pitch frequency states of two speech frames in the adjacent speech frames are different, determine that there is one pitch frequency fluctuation; based on all the pitch frequency fluctuations that occur in the current audio segment, statistically obtain the number of pitch frequency fluctuations in the current audio segment.

[0024] In one embodiment, the determining module is further configured to input the speech information density and the speech information quantity into a trained neural network model, and output a target pause duration through the trained neural network model; the apparatus further includes a training module; the training module is further configured to obtain a sample audio signal; the sample speech interval duration between adjacent sample audio segments in the sample audio signal is within a preset duration range; determine the sample speech information density and the sample speech information quantity of each sample audio segment in the sample audio signal; use the sample speech information density and the sample speech information quantity of the sample audio segment as training inputs, and use the sample speech interval duration corresponding to the corresponding sample audio segment as a training label; train the neural network model based on the training inputs and the training labels until the training stop condition is reached and then stop, to obtain a trained neural network model.

[0025] In one embodiment, the inserting module is further configured to, if the speech interval duration is less than the target pause duration, determine an inserted pause duration based on the target pause duration and the speech interval duration; the sum of the inserted pause duration and the speech interval duration is greater than or equal to the target pause duration; obtain a pause segment of the inserted pause duration, and insert the pause segment between the current audio segment and the subsequent audio segment.

[0026] In one embodiment, the audio signal includes pre-recorded audio content, and the inserting module is further configured to continue to detect the speech interval duration and insert pause segments for subsequent audio segments in the audio signal until the speech interval duration between all adjacent audio segments in the audio signal is adjusted to the target pause duration, to obtain a target audio; the apparatus further includes a playing module; the playing module is configured to play the target audio, and during the playing of the target audio, there is a pause effect with a target pause duration between different content segments of the audio content.

[0027] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0028] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0029] A computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium, the processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the above method.

[0030] The above audio processing method, device, computer device, and storage medium obtain the speech information density and speech information volume of the current audio segment in the audio signal. Then, based on the speech information density and speech information volume, the target pause duration for the current audio segment to pause can be accurately determined. If the speech interval duration between the current audio segment and the subsequent audio segment is less than the target pause duration, a pause segment is inserted between the current audio segment and the subsequent audio segment, and a longer pause can be made after the current audio segment, more accurately regulating the pause duration of the current audio segment, allowing the listener to have more time to understand and absorb the content of the current audio segment, effectively conveying the audio content of the current audio segment, and improving the effectiveness of audio content transmission. Description of the Drawings

[0031] Figure 1 It is an application environment diagram of the audio processing method in an embodiment;

[0032] Figure 2 It is a schematic flowchart of the audio processing method in an embodiment;

[0033] Figure 3 It is a schematic flowchart of the step of determining the current audio segment based on the speech information of each audio frame in an embodiment;

[0034] Figure 4 It is a schematic flowchart of the audio processing method in another embodiment;

[0035] Figure 5 It is a schematic flowchart of the step of detecting the fundamental frequency of each audio frame in the current audio segment to obtain the number of fundamental frequency fluctuations of the current audio segment in an embodiment;

[0036] Figure 6 It is a schematic flowchart of the step of determining the number of fundamental frequency fluctuations of the current audio segment based on the fundamental frequency values between adjacent speech frames in an embodiment;

[0037] Figure 7 It is a schematic flowchart of the training step of the neural network model in an embodiment;

[0038] Figure 8 It is a schematic flowchart of the audio processing method in another embodiment;

[0039] Figure 9 It is a structural block diagram of the audio processing device in an embodiment;

[0040] Figure 10 It is an internal structure diagram of the computer device in an embodiment. Detailed Implementation Modes

[0041] In order to make the objectives, technical solutions and advantages of this application more clear and understandable, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0042] The audio processing method provided by this application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. The terminal 102 and the server 104 can be used alone to execute the audio processing method provided by the embodiments of this application, or can be used in cooperation to execute the audio processing method provided by the embodiments of this application. Taking the terminal 102 executing this audio processing method alone as an example, the server 104 sends an audio signal to the terminal 102. The terminal 102 obtains the speech information density and the speech information amount of the current audio segment in the audio signal; the speech information density is used to measure the frequency of speech information fluctuations; based on the speech information density and the speech information amount, the target pause duration of the current audio segment is determined; the speech interval duration between the current audio segment and the subsequent audio segment in the audio signal is obtained; if the speech interval duration is less than the target pause duration, a pause segment is inserted between the current audio segment and the subsequent audio segment.

[0043] Among them, the terminal 102 can be but is not limited to a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, a smart TV. The server 104 can be an independent physical server, can also be a server cluster or a distributed system composed of multiple physical servers, or can also be a cloud server providing cloud computing services.

[0044] In one embodiment, as Figure 2 shown, an audio processing method is provided. This method is applied to a computer device. The computer device can be Figure 1 the terminal or the server in it. The audio processing method includes the following steps:

[0045] Step S202, obtain the speech information density and the speech information amount of the current audio segment in the audio signal; the speech information density is used to measure the frequency of speech information fluctuations.

[0046] Among them, the audio signal can specifically be the audio signal of an audio program, the audio signal of a recording, the audio signal of a song, etc. In the audio signal, there is one or more audio segments, and there is a pause segment between every two adjacent audio segments for the listener to understand the content of the previous audio segment.

[0047] The speech information density refers to the density of speech information in the current audio segment and is used to measure the frequency of fluctuations in speech information. The speech information volume refers to the amount of information contained in the current audio segment and is used to measure the speech energy in the audio segment, which can be characterized by the number of fluctuations in the fundamental frequency. It can be understood that the higher the speech information density of the current audio segment per unit time, the more speech information it contains; the more the total speech information contained in the current audio segment, the more the speech information volume.

[0048] Specifically, the computer device uses fundamental frequency detection technology to detect the current audio segment in the audio signal and obtains the speech information density and speech information volume of the current audio segment.

[0049] It can be understood that generally, sounds are composed of a series of vibrations with different frequencies and amplitudes emitted by a sounding body. Among these vibrations, there is a vibration with the lowest frequency, and the sound emitted by it is the fundamental tone, and the rest are overtones.

[0050] The speech information density is characterized by the frequency of fluctuations in the fundamental frequency of the current audio segment; the speech information volume is characterized by the number of fluctuations in the fundamental frequency contained in the current audio segment.

[0051] The number of fluctuations in the fundamental frequency refers to the number of times the frequency of the fundamental tone fluctuates in the current audio segment. For example, if the fundamental frequency in the current audio segment changes from f1 to f2 in sequence, and then from f2 to f3, the number of fluctuations in the fundamental frequency of the current audio segment is 2. The frequency of fluctuations in the fundamental frequency refers to the number of fluctuations in the fundamental frequency per unit time. It can be understood that the higher the frequency of fluctuations in the fundamental frequency per unit time, the higher the speech information density of the current audio segment. On the contrary, if the frequency of fluctuations in the fundamental frequency is lower, it means that the speech information density of the current audio segment is lower. The more the number of fluctuations in the fundamental frequency, the more speech information the current audio segment contains, that is, the more the speech information volume.

[0052] Step S204, based on the speech information density and speech information volume, determine the target pause duration of the current audio segment.

[0053] Among them, the target pause duration refers to the duration for which the current audio segment needs to pause.

[0054] It can be understood that the higher the speech information density of the current audio segment, the more speech information it contains per unit time; the more the speech information volume of the current audio segment, the more speech information it contains. Then, after listening to the current audio segment, the listener needs to spend more time to understand and absorb the speech information of the current audio segment. Therefore, the longer the duration for which a pause is needed after the current audio segment.

[0055] Specifically, the computer device inputs the speech information density and the speech information volume into the trained neural network model, and calculates the target pause duration corresponding to the speech information density and the speech information volume through the neural network model.

[0056] Step S206: Obtain the speech interval duration between the current audio segment and the subsequent audio segment in the audio signal.

[0057] Among them, the speech interval duration refers to the interval duration of the speech content between the current audio segment and the subsequent audio segment. Among them, each audio frame in the audio segment between the current audio segment and the subsequent audio segment does not meet the speech activity condition. When an audio frame meets the speech activity condition, it can be determined that the audio frame contains speech information and the audio frame is a speech frame; when an audio frame does not meet the speech activity condition, it can be determined that the audio frame does not contain speech information and the audio frame is a non-speech frame. The speech activity condition can be set as needed. For example, the speech activity condition can be that the speech activity detection algorithm determines that the audio frame is a speech frame, or that the audio frame contains human voices, or that the sound decibel in the audio frame is greater than a specified decibel value, etc., and is not limited thereto.

[0058] Specifically, the computer device determines the end node of the current audio segment and the start node of the subsequent audio segment, and counts the speech interval duration of the audio segment between the end node of the current audio segment and the start node of the subsequent audio segment.

[0059] In one embodiment, the computer device can determine the last speech frame in the current audio segment and the first audio frame of the subsequent audio segment, and use the audio frame interval duration between the last speech frame in the current audio segment and the first audio frame of the subsequent audio segment as the speech interval duration.

[0060] It should be noted that the subsequent audio segment can specifically be the next audio segment after the current audio segment, or the Nth (N is a positive integer greater than 1) audio segment after the current audio segment, etc., and the embodiments of the present application do not limit this.

[0061] Step S208: If the speech interval duration is less than the target pause duration, insert a pause segment between the current audio segment and the subsequent audio segment.

[0062] Among them, a pause segment refers to an audio segment used for pausing after the current audio segment to increase the speech interval duration between the current audio segment and the subsequent audio segment. It should be noted that the duration of the pause segment can be set as needed. For example, the duration of the pause segment can be the difference between the target pause duration and the speech interval duration, or a duration greater than this difference, or a duration less than this difference. It can be understood that regardless of the specific duration of the pause segment, it can increase the speech interval duration between the current audio segment and the subsequent audio segment, enabling the listener to have more time to understand the content of the current audio segment.

[0063] Specifically, if the speech interval duration is less than the target pause duration, the computer device uses comfort noise generation (cng) technology to obtain a pause segment and inserts this pause segment between the current audio segment and the subsequent audio segment. Among them, this pause segment can be a silent segment or a small noise segment that does not contain any speech information.

[0064] It should be noted that this pause segment can be inserted at any position between the current audio segment and the subsequent audio segment. For example, this pause segment can be inserted after the end of the current audio segment, or before the subsequent audio segment, or in the middle of the current audio segment and the subsequent audio segment.

[0065] For the above audio processing method, the speech information density and speech information volume of the current audio segment in the audio signal are obtained. Then, based on the speech information density and speech information volume, the target pause duration required for the current audio segment to pause can be accurately determined. If the speech interval duration between the current audio segment and the subsequent audio segment is less than this target pause duration, a pause segment is inserted between the current audio segment and the subsequent audio segment, which can pause for a longer time after the current audio segment, more accurately regulate the pause duration of the current audio segment, enable the listener to have more time to understand and absorb the content of the current audio segment, effectively convey the audio content of the current audio segment, and improve the effectiveness of audio content transmission.

[0066] In one embodiment, before obtaining the speech information density and speech information volume of the current audio segment in the audio signal, it further includes a determination step of the current audio segment. The determination step of this audio segment includes: sequentially obtaining each audio frame in the audio signal and detecting the speech information of each audio frame; based on the speech information of each audio frame, determining the current audio segment.

[0067] Specifically, the computer device sequentially obtains each audio frame in the audio signal according to the playback order of the audio signal, and uses a voice activity detection (VAD) algorithm to detect the voice information of each audio frame; it determines whether the voice information of each audio frame meets the voice activity condition. If the voice information of the audio frame meets the voice activity condition, the audio frame is a voice frame; if the voice information of the audio frame does not meet the voice activity condition, the audio frame is a non-voice frame. Among them, the computer device can set a specified duration for each interval in the audio signal as an audio frame. The specified duration can be set as needed. For example, the specified duration is 20 ms (milliseconds), that is, the duration of each audio frame is 20 ms, and an audio frame is obtained every 20 ms in the audio signal.

[0068] Further, the voice activity condition can be set to that the audio frame contains voice information. Then, if the audio frame contains voice information, the audio frame is a voice frame; if the audio frame does not contain voice information, the audio frame is a non-voice frame.

[0069] When more than a specified number of consecutive audio frames are all non-voice frames, the last frame among the more than a specified number of consecutive audio frames is the last frame of the current audio segment, that is, the end node of the current audio segment, and the current audio segment is obtained.

[0070] In this embodiment, each audio frame in the audio signal is sequentially obtained, and the voice information of each audio frame is detected. Based on the voice information of each audio frame, the current audio segment can be accurately determined.

[0071] In one embodiment, as Figure 3 shown, determining the current audio segment based on the voice information of each audio frame includes:

[0072] Step S302, taking the first voice frame in the audio signal as the start node of the audio segment, or taking the first voice frame after the end of the previous audio segment as the start node of the audio segment; where the voice frame is an audio frame whose voice information meets the voice activity condition.

[0073] The voice activity condition is a condition for determining whether an audio frame is a voice frame. The voice activity condition can be set as needed. For example, the voice activity condition can be that the voice activity detection algorithm determines that the audio frame is a voice frame, or that the audio frame contains human voices, or that the sound decibel in the audio frame is greater than a specified decibel value, etc., and is not limited thereto.

[0074] Specifically, the computer device sequentially detects whether the speech information of each audio frame satisfies the speech activity condition according to the playing order of the audio signal; when the speech information of the audio frame satisfies the speech activity condition, the audio frame is a speech frame; when the speech information of the audio frame does not satisfy the speech activity condition, the audio frame is a non-speech frame. The computer device takes the first speech frame in the audio signal, or the first speech frame after the end of the previous audio segment, as the start node of the audio segment.

[0075] Step S304, if there are more than a specified number of consecutive non-speech frames after the start node, determine the end node of the audio segment based on the more than a specified number of consecutive non-speech frames; where the non-speech frame is an audio frame whose speech information does not satisfy the speech activity condition.

[0076] The specified number can be set as needed. For example, the specified number is 10 frames, 20 frames, or 50 frames, etc.

[0077] If there are more than a specified number of consecutive non-speech frames after the start node, indicating that the current audio segment has ended, then determine the end node of the audio segment based on the more than a specified number of consecutive non-speech frames. In one implementation, if there are more than a specified number of consecutive non-speech frames after the start node, the computer device takes the last non-speech frame among the more than a specified number of consecutive non-speech frames as the end node of the audio segment. In another implementation, if there are more than a specified number of consecutive non-speech frames after the start node, the computer device takes the first non-speech frame among the more than a specified number of consecutive non-speech frames as the end node of the audio segment. In another implementation, if there are more than a specified number of consecutive non-speech frames after the start node, the computer device randomly determines a non-speech frame from the more than a specified number of consecutive non-speech frames as the end node of the audio segment. In other implementations, the computer device can also use other methods to determine the end node of the audio segment, which is not limited here.

[0078] Step S306, determine the current audio segment based on the start node and the end node.

[0079] Specifically, the computer device forms the current audio segment with all the audio frames between the start node and the end node.

[0080] In another embodiment, the computer device determines each audio frame between the start node and the end node, removes the non-speech frames from each audio frame, and forms the current audio segment with each speech frame.

[0081] In this embodiment, the first speech frame in the audio signal, or the first speech frame after the end of the previous audio segment, is used as the start node of the audio segment; if there are more than a specified number of consecutive non-speech frames after the start node, indicating that the current audio segment has ended, then based on the more than a specified number of consecutive non-speech frames, the end node of the audio segment is determined. Then, based on the start node and the end node, the current audio segment can be accurately determined.

[0082] In one embodiment, if there are more than a specified number of consecutive non-speech frames after the start node, then based on the more than a specified number of consecutive non-speech frames, determining the end node of the audio segment includes: for the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, take the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute until the non-speech count value exceeds the specified number; take the last non-speech frame in the consecutive specified number of non-speech frames as the end node of the audio segment.

[0083] The non-speech count value is a parameter value used to count non-speech frames. The default value of the non-speech count value is zero, that is, the initial value is zero.

[0084] For the audio frames after the start node, if the first non-speech frame is detected, increment the non-speech count value by one, and this non-speech count value is 1; if the next audio frame is also a non-speech frame, continue to increment the non-speech count value by one, and this non-speech count value is 2... If the non-speech count value is less than or equal to the specified number and the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value exceeds the specified number, that is, there are more than a specified number of consecutive non-speech frames after the start node, take the last non-speech frame in the consecutive specified number of non-speech frames as the end node of the audio segment.

[0085] Among them, when the non-speech count value is less than or equal to the specified number and the current audio frame is a speech frame, that is, there are less than or equal to the specified number of consecutive non-speech frames, but this less than or equal to the specified number of consecutive non-speech frames does not end the audio segment, but is a short pause in the middle of the audio segment. Therefore, the non-speech count value is set to zero. For example, a short pause in the middle of a sentence recited in an audio program.

[0086] In this embodiment, for the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, take the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame, and continue to execute until the non-speech count value exceeds the specified number. Then, the last non-speech frame among the consecutive specified number of non-speech frames can be taken as the end node of the audio segment, accurately determining the end node of the audio segment, and thus accurately determining the current audio segment.

[0087] In one embodiment, taking the first speech frame after the end of the previous audio segment as the start node of the audio segment includes: for the currently obtained current audio frame, if the current start marker parameter is the first value and the current audio frame is a speech frame, take the current audio frame as the start node of the audio segment and adjust the start marker parameter from the first value to the second value; wherein, when it is detected that there are consecutive specified number of non-speech frames, the start marker parameter is adjusted from the second value to the first value.

[0088] The start marker parameter is a parameter used to determine whether the current audio frame is the start node. The value of the start marker parameter can be the first value or the second value. When the start marker parameter is the first value, it means that the previous audio segment has ended before the current time point. When the start marker parameter is the second value, it means that the current audio frame is within the current audio segment. Among them, both the first value and the second value can be set as needed. For example, the first value is 0 and the second value is 1.

[0089] When the computer device detects that there are consecutive specified number of non-speech frames, adjusting the start marker parameter from the second value to the first value means that the current audio segment has ended. For the currently obtained current audio frame, if the current start marker parameter is the first value and the current audio frame is a speech frame, it means that the previous audio segment has ended and the obtained current audio frame is a speech frame, which is the first audio frame of the next audio segment. Then, take the current audio frame as the start node of the audio segment and adjust the start marker parameter from the first value to the second value, indicating entering a new audio segment. In this embodiment, for the currently obtained current audio frame, if the current start marker parameter is the first value and the current audio frame is a speech frame, taking the current audio frame as the start node of the audio segment and adjusting the start marker parameter from the first value to the second value can accurately determine that the current speech frame is the first audio frame of a new audio segment, and thus continue to perform audio processing on the new audio segment.

[0090] In one embodiment, if there are consecutive non-speech frames exceeding a specified number after the start node, based on the consecutive non-speech frames exceeding the specified number, determining the end node of the audio segment includes: for the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, take the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute until the non-speech count value is set to zero after exceeding the specified number; take the frame immediately preceding the speech frame that triggers the non-speech count value to be set to zero as the end node of the audio segment.

[0091] For the audio frames after the start node, if the first non-speech frame is detected, increment the non-speech count value by one, and this non-speech count value is 1; if the next audio frame is also a non-speech frame, continue to increment the non-speech count value by one, and this non-speech count value is 2... If the non-speech count value is less than or equal to the specified number and the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, continue to the next audio frame until the non-speech count value is set to zero after exceeding the specified number, and take the frame immediately preceding the speech frame that triggers the non-speech count value to be set to zero as the end node of the audio segment.

[0092] In this embodiment, for the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, take the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute. If the non-speech count value is set to zero after exceeding the specified number, then the frame immediately preceding the speech frame that triggers the non-speech count value to be set to zero can be taken as the end node of the audio segment, accurately determining the end node of the audio segment, and thus accurately determining the current audio segment.

[0093] In one embodiment, the above method further includes: if there are consecutive non-speech frames exceeding a specified number after the start node, take the first speech frame that appears after the consecutive non-speech frames exceeding the specified number as the start node of the next audio segment.

[0094] If there are more than a specified number of consecutive non-speech frames after the start node, it indicates that the current audio segment has ended, and continue to obtain each audio frame in sequence. When the first speech frame appears after more than a specified number of consecutive non-speech frames, then this first speech frame is the start of the next audio segment. As the start node of the next audio segment, the start node of the next audio segment can be accurately determined, and continue to accurately process the next audio segment.

[0095] In another embodiment, as Figure 4 shown, an audio processing method is provided, including the following steps:

[0096] Step S402, obtain each audio frame in the audio signal in sequence, and detect the speech information of each audio frame.

[0097] Step S404, use the first speech frame in the audio signal as the start node of the audio segment, or use the first speech frame after the end of the previous audio segment as the start node of the audio segment; wherein, the speech frame is an audio frame whose speech information meets the speech activity condition.

[0098] Step S406, for the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one.

[0099] Step S408, if the non-speech count value exceeds the specified number, adjust the start flag parameter from the second value to the first value.

[0100] Step S410, if there are more than a specified number of consecutive non-speech frames after the start node, determine the end node of the audio segment based on the more than a specified number of consecutive non-speech frames; wherein, the non-speech frame is an audio frame whose speech information does not meet the speech activity condition.

[0101] Step S412, determine the current audio segment based on the start node and the end node.

[0102] Step S414, obtain the speech information density and the speech information amount of the current audio segment in the audio signal.

[0103] Step S416, determine the target pause duration of the current audio segment based on the speech information density and the speech information amount.

[0104] Step S418, if it is detected that the current audio frame is a speech frame and the current start flag parameter is the second value, use the detected speech frame as the start node of the next audio segment.

[0105] It can be understood that if the current audio frame is the start node of an audio segment, the start marker parameter is adjusted from the first value to the second value. If the non-speech count value exceeds the specified number, indicating that the current audio segment has ended, the start marker parameter is adjusted from the second value to the first value. Then, when it is detected that the current audio frame is a speech frame and the start marker parameter is the first value, it can be considered that the start marker parameter before the current audio frame is the first value, that is, the start marker parameter from after the end of the previous audio segment to before the current audio frame is the first value. Then the current audio frame, which is a speech frame, is also the start node of the next audio segment. Therefore, the start marker parameter can be adjusted from the first value to the second value again, indicating the start of a new audio segment.

[0106] Step S420: Obtain the speech interval duration between the current audio segment and the subsequent audio segment in the audio signal.

[0107] Step S422: If the speech interval duration is less than the target pause duration, insert a pause segment between the current audio segment and the subsequent audio segment.

[0108] In this embodiment, for the audio frames after the start node, if the current audio frame is a non-speech frame, the non-speech count value is incremented by one; if the non-speech count value reaches or exceeds the specified number, the start marker parameter is adjusted from the second value to the first value. Then, if it is detected that the current audio frame is a speech frame and the current start marker parameter is the second value, it can be considered that the current audio frame is the start node of the next audio segment. Then, it can be statistically determined whether the speech interval duration between the current audio segment and the subsequent audio segment is greater than the target pause duration. If it is less than the target pause duration, a pause segment is inserted. The start node of the next audio segment can be accurately determined, and a pause segment is inserted before the start of the next audio segment, and then the next audio segment is processed continuously, which can improve the accuracy of audio processing.

[0109] In one embodiment, obtaining the speech information density and speech information volume of the current audio segment in the audio signal includes: obtaining the current audio segment in the audio signal; performing pitch frequency detection on each audio frame in the current audio segment to obtain the number of pitch frequency fluctuations of the current audio segment; the number of pitch frequency fluctuations characterizes the speech information volume; determining the speech duration of the current audio segment; and determining the speech information density of the current audio segment based on the comparison value between the number of pitch frequency fluctuations and the speech duration.

[0110] It can be understood that generally, sound is composed of a series of vibrations with different frequencies and amplitudes emitted by a sounding body. Among these vibrations, there is a vibration with the lowest frequency, and the sound emitted by it is the fundamental tone, and the rest are overtones.

[0111] The number of fluctuations in the fundamental frequency refers to the number of times the fundamental frequency fluctuates in the current audio segment. For example, if the fundamental frequency in the current audio segment changes from f1 to f2 and then from f2 to f3 in sequence, the number of fluctuations in the fundamental frequency of the current audio segment is 2. It can be understood that the more the number of fluctuations in the fundamental frequency, the more speech information the current audio segment contains, that is, the more speech information volume. The higher the frequency of fluctuations in the fundamental frequency per unit time, the higher the speech information density of the current audio segment. On the contrary, if the frequency of fluctuations in the fundamental frequency is lower, it means that the speech information density of the current audio segment is lower.

[0112] Specifically, the computer device uses the fundamental frequency detection technology to detect the fundamental frequency of each audio frame in the current audio segment of the audio signal, obtains the number of fluctuations in the fundamental frequency of the current audio segment, and uses a timer to count the duration of the current audio segment to obtain the speech duration of the current audio segment.

[0113] In one implementation, the computer device divides the number of fluctuations in the fundamental frequency by the speech duration to obtain the speech information density of the current audio segment. In another implementation, the computer device divides the speech duration by the number of fluctuations in the fundamental frequency to obtain the speech information density of the current audio segment. In other implementations, the computer device can also respectively obtain the weight factors of the number of fluctuations in the fundamental frequency and the speech duration, multiply the number of fluctuations in the fundamental frequency and the speech duration by the corresponding weight factors respectively, and then obtain the comparison value between the two products to determine the speech information density of the current audio segment. The specific method of calculating the speech information density can be set according to needs and is not limited here.

[0114] Specifically, the computer device takes an audio frame as a unit of time, counts the number of audio frames in the current audio segment, and uses the number of all audio frames in the current audio segment as the speech duration. In another embodiment, the computer device uses a timer to count the speech duration of each audio frame in the current audio segment.

[0115] In this embodiment, the computer device detects the fundamental frequency of each audio frame in the current audio segment of the audio signal, can accurately obtain the number of fluctuations in the fundamental frequency of the current audio segment, and this number of fluctuations in the fundamental frequency characterizes the speech information volume; determines the speech duration of the current audio segment; and then based on the comparison value between the number of fluctuations in the fundamental frequency and the speech duration, can accurately determine the speech information density of the current audio segment.

[0116] In one embodiment, as Figure 5 shown, detecting the fundamental frequency of each audio frame in the current audio segment to obtain the number of fluctuations in the fundamental frequency of the current audio segment includes:

[0117] Step S502: Perform pitch frequency detection on each audio frame in the current audio segment to obtain the pitch frequency value of each speech frame; wherein, a speech frame is an audio frame whose speech information meets the speech activity condition.

[0118] The pitch frequency value refers to the lowest frequency value in the audio frame containing speech information, that is, the pitch frequency value.

[0119] It can be understood that if an audio frame does not meet the speech activity condition, that is, the audio frame is a non-speech frame and does not contain speech information, then the pitch frequency value of the audio frame is 0, and the pitch frequency detection of the audio frame can be omitted, saving computer resources and improving the efficiency of pitch frequency detection for the current audio segment.

[0120] The computer device performs pitch frequency detection on each audio frame in the current audio segment of the audio signal, determines each speech frame that meets the speech activity condition, and obtains the lowest frequency value from each speech frame as the pitch frequency value of the corresponding audio frame.

[0121] Step S504: Based on the pitch frequency values between adjacent speech frames, determine the pitch frequency fluctuation times of the current audio segment.

[0122] Specifically, the computer device compares the pitch frequency values between adjacent speech frames. If the comparison result meets the pitch frequency fluctuation condition, it represents a pitch frequency fluctuation until the pitch frequency values between all adjacent speech frames are detected, and the pitch frequency fluctuation times of the current audio segment are obtained.

[0123] Among them, the pitch frequency fluctuation condition is a condition used to judge whether the pitch frequency values between adjacent speech frames constitute a pitch frequency fluctuation. The pitch frequency fluctuation condition can be set as needed. For example, the pitch frequency fluctuation condition can be that the pitch frequency values between adjacent speech frames are different, or the difference between the pitch frequency values between adjacent speech frames is greater than a preset threshold, which is not limited here.

[0124] In this embodiment, perform pitch frequency detection on each audio frame in the current audio segment of the audio signal to obtain the pitch frequency value of each audio frame containing speech information. Then, based on the pitch frequency values between adjacent audio frames containing speech information, the pitch frequency fluctuation times of the current audio segment can be accurately determined.

[0125] In one embodiment, as Figure 6 shown, determining the pitch frequency fluctuation times of the current audio segment based on the pitch frequency values between adjacent speech frames includes:

[0126] Step S602: Based on the pitch frequency values between adjacent speech frames, determine the pitch frequency states corresponding to the adjacent speech frames respectively.

[0127] The pitch frequency state refers to the state of the frequency values between adjacent speech frames. Specifically, the pitch frequency state may include an ascending state, a flat state, and a descending state.

[0128] Specifically, the computer device compares the pitch frequency values between adjacent speech frames. If the difference f(i) - f(i - 1) obtained by subtracting the pitch frequency value f(i - 1) of the previous audio frame i - 1 from the pitch frequency value f(i) of the current audio frame i is greater than the frequency state threshold THRD_F, it is determined that the pitch frequency state between the current audio frame and the previous audio frame is in the ascending state; if the difference f(i - 1) - f(i) obtained by subtracting the pitch frequency value f(i) of the current audio frame i from the pitch frequency value f(i - 1) of the previous audio frame i - 1 is greater than the frequency state threshold THRD_F, it is determined that the pitch frequency state between the current audio frame and the previous audio frame is in the descending state; if the difference between the pitch frequency value f(i) of the current audio frame i and the pitch frequency value f(i - 1) of the previous audio frame i - 1 is less than or equal to the frequency state threshold THRD_F, it is determined that the pitch frequency state between the current audio frame and the previous audio frame is in the flat state. Among them, the frequency state threshold THRD_F can be set as needed.

[0129] Step S604, if the pitch frequency states of two speech frames in adjacent speech frames are different, it is determined that there is a pitch frequency fluctuation.

[0130] Step S606, based on all the pitch frequency fluctuations that occur in the current audio segment, the pitch frequency fluctuation times of the current audio segment are statistically obtained.

[0131] The computer device obtains the pitch frequency states of each in the current audio segment, compares the pitch frequency states, and if the pitch frequency states of two speech frames in adjacent speech frames are different, that is, the adjacent pitch frequency states change, it is determined that there is a pitch frequency fluctuation. Then, the computer device can compare the adjacent pitch frequency states in turn, so as to statistically obtain all the pitch frequency fluctuations that occur in the current audio segment and obtain the pitch frequency fluctuation times of the current audio segment.

[0132] For example, if the current pitch frequency state changes from the ascending state to the descending state, it represents a pitch frequency fluctuation. Another example is that if the current pitch frequency state changes from the flat state to the ascending state, it represents a pitch frequency fluctuation.

[0133] In this embodiment, based on the pitch frequency values between adjacent speech frames, the pitch frequency states corresponding to the adjacent speech frames are determined; if the pitch frequency states of two speech frames in the adjacent speech frames are different, it is determined that there is a pitch frequency fluctuation; then, based on all the pitch frequency fluctuations that occur in the current audio segment, the number of pitch frequency fluctuations in the current audio segment can be accurately counted.

[0134] In one embodiment, based on the speech information density and the speech information quantity, the target pause duration of the current audio segment is determined, including: inputting the speech information density and the speech information quantity into a trained neural network model, and outputting the target pause duration through the trained neural network model.

[0135] The neural network model can specifically be a convolutional neural network model, a residual shrinkage network model, a multi-layer perceptron neural network model, etc.

[0136] Among them, as Figure 7 shown, the training steps of the neural network model include:

[0137] Step S702, obtaining a sample audio signal; the sample speech interval duration between adjacent sample audio segments in the sample audio signal is within a preset duration range.

[0138] The sample audio signal refers to the audio signal used as a sample to train the neural network model. The sample audio segment is an audio segment in the sample audio signal. The preset duration range can be set as needed.

[0139] In the sample audio signal, the sample speech interval duration between adjacent sample audio segments is within a preset duration range, that is, the sample speech interval duration between adjacent sample audio segments is long enough for the listener to understand the content of the previous sample audio segment.

[0140] Step S704, determining the sample speech information density and the sample speech information quantity of each sample audio segment in the sample audio signal.

[0141] Specifically, the computer device uses pitch frequency detection technology to detect the sample speech information density and the sample speech information quantity of each sample audio segment in the sample audio signal.

[0142] Step S706, using the sample speech information density and the sample speech information quantity of the sample audio segment as the training input, and using the sample speech interval duration corresponding to the corresponding sample audio segment as the training label.

[0143] Step S708, training the neural network model based on the training input and the training label until the training stop condition is reached and then stopping, to obtain the trained neural network model.

[0144] The sample speech interval duration corresponding to the sample audio segment is used as the training label, that is, this sample speech interval duration is used as the reasonable pause duration of the corresponding sample audio segment. The training stop condition can be set as needed. For example, the training stop condition can be that the training duration reaches a preset duration, or the number of training times reaches a preset number, or the difference between the training output and the training label is less than a preset threshold, and it is not limited to this. The training output refers to the training output obtained by the training model based on the training input, that is, the training speech interval duration determined based on the sample speech information density and the sample speech information amount of the sample audio segment.

[0145] The computer device uses the sample speech information density and the sample speech information amount of the sample audio segment as the training input, and uses the sample speech interval duration corresponding to the corresponding sample audio segment as the training label. Then, the neural network model determines the training speech interval duration of the corresponding sample audio segment based on the sample speech information density and the sample speech information amount of the sample audio segment through the function func(fc, p), compares this training speech interval duration with the sample speech interval duration, and adjusts the parameters of the neural network model based on the comparison result, and continues the next round of training until it stops when the training stop condition is reached, and a trained neural network model is obtained. Among them, the function func can be implemented by statistical methods or neural network methods.

[0146] In this embodiment, the computer device acquires the sample audio signal, and a more accurate neural network model can be trained based on the sample audio signal. Through this trained neural network model, a more accurate target pause duration can be output based on the speech information density and the speech information amount.

[0147] In one embodiment, if the speech interval duration is less than the target pause duration, a pause segment is inserted between the current audio segment and the subsequent audio segment, including: if the speech interval duration is less than the target pause duration, the inserted pause duration is determined based on the target pause duration and the speech interval duration; the sum of the inserted pause duration and the speech interval duration is greater than or equal to the target pause duration; the pause segment with the inserted pause duration is acquired, and the pause segment is inserted between the current audio segment and the subsequent audio segment.

[0148] The pause segment refers to the audio segment that needs to be inserted between the current audio segment and the subsequent audio segment and is used for pausing. The inserted pause duration refers to the duration of the pause segment.

[0149] Specifically, the computer device obtains the voice interval duration and the target pause duration. If the voice interval duration is less than the target pause duration, the target pause duration is subtracted by the voice interval duration to obtain the inserted pause duration. The comfort noise generation (cng) technology is used to obtain the pause segment of the inserted pause duration, and the pause segment is inserted between the current audio segment and the subsequent audio segment. Among them, the pause segment can be a silent segment or a small noise segment that does not contain any voice information.

[0150] In this embodiment, if the voice interval duration is less than the target pause duration, the inserted pause duration is determined based on the target pause duration and the voice interval duration, and the sum of the inserted pause duration and the voice interval duration is greater than or equal to the target pause duration. Then, the pause segment of the inserted pause duration is inserted between the current audio segment and the subsequent audio segment, which can not only have sufficient pause duration after the current audio segment but also avoid inserting too many audio segments, thus achieving the balance between the accuracy of audio processing and computer resources.

[0151] In one embodiment, the audio signal includes pre-recorded audio content. The above method further includes: continuing to detect the voice interval duration and insert the pause segment for the subsequent audio segment in the audio signal until the voice interval duration between all adjacent audio segments in the audio signal is adjusted to the target pause duration to obtain the target audio; playing the target audio, and there is a pause effect of the target pause duration between different content segments of the audio content during the playing of the target audio.

[0152] The target audio is the audio content obtained after adjusting the voice interval duration of each audio segment in the audio signal to the target pause duration.

[0153] After the computer device finishes processing the current audio segment, it continues to detect the voice interval duration and insert the pause segment for the subsequent audio segment until the voice interval duration between all adjacent audio segments in the audio signal is adjusted to the target pause duration to obtain the target audio. Then, during the playing of the target audio, there is a corresponding target pause duration after each content segment for the listener to understand and absorb the content segment, effectively conveying the audio content of the content segment and improving the effectiveness of audio content conveyance.

[0154] In another embodiment, as Figure 8 shown, the computer device obtains the audio signal and performs vad detection on the current audio segment in the audio signal, that is, voice activity detection, to determine whether the vad value of each audio frame is 0. If the vad of the audio frame is 0, it means that the audio frame is a non-voice frame. If the vad of the audio frame is not 0, it means that the audio frame is a voice frame.

[0155] When the VAD of the audio frame is 0, the non_VAD_cnt is incremented, that is, the non-speech count value is incremented. Determine whether the non-speech count value non_VAD_cnt exceeds the specified quantity THRD. If non_VAD_cnt > THRD, it indicates that the current audio segment has ended, then the start flag parameter Startflag is adjusted from the second value 1 to the first value 0 to end the parameter statistics of the current audio segment. At the same time, the computer device calculates the speech information density p = fc / t based on the number of pitch frequency fluctuations fc and the speech duration t; where the number of pitch frequency fluctuations characterizes the speech information volume; then based on the speech information density and the speech information volume, the target pause duration T of the current audio segment is determined. If non_VAD_cnt <= THRD, it indicates that the current audio segment has not ended, and the parameter statistics continue, including: performing pitch frequency detection on the audio frame to obtain the number of pitch frequency fluctuations fc, and counting the speech duration t of the current audio segment.

[0156] When the VAD is not 0, determine whether the start flag parameter Startflag is the second value 1. If the start flag parameter Startflag is not the second value 1, that is, the start flag parameter Startflag is the first value 0, it means that a new audio segment has not started before the current audio frame, and the current audio frame is a speech frame, indicating that the current audio frame is the first audio frame of a new audio segment. Then, determine whether the non-speech count value non_VAD_cnt is greater than the target pause duration T. If the non-speech count value non_VAD_cnt is greater than the target pause duration T, it means that the duration of the non-speech frames after the current audio segment is long enough for the listener to understand the content of the current audio segment. Therefore, the start flag parameter Startflag is adjusted from the first value 0 to the second value 1, and the parameter statistics of the new audio segment are started. At the same time, the number of pitch frequency fluctuations fc and the speech duration t are cleared, and the parameter statistics for the new audio segment are counted. If the non-speech count value non_VAD_cnt is less than or equal to the target pause duration T, that is, the speech interval duration between the current audio segment and the subsequent audio segment is less than the target pause duration T, then a pause segment is inserted between the current audio segment and the subsequent audio segment, and the start flag parameter Startflag is adjusted from the first value 0 to the second value 1, and the parameter statistics of the new audio segment are started.

[0157] If the VAD is not 0 and the start flag parameter Startflag is the second value 1, it means that the current audio segment has not ended, and the current audio frame is the audio frame in the middle of the current audio segment. Therefore, the non-speech count value non_VAD_cnt is cleared, and the parameter statistics continue.

[0158] In one embodiment, another audio processing method is provided. This method is applied to a computer device, and the computer device can be a terminal or a server, including the following steps:

[0159] Step 1: Obtain each audio frame in the audio signal in sequence, and detect the speech information of each audio frame.

[0160] Step 2: Use the first speech frame in the audio signal as the start node of the audio segment, or for the currently obtained current audio frame, if the current start flag parameter is the first value and the current audio frame is a speech frame, then use the current audio frame as the start node of the audio segment, and adjust the start flag parameter from the first value to the second value; wherein, when it is detected that there are a specified number of consecutive non-speech frames, the start flag parameter is adjusted from the second value to the first value; wherein, a speech frame is an audio frame whose speech information meets the speech activity condition.

[0161] Step 3: For the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one; if the non-speech count value exceeds the specified number, adjust the start flag parameter from the second value to the first value.

[0162] After the computer device executes Step 3, it executes an end node determination step, and this end node determination step is Step 4a or Step 4b.

[0163] Step 4a: For the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one; wherein, a non-speech frame is an audio frame whose speech information does not meet the speech activity condition; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, use the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute until the non-speech count value exceeds the specified number; use the last non-speech frame among the consecutive specified number of non-speech frames as the end node of the audio segment.

[0164] Step 4b: For the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one; wherein, a non-speech frame is an audio frame whose speech information does not meet the speech activity condition; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, use the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute until the non-speech count value is set to zero after exceeding the specified number; use the previous frame before the speech frame that triggers the non-speech count value to be set to zero as the end node of the audio segment.

[0165] Step 5: Determine the current audio segment based on the start node and the end node.

[0166] Step 6: Obtain the current audio segment in the audio signal; perform pitch frequency detection on each audio frame in the current audio segment to obtain the pitch frequency values of each speech frame; where a speech frame is an audio frame whose speech information meets the speech activity condition.

[0167] Step 7: Based on the pitch frequency values between adjacent speech frames, determine the pitch frequency states corresponding to the adjacent speech frames; if the pitch frequency states of two speech frames in adjacent speech frames are different, it is determined that there is a pitch frequency fluctuation; based on all the pitch frequency fluctuations occurring in the current audio segment, count the number of pitch frequency fluctuations in the current audio segment; the number of pitch frequency fluctuations characterizes the speech information volume.

[0168] Step 8: Determine the speech duration of the current audio segment.

[0169] Step 9: Based on the comparison value between the number of pitch frequency fluctuations and the speech duration, determine the speech information density of the current audio segment; the speech information density is used to measure the frequency of speech information fluctuations.

[0170] Step 10: Input the speech information density and the speech information volume into the trained neural network model, and output the target pause duration through the trained neural network model; where the training steps of the neural network model include: obtaining a sample audio signal; the sample speech interval duration between adjacent sample audio segments in the sample audio signal is within a preset duration range; determining the sample speech information density and the sample speech information volume of each sample audio segment in the sample audio signal; using the sample speech information density and the sample speech information volume of the sample audio segment as training inputs, and using the corresponding sample speech interval duration of the sample audio segment as a training label; training the neural network model based on the training inputs and the training labels until the training stop condition is reached and then stopping to obtain the trained neural network model.

[0171] Step 11: If it is detected that the current audio frame is a speech frame and the current start marker parameter is the second value, use the detected speech frame as the start node of the next audio segment.

[0172] Step 12: Obtain the speech interval duration between the current audio segment and the subsequent audio segment in the audio signal.

[0173] Step 13: If the speech interval duration is less than the target pause duration, determine the inserted pause duration based on the target pause duration and the speech interval duration; the sum of the inserted pause duration and the speech interval duration is greater than or equal to the target pause duration; obtain the pause segment of the inserted pause duration and insert the pause segment between the current audio segment and the subsequent audio segment.

[0174] Step 14: Continue to detect the speech interval duration of the subsequent audio segments in the audio signal and insert pause segments until the speech interval duration between all adjacent audio segments in the audio signal is adjusted to the target pause duration, obtaining the target audio; play the target audio, and during the playback of the target audio, there is a pause effect with the target pause duration between different content segments of the audio content.

[0175] For the above audio processing method, the speech information density and speech information volume of the current audio segment in the audio signal are obtained. Then, based on the speech information density and speech information volume, the target pause duration for which the current audio segment needs to pause can be accurately determined. If the speech interval duration between the current audio segment and the subsequent audio segment is less than the target pause duration, a pause segment is inserted between the current audio segment and the subsequent audio segment, so that a longer pause can be made after the current audio segment, and the pause duration of the current audio segment can be more accurately regulated, allowing the listener to have more time to understand and absorb the content of the current audio segment, effectively conveying the audio content of the current audio segment, and improving the effectiveness of audio content transmission.

[0176] The present application also provides an application scenario that applies the above audio processing method. Specifically, the application of the audio processing method in this application scenario is as follows:

[0177] The computer device obtains the audio program specified by the user, obtains the audio signal of the audio program, plays the audio signal, and determines the current audio segment during the playback, calculates the target pause duration of the current audio segment. If the speech pause duration between the current audio segment and the subsequent audio segment is less than the target pause duration, a pause segment is inserted, so that the user has enough time to understand the content of the current audio segment.

[0178] The present application also additionally provides an application scenario that applies the above audio processing method.

[0179] Specifically, the application of the audio processing method in this application scenario is as follows:

[0180] The computer device obtains the audio recorded by the user, obtains the audio signal of the recorded audio, plays the audio signal, and determines the current audio segment during the playback, calculates the target pause duration of the current audio segment. If the speech pause duration between the current audio segment and the subsequent audio segment is less than the target pause duration, a pause segment is inserted, so that the user has enough time to understand the content of the current audio segment.

[0181] It should be understood that although Figures 2 to 8The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 2 to 8 At least some of the steps may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.

[0182] In one embodiment, as Figure 9 shown, an audio processing device is provided. This device can be a software module, a hardware module, or a combination of both to form part of a computer device. Specifically, the device includes: an acquisition module 902, a determination module 904, and an insertion module 906, where:

[0183] The acquisition module 902 is configured to acquire the speech information density and the speech information amount of the current audio segment in the audio signal; the speech information density is used to measure the frequency of speech information fluctuations.

[0184] The determination module 904 is configured to determine the target pause duration of the current audio segment based on the speech information density and the speech information amount.

[0185] The acquisition module 902 is further configured to acquire the speech interval duration between the current audio segment and the subsequent audio segment in the audio signal.

[0186] The insertion module 906 is configured to insert a pause segment between the current audio segment and the subsequent audio segment if the speech interval duration is less than the target pause duration.

[0187] For the above audio processing device, by acquiring the speech information density and the speech information amount of the current audio segment in the audio signal, the target pause duration required for the current audio segment can be accurately determined based on the speech information density and the speech information amount. If the speech interval duration between the current audio segment and the subsequent audio segment is less than the target pause duration, a pause segment is inserted between the current audio segment and the subsequent audio segment, so that a longer pause can be made after the current audio segment, and the pause duration of the current audio segment can be more accurately regulated, allowing the listener to have more time to understand and absorb the content of the current audio segment, effectively conveying the audio content of the current audio segment, and improving the effectiveness of audio content transmission.

[0188] In one embodiment, the determining module 904 is further configured to sequentially obtain each audio frame in the audio signal and detect the speech information of each audio frame; and determine the current audio segment based on the speech information of each audio frame.

[0189] In one embodiment, the determining module 904 is further configured to use the first speech frame in the audio signal as the start node of the audio segment, or use the first speech frame after the end of the previous audio segment as the start node of the audio segment; wherein, the speech frame is an audio frame whose speech information meets the speech activity condition; if there are more than a specified number of consecutive non-speech frames after the start node, then based on the more than a specified number of consecutive non-speech frames, determine the end node of the audio segment; wherein, the non-speech frame is an audio frame whose speech information does not meet the speech activity condition; determine the current audio segment based on the start node and the end node.

[0190] In one embodiment, for the currently obtained current audio frame, if the current start flag parameter is the first value and the current audio frame is a speech frame, the determining module 904 is further configured to use the current audio frame as the start node of the audio segment and adjust the start flag parameter from the first value to the second value; wherein, when it is detected that there are a specified number of consecutive non-speech frames, the start flag parameter is adjusted from the second value to the first value.

[0191] In one embodiment, for the audio frames after the start node, if the current audio frame is a non-speech frame, the determining module 904 is further configured to increment the non-speech count value by one; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, use the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute until the non-speech count value exceeds the specified number; use the last non-speech frame in the consecutive specified number of non-speech frames as the end node of the audio segment.

[0192] In one embodiment, for the audio frames after the start node, if the current audio frame is a non-speech frame, the determining module 904 is further configured to increment the non-speech count value by one; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, use the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute until the non-speech count value is set to zero after exceeding the specified number; use the previous frame before the speech frame that triggers the non-speech count value to be set to zero as the end node of the audio segment.

[0193] In one embodiment, the above-mentioned determination module 904 is further configured to, if there are consecutive non-speech frames exceeding a specified number after the start node, use the first speech frame that appears after the consecutive non-speech frames exceeding the specified number as the start node of the next audio segment.

[0194] In one embodiment, the above-mentioned apparatus further includes an adjustment module. The adjustment module is configured to, for the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one; if the non-speech count value exceeds the specified number, adjust the start flag parameter from the second value to the first value; the above-mentioned determination module 904 is further configured to, if it is detected that the current audio frame is a speech frame and the current start flag parameter is the second value, use the detected speech frame as the start node of the next audio segment.

[0195] In one embodiment, the above-mentioned acquisition module 902 is further configured to acquire the current audio segment in the audio signal; perform pitch frequency detection on each audio frame in the current audio segment to obtain the number of pitch frequency fluctuations in the current audio segment; the number of pitch frequency fluctuations characterizes the speech information amount; determine the speech duration of the current audio segment; based on the comparison value between the number of pitch frequency fluctuations and the speech duration, determine the speech information density of the current audio segment.

[0196] In one embodiment, the above-mentioned acquisition module 902 is further configured to perform pitch frequency detection on each audio frame in the current audio segment in the audio signal to obtain the pitch frequency values of each speech frame; based on the pitch frequency values between adjacent speech frames, determine the number of pitch frequency fluctuations in the current audio segment.

[0197] In one embodiment, the above-mentioned acquisition module 902 is further configured to, based on the pitch frequency values between adjacent speech frames, determine the pitch frequency states corresponding to the adjacent speech frames respectively; if the pitch frequency states of two speech frames in the adjacent speech frames are different, determine that there is one pitch frequency fluctuation; based on all the pitch frequency fluctuations that occur in the current audio segment, statistically obtain the number of pitch frequency fluctuations in the current audio segment.

[0198] In one embodiment, the above-mentioned determination module 904 is further configured to input the voice information density and the voice information quantity into the trained neural network model, and output the target pause duration through the trained neural network model; the above-mentioned apparatus further includes a training module, and the training module is configured to obtain a sample audio signal; the sample voice interval duration between adjacent sample audio segments in the sample audio signal is within a preset duration range; determine the sample voice information density and the sample voice information quantity of each sample audio segment in the sample audio signal; use the sample voice information density and the sample voice information quantity of the sample audio segment as training inputs, and use the sample voice interval duration corresponding to the corresponding sample audio segment as a training label; train the neural network model based on the training inputs and the training labels until the training stop condition is reached and then stop, to obtain the trained neural network model.

[0199] In one embodiment, the above-mentioned insertion module 906 is further configured to, if the voice interval duration is less than the target pause duration, determine the insertion pause duration based on the target pause duration and the voice interval duration; the sum of the insertion pause duration and the voice interval duration is greater than or equal to the target pause duration; obtain the pause segment of the insertion pause duration, and insert the pause segment between the current audio segment and the subsequent audio segment.

[0200] In one embodiment, the audio signal includes pre-recorded audio content; the above-mentioned insertion module is further configured to continue to detect the voice interval duration and insert the pause segment for the subsequent audio segment in the audio signal until the voice interval duration between all adjacent audio segments in the audio signal is adjusted to the target pause duration to obtain the target audio; the above-mentioned apparatus further includes a playback module; the playback module is configured to play the target audio, and during the playback of the target audio, there is a pause effect with the target pause duration between different content segments of the audio content.

[0201] For the specific limitations of the audio processing apparatus, reference can be made to the limitations of the audio processing method in the above text, which will not be elaborated here. Each module in the above audio processing apparatus can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0202] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 10As shown in the figure. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an audio processing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0203] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0204] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in the above method embodiments are implemented.

[0205] In one embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0206] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0207] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0208] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0209] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. An audio processing method, characterized in that, The method includes: Obtaining the speech information density and speech information volume of the current audio segment in the audio signal; the speech information density is used to measure the frequency of speech information fluctuations; the speech information volume is characterized by the number of pitch frequency fluctuations of the current audio segment, and the number of pitch frequency fluctuations characterizes the number of changes in the pitch frequency state of adjacent speech frames in the current audio segment; Inputting the speech information density and the speech information volume into a trained neural network model, and outputting a target pause duration through the trained neural network model; Obtaining the speech interval duration between the current audio segment and the subsequent audio segment in the audio signal; If the speech interval duration is less than the target pause duration, inserting a pause segment between the current audio segment and the subsequent audio segment.

2. The method according to claim 1, wherein Before obtaining the speech information density and the speech information volume of the current audio segment in the audio signal, the method further includes: Sequentially obtaining each audio frame in the audio signal and detecting the speech information of each audio frame; Determining the current audio segment based on the speech information of each audio frame.

3. The method according to claim 2, characterized in that, The determining the current audio segment based on the speech information of each audio frame includes: Taking the first speech frame in the audio signal as the start node of the audio segment, or taking the first speech frame after the end of the previous audio segment as the start node of the audio segment; wherein, the speech frame is an audio frame whose speech information meets the speech activity condition; If there are more than a specified number of consecutive non-speech frames after the start node, determining the end node of the audio segment based on the consecutive non-speech frames that exceed the specified number; wherein, the non-speech frame is an audio frame whose speech information does not meet the speech activity condition; Determining the current audio segment based on the start node and the end node.

4. The method according to claim 3, wherein The taking the first speech frame after the end of the previous audio segment as the start node of the audio segment includes: For the currently obtained current audio frame, if the current start marker parameter is the first value and the current audio frame is a speech frame, taking the current audio frame as the start node of the audio segment and adjusting the start marker parameter from the first value to the second value; wherein, in the case where it is detected that there are a specified number of consecutive non-speech frames, the start marker parameter is adjusted from the second value to the first value.

5. The method according to claim 3, characterized in that, The if there are more than a specified number of consecutive non-speech frames after the start node, determining the end node of the audio segment based on the consecutive non-speech frames that exceed the specified number includes: For the audio frames after the start node, if the current audio frame is a non-speech frame, incrementing the non-speech count value by one; If the current audio frame is a speech frame, setting the non-speech count value to zero; If the non-speech count value is less than or equal to the specified number, taking the next audio frame as the current audio frame and returning to the step of if the current audio frame is a non-speech frame, incrementing the non-speech count value by one and continuing to execute until the non-speech count value exceeds the specified number; Taking the last non-speech frame in the consecutive non-speech frames that exceed the specified number as the end node of the audio segment.

6. The method according to claim 3, wherein If there are consecutive non-speech frames exceeding a specified number after the start node, determining an end node of an audio segment based on the consecutive non-speech frames exceeding the specified number includes: For an audio frame after the start node, if the current audio frame is a non-speech frame, increment a non-speech count value by one; If the current audio frame is a speech frame, set the non-speech count value to zero; If the non-speech count value is less than or equal to the specified number, take the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame, and continue to execute until the non-speech count value is set to zero after exceeding the specified number; Take the previous frame before the speech frame that triggers the non-speech count value to be set to zero as the end node of the audio segment.

7. The method according to claim 3, characterized in that The method further includes: If there are consecutive non-speech frames exceeding a specified number after the start node, take the first speech frame that appears after the consecutive non-speech frames exceeding the specified number as the start node of the next audio segment.

8. The method according to claim 1, wherein Obtaining the speech information density and speech information quantity of the current audio segment in the audio signal includes: Obtaining the current audio segment in the audio signal; Performing pitch frequency detection on each audio frame in the current audio segment to obtain the number of pitch frequency fluctuations of the current audio segment; the number of pitch frequency fluctuations characterizes the speech information quantity; Determining the speech duration of the current audio segment; Based on a comparison value between the number of pitch frequency fluctuations and the speech duration, determining the speech information density of the current audio segment.

9. The method according to claim 8, wherein The performing pitch frequency detection on each audio frame in the current audio segment to obtain the number of pitch frequency fluctuations of the current audio segment includes: Performing pitch frequency detection on each audio frame in the current audio segment to obtain the pitch frequency value of each speech frame; wherein, the speech frame is an audio frame whose speech information meets the speech activity condition; Based on the pitch frequency values between adjacent speech frames, determining the number of pitch frequency fluctuations of the current audio segment.

10. The method according to claim 9, characterized in that, The based on the pitch frequency values between adjacent speech frames, determining the number of pitch frequency fluctuations of the current audio segment includes: Based on the pitch frequency values between adjacent speech frames, determining the respective pitch frequency states corresponding to the adjacent speech frames; If the pitch frequency states of two speech frames in the adjacent speech frames are different, determine that there is one pitch frequency fluctuation; Based on all the pitch frequency fluctuations that occur in the current audio segment, statistically obtaining the number of pitch frequency fluctuations of the current audio segment.

11. The method according to claim 1, characterized in that The training steps of the neural network model include: Obtaining a sample audio signal; the sample speech interval duration between adjacent sample audio segments in the sample audio signal is within a preset duration range; Determining the sample speech information density and sample speech information quantity of each sample audio segment in the sample audio signal; Taking the sample speech information density and sample speech information quantity of the sample audio segment as training inputs, and taking the sample speech interval duration corresponding to the respective sample audio segment as a training label; Train the neural network model based on the training input and the training label until the training stop condition is reached, and then stop to obtain the trained neural network model.

12. The method according to claim 1, wherein The step of inserting a pause segment between the current audio segment and the subsequent audio segment if the voice interval duration is less than the target pause duration includes: If the voice interval duration is less than the target pause duration, determine the inserted pause duration based on the target pause duration and the voice interval duration; the sum of the inserted pause duration and the voice interval duration is greater than or equal to the target pause duration; Obtain the pause segment with the inserted pause duration, and insert the pause segment between the current audio segment and the subsequent audio segment.

13. The method according to any one of claims 1 to 12, characterized in that, The audio signal includes pre-recorded audio content, and the method further includes: Continue to detect the voice interval duration and insert pause segments for the subsequent audio segments in the audio signal until the voice interval duration between all adjacent audio segments in the audio signal is adjusted to the target pause duration to obtain the target audio; Play the target audio, and during the playback of the target audio, there is a pause effect with the target pause duration between different content segments of the audio content.

14. A voice processing device, characterized in that, The device includes: An acquisition module, configured to acquire the voice information density and the voice information volume of the current audio segment in the audio signal; the voice information density is used to measure the frequency of voice information fluctuations; the voice information volume is characterized by the number of pitch frequency fluctuations of the current audio segment, and the number of pitch frequency fluctuations represents the number of times the pitch frequency state of adjacent voice frames in the current audio segment changes; A determination module, configured to input the voice information density and the voice information volume into the trained neural network model, and output the target pause duration through the trained neural network model; The acquisition module is further configured to acquire the voice interval duration between the current audio segment and the subsequent audio segment in the audio signal; An insertion module, configured to insert a pause segment between the current audio segment and the subsequent audio segment if the voice interval duration is less than the target pause duration.

15. The voice processing device according to claim 14, characterized in that, The determination module is further configured to sequentially acquire each audio frame in the audio signal and detect the voice information of each audio frame; based on the voice information of each audio frame, determine the current audio segment.

16. The voice processing device according to claim 15, wherein The determination module is further configured to use the first voice frame in the audio signal as the start node of the audio segment, or use the first voice frame after the end of the previous audio segment as the start node of the audio segment; wherein, the voice frame is an audio frame whose voice information meets the voice activity condition; if there are more than a specified number of consecutive non-voice frames after the start node, determine the end node of the audio segment based on the consecutive non-voice frames that exceed the specified number; wherein, the non-voice frame is an audio frame whose voice information does not meet the voice activity condition; based on the start node and the end node, determine the current audio segment.

17. The voice processing device according to claim 16, wherein The determining module is further configured to, for the currently acquired current audio frame, if the current start marker parameter is the first value and the current audio frame is a speech frame, use the current audio frame as the start node of the audio segment, and adjust the start marker parameter from the first value to the second value; wherein, when it is detected that there are a specified number of consecutive non-speech frames, the start marker parameter is adjusted from the second value to the first value.

18. The voice processing device according to claim 16, wherein The determining module is further configured to, for the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, use the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute until the non-speech count value exceeds the specified number; use the last non-speech frame among the consecutive specified number of non-speech frames as the end node of the audio segment.

19. The voice processing device according to claim 16, wherein The determining module is further configured to, for the audio frames after the start node, if the current audio frame is a non-speech frame, increment the non-speech count value by one; if the current audio frame is a speech frame, set the non-speech count value to zero; if the non-speech count value is less than or equal to the specified number, use the next audio frame as the current audio frame and return to the step of incrementing the non-speech count value by one if the current audio frame is a non-speech frame and continue to execute until the non-speech count value is set to zero after exceeding the specified number; Use the previous frame before the speech frame that triggers the non-speech count value to be set to zero as the end node of the audio segment.

20. The voice processing device according to claim 16, wherein The determining module is further configured to, if there are consecutive non-speech frames exceeding the specified number after the start node, use the first speech frame that appears after the consecutive non-speech frames exceeding the specified number as the start node of the next audio segment.

21. The voice processing device according to claim 14, wherein The acquisition module is further configured to acquire the current audio segment in the audio signal; perform pitch frequency detection on each audio frame in the current audio segment to obtain the number of pitch frequency fluctuations in the current audio segment; the number of pitch frequency fluctuations characterizes the speech information volume; determine the speech duration of the current audio segment; based on the comparison value between the number of pitch frequency fluctuations and the speech duration, determine the speech information density of the current audio segment.

22. The voice processing device according to claim 21, characterized in that, The acquisition module is further configured to perform pitch frequency detection on each audio frame in the current audio segment to obtain the pitch frequency values of each speech frame; based on the pitch frequency values between adjacent speech frames, determine the number of pitch frequency fluctuations in the current audio segment.

23. The voice processing device according to claim 22, wherein The acquisition module is further configured to, based on the pitch frequency values between adjacent speech frames, determine the pitch frequency states corresponding to the adjacent speech frames respectively; if the pitch frequency states of two speech frames in adjacent speech frames are different, determine that there is one pitch frequency fluctuation; based on all the pitch frequency fluctuations that occur in the current audio segment, count and obtain the number of pitch frequency fluctuations in the current audio segment.

24. The voice processing device according to claim 14, wherein The device further includes a training module, configured to obtain sample audio signals; the sample speech interval duration between adjacent sample audio segments in the sample audio signals is within a preset duration range; determine the sample speech information density and the sample speech information quantity of each sample audio segment in the sample audio signals; use the sample speech information density and the sample speech information quantity of the sample audio segments as training inputs, and use the sample speech interval duration corresponding to the respective sample audio segments as training labels; train a neural network model based on the training inputs and the training labels until the training stop condition is reached and then stop, to obtain a trained neural network model.

25. The voice processing device according to claim 14, characterized in that, The insertion module is further configured to, if the speech interval duration is less than the target pause duration, determine an inserted pause duration based on the target pause duration and the speech interval duration; the sum of the inserted pause duration and the speech interval duration is greater than or equal to the target pause duration; obtain a pause segment of the inserted pause duration, and insert the pause segment between the current audio segment and the subsequent audio segment.

26. The voice processing device according to any one of claims 14 to 25, characterized in that The audio signal includes pre-recorded audio content, and the insertion module is further configured to continue to detect the speech interval duration and insert pause segments for subsequent audio segments in the audio signal until the speech interval duration between all adjacent audio segments in the audio signal is adjusted to the target pause duration, to obtain a target audio; the device further includes a playback module; the playback module is configured to play the target audio, and during the playback of the target audio, there is a pause effect with a target pause duration between different content segments of the audio content.

27. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 13.

28. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 13.

29. A computer program product, the computer program product includes computer instructions, and when the computer instructions are executed by the processor, it implements the steps of the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Voice processing device and voice processing method

    CN104183246A

  • Method, device and equipment for voice segmentation and storage medium

    CN110675861A

  • Data transmission method and device, terminal and storage medium

    CN110890945A

  • Voice endpoint detection method and device, electronic equipment and readable storage medium

    CN112992191A