Sound signal processing method and device, computer equipment and storage medium

By rectifying the sound input signal and adjusting the gain, combined with the screening process of preset rectifying thresholds and level thresholds, the distortion problems such as cutting the top and breaking sound in sound signal processing are solved, and the sound quality and user's listening experience are improved.

CN120186533APending Publication Date: 2025-06-20GUANGDONG DINGCHUANG SMART MANUFACTURING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150744.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, the sound signal processing process is prone to distortion phenomena such as top-cutting and sound breaking, resulting in poor sound quality.

Method used

By obtaining the sound input signal, rectifying it, and obtaining the sound rectification value; comparing the sound rectification value based on multiple preset rectification thresholds, and performing gain adjustment processing on the sound input signal to obtain the sound gain signal corresponding to each preset rectification threshold; filtering all sound gain signals based on the preset level threshold and the preset rectification threshold, and determining the appropriate sound gain signal as the sound output signal.

Benefits of technology

It effectively improves the range of gain adjustment, reduces the probability of top-cut distortion caused by mismatch of fixed parameter settings, prevents sound breakage, and reduces the risk of sound quality damage caused by gain adjustment on the basis of fidelity, improves the quality of the sound output signal, and improves the user's listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186533A_ABST
    Figure CN120186533A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio processing, and discloses a sound signal processing method and device, computer equipment and a storage medium. The sound signal processing method comprises the following steps: acquiring a sound input signal, and performing rectification processing on the sound input signal to obtain a sound rectification value; performing comparison processing on the sound rectification value based on a plurality of preset rectification thresholds, and performing gain adjustment processing on the sound input signal to obtain sound gain signals corresponding to the preset rectification thresholds; and screening all the sound gain signals according to a preset level threshold value and a preset rectification threshold value, and determining the screened sound gain signals as sound output signals. Compared with the prior art, the sound rectification value is compared and subjected to gain adjustment processing based on the multiple preset rectification threshold values, the audio input gain adjustment range required by all scenes can be covered, the clipping distortion probability is reduced, sound cracking is effectively prevented, the quality of a sound output signal is improved, and the hearing experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and particularly to a method and apparatus for processing sound signals, a computer device, and a storage medium. Background Art

[0002] The main task of a microphone is to convert sound into an electrical signal. Ideally, the sound signal collected by the microphone should be a sine wave, which is a continuous and smooth waveform that can accurately reflect the frequency and amplitude of the sound. However, in actual applications, when the microphone picks up human voices, since the human voice is sometimes loud and sometimes soft, it is necessary to process the sound through software or algorithms.

[0003] In the prior art, when processing sound based on software or algorithms, parameter mismatches are likely to occur, resulting in distortion phenomena such as clipping and popping. For example, when the sound pressure collected by the microphone exceeds the maximum range it can withstand, the peak of the waveform may exceed the set fixed parameter upper limit (such as 1V), resulting in clipping. Clipping means that the peak of the waveform is truncated, forming a square top, and the waveform will be distorted, changing from a sine wave to a square wave. Clipping distortion makes the sound sharp and harsh, and even causes popping, seriously affecting the sound quality and making the sound sound unnatural and unharmonious. In addition, clipping distortion may also cause other types of distortion (such as harmonic distortion, intermodulation distortion, etc.), further reducing the user's listening experience. Summary of the Invention

[0004] Based on this, it is necessary to provide a method and apparatus for processing sound signals, a computer device, and a storage medium for the above technical problems, so as to solve the problem that the existing sound signal processing process is prone to distortion and popping, resulting in poor sound quality.

[0005] A method for processing a sound signal includes: Obtaining a sound input signal, rectifying the sound input signal to obtain a sound rectification value; Comparing the sound rectification value based on a plurality of preset rectification thresholds, and performing gain adjustment processing on the sound input signal to obtain sound gain signals corresponding to the respective preset rectification thresholds; Screening all the sound gain signals according to a preset level threshold and the preset rectification thresholds, and determining the screened sound gain signals as sound output signals.

[0006] A sound signal processing apparatus includes: A rectification module, configured to obtain a sound input signal, rectify the sound input signal to obtain a sound rectification value; A gain adjustment module, configured to compare and process the sound rectification value based on a plurality of preset rectification thresholds, and perform gain adjustment processing on the sound input signal to obtain sound gain signals corresponding to the respective preset rectification thresholds; A comparison module, configured to screen all the sound gain signals according to a preset level threshold and the preset rectification thresholds, and determine the screened sound gain signals as sound output signals.

[0007] A computer device includes a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the above-mentioned sound signal processing method is implemented.

[0008] A computer-readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the above-mentioned sound signal processing method.

[0009] In the above-mentioned sound signal processing method, apparatus, computer device, and storage medium, the sound signal processing method obtains a sound input signal, performs rectification processing on the sound input signal to obtain a sound rectification value; compares and processes the sound rectification value based on a plurality of preset rectification thresholds, and performs gain adjustment processing on the sound input signal to obtain sound gain signals corresponding to the respective preset rectification thresholds; screens all the sound gain signals according to a preset level threshold and the preset rectification thresholds, and determines the screened sound gain signals as sound output signals. After obtaining the sound rectification value, the present application compares and processes the sound rectification value and performs gain adjustment processing based on a plurality of preset rectification thresholds, which can cover the audio input gain adjustment range required by all scenarios, effectively improve the range breadth of gain adjustment, reduce the probability of clipping distortion caused by mismatched fixed parameter settings, and thus effectively prevent sound distortion. In addition, the present application compares whether there is distortion according to the gain adjustment result, so as to screen out appropriate sound gain signals as sound output signals, which can reduce the risk of sound quality degradation caused by gain adjustment while ensuring fidelity, improve the quality of the sound output signal, and improve the user's listening experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1It is a schematic flowchart of a method for processing sound signals in an embodiment of the present invention; Figure 2 It is a schematic structural diagram of a sound signal processing device in an embodiment of the present invention; Figure 3 It is a schematic diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0013] The sound signal processing method provided in this embodiment can be applied to an application scenario where a sound collection device is used to process the collected sound signal and convert it into other signals. Among them, the sound collection device can be a microphone, and the microphone includes a sound signal processing system, which can be regarded as a server for implementing the sound signal processing process. The sound signal processing system includes a rectifier control part, a gain amplifier control part, and a comparison module control part with different functions. Based on the mutual cooperation among the three, the sound input signal is converted into a sound output signal. Specifically, the rectifier control part is used to obtain the sound input signal and perform rectification processing on the sound input signal to obtain a sound rectification value; the gain amplifier control part is used to perform comparison processing and gain adjustment processing on the sound rectification value based on multiple preset rectification thresholds to obtain sound gain signals corresponding to each preset rectification threshold; the comparison module control part is used to perform screening processing on all the sound gain signals according to a preset level threshold and a preset rectification threshold, and determine the screened sound gain signal as the sound output signal.

[0014] In one embodiment, as Figure 1 shown, a method for processing sound signals is provided, including the following steps S10 - S30: S10. Obtain a sound input signal, and perform rectification processing on the sound input signal to obtain a sound rectification value.

[0015] Understandably, the voice input signal refers to the acoustic wave signal that needs to be converted and processed. For example, the voice input signal can be an analog or digital signal captured by a voice sensor such as a microphone for characterizing sound. After obtaining the voice input signal, since the voice input signal is an AC signal with a sine waveform, rectification processing is required to obtain the voice rectification value. The voice rectification value refers to the result of rectifying the input signal for measuring the effective value of the voice signal. Rectification processing is a process of converting an AC signal into a DC signal or a signal form similar to DC.

[0016] In one embodiment, in step S10, that is, obtaining the voice input signal and performing rectification processing on the voice input signal to obtain the voice rectification value, includes: S101. Obtain the voice input signal within a preset acquisition period, perform root mean square (RMS) calculation processing on the voice input signal, and determine the calculated RMS value as the voice rectification value.

[0017] Understandably, in audio engineering, the root mean square (RMS) value is crucial for evaluating audio quality, optimizing the audio experience, and controlling audio levels. The RMS value is a method for measuring the effective value of an AC signal and represents the average power of the signal over time. Compared with the peak value or average value, the RMS value can better reflect the actual energy or power of the signal because it takes into account all the amplitudes of the signal and performs a weighted average according to the length of time they appear. The RMS value reduces the influence of random fluctuations on the measurement of signal strength and has a certain noise suppression effect. In addition, the RMS value also affects the dynamic range of the audio, that is, the difference between the loudest and weakest sounds. RMS rectification can be used as a reference for dynamic range compression to ensure that the signal maintains an appropriate balance between strong and weak sounds. Subsequently, by adjusting the gain to optimize the RMS value, the user's listening experience can be improved.

[0018] In one embodiment, an RMS rectifier is used to implement the rectification processing of the voice input signal. The RMS rectifier control part in the voice signal processing system performs root mean square calculation processing on the voice input signal according to a preset acquisition period, and determines the calculated RMS value as the voice rectification value. The preset acquisition period is a fixed time period preset for capturing and analyzing the voice signal. For example, the default value of the preset acquisition period can be set to 1 second, or it can be adjusted according to needs. In practical applications, the captured voice signal is divided into multiple frames according to the preset acquisition period (each frame contains a certain number of sampling points), and then the RMS value is calculated for each frame respectively. Finally, the voice rectification values of the voice signal in different time periods are obtained. The calculation formula of the RMS value is , where represents the voice input signal at time Indicates a preset acquisition period.

[0019] In this embodiment, by performing root-mean-square calculation processing on the sound input signal and using the obtained root-mean-square value as the sound rectification value, an effective method is adopted to measure the intensity of the sound signal, suppress unwanted noise, which helps to meet subsequent specific processing requirements, thereby improving the sound quality.

[0020] S20. Compare and process the sound rectification value based on multiple preset rectification thresholds, and perform gain adjustment processing on the sound input signal to obtain sound gain signals corresponding to the respective preset rectification thresholds.

[0021] Understandably, the preset rectification threshold refers to the amplitude of the rectification signal that is preset and acceptable to the system without distortion. The RMS value is the root-mean-square result of the system's amplitude calculation for the audio segment within a preset acquisition period. By setting in advance a sound rectification value threshold close to 0 dbfs, such as -3 db. In digital audio processing, 0 dbfs is an important reference point that defines the maximum dynamic range of the signal. 0 dbfs represents the maximum possible peak level of the digital audio signal, corresponding to the full-scale value of the digital audio system. Different gain adjustment processes can be performed on the sound input signal based on different comparison results between the sound rectification value and the preset rectification threshold. The gain adjustment process is a process of amplifying or reducing the sound input signal. For example, when the preset rectification threshold is -3 db, signals not exceeding -3 db are enhanced, and signals exceeding -3 db are attenuated to prevent clipping distortion. The sound gain signal refers to the result of performing gain adjustment processing on the sound input signal.

[0022] In one embodiment, a Voltage Controlled Amplifier (VCA) is used as a gain amplifier to implement the comparison processing and gain adjustment processing of the sound rectification value. The VCA is a variable gain amplifier with controllable gain. When the intensity of the input signal exceeds the set threshold, the VCA will automatically reduce the gain of the signal, thereby compressing the audio signal, which helps to prevent clipping distortion and ensure the quality and stability of the audio signal. Since there are multiple preset rectification thresholds set in the sound signal processing system, for each sound rectification value, it is necessary to compare it with multiple preset rectification thresholds one by one, and perform gain adjustment processing on the sound input signal according to the comparison result to obtain sound gain signals corresponding to each preset rectification threshold. Among them, for each preset rectification threshold, after the sound rectification value exceeding the preset rectification threshold is detected by comparison, the VCA is used to attenuate and adjust the sound input signal, and after the sound rectification value lower than the preset rectification threshold is detected by comparison, the VCA is used to enhance and adjust the sound input signal.

[0023] In one embodiment, in step S20, that is, based on multiple preset rectification thresholds, comparing the sound rectification value and performing gain adjustment processing on the sound input signal to obtain sound gain signals corresponding to each preset rectification threshold includes: S201. Compare each preset rectification threshold with the sound rectification value to obtain rectification comparison results corresponding to each preset rectification threshold; S202. According to the rectification comparison results, look up the gain adjustment coefficients corresponding to the preset rectification thresholds in the preset rectification gain association table, and perform gain processing on the sound input signal according to the gain adjustment coefficients to obtain sound gain signals corresponding to each preset rectification threshold.

[0024] It can be understood that since there are multiple preset rectification thresholds set in the sound signal processing system, for each sound rectification value, it is necessary to compare it with multiple preset rectification thresholds one by one to obtain rectification comparison results corresponding to each preset rectification threshold. The rectification comparison result is a result used to represent the magnitude relationship between the preset rectification threshold and the sound rectification value. The preset rectification gain association table is a pre-established data table used to record the mapping relationship between each preset rectification threshold, different rectification comparison results, and different gain adjustment coefficients. The gain adjustment coefficient is a coefficient used to amplify or reduce the sound input signal. According to the rectification comparison results, look up the gain adjustment coefficients corresponding to the preset rectification thresholds in the preset rectification gain association table, and perform gain processing on the sound input signal according to the gain adjustment coefficients to obtain sound gain signals corresponding to each preset rectification threshold.

[0025] In one embodiment, four preset rectification thresholds of -3dB, -4dB, -5dB, and -6dB are set in the sound signal processing system, and a VCA is used as the gain amplifier. The preset rectification gain association table contains different gain adjustment coefficients corresponding to different preset rectification thresholds under different rectification comparison results. For example, not only are the gain adjustment coefficients corresponding to the RMS value less than -3dB different from those corresponding to the RMS value greater than -3dB, but the gain adjustment coefficients corresponding to the RMS value less than -3dB are also different from those corresponding to the RMS value less than -4dB. The VCA compares the RMS value calculated by the -3dB and RMS rectifier, and looks up the gain adjustment coefficient corresponding to -3dB from the preset rectification gain association table according to the comparison result, and uses this gain adjustment coefficient to compress the peak value of the sound input signal corresponding to the RMS value to obtain the first gain audio (the sound gain signal corresponding to -3dB). Then, the VCA processes in the same steps according to the three preset rectification thresholds of -4dB, -5dB, and -6dB in turn, and obtains the second gain audio (the sound gain signal corresponding to -4dB), the third gain audio (the sound gain signal corresponding to -5dB), and the fourth gain audio (the sound gain signal corresponding to -6dB) respectively.

[0026] In another embodiment, four preset rectification thresholds of -3dB, -4dB, -5dB, and -6dB are set in the sound signal processing system, and a VCA is used as the gain amplifier. The preset rectification gain association table only records the mapping relationship of different gain adjustment coefficients corresponding to different rectification comparison results, and the same gain adjustment coefficient corresponding to different preset rectification thresholds under the same rectification comparison result. For example, the gain adjustment coefficients corresponding to the RMS value less than -3dB are different from those corresponding to the RMS value greater than -3dB, but the gain adjustment coefficients corresponding to the RMS value less than -3dB are the same as those corresponding to the RMS value less than -4dB, -5dB, and -6dB. By adopting the steps of the above embodiment, the first gain audio (the sound gain signal corresponding to -3dB), the second gain audio (the sound gain signal corresponding to -4dB), the third gain audio (the sound gain signal corresponding to -5dB), and the fourth gain audio (the sound gain signal corresponding to -6dB) can also be obtained.

[0027] In this embodiment, by comparing the sound rectification value with multiple preset rectification thresholds and looking up the corresponding gain adjustment coefficient according to the comparison result, the rationality of the gain adjustment process for the sound input signal can be improved. At the same time, based on the comparison and gain adjustment processes of the sound rectification value with multiple preset rectification thresholds, the audio input gain adjustment range required by all scenarios can be covered, effectively improving the breadth of the gain adjustment range.

[0028] In one embodiment, before step S202, that is, before finding the gain adjustment coefficient corresponding to the preset rectification threshold from the preset rectification gain association table according to the rectification comparison result, the following steps are included: S2021. Obtain historical rectification data within a preset historical period and the historical gain coefficient corresponding to the historical rectification data; S2022. Perform interval classification processing on the historical rectification data according to a preset classification rule to obtain multiple preset rectification thresholds, and determine the gain adjustment coefficient corresponding to each preset rectification threshold according to the historical gain coefficient. The gain adjustment coefficient includes a standard gain coefficient and a compression gain coefficient; S2023. Associate each preset rectification threshold with the gain adjustment coefficient corresponding to the preset rectification threshold to generate a rectification gain association data group; S2024. Generate a preset rectification gain association table according to all the rectification gain association data groups.

[0029] Understandably, before finding the gain adjustment coefficient corresponding to the preset rectification threshold from the preset rectification gain association table according to the rectification comparison result, it is necessary to first establish a preset rectification gain association table based on historical data.

[0030] First, obtain historical rectification data within a preset historical period and the historical gain coefficient corresponding to the historical rectification data. The preset historical period is a specific time period preset for collecting and analyzing historical data, such as the past three months. Multiple historical sound input signals can be collected within the preset historical period, as well as the historical rectification data, historical gain coefficient, and historical sound output signal corresponding to each historical sound input signal. The historical rectification data refers to the result of rectifying the historical sound input signal. The historical rectification data includes multiple historical sound rectification values obtained by rectifying the historical sound input signal. The historical gain coefficient refers to the actual gain adjustment coefficient used when converting the historical sound input signal to the historical sound output signal.

[0031] Secondly, perform interval classification processing on the historical rectification data according to the preset classification rules to obtain multiple preset rectification thresholds. The preset classification rules refer to the conditions preset for dividing all historical rectification data into intervals. For example, when the historical rectification data includes -2.5db, -2.8db, -3.2db, -3.9db, -4.1db, -4.5db, -5.3db, -5.8db, and -6.1db, perform interval classification processing according to the rounding division interval rule, and determine the integer boundary values -3db, -4db, -5db, and -6db of the interval as the preset rectification thresholds. In addition, after determining the preset rectification thresholds, perform corresponding statistics on the historical gain coefficients to determine the gain adjustment coefficients corresponding to each preset rectification threshold. The gain adjustment coefficients include the standard gain coefficient and the compression gain coefficient. The standard gain coefficient refers to the coefficient for performing gain adjustment processing on the sound input signal when the sound rectification value is less than the preset rectification threshold, and the compression gain coefficient refers to the coefficient for performing gain adjustment processing on the sound input signal when the sound rectification value is greater than or equal to the preset rectification threshold. For example, when the preset rectification threshold is -3db, determine the standard gain coefficient corresponding to -3db based on the historical gain coefficients actually used for the historical sound input signals corresponding to the historical rectification data (-2.5db, -2.8db) less than -3db, and determine the compression gain coefficient corresponding to -3db based on the historical gain coefficients actually used for the historical sound input signals corresponding to the historical rectification data (-3.2db, -3.9db, -4.1db, -4.5db, -5.3db, -5.8db, -6.1db) greater than or equal to -3db. Each preset rectification threshold corresponds to a standard gain coefficient and a compression gain coefficient. The standard gain coefficients and compression gain coefficients between different preset rectification thresholds can be the same or different. For example, the standard gain coefficient corresponding to the RMS value less than -3db and the standard gain coefficient corresponding to the RMS value less than -4db may be the same or different.

[0032] Finally, associate each preset rectification threshold with the gain adjustment coefficient corresponding to it to generate a rectification gain association data group. The rectification gain association data group is a data group used to represent the mapping relationship between a preset rectification threshold and a set of standard gain coefficients and compression gain coefficients. Combining all the rectification gain association data groups can generate a preset rectification gain association table.

[0033] This embodiment starts from obtaining historical data, and through data classification and processing, and data association, finally generates a preset rectification gain association table, improving the orderliness and data availability between the preset rectification thresholds and the gain adjustment coefficients. At the same time, it can ensure that the corresponding gain adjustment coefficients can be quickly found based on the preset rectification gain association table in the follow-up, so as to achieve precise control of the rectification process.

[0034] It should be noted that the above steps are only an example of selecting the corresponding gain adjustment coefficient according to the preset rectification threshold. However, the present invention is not limited thereto. In addition to the method of establishing a mapping table, any other method can be used to establish the corresponding relationship.

[0035] In one embodiment, in step S201, that is, comparing each of the preset rectification thresholds with the sound rectification value to obtain a rectification comparison result corresponding to each of the preset rectification thresholds, includes: S2011. Comparing each of the preset rectification thresholds with the sound rectification value to determine whether the sound rectification value is less than the preset rectification threshold; S2012. If the sound rectification value is less than the preset rectification threshold, determining that the rectification comparison result corresponding to the preset rectification threshold is that the rectification result meets the standard; S2013. If the sound rectification value is greater than or equal to the preset rectification threshold, determining that the rectification comparison result corresponding to the preset rectification threshold is that the rectification result exceeds the standard.

[0036] It can be understood that when comparing each preset rectification threshold with the sound rectification value, it is necessary to determine whether the sound rectification value is less than the preset rectification threshold, and the size relationship between the preset rectification threshold and the sound rectification value is characterized by the rectification comparison result. The rectification comparison result includes that the rectification result meets the standard and the rectification result exceeds the standard. The rectification result meeting the standard is a result used to characterize that the sound rectification value is less than the preset rectification threshold, and the rectification result exceeding the standard is a result used to characterize that the sound rectification value is greater than the preset rectification threshold.

[0037] In one embodiment, the rectification result meeting the standard and the rectification result exceeding the standard are represented by the Boolean values "1" and "0". When the sound rectification value is less than the preset rectification threshold, the rectification comparison result is that the rectification result meets the standard, and the obtained Boolean value is 1. When the sound rectification value is greater than or equal to the preset rectification threshold, the rectification comparison result is that the rectification result exceeds the standard, and the obtained Boolean value is 0.

[0038] This embodiment distinguishes different rectification comparison results by judging the size relationship between the sound rectification value and the preset rectification threshold, ensuring the accuracy of the rectification comparison result and providing a basis for the gain adjustment process.

[0039] In one embodiment, in step S202, that is, finding out the gain adjustment coefficient corresponding to the preset rectification threshold from the preset rectification gain association table according to the rectification comparison result, includes: S2025. When the rectification comparison result is that the rectification result meets the standard, finding out the standard gain coefficient corresponding to the preset rectification threshold from the preset rectification gain association table and determining the standard gain coefficient as the gain adjustment coefficient; S2026. When the rectification comparison result indicates that the rectification result exceeds the standard, look up the compression gain coefficient corresponding to the preset rectification threshold in the preset rectification gain association table, and determine the compression gain coefficient as the gain adjustment coefficient.

[0040] Understandably, in one embodiment, the preset rectification gain association table forms rectification gain association data groups including "preset rectification threshold ~ rectification result meets the standard ~ standard gain coefficient" and "preset rectification threshold ~ rectification result exceeds the standard ~ compression gain coefficient" with the preset rectification threshold as a unit, so as to facilitate looking up the gain adjustment coefficient corresponding to the preset rectification threshold from the preset rectification gain association table according to the rectification comparison result. Specifically, when the preset rectification threshold is -3db, the corresponding rectification gain association data groups are "-3db ~ 1 ~ A" and "-3db ~ 0 ~ B". Among them, -3db represents the value of the preset rectification threshold, 1 represents that the rectification result meets the standard, 0 represents that the rectification result exceeds the standard, A represents the standard gain coefficient corresponding to -3db, and B represents the compression gain coefficient corresponding to -3db. For example, A is 1.2 and B is 0.8.

[0041] In this embodiment, according to different preset rectification thresholds and rectification comparison results, a suitable gain adjustment coefficient is intelligently selected from the preset rectification gain association table, thereby realizing precise control of the signal strength and improving the flexibility of the system.

[0042] S30. Screen all the voice gain signals according to the preset level threshold and the preset rectification threshold, and determine the screened voice gain signal as the voice output signal.

[0043] Understandably, the preset level threshold refers to the maximum possible peak level of the digital audio signal during audio processing, that is, 0dbfs, corresponding to the full-scale value of the digital audio system. 0dbfs is the highest point of the level, and exceeding this value will cause signal overload. There are multiple preset rectification thresholds set in the voice signal processing system, and each preset rectification threshold corresponds to a voice gain signal. The voice gain signals corresponding to some preset rectification thresholds may exceed 0dbfs, so it is necessary to screen out the voice gain signals that exceed 0dbfs. After the 0dbfs screening, if only one voice gain signal remains, directly determine this one voice gain signal as the voice output signal. If two or more voice gain signals still remain, it is necessary to further screen according to the magnitude relationship of the preset rectification thresholds, and finally screen out one voice gain signal and determine it as the voice output signal.

[0044] In one embodiment, in step S30, that is, screening all the voice gain signals according to the preset level threshold and the preset rectification threshold, and determining the screened voice gain signal as the voice output signal, includes: S301. Obtain the peak value of the gain signal for each of the said voice gain signals; S302. Select from all the said voice gain signals the voice gain signals whose peak values of the gain signals are less than or equal to a preset level threshold, and determine the selected voice gain signals as candidate output signals; S303. Sort all the said candidate output signals according to the magnitude of a preset rectification threshold, and determine the candidate output signal with the smallest preset rectification threshold as the voice output signal.

[0045] Understandably, when performing screening processing on all voice gain signals, first, it is necessary to obtain the peak value of the gain signal for each voice gain signal. The peak value of the gain signal refers to the peak value of the curve in the voice gain signal. Then, judge one by one the magnitude relationship between each peak value of the gain signal and the preset level threshold, screen out the voice gain signals whose peak values of the gain signals are greater than the preset level threshold, and retain the voice gain signals whose peak values of the gain signals are less than or equal to the preset level threshold to be determined as candidate output signals. The candidate output signal refers to the voice gain signal after being screened out by judging the magnitude of the preset level threshold. Next, judge whether the number of candidate output signals is greater than one. If the number of candidate output signals is only one, directly determine the candidate output signal as the voice output signal. Finally, if the number of candidate output signals is greater than one, it is necessary to perform secondary screening on the candidate output signals according to the preset rectification threshold, sort all the candidate output signals according to the magnitude of the preset rectification threshold, and determine the candidate output signal with the smallest preset rectification threshold as the voice output signal. For example, when the preset rectification thresholds corresponding to the candidate output signals are -3dB, -5dB, and -6dB, determine the candidate output signal corresponding to -6dB as the voice output signal, that is, the voice gain signal obtained by processing the voice input signal based on the gain adjustment coefficient of -6dB.

[0046] Based on screening the voice gain signals by the preset level threshold, this embodiment further performs secondary screening on the selected candidate output signals. On the premise of ensuring that the output voice signal meets the intensity requirements, select the gain audio with the lowest compression degree for output, reduce the deterioration of the audio quality caused by compression, and effectively improve the quality of the input audio.

[0047] In summary, in this embodiment, by obtaining a voice input signal, rectifying the voice input signal to obtain a voice rectification value; comparing the voice rectification value based on multiple preset rectification thresholds, and performing a gain adjustment process on the voice input signal to obtain voice gain signals corresponding to the respective preset rectification thresholds; screening all the voice gain signals according to a preset level threshold and a preset rectification threshold, and determining the screened voice gain signal as a voice output signal. After obtaining the voice rectification value in this embodiment, comparing and adjusting the gain of the voice rectification value based on multiple preset rectification thresholds can cover the audio input gain adjustment range required by all scenarios, effectively improve the breadth of the gain adjustment range, reduce the probability of clipping distortion caused by mismatched fixed parameter settings, and thus effectively prevent voice distortion. In addition, in this embodiment, it is determined whether there is distortion according to the gain adjustment result, so as to screen out a suitable voice gain signal as the voice output signal, which can reduce the risk of impaired sound quality caused by gain adjustment while ensuring fidelity, improve the quality of the voice output signal, and improve the user's listening experience.

[0048] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0049] In one embodiment, a voice signal processing device is provided, which corresponds one-to-one to the voice signal processing method in the above embodiment. As Figure 2 shown, the voice signal processing device includes a rectification module 10, a gain adjustment module 20, and a comparison module 30, which respectively correspond to the rectifier control part, the gain amplifier control part, and the comparison module control part with different functions in the voice signal processing system. The detailed description of each functional module is as follows: The rectification module 10 is configured to obtain a voice input signal and perform a rectification process on the voice input signal to obtain a voice rectification value; The gain adjustment module 20 is configured to compare the voice rectification value based on multiple preset rectification thresholds, and perform a gain adjustment process on the voice input signal to obtain voice gain signals corresponding to the respective preset rectification thresholds; The comparison module 30 is configured to screen all the voice gain signals according to a preset level threshold and the preset rectification threshold, and determine the screened voice gain signal as a voice output signal.

[0050] In one embodiment, the rectification module 10 includes: The root mean square calculation processing unit is used to obtain the sound input signal within a preset acquisition period, perform root mean square calculation processing on the sound input signal, and determine the calculated root mean square value as the sound rectification value.

[0051] In one embodiment, the gain adjustment module 20 includes: The rectification comparison unit is used to compare each of the preset rectification thresholds with the sound rectification value to obtain rectification comparison results corresponding to the respective preset rectification thresholds; The gain adjustment processing unit is used to look up the gain adjustment coefficient corresponding to the preset rectification threshold from the preset rectification gain association table according to the rectification comparison result, and perform gain processing on the sound input signal according to the gain adjustment coefficient to obtain sound gain signals corresponding to the respective preset rectification thresholds.

[0052] In one embodiment, the gain adjustment module 20 further includes: The historical data acquisition unit is used to acquire historical rectification data within a preset historical period and the historical gain coefficients corresponding to the historical rectification data; The classification processing unit is used to perform interval classification processing on the historical rectification data according to a preset classification rule to obtain a plurality of preset rectification thresholds, and determine the gain adjustment coefficients corresponding to the respective preset rectification thresholds according to the historical gain coefficients, where the gain adjustment coefficients include a standard gain coefficient and a compression gain coefficient; The association processing unit is used to associate each of the preset rectification thresholds with the gain adjustment coefficient corresponding to the preset rectification threshold to generate a rectification gain association data group; The association table generation unit is used to generate a preset rectification gain association table according to all the rectification gain association data groups.

[0053] In one embodiment, the gain adjustment module 20 further includes: The rectification value comparison unit is used to compare each of the preset rectification thresholds with the sound rectification value to determine whether the sound rectification value is less than the preset rectification threshold; The rectification value compliance determination unit is used to determine that the rectification comparison result corresponding to the preset rectification threshold is rectification result compliance if the sound rectification value is less than the preset rectification threshold; The rectification value exceeding determination unit is used to determine that the rectification comparison result corresponding to the preset rectification threshold is rectification result exceeding if the sound rectification value is greater than or equal to the preset rectification threshold.

[0054] In one embodiment, the gain adjustment module 20 further includes: A compliance gain coefficient determination unit, configured to, when the rectification comparison result indicates that the rectification result is compliant, look up a standard gain coefficient corresponding to the preset rectification threshold from a preset rectification gain association table, and determine the standard gain coefficient as the gain adjustment coefficient; An over-standard gain coefficient determination unit, configured to, when the rectification comparison result indicates that the rectification result is over-standard, look up a compression gain coefficient corresponding to the preset rectification threshold from a preset rectification gain association table, and determine the compression gain coefficient as the gain adjustment coefficient.

[0055] In one embodiment, the comparison module 30 includes: A gain signal peak acquisition unit, configured to acquire the gain signal peak of each of the voice gain signals; A candidate output signal determination unit, configured to select, from all the voice gain signals, the voice gain signals whose gain signal peaks are less than or equal to a preset level threshold, and determine the selected voice gain signals as candidate output signals; A voice output signal determination unit, configured to sort all the candidate output signals according to the magnitude of the preset rectification threshold, and determine the candidate output signal with the smallest preset rectification threshold as the voice output signal.

[0056] For the specific limitations on the voice signal processing device, reference may be made to the limitations on the voice signal processing method in the foregoing text, which will not be elaborated here. Each module in the foregoing voice signal processing device can be implemented in whole or in part by software, hardware, and their combination. The foregoing modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so as to be called by the processor to execute the operations corresponding to the foregoing modules.

[0057] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structural diagram may be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium and an internal memory. The readable storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The database of the computer device is used to store the data involved in the voice signal processing method. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, a voice signal processing method is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.

[0058] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored on the memory and executable on the processor. When the processor executes the computer-readable instructions, the following steps are implemented: Obtain a voice input signal, perform rectification processing on the voice input signal to obtain a voice rectification value; Based on multiple preset rectification thresholds, perform comparison processing on the voice rectification value, and perform gain adjustment processing on the voice input signal to obtain voice gain signals corresponding to the respective preset rectification thresholds; Perform screening processing on all the voice gain signals according to a preset level threshold and the preset rectification thresholds, and determine the screened voice gain signals as voice output signals.

[0059] In one embodiment, one or more computer-readable storage media storing computer-readable instructions are provided. The readable storage media provided in this embodiment include non-volatile readable storage media and volatile readable storage media. Computer-readable instructions are stored on the readable storage media. When the computer-readable instructions are executed by one or more processors, the following steps are implemented: Obtain a voice input signal, perform rectification processing on the voice input signal to obtain a voice rectification value; Based on multiple preset rectification thresholds, perform comparison processing on the voice rectification value, and perform gain adjustment processing on the voice input signal to obtain voice gain signals corresponding to the respective preset rectification thresholds; Perform screening processing on all the voice gain signals according to a preset level threshold and the preset rectification thresholds, and determine the screened voice gain signals as voice output signals.

[0060] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0061] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0062] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A sound signal processing method, characterized in that: include: Acquire a sound input signal, perform rectification processing on the sound input signal, and obtain a sound rectification value; Comparing the sound rectification value based on a plurality of preset rectification thresholds, and performing gain adjustment processing on the sound input signal to obtain a sound gain signal corresponding to each of the preset rectification thresholds; All the sound gain signals are screened according to the preset level threshold and the preset rectification threshold, and the screened sound gain signals are determined as sound output signals.

2. The sound signal processing method according to claim 1, characterized in that: The step of acquiring a sound input signal and rectifying the sound input signal to obtain a sound rectification value includes: A sound input signal within a preset collection period is acquired, a root mean square calculation process is performed on the sound input signal, and the calculated root mean square value is determined as the sound rectification value.

3. The sound signal processing method according to claim 1, characterized in that: The comparing and processing the sound rectification values ​​based on a plurality of preset rectification thresholds, and performing gain adjustment processing on the sound input signal to obtain a sound gain signal corresponding to each of the preset rectification thresholds, comprises: Comparing each of the preset rectification thresholds with the sound rectification value to obtain a rectification comparison result corresponding to each of the preset rectification thresholds; According to the rectification comparison result, the gain adjustment coefficient corresponding to the preset rectification threshold is found from the preset rectification gain association table, and the sound input signal is gain processed according to the gain adjustment coefficient to obtain the sound gain signal corresponding to each preset rectification threshold.

4. The sound signal processing method according to claim 3, characterized in that: Before searching the gain adjustment coefficient corresponding to the preset rectification threshold from the preset rectification gain association table according to the rectification comparison result, the method includes: Acquire historical rectification data within a preset historical period and historical gain coefficients corresponding to the historical rectification data; Performing interval classification processing on the historical rectification data according to preset classification rules to obtain a plurality of preset rectification thresholds, and determining a gain adjustment coefficient corresponding to each of the preset rectification thresholds according to the historical gain coefficients, wherein the gain adjustment coefficients include a standard gain coefficient and a compression gain coefficient; Associating each of the preset rectification thresholds with a gain adjustment coefficient corresponding to the preset rectification threshold to generate a rectification gain association data group; A preset rectification gain association table is generated according to all the rectification gain association data groups.

5. The sound signal processing method according to claim 3, characterized in that: The step of comparing each of the preset rectification thresholds with the sound rectification value to obtain a rectification comparison result corresponding to each of the preset rectification thresholds includes: Comparing each of the preset rectification thresholds with the sound rectification value to determine whether the sound rectification value is less than the preset rectification threshold; If the sound rectification value is less than the preset rectification threshold, determining that the rectification comparison result corresponding to the preset rectification threshold is a rectification result that meets the standard; If the sound rectification value is greater than or equal to the preset rectification threshold, it is determined that the rectification comparison result corresponding to the preset rectification threshold is that the rectification result exceeds the standard.

6. The sound signal processing method according to claim 5, characterized in that: The gain adjustment coefficient includes a standard gain coefficient and a compression gain coefficient; The step of searching a preset rectification gain association table for a gain adjustment coefficient corresponding to the preset rectification threshold according to the rectification comparison result includes: When the rectification comparison result is that the rectification result meets the standard, a standard gain coefficient corresponding to the preset rectification threshold is found from a preset rectification gain association table, and the standard gain coefficient is determined as the gain adjustment coefficient; When the rectification comparison result is that the rectification result exceeds the standard, a compression gain coefficient corresponding to the preset rectification threshold is found from a preset rectification gain association table, and the compression gain coefficient is determined as a gain adjustment coefficient.

7. The sound signal processing method according to claim 1, characterized in that: The step of screening all the sound gain signals according to the preset level threshold and the preset rectification threshold, and determining the screened sound gain signals as the sound output signals, comprises: Obtaining a gain signal peak value of each of the sound gain signals; Selecting a sound gain signal whose gain signal peak value is less than or equal to a preset level threshold from all the sound gain signals, and determining the selected sound gain signal as a candidate output signal; All the candidate output signals are sorted according to the size of the preset rectification threshold, and the candidate output signal with the smallest preset rectification threshold is determined as the sound output signal.

8. A sound signal processing device, characterized in that: include: A rectification module, used for acquiring a sound input signal, rectifying the sound input signal, and obtaining a sound rectification value; A gain adjustment module, used for comparing the sound rectification value based on a plurality of preset rectification thresholds, and performing gain adjustment processing on the sound input signal to obtain a sound gain signal corresponding to each of the preset rectification thresholds; The comparison module is used to screen all the sound gain signals according to a preset level threshold and the preset rectification threshold, and determine the screened sound gain signals as sound output signals.

9. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that: When the processor executes the computer-readable instructions, the sound signal processing method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing computer-readable instructions, characterized in that: When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to perform the sound signal processing method according to any one of claims 1 to 7.