A method and system for voice noise analysis

By extracting and analyzing noise audio clips that contain only noise, the problem of difficulty in performing voice noise evaluation without reference audio is solved, and objective evaluation and accurate reflection of voice data noise levels are achieved.

CN114639390BActive Publication Date: 2025-06-13DMAI (GUANGZHOU) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011499230.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-15
Publication Date
2025-06-13
Estimated Expiration
2040-12-15

AI Technical Summary

Technical Problem

The prior art is difficult to objectively evaluate speech noise without reference audio, especially in teaching scenarios, where the lack of reference audio results in noise evaluation being unable to be performed.

Method used

By obtaining the voice data to be analyzed, the noise audio clips containing only noise are extracted, and the noise intensity level of the noise audio clip is determined based on the noise intensity index and the preset noise intensity division level, and the noise level of the voice data is evaluated based on the distribution of these levels.

Benefits of technology

It realizes objective evaluation of the noise level of the speech data analysis, avoids the influence of normal speech, does not require reference audio, and has a wider range of applications, which can accurately reflect the noise situation in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114639390B_ABST
    Figure CN114639390B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for voice noise analysis. The method includes: obtaining voice data to be analyzed; extracting noise audio segments containing only noise from the voice data to be analyzed; determining the noise intensity level corresponding to each noise audio segment based on the noise intensity index of each noise audio segment and a preset noise intensity division level; and determining the noise level evaluation result of the voice data to be analyzed according to the distribution of the noise intensity levels corresponding to each noise audio segment. By calculating the noise intensity index of the noise audio segments containing only noise for separate analysis, and then determining the noise level evaluation result of the entire voice data to be analyzed according to the distribution of the noise intensity levels of all the noise audio segments, the influence of normal speech is avoided, an objective evaluation of the noise level of the voice data to be analyzed is achieved, and there is no need to refer to an audio, so the application range is wider and the noise situation in various scenarios can be accurately reflected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice signal processing, and in particular to a method and system for voice noise analysis. Background Art

[0002] With the rapid development of the mobile Internet, the applications of communication software are becoming more and more extensive. For example, more and more teachers use instant messaging software to provide online teaching guidance to students to replace the traditional face-to-face teaching method. However, when using communication software, noise will seriously affect the quality of communication audio. In places with high requirements for noise, such as when students listen to the audio courses recorded by teachers online through communication software, the audio noise in the audio courses should be as small as possible to improve the teaching effect. However, due to the huge number of online teaching audios, the traditional method of relying on manual analysis of the noise in each class has a huge workload and the analysis results are highly subjective.

[0003] In the prior art, the index evaluation methods for objectively evaluating the noise situation (such as signal-to-noise ratio, segmental signal-to-noise ratio, etc.) require a reference audio that is exactly the same as the voice content with strict time alignment when measuring the noise situation of an audio. In the teaching scenario or other situations where a reference audio cannot be obtained, the existing noise evaluation methods will not be able to evaluate the noise. Therefore, how to achieve an objective evaluation of voice noise without a reference audio is an urgent problem to be solved. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method and system for voice noise analysis to overcome the problem in the prior art that it is difficult to objectively evaluate voice noise without a reference audio.

[0005] Embodiments of the present invention provide a method for voice noise analysis, including:

[0006] Obtaining the voice data to be analyzed;

[0007] Extracting a noise audio segment that only contains noise from the voice data to be analyzed;

[0008] Based on the noise intensity index of each noise audio segment and a preset noise intensity division level, determining the noise intensity level corresponding to each noise audio segment;

[0009] Determining the evaluation result of the noise level of the voice data to be analyzed according to the distribution of the noise intensity levels corresponding to each noise audio segment.

[0010] Optionally, the extracting a noise audio segment that only contains noise from the voice data to be analyzed includes:

[0011] Divide the speech data to be analyzed into multiple audio segments based on the total duration of the speech data to be analyzed and a preset extraction duration period;

[0012] Convert each audio segment into a magnitude spectrum;

[0013] Input the magnitude spectrum corresponding to each audio segment into a preset noise classification model to obtain the probability of each audio segment containing only noise;

[0014] Screen out the noise audio segments containing only noise from the audio segments based on a preset probability threshold.

[0015] Optionally, the determining the noise intensity level corresponding to each of the noise audio segments based on the noise intensity index of each noise audio segment and a preset noise intensity division level includes:

[0016] Calculate the noise intensity index corresponding to each noise audio segment respectively;

[0017] Obtain the noise intensity index ranges corresponding to different noise intensity levels in the preset noise intensity division level;

[0018] Determine the current noise intensity index range corresponding to the current noise audio segment according to the noise intensity index corresponding to the current noise audio segment;

[0019] Determine the noise intensity level corresponding to the current noise audio segment as the noise intensity level of the current noise audio segment.

[0020] Optionally, the determining the noise level of the speech data to be analyzed according to the distribution of the noise intensity levels corresponding to each of the noise audio segments includes:

[0021] Obtain the proportion of different noise intensity levels in each of the noise audio segments;

[0022] Determine the evaluation result of the noise level of the speech data to be analyzed according to the proportion of different noise intensity levels and a preset proportion evaluation index.

[0023] Optionally, the noise intensity levels include: high-intensity noise level, medium-intensity noise level, and low-intensity noise level.

[0024] Optionally, the determining the evaluation result of the noise level of the speech data to be analyzed according to the proportion of different noise intensity levels and a preset proportion evaluation index includes:

[0025] Obtain the proportion of the high-intensity noise level;

[0026] Determine the noise level evaluation result of the speech data to be analyzed according to the relationship between the proportion of the high-intensity noise level and the preset high-intensity noise level proportion range in the preset proportion evaluation index.

[0027] Optionally, the noise level evaluation result includes: low noise level, moderate noise level, and high noise level, where

[0028] When the proportion of the high-intensity noise level is less than the minimum value of the preset high-intensity noise level proportion range in the preset proportion evaluation index, it is determined that the noise level evaluation result is a low noise level;

[0029] When the proportion of the high-intensity noise level is within the preset high-intensity noise level proportion range in the preset proportion evaluation index, it is determined that the noise level evaluation result is a moderate noise level;

[0030] When the proportion of the high-intensity noise level is greater than the maximum value of the preset high-intensity noise level proportion range in the preset proportion evaluation index, it is determined that the noise level evaluation result is a high noise level.

[0031] The embodiment of the present invention also provides a speech noise analysis system, including:

[0032] An acquisition module, configured to acquire speech data to be analyzed;

[0033] A first processing module, configured to extract a noise audio segment containing only noise from the speech data to be analyzed;

[0034] A second processing module, configured to determine the noise intensity level corresponding to each noise audio segment based on the noise intensity index of each noise audio segment and a preset noise intensity division level;

[0035] A third processing module, configured to determine the noise level evaluation result of the speech data to be analyzed according to the distribution of the noise intensity levels corresponding to each noise audio segment.

[0036] The embodiment of the present invention also provides an electronic device, including: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the speech noise analysis method provided by the embodiment of the present invention.

[0037] The embodiment of the present invention also provides a computer-readable storage medium, the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the speech noise analysis method provided by the embodiment of the present invention.

[0038] The technical solution of the present invention has the following advantages:

[0039] An embodiment of the present invention provides a method and system for voice noise analysis. The method includes: obtaining voice data to be analyzed; extracting a noise audio segment containing only noise from the voice data to be analyzed; determining the noise intensity level corresponding to each noise audio segment based on the noise intensity index of each noise audio segment and a preset noise intensity division level; and determining the noise level evaluation result of the voice data to be analyzed according to the distribution of the noise intensity levels corresponding to each noise audio segment. Thus, by calculating the noise intensity index of the noise audio segment containing only noise, the noise intensity level of each noise audio segment is analyzed separately, and then the noise level evaluation result of the entire voice data to be analyzed is determined according to the distribution of the noise intensity levels of all noise audio segments, avoiding the influence of normal voice in the voice data to be analyzed, realizing an objective evaluation of the noise level of the voice data to be analyzed, and without the need for a reference audio, having a wider application range and being able to accurately reflect the noise situation in various scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is a flowchart of the voice noise analysis method in the embodiment of the present invention;

[0042] Figure 2 It is a schematic diagram of the process of inputting the amplitude spectrum corresponding to each audio segment into a preset noise classification model to obtain the probability of each audio segment containing only noise in the embodiment of the present invention;

[0043] Figure 3 It is a schematic diagram of the structure of the voice noise analysis system in the embodiment of the present invention;

[0044] Figure 4 It is a schematic diagram of the structure of an electronic device in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0046] In different embodiments of the present invention described below, the technical features involved can be combined with each other as long as they do not conflict with each other.

[0047] With the rapid development of the mobile Internet, online education has gradually replaced the traditional education method. Currently, more and more teachers use instant messaging software to provide teaching guidance to students, which makes it more convenient to intelligently analyze the classroom situation. Noise is a factor that affects the quality of students' classroom learning. Therefore, it is necessary to detect noise to create a quiet environment for students and ensure the learning effect. However, currently, a large amount of online teaching audio and video is generated every day. Manually analyzing the noise situation in each class is a huge workload and the analysis results are highly subjective. Therefore, objective intelligent analysis of audio noise is particularly necessary.

[0048] Currently, the indicators for objectively evaluating the noise situation (such as signal-to-noise ratio, segmental signal-to-noise ratio, etc.) have great limitations in application. That is, when measuring the noise situation of a teaching audio, a reference audio with exactly the same speech content that is strictly time-aligned with it is required, which is extremely difficult to obtain in the teaching scenario. Therefore, how to achieve noise assessment without a reference is an urgent problem to be solved.

[0049] The embodiment of the present invention provides a method for analyzing voice noise, which can be applied to noise analysis of an online teaching platform, such as Figure 1 shown, the method for analyzing voice noise mainly includes the following steps:

[0050] Step S101: Obtain the voice data to be analyzed. Specifically, the voice data to be analyzed is audio data containing noise, for example: teaching audio recorded on an online teaching platform or audio data extracted from a teaching video containing voice data, etc. The acquisition method of the voice data to be analyzed can be directly downloading the audio data or extracting it from a preset voice database to be analyzed, etc. The present invention is not limited thereto.

[0051] Step S102: Extract a noise audio segment containing only noise from the voice data to be analyzed. Specifically, since the voice data to be analyzed containing noise contains both normal speech and noise, in order to avoid the need for a reference audio (i.e., normal speech) to evaluate the noise, by extracting a noise audio segment containing only noise, the magnitude of the noise can be directly and intuitively measured by extracting noise intensity indicators such as the volume or energy of the audio.

[0052] Step S103: Based on the noise intensity index of each noise audio segment and the preset noise intensity grading levels, determine the noise intensity level corresponding to each noise audio segment. Specifically, there are significant differences in the noise energy or volume among different noise audio segments. By grading each noise audio segment, it is possible to more intuitively compare each noise audio segment, facilitating the subsequent assessment of the noise level of the entire speech data to be analyzed.

[0053] Step S104: Determine the assessment result of the noise level of the speech data to be analyzed according to the distribution of the noise intensity levels corresponding to each noise audio segment. Specifically, since a complete speech data to be analyzed contains many noise audio segments, to improve the accuracy of the noise assessment of the entire speech data to be analyzed, the noise level assessment result is obtained by considering the distribution of the noise intensity levels of all noise audio segments, achieving an objective noise assessment of the speech data to be analyzed.

[0054] Through the above steps S101 to S104, the speech noise analysis method provided by the embodiment of the present invention separately analyzes the noise intensity levels of each noise audio segment by calculating the noise intensity index of the noise audio segment containing only noise, and then determines the assessment result of the noise level of the entire speech data to be analyzed according to the distribution of the noise intensity levels of all noise audio segments, avoiding the influence of normal speech in the speech data to be analyzed, achieving an objective assessment of the noise level of the speech data to be analyzed, and without the need to refer to an audio, having a wider application range and being able to accurately reflect the noise conditions in various scenarios.

[0055] Specifically, in one embodiment, the above step S102 specifically includes the following steps:

[0056] Step S201: Based on the total duration of the speech data to be analyzed and the preset extraction duration period, divide the speech data to be analyzed into multiple audio segments. Specifically, divide the speech data to be analyzed into several audio segments with relatively short and equal lengths according to the time axis of the total duration. The preset extraction duration period can be flexibly set according to the total duration and the accuracy requirements of noise analysis, such as 1 s, 3 s, etc., and the present invention is not limited thereto.

[0057] Step S202: Convert each audio segment into a magnitude spectrum. Specifically, each audio segment is converted into a magnitude spectrum by operations such as framing, Fourier transform, magnitude calculation, and magnitude normalization for each audio segment.

[0058] Step S203: Input the amplitude spectrum corresponding to each audio segment into a preset noise classification model to obtain the probability that each audio segment contains only noise. The preset noise classification model is a classification model established in advance. The input of this classification model is an audio segment, and the output is the probability of predicting that the audio segment contains only noise. After training this classification model with a large number of known audio segments, the model is obtained.

[0059] In the embodiment of the present invention, as Figure 2 shown, the classification model uses mobilenet-v2 as the backbone network to further obtain several deep features of the audio, and then aggregates these deep features to obtain the dense features of the audio, and finally sends them to the classifier for classification. Among them, the backbone network mobilenet-v2 uses depthwise separable convolution instead of traditional convolution, and the inference speed is faster. It has been widely used in the industry and will not be introduced in depth here; in the feature aggregation stage, a more effective feature aggregation method NetVLAD Pooling is used. Assume that the deep features obtained by the backbone network are {x 1 ,x 2 ,…,x T}, and the intermediate output of NetVLADPooling is a K×D matrix V, where K represents the predefined number of clusters, and D represents the dimension size of each cluster center. Then each row of the matrix V is obtained by the following formula:

[0060]

[0061] where {w k}, {b k}, {c k} are training parameters and are trained together with the classification model. After performing L2 regularization on the matrix V and splicing them together, they are the features aggregated by NetVLAD Pooling, and then sent to the fully connected layer for binary classification. The entire classification model uses the binary cross-entropy loss function as the objective for training.

[0062] Step S204: Screen out the noise audio segments that contain only noise from the audio segments based on a preset probability threshold. Specifically, compare and judge the probability of the audio segment obtained in the above step S203 with the preset probability threshold. If the probability value exceeds the threshold, it means that the audio segment contains only noise; otherwise, the audio also contains non-noise such as human voices. Then retain all the audio segments judged to contain only noise, that is, the noise audio segments, and discard other audio segments.

[0063] Specifically, in one embodiment, the above step S103 specifically includes the following steps:

[0064] Step S301: Calculate the noise intensity index corresponding to each noisy audio segment respectively. Specifically, the noise intensity index can be an index that can reflect the magnitude of the noise, such as the energy or volume of the noise. In the embodiments of the present invention, the noise energy index is adopted, that is, calculate the energy of the noisy audio segment. Suppose the noisy audio segment is represented by a={a 1 ,a 2 ,…,a N}, and N is the number of sample points included in the audio, then the energy of the audio is calculated by the following method:

[0065]

[0066] where energy represents the energy of the audio, t represents the duration of the audio, N is the number of sample points included in the audio, and a 1 ,a 2 ,…,a N represent the energy values of each sample point of the audio.

[0067] Step S302: Obtain the noise intensity index ranges corresponding to different noise intensity levels in the preset noise intensity division levels. Specifically, in the embodiments of the present invention, taking the preset noise intensity division levels including low-intensity noise level, medium-intensity noise level and high-intensity noise level as an example, the division basis is the noise intensity index range. In the embodiments of the present invention, by presetting two energy thresholds T l , T h , the noise with energy less than the low threshold is low-intensity noise, the noise between the low threshold and the high threshold is medium-intensity noise, and the noise greater than the high threshold is classified as high-intensity noise.

[0068] Step S303: Determine the current noise intensity index range corresponding to the current noisy audio segment according to the noise intensity index corresponding to the current noisy audio segment. Determine the current noise intensity index range to which it belongs through the relationship between the noise energy value calculated in the above step S301 and the two energy thresholds T l , T h in the above step S302.

[0069] Step S304: Determine the noise intensity level corresponding to the current noise intensity index range as the noise intensity level of the current noisy audio segment. Specifically, suppose the energy value corresponding to the current noisy audio segment is A, and T l < A < T h , then the current noise intensity index range corresponding to the noisy audio segment corresponds to the medium-intensity noise level, and the noise level of the noisy audio segment is determined as the medium-intensity noise level.

[0070] Specifically, in one embodiment, the above step S104 specifically includes the following steps:

[0071] Step S401: Obtain the proportion of different noise intensity levels in each noise audio segment. Specifically, by calculating the proportion of the noise audio segments belonging to the low-intensity noise level in the total number of noise audio segments in the entire speech data to be analyzed, as well as the corresponding proportions of the noise audio segments of the medium-intensity noise level and the high-intensity noise level.

[0072] Step S402: Determine the noise level evaluation result of the speech data to be analyzed according to the proportion of different noise intensity levels and the preset proportion evaluation index. Specifically, the preset proportion evaluation index can be set according to actual needs. For example, the noise intensity level with the largest proportion is used as the noise level evaluation result of the speech data to be analyzed. Or, weights can also be set for the proportions of different noise intensity levels, and then the weighted proportions are compared, and the noise intensity level with the largest weighted proportion is used as the noise level evaluation result of the speech data to be analyzed, etc. The present invention is not limited thereto.

[0073] In the embodiment of the present invention, in the case of paying more attention to the high-intensity noise that affects students' learning, in order to improve the sensitivity of the noise level evaluation result to high-intensity noise, the proportion of the high-intensity noise level is obtained by comprehensively considering the noise level results of each noise audio segment. According to the relationship between the proportion of the high-intensity noise level and the preset high-intensity noise level proportion range in the preset proportion evaluation index, the noise level evaluation result of the speech data to be analyzed is determined. When the proportion of the high-intensity noise level is less than the minimum value of the preset high-intensity noise level proportion range in the preset proportion evaluation index, it is determined that the noise level evaluation result is low noise level; when the proportion of the high-intensity noise level is within the preset high-intensity noise level proportion range in the preset proportion evaluation index, it is determined that the noise level evaluation result is medium noise level; when the proportion of the high-intensity noise level is greater than the maximum value of the preset high-intensity noise level proportion range in the preset proportion evaluation index, it is determined that the noise level evaluation result is high noise level. In practical applications, the preset high-intensity noise level proportion range can also be determined by setting two thresholds T L 、T H . If it is less than T L , it indicates that the overall noise level of the speech data to be analyzed is low. If it is between T L and T H , it indicates that the overall noise level of the speech data to be analyzed is medium. If it is greater than T H , it indicates that the overall noise level of the speech data to be analyzed is high.

[0074] By performing the above steps, the voice noise analysis method provided by the embodiments of the present invention obtains the voice data to be analyzed; extracts the noise audio segments containing only noise from the voice data to be analyzed; determines the noise intensity levels corresponding to the respective noise audio segments based on the noise intensity indexes of each noise audio segment and the preset noise intensity classification levels; and determines the noise level evaluation result of the voice data to be analyzed according to the distribution of the noise intensity levels corresponding to the respective noise audio segments. Thus, by calculating the noise intensity indexes of the noise audio segments containing only noise, the noise intensity levels of the respective noise audio segments are analyzed separately, and then the noise level evaluation result of the entire voice data to be analyzed is determined according to the distribution of the noise intensity levels of all the noise audio segments, avoiding the influence of the normal voice in the voice data to be analyzed, achieving an objective evaluation of the noise level of the voice data to be analyzed, and without the need to refer to an audio, having a wider application range and being able to accurately reflect the noise conditions in various scenarios.

[0075] The embodiments of the present invention also provide a voice noise analysis system, as Figure 3 shown. The voice noise analysis system includes:

[0076] An acquisition module 101, configured to acquire the voice data to be analyzed. For the detailed content, refer to the relevant description of step S101 in the above method embodiment, and details will not be described herein again.

[0077] A noise extraction module 102, configured to extract the noise audio segments containing only noise from the voice data to be analyzed. For the detailed content, refer to the relevant description of step S102 in the above method embodiment, and details will not be described herein again.

[0078] A noise estimation module 103, configured to determine the noise intensity levels corresponding to the respective noise audio segments based on the noise intensity indexes of each noise audio segment and the preset noise intensity classification levels. For the detailed content, refer to the relevant description of step S103 in the above method embodiment, and details will not be described herein again.

[0079] A noise statistics module 104, configured to determine the noise level evaluation result of the voice data to be analyzed according to the distribution of the noise intensity levels corresponding to the respective noise audio segments. For the detailed content, refer to the relevant description of step S104 in the above method embodiment, and details will not be described herein again.

[0080] Through the collaborative cooperation of the above-mentioned various components, the voice noise analysis system provided by the embodiment of the present invention obtains the voice data to be analyzed; extracts the noise audio segments containing only noise from the voice data to be analyzed; determines the noise intensity levels corresponding to the respective noise audio segments based on the noise intensity indicators of each noise audio segment and the preset noise intensity division levels; and determines the noise level evaluation result of the voice data to be analyzed according to the distribution of the noise intensity levels corresponding to the respective noise audio segments. Thus, by calculating the noise intensity indicators of the noise audio segments containing only noise, the noise intensity levels of the respective noise audio segments are analyzed individually, and then the noise level evaluation result of the entire voice data to be analyzed is determined according to the distribution of the noise intensity levels of all the noise audio segments, avoiding the influence of the normal voice in the voice data to be analyzed, achieving an objective evaluation of the noise level of the voice data to be analyzed, and without the need to refer to an audio, having a wider application range and being able to accurately reflect the noise conditions in various scenarios.

[0081] According to an embodiment of the present invention, there is also provided an electronic device, as Figure 4 shown. The electronic device may include a processor 901 and a memory 902, where the processor 901 and the memory 902 may be connected through a bus or other means. Figure 4 Taking connection through a bus as an example.

[0082] The processor 901 may be a central processing unit (CPU). The processor 901 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., or a combination of the above types of chips.

[0083] The memory 902, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the method in the method embodiment of the present invention. The processor 901 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 902, that is, implements the method in the above method embodiment.

[0084] The memory 902 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created by the processor 901 and the like. In addition, the memory 902 may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 902 may optionally include a memory remotely disposed relative to the processor 901, and these remote memories may be connected to the processor 901 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0085] One or more modules are stored in the memory 902 and, when executed by the processor 901, perform the methods in the above method embodiments.

[0086] For specific details of the above electronic device, reference may be made to the corresponding related descriptions and effects in the above method embodiments for understanding, and details are not described herein again.

[0087] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. Among them, the storage medium may be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium may also include a combination of the above types of memories.

[0088] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A method for voice noise analysis, characterized in that, it includes: Obtain the voice data to be analyzed; Extract a noise audio segment containing only noise from the voice data to be analyzed; The extracting a noise audio segment containing only noise from the voice data to be analyzed includes: dividing the voice data to be analyzed into multiple audio segments based on the total duration of the voice data to be analyzed and a preset extraction duration period; converting each audio segment into a magnitude spectrum; inputting the magnitude spectrum corresponding to each audio segment into a preset noise classification model to obtain the probability of each audio segment containing only noise; screening out the noise audio segments containing only noise from the audio segments based on a preset probability threshold, and the preset noise classification module is trained using known audio segments; Determine the noise intensity level corresponding to each noise audio segment based on the noise intensity index of each noise audio segment and a preset noise intensity classification level; The determining the noise intensity level corresponding to each noise audio segment based on the noise intensity index of each noise audio segment and a preset noise intensity classification level includes: calculating the noise intensity index corresponding to each noise audio segment respectively; obtaining the noise intensity index ranges corresponding to different noise intensity levels in the preset noise intensity classification level; determining the current noise intensity index range corresponding to the current noise audio segment according to the noise intensity index corresponding to the current noise audio segment; determining the noise intensity level corresponding to the current noise audio segment as the noise intensity level corresponding to the current noise intensity index range; Determine the noise level evaluation result of the voice data to be analyzed according to the distribution of the noise intensity levels corresponding to each noise audio segment; The determining the noise level of the voice data to be analyzed according to the distribution of the noise intensity levels corresponding to each noise audio segment includes: obtaining the proportion of different noise intensity levels in each noise audio segment; determining the noise level evaluation result of the voice data to be analyzed according to the proportion of different noise intensity levels and a preset proportion evaluation index.

2. The method according to claim 1, characterized in that, the noise intensity levels include: high-intensity noise level, medium-intensity noise level, and low-intensity noise level.

3. The method according to claim 2, characterized in that, the determining the noise level evaluation result of the voice data to be analyzed according to the proportion of different noise intensity levels and a preset proportion evaluation index includes: Obtain the proportion of the high-intensity noise level; Determine the noise level evaluation result of the voice data to be analyzed according to the relationship between the proportion of the high-intensity noise level and the preset high-intensity noise level proportion range in the preset proportion evaluation index.

4. The method according to claim 3, characterized in that, the noise level evaluation result includes: low noise level, moderate noise level, and high noise level, where when the proportion of the high-intensity noise level is less than the minimum value of the preset high-intensity noise level proportion range in the preset proportion evaluation index, it is determined that the noise level evaluation result is a low noise level; When the proportion of the high-intensity noise level is within the preset proportion range of the high-intensity noise level in the preset proportion evaluation index, it is determined that the noise level evaluation result is medium noise level; When the proportion of the high-intensity noise level is greater than the maximum value of the preset proportion range of the high-intensity noise level in the preset proportion evaluation index, it is determined that the noise level evaluation result is high noise level.

5. A voice noise analysis system, Characterized in that, Comprising: An acquisition module, configured to acquire voice data to be analyzed; A noise extraction module, configured to extract a noise audio segment containing only noise from the voice data to be analyzed; The extracting a noise audio segment containing only noise from the voice data to be analyzed includes: dividing the voice data to be analyzed into multiple audio segments based on the total duration of the voice data to be analyzed and a preset extraction duration period; converting each audio segment into a magnitude spectrum; inputting the magnitude spectrum corresponding to each audio segment into a preset noise classification model to obtain the probability of each audio segment containing only noise; screening out the noise audio segments containing only noise from the audio segments based on a preset probability threshold, and the preset noise classification module is trained using known audio segments; A noise estimation module, configured to determine the noise intensity level corresponding to each of the noise audio segments based on the noise intensity index of each noise audio segment and a preset noise intensity division level; The determining the noise intensity level corresponding to each of the noise audio segments based on the noise intensity index of each noise audio segment and a preset noise intensity division level includes: respectively calculating the noise intensity index corresponding to each noise audio segment; obtaining the noise intensity index ranges corresponding to different noise intensity levels in the preset noise intensity division level; determining the current noise intensity index range corresponding to the current noise audio segment according to the noise intensity index corresponding to the current noise audio segment; determining the noise intensity level corresponding to the current noise intensity index range as the noise intensity level of the current noise audio segment; A noise statistics module, configured to determine the noise level evaluation result of the voice data to be analyzed according to the distribution of the noise intensity levels corresponding to each of the noise audio segments; The determining the noise level of the voice data to be analyzed according to the distribution of the noise intensity levels corresponding to each of the noise audio segments includes: obtaining the proportion of different noise intensity levels in each of the noise audio segments; determining the noise level evaluation result of the voice data to be analyzed according to the proportion of different noise intensity levels and a preset proportion evaluation index.

6. An electronic device, Characterized in that, Comprising: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method according to any one of claims 1-4.

7. A computer-readable storage medium, Characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • System and method for noise activity detection

    CN101821971A

  • Audio processing method and device, computer equipment and storage medium

    CN111986691A

  • Ambient noise estimation device, sound volume adjusting device, ambient noise estimation method, and ambient noise estimation program

    JP2014030140A