An artificial intelligence-based noise processing method and system

By extracting and analyzing feature information from historical speech noise processing big data, and combining smoothing processing and semantic analysis, the problem of incomplete noise removal in noise processing is solved, and the efficient, accurate extraction and clear presentation of target speech information are achieved.

CN119694327BActive Publication Date: 2026-05-15SHENZHEN HANFEIKE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN HANFEIKE TECH CO LTD
Filing Date
2024-12-03
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively remove background noise when the difference between the ambient noise and the target speech is small, resulting in poor noise reduction and hindering the acquisition of target speech information.

Method used

By extracting background noise features from historical speech noise processing big data under different environments, background noise is filtered out using big data feature data, and target speech information is deeply extracted by combining smoothing processing and semantic analysis.

Benefits of technology

It achieves efficient and accurate removal of most noise in different environments, ensuring the clarity and accuracy of the target speech information and avoiding the impact of noise on information transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119694327B_ABST
    Figure CN119694327B_ABST
Patent Text Reader

Abstract

The application provides an artificial intelligence-based noise processing method and system, and relates to the technical field of noise processing. The method comprises the following steps: obtaining historical voice noise processing data, performing noise feature extraction based on an environmental background to form background noise feature data; performing de-noising processing on the voice information to be processed according to the background noise feature data to form initial de-noised voice information; and performing noise screening analysis based on smoothing and semantics on the initial de-noised voice information to extract target voice information. The method can more effectively identify and screen most of the noise in the recorded voice by using big data to establish the feature data of the environmental background noise, so that clear target voice information can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of noise processing technology, and more specifically, to a noise processing method and system based on artificial intelligence. Background Technology

[0002] Speech acquisition involves extracting the target speech from a complex background environment. Due to the presence of background noise, a lot of noise is mixed in, making the original recorded speech quite mixed and unable to clearly capture the target speech information. Therefore, noise reduction processing is needed to clearly highlight the target speech.

[0003] Currently, most speech noise reduction processing is based on the large difference in sound characteristics between the background noise and the target speech. However, when the difference is relatively small, the noise reduction effect is no longer so effective, and the noise reduction speech still has noise that affects the acquisition of the target speech, so it cannot achieve a significant noise reduction effect.

[0004] Therefore, designing an artificial intelligence-based noise processing method and system that utilizes big data to establish characteristic data of environmental background noise, thereby enabling more effective identification and removal of noise in recorded speech and ensuring the acquisition of clear and distinct target speech information, is an urgent problem to be solved. Summary of the Invention

[0005] The purpose of this invention is to provide an artificial intelligence-based noise processing method. By extracting feature information of background noise in different environments from historical speech noise processing big data, the method achieves background noise removal of the speech to be processed using big data feature data. Since the extracted feature data comes from big data, it can avoid the noise information loss that exists in a single denoising feature to a certain extent, so that the noise feature data in different environments can basically cover most of the noise in the corresponding environmental scenarios. At the same time, on the basis of removing most of the environmental noise, the method uses smoothing processing and semantic analysis of speech information to achieve deep extraction of different types of target speech information, completely extracting them from the still remaining small amount of noise information, ensuring the accuracy of the extracted target speech information, making the formed speech information clearer and more concise, and effectively avoiding the impact of noise on its information transmission.

[0006] The present invention also aims to provide an artificial intelligence-based noise processing system. This system is configured to collect environmental background noise feature information by utilizing historical speech noise processing big data to screen out environmental background noise of the speech information to be processed, and to accurately and reasonably extract the target speech information. This system achieves more efficient and accurate noise reduction processing, and is an important material basis for realizing noise reduction processing of the speech to be processed.

[0007] In a first aspect, the present invention provides a noise processing method based on artificial intelligence, comprising acquiring historical speech noise processing data, extracting noise features based on environmental background to form background noise feature data; performing denoising processing on the speech information to be processed based on the background noise feature data to form initial denoised speech information; and performing noise filtering analysis based on smoothing and semantics on the initial denoised speech information to extract target speech information.

[0008] In this invention, the method extracts feature information of background noise in different environments from historical speech noise processing big data, thereby achieving background noise removal of the speech to be processed using big data feature data. Since the extracted feature data comes from big data, it can avoid the loss of noise information in a single denoising feature to a certain extent, so that the noise feature data in different environments can basically cover most of the noise in the corresponding environmental scene. At the same time, on the basis of removing most of the environmental noise, it uses smoothing processing and semantic analysis of speech information to achieve deep extraction of different types of target speech information, completely extracting it from the still remaining small amount of noise information, ensuring the accuracy of the extracted target speech information, making the formed speech information clearer and more concise, and effectively avoiding the impact of noise on its information transmission.

[0009] One possible approach is to acquire historical speech noise processing data, perform noise feature extraction based on environmental background, and form background noise feature data. This includes: performing volume analysis on different historical speech processing information in the historical speech noise processing data for the target information of the historical speech, forming noise volume differentiation feature data; and extracting background noise features under different environmental conditions within the volume range of the target information of the historical speech for different historical speech processing information, forming background noise feature data for the target volume range.

[0010] In this invention, the extraction of environmental background noise feature information from historical speech noise processing data mainly considers two aspects. Firstly, in the initial stages of historical noise removal, the environmental noise contained in the speech information to be processed will have a relatively large volume of noise that differs significantly from the sound features of the target speech information. Therefore, initial screening can be performed. Since this significant feature gap is frequently found in the collected speech data, historical noise feature data can be used to fully extract this feature information, enabling the rapid removal of noise with a large difference in feature information from the target speech information. Here, the feature information of the significantly different noise is mainly considered in terms of noise volume. It is understood that the target speech information to be collected must be clearly and completely captured during the acquisition process, and clear and complete capture requires a certain range of volume. Therefore, by analyzing the volume, noise covered by volumes that do not correspond to the volume range of the target speech information can be screened out. Secondly, considering that the forms of noise exist in different background environments are different, there are significant differences not only in non-steady-state noise but also in steady-state noise. Therefore, extracting noise feature information based on different environmental backgrounds can ensure that the extracted noise feature data is more targeted, so that subsequent noise reduction processing can more effectively remove noise after reference.

[0011] As one possible implementation, volume analysis is performed on different historical speech processing information in historical speech noise processing data to target historical speech information, forming noise volume distinguishing feature data. This includes: determining the volume range of historical speech target information corresponding to different historical speech processing information in historical speech noise processing data, forming historical target speech volume ranges corresponding to different historical speech processing information; and performing a union operation on all historical target speech volume ranges to form noise volume distinguishing feature ranges.

[0012] In this invention, for the collection of noise feature information that has obvious differences in sound feature information with the target speech information, the aim is to cover these noise feature data as much as possible under big data, so as to effectively remove this part of the noise in the future. After obtaining the noise feature information with obvious feature differences in the corresponding speech information to be processed based on the volume range of the historical speech target information, the union operation is performed on all the noise feature information obtained to ensure that all this type of noise feature information in the historical data can be covered.

[0013] As one possible implementation, background noise features under different environmental conditions within the volume range of the target historical speech information are extracted from different historical speech processing information to form background noise feature data for the target volume range. This includes: extracting corresponding speech information within the volume range of the target historical speech information from different historical speech processing information to form corresponding volume range historical speech information; clustering the historical speech information of different volume ranges based on different background environment types to form a set of historical speech information for each environment type volume range; and performing a union operation on the sound wave information of all historical speech information in the set of historical speech information for different environment types to form background noise feature data for the target volume range corresponding to the background environment.

[0014] In this invention, for noise with the same volume range as the target speech information, the remaining noise information is obtained by utilizing the historical speech target information already labeled in the historical speech processing information. Since the noise feature information generated by different environmental backgrounds varies significantly, the obtained noise information is clustered only under the same background environment. The definition of the background environment can be made according to actual needs. Clustering of noise feature information based on different environmental backgrounds can further improve the targeting of noise feature information and enhance the accuracy of subsequent noise reduction processing of the speech information to be processed. It should be noted that for this type of noise feature information, sound wave information is used as the aspect of feature information acquisition. This ensures that the acquired feature information is multi-dimensional, such as frequency, period interval, amplitude, etc., which can be obtained from sound wave information. Comprehensive feature information from multiple aspects is more recognizable, making subsequent noise reduction using feature information more accurate.

[0015] As one possible implementation, the speech information to be processed is denoised based on background noise feature data to form initial denoised speech information, including: extracting speech information within the noise volume distinction feature range based on the noise volume distinction feature range to form target processing interval speech information; and denoising the target processing interval speech information based on background noise feature data of different target volume intervals to form initial denoised speech information.

[0016] In this invention, after extracting noise feature information based on historical speech processing information, the extracted noise feature information can be used for noise reduction of the speech information to be processed. Depending on the form of the extracted background noise feature data, the noise reduction of the speech information to be processed includes two sequential steps. First, noise reduction is performed on the speech information based on the volume feature range by using noise volume to distinguish feature ranges, that is, removing noise portions where the feature information differs significantly from the feature information of the target language information. Second, background noise feature data of the target volume range is used to remove most of the environmental background noise from the speech information in the target processing range. Of course, before removing noise, the environmental background of the speech information in the target processing range needs to be determined, and then the corresponding background noise feature data of the target volume range is extracted for noise reduction.

[0017] As one possible implementation, background noise denoising is performed on the speech information in the target processing range based on background noise feature data of different target volume ranges to form initial denoised speech information. This includes: extracting the sound wave information of the speech information in the target processing range and performing matching analysis with background noise feature data of different target volume ranges to determine the background noise feature data of the target volume range with the most matching sound waves, and using the background environment corresponding to the background noise feature data of the target volume range as the background environment corresponding to the speech information in the target processing range; removing the speech information corresponding to all sound wave information that matches the background noise feature data of the target volume range from the speech information in the target processing range to form initial denoised speech information.

[0018] In this invention, background noise feature data of the corresponding target volume range is used to perform background noise denoising on the speech information of the target processing range. The specific processing method is to determine whether there is noise feature information in the speech information of the target processing range that matches the background noise feature data of the target volume range. Only if they match can it be determined that the matched speech information belongs to environmental noise. This also avoids the loss of target speech information due to the removal of target speech information to a certain extent.

[0019] As one possible implementation, noise filtering analysis based on smoothing and semantics is performed on the initial denoised speech information to extract the target speech information, including: smoothing noise filtering on the initial denoised speech information to determine the smoothed speech information to be confirmed, and performing music-based judgment analysis on the smoothed speech information to be confirmed to form smoothed speech type labeling information; semantic noise filtering analysis is performed on the initial denoised speech information to form the target language speech information.

[0020] In this invention, after denoising the speech information to be processed using background noise feature data, it is considered that the remaining speech information may still contain some new noise information not yet covered by big data, as well as noise whose feature information is very close to the target speech feature information. Therefore, reasonable filtering is required to extract the target speech information. The extraction of target speech information is carried out from two aspects. First, if the target speech information is speech with obvious regularity, such as different types of music, then since such regular speech is usually smoothed, its smoothness is higher than that of noise. Therefore, smoothness-based analysis can extract speech with obvious regularity. Second, if the target speech is language information, then the target speech has obvious semantic information compared to noise, and more language information can be obtained compared to human voices mixed in with noise. Therefore, this type of target speech information can be extracted by analyzing the speech information.

[0021] One possible implementation involves smoothing and filtering out noise from the initial denoised speech information to identify the speech information to be smoothed and confirmed. Then, based on musical characteristics, the speech information to be smoothed and confirmed is analyzed to form smoothed speech type labeling information. This includes: extracting the spectra of all speech in the initial denoised speech information to form the original speech spectrum of the corresponding speech; performing derivatives on the spectral lines in the original speech spectrum to form the corresponding spectral line derivative information; identifying the speech information corresponding to the spectral line with the fewest non-differentiable position points in the derivative information of different spectral lines as the speech information to be smoothed and confirmed; setting a frequency change threshold, and performing the following analysis on the speech information to be smoothed and confirmed: if the maximum frequency difference of the spectral lines corresponding to the speech information to be smoothed and confirmed does not exceed the frequency change threshold, then the speech information to be smoothed and confirmed is labeled as non-musical smoothed speech information; if the maximum frequency difference of the spectral lines corresponding to the speech information to be smoothed and confirmed exceeds the frequency change threshold, then the speech information to be smoothed and confirmed is labeled as musical smoothed speech information.

[0022] In this invention, the extraction of regular target speech information is mainly achieved by performing smoothness analysis on the spectra of all speech in the initial denoised speech information. If the target speech is regular, due to its high smoothness, the derivative information formed after spectral derivativeization will generally not have many non-differentiable positions. Therefore, the speech information with the fewest statistically significant non-differentiable positions is most likely to be the regular target speech information. Of course, to determine whether there is such obvious regularity, it is also necessary to judge the range of frequency changes, since there is a significant difference in the range of frequency changes between a single tone and a musical piece.

[0023] As one possible approach, semantic noise removal analysis is performed on the initial denoised speech information to form target language speech information. This includes: extracting language phrases from different speech information in the initial denoised speech information to form corresponding speech phrase sets; and determining the speech information with the largest number of language phrases in all speech phrase sets as the target language speech information.

[0024] In this invention, the noise removal of the initial denoised speech information through semantic analysis mainly involves extracting the language phrases conveyed by different speech information. Considering that if the target speech is a human voice, the phrases are coherent and obvious, the target speech information can be determined by the size of the number of phrases.

[0025] Secondly, the present invention provides an artificial intelligence-based noise processing system, configured to: acquire historical speech noise processing data, extract noise features based on environmental background to form background noise feature data; perform denoising processing on the speech information to be processed based on the background noise feature data to form initial denoised speech information; and perform noise filtering analysis based on smoothing and semantics on the initial denoised speech information to extract target speech information.

[0026] In this invention, the system is configured to collect environmental background noise feature information by utilizing historical speech noise processing big data to achieve the screening of environmental background noise in the speech information to be processed, and to accurately and reasonably extract the target speech information. This system achieves more efficient and accurate noise reduction processing, and is an important material basis for realizing noise reduction processing of the speech to be processed.

[0027] The beneficial effects of the noise processing method and system based on artificial intelligence provided by this invention are as follows:

[0028] This method extracts feature information of background noise in different environments from historical speech noise processing big data, thereby filtering out background noise from the speech to be processed using big data feature data. Since the extracted feature data comes from big data, it can avoid the loss of noise information in individual denoising features to a certain extent, so that the noise feature data in different environments can basically cover most of the noise in the corresponding environmental scenarios. At the same time, on the basis of removing most of the environmental noise, it uses smoothing processing and semantic analysis of speech information to achieve deep extraction of different types of target speech information, completely extracting them from the small amount of noise information that still exists, ensuring the accuracy of the extracted target speech information, making the formed speech information clearer and more concise, and effectively avoiding the impact of noise on its information transmission.

[0029] This system is configured to collect environmental background noise feature information using historical speech noise processing big data to screen out environmental background noise from the speech information to be processed, and to accurately and reasonably extract the target speech information. It achieves more efficient and accurate noise reduction processing, and is an important material basis for realizing noise reduction processing of speech. Attached Figure Description

[0030] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 A flowchart illustrating the steps of an artificial intelligence-based noise processing method provided in an embodiment of the present invention;

[0032] Figure 2 A schematic diagram of volume range feature extraction based on artificial intelligence provided in an embodiment of the present invention;

[0033] Figure 3 This is a schematic diagram of an artificial intelligence-based smoothing and denoising sound wave provided for an embodiment of the present invention. Detailed Implementation

[0034] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.

[0035] Speech acquisition involves extracting the target speech from a complex background environment. Due to the presence of background noise, a lot of noise is mixed in, making the original recorded speech quite mixed and unable to clearly capture the target speech information. Therefore, noise reduction processing is needed to clearly highlight the target speech.

[0036] Currently, most speech noise reduction processing is based on the large difference in sound characteristics between the background noise and the target speech. However, when the difference is relatively small, the noise reduction effect is no longer so effective, and the noise reduction speech still has noise that affects the acquisition of the target speech, so it cannot achieve a significant noise reduction effect.

[0037] refer to Figures 1-3This invention provides an artificial intelligence-based noise processing method. This method extracts feature information of background noise from historical speech noise processing big data under different environments, enabling the filtering of background noise from the speech to be processed using big data feature data. Since the extracted feature data comes from big data, it can avoid the loss of noise information present in individual denoising features to a certain extent, allowing noise feature data from different environments to basically cover most noise in the corresponding environmental scenarios. Simultaneously, based on removing most environmental noise, it utilizes smoothing processing and semantic analysis of speech information to achieve deep extraction of different types of target speech information, completely extracting it from the remaining small amount of noise information, ensuring the accuracy of the extracted target speech information, making the resulting speech information clearer and more concise, and effectively avoiding the impact of noise on its information transmission.

[0038] The noise processing method based on artificial intelligence specifically includes the following steps:

[0039] S1: Obtain historical speech noise processing data, extract noise features based on environmental background, and form background noise feature data.

[0040] Acquire historical speech noise processing data, perform noise feature extraction based on environmental background, and form background noise feature data, including: performing volume analysis on different historical speech processing information in the historical speech noise processing data for the target information of the historical speech, forming noise volume differentiation feature data; extracting background noise features under different environmental conditions within the volume range of the target information of the historical speech for different historical speech processing information, forming background noise feature data of the target volume range.

[0041] Extracting environmental background noise features from historical speech noise processing data mainly considers two aspects. First, the environmental noise initially present in the speech data to be processed during historical noise removal often exhibits a significant difference in sound characteristics compared to the target speech data. This initial screening can be performed. Since this significant feature difference is frequently found in the collected speech data, historical noise feature data can be used to fully extract this feature information, enabling rapid removal of noise with significant discrepancies between its features and the target speech data. Here, the feature information for noise with significant discrepancies primarily focuses on noise volume. It's understood that the target speech data must be clearly and completely captured during the acquisition process, requiring a certain volume range. Therefore, by analyzing volume, noise within the volume range that does not correspond to the target speech data can be filtered out. Secondly, considering that the forms of noise exist in different background environments are different, there are significant differences not only in non-steady-state noise but also in steady-state noise. Therefore, extracting noise feature information based on different environmental backgrounds can ensure that the extracted noise feature data is more targeted, so that subsequent noise reduction processing can more effectively remove noise after reference.

[0042] The volume analysis of different historical speech processing information in historical speech noise processing data is performed on the target information of historical speech to form noise volume discrimination feature data, including: determining the volume range of the target information of historical speech corresponding to different historical speech processing information in historical speech noise processing data, forming the volume range of the target speech corresponding to different historical speech processing information; performing a union operation on all the volume ranges of the target speech to form the noise volume discrimination feature range.

[0043] For the collection of noise feature information that has obvious differences in sound feature information with the target speech information, we consider to cover these noise feature data as much as possible under big data, so as to effectively remove this part of the noise in the future. After obtaining the noise feature information with obvious feature differences in the corresponding speech information to be processed based on the volume range of the historical speech target information, we perform a union operation on all the noise feature information obtained to ensure that all this type of noise feature information in the historical data can be covered.

[0044] Background noise features under different environmental conditions within the volume range of the target historical speech information are extracted from different historical speech processing information to form background noise feature data for the target volume range. This includes: extracting the corresponding speech information within the volume range of the target historical speech information from different historical speech processing information to form corresponding volume range historical speech information; clustering the historical speech information of different volume ranges based on different background environment types to form a set of historical speech information for each environment type volume range; and performing a union operation on the sound wave information of all historical speech information in the set of historical speech information for different environment types to form background noise feature data for the target volume range corresponding to the background environment.

[0045] For noise with the same volume range as the target speech information, the remaining noise information is obtained by utilizing the historical speech target information already labeled in the historical speech processing information. Since the noise feature information generated by different environmental backgrounds varies significantly, the obtained noise information is only clustered under the same background environment. The definition of the background environment can be determined according to actual needs. Clustering of noise feature information based on different environmental backgrounds can further improve the targeting of noise feature information and enhance the accuracy of subsequent noise reduction processing of the speech information to be processed. It should be noted that for this type of noise feature information, sound wave information is used as the aspect of feature acquisition. This ensures that the acquired feature information is multi-dimensional, such as frequency, period interval, and amplitude that can be obtained from sound wave information. Comprehensive feature information from multiple aspects is more recognizable, making subsequent noise reduction using feature information more accurate.

[0046] S2: Based on the background noise feature data, perform noise reduction processing on the speech information to be processed to form the initial denoised speech information.

[0047] Based on background noise feature data, the speech information to be processed is denoised to form initial denoised speech information, including: extracting speech information within the noise volume distinction feature range based on the noise volume distinction feature range to form speech information of the target processing interval; and denoising the speech information of the target processing interval based on background noise feature data of different target volume intervals to form initial denoised speech information.

[0048] After extracting noise features based on historical speech processing information, the extracted noise features can be used for noise reduction of the speech information to be processed. Depending on the form of the extracted background noise feature data, the noise reduction of the speech information to be processed involves two sequential steps. First, noise reduction is performed on the speech information based on the volume feature range, i.e., removing noise portions where the feature information differs significantly from the target language information. Second, background noise feature data within the target volume range is used to remove most of the environmental background noise from the speech information within the target processing range. Of course, before this removal, the environmental background of the speech information within the target processing range needs to be determined, and then the corresponding background noise feature data for the target volume range is extracted for noise reduction.

[0049] Based on background noise feature data of different target volume ranges, background noise denoising processing is performed on the speech information of the target processing range to form initial denoised speech information. This includes: extracting the sound wave information of the speech information of the target processing range and performing matching analysis with background noise feature data of different target volume ranges to determine the background noise feature data of the target volume range with the most matching sound waves, and using the background environment corresponding to the background noise feature data of the target volume range as the background environment corresponding to the speech information of the target processing range; removing the speech information corresponding to all sound wave information that matches the background noise feature data of the target volume range from the speech information of the target processing range to form initial denoised speech information.

[0050] The background noise feature data of the target volume range is used to denoise the speech information in the target processing range. The specific processing method is to determine whether there is noise feature information in the speech information in the target processing range that matches the background noise feature data of the target volume range. Only if they match can it be determined that the matched speech information belongs to environmental noise. This also avoids the loss of target speech information due to the removal of target speech information to a certain extent.

[0051] S3: Perform noise filtering analysis based on smoothing and semantics on the initial denoised speech information to extract the target speech information.

[0052] The initial denoised speech information is subjected to noise removal analysis based on smoothing and semantics to extract target speech information, including: smoothing noise removal of the initial denoised speech information to determine the smoothed speech information to be confirmed, and performing music characteristic-based judgment analysis on the smoothed speech information to be confirmed to form smoothed speech type labeling information; semantic noise removal analysis of the initial denoised speech information to form target language speech information.

[0053] After denoising the speech information using background noise feature data, it's considered that the remaining speech information may still contain new noise information not yet covered by big data, as well as noise whose feature information is very similar to the target speech feature information. Therefore, reasonable filtering is needed to extract the target speech information. This extraction is approached from two aspects. First, if the target speech information is a type of speech with obvious regularity, such as different types of music, then this type of regular speech is usually smoothed, resulting in a higher smoothness compared to noise. Therefore, smoothness-based analysis can extract speech with obvious regularity. Second, if the target speech is language information, then the target speech has significant semantic information compared to noise, and more language information can be obtained from human voices mixed in with noise. Therefore, this type of target speech information can be extracted through speech information analysis.

[0054] The initial denoised speech information is smoothed and noise is filtered out to identify the speech information to be smoothed. Based on musical characteristics, the speech information to be smoothed is analyzed to form smoothed speech type labeling information. This includes: extracting the spectra of all speech in the initial denoised speech information to form the original speech spectrum of the corresponding speech; performing derivatives on the spectral lines in the original speech spectrum to form the corresponding spectral line derivative information; identifying the speech information corresponding to the spectral line with the fewest non-differentiable position points in the derivative information of different spectral lines as the speech information to be smoothed; setting a frequency change threshold, the speech information to be smoothed is analyzed as follows: if the maximum frequency difference of the spectral lines corresponding to the speech information to be smoothed does not exceed the frequency change threshold, the speech information to be smoothed is labeled as non-musical smoothed speech information; if the maximum frequency difference of the spectral lines corresponding to the speech information to be smoothed exceeds the frequency change threshold, the speech information to be smoothed is labeled as musical smoothed speech information.

[0055] The extraction of regular target speech information is mainly achieved by collecting the spectrum of all speech in the initial denoised speech information and performing smoothness analysis. If the target speech is regular, due to its high smoothness, the derivative information formed after spectral derivativeization will generally not have many non-differentiable points. Therefore, the speech information with the fewest statistically significant non-differentiable points is most likely to be the regular target speech information. Of course, to determine whether there is such obvious regularity, it is also necessary to judge the magnitude range of frequency changes, since there is a significant difference in the magnitude range of frequency changes between a single tone and a concerto.

[0056] Semantic noise removal analysis is performed on the initial denoised speech information to form target language speech information, including: extracting language phrases from different speech information in the initial denoised speech information to form corresponding speech phrase sets; and determining the speech information with the largest number of language phrases in all speech phrase sets as the target language speech information.

[0057] Noise removal in semantic analysis of initial denoised speech information mainly involves extracting language phrases conveyed by different speech information. If the target speech is human voice, then the phrases are coherent and obvious, so the target speech information can be determined by the size of the number of phrases.

[0058] The present invention also provides an artificial intelligence-based noise processing system, which is configured to: acquire historical speech noise processing data, extract noise features based on environmental background to form background noise feature data; perform denoising processing on the speech information to be processed based on the background noise feature data to form initial denoised speech information; and perform noise filtering analysis based on smoothing and semantics on the initial denoised speech information to extract target speech information.

[0059] This system is configured to collect environmental background noise feature information using historical speech noise processing big data to screen out environmental background noise from the speech information to be processed, and to accurately and reasonably extract the target speech information. It achieves more efficient and accurate noise reduction processing, and is an important material basis for realizing noise reduction processing of speech.

[0060] In summary, the beneficial effects of the artificial intelligence-based noise processing method and system provided in the embodiments of the present invention are as follows:

[0061] This method extracts feature information of background noise in different environments from historical speech noise processing big data, thereby filtering out background noise from the speech to be processed using big data feature data. Since the extracted feature data comes from big data, it can avoid the loss of noise information in individual denoising features to a certain extent, so that the noise feature data in different environments can basically cover most of the noise in the corresponding environmental scenarios. At the same time, on the basis of removing most of the environmental noise, it uses smoothing processing and semantic analysis of speech information to achieve deep extraction of different types of target speech information, completely extracting them from the small amount of noise information that still exists, ensuring the accuracy of the extracted target speech information, making the formed speech information clearer and more concise, and effectively avoiding the impact of noise on its information transmission.

[0062] This system is configured to collect environmental background noise feature information using historical speech noise processing big data to screen out environmental background noise from the speech information to be processed, and to accurately and reasonably extract the target speech information. It achieves more efficient and accurate noise reduction processing, and is an important material basis for realizing noise reduction processing of speech.

[0063] In the embodiments of this application, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain piece of information is called the information to be instructed. In the specific implementation process, there are many ways to instruct the information to be instructed, such as, but not limited to, directly instructing the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly instruct the information to be instructed by instructing other information, where there is a relationship between the other information and the information to be instructed. It can also instruct only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction of specific information can be achieved by using a pre-agreed (e.g., protocol-defined) arrangement of various pieces of information, thereby reducing instruction overhead to some extent. At the same time, common parts of various pieces of information can be identified and uniformly indicated to reduce the instruction overhead caused by individually indicating the same information.

[0064] Furthermore, the specific indication method can also be any existing indication method, such as, but not limited to, the above-mentioned indication methods and their various combinations. Specific details of various indication methods can be found in existing technologies, and will not be repeated here. As described above, for example, when multiple pieces of information of the same type need to be indicated, the indication methods for different pieces of information may differ. In the specific implementation process, the required indication method can be selected according to specific needs. This application embodiment does not limit the selected indication method; therefore, the indication methods involved in this application embodiment should be understood to cover various methods that enable the party to be indicated to obtain the information to be indicated.

[0065] It should be understood that the information to be indicated can be sent as a whole or divided into multiple sub-information messages sent separately, and the sending period and / or timing of these sub-information messages can be the same or different. The specific sending method is not limited in this application embodiment. The sending period and / or timing of these sub-information messages can be predefined, for example, according to a protocol, or configured by the sending device by sending configuration information to the receiving device.

[0066] "Predefined" or "pre-configured" can be achieved by pre-saving corresponding codes, tables, or other means that can be used to indicate relevant information in the device. This application does not limit the specific implementation method. "Saving" can refer to saving in one or more memories. These memories can be separate installations or integrated into the encoder, decoder, processor, or communication device. Alternatively, some memories can be separately installed, while others are integrated into the decoder, processor, or communication device. The type of memory can be any form of storage medium, and this application does not limit this.

[0067] The “protocol” mentioned in the embodiments of this application may refer to a protocol family in the field of communication, a standard protocol with a similar protocol family frame structure, or a related protocol applied to future communication systems. The embodiments of this application do not specifically limit this.

[0068] In the embodiments of this application, descriptions such as "when," "under the circumstances," "if," and "if" all refer to the device making corresponding processing under certain objective circumstances, and are not limited to a specific time. They do not require the device to make a judgment action during implementation, nor do they imply any other limitations.

[0069] In the description of the embodiments of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can represent A or B. "And / or" in the embodiments of this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. Furthermore, in the description of the embodiments of this application, unless otherwise stated, "multiple" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. Additionally, to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or order of execution, and that "first," "second," etc., are not necessarily different. Furthermore, in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate that something is being used as an example, illustration, or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.

[0070] It should be understood that the processor in the embodiments of this application can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0071] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0072] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0073] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0074] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0075] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0076] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0077] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0078] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0079] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0080] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0081] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0082] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A noise processing method based on artificial intelligence, characterized in that, include: Acquire historical speech noise processing data, perform noise feature extraction based on environmental background, and form background noise feature data; Based on the background noise feature data, the speech information to be processed is denoised to form initial denoised speech information. The initial denoised speech information is subjected to noise filtering analysis based on smoothing and semantics to extract the target speech information; This includes acquiring historical speech noise processing data, extracting noise features based on environmental background, and forming background noise feature data, including: The volume of different historical speech processing information in the historical speech noise processing data is analyzed for the target information of the historical speech to form noise volume distinguishing feature data: Determine the volume range of the historical speech target information corresponding to different historical speech processing information in the historical speech noise processing data, and form the historical target speech volume range corresponding to different historical speech processing information; Perform a union operation on all the historical target speech volume ranges to form a noise volume distinguishing feature range; Background noise features are extracted from different historical speech processing information under different environmental conditions within the volume range of the target historical speech information to form background noise feature data for the target volume range: For different historical speech processing information, extract the corresponding speech information within the volume range of the historical speech target information to form the corresponding volume range historical speech information; Cluster the historical voice information of different volume ranges based on different background environment types to form a set of historical voice information of volume range by environment type. The sound wave information of all the historical voice information of the volume range in the historical voice information set of different environment types is combined to form the background noise feature data of the target volume range corresponding to the background environment.

2. The noise processing method based on artificial intelligence according to claim 1, characterized in that, The step of denoising the speech information to be processed based on the background noise feature data to form initial denoised speech information includes: Based on the noise volume differentiation feature range, the speech information to be processed is extracted within the noise volume differentiation feature range to form speech information of the target processing interval. Based on the background noise feature data of different target volume ranges, the speech information of the target processing range is subjected to background noise denoising processing to form initial denoised speech information.

3. The noise processing method based on artificial intelligence according to claim 2, characterized in that, The step of denoising the speech information in the target processing range based on background noise feature data of different target volume ranges to form initial denoised speech information includes: Extract the sound wave information of the speech information in the target processing interval, and perform matching analysis with the background noise feature data of different target volume intervals to determine the background noise feature data of the target volume interval with the most matching sound waves, and take the background environment corresponding to the background noise feature data of the target volume interval as the background environment corresponding to the speech information in the target processing interval. Remove all audio information in the target processing range that matches the background noise feature data of the target volume range from the audio information to form the initial denoised audio information.

4. The noise processing method based on artificial intelligence according to claim 3, characterized in that, The step of performing noise removal analysis based on smoothing and semantics on the initial denoised speech information to extract the target speech information includes: The initial denoised speech information is subjected to smooth noise removal to determine the smooth speech information to be confirmed, and the smooth speech information to be confirmed is subjected to judgment and analysis based on music characteristics to form smooth speech type labeling information. The initial denoised speech information is subjected to semantic noise removal analysis to form target language speech information.

5. The noise processing method based on artificial intelligence according to claim 4, characterized in that, The process of smoothing and filtering noise from the initial denoised speech information to determine the smoothed speech information to be confirmed, and then performing music-based judgment and analysis on the smoothed speech information to be confirmed to form smoothed speech type labeling information, includes: Extract the spectrum of all speech in the initial denoised speech information to form the original speech spectrum of the corresponding speech; The spectral lines in the original speech spectrum are derivatized to form the corresponding spectral line derivative information; The speech information corresponding to the spectrum line with the minimum number of non-differentiable position points in the derivative information of different spectrum lines is determined as the smoothed speech information to be confirmed. By setting a frequency variation threshold, the smoothed unconfirmed voice information is analyzed as follows: If the maximum frequency difference of the spectral lines corresponding to the smoothed speech information to be confirmed does not exceed the frequency change threshold, then the smoothed speech information to be confirmed is labeled as non-musical smoothed speech information. If the maximum frequency difference of the spectral lines corresponding to the smoothed unconfirmed speech information exceeds the frequency change threshold, then the smoothed unconfirmed speech information is labeled as music smoothed speech information.

6. The noise processing method based on artificial intelligence according to claim 5, characterized in that, The step of performing semantic noise removal analysis on the initial denoised speech information to form target language speech information includes: Language phrases are extracted from different speech information in the initial denoised speech information to form a corresponding speech phrase set; The speech information with the largest number of language phrases in the set of all the speech phrases is determined as the target language speech information.

7. A noise processing system based on artificial intelligence, characterized in that, Configured as: Acquire historical speech noise processing data, perform noise feature extraction based on environmental background, and form background noise feature data; Based on the background noise feature data, the speech information to be processed is denoised to form initial denoised speech information. The initial denoised speech information is subjected to noise filtering analysis based on smoothing and semantics to extract the target speech information; This includes acquiring historical speech noise processing data, extracting noise features based on environmental background, and forming background noise feature data, including: The volume of different historical speech processing information in the historical speech noise processing data is analyzed for the target information of the historical speech to form noise volume distinguishing feature data: Determine the volume range of the historical speech target information corresponding to different historical speech processing information in the historical speech noise processing data, and form the historical target speech volume range corresponding to different historical speech processing information; Perform a union operation on all the historical target speech volume ranges to form a noise volume distinguishing feature range; Background noise features are extracted from different historical speech processing information under different environmental conditions within the volume range of the target historical speech information to form background noise feature data for the target volume range: For different historical speech processing information, extract the corresponding speech information within the volume range of the historical speech target information to form the corresponding volume range historical speech information; Cluster the historical voice information of different volume ranges based on different background environment types to form a set of historical voice information of volume range by environment type. The sound wave information of all the historical voice information of the volume range in the historical voice information set of different environment types is combined to form the background noise feature data of the target volume range corresponding to the background environment.