Intelligent ward intercom call system with background noise suppression function

By processing audio data from intercom devices in frames and analyzing frequency domain features, crosstalk audio when multiple devices are used in the ward is distinguished and removed, thus solving the problem of aliasing of voice information and improving the accuracy and clarity of noise reduction.

CN120877758BActive Publication Date: 2025-12-09HUNAN SHANGYIKANG MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511404583.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-12-09
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

In a hospital ward setting, when multiple patients use intercom devices simultaneously, the electronic prompts from other intercom devices can cause crosstalk to the audio of one intercom device, resulting in aliasing of voice information. Existing methods cannot effectively distinguish between real human voices and electronic prompts, reducing the accuracy of noise reduction.

Method used

The audio data acquisition module performs frame-by-frame processing, and the audio segment division module analyzes the amplitude energy difference characteristics. The crosstalk index analysis module distinguishes effective information segments through frequency domain distribution and phase information. The audio denoising module performs targeted denoising processing to filter and remove crosstalk audio segments.

Benefits of technology

It improves the accuracy of background noise suppression, ensures the clarity of key information, avoids misjudgments caused by global processing, and enhances the noise reduction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877758B_ABST
    Figure CN120877758B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of audio denoising, and particularly relates to an intercom calling system with background noise suppression function for intelligent ward. The system comprises: an audio data acquisition module, which is used for acquiring initial audio data and performing frame processing to obtain audio frames; an audio segment division module, which is used for segmenting the initial audio data based on amplitude energy difference characteristics between the audio frames to obtain audio information segments; a crosstalk index analysis module, which is used for analyzing superposition conditions such as prompt tones in other intercom devices that may be contained in the audio information segments, analyzing frequency domain information and phase information of the audio information segments, distinguishing crosstalk audio segments and normal audio segments, and calculating crosstalk indexes of each crosstalk audio segment; and an audio denoising module, which is used for performing targeted denoising on the crosstalk audio segments according to the crosstalk indexes, improving the accuracy of background noise suppression, and effectively improving the denoising effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of audio denoising, and particularly relates to an intercom calling system with background noise suppression function for smart ward. BACKGROUND

[0002] The prior art usually uses a multi-microphone array technology on the whole audio, enhances the sound waves in the target direction and weakens the noise in other directions through phase adjustment, so as to improve the noise suppression effect. However, not all of the whole audio contains valid information, and in the actual ward scene, when multiple patients use the intercom equipment at the same time, the electronic prompt sound of other intercom equipment will cause crosstalk to the audio of a certain intercom equipment, resulting in the aliasing of voice information. The existing method performs global processing, which not only enhances useless calculation, but also cannot distinguish the real human voice and the electronic prompt sound in the superposition situation, so that the accuracy of the denoising effect is greatly reduced. SUMMARY

[0003] In order to solve the technical problem that in the actual ward scene, when multiple patients use the intercom equipment at the same time, the electronic prompt sound of other intercom equipment will cause crosstalk to the audio of a certain intercom equipment, resulting in the aliasing of voice information, and the existing method performs global processing, which not only enhances useless calculation, but also cannot distinguish the real human voice and the electronic prompt sound in the superposition situation, so that the accuracy of the denoising effect is greatly reduced, the purpose of the present application is to provide an intercom calling system with background noise suppression function for smart ward, and the technical scheme is as follows:

[0004] The present application provides an intercom calling system with background noise suppression function for smart ward, which comprises:

[0005] An audio data acquisition module is arranged in the ward to acquire initial audio data recorded when each intercom equipment is used and to perform frame processing to obtain audio frames.

[0006] An audio segment division module is arranged to analyze the amplitude energy difference characteristics between the audio frames, determine the segmentation points, divide the initial audio data, and obtain audio information segments.

[0007] A crosstalk index analysis module is arranged to determine the effective index of the audio information segment according to the amplitude variation characteristics and the frequency domain distribution characteristics in the audio information segment, compare the frequency domain characteristic differences and the phase information between the audio frames in the effective information segment, combine the effective index, determine the crosstalk probability for distinguishing the crosstalk audio segment and the normal audio segment, analyze the fundamental frequency difference characteristics and the confusion degree of the frequency domain distribution between the audio frames in the crosstalk audio segment, and combine the crosstalk probability to determine the crosstalk index of each crosstalk audio segment.

[0008] An audio denoising module is configured to perform denoising processing on the crosstalk audio segment based on the crosstalk index, so as to obtain denoised audio data.

[0009] Further, the method for obtaining the audio information segment comprises:

[0010] Obtaining the root mean square value of all amplitudes in each audio frame as the short-time energy value of each audio frame.

[0011] In all audio frames, the absolute value of the difference between the short-time energy values in each adjacent two audio frames is taken as the energy mutation value of the latter one of the two adjacent audio frames.

[0012] The average of the energy mutation values of all audio frames is taken as the energy change reference value, and the starting point of the audio frame with an energy mutation value greater than the energy change reference value is taken as the segmentation point.

[0013] In the initial audio data, the initial audio data is segmented based on the segmentation point, so as to obtain all audio information segments.

[0014] Further, the method for obtaining the effective index comprises:

[0015] Obtaining the zero-crossing rate of each audio information segment.

[0016] In each audio information segment, a frequency spectrum is obtained based on the short-time Fourier transform, and the fundamental frequency and harmonics are extracted therefrom, in all harmonics, the harmonics with a frequency being an integer multiple of the fundamental frequency are taken as effective harmonics, and the number of effective harmonics is taken as a quantity factor.

[0017] The product of the centroid of the frequency spectrum of each audio information segment and the zero-crossing rate is negatively correlated and normalized, and the value after the mapping is taken as an effective factor of each audio information segment.

[0018] The product of the effective factor and the quantity factor of each audio information segment is normalized, and the value after the normalization is taken as the effective index of each audio information segment.

[0019] Further, the method for obtaining the effective information segment comprises:

[0020] In all audio information segments, the audio information segment with an effective index greater than a preset effective threshold is taken as an effective information segment.

[0021] Further, the method for obtaining the crosstalk probability comprises:

[0022] When there are at least two complete audio frames in each valid information segment, in each valid information segment, phase angles of each audio frame are obtained based on a short-time Fourier transform, absolute values of differences between phase angles of each adjacent two audio frames are calculated as instantaneous phase differences, and a variance of all instantaneous phase differences is taken as a first crosstalk factor of each valid information segment;

[0023] In each valid information segment, harmonics of each audio frame are extracted based on a short-time Fourier transform, peak points of all harmonics are connected to form a spectral envelope line, a slope value of the spectral envelope line corresponding to each audio frame is obtained as a change value, in adjacent two audio frames, a normalized value of a difference between the change value of the former audio frame and the change value of the latter audio frame is taken as a change difference value, and a maximum value in all change difference values is taken as a second crosstalk factor of the valid information segment;

[0024] A normalized value of a ratio of a product of the first crosstalk factor and the second crosstalk factor of each valid information segment to the effective index of each valid information segment is taken as a crosstalk probability of each valid information segment.

[0025] When there are no two or more complete audio frames in the valid information segment, the crosstalk probability of the valid information segment is a preset value.

[0026] Further, the distinguishing of the crosstalk audio segment and the normal audio segment comprises:

[0027] In all valid information segments, a valid information segment with a crosstalk probability greater than a preset target threshold value is taken as a crosstalk audio segment, and a valid information segment with a crosstalk probability less than or equal to the target threshold value is taken as a normal audio segment.

[0028] Further, the crosstalk index acquisition method comprises:

[0029] In each crosstalk audio segment, a fundamental frequency of each audio frame is obtained based on a short-time Fourier transform, and all fundamental frequencies of the audio frames are arranged in time sequence to obtain a sorting sequence.

[0030] A first difference sequence of the fundamental frequencies in the sorting sequence is calculated, and a ratio of the crosstalk probability of each crosstalk audio segment to a maximum value in the first difference sequence is taken as a first crosstalk parameter.

[0031] Harmonics in each audio frame are extracted based on a short-time Fourier transform, in all harmonics, harmonics with frequencies that are not integer multiples of the fundamental frequency are taken as disturbing harmonics, and a value obtained by negatively correlating the number of the disturbing harmonics is taken as a second crosstalk parameter.

[0032] A normalized value of a product of the first crosstalk parameter and the second crosstalk parameter is taken as a crosstalk index of each crosstalk audio segment.

[0033] Further, the method for obtaining the de-noised audio data comprises:

[0034] Distinguishing the target audio segment and the non-target audio segment in all crosstalk audio segments by using the crosstalk index;

[0035] Taking the sum of the crosstalk index of each target audio segment and a preset parameter as an adjustment degree value, taking the product of the adjustment degree value and a preset attenuation depth value as an adjustment attenuation depth, and taking the adjustment attenuation depth of each target audio segment as an input of an IIR comb filter, so as to obtain a filtered audio segment;

[0036] In the initial audio data, taking the mean of the d-vectors of all normal audio segments as a comparison vector, and respectively obtaining the d-vector of each filtered audio segment and the d-vector of each non-target audio segment;

[0037] Taking the cosine similarity between the d-vector of each filtered audio segment and the comparison vector as a first similarity value, and taking the cosine similarity between the d-vector of each non-target audio segment and the comparison vector as a second similarity value;

[0038] Taking the filtered audio segment corresponding to the first similarity value greater than a preset similarity threshold and the non-target audio segment corresponding to the second similarity value greater than the preset similarity threshold as a to-be-processed audio segment, and taking the filtered audio segment corresponding to the first similarity value less than or equal to the preset similarity threshold and the non-target audio segment corresponding to the second similarity value less than or equal to the preset similarity threshold as a non-to-be-processed audio segment;

[0039] Taking the crosstalk index of each to-be-processed audio segment as a noise reduction weight, and using weighted spectral subtraction to subtract the amplitude spectrum of each to-be-analyzed audio segment, so as to obtain a de-noised audio segment;

[0040] In the initial audio data, replacing the corresponding to-be-processed audio segment with the de-noised audio segment and replacing the corresponding target audio segment with the non-to-be-processed audio segment, so as to obtain the de-noised audio data corresponding to the initial audio data.

[0041] Further, the distinguishing the target audio segment and the non-target audio segment comprises:

[0042] In all crosstalk audio segments, taking the crosstalk audio segment with the crosstalk index greater than a preset crosstalk threshold as the target audio segment, and taking the crosstalk audio segment with the crosstalk index less than or equal to the preset crosstalk threshold as the non-target audio segment.

[0043] Further, the preset attenuation depth value is 25.

[0044] The present application has the following beneficial effects:

[0045] Firstly in the audio data acquisition module, in the ward, the initial audio data recorded when each patient uses the intercom device is acquired and frame processing is performed to obtain audio frames. In the process of recording audio by the intercom system of the ward, when there is a sudden change in amplitude energy in a frame and the mutation degree is large, the possibility of change in tone color or content at the frame is large. Therefore, in the audio segment division module, the segment point detection is performed based on the amplitude energy difference characteristics between the audio frames, so that the audio information segments are segmented. In the process of recording audio by the intercom system of the smart ward, the information segments containing important information need to be focused on to ensure the intelligibility, and the degree of effective information contained in different audio information segments is different. Therefore, in the crosstalk index analysis module, the effective information segments containing patient voice are accurately identified by combining the amplitude variation characteristics and frequency domain distribution characteristics in the audio information segments, so as to avoid the misjudgment of key information caused by global processing. Further, since the phenomenon of superposition of patient voice and electronic prompt sound of other intercom devices may occur in part of the effective information segments, in this module, the voice and the electronic prompt sound of the rest of the intercom devices are further distinguished. The frequency domain feature difference, distribution and phase information between the audio frames in the effective information segments are jointly analyzed, so that the crosstalk audio segments can be screened out in the effective information segments and the crosstalk index of each crosstalk audio segment can be calculated to represent the signal aliasing problem when multiple devices are used in parallel. Finally, in the audio denoising module, the crosstalk audio segments are subjected to targeted denoising processing based on the crosstalk index, the accuracy of background noise suppression is improved, and the denoising effect is effectively improved. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art and the advantages thereof, a brief introduction will be given to the drawings needed in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on these drawings.

[0047] Figure 1 is a system block diagram of an intercom calling system with background noise suppression function for smart ward provided by an embodiment of the present application;

[0048] Figure 2 is a system structure schematic diagram of an intercom calling system with background noise suppression function for smart ward provided by an embodiment of the present application. DETAILED DESCRIPTION

[0049] In order to further clarify the technical means and effects taken by the present application to achieve the predetermined inventive purpose, the specific implementation, structure, features and effects of the intelligent ward intercom calling system with background noise suppression function according to the present application are described in detail below in combination with the drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.

[0051] The specific scheme of the intelligent ward intercom calling system with background noise suppression function provided by the present application is described below in combination with the drawings.

[0052] Please refer to Figure 1 , which shows the system block diagram of the intelligent ward intercom calling system with background noise suppression function provided by one embodiment of the present application. The system comprises an audio data acquisition module 101, an audio segment division module 102, a crosstalk index analysis module 103 and an audio denoising module 104.

[0053] The audio data acquisition module 101 is used to acquire the initial audio data recorded when the intercom equipment is used in the ward and perform frame division processing to obtain audio frames.

[0054] In the actual intelligent ward environment, the user of the intercom system is usually a patient or a caregiver. The user sends an intercom request and inputs voice information. The intercom system transmits it to the nurse station, and the medical staff in the nurse station answers and gives voice instructions. In this process, there may be multiple patients using intercom equipment at the same time, so that when a patient uses intercom equipment, the recorded audio may contain the output audio of the remaining intercom equipment. In the embodiment of the present application, the main purpose is to distinguish the human voice in the audio recorded by each intercom equipment and the electronic prompt tone output by the intercom calling equipment, so as to carry out targeted denoising, improve the accuracy of denoising, and ensure that the final denoised audio data is clearer and more complete.

[0055] Firstly, when the user sends an intercom request in the ward, such as pressing the recording button of the intercom equipment to start the intercom equipment, the intercom equipment starts recording, thereby acquiring the initial audio data recorded when each intercom equipment is used. The initial audio data is frame-processed. Specifically, the frame length is 25 ms and the frame shift is 10 ms, thereby obtaining all the audio frames.

[0056] It should be noted that the collection and acquisition of personal information data in the embodiments of the present application are all authorized by relevant users, the process does not violate relevant laws and regulations, and does not violate public order and good customs; in the frame processing process, the frame length and the frame shift can be adjusted according to the implementation scene, which is not limited here.

[0057] The audio segment division module 102 is configured to analyze the amplitude energy difference characteristics between audio frames, determine a segmentation point, divide the initial audio data, and obtain the audio information segment.

[0058] When the amplitude energy of an audio frame suddenly changes and the change degree is large, it is considered that the tone or content at the audio frame has a large possibility of change (i.e., there may be burst noise, for example, there is a situation that the remaining patients in the same ward use the intercom device), and the possibility of being an event boundary is large, so in this module, the amplitude energy difference characteristics between short-time audio frames can be analyzed to determine a segmentation point for cutting the continuous initial audio data into discrete audio information segments.

[0059] Preferably, in an embodiment of the present application, the method for obtaining the audio information segment comprises:

[0060] The root mean square value can reflect the fluctuation level of the audio frame energy, so the root mean square value of all amplitudes in each audio frame is obtained as the short-time energy value of each audio frame.

[0061] Then in all audio frames, the absolute value of the difference between the short-time energy values in each adjacent two audio frames is taken as the energy mutation value of the latter audio frame in the adjacent two audio frames, and the greater the energy mutation value, the more obvious the energy mutation between the adjacent two audio frames.

[0062] At this point, the energy mutation value corresponding to each audio frame can be obtained (the energy mutation value of the first audio frame is set to be the same as that of the second audio frame), and the greater the value, the more obvious the energy mutation at the audio frame, so the average of the energy mutation values of all audio frames is taken as the energy change reference value, and the starting point of the audio frame with the energy mutation value greater than the energy change reference value is taken as the segmentation point.

[0063] Finally, the initial audio data is segmented based on the segmentation point in the initial audio data, thereby obtaining all the audio information segments.

[0064] At this time, the segmentation of the initial audio data can be completed, and each audio information segment will contain an audio frame, and it should be noted that due to the frame shift in the division of the audio frame, there is an overlapping part between the audio frames, so in the audio information segment divided based on the audio frame, there may be two cases of incomplete audio frames and complete audio frames.

[0065] The crosstalk index analysis module 103 is configured to determine an effective index of the audio information segment according to the amplitude variation feature and the frequency domain distribution feature in the audio information segment, to filter effective information segments; in the effective information segments, to compare the frequency domain feature difference and the phase information between the audio frames, and to determine a crosstalk probability for distinguishing crosstalk audio segments and normal audio segments in combination with the effective index; and in the crosstalk audio segments, to analyze the fundamental frequency difference feature and the confusion degree of the frequency domain distribution between the audio frames, and to determine a crosstalk index of each crosstalk audio segment in combination with the crosstalk probability.

[0066] In a general ward, taking an audio segment obtained by the intercom system when a single user sends an intercom request and inputs voice information as an example, if the audio contains clear voice information of the user and can reflect the nursing needs of the user, it is considered that the degree of containing effective information in the audio is relatively high, and vice versa. Therefore, when the degree of containing effective information in the audio is relatively high, the waveform variation is more gentle (the zero-crossing rate is lower), the energy distribution is concentrated (the spectral centroid is biased to low frequency), and the harmonic structure is regular. Therefore, the degree of containing effective information in each audio information segment is quantified based on the foregoing logic to obtain an effective index, so as to filter effective information segments from all audio information segments.

[0067] Preferably, in an embodiment of the present application, the method for obtaining the effective index comprises:

[0068] The zero-crossing rate of each audio information segment is obtained. The zero-crossing rate specifically refers to the number of times that the audio signal crosses the zero level per unit time, that is, the number of times that the amplitude signs of adjacent two time instants are opposite. The ratio of the number of times to the length of the audio information segment is taken as the zero-crossing rate. The greater the zero-crossing rate, the more complex the variation of the audio information segment, and the lower the degree of containing effective information.

[0069] In each audio information segment, the frequency spectrum is obtained based on the short-time Fourier transform, and the fundamental frequency and the harmonics are extracted. In an ideal case, the frequency of the harmonics is an integer multiple of the fundamental frequency. However, due to the influence of noise and the like, the harmonic frequency may deviate slightly from the integer multiple. Therefore, among all the harmonics, the harmonics with a frequency that is an integer multiple of the fundamental frequency are taken as effective harmonics, and the number of the effective harmonics is taken as a quantity factor. The greater the quantity factor, the more effective information the audio information segment contains.

[0070] The centroid of the frequency spectrum of each audio information segment is obtained. The smaller the centroid, the more the frequency is concentrated in the low frequency region, and the more effective information is contained. Therefore, the product of the centroid of the frequency spectrum and the zero-crossing rate is negatively correlated and normalized to correct the logical relationship, and is taken as an effective factor of each audio information segment. The greater the effective factor, the higher the effective degree of the audio information segment. The negative correlation and normalization processing can be performed by the formula wherein, represents an exponential function with a natural constant e as a base, and x represents an independent variable.

[0071] Finally, the product of the effective factor and the quantity factor of each audio information segment is normalized as the effective index of each audio information segment. Based on the foregoing analysis and calculation, the greater the effective index, the higher the possibility of containing patient voice information in the audio information segment, that is, the greater the effective degree. In the subsequent process, in order to obtain clearer audio data, the audio information segment with a greater effective index needs to be given priority to denoising processing. The normalization is a technology known to those skilled in the art, and the selection of the normalization function can be linear normalization or standard normalization, and the specific normalization method is not limited herein.

[0072] It should be noted that the short-time Fourier transform in the embodiment of the present application is a known technology, and the specific process is not described herein.

[0073] At this point, the effective index of each audio information segment can be obtained, and the effective information segment can be screened from all audio information segments based on the index.

[0074] Preferably, in an embodiment of the present application, the method for obtaining the effective information segment comprises:

[0075] Based on the foregoing steps, the greater the effective index, the more effective information the audio information segment contains, and the more attention it needs in the subsequent process. Therefore, in all audio information segments, the audio information segment with an effective index greater than a preset effective threshold is regarded as an effective information segment.

[0076] It should be noted that the preset effective threshold is 0.5, and the specific value can be adjusted according to the implementation scene, which is not limited herein.

[0077] In the smart ward, when a patient (referred to as a target patient) uses the intercom system to input audio information in real time, if other patients in the same ward also use the intercom system at the same time, the audio data recorded by the target patient may contain interference of the audio output by the rest of the intercom system, for example, there may be electronic voice prompts (such as “please press the prompt tone operation”, “the call has been connected”, etc.) and audio processed by the intercom device of the medical staff output by the rest of the intercom system, which may interfere with the voice information recognition of the target patient. Therefore, in the effective information segment, the real human voice and the audio output by the rest of the intercom system need to be further distinguished.

[0078] First, in the valid information segment, distinguish the crosstalk audio segment and the normal audio segment: the output electronic audio in the intercom system (including the synthesized electronic medical prompt sound and the sound of medical staff after being processed by the intercom device) may be attenuated due to the bandwidth limitation of the intercom device, and there may be phase anomaly of the audio (the greater the fluctuation of the instantaneous phase difference), due to the noise introduced by the audio output by the rest of the intercom system, the degree of containing valid information in the audio segment interfered by the noise is lower than the real human voice audio of the target patient input in real time. The original human voice has rich frequency components and high energy, while the human voice feature after being processed by the intercom device is significantly weakened in the high frequency part, the overall frequency range is narrowed, and the energy is more concentrated in the low frequency band, that is, after being processed by the intercom device, there are phenomena such as signal energy weakening and high frequency information loss.

[0079] Based on the above analysis, the frequency domain feature difference and the phase information between the audio frames can be compared in the valid information segment, and the crosstalk probability can be determined for distinguishing the crosstalk audio segment and the normal audio segment by combining the effective index. Preferably, in an embodiment of the present application, the method for obtaining the crosstalk probability comprises:

[0080] Based on the foregoing analysis, if the electronic medical prompt sound of other intercom devices and the sound of medical staff after being processed by the intercom device are mixed in a certain valid information segment, then due to the signal superposition of the audio of the rest of the intercom devices after being reflected in the ward and the current audio data, the instantaneous phase difference will have obvious fluctuation, and due to the bandwidth limitation of the intercom device, high frequency truncation phenomenon is easy to occur when multiple intercom devices work at the same time. Therefore, in each valid information segment, the frequency domain features and the phase information of the audio frames are compared, and the crosstalk probability of each valid information segment can be determined by combining the effective index, which is used to represent the possibility of containing the output audio of other intercom devices.

[0081] As can be known in the audio segment division module 102, each audio information segment may contain complete audio frames or incomplete audio frames, so when the valid information segment contains at least two complete audio frames, in each valid information segment, the phase angle of each audio frame is obtained based on the short-time Fourier transform, and the absolute value of the difference between the phase angles of each adjacent two audio frames is calculated as the instantaneous phase difference. The variance of all instantaneous phase differences is taken as the first crosstalk factor of each valid information segment. The greater the first crosstalk factor, the more obvious the fluctuation of the phase difference between the audio frames in the valid information segment, which is consistent with the characteristics of mixing the electronic medical prompt sound of other intercom devices and the sound of medical staff after being processed by the intercom device, so the crosstalk probability will be greater.

[0082] Then, in each valid information segment, the harmonics of each audio frame are extracted based on short-time Fourier transform, and the peak points of all the harmonics are connected to form a spectral envelope line, and the slope value of the spectral envelope line corresponding to each audio frame is obtained as a change value, which can represent the frequency component information of each audio frame in the valid information segment. In adjacent two audio frames, the difference between the change value of the former audio frame and the change value of the latter audio frame is normalized as a change difference value, and the greater the change difference value, the more likely the audio data is truncated at high frequency. At this time, there is a change difference value between each adjacent two audio frames, and the maximum value of all the change difference values is taken as a second crosstalk factor of the valid information segment. The greater the second crosstalk factor, the more likely the high frequency truncation occurs in the valid information segment, and the greater the probability of crosstalk also occurs. Since the difference between the change values can be positive or negative, the normalization can use the function .

[0083] Finally, based on the foregoing analysis, the greater the first crosstalk factor, the greater the probability of crosstalk in the valid information segment. Similarly, the greater the second crosstalk factor, the greater the probability of crosstalk in the valid information segment. The smaller the effective index of the valid information segment, the less effective information it contains, and the attention needs to be appropriately reduced, that is, the probability of subsequent operation needs to be reduced. Therefore, the product of the first crosstalk factor and the second crosstalk factor of each valid information segment is normalized with the ratio of the effective index of each valid information segment, and the normalized value is taken as the crosstalk probability of each valid information segment. At this time, the greater the crosstalk probability, the more effective information the valid information segment contains, and the more likely it is mixed with other intercom device electronic medical prompt tones and medical staff voices processed by the intercom device, and more subsequent processing and analysis are needed. The normalization is a well-known technical means to those skilled in the art, and the selection of the normalization function can be linear normalization or standard normalization, and the specific normalization method is not limited herein.

[0084] When there are no two or more complete audio frames in the valid information segment, the analysis value is considered to be low in the embodiment of the present application, and therefore the crosstalk probability of such valid information segment is directly set to a preset value, wherein the preset value is 0.

[0085] At this point, the crosstalk probability of each valid information segment can be obtained, and the crosstalk audio segment and the normal audio segment can be distinguished based on the index.

[0086] Preferably, in an embodiment of the present application, the crosstalk audio segment and the normal audio segment are distinguished, comprising:

[0087] The greater the crosstalk probability of the effective information segment is, the higher the possibility of crosstalk is, and the more attention needs to be paid to subsequent denoising processing, so in all effective information segments, the effective information segment with a crosstalk probability greater than a preset target threshold is regarded as a crosstalk audio segment, and the effective information segment with a crosstalk probability less than or equal to the preset target threshold is regarded as a normal audio segment.

[0088] It should be noted that in this embodiment of the present application, the preset target threshold is 0.5, and the specific value can be adjusted according to the implementation scene, which is not limited herein.

[0089] In the smart ward, when the patient records the audio using the intercom system, there may be electronic voice prompts in the intercom system and audio of medical staff processed by the intercom device, and the interference degrees of the target patient voice information are different, so they are distinguished; the audio of the medical staff processed by the intercom device usually still retains the characteristics of continuous distribution of real human voice harmonics and the slight physiological fluctuations of the fundamental frequency, but after being processed by the intercom device, there may be signal distortion, so that there are non-integer multiple harmonics, and the electronic medical prompt tone of the intercom system is usually synthesized by an algorithm, and the high-frequency truncation and phase fluctuation are more obvious, and the harmonics are strictly integer multiples, and the fundamental frequency is completely stable. Therefore, in the crosstalk audio segment, the difference characteristics of the fundamental frequency between the audio frames in the frequency domain and the degree of confusion of the frequency domain distribution are analyzed, and the crosstalk probability is combined to determine the crosstalk index of each crosstalk audio segment, which is helpful for further dividing the crosstalk audio segment.

[0090] Preferably, in an embodiment of the present application, the method for obtaining the crosstalk index comprises:

[0091] If the crosstalk audio segment contains the electronic prompt tone output by the other intercom device, the fundamental frequency will be more stable and the number of non-integer multiple harmonics will be less (different from the audio of the medical staff processed by the intercom device, the electronic medical prompt tone in the intercom device is usually synthesized by an algorithm, and usually does not have the fundamental frequency fluctuation of natural human voice tone), and the high-frequency truncation and phase fluctuation are more obvious (i.e. the crosstalk probability is greater), so in this step, first, in each crosstalk audio segment, the fundamental frequency of each audio frame is obtained based on short-time Fourier transform, and the fundamental frequencies of all audio frames are arranged in time sequence to obtain a sorting sequence.

[0092] The first-order difference sequence of the fundamental frequency in the sorting sequence is calculated, the values in the first-order difference sequence can reflect the difference of the fundamental frequency of the audio frames in each audio segment, and the greater the value is, the greater the difference is, which can be regarded as the more obvious the fluctuation is, so the ratio of the crosstalk probability of each crosstalk audio segment to the maximum value in the first-order difference sequence is regarded as a first crosstalk parameter, and the greater the first crosstalk parameter is, the higher the possibility of the crosstalk audio segment containing the electronic prompt tone of the other intercom device is.

[0093] Then the harmonics in each audio frame are extracted based on short-time Fourier transform, among all the harmonics, the harmonics whose frequencies are not integer times of the fundamental frequency are regarded as disturbing harmonics, the more the number of disturbing harmonics, the lower the possibility of containing the electronic prompt tone of other intercom equipment, so the number of disturbing harmonics is negatively correlated and mapped, the logical relationship is corrected, and the second crosstalk parameter is obtained, the larger the second crosstalk parameter, the higher the possibility of containing the electronic prompt tone of other intercom equipment in the crosstalk audio segment. The negative correlation mapping processing here adopts the formula , wherein, represents the exponential function with the natural constant e as the base, and x represents the independent variable.

[0094] Finally, the product of the first crosstalk parameter and the second crosstalk parameter is normalized, and the value after normalization is taken as the crosstalk index of each crosstalk audio segment, the larger the crosstalk index, the higher the possibility of containing the electronic prompt tone of other intercom equipment in the crosstalk audio segment.

[0095] The audio denoising module 104 is configured to perform denoising processing on the initial audio data based on the crosstalk index to obtain denoised audio data.

[0096] The crosstalk index of each crosstalk audio can be calculated in the crosstalk index analysis module 103, in this module, the electronic prompt tone and the audio of the medical staff processed by the intercom equipment can be further distinguished in the crosstalk audio segment, so that targeted denoising processing is performed based on the crosstalk index, thereby obtaining denoised audio data.

[0097] Preferably, in an embodiment of the present application, the method for obtaining denoised audio data comprises:

[0098] Based on the foregoing modules, the effective information segment can be first screened out in the initial audio data, and then the crosstalk audio segment (containing the electronic prompt tone and the medical staff sound) and the normal audio segment can be distinguished in the effective information segment. Here, the target audio segment (containing the electronic prompt tone) and the non-target audio segment can be distinguished in the crosstalk audio segment: in all crosstalk audio segments, the crosstalk audio segment with a crosstalk index greater than a preset crosstalk threshold is regarded as a target audio segment, and the crosstalk audio segment with a crosstalk index less than or equal to the preset crosstalk threshold is regarded as a non-target audio segment.

[0099] It should be noted that the preset crosstalk threshold in this embodiment of the present application is 0.5, and the specific value can be adjusted according to the implementation scene, which is not limited here.

[0100] Since the electronic medical prompt tone usually does not have the natural human voice tone as the medical staff voice output by the talkback system, the IIR comb filter can be used for dynamic noise reduction to reduce the interference on the voice: taking the sum of the crosstalk index of each target audio segment and a preset parameter as an adjustment degree value, taking the product of the adjustment degree value and a preset attenuation depth value as an adjustment attenuation depth, and taking the adjustment attenuation depth of each target audio segment as the input of the IIR comb filter, so as to obtain a filtered audio segment. In the embodiment of the present application, the preset parameter is set to 1, and the preset attenuation depth value is 25.

[0101] In the initial audio data, the mean value of the d-vectors of all normal audio segments is obtained as a comparison vector, and the d-vector of each filtered audio segment and each non-target audio segment is obtained.

[0102] The cosine similarity between the d-vector of each filtered audio segment and the comparison vector is calculated as a first similarity value, and the cosine similarity between the d-vector of each non-target audio segment and the comparison vector is calculated as a second similarity value.

[0103] The filtered audio segment corresponding to the first similarity value greater than a preset similarity threshold and the non-target audio segment corresponding to the second similarity value greater than the preset similarity threshold are taken as a to-be-processed audio segment, and the filtered audio segment corresponding to the first similarity value less than or equal to the preset similarity threshold and the non-target audio segment corresponding to the second similarity value less than or equal to the preset similarity threshold are taken as a non-to-be-processed audio segment; the to-be-processed audio segment is considered to be audio data that greatly differs from the normal audio segment, which may contain more noise interference, and thus needs to be further processed.

[0104] The crosstalk index of each to-be-processed audio segment is taken as a noise reduction weight, and the weighted spectral subtraction is used to weight and subtract the amplitude spectrum of each to-be-analyzed audio segment, so as to obtain a denoised audio segment.

[0105] At this point, the effective information segments extracted from the initial audio data can be processed by different degrees of noise reduction.

[0106] Finally, in the initial audio data, the to-be-processed audio segment is replaced by the denoised audio segment, and the target audio segment is replaced by the non-to-be-processed audio segment, so as to obtain the denoised audio data corresponding to the initial audio data.

[0107] It should be noted that in the embodiment of the present application, the IIR comb filter, the acquisition of the d-vector, the calculation of the cosine similarity, and the weighted spectral subtraction are all known technologies, and the specific process is not described here, the preset similarity threshold is 0.7, and the specific value can be adjusted according to the implementation scene, which is not limited here.

[0108] In summary, first in the audio data acquisition module, in the ward, the initial audio data recorded when each patient uses the intercom device is acquired and frame processing is performed to obtain audio frames. In the process of recording audio by the intercom system of the ward, when there is a sudden change in amplitude energy in a frame and the mutation degree is large, the possibility of change in tone or content at the frame is large, so in the audio segment division module, the segment point detection is performed based on the amplitude energy difference characteristics between the audio frames, so as to obtain the audio information segment. In the process of recording audio by the intercom system of the smart ward, the information segment containing important information needs to be focused on to ensure its intelligibility, and the degree of effective information contained in different audio information segments is different, so in the crosstalk index analysis module, the amplitude change characteristics and frequency domain distribution characteristics in the audio information segment are combined to accurately identify the effective information segment containing patient voice, avoiding the misjudgment of key information caused by global processing. Further, since the phenomenon of superposition of patient voice and other intercom device electronic prompt sound may occur in part of the effective information segment, in this module, the voice and the electronic prompt sound of the rest of the intercom device are further distinguished, mainly by jointly analyzing the frequency domain feature difference, distribution and phase information between the audio frames in the effective information segment, the system can screen out the crosstalk audio segment in the effective information segment and calculate the crosstalk index of each crosstalk audio segment, which is used to represent the signal aliasing problem when multiple devices are used in parallel. Finally, in the audio denoising module, the crosstalk audio segment is processed for targeted denoising based on the crosstalk index, which improves the accuracy of background noise suppression and effectively improves the denoising effect.

[0109] It should be noted that the system provided in the above embodiments is only exemplified by the division of the above functional modules, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above.

[0110] Please refer to Figure 2A system structure schematic diagram of an intercom calling system with background noise suppression function for smart ward provided by one embodiment of the present application is shown, and the system structure schematic diagram comprises a processor 200, a memory 201, a bus 202 and a communication interface 203, the processor 200, the communication interface 203 and the memory 201 are connected through the bus 202; wherein the memory 201 can contain a high-speed random access memory, the bus 202 can be an ISA bus, a PCI bus or an EISA bus, etc., the processor 200 can be an integrated circuit chip with signal processing capability; the memory 201 stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to realize steps in the intercom calling system with background noise suppression function for smart ward.

[0111] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.

[0112] Each embodiment in the specification is described in a progressive manner, and the same and similar parts between each embodiment can be referred to each other, and each embodiment mainly describes the difference from other embodiments.

[0113] The above-mentioned is only the preferred embodiment of the present application, and does not limit the present application, and any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included in the protection scope of the present application.

Claims

1. An intercom call system for a smart ward with background noise suppression function, characterized by, The system comprises: An audio data acquisition module, configured to acquire initial audio data recorded when each intercom device is used in a ward and perform frame processing to obtain audio frames; An audio segment division module, configured to analyze amplitude energy difference characteristics between the audio frames, determine a segment point, divide the initial audio data, and obtain an audio information segment; A crosstalk index analysis module, configured to determine an effective index of the audio information segment according to amplitude variation characteristics and frequency domain distribution characteristics in the audio information segment, use the effective index to screen an effective information segment, compare frequency domain characteristic differences and phase information between the audio frames in the effective information segment, and combine the effective index to determine a crosstalk probability for distinguishing a crosstalk audio segment and a normal audio segment, analyze fundamental frequency difference characteristics and confusion degrees of frequency domain distribution between the audio frames in the crosstalk audio segment, and combine the crosstalk probability to determine a crosstalk index of each crosstalk audio segment; An audio denoising module, configured to perform denoising processing on the crosstalk audio segment based on the crosstalk index, and obtain denoised audio data. The crosstalk index acquisition method comprises: In each crosstalk audio segment, a fundamental frequency of each audio frame is acquired based on a short-time Fourier transform, all the fundamental frequencies of the audio frames are arranged in time sequence to obtain a sorting sequence; A first-order difference sequence of the fundamental frequencies in the sorting sequence is calculated, and a ratio of the crosstalk probability of each crosstalk audio segment to a maximum value in the first-order difference sequence is taken as a first crosstalk parameter; Harmonics in each audio frame are extracted based on the short-time Fourier transform, in all the harmonics, harmonics with frequencies that are not integer multiples of the fundamental frequency are taken as disturbing harmonics, and a value obtained by negatively correlating the number of the disturbing harmonics is taken as a second crosstalk parameter; A value obtained by normalizing a product of the first crosstalk parameter and the second crosstalk parameter is taken as the crosstalk index of each crosstalk audio segment.

2. The intercom call system for a smart ward with background noise suppression function according to claim 1, characterized in that, The audio information segment acquisition method comprises: Root mean square values of all amplitudes in each audio frame are acquired as short-time energy values of each audio frame; In all the audio frames, absolute values of differences between the short-time energy values in every adjacent two audio frames in time sequence are taken as energy mutation values of a latter audio frame in the adjacent two audio frames; A mean value of the energy mutation values of all the audio frames is taken as an energy change reference value, and a starting point of an audio frame with an energy mutation value greater than the energy change reference value is taken as a segment point; In the initial audio data, the initial audio data is segmented based on the segment point, and all the audio information segments are obtained.

3. The intercom call system for a smart ward with background noise suppression function according to claim 1, characterized in that, The effective index acquisition method comprises: A zero-crossing rate of each audio information segment is acquired; In each audio information segment, a spectrogram is obtained based on the short-time Fourier transform, and a fundamental frequency and harmonics in the spectrogram are extracted, in all the harmonics, harmonics with frequencies that are integer multiples of the fundamental frequency are taken as effective harmonics, and a number of the effective harmonics is taken as a quantity factor; A value obtained by negatively correlating and normalizing a product of a centroid of the spectrogram and the zero-crossing rate of each audio information segment is taken as an effective factor of each audio information segment; A value obtained by normalizing a product of the effective factor and the quantity factor of each audio information segment is taken as the effective index of each audio information segment.

4. The intercom call system for a smart ward with background noise suppression function according to claim 1, characterized in that, The effective information segment acquisition method comprises: In all audio information segments, an audio information segment with an effective index greater than a preset effective threshold is regarded as an effective information segment.

5. The intercom call system for a smart ward with background noise suppression function according to claim 1, characterized in that, The method for obtaining the crosstalk probability comprises: When at least two complete audio frames are contained in the effective information segment, in each effective information segment, a phase angle of each audio frame is obtained based on a short-time Fourier transform, an absolute value of a difference between phase angles of each two adjacent audio frames is calculated as an instantaneous phase difference, and a variance of all instantaneous phase differences is taken as a first crosstalk factor of each effective information segment; In each effective information segment, a harmonic of each audio frame is extracted based on the short-time Fourier transform, peak points of all harmonics are connected to form a spectral envelope line, a slope value of the spectral envelope line corresponding to each audio frame is obtained as a change value, in adjacent two audio frames, a normalized value of a difference between the change value of a previous audio frame and the change value of a subsequent audio frame is taken as a change difference value, and a maximum value in all change difference values is taken as a second crosstalk factor of the effective information segment; A product of the first crosstalk factor and the second crosstalk factor of each effective information segment is normalized by a ratio of the effective index of each effective information segment, and a normalized value is taken as a crosstalk probability of each effective information segment. When two or more complete audio frames do not exist in the effective information segment, the crosstalk probability of the effective information segment is a preset value.

6. The intercom call system for a smart patient room with background noise suppression according to claim 1, wherein, The method for distinguishing the crosstalk audio segment and the normal audio segment comprises: In all effective information segments, an effective information segment with a crosstalk probability greater than a preset target threshold is regarded as a crosstalk audio segment, and an effective information segment with a crosstalk probability less than or equal to the target threshold is regarded as a normal audio segment.

7. The intercom call system for a smart room with background noise suppression function according to claim 1, characterized in that, The method for obtaining the denoised audio data comprises: The crosstalk index is used to distinguish a target audio segment and a non-target audio segment in all crosstalk audio segments; A sum of the crosstalk index of each target audio segment and a preset parameter is taken as an adjustment degree value, a product of the adjustment degree value and a preset attenuation depth value is taken as an adjustment attenuation depth, and the adjustment attenuation depth of each target audio segment is taken as an input of an IIR comb filter, so that a filtered audio segment is obtained; In the initial audio data, a mean value of d-vectors of all normal audio segments is taken as a comparison vector, and d-vectors of each filtered audio segment and each non-target audio segment are obtained respectively; A cosine similarity between the d-vector of each filtered audio segment and the comparison vector is calculated as a first similarity value, and a cosine similarity between the d-vector of each non-target audio segment and the comparison vector is calculated as a second similarity value; Filtered audio segments corresponding to the first similarity values greater than a preset similarity threshold and non-target audio segments corresponding to the second similarity values greater than the preset similarity threshold are taken as to-be-processed audio segments, and filtered audio segments corresponding to the first similarity values less than or equal to the preset similarity threshold and non-target audio segments corresponding to the second similarity values less than or equal to the preset similarity threshold are taken as non-to-be-processed audio segments; The crosstalk index of each to-be-processed audio segment is taken as a noise reduction weight, and a weighted spectral subtraction is used to subtract the amplitude spectrum of each to-be-analyzed audio segment, so that a denoised audio segment is obtained. In the initial audio data, the corresponding to-be-processed audio segment is replaced by the de-noised audio segment, and the corresponding target audio segment is replaced by the non-to-be-processed audio segment, so as to obtain de-noised audio data corresponding to the initial audio data.

8. The intercom call system for a smart room with background noise suppression function according to claim 7, characterized in that, The distinguishing of the target audio segment and the non-target audio segment comprises: In all crosstalk audio segments, the crosstalk audio segment with a crosstalk index greater than a preset crosstalk threshold is taken as a target audio segment, and the crosstalk audio segment with a crosstalk index less than or equal to the preset crosstalk threshold is taken as a non-target audio segment.

9. The intercom call system for a smart room with background noise suppression function according to claim 7, characterized in that, The preset attenuation depth value is 25.

Citation Information

Patent Citations

  • Speaker segmentation in noisy conversational speech

    US8543402B1

  • Method for processing multichannel acoustic signal, system thereof, and program

    WO2010092914A1