Acoustic signal analysis system, sound quality improvement system, acoustic signal analysis method, sound quality improvement method and program

The acoustic signal analysis system uses normalized and scaled non-negative matrix factorization to separate overlapping phonemes and enhance high-frequency components of bone-conducted sound, addressing the limitations of existing methods in speech and bone-conducted sound processing.

JP2026043415APending Publication Date: 2026-03-12HIROSHIMA CITY UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing methods like Nonnegative Matrix Factorization (NMF) struggle to accurately separate overlapping phonemes in human speech and enhance high-frequency components of bone-conducted sound, which are easily lost in noisy environments.

Method used

An acoustic signal analysis system that performs non-negative matrix factorization on a spectrogram matrix, normalizes the basis matrix using the Euclidean norm, and scales the activation matrix to separate speech into individual components, while enhancing high-frequency components of bone-conducted sound.

Benefits of technology

Accurately separates speech into individual components even when they overlap in time and enhances high-frequency components of bone-conducted sound, improving sound quality in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026043415000001_ABST
    Figure 2026043415000001_ABST
Patent Text Reader

Abstract

Provided is an acoustic signal analysis system and the like that can accurately separate a sound into its individual components even if the components of the constituent units overlap in time. [Solution] An approximation calculation unit 31 performs nonnegative matrix factorization on a spectrogram matrix obtained by short-time Fourier transform of an acoustic signal of a speech sound to estimate a basis matrix indicating a spectral pattern and an activation matrix indicating time changes in signal strength. A normalization unit 32 normalizes the basis matrix using the Euclidean norm of the basis matrix and scales the activation matrix using the Euclidean norm. A separation unit 33 separates the acoustic signal into basis components based on the normalized basis matrix and the scaled activation matrix.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an acoustic signal analysis system, a sound quality improvement system, an acoustic signal analysis method, a sound quality improvement method, and a program. [Background technology]

[0002] Nonnegative Matrix Factorization (NMF) is a matrix decomposition method for nonnegative matrices (see, for example, Non-Patent Document 1). When NMF is applied to the spectrogram of an acoustic signal, it can be decomposed into two matrices: a basis matrix and an activation matrix. The basis matrix is ​​a matrix that contains frequency information such as pitch and timbre, and the activation matrix is ​​a matrix that contains time information such as the timing and strength of pronunciation.

[0003] NMF has been actively studied for tasks targeting musical instrument signals, such as automatic music transcription and sound source separation (see, for example, Non-Patent Documents 2 and 3). According to NMF, a good approximation of the decomposed signal can be obtained by performing a low-rank matrix decomposition of the musical instrument signal using NMF. However, compared to musical instrument signals, the spectral pattern of speech signals changes continuously depending on the context, making low-rank approximate decomposition difficult. Therefore, NMF is considered unsuitable for modeling speech signals.

[0004] On the other hand, bone-conducted sound is sound that is transmitted through the body, such as the skin, muscles, or bones, during speech (see, for example, Non-Patent Document 4). It is expected that the use of bone-conducted sound will lead to the realization of a useful voice interface. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] DD Lee and HS Seung, “Learning the parts of objects by non-negative matrix factorization,” Nature, vol. 401, no. 6755, pp. 788-791, 1999. [Non-patent document 2] P. Smaragdis and JC Brown, “Non-negative matrix factorization for polyphonic music transcription,” Proc. WASPAA, pp. 177-180, 2003. [Non-patent document 3] A. Ozerov and C. F▲e▼votte, “Multichannel nonnegative matrix factorization in convolutive mixtures for audio source separation,” IEEE Trans. ASLP, vol. 18, no. 3, pp. 550-563, 2010. [Non-patent document 4] Shunsuke Ishimitsu et al., "Study on a bone conduction sound recognition system for marine engine operation support," Journal of the Japan Society of Marine Engineering, Vol. 39, No. 4, 2004. Summary of the Invention [Problem to be solved by the invention]

[0006] The smallest structural unit of human speech is the phoneme, and most Japanese words are expressed with one or two phonemes. However, consecutive phonemes are not clearly separated in time but overlap, and overlapping phonemes are difficult to separate.

[0007] NMF can separate sound into components of each unit of sound, such as phonemes, that overlap in time, and each component has its own frequency and time-varying characteristics. However, NMF does not accurately separate the signal strength ratio of each unit of sound.

[0008] However, compared to human speech, bone-conducted sound propagating through the body is characterized by a tendency for high-frequency components to attenuate, making it difficult to pick up and play back bone-conducted sound with a microphone.

[0009] The present invention has been made in light of the above-mentioned circumstances, and aims to provide an acoustic signal analysis system, an acoustic signal analysis method, and a program that can accurately separate speech into individual components even when the constituent components overlap in time. The present invention also provides a sound quality improvement system, a sound quality improvement method, and a program that can enhance the high-frequency components of bone-conducted sound, which are easily lost in noisy environments, and improve the sound quality of the sound. [Means for solving the problem]

[0010] In order to achieve the above object, an acoustic signal analysis system according to a first aspect of the present invention comprises: an approximation calculation unit that performs non-negative matrix factorization on a spectrogram matrix obtained by short-time Fourier transform of an acoustic signal of a speech sound to estimate a basis matrix indicating a spectral pattern and an activation matrix indicating a time change in signal intensity; a normalization unit that normalizes the basis matrix using a Euclidean norm of the basis matrix and scales the activation matrix with the Euclidean norm; a separation unit that separates the acoustic signal into components for each basis based on a normalized basis matrix and a scaled activation matrix; Equipped with.

[0011] A display unit is provided for displaying the magnitudes of the components for each separated basis so that they can be compared with each other. This may also be the case.

[0012] a feature extraction unit that extracts an acoustic feature of the component in the acoustic signal based on the component for each separated basis; This may also be the case.

[0013] The separation unit is identifying, from among syllables or phonemes included in the utterance, syllables or phonemes corresponding to the basis based on the magnitude of a component of each basis or a position in a time axis direction in the acoustic signal; This may also be the case.

[0014] The approximation calculation unit performing non-negative matrix factorization on the spectrogram matrix to generate a basis matrix and an activation matrix whose basis number is 1; and calculating the maximum basis number based on the number of peaks of the waveform represented by the activation matrix; The number of bases is reduced by deleting one of the two columns in which the similarity of the pattern of time change in signal strength indicated by the activation vector of the activation matrix is ​​equal to or greater than a threshold value; performing non-negative matrix factorization on the spectrogram matrix to estimate a reduced basis matrix and an activation matrix; This may also be the case.

[0015] A sound quality improving system according to a second aspect of the present invention comprises: a differential processing unit that calculates a differential signal that indicates a time change in the signal waveform of bone-conducted sound that is conducted within the body of a speaker; a first estimation unit that performs non-negative matrix factorization on a difference signal of bone-conducted sound that does not include speech sound to estimate a first basis matrix that indicates a spectral pattern of noise; a second estimation unit that performs non-negative matrix factorization on a difference signal of bone-conducted sound including noise and speech sound, and estimates, based on the first basis matrix, a first activation matrix that indicates a time change in signal strength of the noise, a second basis matrix that indicates a spectral pattern of high-frequency components of the speech sound, and a second activation matrix that indicates a time change in signal strength of the high-frequency components of the speech sound; Equipped with.

[0016] The first estimation unit normalizing the first basis matrix using a Euclidean norm of the first basis matrix; The second estimation unit Scale the first activation matrix by multiplying it by the Euclidean norm of the first basis matrix; normalizing the second basis matrix using the Euclidean norm of the second basis matrix, and adjusting the scale of the second activation matrix with the Euclidean norm of the second basis matrix; This may also be the case.

[0017] An acoustic signal analysis method according to a third aspect of the present invention comprises: An acoustic signal analysis method executed by an information processing device, comprising: an approximation step of performing non-negative matrix factorization on a spectrogram matrix obtained by short-time Fourier transform of an acoustic signal of a speech sound to estimate a basis matrix indicating a spectral pattern and an activation matrix indicating a time change in signal intensity; a normalization step of normalizing the basis matrix using the Euclidean norm of the basis matrix and scaling the activation matrix with the Euclidean norm; a separation step of separating the acoustic signal into basis-specific components based on a normalized basis matrix and a scaled activation matrix; Includes.

[0018] A sound quality improving method according to a fourth aspect of the present invention comprises: A sound quality improvement method executed by an information processing device, a differential processing step of calculating a differential signal indicating a time change in the signal waveform of bone-conducted sound transmitted through the body of a speaker; a first estimation step of performing non-negative matrix factorization on a difference signal of bone-conducted sound that does not include speech sound to estimate a first basis matrix that indicates a spectral pattern of noise; a second estimation step of performing non-negative matrix factorization on the difference signal of the bone-conducted sound including noise and speech sound, and estimating, based on the first basis matrix, a first activation matrix indicating a time change in the signal strength of the noise, a second basis matrix indicating a spectral pattern of high-frequency components of the speech sound, and a second activation matrix indicating a time change in the signal strength of the high-frequency components of the speech sound; Includes.

[0019] A program according to a fifth aspect of the present invention comprises: Computer, an approximation calculation unit that performs non-negative matrix factorization on a spectrogram matrix obtained by short-time Fourier transform of an acoustic signal of a speech sound to estimate a basis matrix indicating a spectral pattern and an activation matrix indicating a time change in signal intensity; a normalization unit that normalizes the basis matrix using a Euclidean norm of the basis matrix and scales the activation matrix with the Euclidean norm; a separation unit that separates the acoustic signal into components for each basis based on the normalized basis matrix and the scaled activation matrix; Function as.

[0020] A program according to a sixth aspect of the present invention comprises: Computer, a differential processing unit that calculates a differential signal indicating a time change in the signal waveform of bone-conducted sound that is conducted within the body of the speaker; a first estimation unit that performs non-negative matrix factorization on a difference signal of bone-conducted sound that does not include speech sound to estimate a first basis matrix that indicates a spectral pattern of noise; a second estimation unit that performs non-negative matrix factorization on a difference signal of bone-conducted sound including noise and speech sound, and estimates, based on the first basis matrix, a first activation matrix indicating a time change in signal strength of the noise, a second basis matrix indicating a spectral pattern of high-frequency components of the speech sound, and a second activation matrix indicating a time change in signal strength of the high-frequency components of the speech sound; Function as. [Effects of the Invention]

[0021] According to the present invention, it is possible to accurately separate speech into individual components even when the constituent components overlap in time. Furthermore, according to the present invention, it is possible to enhance the high-frequency components of bone-conducted sound, which are easily lost in noisy environments, thereby improving the sound quality. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a block diagram showing a functional configuration of an acoustic signal analysis system according to a first embodiment of the present invention. [Figure 2] (A) is a schematic diagram of a matrix defined by nonnegative matrix factorization. (B) is a schematic diagram of a matrix defined by nonnegative matrix factorization when the basis number is 2. [Figure 3] (A) is a diagram showing an example of a signal waveform of speech data. (B) is a diagram showing an example of a spectral pattern corresponding to a basis vector for each phoneme of a basis matrix obtained from the speech data of (A). (C) is a diagram showing an example of a change over time in the absolute value of signal intensity corresponding to an activation vector for each phoneme of an activation matrix obtained from the speech data of (A). [Figure 4] 10A and 10B are diagrams showing an example of the change over time in the absolute value of the signal strength indicated by the activation vector of each phoneme in the activation matrix. [Figure 5] 2 is a table showing the ratio of the magnitude of components for each sound source between conventional NMF and NMF in the acoustic signal analysis system of FIG. 1. [Figure 6] (A) is a diagram showing an example of a signal waveform of speech data. (B) is a diagram showing temporal changes in the absolute value of signal strength for each phoneme included in the speech data of (A) using an unnormalized activation matrix. (C) is a diagram showing temporal changes in the absolute value of signal strength for each phoneme included in the speech data of (A) using a normalized activation matrix. [Figure 7] FIG. 2 is a block diagram showing the hardware configuration of the acoustic signal analysis system of FIG. 1. [Figure 8] 2 is a flowchart of an analysis process of the acoustic signal analysis system of FIG. 1. [Figure 9] FIG. 10 is a block diagram showing a functional configuration of an acoustic signal analysis system according to a second embodiment of the present invention. [Figure 10] 10A and 10B are schematic diagrams showing a process for determining a basis number. [Figure 11] FIG. 10 is a block diagram showing the functional configuration of a sound quality improving system according to a third embodiment of the present invention. [Figure 12] FIG. 10 is a schematic diagram illustrating the processing of an estimation unit. [Figure 13] 12 is a flowchart of the sound quality improvement process of the sound quality improvement system of FIG. DETAILED DESCRIPTION OF THE INVENTION

[0023] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each drawing, the same or equivalent parts are denoted by the same reference numerals. In the following embodiments, the terms "have," "include," or "contain" also mean "consist of" or "consist of."

[0024] Embodiment 1 First, a description will be given of a first embodiment of the present invention. An acoustic signal analysis system 1A according to this embodiment shown in Fig. 1 analyzes an acoustic signal of a speech sound of a subject.

[0025] 1, the acoustic signal analysis system 1A includes a terminal device 2 as an information processing device, and a server device 3. The terminal device 2 may be a mobile terminal or a smartphone, or may be a tablet, a wearable device, or a personal computer. The server device 3 is a computer communicatively connected to the terminal device 2 via a wireless or wired communication network (not shown). The server device 3 can be connected to a plurality of terminal devices 2 via the communication network.

[0026] [Terminal device: Audio input section] The terminal device 2 includes a voice input unit 20 and a display unit 21. The voice input unit 20 is, for example, a microphone, and inputs voice data including the subject's vocalizations. The subject utters a monosyllabic sound of a pre-specified type, such as "o," "sa," or "so." A monosyllabic is a syllable that has one vowel as the tonic, and is composed of that vowel alone or with one or more consonants before and after the vowel. The subject inputs voice data including this monosyllabic sound to the voice input unit 20. Note that the input voice data is not limited to monosyllabic data, and may be voice data including multiple phonemes.

[0027] The voice input unit 20 converts the input analog voice data into digital data. The voice input unit 20 transmits the input voice data (digital data) to the server device 3. When transmitting, identification information of the subject who spoke and information indicating the content of the utterance are added to the voice data as header information.

[0028] The display unit 21 displays information relating to acoustic features of the voice data including the voice of the subject transmitted from the server device 3. The operation of the display unit 21 will be described later.

[0029] [Server device] The server device 3 includes a Fourier transform unit 30, an approximation calculation unit 31, a normalization unit 32, a separation unit 33, and a feature extraction unit .

[0030] The Fourier transform unit 30 acquires voice data including monosyllabic utterances of the subject, which is input to the voice input unit 20 of the terminal device 2 and transmitted from the terminal device 2. The Fourier transform unit 30 performs a short-time Fourier transform on the voice data, which is an acoustic signal of the utterance. The short-time Fourier transform results in a spectrogram matrix Z shown in FIG. 2(A). The spectrogram matrix Z is a matrix with I rows and J columns (I and J are natural numbers). I indicates the frame length in the frequency direction, and J indicates the frame length in the time axis direction. The column vectors of the spectrogram matrix Z indicate the spectral components in that time period. That is, in the spectrogram matrix Z, each element represents the magnitude of a component of a specific frequency in a specific time period, and the values ​​of all elements are non-negative.

[0031] As shown in Fig. 2(A), the approximation calculation unit 31 performs non-negative matrix factorization (NMF) on the spectrogram matrix Z obtained by the Fourier transform unit 30 to estimate a basis matrix T and an activation matrix V. The basis matrix T represents a spectral pattern, and the activation matrix V represents a time change in signal strength. The basis matrix T is also called a basis function, and the activation matrix V is also called a coefficient matrix.

[0032] The basis matrix T is a matrix with I rows and K columns (K is a natural number). I indicates the frame length in the frequency direction, and K indicates the basis number. The basis number K is usually set so that min(I,J)≫K. The basis number K can be determined, for example, based on information indicating the content of the utterance sent from the terminal device 2 along with the voice data. For example, if the utterance includes one vowel and one consonant, the basis number K can be set to 2. Figure 2(B) shows the basis matrix T and activation matrix V when the basis number K=2. The column vectors of the basis matrix T are denoted by the basis vector t k (k=1~K). The basis vector t k represents the spectral pattern for each basis. The basis matrix T is the basis vector t kare arranged in the row direction. As shown in Figure 2(B), when the number of bases K is 2, the basis matrix T is composed of spectral patterns of two bases arranged in the row direction. The values ​​of all elements of the basis matrix T are non-negative.

[0033] The activation matrix V is a matrix with K rows and J columns. The column vectors of the activation matrix V are denoted as the activation vector v k (k=1~K). Activation vector v k represents the time variation of the signal intensity of the spectral pattern for each basis. The activation matrix V is the activation vector v k As shown in Figure 2(B), when the number of bases K is 2, the activation matrix V is composed of the spectral patterns of the two bases arranged in the column direction. The values ​​of all elements of the activation matrix V are non-negative.

[0034] The approximation calculation unit 31 estimates the basis matrix T and the activation matrix V from the spectrogram matrix Z. The basis matrix T and the activation matrix V are estimated as solutions to the minimization problem of the following equations.

number

[0035] Furthermore, D(Z||TV) is a similarity function between two matrices. In this embodiment, we will explain NMF based on the generalized Kullback-Leibler pseudodistance (KL-NMF). The minimization problem of KL-NMF is solved by alternately optimizing the basis matrix T and the activation matrix V using the following update formula:

number

[0036] The normalization unit 32 normalizes the basis matrix T using the Euclidean norm of the basis matrix T, and also normalizes the basis matrix T by multiplying the activation matrix V by the corresponding Euclidean norm, thereby adjusting the scale of the activation matrix V. The normalization unit 32 normalizes the basis matrix T using the Euclidean norm to match the scale of the basis matrix T, and applies the calculated scale to the activation matrix. N ] T ∈R N The Euclidean norm of is expressed as follows:

number

number

[0037] Furthermore, the normalization unit 32 calculates the basis vector t k Euclidean norm of ||t k || is the activation vector v k and the scale-unified basis vector t' is given as k The basis matrix T' and activation vector v' are composed of kWhen the number of bases K is 2, the basis vectors of the basis matrix T' are t'1 and t'2, and the activation vectors of the activation matrix V' are v'1 and v'2.

[0038] The separation unit 33 separates the spectrogram matrix Z of the audio signal into components for each basis based on the basis matrix T' and the activation matrix V'. For example, when the number of bases K is 2, the basis matrix T' is separated into basis vectors t'1 and t'2, and the activation matrix V' is separated into activation vectors v'1 and v'2. The separation unit 33 multiplies the basis vector t'1 by the activation vector v'1 to generate an audio signal corresponding to a certain basis component t'1v'1, and multiplies the basis vector t'2 by the activation vector v'2 to generate an audio signal corresponding to a certain basis component t'2v'2.

[0039] Here, it is assumed that the speech data is a monosyllabic speech signal consisting of a consonant z and a vowel a. In this case, speech data for / za / is acquired as shown in FIG. 3(A). A Fourier transform unit 30 generates a spectrogram matrix Z based on this speech data. An approximation calculation unit 31 performs NMF to generate a basis matrix T and an activation matrix V that approximate the spectrogram matrix Z. Furthermore, a normalization unit 32 normalizes the basis matrix T to generate a basis matrix T', and adjusts the scale of the activation matrix V to generate an activation matrix V'.

[0040] In this way, the separation unit 33 can decompose the basis matrix T' into basis vectors t'1 and t'2, and separate the activation matrix V' into activation vectors v'1 and v'2. Furthermore, the separation unit 33 identifies syllables or phonemes included in the utterance that correspond to the bases, based on the magnitude or position of each base component in the acoustic signal along the time axis.

[0041] For example, as shown in Fig. 4(A), when the speech data includes one consonant and one vowel, the separation unit 33 identifies the one with the larger absolute value of the signal intensity as the time waveform of the consonant, and the one with the smaller absolute value as the time waveform of the vowel, in the absolute values ​​of the time deformation of the signal intensity corresponding to the activation vectors v'1 and v'2. Alternatively, as shown in Fig. 4(B), the separation unit 33 identifies the one with the earlier rising edge as the time waveform of the consonant, and the one with the later rising edge as the time waveform of the vowel, in the time deformation of the absolute values ​​of the signal intensity corresponding to the activation vectors v'1 and v'2.

[0042] In this way, the signal intensity of each phoneme can be accurately detected by normalizing the basis vectors t1 and t2 and adjusting the scales of the activation vectors v1 and v2 in the normalization unit 32. Therefore, each phoneme (e.g., consonant, vowel) included in the speech data can be associated with the normalized basis vectors t'1 and t'2 and activation vectors v'1 and v'2.

[0043] For example, as shown in Figure 5, when conventional NMF is performed on a sound composed of three sine waves with different pitches and amplitudes, C4 (261.63 Hz: amplitude 1.00), G4 (392 Hz: amplitude 3.00), and E4 (329.62 Hz: amplitude 2.00), with a basis number K of 3, the amplitudes of the separated C4, G4, and E4 are 1.00, 1.27, and 1.21, respectively, which are different from the original amplitudes. In contrast, when the basis matrix T is normalized, the amplitudes of the separated C4, G4, and E4 are 1.00, 3.02, and 2.02, respectively, and the absolute values ​​of the amplitudes of the respective sine waves are almost the same as the absolute value of the amplitude of the synthesized sound.

[0044] Furthermore, consider the case where speech data of a speech consisting of one consonant and one vowel is obtained, as shown in Figure 6(A). For this speech, if the basis matrix T is not normalized and the activation matrix V is not scaled, the ratio of the absolute values ​​of the signal intensities of the consonants and vowels (amplitude ratio) represented by the activation matrix V will be approximately 1:1, as shown in Figure 6(B), and will not represent the ratio of the absolute values ​​of the signal intensities of the consonants and vowels contained in the original waveform. On the other hand, by normalizing the basis matrix T and scaling the activation matrix V as in the acoustic signal analysis system 1A according to this embodiment, the amplitude ratio of the absolute values ​​of the signal intensities of the consonants and vowels represented by the activation matrix V' will accurately represent the amplitude ratio of the absolute values ​​of the signal intensities of the actual consonants and vowels contained in the original signal waveform.

[0045] The feature extraction unit 34 extracts acoustic features of the corresponding components in the audio signal based on the components for each basis. In this way, acoustic features for each basis, i.e., for each phoneme, are calculated. Examples of such acoustic features include Mel-Frequency Cepstrum Coefficients (MFCCs). In addition, various parameters that are acoustic features of speech data in a wide variety of analytical processes, such as bandpass filter analysis, linear prediction analysis, cepstrum analysis, average power analysis, and formant frequency analysis, can be included as acoustic features that serve as indicators for diagnosing cardiovascular diseases. Such acoustic features provide useful information in various fields, including medicine.

[0046] The separating unit 33 can transmit the separated audio data of the components for each basis to the terminal device 2, and the feature extracting unit 34 can transmit the calculated acoustic feature data to the terminal device 2.

[0047] [Terminal device: display unit] The display unit 21 displays the signal waveform of the acoustic signal in such a way that the magnitude of the components for each base can be distinguished from one another. Specifically, as shown in Fig. 6(C), the components of the signal waveform are displayed so as to accurately reflect the magnitude of the components for each base, i.e., for each phoneme. The display unit 21 can also receive and display acoustic features extracted by the feature extraction unit 34.

[0048] [Hardware configuration] The acoustic signal analysis system 1A shown in Fig. 1 is realized, for example, by a terminal device 2 and a server device 3 having the hardware configuration shown in Fig. 7 executing a software program. Specifically, the server device 3 includes a CPU (Central Processing Unit) 41, which is a processor that controls the entire device, a main memory 42 such as a RAM (Random Access Memory), an external memory 43 configured from a non-volatile memory such as a flash memory or a hard disk, a communication interface 46 that performs data communication with the terminal device 2, and an internal bus 48 that connects these.

[0049] The program 49 is loaded from the external memory 43 into the main memory 42 and executed by the CPU 41. This realizes the functions of the server device 3. When executing the program 49, the CPU 41 performs data communication with an external computer via the communication interface 46 as necessary.

[0050] The functions of the server device 3 can be implemented in a computer system consisting of one or more computers, each including one or more processors and one or more storage devices, including a non-transitory storage medium. The multiple computers realize the functions of the server device 3 while communicating via an interconnected communication network. For example, some of the functions of the server device 3 may be implemented in one computer, and other parts may be implemented in other computers. The functions of the server device 3 may also be realized by a cloud computer.

[0051] Similarly, the terminal device 2 shown in Fig. 1 is realized by a computer having the hardware configuration shown in Fig. 7 executing a software program. Specifically, the terminal device 2 includes a CPU (Central Processing Unit) 51, a main memory 52, an external memory 53 for storing a program 59, an operation unit 54 including devices such as a keyboard and a mouse, a display 55 including a display device such as a CRT (Cathode Ray Tube) or an LCD monitor, a communication interface 56 for data communication with other computers, a microphone 57 for inputting voice, and an internal bus 58 connecting these. The microphone 57 corresponds to the voice input unit 20 described above.

[0052] The program 59 is loaded from the external memory 53 into the main memory 52 and executed by the CPU 51. The execution content of the program 59 is controlled by operation input from the operation unit 54, and data communication with an external computer is performed via the communication interface 56 as necessary, and the signal waveform, acoustic feature quantities, etc. of the acoustic signal are displayed on the display 55. In this way, the functions of the terminal device 2 are realized.

[0053] [Acoustic signal analysis system operation] Next, the operation of the acoustic signal analysis system 1A according to this embodiment shown in FIG. 1, that is, the analysis process (acoustic signal analysis method), will be described.

[0054] (Terminal device processing) 8, in the terminal device 2, the voice input unit 20 inputs voice data (step S1). The voice input unit 20 performs A / D conversion on the input voice data and transmits the digital voice data to the server device 3.

[0055] (Server device processing) In the server device 3, the Fourier transform unit 30 waits until it receives voice data (step S11; No). When it receives the voice data (step S11; Yes), the Fourier transform unit 30 performs a short-time Fourier transform on the received voice data (step S12). This generates a spectrogram matrix Z of the voice data (see FIG. 2(A)).

[0056] Next, the approximation calculation unit 31 performs NMF on the spectrogram matrix Z to perform approximation calculation to estimate the basis matrix T and the activation matrix V (step S13; approximation calculation step). As a result, the non-normalized basis matrix T and activation matrix V are generated.

[0057] Next, the normalization unit 32 normalizes the basis matrix T using the Euclidean norm ∥T∥ of the basis matrix T, and also normalizes the activation matrix V by multiplying the Euclidean norm ∥T∥ by the activation matrix V, thereby adjusting the scale of the activation matrix (step S14; normalization step). As a result, a normalized basis matrix T' and a scale-adjusted activation matrix V' are obtained. The basis vector t' of the basis matrix T' is k represents the spectral pattern of each base phoneme, and the activation vector v' of the activation matrix V' k represents the time change pattern of the signal strength of each base phoneme.

[0058] Next, the separation unit 33 separates the acoustic signal into components for each basis based on the normalized basis matrix T' and the scale-adjusted activation matrix V', and the feature extraction unit 34 extracts acoustic features of the components in the acoustic signal based on the separated components for each basis (step S15; separation step, acoustic feature extraction step). For example, when the number of bases K is 2, the separation unit 33 generates an acoustic signal of the first phoneme from the basis vector t'1 of the basis matrix T' and the activation vector v'1 of the activation matrix V', and generates an acoustic signal of the second phoneme from the basis vector t'2 of the basis matrix T' and the activation vector v'2 of the activation matrix V'.

[0059] The separation unit 33 or the feature extraction unit 34 transmits the normalized basis matrix T' and the scale-adjusted activation matrix V' or the acoustic signals separated for each basis or the acoustic features of the acoustic signals for each basis as processing results to the terminal device 2 (step S16). After that, the server device 3 waits to receive speech data (step S11; No).

[0060] (Terminal device processing part 2) Meanwhile, the terminal device 2 waits for the processing result to be received (step S2; No). When the processing result is received (step S2; Yes), the display unit 21 displays the processing result (step S3). After executing step S3, the terminal device 2 ends the processing.

[0061] As described above in detail, the acoustic signal analysis system 1A according to this embodiment uses NMF to decompose the spectrogram matrix Z of an acoustic signal into a basis matrix T indicating its frequency components and an activation matrix V indicating changes in signal strength over time, and then normalizes the basis matrix T to adjust the scale of the activation matrix V. This makes it possible to accurately separate speech into its individual components even if the constituent components overlap in time.

[0062] Embodiment 2 Next, a second embodiment of the present invention will be described. As shown in Fig. 9, an acoustic signal analysis system 1B according to this embodiment differs from the acoustic signal analysis system 1A in Fig. 1 in that it includes an approximation calculation unit 35 instead of the approximation calculation unit 31.

[0063] Similar to the approximation calculation unit 31, the approximation calculation unit 35 performs NMF on the spectrogram matrix Z obtained by short-time Fourier transform of the acoustic signal of the vocalization to estimate the basis matrix T and the activation matrix V. However, before this estimation, the approximation calculation unit 35 optimizes the number of bases K of the basis matrix T and the activation matrix V based on the obtained speech data, and performs the above estimation using the optimized number of bases K.

[0064] 10(A), the approximation calculation unit 35 first performs NMF on the spectrogram matrix Z generated by the Fourier transform unit 30 to estimate a basis matrix T(I×1) and an activation matrix V(1×J) in which the basis number K is 1. Then, the approximation calculation unit 35 calculates the maximum basis number KMAX based on the number of peaks (the number of peaks whose signal level is equal to or higher than a predetermined level) of the signal waveform represented by the activation matrix V(1×J).

[0065] Furthermore, the approximation calculation unit 35 calculates the activation vector v of the activation matrix V(1×J) k Reduce the number of bases, K, based on the similarity between the columns of the activation vector v k The similarity between the two activation vectors v is determined by comparing the occurrence times and magnitudes of the multiple peaks in the signal waveforms represented by those vectors. For example, the similarity can be determined by the correlation coefficient calculated when the multiple peaks are aligned along the time axis and the two vectors are superimposed. If the similarity is above a predetermined level, the two activation vectors v k For example, as shown in FIG. 10(B), in the activation matrix V(1×J), the activation vectors v1 to v KMAX In the figure, the activation vector v is similar to m , vn (m is between 1 and M, and n is between 1 and N), delete one of them and find the corresponding basis vector t m , t n In this way, the approximation calculation unit 35 performs clustering on the basis matrix T and the activation matrix V.

[0066] As a result of the above clustering, the final reduced basis number becomes K. The approximation calculation unit 35 is the same as the approximation calculation unit 31 according to the first embodiment in that it performs NMF using the reduced basis number K to estimate a normalized basis matrix T' and a scale-adjusted activation matrix V'.

[0067] In this way, in this embodiment, the basis number K contained in an audio signal can be estimated from the signal. The basis number K basically matches the number of phonemes contained in the audio signal.

[0068] In NMF, the basis vector t k Therefore, the approximation calculation unit 35 calculates the basis vector t' of the basis matrix T' by k In this case, the basis vector t' in the basis matrix T' may be k According to the rearrangement of the activation matrix V', the activation vector v' k The basis vector t' where the height of the main frequency is unclear can also be rearranged. k is placed at the bottom layer. If the audio signal represents a piece of music, the rearranged activation matrix V' will be like a musical score that represents the time evolution of the musical scale.

[0069] Embodiment 3 Next, a third embodiment of the present invention will be described. A sound quality improving system 1C according to this embodiment shown in Fig. 11 improves the sound quality of bone-conducted sound that is transmitted within the body of a speaker.

[0070] As shown in FIG. 11, the sound quality improving system 1C includes a bone conduction microphone 60, a differential processing unit 61, an estimation unit 62, and an integration processing unit 63.

[0071] The bone conduction microphone 60 is attached to the head of the speaker. The bone conduction microphone 60 receives bone conduction sound that is transmitted through the body of the speaker, performs A / D conversion, and outputs the converted sound.

[0072] The differential processing unit 61 calculates a differential signal D(t) that indicates a change over time in the signal waveform of the signal S(t) of bone-conducted sound that is transmitted through the speaker's body and that is received by the bone-conducting microphone 60. This differential signal D(t) can be considered as a differential signal of S(t), and is a signal that includes high-frequency components of the bone-conducted sound.

[0073] As shown in FIG. 12, the estimation unit 62 includes a Fourier transform unit 71, a first estimation unit 72, a second estimation unit 73, and an inverse Fourier transform unit 74.

[0074] The Fourier transform unit 71 performs a Fourier transform on the difference signal D(t) output from the difference processing unit 61. This results in a spectrogram matrix S of the difference signal D(t). Note that in the spectrogram matrix S, a power spectrogram may be used instead of an amplitude spectrogram.

[0075] The differential signal D(t) is a signal that contains a mixture of noise and the sound of the speaker's voice transmitted through the body, and is a signal that represents only noise while the speaker is not speaking. Therefore, based on the differential signal D(t) that represents only noise during periods when the speaker is not speaking, a model is created using the noise spectrogram matrix S, a first basis matrix F that represents the noise spectral pattern, and a first activation matrix W that represents the time change in the noise signal strength, as shown in the following equation. S=FW …(5)

[0076] The first estimation unit 72 estimates the first basis matrix F and the first activation matrix W by performing an approximation operation so as to minimize a distance function between the spectrogram matrix S and the product of the first basis matrix F and the first activation matrix W. That is, the first estimation unit 72 performs NMF using the noise differential signal D(t) as a teacher signal to estimate the first basis matrix F of the noise. The first estimation unit 72 normalizes the first basis matrix F by its Euclidean norm. The normalized first basis matrix F is output to the second estimation unit 73.

[0077] The differential signal D(t) is a differential signal of body-conducted sound, which is a mixture of noise and high-frequency components of speech-conducted sound when a speaker is speaking. Therefore, this differential signal D(t) is modeled as follows using a spectrogram matrix S, a first basis matrix F and a first activation matrix W of noise, a second basis matrix T indicating the spectral pattern of the high-frequency components of speech sound, and a second activation matrix V indicating the time change in signal strength of the high-frequency components of speech sound: S=FW+TV …(6)

[0078] The second estimation unit 73 performs NMF on the spectrum matrix of the difference signal D(t) and estimates a first activation matrix W corresponding to noise, and a second basis matrix T and a second activation matrix V corresponding to high-frequency components of body-conducted sound due to speech, based on the first basis matrix F estimated by the first estimation unit 72. The second estimation unit 73 adjusts the scale of the first activation matrix W by multiplying the first activation matrix W by the Euclidean norm of the first basis matrix F. Furthermore, the second estimation unit 73 normalizes the second basis matrix T using the Euclidean norm of the second basis matrix T, and adjusts the scale of the second activation matrix V using the Euclidean norm of the second basis matrix T.

[0079] The second estimation unit 73 generates the product of a second basis matrix T corresponding to the high-frequency components of the body-conducted sound generated by the speaker and a second activation matrix V.

[0080] The inverse Fourier transform unit 74 generates a time series signal by performing an inverse frequency transform on at least one of the products of the second basis matrix T and the second activation matrix V. That is, the inverse Fourier transform unit 74 generates and outputs a difference signal D'(t) based on the product of the estimated second basis matrix T and the second activation matrix V.

[0081] 11, the integration processing unit 63 integrates the difference signal D'(t) output from the estimation unit 62 and outputs the body-conducted sound signal S'(t). The body-conducted sound signal S'(t) is a signal in which noise is suppressed and the high-frequency components of the body-conducted sound caused by speech, which is easily attenuated within the body, are emphasized.

[0082] The hardware of the sound quality improvement system 1C is, for example, almost the same as the configuration of the terminal device 2 in Fig. 7. In the sound quality improvement system 1C, the microphone 57 is a bone conduction microphone 60.

[0083] Next, the operation of the sound quality improvement system 1C, that is, the sound quality improvement process (sound quality improvement method) will be described.

[0084] 13, the bone conduction microphone 60 receives an input of a signal of the speaker's body-conducted sound (step S21). Subsequently, the differential processing unit 61 performs differential processing to calculate a differential signal that indicates a time change in the signal waveform of the bone-conducted sound that is transmitted through the speaker's body (step S22; differential processing step).

[0085] Furthermore, the Fourier transform unit 71 of the estimation unit 62 performs a Fourier transform on the difference signal D(t) to generate a spectrogram matrix S (step S23).

[0086] The first estimation unit 72 of the estimation unit 62 performs NMF on the bone-conducted sound difference signal D(t) that does not include speech sounds, i.e., the difference signal D(t) when speech is not occurring, to estimate a first basis matrix F that represents the noise spectrum pattern (step S24; first estimation step). Whether speech is occurring can be determined by whether the signal strength is equal to or less than a threshold. At this time, the first basis matrix F is also normalized.

[0087] Next, the second estimation unit 73 of the estimation unit 62 performs NMF on the difference signal of the bone-conducted sound including noise and speech sound, and performs speech sound estimation to estimate a first activation matrix W corresponding to the noise and a second basis matrix T and a second activation matrix V corresponding to the speech sound based on the first basis matrix F, and calculates the product of the second basis matrix T and the second activation matrix V (step S25; second estimation step). At this time, the first activation matrix W is scaled, the second basis matrix T is normalized, and the second activation matrix V is scaled.

[0088] Next, the inverse Fourier transform unit 74 of the estimation unit 62 performs an inverse Fourier transform on the product of the basis matrix F and the activation matrix V to generate a difference signal D'(t) (step S26).

[0089] The integration processing unit 63 performs integration processing on the difference signal D'(t) and outputs a body-conducted sound signal S'(t) (step S27). The signal S'(t) is a signal in which noise is suppressed and the high-frequency components of the speech sound are emphasized.

[0090] As explained in detail above, according to the sound quality improving system 1C of this embodiment, NMF can be used to separate bone-conducted sound within the body into noise and body-conducted sound due to speech, thereby emphasizing the high-frequency components of bone-conducted sound that are easily lost in a noisy environment, thereby improving the sound quality.

[0091] In the above embodiment, the case where the audio signal is separated on a phoneme-by-phoneme basis has been described. However, this is not limiting. By setting the basis number K to a number corresponding to the words contained in the audio signal, the audio signal can also be separated on a word-by-word basis.

[0092] The hardware and software configurations of the server device 3 and the terminal device 2 are merely examples and can be changed or modified as desired, as can the hardware of the sound quality improvement system 1C according to the third embodiment.

[0093] The core processing components of the server device 3 and the terminal device 2, which are composed of CPUs 41, 51, main memories 42, 52, external memories 43, 53, operation unit 54, display 55, communication interfaces 46, 56, microphone 57, and internal buses 48, 58, can be realized using an ordinary computer system rather than a dedicated system. For example, a computer program for executing the above operations may be stored and distributed on a computer-readable recording medium (such as a flexible disk, CD-ROM, or DVD-ROM), and the computer program may be installed on a computer to configure the server device 3 and the terminal device 2 that execute the above processing. Alternatively, the computer program may be stored in a storage device of a server device on a communication network such as the Internet, and the server device 3 and the terminal device 2 may be configured by downloading the computer program to an ordinary computer system.

[0094] When the functions of the server device 3 and the terminal device 2 are realized by sharing the functions of an OS (operating system) and an application program, or by cooperation between the OS and the application program, only the application program portion may be stored on a recording medium or storage device.

[0095] The configurations shown in the above-described embodiments may be partially replaced or modified, and configurations of different embodiments may be combined. The dimensions, scales, and aspect ratios in the drawings may also be changed as appropriate.

[0096] This invention allows various embodiments and modifications without departing from the broad spirit and scope of this invention. Furthermore, the above-described embodiments are intended to explain this invention and do not limit the scope of this invention. That is, the scope of this invention is defined by the claims, not the embodiments. Various modifications made within the scope of the claims and the meaning of the invention equivalent thereto are considered to be within the scope of this invention. [Industrial Applicability]

[0097] The present invention can be applied to the audio separation of speech sounds by constituent units, and also to the improvement of the sound quality of bone-conducted sound. [Explanation of symbols]

[0098] 1A, 1B Acoustic signal analysis system, 1C Sound quality improvement system, 2 Terminal device, 3 Server device, 20 Voice input unit, 21 Display unit, 30 Fourier transform unit, 31 Approximation calculation unit, 32 Normalization unit, 33 Separation unit, 34 Feature extraction unit, 35 Approximation calculation unit, 41 CPU, 42 Main memory, 43 External memory, 46 Communication interface (I / F), 48 Internal bus, 49 Program, 51 CPU, 52 Main memory, 53 External memory, 54 Operation unit, 55 Display, 56 Communication interface (I / F), 57 Microphone, 58 Internal bus, 59 Program, 60 Bone conduction microphone, 61 Difference processing unit, 62 Estimation unit, 63 Integration processing unit, 71 Fourier transform unit, 72 First estimation unit, 73 Second estimation unit, 74 Inverse Fourier transform unit

Claims

1. an approximation calculation unit that performs non-negative matrix factorization on a spectrogram matrix obtained by short-time Fourier transform of an acoustic signal of a speech sound to estimate a basis matrix indicating a spectral pattern and an activation matrix indicating a time change in signal intensity; a normalization unit that normalizes the basis matrix using a Euclidean norm of the basis matrix and scales the activation matrix with the Euclidean norm; a separation unit that separates the acoustic signal into components for each basis based on a normalized basis matrix and a scaled activation matrix; An acoustic signal analysis system comprising:

2. A display unit is provided for displaying the magnitudes of the components for each separated basis so that they can be compared with each other. The acoustic signal analysis system according to claim 1 .

3. a feature extraction unit that extracts an acoustic feature of the component in the acoustic signal based on the component for each separated basis; The acoustic signal analysis system according to claim 1 .

4. The separation unit is identifying, from among syllables or phonemes included in the utterance, syllables or phonemes corresponding to the basis based on the magnitude of a component of each basis or a position in a time axis direction in the acoustic signal; The acoustic signal analysis system according to claim 1 .

5. The approximation calculation unit performing non-negative matrix factorization on the spectrogram matrix to generate a basis matrix and an activation matrix whose basis number is 1; and calculating the maximum basis number based on the number of peaks of the waveform represented by the activation matrix; Reducing the number of bases by deleting one of two columns in which the similarity of the pattern of time change in signal strength indicated by the activation vector of the activation matrix is ​​equal to or greater than a threshold; performing non-negative matrix factorization on the spectrogram matrix to estimate a reduced basis matrix and an activation matrix; The acoustic signal analysis system according to claim 1 .

6. a differential processing unit that calculates a differential signal that indicates a time change in the signal waveform of bone-conducted sound that is conducted within the body of a speaker; a first estimation unit that performs non-negative matrix factorization on a difference signal of bone-conducted sound that does not include speech sound to estimate a first basis matrix that indicates a spectral pattern of noise; a second estimation unit that performs non-negative matrix factorization on a difference signal of bone-conducted sound including noise and speech sound, and estimates, based on the first basis matrix, a first activation matrix indicating a time change in signal strength of the noise, a second basis matrix indicating a spectral pattern of high-frequency components of the speech sound, and a second activation matrix indicating a time change in signal strength of the high-frequency components of the speech sound; A sound quality improvement system.

7. The first estimation unit normalizing the first basis matrix using the Euclidean norm of the first basis matrix; The second estimation unit Scale the first activation matrix by multiplying it by the Euclidean norm of the first basis matrix; normalizing the second basis matrix using the Euclidean norm of the second basis matrix, and adjusting the scale of the second activation matrix with the Euclidean norm of the second basis matrix; The sound quality improving system according to claim 6.

8. An acoustic signal analysis method executed by an information processing device, comprising: an approximation step of performing non-negative matrix factorization on a spectrogram matrix obtained by short-time Fourier transform of an acoustic signal of a speech sound to estimate a basis matrix indicating a spectral pattern and an activation matrix indicating a time change in signal intensity; a normalization step of normalizing the basis matrix using the Euclidean norm of the basis matrix and scaling the activation matrix with the Euclidean norm; a separation step of separating the acoustic signal into basis-specific components based on a normalized basis matrix and a scaled activation matrix; An acoustic signal analysis method comprising:

9. A sound quality improvement method executed by an information processing device, a differential processing step of calculating a differential signal indicating a time change in the signal waveform of bone-conducted sound transmitted through the body of a speaker; a first estimation step of estimating a first basis matrix representing a spectral pattern of noise by performing non-negative matrix factorization on a difference signal of bone-conducted sound that does not include speech sound; a second estimation step of performing non-negative matrix factorization on a difference signal of bone-conducted sound including noise and speech sound, and estimating, based on the first basis matrix, a first activation matrix indicating a time change in signal strength of the noise, a second basis matrix indicating a spectral pattern of high-frequency components of the speech sound, and a second activation matrix indicating a time change in signal strength of the high-frequency components of the speech sound; Sound quality improvement methods including.

10. Computer, an approximation calculation unit that performs non-negative matrix factorization on a spectrogram matrix obtained by short-time Fourier transform of an acoustic signal of a speech sound to estimate a basis matrix indicating a spectral pattern and an activation matrix indicating a time change in signal intensity; a normalization unit that normalizes the basis matrix using a Euclidean norm of the basis matrix and scales the activation matrix with the Euclidean norm; a separation unit that separates the acoustic signal into components for each basis based on the normalized basis matrix and the scaled activation matrix; A program that functions as a

11. Computer, a differential processing unit that calculates a differential signal indicating a time change in the signal waveform of bone-conducted sound that is conducted within the body of the speaker; a first estimation unit that performs non-negative matrix factorization on a difference signal of bone-conducted sound that does not include speech sound to estimate a first basis matrix that indicates a spectral pattern of noise; a second estimation unit that performs non-negative matrix factorization on a difference signal of bone-conducted sound including noise and speech sound, and estimates, based on the first basis matrix, a first activation matrix indicating a time change in signal strength of the noise, a second basis matrix indicating a spectral pattern of high-frequency components of the speech sound, and a second activation matrix indicating a time change in signal strength of the high-frequency components of the speech sound; A program that functions as a