Audio processing methods and audio processing systems

By generating observation and output envelopes and using a hybrid matrix to process instrument sound signals, the processing load problem of estimating spill tone transmission characteristics in multi-instrument performances is solved, achieving more efficient estimation of sound source sound levels and accurate control of target sound levels.

CN114402387BActive Publication Date: 2025-11-14YAMAHA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080064954.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-27
Filing Date
2020-09-23
Publication Date
2025-11-14
Estimated Expiration
2040-09-23

AI Technical Summary

Technical Problem

Existing technologies, when processing multiple instrument performances, suffer from an excessive processing load due to the presumed transmission characteristics of spilled sounds, and cannot effectively reduce the processing load on the sound level of each sound source.

Method used

By generating observation and output envelopes, signal processing is performed using a mixing matrix, reducing the processing load for estimating the sound level of each sound source. The mixing matrix and coefficient matrix are generated using a non-negative matrix decomposition method, thus reducing processing complexity.

Benefits of technology

It effectively reduces the estimation processing load of the sound level of each sound source, and improves the accuracy and processing efficiency of the level and time changes of the target sound.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114402387B_ABST
    Figure CN114402387B_ABST
Patent Text Reader

Abstract

The audio processing system acquires multiple observation envelopes, including a first observation envelope and a second observation envelope. The first observation envelope represents the contour of a first sound signal generated by sound pickup near a first sound source, containing a first target sound from the first sound source and a second overflow sound from a second sound source. The second observation envelope represents the contour of a second sound signal generated by sound pickup near a second sound source, containing a second target sound from the second sound source and a first overflow sound from the first sound source. Using a mixing matrix that includes the mixing ratio of the second overflow sound of the first sound signal and the mixing ratio of the first overflow sound of the second sound signal, the system generates multiple output envelopes, including a first output envelope and a second output envelope, based on the multiple observation envelopes. The first output envelope represents the contour of the first target sound of the first observation envelope, and the second output envelope represents the contour of the second target sound of the second observation envelope.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a technique for processing sound signals that have been picked up from sound sources such as musical instruments. Background Technology

[0002] For example, when recording the sounds of multiple instruments, sometimes a separate pickup device is set up for each instrument. The sound picked up by the pickup device mainly includes the sound from the instrument on which the pickup device is set, but it also includes sound from other instruments (so-called spill sounds). Patent Document 1 discloses a structure in which the transmission characteristics of spill sounds occurring between multiple sound sources are estimated, and spill sounds from other sound sources are removed from the sound picked up by the pickup device.

[0003] Patent Document 1: Japanese Patent Application Publication No. 2013-66079 Summary of the Invention

[0004] However, the technology in Patent Document 1 suffers from a large processing load for estimating the transmission characteristics of spilled sounds occurring between each sound source. Furthermore, it is conceivable that it is unnecessary to separate the sound from each individual sound source; obtaining the sound level of each source is sufficient. Considering the above, one objective of the present invention is to reduce the processing load for obtaining the sound level of each sound source.

[0005] To address the above-mentioned issues, one aspect of the present invention relates to an audio processing method that obtains a plurality of observation envelopes including a first observation envelope and a second observation envelope. The first observation envelope represents the contour of a first sound signal generated by sound pickup near a first sound source and including a first target sound from the first sound source and a second overflow sound from a second sound source. The second observation envelope represents the contour of a second sound signal generated by sound pickup near a second sound source and including a second target sound from the second sound source and a first overflow sound from the first sound source. Using a mixing matrix comprising the mixing ratio of the second overflow sound of the first sound signal and the mixing ratio of the first overflow sound of the second sound signal, a plurality of output envelopes including a first output envelope and a second output envelope are generated based on the plurality of observation envelopes. The first output envelope represents the contour of the first target sound of the first observation envelope, and the second output envelope represents the contour of the second target sound of the second observation envelope.

[0006] One aspect of the present invention relates to an audio processing system comprising: an envelope acquisition unit that acquires a plurality of observation envelopes including a first observation envelope and a second observation envelope, wherein the first observation envelope represents the outline of a first sound signal generated as a signal picked up near a first sound source and including a first target tone from the first sound source and a second overflow tone from a second sound source, and the second observation envelope represents the outline of a first sound signal generated as a signal picked up near a second sound source and including a second target tone from the second sound source and a second overflow tone from the first sound source. The contour of a second sound signal of a first overflow tone from a sound source; and a signal processing unit that, using a mixing matrix comprising the mixing ratio of the second overflow tone of the first sound signal and the mixing ratio of the first overflow tone of the second sound signal, generates a plurality of output envelopes comprising a first output envelope and a second output envelope based on the plurality of observation envelopes, wherein the first output envelope represents the contour of the first target tone of the first observation envelope, and the second output envelope represents the contour of the second target tone of the second observation envelope. Attached Figure Description

[0007] Figure 1 This is a block diagram illustrating the structure of an audio system.

[0008] Figure 2 This is a block diagram illustrating the structure of an audio processing system.

[0009] Figure 3 This is a block diagram illustrating the functional structure of the control device.

[0010] Figure 4 This is an explanatory diagram of the observed envelope.

[0011] Figure 5 This is an explanatory diagram of the presumption processing performed by the presumption processing department.

[0012] Figure 6 This is a flowchart illustrating the specific process of the presumption process.

[0013] Figure 7 This is a flowchart illustrating the specific process of learning and processing.

[0014] Figure 8 This is a schematic diagram of image analysis.

[0015] Figure 9 This is a schematic diagram of image analysis.

[0016] Figure 10 This is a schematic diagram of image analysis.

[0017] Figure 11 This is a schematic diagram of image analysis.

[0018] Figure 12 This is an explanatory diagram of the gate processing performed by the audio processing department.

[0019] Figure 13 This is an explanatory diagram of the compression process performed by the audio processing department.

[0020] Figure 14 This is a flowchart illustrating the overall operation of the audio processing system.

[0021] Figure 15 This is an explanatory diagram of the presumed process in the second embodiment.

[0022] Figure 16 This is an explanatory diagram of the presumed processing in the third embodiment.

[0023] Figure 17 This is a schematic diagram of the analytical image of the variant example. Detailed Implementation

[0024] A: Implementation Method 1

[0025] Figure 1 This is a block diagram illustrating the structure of the sound system 100 according to the first embodiment of the present invention. The sound system 100 is a recording system for music production that picks up and processes sounds generated from N sound sources S[1] to S[N] (N is a natural number of 2 or more). Each sound source S[n] (n = 1 to N) is, for example, a musical instrument that produces sound by playing. For example, multiple percussion instruments constituting a drum kit (e.g., cymbals, kick drums, snare drums, hi-hats, and bass drums) are respectively equivalent to sound sources S[n]. The N sound sources S[1] to S[N] are arranged adjacent to each other in a sound space. In addition, a combination of two or more instruments may be used as a sound source S[n].

[0026] The audio system 100 has N pickup devices D[1] to D[N], an audio processing system 10, and a playback device 20. Each pickup device D[n] is connected to the audio processing system 10 via wired or wireless means. Similarly, the playback device 20 is connected to the audio processing system 10 via wired or wireless means. Alternatively, the audio processing system 10 and the playback device 20 can be integrated into one unit.

[0027] N pickup devices D[1] to D[N] correspond to any one of the N sound sources S[1] to S[N]. That is, the N pickup devices D[1] to D[N] and the N sound sources S[1] to S[N] correspond to each other in a one-to-one manner. Each pickup device D[n] is a microphone that picks up ambient sound. For example, the pickup device D[n] is a directional microphone that points to the sound source S[n]. The pickup device D[n] generates a sound signal A[n] that represents the waveform of the ambient sound. The N channels of sound signals A[1] to A[N] are supplied in parallel to the sound processing system 10.

[0028] Each pickup device D[n] is positioned near a sound source S[n] with the purpose of picking up the sound (hereinafter referred to as the "target sound") originating from the sound source S[n]. Therefore, the target sound from the sound source S[n] predominantly reaches the pickup device D[n]. However, since the sound sources S[n] are positioned adjacent to each other, sounds originating from sound sources S[n'] (n' = 1 to N, n' ≠ n) other than the sound source S[n] corresponding to that pickup device D[n] (hereinafter referred to as "spill sound") also reach each pickup device D[n]. That is, the sound signal A[n] generated by the pickup device D[n] contains not only the predominantly target sound component from the sound source S[n], but also spill sound components from other sound sources S[n'] located around the sound source S[n]. Furthermore, for convenience, the illustration of the A / D converter that converts each sound signal A[n] from analog to digital is omitted.

[0029] The audio processing system 10 is a computer system for processing N channels of audio signals A[1] to A[N]. Specifically, the audio processing system 10 generates multiple channels of audio signals B by processing the N channels of audio signals A[1] to A[N]. The playback device 20 plays the sound represented by the audio signal B. Specifically, the playback device 20 has: a D / A converter that converts the audio signal B from digital to analog; an amplifier that amplifies the audio signal B; and a playback device that plays the sound corresponding to the audio signal B.

[0030] Figure 2 This is a block diagram illustrating the structure of the audio processing system 10. The audio processing system 10 is implemented by a computer system having a control device 11, a storage device 12, a display device 13, an operation device 14, and a communication device 15. In addition, the audio processing system 10 can be implemented by multiple devices configured separately from each other, in addition to being implemented by a single device.

[0031] The control device 11 consists of one or more processors that control the various elements of the audio processing system 10. For example, the control device 11 consists of one or more processors such as CPU (Central Processing Unit), SPU (Sound Processing Unit), DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit). The communication device 15 communicates with N pickup devices D[1] to D[N] and the playback device 20. For example, the communication device 15 has an input port connected to each pickup device D[n] and an output port connected to the playback device 20.

[0032] Display device 13 displays an image directed by control device 11. Display device 13 is, for example, a liquid crystal display panel or an organic EL display panel. Operation device 14 receives operations from the user. Operation device 14 is, for example, a touch panel that detects contact with the display surface of display device 13, or an operation device operated by the user.

[0033] Storage device 12 is a single or multiple memory that stores the program executed by control device 11 and the data used by control device 11. Specifically, storage device 12 stores estimation processing program P1, learning processing program P2, display control program P3, and sound processing program P4. Storage device 12 is constructed from known recording media such as magnetic recording media or semiconductor recording media. In addition, storage device 12 can be constructed by a combination of various recording media. Furthermore, a removable recording medium that can be attached to or detached from sound processing system 10, or an external recording medium (e.g., a network hard drive) that can communicate with sound processing system 10 can also be used as storage device 12.

[0034] Figure 3 This is a block diagram illustrating the functional structure of the audio processing system 10. The control device 11 performs multiple functions (estimation processing unit 31, learning processing unit 32, display control unit 33, and audio processing unit 34) by executing programs stored in the storage device 12. The functions performed by the control device 11 will be described in detail below.

[0035] [1] Estimated processing unit 31

[0036] The control device 11 functions as an estimation processing unit 31 by executing the estimation processing program P1. The estimation processing unit 31 analyzes the audio signals A[1] to A[N] of N channels. Specifically, the estimation processing unit 31 includes an envelope acquisition unit 311 and a signal processing unit 312.

[0037] The envelope acquisition unit 311 generates an observation envelope Ex[n] (Ex[1] to Ex[N]) for each of the N channel audio signals A[1] to A[N]. The observation envelope Ex[n] of each audio signal A[n] is a signal representing the time region of the waveform contour of the audio signal A[n] on the time axis.

[0038] Figure 4 This is an explanatory diagram of the observation envelope Ex[n]. For each specified length period (hereinafter referred to as "analysis period") Ta on the time axis, N channels of observation envelopes Ex[1] to Ex[N] are generated. Each analysis period Ta consists of M unit periods Tu[1] to Tu[M] on the time axis (M is a natural number greater than 2). Each unit period Tu[m] (m = 1 to M) is a period of time corresponding to the U signal values ​​(samples) constituting the sound signal A[n]. The envelope acquisition unit 311 calculates the level x[n,m] of the observation envelope Ex[n] based on the sound signal A[n] for each unit period Tu[m]. The observation envelope Ex[n] of the nth channel of one analysis period Ta is represented by the time sequence of the M levels x[n,1] to x[n,M] within that analysis period Ta. Any one level x[n,m] of the observation envelope Ex[n] can be represented by, for example, the following equation (1).

[0039] [Formula (1]

[0040]

[0041] In equation (1), the notation a[n,u] refers to the u-th (u = 1 to U) signal value among the U signal values ​​a[n,1] to a[n,U] of the sound signal A[n] in the n-th channel within a unit period Tu[m]. As understood from equation (1), each level x[n,m] of the observation envelope Ex[n] is a non-negative effective value equivalent to the root mean square (RMS) of the sound signal A[n]. As understood from the above explanation, the envelope acquisition unit 311 generates a level x[n,m] for each of the N channels for each unit period Tu[m], and uses the time sequence (levels x[n,1] to x[n,M]) of M such levels x[n,m] as the observation envelope Ex[n]. That is, the observation envelope Ex[n] of each channel is represented by an M-dimensional vector with M levels x[n,1] to x[n,M] as elements.

[0042] Figure 5 This is an explanatory diagram of the operation of the estimation processing unit 31. The observation envelope Ex[n] described above is generated for each of the N channels of audio signals A[1] to A[N]. Therefore, a non-negative matrix (hereinafter referred to as the "observation matrix") X, in which the N observation envelopes Ex[1] to Ex[N] are arranged vertically in N rows and M columns, is generated for each resolution period Ta. The element in the nth row and mth column of the observation matrix X is the mth level x[n,m] of the observation envelope Ex[n] of the nth channel. Furthermore, in the following figures, the case where the total number of channels N of the audio signal A[n] is 3 is illustrated.

[0043] Figure 3 The signal processing unit 312 generates output envelopes Ey[1] to Ey[N] of N channels based on the observation envelopes Ex[1] to Ex[N] of N channels. Figure 5 As illustrated, the output envelope Ey[n] corresponding to the observation envelope Ex[n] emphasizes (ideally extracts) the time-region signal of the target tone from the sound source S[n] that reflects the observation envelope Ex[n]. That is, in the output envelope Ey[n], the level of the overflow tone from each sound source S[n'] other than the sound source S[n] is reduced (ideally removed). As understood from the above description, the output envelope Ey[n] represents the temporal change in the level of the target tone originating from the sound source S[n]. Therefore, according to the first embodiment, there is an advantage that the user can accurately grasp the temporal change in the level of the target tone from each sound source S[n].

[0044] The signal processing unit 312 generates the output envelopes Ey[1] to Ey[N] of the N channels for each analysis period Ta based on the observation envelopes Ex[1] to Ex[N] of each analysis period Ta. That is, the output envelopes Ey[1] to Ey[N] of the N channels are generated for each analysis period Ta. The output envelope Ey[n] of the nth channel of one analysis period Ta is represented by the time series of M levels y[n,1] to y[n,M] corresponding to different unit periods Tu[m] within that analysis period Ta. That is, each output envelope Ey[n] is represented by an M-dimensional vector with the M levels y[n,1] to y[n,M] as elements. The output envelopes Ey[1] to Ey[N] of the N channels generated by the signal processing unit 312 constitute an N-row M-column non-negative matrix (hereinafter referred to as the "coefficient matrix") Y. The element in the nth row and mth column of the coefficient matrix Y (activation matrix) is the mth level y[n,m] of the output envelope Ey[n].

[0045] During one analysis period Ta, the signal processing unit 312 generates a coefficient matrix Y based on the observation matrix X by using non-negative matrix factorization (NMF) of the existing mixing matrix Q (basis matrix). The mixing matrix Q is an N-row N-column square matrix with multiple mixing ratios q[n1,n2] (n1 = 1 to N, n2 = 1 to N) arranged in sequence. The mixing matrix Q is generated in advance by machine learning and stored in the storage device 12. The diagonal elements of the mixing matrix Q, i.e., each mixing ratio q[n,n] (n1 = n2 = n), are set to a reference value (specifically, 1).

[0046] Each observation envelope Ex[n] is represented by the following formula (2).

[0047] Ex[n]≒q[n,1]Ey[1]+q[n,2]Ey[2]+…+q[n,N]Ey[N](2)

[0048] That is, the N mixing ratios q[n,1] to q[n,N] corresponding to the observation envelope Ex[n] are equivalent to the weighted values ​​of each output envelope Ey[n] when the observation envelope Ex[n] is approximately represented by the weighted sum of the output envelopes Ey[1] to Ey[N] of N channels.

[0049] That is, the mixing ratios q[n1,n2] of the mixing matrix Q are indicators of the degree to which spillover sound from sound source S[n2] is mixed into the sound signal A[n1] (observation envelope Ex[n1]). The mixing ratio q[n1,n2] is also referred to as an indicator related to the arrival rate (or attenuation rate) of spillover sound arriving from sound source S[n2] for the pickup device D[n1]. Specifically, the mixing ratio q[n1,n2], when the volume of the target sound picked up by the pickup device D[n1] from sound source S[n1] is set to 1 (reference value), is the ratio (intensity ratio) of the volume of spillover sound picked up by the pickup device D[n1] from other sound sources S[n2]. Therefore, the product of the mixing ratio q[n1,n2] and the level y[n2,m] of the output envelope Ey[n2], q[n1,n2]y[n2,m], is equivalent to the volume of the overflow sound from the sound source S[n2] to the pickup device D[n1].

[0050] For example, Figure 5 The mixing ratio q[1,2] of the mixing matrix Q is 0.1, which means that the overflow sound from the sound source S[2] is mixed in the sound signal A[1] (observation envelope Ex[1]) at a ratio of 0.1 relative to the target sound from the sound source S[1]. Similarly, the mixing ratio q[1,3] is 0.2, which means that the overflow sound from the sound source S[3] is mixed in the sound signal A[1] (observation envelope Ex[1]) at a ratio of 0.2 relative to the target sound from the sound source S[1]. Likewise, for example, the mixing ratio [3,1] is 0.2, which means that the overflow sound from the sound source S[1] is mixed in the sound signal A[3] (observation envelope Ex[3]) at a ratio of 0.2 relative to the target sound from the sound source S[3]. That is, the larger the mixing ratio q[n1,n2], the larger the overflow sound from the sound source S[n2] to the pickup device D[n1].

[0051] The signal processing unit 312 in the first embodiment repeatedly updates the coefficient matrix Y in a manner that makes the product QY of the mixing matrix Q and the coefficient matrix Y close to the observation matrix X. For example, the signal processing unit 312 calculates the coefficient matrix Y in a manner that minimizes the evaluation function F(X|QY) representing the distance between the observation matrix X and the product QY. The evaluation function F(X|QY) is, for example, any distance norm such as Euclidean distance, KL (Kullback-Leibler) divergence, Itakura Saito distance, or β divergence.

[0052] Focusing on any two sound sources S[k1] and S[k2] from N sound sources S[1] to S[N] (k1 = 1 to N, k2 = 1 to N, k1 ≠ k2). The observation envelopes Ex[1] to Ex[N] of the N channels include the observation envelope Ex[k1] and the observation envelope Ex[k2]. The observation envelope Ex[k1] is the contour of the sound signal A[k1] picked up from the target sound source S[k1]. The observation envelope Ex[k1] is an example of the "first observation envelope", the sound source S[k1] is an example of the "first sound source", and the sound signal A[k1] is an example of the "first sound signal". On the other hand, the observation envelope Ex[k2] is the contour of the sound signal A[k2] picked up from the target sound source S[k2]. The observation envelope Ex[k2] is an example of the "second observation envelope", the sound source S[k2] is an example of the "second sound source", and the sound signal A[k2] is an example of the "second sound signal".

[0053] The mixing matrix Q contains mixing ratios q[k1,k2] and q[k2,k1]. Mixing ratio q[k1,k2] is the mixing ratio of the overflow sound from sound source S[k2] in the sound signal A[k1] (observation envelope Ex[k1]), and mixing ratio q[k2,k1] is the mixing ratio of the overflow sound from sound source S[k1] in the sound signal A[k2] (observation envelope Ex[k2]). The output envelopes Ey[1] to Ey[N] of the N channels contain output envelopes Ey[k1] and Ey[k2]. Output envelope Ey[k1] is an example of the "first output envelope", which refers to the signal representing the contour of the target sound from sound source S[k1] in the observation envelope Ex[k1]. On the other hand, the output envelope Ey[k2] is an example of the "second output envelope", which refers to the signal representing the contour of the target tone from the sound source S[k2] of the observation envelope Ex[k2].

[0054] Figure 6 This is a flowchart illustrating the specific process of the control device 11 generating the coefficient matrix Y (hereinafter referred to as "presumption process") Sa. The prediction process Sa begins with an instruction from the user to the operating device 14 and executes in parallel the sounds based on N sound sources S[1] to S[N]. For example, the user of the audio system 100 plays a musical instrument as a sound source S[n]. The prediction process Sa is executed in parallel with the performance by multiple users. The prediction process Sa is executed for each resolution period Ta.

[0055] If estimation processing Sa begins, the envelope acquisition unit 311 generates observation envelopes Ex[1] to Ex[N] (i.e., observation matrix X) (Sa1) for N channels based on the audio signals A[1] to A[N] of N channels. Specifically, the envelope acquisition unit 311 calculates the level x[n,m] of each observation envelope Ex[n] through the aforementioned formula (1).

[0056] The signal processing unit 312 initializes the coefficient matrix Y (Sa2). For example, the signal processing unit 312 sets the observation matrix X of the previous analysis period Ta as the initial value of the coefficient matrix Y for the current analysis period Ta. However, the method for initializing the coefficient matrix Y is not limited to the above examples. For example, the signal processing unit 312 may also set the observation matrix X generated for the current analysis period Ta as the initial value of the coefficient matrix Y for the current analysis period Ta. Alternatively, the signal processing unit 312 may set the observation matrix X of the previous analysis period Ta, or a matrix obtained by adding random numbers to each element of the coefficient matrix Y, as the initial value of the coefficient matrix Y for the current analysis period Ta.

[0057] The signal processing unit 312 calculates an evaluation function F(X|QY) that represents the distance between the product QY of the existing mixing matrix Q and the current coefficient matrix Y and the observation matrix X of the current analysis period Ta (Sa3). The signal processing unit 312 determines whether a predetermined termination condition is met (Sa4). The termination condition is, for example, that the evaluation function F(X|QY) is less than a predetermined threshold, or that the number of times the coefficient matrix Y has been updated has reached a predetermined threshold.

[0058] If the termination condition is not met (Sa4: NO), the signal processing unit 312 updates the coefficient matrix Y in a way that reduces the evaluation function F(X|QY) (Sa5). The calculation of the evaluation function F(X|QY) (Sa3) and the updating of the coefficient matrix Y (Sa5) are repeated until the termination condition is met (Sa4: YES). The coefficient matrix Y is determined by the values ​​at the stage where the termination condition is met (Sa4: YES).

[0059] The generation of the observation envelopes Ex[1] to Ex[N] of N channels (Sa1) and the generation of multiple output envelopes Ey[1] to Ey[N] (Sa2 to Sa5) are performed in parallel with the pickups from N sound sources S[1] to S[N] for each resolution period Ta.

[0060] As understood from the above description, in the first embodiment, the output envelope Ey[n] is generated by processing the observation envelope Ex[n] representing the contour of each sound signal A[n]. Therefore, compared with the structure that analyzes each sound signal A[n], the load of the estimation processing Sa that estimates the level of the target tone (output envelope Ey[n]) of each sound source S[n] can be reduced.

[0061] [2] Learning Processing Department 32

[0062] like Figure 3 As illustrated, the control device 11 functions as a learning processing unit 32 by executing the learning processing program P2. The learning processing unit 32 generates a mixing matrix Q used in the estimation process Sa. The mixing matrix Q is generated (or trained) at any point in time before the estimation process Sa is executed. Specifically, in addition to generating a new initial mixing matrix Q, the generated mixing matrix Q is also trained (retrained). The learning processing unit 32 includes an envelope acquisition unit 321 and a signal processing unit 322.

[0063] The envelope acquisition unit 321 generates an observation envelope Ex[n] (Ex[1] to Ex[N]) for each of the N channel audio signals A[1] to A[N] prepared for training. The duration of the training audio signal A[n] is equivalent to the duration of M unit periods Tu[1] to Tu[M] (i.e., the duration of the resolution period Ta). That is, an observation matrix X of N rows and M columns containing the observation envelopes Ex[1] to Ex[N] of the N channels is generated. The operation performed by the envelope acquisition unit 321 is the same as the operation performed by the envelope acquisition unit 311.

[0064] The signal processing unit 322 generates a mixing matrix Q and output envelopes Ey[1] to Ey[N] of N channels based on the observation envelopes Ex[1] to Ex[N] of the N channels during the analysis period Ta. That is, the mixing matrix Q and the coefficient matrix Y are generated based on the observation matrix X. The process of updating the mixing matrix Q using the observation envelopes Ex[1] to Ex[N] of the N channels is set as one full sample training epoch, and this full sample training is repeated multiple times until the specified termination condition is met, thereby determining the mixing matrix Q used in the estimation process Sa. The termination condition may also be different from the termination condition of the aforementioned estimation process Sa. The mixing matrix Q generated by the signal processing unit 322 is stored in the storage device 12.

[0065] The signal processing unit 322 generates a mixture matrix Q and a coefficient matrix Y based on the observation matrix X through nonnegative matrix factorization. That is, for each training iteration of all samples, the signal processing unit 322 updates the coefficient matrix Y in a manner that makes the product QY of the mixture matrix Q and the coefficient matrix Y close to the observation matrix X. The signal processing unit 322 repeatedly updates the coefficient matrix Y across multiple training iterations of all samples, calculating the coefficient matrix Y in a manner that gradually decreases the evaluation function F(X|QY) representing the distance between the observation matrix X and the product QY.

[0066] Figure 7 This is a flowchart illustrating the specific process of the control device 11 generating (i.e., training) the mixing matrix Q (hereinafter referred to as "learning process") Sb. The learning process Sb is initiated by the user's instruction to the operating device 14. For example, before the formal performance of the presumed process Sa begins (e.g., during a rehearsal), the performer plays the instrument that is the sound source S[n]. The user of the sound system 100 obtains the sound signals A[1] to A[N] of N channels for training by picking up the played sounds.

[0067] Furthermore, if the pickup conditions, such as the position of the sound source S[n], the position of the pickup device D[n], or the relative positional relationship between the sound source S[n] and the pickup device D[n], change, the degree of overflow sound from other sound sources S[n'] to each pickup device D[n] will also change. Therefore, whenever the pickup conditions change, the learning process Sb is executed according to the user's instructions, thereby updating the mixing matrix Q.

[0068] Furthermore, if a change in pickup conditions or an error in the estimation result is detected during the estimation process Sa performed in parallel with the performance of each instrument, the user instructs the sound system 100 to retrain the mixing matrix Q. Based on the user's instruction, the sound system 100 performs the estimation process Sa using the mixing matrix Q at the current time point while recording the current performance, thereby obtaining a training sound signal A[n]. The learning processing unit 32 retrains the mixing matrix Q using the learning process Sb utilizing the training sound signal A[n]. The estimation processing unit 31 uses the retrained mixing matrix Q in the estimation process Sa for future performances. That is, the mixing matrix Q is updated midway through the performance.

[0069] If learning to process Sb begins, the envelope acquisition unit 321 generates observation envelopes Ex[1] to Ex[N] (Sb1) for N channels based on the audio signals A[1] to A[N] used for training. Specifically, the envelope acquisition unit 321 calculates the level x[n,m] of each observation envelope Ex[n] through the aforementioned formula (1).

[0070] The signal processing unit 322 initializes the mixing matrix Q and the coefficient matrix Y (Sb2). For example, the signal processing unit 322 sets the diagonal element (q[n,n]) to 1 and sets all other elements to random numbers. Furthermore, the method for initializing the mixing matrix Q is not limited to the above examples. For example, the mixing matrix Q generated by the previous learning process Sb can be used as the initial mixing matrix Q for the current learning process Sb during retraining. Additionally, the signal processing unit 322 sets the observation matrix X as the initial value of the coefficient matrix Y, for example. Furthermore, the method for initializing the coefficient matrix Y is not limited to the above examples. For example, if the same sound signal A[n] was used in the previous learning process Sb, the signal processing unit 322 can also use the coefficient matrix Y generated by that learning process Sb as the initial value of the coefficient matrix Y for the current learning process Sb. Additionally, the signal processing unit 322 can also set the matrix obtained by adding random numbers to each element of the observation matrix X or the coefficient matrix Y as illustrated above as the initial value of the coefficient matrix Y for the current analysis period Ta.

[0071] The signal processing unit 322 calculates the evaluation function F(X|QY), which represents the distance between the product QY of the mixing matrix Q and the coefficient matrix Y and the observation matrix X of the current analysis period Ta (Sb3). The signal processing unit 322 determines whether a predetermined termination condition is met (Sb4). The termination condition for the learning process Sb is, for example, that the evaluation function F(X|QY) is less than a predetermined threshold, or that the number of times the coefficient matrix Y has been updated has reached a predetermined threshold.

[0072] If the termination condition is not met (Sb4: NO), the signal processing unit 322 updates the mixing matrix Q and the coefficient matrix Y in a way that reduces the evaluation function F(X|QY) (Sb5). The update of the mixing matrix Q and the coefficient matrix Y (Sb5) and the calculation of the evaluation function F(X|QY) (Sb3) are used as one full-sample training iteration until the termination condition is met (Sb4: YES), and this full-sample training is repeated. The mixing matrix Q is determined by the value at the stage where the termination condition is met (Sb4: YES).

[0073] As understood from the above description, in the first embodiment, the mixing matrix Q, which contains the mixing ratio q[n,n'] of the overflow tones from other sound sources S[n'] of each sound signal A[n] (observation envelope Ex[n]), is generated in advance based on the observation envelopes Ex[1] to Ex[N] of the N channels used for training. The mixing matrix Q represents the degree to which the sound signal A[n] corresponding to each sound source S[n] contains overflow tones from other sound sources S[n'] (the degree of sound overflow). Here, the observation envelope Ex[n], which represents the contour of the sound signal A[n], is processed, thereby reducing the load of the learning processing Sb for generating the mixing matrix Q compared to the structure that processes the sound signal A[n].

[0074] Furthermore, the difference between estimation processing Sa and learning processing Sb is that in estimation processing Sa, the mixing matrix Q is fixed, while in learning processing Sb, the mixing matrix Q and the coefficient matrix Y are updated together. That is, estimation processing Sa and learning processing Sb are common to each other except for the point where the mixing matrix Q is not updated. Therefore, the function of learning processing unit 32 can also be used as estimation processing unit 31. That is, in learning processing Sb performed by learning processing unit 32, the mixing matrix Q is fixed, and the observation envelope Ex[n] spanning M unit periods Tu[m] is processed in a concentrated manner, thereby realizing estimation processing Sa. In the foregoing example, estimation processing unit 31 and learning processing unit 32 were described as separate independent elements, but estimation processing unit 31 and learning processing unit 32 can also be mounted as a single element in audio processing system 10.

[0075] [3] Display control unit 33

[0076] like Figure 3 As illustrated, the control device 11 functions as the display control unit 33 by executing the display control program P3. The display control unit 33 displays an image (hereinafter referred to as "analyzed image") Z representing the processing result obtained through the estimation process Sa or the learning process Sb on the display device 13. Specifically, the display control unit 33 displays any one of the multiple analytical images Z (Za to Zd) on the display device 13 in correspondence with, for example, instructions from the user to the operation device 14. The display of the analytical image Z by the display device 13 is initiated by the user's instructions to the operation device 14 and is performed in parallel with the sound production based on N sound sources S[1] to S[N]. That is, the user of the audio system 100 can visually confirm the analytical image Z in real time in parallel with the sound production based on N sound sources S[1] to S[N] (e.g., the playing of a musical instrument). Furthermore, the values ​​of each analytical image Z are displayed, for example, in decibels.

[0077] [3A] Analyzing image Za

[0078] Figure 8 This is a schematic diagram of the analytical image Za. The analytical image Za contains N unit images Ga[1] to Ga[N] corresponding to different channels (CH). Each unit image Ga[n] is an image representing the volume. Specifically, each unit image Ga[n] is a strip-shaped image covering the entire range representing the lower end of the minimum value Lmin and the upper end of the maximum value Lmax. The minimum value Lmin represents silence (-∞ dB). Furthermore, the analytical image Za is an example of the "fourth image".

[0079] A unit image Ga[n] corresponding to any one sound source S[n] is an image representing the level x[n,m] of the observation envelope Ex[n] and the level y[n,m] of the output envelope Ey[n] at a point in time on the time axis. Specifically, each unit image Ga[n] includes a range Ra and a range Rb. The range Ra and the range Rb are displayed in different forms. Furthermore, in this specification, the “form” of an image refers to the characteristics of an image that can be visually discerned by an observer. For example, in addition to the three attributes of color, namely hue (tone), chroma, and brightness (grayscale), size and image content (e.g., pattern or shape) are also included in the concept of “form”.

[0080] The upper end of the range Ra of the unit image Ga[n] represents the level y[n,m] of the output envelope Ey[n,m]. On the other hand, the upper end of the range Rb represents the level x[n,m] of the observation envelope Ex[n]. Therefore, the range Ra refers to the level of the target tone picked up by the pickup device D[n] from the sound source S[n], and the range Rb refers to the increase in level caused by the spilled tones picked up by the pickup device D[n] from the other (N-1) sound sources S[n']. Since the levels of the target tone and spilled tones for the pickup device D[n] vary over time, each unit image Ga[n] changes constantly with the passage of time (specifically, the progression of the performance).

[0081] As understood from the above explanation, by visually verifying the analytical image Za, the user can visually compare the degree of overflow sound relative to the target sound reaching the pickup device D[n] for each pickup device D[n] (each channel). For example, according to Figure 8 The illustrated analytical image Za can detect overflow sounds of the same level as the target sound reaching the pickup device D[1], and overflow sounds of a level significantly lower than the target sound reaching the pickup device D[2]. Furthermore, if the overflow sound from the pickup device D[n] is significant, the user can adjust the position or orientation of the pickup device D[n]. After adjusting the pickup device D[n], the aforementioned learning process Sb is performed.

[0082] [3B] Analyzing image Zb

[0083] Figure 9 This is a schematic diagram of the analytical image Zb. The analytical image Zb contains N unit images Gb[1] to Gb[N] corresponding to different channels (CH). Each channel corresponds to a sound source S[n], so the N unit images Gb[1] to Gb[N] are also referred to as images corresponding to different sound sources S[n]. Each unit image Gb[n] is, like the unit image Ga[n], a strip-shaped image covering the entire range representing the lower end of the minimum value Lmin and the upper end of the maximum value Lmax. Furthermore, the analytical image Zb is an example of the "first image".

[0084] The user can select any one of the N sound sources S[1] to S[N] by appropriately operating the operating device 14. Hereinafter, the sound source S[n] selected by the user from the N sound sources S[1] to S[N] is labeled as the first sound source S[k1], and the (N-1) sound sources S[n] other than the first sound source S[k1] are labeled as the second sound sources S[k2]. Figure 9 In the example, the following situation is illustrated: sound source S[1] is selected as the first sound source S[k1], and sound sources S[2] and S[3] are the second sound sources S[k2]. The shape of the unit image Gb[k1] corresponding to the first sound source S[k1] in the N unit images Gb[1] to Gb[N] is the same as the unit image Ga[n] of the analytical image Za. That is, the unit image Gb[k1] represents the level x[k1,m] of the observation envelope Ex[k1] and the level y[k1,m] of the output envelope Ey[k1].

[0085] Among the N unit images Gb[1] to Gb[N], the unit image Gb[k2] corresponding to each second sound source S[k2] represents the level (hereinafter referred to as "overflow amount") Lb[k2) of the overflow tone from the second sound source S[k2] in the observation envelope Ex[k1] of the first sound source S[k1]. The overflow amount Lb[k2] represents the level of the overflow tone from the second sound source S[k2] to the pickup device D[k1]. Specifically, it is displayed in the unit image Gb[k2] within the range Rb. The upper end of the range Rb of the unit image Gb[k2] represents the overflow amount Lb[k2]. The display control unit 33 calculates the overflow amount Lb[k2] (Lb[k2] = q[k1,k2]y[k2,m]) by multiplying the mixing ratio q[k1,k2] of the mixing matrix Q and the level y[k2,m] of the output envelope Ey[k2].

[0086] For example, Figure 9The overflow amount Lb[2] refers to the level of the overflow sound from the sound source S[2] for the pickup device D[1]. It is calculated by multiplying the mixing ratio q[1,2] of the mixing matrix Q and the level y[2,m] of the output envelope Ey[2], which is (Lb[2]=q[1,2]y[2,m]). In addition, Figure 9 The overflow amount Lb[3] represents the level of the overflow tone from the sound source S[3] for the pickup device D[1], which is calculated by multiplying the mixing ratio q[1,3] of the mixing matrix Q and the level y[3,m] of the output envelope Ey[3], and is (Lb[3]=q[1,3]y[3,m]).

[0087] As understood from the above explanation, the total overflow Lb[k2] of the (N-1) second sound sources S[k2] is equivalent to the total level of the overflow tones from the (N-1) second sound sources S[k2] to the pickup device D[k1] (i.e., the range Rb of the unit image Gb[k1]). The level of the overflow tones from the pickup device D[k1] varies over time, therefore the unit image Gb[k1] and each unit image Gb[k2] change constantly over time (specifically, as the performance progresses).

[0088] As understood from the above explanation, by visually verifying the analyzed image Zb, the user can visually grasp the extent to which the overflow sounds from each of the second sound sources S[k2] affect the sound signal A[k1] picked up from the target sound source S[k1]. For example, according to Figure 9 The illustrated analytical image Zb shows that the level of the overflow sound arriving from sound source S[2] for the pickup device D[1] is greater than the level of the overflow sound arriving from sound source S[3]. Furthermore, when the level of overflow sound from the second sound source S[k2] is high, the user can adjust the position or direction of each pickup device D[n] to reduce the overflow sound from the second sound source S[k2]. The aforementioned learning process Sb is performed after adjusting the pickup devices D[n].

[0089] [3C] Analyze image Zc

[0090] Figure 10 This is a schematic diagram of the analytical image Zc. The analytical image Zc contains N unit images Gc[1] to Gc[N] corresponding to different channels (CH). The N unit images Gc[1] to Gc[N] are also referred to as images corresponding to different sound sources S[n]. Each unit image Gc[n] is, like the unit image Ga[n], a strip-shaped image covering the entire range representing the lower end of the minimum value Lmin and the upper end of the maximum value Lmax. Furthermore, the analytical image Zc is an example of a "second image".

[0091] The user can select any one of the N sound sources S[1] to S[N] as the first sound source S[k1] by appropriately operating the operating device 14. The (N-1) sound sources S[n] other than the first sound source S[k1] among the N sound sources S[1] to S[N] are the second sound sources S[k2]. Figure 10 In the example, the following situation is illustrated: sound source S[2] is selected as the first sound source S[k1], and sound sources S[1] and S[3] are the second sound sources S[k2]. The shape of the unit image Gc[k1] corresponding to the first sound source S[k1] in the N unit images Gc[1]~Gc[N] is the same as the unit image Ga[n] of the analytical image Za. That is, the unit image Gc[k1] represents the level x[k1,m] of the observation envelope Ex[k1] and the level y[k1,m] of the output envelope Ey[k1].

[0092] Among the N unit images Gc[1] to Gc[N], the unit image Gc[k2] corresponding to each second sound source S[k2] represents the overflow amount Lc[k1] of the observation envelope Ex[k2] of the second sound source S[k2] from the first sound source S[k1]. The overflow amount Lc[k2] represents the level of the overflow tone from the first sound source S[k1] to each pickup device D[k2]. Specifically, the unit image Gc[k2] is displayed in the range Rb. The upper end of the range Rb of the unit image Gc[k2] represents the overflow amount Lc[k2]. The display control unit 33 calculates the overflow amount Lc[k2] (Lc[k2] = q[k2,k1]y[k1,m]) by multiplying the mixing ratio q[k2,k1] of the mixing matrix Q and the level y[k1,m] of the output envelope Ey[k1].

[0093] For example, Figure 10 The overflow amount Lc[1] represents the level of the overflow tone from the sound source S[2] for the pickup device D[1]. It is calculated by multiplying the mixing ratio q[1,2] of the mixing matrix Q and the level y[2,m] of the output envelope Ey[2], and is (Lc[1]=q[1,2]y[2,m]). In addition, Figure 10 The overflow amount Lc[3] represents the level of the overflow sound from the sound source S[2] for the pickup device D[3]. It is calculated by multiplying the mixing ratio q[3,2] of the mixing matrix Q and the level y[2,m] of the output envelope Ey[2], which is (Lc[3]=q[3,2]y[2,m]).

[0094] Since the level of the overflow tone for the pickup device D[k1] varies over time, the unit image Gc[k1] and each unit image Gc[k2] change constantly over time (specifically, as the performance progresses).

[0095] As understood from the above explanation, by visually verifying the analyzed image Zc, the user can visually grasp the extent to which the overflow sound from the first sound source S[k1] affects the sound signal A[k2] picked up from each of the second sound sources S[k2]. For example, according to Figure 10 The illustrated analytical image Zc shows that the level of the overflow sound arriving from the sound source S[2] for the pickup device D[1] is less than the level of the overflow sound arriving from the sound source S[2] for the pickup device D[3].

[0096] [3D] Image Zd

[0097] Figure 11 This is a schematic diagram of the analytical image Zd. The analytical image Zd is an image representing the mixing matrix Q. Specifically, the analytical image Zd, like the mixing matrix Q, contains N rows and N columns of N elements arranged in a matrix. 2 Each unit image is Gd[1,1]~Gd[N,N].

[0098] Any one unit image Gd[n1,n2] of the analytical image Zd represents the mixing ratio q[n1,n2] of the mixing matrix Q located in the n1th row and n2th column. Specifically, the unit image Gd[n1,n2] is displayed in a form (e.g., hue or brightness) corresponding to the mixing ratio q[n1,n2]. For example, imagine a structure in which the unit image Gd[n1,n2] is displayed with a longer wavelength hue as the mixing ratio q[n1,n2] is larger, or a structure in which the unit image Gd[n1,n2] is displayed with a higher brightness (lighter grayscale) as the mixing ratio q[n,n'] is larger. That is, the analytical image Zd is an image in which the mixing ratio q[n,n'] of the target sound from the sound source S[n] and the overflow sound from the other sound sources S[n'] are arranged for each of the N sound sources S[1] to S[N]. The analytical image Zd is an example of a "third image".

[0099] As understood from the above description, the user can visually grasp the degree to which sound source S[n] affects sound source S[n'] in relation to any combination of any two sound sources (S[n], S[n']) among N sound sources S[1] to S[N].

[0100] [4] Audio processing unit 34

[0101] like Figure 3As illustrated, the control device 11 functions as the audio processing unit 34 by executing the audio processing program P4. The audio processing unit 34 generates audio signals B[n] (B[1] to B[N]) by performing audio processing on each of the N channel audio signals A[1] to A[N]. Specifically, the audio processing unit 34 performs audio processing on the audio signal A[n] corresponding to the level y[n,m] of the output envelope Ey[n] generated by the estimation processing unit 31. As described above, the output envelope Ey[n] is the envelope representing the contour of the target sound from the sound source S[n] of the audio signal A[n]. Specifically, the audio processing unit 34 performs audio processing on each of the multiple processing periods H set for the audio signal A[n], corresponding to the level y[n,m] of the output envelope Ey[n].

[0102] For example, focusing on any two sound sources S[k1] and S[k2] from N sound sources S[1] to S[N]. The sound processing unit 34 performs sound processing corresponding to the level y[k1,m] of the output envelope Ey[k1] for the sound signal A[k1], and performs sound processing corresponding to the level y[k2,m] of the output envelope Ey[k2] for the sound signal A[k2].

[0103] The audio processing unit 34 generates an audio signal B based on the audio signals B[1] to B[N] from N channels. Specifically, the audio processing unit 34 generates an audio signal B by mixing the audio signals from the N channels by multiplying each of the audio signals B[1] to B[N] by a coefficient. The coefficients (i.e., weighting values) of each audio signal B[n] are set, for example, according to instructions from the user to the operating device 14.

[0104] The audio processing unit 34 performs audio processing including dynamic control of the volume of the audio signal A[n]. The dynamic control includes effects processing such as gate processing and compression processing. The user can select the type of audio processing by operating the operating device 14 appropriately. The type of audio processing can be selected individually for each of the N channels of audio signals A[1] to A[N], or it can be selected collectively for all N channels of audio signals A[1] to A[N].

[0105] [4A] Gate processing

[0106] Figure 12This is an explanatory diagram of the gating process in audio processing. When the user selects the gating process, the audio processing unit 34 sets the processing period H to the variable length during which the level y[n,m] of the output envelope Ey[n] is less than a predetermined threshold yTH1. The threshold yTH1 is, for example, a variable value corresponding to an instruction from the user to the operation device 14. However, the threshold yTH1 may also be fixed to a predetermined value.

[0107] The sound processing unit 34 reduces the volume of the sound signal A[n] during each processing period H. Specifically, the sound processing unit 34 sets the level of the sound signal A[n] to zero (i.e., mutes it) during the processing period H. According to the gating process illustrated above, the overflow sound of the sound signal A[n] from other sound sources S[n'] can be effectively reduced.

[0108] [4B] Compression Processing

[0109] Figure 13 This is an explanatory diagram of compression processing in audio processing. When the user selects compression processing, the audio processing unit 34 reduces the gain of the audio signal A[n] of the nth channel during the processing period H when the level y[n,m] of the output envelope Ey[n] of the nth channel is greater than a predetermined threshold yTH2. The threshold yTH2 is, for example, a variable value corresponding to the instruction from the user to the operation device 14. However, the threshold yTH2 may also be fixed to a predetermined value.

[0110] The sound processing unit 34 reduces the volume of the sound signal A[n] during each processing period H. Specifically, the sound processing unit 34 reduces the signal value by decreasing the gain during each processing period H of the sound signal A[n]. The degree (ratio) of reducing the gain of the sound signal A[n] is set, for example, according to the user's instruction to the operation device 14. As mentioned above, the output envelope Ey[n] is a signal representing the contour of the target tone from the sound source S[n]. Therefore, by reducing the volume of the sound signal A[n] during the processing period H when the level y[n,m] of the output envelope Ey[n] is greater than the threshold yTH2, the change in the volume of the target tone of the sound signal A[n] can be effectively controlled.

[0111] Figure 14 This is a flowchart illustrating the overall operation performed by the control device 11 of the sound processing system 10. For example, it is performed in parallel with the pronunciation of N sound sources S[1] to S[N] for each resolution period Ta. Figure 14 The processing.

[0112] The control device 11 (estimation processing unit 31) generates the output envelopes Ey[1] to Ey[N] of N channels based on the aforementioned estimation processing Sa, according to the observation envelopes Ex[1] to Ex[N] of N channels and the mixing matrix Q (S1). Specifically, firstly, the control device 11 generates the observation envelopes Ex[1] to Ex[N] based on the audio signals A[1] to A[N] of N channels. Secondly, the control device 11 generates the observation envelopes Ex[1] to Ex[N] based on the audio signals A[1] to A[N] of N channels. Figure 6 The estimated processing Sa generates N channels of output envelopes Ey[1]~Ey[N].

[0113] The control device 11 (display control unit 33) displays the resolved image Z on the display device 13 (S2). For example, the control device 11 displays the resolved image Za corresponding to the observation envelopes Ex[1] to Ex[N] of N channels and the output envelopes Ey[1] to Ey[N] of N channels on the display device 13. In addition, the control device 11 displays the resolved image Zb or the resolved image Zc corresponding to the mixing matrix Q and the output envelopes Ey[1] to Ey[N] of N channels on the display device 13. The control device 11 displays the resolved image Zd corresponding to the mixing matrix Q on the display device 13. The resolved image Z is updated sequentially for each resolution period Ta.

[0114] The control device 11 (sound processing unit 34) performs sound processing (S3) corresponding to the level y[n,m] of the output envelope Ey[n] for each of the N channel sound signals A[1] to A[N]. Specifically, the control device 11 performs sound processing for each processing period H set for the sound signal A[n] based on the level y[n,m] of the output envelope Ey[n].

[0115] As explained above, in the first embodiment, acoustic processing corresponding to the level y[n,m] of the output envelope Ey[n] representing the profile of the target tone from the sound source S[n] and representing the observation envelope Ex[n] is performed on the sound signal A[n]. Therefore, the influence of the overflow tone contained in the sound signal A[n] can be reduced, and appropriate acoustic processing is performed on the sound signal A[n].

[0116] B: Implementation Method 2

[0117] The second embodiment will be described. Furthermore, in the embodiments illustrated below, elements that function the same as those in the first embodiment will retain the reference numerals used in the description of the first embodiment. Detailed descriptions of each will be omitted as appropriate.

[0118] In the first embodiment, estimation process Sa is performed for each parsing period Ta that includes multiple unit periods Tu[m] (Tu[1] to Tu[M]). In the second embodiment, estimation process Sa is performed for each unit period Tu[m]. That is, the second embodiment limits the number M of unit periods Tu[m] included in one parsing period Ta of the first embodiment to one.

[0119] Figure 15 This is an explanatory diagram of the estimation process Sa in the second embodiment. In the second embodiment, for each unit period Tu[i] (i is a natural number) on the time axis, N channels of levels x[1,i] to x[N,i] are generated. The observation matrix X is a non-negative matrix with the levels x[1,i] to x[N,i] of the N channels corresponding to one unit period Tu[i] arranged in N rows and 1 columns vertically. Therefore, the time series of the observation matrix X for multiple unit periods Tu[i] is equivalent to the observation envelopes Ex[1] to Ex[N] of N channels. That is, the observation envelope Ex[n] of the nth channel is represented by the time series of the levels x[n,i] of multiple unit periods Tu[i]. Similarly, the coefficient matrix Y is a non-negative matrix with the levels y[1,i] to y[N,i] of the N channels corresponding to one unit period Tu[i] arranged in N rows and 1 columns vertically. Therefore, the time series of the coefficient matrix Y of multiple unit periods Tu[i] is equivalent to the output envelopes Ey[1] to Ey[N] of N channels. The mixing matrix Q is the same as in the first embodiment, and is an N-row N-column square matrix with multiple mixing ratios q[n1,n2] arranged in order.

[0120] In the first embodiment, the process is performed for each parsing period Ta, which contains M unit periods Tu[1] to Tu[M]. Figure 6 The estimation process Sa is performed for each unit period Tu[i]. That is, the estimation process Sa is performed in real time in parallel with the pronunciation of N sound sources S[1] to S[N]. Furthermore, the content of the estimation process Sa is the same as in the first embodiment. On the other hand, the learning process Sb is performed in the same way as in the first embodiment for one analytical period Ta, which includes M unit periods Tu[1] to Tu[m]. That is, in the second embodiment, the estimation process Sa is a real-time process that calculates the level y[n,i] of each unit period Tu[i], while the learning process Sb is a non-real-time process that calculates the output envelope Ey[n] of the entire multiple unit periods Tu[1] to Tu[M].

[0121] As understood from the above description, according to the second embodiment, the delay of the output envelope Ey[n] relative to the pronunciation of the N sound sources S[1] to S[N] is reduced. That is, each output envelope Ey[n] can be generated in real time in parallel with the pronunciation of the N sound sources S[1] to S[N].

[0122] Figure 14 The illustrated processing (S1 to S3) is performed for each unit period Tu[i]. Therefore, the control device 11 (display control unit 33) updates the analytical image Z (Za, Zb, Zc, Zd) displayed on the display device 13 for each unit period Tu[i] (S2). That is, the analytical image Z is updated in real time in parallel with the pronunciation of the N sound sources S[1] to S[N]. As understood from the above description, according to the second embodiment, the analytical image Z is updated without delay relative to the pronunciation of the N sound sources S[1] to S[N]. Therefore, the user can visually confirm the changes in the overflow tone of each channel in real time. For example, in the analytical image Za, the level x[n,i] of the observation envelope Ex[n] and the level y[n,i] of the output envelope Ey[n] for one unit period Tu[i] are displayed on the display device 13 for each channel, and the analytical image Za is updated sequentially for each unit period Tu[i].

[0123] Furthermore, the control device 11 (sound processing unit 34) performs sound processing (S3) on the sound signal A[n] for each unit period Tu[i]. Therefore, it is possible to process each sound signal A[n] without delay relative to the pronunciation of N sound sources S[1] to S[N].

[0124] C: Third Implementation Method

[0125] Figure 16 This is an explanatory diagram of the estimation process Sa in the third embodiment. The envelope acquisition unit 311 of the estimation processing unit 31 in the first embodiment generates observation envelopes Ex[1] to Ex[N] for N channels corresponding to different sound sources S[n]. The envelope acquisition unit 311 of the third embodiment generates observation envelopes Ex[n] (Ex[n]_L, Ex[n]_M, Ex[n]_H) for each channel, corresponding to different frequency bands. Observation envelope Ex[n]_L corresponds to the low-frequency band, observation envelope Ex[n]_M corresponds to the mid-frequency band, and observation envelope Ex[n]_H corresponds to the high-frequency band. The low-frequency band is located on the low-frequency side of the mid-frequency band, and the high-frequency band is located on the high-frequency side of the mid-frequency band. Specifically, the low-frequency band is the frequency band below the lower end value of the mid-frequency band, and the high-frequency band is the frequency band above the upper end value of the mid-frequency band. Furthermore, the total number of frequency bands for calculating the observation envelope Ex[n] is not limited to 3 and is arbitrary. Furthermore, the low-frequency band, mid-frequency band, and high-frequency band can also locally overlap with each other.

[0126] The envelope acquisition unit 311 divides each audio signal A[n] into three frequency bands: low-frequency band, mid-frequency band, and high-frequency band. Using the same method as in the first embodiment, it generates an observation envelope Ex[n] (Ex[n]_L, Ex[n]_M, Ex[n]_H) for each frequency band. As understood from the above description, the observation matrix X is a 3N x M non-negative matrix arranging the observation envelopes Ex[n] (Ex[n]_L, Ex[n]_M, Ex[n]_H) of the three systems across N channels. Furthermore, the mixing matrix Q is a 3N x 3N square matrix arranging the three elements corresponding to different frequency bands across N channels.

[0127] The signal processing unit 312 generates output envelopes Ey[n] (Ey[n]_L, Ey[n]_M, Ey[n]_H) for each of the N channels, corresponding to different frequency bands. Output envelope Ey[n]_L corresponds to the low-frequency band, output envelope Ey[n]_M corresponds to the mid-frequency band, and output envelope Ey[n]_H corresponds to the high-frequency band. Therefore, the coefficient matrix Y is a 3N x M non-negative matrix arranging the output envelopes Ey[n] (Ey[n]_L, Ey[n]_M, Ey[n]_H) across the N channels. The signal processing unit 312 generates the coefficient matrix Y based on the observation matrix X by utilizing the non-negative matrix decomposition of the known mixing matrix Q.

[0128] In the above explanation, we focus on the estimation process Sa, but the same applies to the learning process Sb. Specifically, the envelope acquisition unit 321 of the learning processing unit 32 generates observation envelopes Ex[n] (Ex[n]_L, Ex[n]_M, Ex[n]_H) of the three systems corresponding to different frequency bands based on the audio signals A[n] of each of the N channels. That is, the envelope acquisition unit 321 generates an observation matrix X of 3N rows and N columns that arranges the observation envelopes Ex[n] (Ex[n]_L, Ex[n]_M, Ex[n]_H) of the three systems across N channels. The mixing matrix Q is a square matrix of 9 rows and 9 columns that arranges the three elements corresponding to different frequency bands across N channels. The coefficient matrix Y is a non-negative matrix of 3N rows and N columns that arranges the output envelopes Ey[n] (Ey[n]_L, Ey[n]_M, Ey[n]_H) of the three systems corresponding to different frequency bands across N channels. The signal processing unit 322 generates a mixture matrix Q and a coefficient matrix Y based on the observation matrix X through non-negative matrix decomposition.

[0129] In the third embodiment, the same effect as in the first embodiment is achieved. Furthermore, the third embodiment has the advantage that the observation envelope Ex[n] and output envelope Ey[n] of each channel are separated into multiple frequency bands, thereby enabling the generation of observation envelope Ex[n] and output envelope Ey[n] that accurately reflect the target tone of the sound source S[n]. In addition, in Figure 16 The example shown is based on the first embodiment, but the structure of the third embodiment is also applicable in the second embodiment, which performs the presumed process Sa for each unit period Tu[i].

[0130] D: Variation Example

[0131] The following examples illustrate specific variations of the methods illustrated above. Within the bounds of non-contradiction, two or more methods arbitrarily selected from the examples may be combined.

[0132] (1) In the aforementioned methods, the observation envelopes Ex[n] of each sound signal A[n] are generated by the calculation of the aforementioned formula (1), but the method by which the envelope acquisition unit 311 or the envelope acquisition unit 321 generates the observation envelopes Ex[n] is not limited to the examples above. For example, the observation envelopes Ex[n] can also be constructed by curves or straight lines of the sound signal A[n] that decay over time from each peak on the positive side. Alternatively, the observation envelopes Ex[N] can be generated by smoothing the positive side components of the sound signal A[n].

[0133] (2) In the aforementioned embodiments, the envelope acquisition unit 311 and envelope acquisition unit 321 of the audio processing system 10 generate an observation envelope Ex[n] based on each audio signal A[n]. However, the envelope acquisition unit 311 or envelope acquisition unit 321 may also receive an observation envelope Ex[n] generated by an external device. That is, the envelope acquisition unit 311 or envelope acquisition unit 321 includes both the element of generating an observation envelope Ex[n] through processing the audio signal A[n] and the element of receiving an observation envelope Ex[n] generated by an external device.

[0134] (3) In the aforementioned methods, nonnegative matrix decomposition is exemplified, but the method for generating N-channel output envelopes Ey[1] to Ey[N] based on the observation envelopes Ex[1] to Ex[N] of N channels is not limited to the above examples. For example, each output envelope Ey[n] can also be generated using the non-negative least squares (NNLS) method. That is, an arbitrary optimization method can also be used to approximate the observation matrix X using the mixing matrix Q and the coefficient matrix Y.

[0135] (4) In the aforementioned methods, an analytical image Za representing the level x[n,m] of the observation envelope Ex[n] and the level y[n,m] of the output envelope Ey[n] at one time point on the time axis is shown, but the content of the analytical image Za is not limited to the above examples. For example, it can also be as follows: Figure 17 As illustrated, the display control unit 33 displays a resolution image Za on the display device 13, which has the observation envelope Ex[n] and the output envelope Ey[n] configured based on a common time axis. The difference between the observation envelope Ex[n] and the output envelope Ey[n] corresponds to the volume of the overflow sound from the sound source S[n'] other than the sound source S[n] to the pickup device D[n]. As understood from the above example, the resolution image Za (the fourth image) is a comprehensive representation of the level x[n,m] of the sound source S[n] and the observation envelope Ex[n], and the level y[n,m] of the output envelope Ey[n] of the sound source S[n].

[0136] (5) In the foregoing embodiments, the structure in which the audio processing unit 34 performs gating or compression processing on the audio signal A[n] is illustrated, but the content of the audio processing performed by the audio processing unit 34 is not limited to the above examples. In addition to gating or compression processing, dynamic control such as limiting processing, expanding processing, or maximizing processing may also be performed by the audio processing unit 34. Limiting processing is, for example, setting the volume of the audio signal A[n] to the predetermined value for each processing period H where the level y[n,m] of the output envelope Ey[n] is greater than a threshold. Expanding processing is reducing the volume of the audio signal A[n] during each processing period H. Maximizing processing is increasing the volume of the audio signal A[n] during each processing period H. In addition, audio processing is not limited to dynamic control that controls the volume of the audio signal A[n]. For example, various sound processing operations, such as distortion processing that generates waveform distortion in H during each processing period of sound signal A[n], or reverberation processing that imparts echo to H during each processing period of sound signal A[n], are performed by the sound processing unit 34.

[0137] (6) The audio processing system 10 can also be implemented by a server device that communicates with terminal devices such as mobile phones or smartphones. For example, the audio processing system 10 generates output envelopes Ey[1] to Ey[N] of N channels by performing estimation processing Sa or learning processing Sb on the audio signals A[1] to A[N] received from the terminal device. In addition, in the structure that transmits the observation envelopes Ex[1] to Ex[N] of N channels from the terminal device, the envelope acquisition unit 311 or the envelope acquisition unit 321 receives the observation envelopes Ex[1] to Ex[N] of N channels from the terminal device.

[0138] The display control unit 33 of the audio processing system 10 generates image data representing the analytical image Z corresponding to the observation envelopes Ex[1] to Ex[N] of N channels, the mixing matrix Q, and the output envelopes Ey[1] to Ey[N] of N channels. This image data is then sent to the terminal device, thereby displaying the analytical image Z on the terminal device. The audio processing unit 34 of the audio processing system 10 sends the audio signal B generated by audio processing of each audio signal A[n] to the terminal device.

[0139] (7) Among the foregoing embodiments, an audio processing system 10 having an estimation processing unit 31, a learning processing unit 32, a display control unit 33, and an audio processing unit 34 is shown as an example. Some elements of the audio processing system 10 may be omitted. For example, in a structure that supplies a mixing matrix Q generated by an external device to the audio processing system 10, the learning processing unit 32 may be omitted. One or both of the display control unit 33 and the audio processing unit 34 may also be omitted. Furthermore, a device having a learning processing unit 32 that generates the mixing matrix Q is also referred to as a machine learning device. A system having a display control unit 33 that displays the analytical image Z is also referred to as a display control system.

[0140] (8) The functions of the audio processing system 10 described above are achieved through the coordinated operation of one or more processors constituting the control device 11 and the programs (P1 to P4) stored in the storage device 12, as described above. The program involved in this invention can be provided and installed in a computer in the form of a computer-readable recording medium. The recording medium is, for example, a non-transitory recording medium, preferably an optical recording medium (optical disc) such as a CD-ROM, and also includes any known form of recording medium such as a semiconductor recording medium or a magnetic recording medium. Furthermore, as a non-transitory recording medium, it includes any recording medium other than a transient propagating signal, and may also include volatile recording media. In addition, in a structure in which a transmission device transmits a program via a communication network, the storage device 12 that stores the program in the transmission device is equivalent to the aforementioned non-transitory recording medium.

[0141] E: Appendix

[0142] Based on the examples above, one can, for instance, grasp the following structure.

[0143] [Method A]

[0144] In the technology of Patent Document 1, there is a problem of a large processing load for estimating the transmission characteristics of spilled sound occurring between each sound source. Furthermore, it is conceivable that it is not necessary to separate the sound of each sound source itself, and that it is sufficient to obtain the sound level of each sound source. Considering the above, one aspect of the present invention (A) aims to reduce the processing load for obtaining the sound level of each sound source.

[0145] One aspect (A1) of the present invention relates to an audio processing method that obtains a plurality of observation envelopes including a first observation envelope and a second observation envelope. The first observation envelope represents the contour of a first sound signal generated as a signal picked up near a first sound source and including a first target sound from the first sound source and a second overflow sound from a second sound source. The second observation envelope represents a signal generated as a signal picked up near the second sound source and including a second target sound from the second sound source and a first overflow sound from the first sound source. The contour of the second sound signal is obtained by using a mixing matrix comprising the mixing ratio of the second overflow tone of the first sound signal (first observation envelope) and the mixing ratio of the first overflow tone of the second sound signal (second observation envelope). Based on the plurality of observation envelopes, a plurality of output envelopes comprising a first output envelope and a second output envelope are generated. The first output envelope represents the contour of the first target tone of the first observation envelope, and the second output envelope represents the contour of the second target tone of the second observation envelope.

[0146] In the above method, multiple output envelopes are generated, including a first output envelope representing the contour of the first target tone of the first observation envelope and a second output envelope representing the contour of the second target tone of the second observation envelope. Therefore, the temporal changes in the sound levels of the first and second sound sources can be accurately determined. Furthermore, processing the observation envelope representing the contour of the sound signal reduces the processing load compared to structures that process the sound signal itself.

[0147] "Acquiring the observation envelope" includes both the action of generating the observation envelope through signal processing of the sound signal and the action of receiving the observation envelope generated by other devices. Furthermore, "the first output envelope representing the contour of the first target tone of the first observation envelope" refers to the envelope that suppresses (ideally removes) the overflow sound from sound sources other than the first sound source. The same applies to the second observation envelope and the second output envelope.

[0148] In a specific example of method A1 (method A2), during the generation of the plurality of output envelopes, a pre-prepared non-negative mixing matrix and a non-negative coefficient matrix representing the plurality of output envelopes are generated by non-negative matrix decomposition of the observation matrices representing the plurality of observation envelopes. This method has the advantage that, by non-negative matrix decomposition of the observation matrices representing the plurality of observation envelopes, a non-negative coefficient matrix representing the plurality of output envelopes can be easily generated.

[0149] In a specific example of method A1 or method A2 (method A3), the acquisition of the plurality of observation envelopes and the generation of the plurality of output envelopes are performed sequentially and in parallel with the pickup of the first sound source and the second sound source for each of the plurality of resolution periods on the time axis. In the above methods, the acquisition of the plurality of observation envelopes and the generation of the plurality of output envelopes are performed sequentially and in parallel with the pickup of the first sound signal and the second sound signal. Therefore, it is possible to monitor the temporal changes in the sound levels of the first sound source and the second sound source in real time.

[0150] In a specific example of method A3 (method A4), each of the plurality of resolution periods is a unit period for calculating one level of each of the plurality of observation envelopes. According to the above method, the delay of the first and second output envelopes relative to the phonation produced by the first and second sound sources can be sufficiently reduced.

[0151] In a specific example of mode A4 (mode A5), for each unit period, the level of the first observation envelope and the level of the first output envelope during that unit period are displayed on a display device. According to the above method, the user can visually confirm the relationship between the level of the first observation envelope and the level of the first output envelope without delay relative to the sounds produced by the first and second sound sources.

[0152] One aspect (aspect A6) of the present invention relates to an audio processing method that obtains a plurality of observation envelopes including a first observation envelope and a second observation envelope. The first observation envelope represents the contour of a first sound signal generated as a signal picked up near a first sound source and including a first target tone from the first sound source and a second spill tone from a second sound source. The second observation envelope represents the contour of a first sound signal generated as a signal picked up near the second sound source and including a second target tone from the second sound source and a second spill tone from the second sound source. The contour of the second sound signal of the first overflow tone of the sound source is generated based on the plurality of observation envelopes, which includes a mixing matrix, a first output envelope, and a second output envelope. The mixing matrix includes the mixing ratio of the second overflow tone of the first sound signal and the mixing ratio of the first overflow tone of the second sound signal. The first output envelope represents the contour of the first target tone of the first observation envelope, and the second output envelope represents the contour of the second target tone of the second observation envelope.

[0153] In the above method, a mixing matrix is ​​generated based on multiple observation envelopes. This mixing matrix contains the mixing ratio of the second spillover tone of the first sound signal and the mixing ratio of the first spillover tone of the second sound signal. Therefore, it is possible to evaluate the degree to which spillover tone from other sound sources is included in the sound signal corresponding to each sound source (the degree of sound spillover). Furthermore, processing is performed on the observation envelope representing the contour of the sound signal, thus reducing the processing load compared to structures that process the sound signal itself.

[0154] One aspect (aspect A7) of the present invention relates to an audio processing system comprising: an envelope acquisition unit that acquires a plurality of observation envelopes including a first observation envelope and a second observation envelope, wherein the first observation envelope represents the outline of a first sound signal generated as a signal picked up near a first sound source and including a first target tone from the first sound source and a second overflow tone from a second sound source, and the second observation envelope represents the outline of a first sound signal generated as a signal picked up near a second sound source and including a second target tone from the second sound source and a second overflow tone from a second sound source. The first sound source has a first overflow tone of a second sound signal; and a signal processing unit uses a mixing matrix that includes the mixing ratio of the first sound signal and the second overflow tone, and the mixing ratio of the second sound signal and the first overflow tone, to generate a plurality of output envelopes including a first output envelope and a second output envelope based on the plurality of observation envelopes, wherein the first output envelope represents the outline of the first target tone of the first observation envelope, and the second output envelope represents the outline of the second target tone of the second observation envelope.

[0155] Furthermore, one aspect of the present invention (aspect A8) involves a program that causes a computer to function as an envelope acquisition unit, which acquires a plurality of observation envelopes including a first observation envelope and a second observation envelope. The first observation envelope represents the outline of a first sound signal generated as a signal picked up near a first sound source and including a first target sound from the first sound source and a second overflow sound from a second sound source. The second observation envelope represents a signal generated as a signal picked up near a second sound source and including a second target sound from the second sound source. The signal processing unit generates a plurality of output envelopes, including a first output envelope and a second output envelope, based on the plurality of observation envelopes, using a mixing matrix comprising the mixing ratio of the first sound signal and the mixing ratio of the second overflow sound from the first sound source; the first output envelope representing the outline of the first target sound of the first observation envelope, and the second output envelope representing the outline of the second target sound of the second observation envelope.

[0156] [Method B]

[0157] In music production involving processes such as mixing, users need to consider the impact of spillover sound on the sound picked up by each pickup device. However, in the technology of Patent Document 1, the user cannot grasp the impact of spillover sound on the sound from each sound source. Considering the above, one aspect of the present invention (Aspect B) aims to enable visual understanding of the impact of spillover sound from other sound sources on the sound from each sound source.

[0158] One aspect (Aspect B1) of the present invention relates to a display control method, which, for each of a plurality of different sound sources, acquires an observation envelope representing the contour of a sound signal picked up from that sound source, a mixing ratio of the overflow sound from other sound sources relative to the sound from that sound source in the observation envelope (sound signal), and an output envelope representing the contour of the sound from that sound source in the observation envelope. For each of one or more second sound sources other than a first sound source among the plurality of sound sources, corresponding to the mixing ratio and the output envelope acquired for each of the plurality of sound sources, displays a first image representing the level of the second overflow sound of the observation envelope of the first sound source on a display device.

[0159] In the above method, for each second sound source, a first image is displayed on the display device, and the first image represents the level of the second spill tone in the observation envelope of the first sound source. Therefore, the user can visually grasp the degree to which each second spill tone affects the sound signal picked up from the first target sound.

[0160] Furthermore, "acquiring the observation envelope" includes both the action of generating the observation envelope through signal processing of the sound signal and the action of receiving the observation envelope generated by other devices. Similarly, "acquiring the mixing ratio" and "acquiring the output envelope" also include both the action of generating the envelope through signal processing and the action of receiving the envelope from other devices. Additionally, "the output envelope representing the contour of the sound from the sound source that constitutes the observation envelope" refers to the envelope that suppresses (ideally removes) spillover sounds from sound sources other than the sound source that constitute the observation envelope.

[0161] One aspect (Aspect B2) of the present invention relates to a display control method, which, for each of a plurality of different sound sources, obtains an observation envelope representing the contour of a sound signal picked up from that sound source, a mixing ratio of the overflow sound from other sound sources relative to the sound from that sound source in the observation envelope (sound signal), and an output envelope representing the contour of the sound from that sound source in the observation envelope. For each of one or more second sound sources other than a first sound source among the plurality of sound sources, a second image representing the level of the first overflow sound of the observation envelope of the second sound source is displayed on a display device, corresponding to the mixing ratio and the output envelope obtained for each of the plurality of sound sources.

[0162] In the above method, for each second sound source, a second image is displayed on the display device. This second image represents the level of the first spill tone in the observation envelope of the second sound source. Therefore, the user can visually grasp the degree to which the first spill tone affects the sound signal picked up from each second target sound.

[0163] In a specific example of method B1 or method B2 (method B3), a third image is displayed on the display device for each of the plurality of sound sources. This third image arranges the mixing ratio of the sound from that sound source and the overflow sound from other sound sources. In the above methods, a third image arranging the mixing ratio of the sound from that sound source and the overflow sound from other sound sources is displayed for each of the plurality of sound sources. Therefore, for any combination of any two sound sources, the user can visually grasp the degree to which one sound source influences the other.

[0164] In a specific example (method B4) of any of methods B1 to B3, a fourth image is displayed on the display device for one of the plurality of sound sources. This fourth image represents the level of the observation envelope and the level of the output envelope of the sound source. In the above methods, a fourth image representing the level of the observation envelope and the level of the output envelope is displayed for one of the plurality of sound sources. Therefore, it is possible to visually compare the level of sound from one sound source with the level of spilled sound from other sound sources.

[0165] In a specific example of mode B4 (mode B5), for each unit period in which one level of the observation envelope is calculated, the level of the observation envelope and the level of the output envelope during that unit period are displayed on a display device. According to the above method, the user can visually confirm the relationship between the level of the first observation envelope and the level of the first output envelope without delay relative to the sound produced by the sound source.

[0166] One aspect (Aspect B6) of the present invention relates to a display control system comprising: an estimation processing unit that, for each of a plurality of different sound sources, acquires an observation envelope representing the contour of a sound signal picked up from that sound source, a mixing ratio of spill sounds from other sound sources in the observation envelope (sound signal) relative to the sound from that sound source, and an output envelope representing the contour of the sound from that sound source in the observation envelope; and a display control unit that, for each of one or more second sound sources other than a first sound source among the plurality of sound sources, displays a first image on a display device corresponding to the mixing ratio and the output envelope acquired for each of the plurality of sound sources, the first image representing the level of the second spill sound in the observation envelope of the first sound source.

[0167] One aspect (Aspect B7) of the present invention relates to a display control system comprising: an estimation processing unit that, for each of a plurality of different sound sources, acquires an observation envelope representing the contour of a sound signal picked up from that sound source, a mixing ratio of spill sounds from other sound sources in the observation envelope (sound signal) relative to the sound from that sound source, and an output envelope representing the contour of the sound from that sound source in the observation envelope; and a display control unit that, for each of one or more second sound sources other than a first sound source among the plurality of sound sources, displays a second image on a display device corresponding to the mixing ratio and the output envelope acquired for each of the plurality of sound sources, the second image representing the level of the first spill sound in the observation envelope of the second sound source.

[0168] One aspect of the present invention (Aspect B8) involves a program that causes a computer to function as a unit comprising: an estimation processing unit that, for each of a plurality of different sound sources, acquires an observation envelope representing the contour of a sound signal picked up from that sound source, a mixing ratio of the spilled sound from other sound sources relative to the sound from that sound source in the observation envelope (sound signal), and an output envelope representing the contour of the sound from that sound source in the observation envelope; and a display control unit that, for each of one or more second sound sources other than a first sound source among the plurality of sound sources, displays a first image on a display device corresponding to the mixing ratio and the output envelope acquired for each of the plurality of sound sources, the first image representing the level of the second spilled sound in the observation envelope of the first sound source.

[0169] One aspect of the present invention (Aspect B9) involves a program that causes a computer to function as the following functional units: an estimation processing unit that, for each of a plurality of different sound sources, acquires an observation envelope representing the contour of a sound signal picked up from that sound source, a mixing ratio of the overflow sound from other sound sources in the observation envelope (sound signal) relative to the sound from that sound source, and an output envelope representing the contour of the sound from that sound source in the observation envelope; and a display control unit that, for each of one or more second sound sources other than a first sound source among the plurality of sound sources, displays a second image on a display device corresponding to the mixing ratio and the output envelope acquired for each of the plurality of sound sources, the second image representing the level of the first overflow sound in the observation envelope of the second sound source.

[0170] [Method C]

[0171] However, sometimes various sound processing techniques, such as applying effects to the sound signal in accordance with its level, are performed. For example, it is conceivable to perform gating processing to cancel out regions where the sound signal level is below a threshold, or to perform compression processing to suppress regions where the sound signal level is above a threshold. When the sound signal contains spill noise, it may be impossible to properly perform sound processing for the sound from a specific sound source. In view of the above, one aspect of the present invention (aspect C) aims to perform appropriate sound processing on the sound signal by reducing the influence of spill noise.

[0172] One aspect (Aspect C1) of the present invention relates to an audio processing method, which obtains an observation envelope representing the contour of an audio signal picked up from a sound source, generates an output envelope representing the contour of the sound from the sound source based on the observation envelope, and performs audio processing corresponding to the level of the output envelope for the audio signal.

[0173] Based on the above method, acoustic processing corresponding to the level of the output envelope representing the contour of the sound from the sound source is performed on the sound signal, thereby reducing the influence of spilled sound contained in the sound signal and performing appropriate acoustic processing on the sound signal.

[0174] Furthermore, "acquiring the observation envelope" includes both the action of generating the observation envelope through signal processing of the sound signal and the action of receiving the observation envelope generated by other devices. In addition, "the output envelope representing the contour of the sound from the sound source of the observation envelope" refers to the envelope that suppresses (ideally removes) the overflow sound from sound sources other than the sound source of the observation envelope.

[0175] In a specific example of method C1 (method C2), the audio processing includes dynamic control of the volume during periods in the audio signal corresponding to the level of the output envelope. In a specific example of method C2 (method C3), the dynamic control includes gating processing to cancel periods in the audio signal where the level of the output envelope is less than a threshold. According to the above methods, the volume of overflow sounds other than the target sound in the audio signal can be effectively reduced. Furthermore, in a specific example of method C2 or method C3 (method C4), the dynamic control includes compression processing to reduce the volume exceeding a predetermined value during periods in the audio signal where the level of the output envelope is greater than a threshold. According to the above methods, the volume of the audio signal can be effectively reduced.

[0176] In a specific example (method C5) of any of methods C1 to C4, the level of the observation envelope is acquired sequentially for each unit period during the acquisition of the observation envelope, and one level of the output envelope is generated for each unit period during the generation of the output envelope. According to the above method, the delay of the output envelope relative to the phonation performed by the sound source can be sufficiently reduced.

[0177] One aspect (aspect C6) of the present invention relates to an audio processing method that obtains a plurality of observation envelopes including a first observation envelope and a second observation envelope. The first observation envelope represents the contour of a first sound signal, which is a signal generated by pickup near a first sound source and includes a first target sound from the first sound source and a second overflow sound from a second sound source. The second observation envelope represents the contour of a second sound signal, which is a signal generated by pickup near a second sound source and includes a second target sound from the second sound source and a first overflow sound from the first sound source. The method utilizes the first sound signal (the first observation envelope)... A mixing matrix is ​​generated based on the mixing ratio of the second overflow tone of the first audio signal (the second observation envelope) and the mixing ratio of the first overflow tone of the second audio signal (the second observation envelope). Multiple output envelopes, including a first output envelope and a second output envelope, are generated according to the multiple observation envelopes. The first output envelope represents the outline of the first target tone of the first observation envelope, and the second output envelope represents the outline of the second target tone of the second observation envelope. Audio processing corresponding to the level of the first output envelope is performed on the first audio signal, and audio processing corresponding to the level of the second output envelope is performed on the second audio signal.

[0178] Based on the above method, acoustic processing corresponding to the level of the first output envelope representing the contour of the first target tone of the first observation envelope is performed on the first audio signal, and acoustic processing corresponding to the level of the second output envelope representing the contour of the second target tone of the second observation envelope is performed on the second audio signal. Therefore, appropriate acoustic processing can be performed to reduce the influence of spilled tones contained in the first and second audio signals respectively.

[0179] One aspect (Aspect C7) of the present invention relates to an audio processing system comprising: an envelope acquisition unit that acquires an observation envelope representing the contour of an audio signal picked up from a sound source; a signal processing unit that generates an output envelope representing the contour of the sound from the sound source based on the observation envelope; and an audio processing unit that performs audio processing on the audio signal corresponding to the level of the output envelope.

[0180] One aspect of the present invention (aspect C8) involves a program that causes a computer to function as the following functional units: an envelope acquisition unit that acquires an observation envelope representing the contour of a sound signal picked up from a sound source; a signal processing unit that generates an output envelope representing the contour of the sound from the sound source based on the observation envelope; and an audio processing unit that performs audio processing on the sound signal corresponding to the level of the output envelope.

[0181] Explanation of the label

[0182] 100…sound system, 10…sound processing system, 20…playback device, D[n](D[1]~D[N])…pickup device, 11…control device, 12…storage device, 13…display device, 14…operation device, 15…communication device, 31…estimation processing unit, 311…envelope acquisition unit, 312…signal processing unit, 32…learning processing unit, 321…envelope acquisition unit, 322…signal processing unit, 33…display control unit, 34…sound processing unit, Z(Za, Zb, Zc, Zd)…image analysis.

Claims

1. A sound processing method, implemented by a computer, wherein, Multiple observation envelopes, including a first observation envelope and a second observation envelope, are obtained. The first observation envelope represents the contour of a first sound signal generated by sound pickup near a first sound source, containing a first target sound from the first sound source and a second overflow sound from a second sound source. The second observation envelope represents the contour of a second sound signal generated by sound pickup near a second sound source, containing a second target sound from the second sound source and a first overflow sound from the first sound source. Using a mixing matrix comprising the mixing ratio of the second overflow tone of the first sound signal and the mixing ratio of the first overflow tone of the second sound signal, a plurality of output envelopes comprising a first output envelope and a second output envelope are generated based on the plurality of observation envelopes. The first output envelope represents the outline of the first target tone of the first observation envelope, and the second output envelope represents the outline of the second target tone of the second observation envelope.

2. The sound processing method according to claim 1, wherein, In the generation of the plurality of output envelopes, a nonnegative coefficient matrix representing the plurality of output envelopes is generated by applying a nonnegative matrix decomposition of the mixture matrix to the nonnegative observation matrix representing the plurality of observation envelopes, wherein the mixture matrix is ​​generated through a learning process.

3. The sound processing method according to claim 1 or 2, wherein, The acquisition of the plurality of observation envelopes and the generation of the plurality of output envelopes are performed sequentially in parallel with the pickups from the first sound source and the second sound source for each of the plurality of resolution periods on the time axis.

4. The sound processing method according to claim 3, wherein, Each of the plurality of resolution periods is a unit period for calculating one level of each of the plurality of observation envelopes.

5. The sound processing method according to claim 1, wherein, Corresponding to the mixing matrix and the plurality of output envelopes, an image representing the level of the second overflow tone of the first observation envelope is displayed on the display device.

6. The sound processing method according to claim 1, wherein, The plurality of observation envelopes includes a third observation envelope, which represents the contour of the third sound signal generated by pickup near the third sound source. The first sound signal includes a third overflow tone from the third sound source. Corresponding to the mixing matrix and the plurality of output envelopes, a first image is displayed on the display device, the first image representing the level of the second overflow tone of the first observation envelope and the level of the third overflow tone of the first observation envelope from the third sound source.

7. The sound processing method according to claim 1, wherein, The plurality of observation envelopes includes a third observation envelope, which represents the contour of the third sound signal generated by the pickup of sound from the third sound source. Corresponding to the mixing matrix and the plurality of output envelopes, a second image is displayed on the display device, the second image representing the level of the first overflow tone of the second observation envelope and the level of the overflow tone from the first sound source of the third observation envelope.

8. The sound processing method according to claim 1, wherein, A third image is displayed on the display device, the third image arranging the mixing ratio of the first target sound and the second overflow sound, and the mixing ratio of the second target sound and the first overflow sound.

9. The sound processing method according to claim 1, wherein, A fourth image is displayed on the display device, the fourth image representing the level of the first observation envelope and the level of the first output envelope.

10. The sound processing method according to claim 9, wherein, For each unit period in which a level of the first observation envelope is calculated, the level of the first observation envelope and the level of the first output envelope during that unit period are displayed on the display device.

11. The sound processing method according to claim 1, wherein, For the first audio signal, perform audio processing corresponding to the level of the first output envelope.

12. The sound processing method according to claim 11, wherein, The audio processing includes dynamic control of the volume of the first audio signal during a period that corresponds to the level of the first output envelope.

13. The sound processing method according to claim 12, wherein, The dynamic control includes gating a noise cancellation process during the period when the level of the first output envelope in the first audio signal is less than a threshold.

14. The sound processing method according to claim 12 or 13, wherein, The dynamic control includes compression processing that reduces the volume above a predetermined value during periods when the level of the first output envelope in the first audio signal is greater than a threshold.

15. The sound processing method according to any one of claims 11 to 13, wherein, In acquiring the plurality of observation envelopes, the levels of each observation envelope are acquired sequentially for each unit period. In the generation of the plurality of output envelopes, one level of each output envelope is generated for each unit period.

16. The sound processing method according to claim 14, wherein, In acquiring the plurality of observation envelopes, the levels of each observation envelope are acquired sequentially for each unit period. In the generation of the plurality of output envelopes, one level of each output envelope is generated for each unit period.

17. An audio processing system, comprising: An envelope acquisition unit acquires multiple observation envelopes including a first observation envelope and a second observation envelope. The first observation envelope represents the outline of a first sound signal generated by sound pickup near a first sound source and including a first target sound from the first sound source and a second overflow sound from a second sound source. The second observation envelope represents the outline of a second sound signal generated by sound pickup near a second sound source and including a second target sound from the second sound source and a first overflow sound from the first sound source. The signal processing unit uses a mixing matrix that includes the mixing ratio of the second overflow tone of the first sound signal and the mixing ratio of the first overflow tone of the second sound signal to generate multiple output envelopes that include a first output envelope and a second output envelope based on the multiple observation envelopes. The first output envelope represents the outline of the first target tone of the first observation envelope, and the second output envelope represents the outline of the second target tone of the second observation envelope.

Citation Information

Patent Citations

  • Covering sound elimination device

    JP2013066079A

  • Signal processing apparatus for dereverberating a number of input audio signals

    CN106233382A

  • Acoustic analyzer

    JP2014134688A