Sound signal processing method and sound signal processing device
The sound signal processing method automatically determines parameters with low computational load, replicating manual adjustments by skilled mixer engineers, addressing the high computational demands of existing methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- YAMAHA CORP
- Filing Date
- 2022-07-01
- Publication Date
- 2026-07-29
AI Technical Summary
Existing audio signal processing methods require high computational load to automatically determine parameters similar to those manually adjusted by mixing device operators or mixer engineers.
A sound signal processing method that selects a channel, identifies setting data based on time-series volume data, and outputs it for automatic parameter adjustment with low computational load, using artificial intelligence to replicate manual adjustments made by skilled engineers.
Achieves automatic parameter determination with low computational effort, replicating manual adjustments of mixing device operators, such as gain and volume balancing, similar to skilled mixer engineers.
Smart Images

Figure 0007896387000001 
Figure 0007896387000002 
Figure 0007896387000003
Abstract
Description
Technical Field
[0001] One embodiment of the present invention relates to a sound signal processing method and a sound signal processing apparatus.
Background Art
[0002] Patent Document 1 describes inferring the type of a musical instrument by analyzing the sound emitted from the musical instrument and displaying an icon indicating the inferred musical instrument on a display of a tablet. [[ID=*]] [[ID=*]]
[0003] [[ID=*]] Patent Document 2 describes a tablet and a microphone. The tablet identifies the type of a musical instrument by analyzing an audio signal input by the microphone. [[ID=*]] [[ID=*]]
[0004] [[ID=*]] Patent Document 3 describes an equalizer setting device. The equalizer setting device graphically displays the setting state of the frequency characteristics in the equalizer. The equalizer setting device displays an element indicating a sound range corresponding to the category set in the signal processing channel. [[ID=*]]
Prior Art Documents
Patent Documents
[0005] [[ID=*]] [[ID=*]]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0006] [[ID=*]] There is a need for an audio signal processing method that can automatically determine parameters with low computational load, similar to those that a mixing device operator would manually adjust, or an audio signal processing method that can automatically determine parameters similar to those that a mixer engineer would manually balance between channels.
[0007] One embodiment of the present invention aims to provide an audio signal processing method that can automatically determine parameters similar to those that an operator of a mixing device would manually adjust, with a low computational load, or an audio signal processing method that can automatically determine parameters similar to those that an operator of a mixing device would manually balance the volume between channels. [Means for solving the problem]
[0008] The sound signal processing method relating to the present invention is The mixing device accepts an operation to select at least one first channel from among multiple channels provided in the device. The audio signal from the selected channel 1 is input, Based on the time-series volume data derived from the input sound signal, or data relating to a second channel different from the first channel among the multiple channels, setting data for setting the mixing device is identified. The identified setting data is output. [Effects of the Invention]
[0009] According to one embodiment of the sound signal processing method of this invention, the same parameters as when an operator of a mixing device manually adjusts the parameters can be automatically determined with a low computational load, or the same parameters as when an operator of a mixing device manually balances the volume between channels can be automatically determined. [Brief explanation of the drawing]
[0010] [Figure 1]Figure 1 is a block diagram showing the configuration of the mixing device 1a. [Figure 2] Figure 2 shows the external appearance of the mixing device 1a. [Figure 3] Figure 3 is a block diagram of the signal processing performed in the mixing device 1a. [Figure 4] Figure 4 is a block diagram showing the processing configuration of input patch 21a, input channel 22a, mixing bus 23a, output channel 24a, and output patch 25a. [Figure 5] Figure 5 is a flowchart showing an example of the processing performed by the mixing device 1a. [Figure 6] Figure 6 shows the external appearance of the mixing device 1b according to Modification 1 of the First Embodiment. [Figure 7] Figure 7 shows the external appearance of the mixing device 1c according to a modified example 2 of the first embodiment. [Figure 8] Figure 8 is a flowchart showing an example of the processing of the mixing apparatus 1d according to Modification 3 of the First Embodiment. [Figure 9] Figure 9 shows an image displayed on the screen 16a of the mixing device 1f according to Modification 1 of the second embodiment. [Modes for carrying out the invention]
[0011] [First Embodiment] The mixing apparatus 1a that performs the sound signal processing method according to the first embodiment will be described below with reference to the figures. Figure 1 is a block diagram showing the configuration of the mixing apparatus 1a. Figure 2 is a diagram showing the external appearance of the mixing apparatus 1a.
[0012] The mixing device 1a is an example of an audio signal processing device. The mixing device 1a executes signal processing such as level adjustment of an audio signal or mixing of audio signals. As shown in FIG. 1, the mixing device 1a includes an audio interface 11, a network interface 12, a flash memory 13, a RAM (Random Access Memory) 14, a CPU (Central Processing Unit) 15, a display 16, a user interface 17, a DSP (Digital Signal Processor) 18, and a bus 19. The audio interface 11, the network interface 12, the flash memory 13, the RAM 14, the CPU 15, the display 16, the user interface 17, and the DSP 18 are connected to each other via the bus 19.
[0013] The audio interface 11 receives an audio signal from an audio device such as a microphone or an electronic musical instrument via, for example, an audio cable. The audio interface 11 transmits the audio signal subjected to signal processing to an audio device such as a speaker via, for example, an audio cable.
[0014] The network interface 12 communicates with another device (for example, a PC or the like) different from the mixing device 1a via a communication line. The communication line is, for example, the Internet or a LAN (Local Area Network). The network interface 12 and another device such as a PC communicate wirelessly or by wire. Note that the network interface 12 may transmit and receive an audio signal via the network in accordance with a standard such as Dante (registered trademark).
[0015] The flash memory 13 stores various programs. The various programs are, for example, a program for operating the mixing device 1a or a program for executing sound processing according to the sound signal processing method of the present invention. Note that the flash memory 13 does not necessarily have to store various programs. The various programs may be stored in other devices such as a server. In this case, the mixing device 1a receives various programs from other devices such as a server.
[0016] The RAM 14 reads out the program stored in the flash memory 13 and temporarily stores it.
[0017] The CPU 15 (an example of a processor) executes various processes by reading out the program stored in the flash memory 13 into the RAM 14. The various processes are, for example, a process of converting an analog sound signal into a digital sound signal, sound processing according to the sound signal processing method of the present invention, etc. The CPU 15 converts an analog sound signal into a digital sound signal based on a preset sampling frequency and quantization bit number. The sampling frequency is, for example, 48 kHz, and the quantization bit number is, for example, 24 bit.
[0018] The DSP 18 performs signal processing on the sound signal received via the audio interface 11 or the network interface 12. The signal processing is acoustic processing such as mixing or effects. The DSP 18 performs signal processing based on the current data stored in the RAM 14. The current data is the current various parameter values of the sound signal processing (such as gain adjustment, effect processing, and mixing processing) executed by the DSP 18. The various parameter values are changed by the user's operation via the user interface 17. When the CPU 15 receives the user's operation via the user interface 17, the current data is updated. The sound signal after the signal processing is transmitted to the audio interface 11 via the bus 19. Note that the DSP 18 may be composed of a plurality of DSPs.
[0019] The display unit 16 displays various information based on the control of the CPU 15. For example, the display unit 16 displays the level of the sound signal. The display unit 16 includes a screen 16a and a meter 16b. Multiple meters 16b are provided for each channel strip. In the example shown in Figure 2, the meter 16b includes eight meters 16b1 to 16b8 corresponding to eight input channel strips and two meters 16b9 and 16b10 corresponding to two output channel strips. Note that the number of channel strips and the number of meters 16b are not limited to 10.
[0020] The screen 16a is, for example, a liquid crystal display. The screen 16a displays an image based on the control of the CPU 15.
[0021] The meter 16b consists of multiple LEDs for displaying the level of the sound signal. The CPU 15 turns on or off the multiple LEDs of the meter 16b based on the level of the sound signal. For example, in the example shown in Figure 2, each of the meters 16b1 to 16b10 consists of 12 LEDs arranged vertically on the page. Each of the 12 LEDs is assigned a corresponding level value. For example, in Figure 2, if the level of the sound signal is silent, i.e., -∞dB, the CPU 15 turns off all 12 LEDs. If the level of the sound signal is at its maximum, i.e., 0dB, the CPU 15 turns on all 12 LEDs. Also, for example, if the level of the sound signal is -12dB, the CPU 15 turns on the bottom four LEDs. This allows the user to visually know the level of the sound signal input to each input channel. Note that the meter 16b is not limited to LEDs; it may also be an image displayed on the screen 16a.
[0022] The user interface 17 is an example of a set of controls that receive operations on the mixing device 1a from the user of the mixing device 1a (hereinafter referred to as "user"). The user interface 17 includes, for example, a knob 17a, a fader 17b, an increase / decrease button 17c, a store button 17d, a recall button 17e, and a touch panel 17f, as shown in Figure 2. A knob 17a and a fader 17b are provided for each channel strip.
[0023] The touch panel 17f is stacked on the screen 16a. The touch panel 17f accepts touch operations from the user.
[0024] Knob 17a accepts adjustment of the gain of audio signals input to multiple channels. In the example shown in Figure 2, knob 17a includes eight knobs 17a1 to 17a8 corresponding to eight input channel strips and two knobs 17a9 and 17a10 corresponding to two output channel strips. Note that the number of knobs 17a is not limited to 10.
[0025] The fader 17b accepts the level adjustment amount of audio signals input to multiple channels. The user adjusts the amount of audio signal sent from each input channel to the output channel by sliding the fader 17b. In the example shown in Figure 2, the fader 17b includes eight faders 17b1 to 17b8 corresponding to eight input channel strips and two faders 17b9 and 17b10 corresponding to two output channel strips. Note that the number of faders 17b is not limited to 10.
[0026] The store button 17d is a button that instructs the user to store scene memory data (scene data). By operating (pressing) the store button 17d, the user can store the current data as a single scene data in the flash memory 13.
[0027] The increase / decrease button 17c is a button that accepts the operation of selecting the scene memory to be saved and recalled from among multiple scene memories.
[0028] The recall button 17e is a button that accepts an instruction (scene recall) to retrieve scene data stored in the flash memory 13 as current data into the RAM 14. By operating (pressing) the recall button 17e, the user can retrieve the necessary scene memory data and thereby recall the settings of various parameters.
[0029] The functions of the increase / decrease button 17c, the store button 17d, and the recall button 17e may be configured using a GUI (Graphical User Interface) with a touch panel 17f.
[0030] The signal processing performed by the mixing device 1a will be explained below with reference to the figures. Figure 3 is a block diagram of the signal processing performed by the mixing device 1a. Figure 4 is a block diagram showing the processing configuration of input patch 21a, input channel 22a, mixing bus 23a, output channel 24a, and output patch 25a. Note that Figure 4 only shows the signal processing on input channel 1, and the description of the signal processing on input channels 2-32 is omitted.
[0031] As shown in Figure 3, in the mixing device 1a, signal processing is functionally performed by input patch 21a, input channel 22a, mixing bus 23a, output channel 24a, and output patch 25a.
[0032] Input patch 21a receives audio signals from multiple input ports (e.g., analog or digital ports) in the audio interface 11. Input patch 21a assigns one of the multiple input ports to at least one of the multiple input channels included in input channel 22a (e.g., a total of 32 channels from input channels 1 to 32). In this way, input patch 21a sends audio signals to each input channel of input channel 22a.
[0033] Each input channel can be arbitrarily assigned a corresponding control. If the mixing device 1a has, for example, eight knobs 17a and eight faders 17b, then input channels 1-8 can be assigned to each of the eight knobs 17a and eight faders 17b. For example, input channel 1 is assigned knob 17a1 and fader 17b1 as shown in Figure 2. In this case, the user can adjust the gain of the audio signal input to input channel 1 by operating knob 17a1. Similarly, the user can adjust the feed amount of the audio signal output from input channel 1 by operating fader 17b1.
[0034] The following explanation will use signal processing in input channel 1 as an example. As shown in Figure 4, input channel 1 functionally includes a head amplifier (HA) 220, a signal processing block 221, a fader section (FADER) 222, a pan section (PAN) 223, and a send section (SEND) 224.
[0035] The head amplifier 220 adjusts the gain of the audio signal input to input channel 1. The head amplifier 220 then transmits the gain-adjusted audio signal to the signal processing block 221.
[0036] The signal processing block 221 performs signal processing such as equalization or compression on the audio signal whose gain has been adjusted by the head amplifier 220.
[0037] The fader unit 222 adjusts the level of the audio signal processed by the signal processing block 221 based on the feed amount set by the fader 17b, which is an operator.
[0038] The mixing bus 23a includes the stereo bus 231 and the MIX bus 232. The stereo bus 231 is a two-channel bus that serves as the master output. The pan unit 223 adjusts the balance of the audio signals supplied to each of the two channels of the stereo bus 231. As shown in Figure 4, the pan unit 223 outputs the balanced audio signals to the stereo bus 231. The stereo bus 231 is connected to the output channel 24a. The stereo bus 231 transmits the audio signals received from the pan unit 223 to the output channel 24a.
[0039] The MIX bus 232 includes multiple channels (for example, 48 channels as shown in Figure 3 or Figure 4). The sender 224 switches whether or not to supply an audio signal to each channel of the MIX bus 232 based on user operation. The sender 224 also adjusts the level of the audio signal supplied to each channel of the MIX bus 232 based on the send amount set by the user. As shown in Figure 4, the sender 224 outputs the level-adjusted audio signal to the MIX bus 232. The MIX bus 232 is connected to output channel 24a. The MIX bus 232 transmits the audio signal input from the sender 224 to output channel 24a.
[0040] Output channel 24a has multiple channels. Each channel of output channel 24a performs various signal processing on the audio signal received from mixing bus 23a. Each channel of output channel 24a transmits the processed audio signal to output patch 25a.
[0041] Output patch 25a assigns one of the multiple output ports (analog output ports or digital output ports) to at least one of the multiple channels included in output channel 24a. This allows the processed audio signal to be sent to the audio interface 11.
[0042] The processing performed by the input patch 21a, input channel 22a, mixing bus 23a, output channel 24a, and output patch 25a described above is carried out based on the values of various parameters.
[0043] In the above process, the CPU 15 illuminates the meter 16b1 corresponding to input channel 1 based on the level (dB) of the sound signal that has been gain-adjusted in the head amplifier 220. The CPU 15 generates meter data to control the meter 16b1 based on multiple samples of the sound signal over a predetermined time (for example, 1 / 60th of a second). The meter data is an example of time-series volume data. The CPU 15 controls the meter 16b1 based on the meter data to display the level of the sound signal.
[0044] For example, the sampling frequency of the audio signal is 48 kHz, while the sampling frequency of the meter data is lower at 60 Hz. For instance, CPU 15 acquires 800 samples of the audio signal corresponding to 1 / 60th of a second. CPU 15 then generates 1 sample of meter data by, for example, averaging the 800 samples of the audio signal to reduce the sampling frequency.
[0045] For example, the quantization bit count of an audio signal is 24 bits, while the quantization bit count of meter data is 4 bits, which is necessary to turn 12 LEDs on or off, and is smaller than the quantization bit count of the audio signal. The CPU 15 reduces the quantization bit count of the audio signal, which is quantized with 24 bits (approximately 16.77 million gradations), and rounds it down to 4-bit meter data with 12 gradations.
[0046] The following describes the audio signal processing (hereinafter referred to as processing P) with reference to the diagram. Figure 5 is a flowchart showing an example of the processing of the mixing device 1a.
[0047] The mixing device 1a starts the operation shown in Figure 5 when, for example, the program related to process P is executed (Figure 5: START).
[0048] After the start of process P, the CPU 15 accepts an operation to select one of the multiple input channels (the first channel) (Figure 5: step S11). For example, the CPU 15 displays buttons on the screen 16a corresponding to at least one of the input channels 1-32 in input channel 22a (hereinafter referred to as selection buttons). The user selects input channel 1 by, for example, touching the selection button corresponding to input channel 1. The following describes the case where input channel 1 is selected.
[0049] Next, the CPU 15 receives the audio signal from the selected input channel 1 (first channel) (Figure 5: Step S12).
[0050] Next, the CPU 15 identifies setting data to be set in the mixing device 1a (Figure 5: Step S13). In this embodiment, the setting data includes the gain value of the head amplifier 220. Based on the time-series volume data, the CPU 15 identifies the gain value to be set for the head amplifier 220 (an example of setting data). For example, the CPU 15 identifies the gain of the head amplifier 220 such that clipping does not occur in the sound input to the selected input channel (for example, the gain of the head amplifier 220 is identified so that the peak level of the input sound signal does not exceed -6dB).
[0051] In this embodiment, the CPU 15 identifies the setting data by processing with artificial intelligence, such as a neural network (DNN (Deep Neural Network)). A skilled mixer engineer looks at the meter 16b1 and sets the gain of the head amplifier 220 so that the peak level of the input sound signal does not exceed -6dB. In other words, a skilled mixer engineer sets the gain of the head amplifier 220 based on meter data. Therefore, there is a correlation between the time-series volume data and the gain of the head amplifier 220. For this reason, the CPU 15 can train a predetermined model to learn the relationship between the time-series volume data and the gain of the head amplifier 220. The CPU 15 identifies the setting data using a first trained model that has learned the relationship between the time-series volume data (meter data) and the setting data (gain of the head amplifier 220).
[0052] A skilled mixer engineer adjusts the gain of the head amplifier 220 by looking at meter data for approximately 3 samples (1 / 60 sec x 3 samples ≈ 50 msec). After a predetermined model learns the relationship between time-series volume data and the gain of the head amplifier 220, the CPU 15 can identify setting data during the artificial intelligence execution phase by using, for example, 3 samples of meter data (by acquiring time-series volume data for approximately 1 / 60 sec x 3 samples ≈ 50 msec).
[0053] Furthermore, during the execution phase of the artificial intelligence, the CPU 15 may use a number of samples corresponding to the volume indicators used by skilled mixer engineers when setting the gain of the head amplifier 220. A skilled mixer engineer adjusts the gain of the head amplifier 220 using, for example, a VU meter or a loudness meter as an indicator. The VU meter shows the average volume over a period of 300 msec. The CPU 15 may, for example, acquire meter data for multiple samples over a period of 300 msec and identify the setting data by using the acquired meter data for multiple samples. The loudness meter shows the loudness value over a period of 400 msec (momentary loudness), or the loudness value over a period of 3 seconds (short-term loudness), etc. The CPU 15 may, for example, acquire meter data for multiple samples over a period of 400 msec or 3 seconds and identify the setting data by using the acquired meter data for multiple samples. Furthermore, since the VU meter and loudness meter are just examples of volume indicators, other meters may also be used as volume indicators.
[0054] As described above, a skilled engineer adjusts the gain of the head amplifier 220 to a range of approximately 50 msec to 3 sec, using meters such as VU meters or loudness meters as a reference, to achieve an appropriate volume. The CPU 15 replicates this adjustment by a skilled mixer engineer and adjusts the gain of the head amplifier 220 to a range of approximately 50 msec to 3 sec to achieve an appropriate volume. In this way, the CPU 15 can automatically adjust the gain of the head amplifier 220 in the same way that a skilled mixer engineer would manually adjust the gain.
[0055] The CPU 15 outputs the identified setting data (gain of the head amplifier 220) to the RAM 14 (Figure 5: Step S14). The CPU 15 updates the gain value of the head amplifier 220 in the current data of the RAM 14 with the identified gain value. Based on this, the DSP 18 performs signal processing on the audio signal.
[0056] The execution of process P is completed by performing the processes from steps S11 to S14 described above (Figure 5: END).
[0057] [effect] A skilled mixer engineer adjusts the gain of the head amplifier 220 based on the display of meter 16b to prevent clipping. The mixing device 1a replicates this adjustment method of a skilled mixer engineer, for example, using artificial intelligence. The mixing device 1a can automatically adjust the gain of the head amplifier 220 in the same way that a skilled mixer engineer would manually adjust the gain.
[0058] The time-series volume data may be audio signals with multiple samples, but as described above, it is preferable that it be meter data. The sampling frequency of the meter data is significantly lower than the sampling frequency of the audio signal. Also, the number of quantization bits of the meter data is significantly lower than the number of quantization bits of the audio signal. Therefore, by performing the learning and execution stages with meter data, the mixing device 1a can determine the gain of the head amplifier 220 with significantly less computational effort than by performing the learning and execution stages of a predetermined model using audio signals.
[0059] [Modification 1 of the First Embodiment] The mixing apparatus 1b according to Modification 1 of the First Embodiment will be described below with reference to the figures. Figure 6 shows the external appearance of the mixing apparatus 1b according to Modification 1 of the First Embodiment.
[0060] The screen 16a of the mixing device 1b displays the gain of the identified head amplifier 220. For example, the CPU 15 of the mixing device 1b displays an image resembling the knob 17a on the screen 16a, as shown in Figure 6. For example, the CPU 15 displays on the screen 16a an image of the knob 40 (an image resembling the knob 17a1) that indicates the identified gain, and an image showing the range of the identified gain ±α (where α is an arbitrary value). The user adjusts the knob 17a1 by referring to the knob 40 displayed on the screen 16a. The user can arbitrarily decide whether or not to set the gain of the head amplifier 220 to the gain of the head amplifier 220 identified by the trained model.
[0061] The CPU 15 may display an image showing only the gain of the identified head amplifier 220, without displaying the range of the identified gain ±α. For example, it may display a text message indicating the value of the identified gain (e.g., a text message such as "Please set knob 17a1 to -3dB").
[0062] [Modification 2 of the First Embodiment] The mixing apparatus 1c according to Modification 2 of the First Embodiment will now be described with reference to the figures. Figure 7 shows the external appearance of the mixing apparatus 1c according to Modification 2 of the First Embodiment.
[0063] The CPU 15 of the mixing device 1c displays the gain of the identified head amplifier 220 and then accepts an operation to set that gain as the current data. For example, as shown in Figure 7, the CPU 15 displays a button Y on the screen 16a that accepts an operation to set the gain of the identified head amplifier 220 as the current data. If the user wants to set the gain of the head amplifier 220 displayed on the screen 16a as the current data, they touch button Y. In this case, the CPU 15 updates the current data with the gain of the identified head amplifier 220. On the other hand, if the CPU 15 detects an operation of button N displayed on the screen 16a, it does not update the current data with the gain of the identified head amplifier 220. In this way, the user can check the details of the gain of the identified head amplifier 220 and then decide whether or not to set it as that gain.
[0064] [Modification 3 of the First Embodiment] The mixing apparatus 1d according to Modification 3 of the First Embodiment will be described below with reference to the figures. Figure 8 is a flowchart showing an example of the processing of the mixing apparatus 1d according to Modification 3 of the First Embodiment.
[0065] The CPU 15 of the mixing device 1d determines whether or not it has received an operation from the user (for example, an operation related to adjusting the gain of the head amplifier 220) (Figure 8: Step S21). If the CPU 15 has received an operation from the user (Figure 8: Step S21 Yes), it stops outputting the identified setting data (for example, the gain of the head amplifier 220) to the RAM 14 (Figure 8: Step S22). The CPU 15 then sets the setting data based on the operation received from the user into the mixing device 1d (Figure 8: Step S23). This prevents manual adjustment of various parameters by the user from being hindered by automatic adjustment by the mixing device 1d.
[0066] On the other hand, in step S21, if the CPU 15 has not received any input from the user (Figure 8: Step S21 No), it resumes outputting the identified setting data (Figure 8: Step S24). As a result, the mixing device 1d automatically adjusts the gain of the head amplifier 220 when there is no input from the user. The mixing device 1d can appropriately switch whether or not to adjust the gain of the head amplifier 220 depending on whether or not there is input from the user.
[0067] [Second Embodiment] Hereinafter, the mixing device 1e according to the second embodiment will be described with reference to Figure 2. In the description, the case in which input channel 1 is selected by the user will be used as an example, similar to the first embodiment.
[0068] The CPU 15 of the mixing device 1e identifies setting data (fader value of the first channel) based on the fader value (data related to the second channel) of input channels 2-32 (second channel), which are different from input channel 1 (first channel). In this embodiment, the setting data is the level adjustment amount accepted by the fader 17b. For example, the CPU 15 adjusts the fader value of input channel 1 based on the fader values of input channels 2-8 so that the volume balance in the mixing of input channels 1-8 is appropriate.
[0069] In this embodiment, the CPU 15 identifies the setting data using a second trained model that has learned the relationship between the data (fader values) related to input channels 2-32 (second channel) and the setting data (fader values of the first channel). An experienced mixer engineer looks at the fader values 17b2-17b8 of input channels 2-32 and adjusts the fader value 17b1 of input channel 1 while considering the volume balance. Therefore, there is a correlation between the data related to input channels 2-32 (fader values 17b2-17b8) and the fader value 17b1. For this reason, the CPU 15 can train a predetermined model to learn the relationship between the data related to input channels 2-32 and the fader value 17b1.
[0070] [effect] For example, when setting the value of fader 17b1, a skilled mixer engineer will adjust it while considering the volume balance, referencing the values of other faders 17b2-17b8. The mixing device 1e replicates the same volume balance adjustments between channels that are performed by a skilled mixer engineer (for example, the kind of adjustments that are made over time by a skilled mixer engineer when creating a CD audio source). As a result, the mixing device 1e can automatically adjust the volume balance between channels in the same way that a skilled mixer engineer would manually adjust the volume balance between channels.
[0071] [Modification 1 of the second embodiment] The mixing device 1f according to Modification 1 of the Second Embodiment will be described below with reference to the figures. Figure 9 shows an image displayed on the screen 16a of the mixing device 1f according to Modification 1 of the Second Embodiment.
[0072] In this modified example, the data relating to input channels 2-32, which are different from input channel 1, is data relating to the type of sound source (instrument name) (text data, etc.). For example, a skilled mixer engineer may adjust the balance between channels according to the type of instrument corresponding to each of the input channels 1-32. The CPU 15 reproduces this kind of adjustment made by a skilled mixer engineer. As an example, the CPU 15 adjusts the fader value of input channel 1 so that the volume relating to input channel 1 (sound source: vocals) is greater than the volume relating to input channel 2 (sound source: guitar). In this way, the mixing device 1e can reproduce the channel balance adjustment made by a skilled mixer engineer.
[0073] [Modification 2 of the second embodiment] The following description of the mixing device 1g according to Modification 2 of the Second Embodiment will be made with reference to Figure 2. When a scene recall is performed, the CPU 15 of the mixing device 1g updates the current data with the read scene data. For this reason, it is preferable for the CPU 15 to re-adjust the volume balance between channels using the learned model. Accordingly, when the CPU 15 receives a scene recall (when the recall button 17e shown in Figure 2 is pressed by the user), it identifies the level adjustment amount (setting data) for input channel 1 based on the read scene data. The scene data contains useful information for adjusting the volume balance between channels, such as instrument information. Therefore, by using the read scene data, the mixing device 1g can more accurately reproduce the volume balance adjustment between channels performed by a skilled mixer engineer.
[0074] [Modification 3 of the second embodiment] The mixing device 1h according to Modification 3 of the Second Embodiment will be described below with reference to Figure 9. The mixing device 1h uses data relating to input channels to which no controls (knobs 17a, faders 17b) are assigned when identifying setting data.
[0075] In the example shown in Figure 9, the CPU 15 of the mixing device 1h assigns input channels 1-8 (first channel group) to the control. On the other hand, the mixing device 1h does not assign input channels 9-32 (second channel group) to the control. Even in this case, the mixing device 1h adjusts the fader value of input channel 1 based on the fader values of input channels 2-32, for example. This allows the mixing device 1h to adjust the volume balance between channels by considering the fader values of the undisplayed channels (input channels 9-32) without having to check those fader values.
[0076] When a mixing device has a vast number of channels, the number of input channels not assigned to any controls becomes enormous. In this case, it becomes difficult for the user to consider the fader values of all input channels. However, the mixing device 1h automatically sets the fader value of the selected input channel 1 based on the fader values of all input channels.
[0077] The mixing device 1e may have a mode (split mode) that divides input channels 1-32 into a first input channel 1-16 (first channel group), which is a first signal processing system, and a second input channel 1-16, which is a second signal processing system. In this case, the CPU 15 of the mixing device 1h assigns, for example, the first input channel 1-8 (first channel group) of the first input channel 1-16 of the first signal processing system to an operator, while not assigning the first input channels 9-16 of the first signal processing system and the second input channel 1-16 (second channel group) of the second signal processing system to an operator. At this time, the user selects, for example, input channel 1 of the first signal processing system. The mixing device 1h determines the fader value of the first input channel 1 of the first signal processing system based on the fader values of the first input channels 2-16 of the first signal processing system and the fader values of the second input channels 1-16 of the second signal processing system. However, the audio signal input to the first input channel 1 of the first signal processing system and the signal input to the second input channel 1 of the second signal processing system are the same. Therefore, the mixing device 1h does not necessarily have to use the fader value of the second input channel 1 of the second signal processing system when determining the fader value of input channel 1.
[0078] The configurations of the mixing devices 1a to 1h may be combined in any way. [Explanation of Symbols]
[0079] 1a~1h: Mixing device 11: Audio Interface 12: Network Interface 13: Flash memory 14: RAM 15:CPU 16: Display 16a: Screen 16b, 16b1, 16b1-16b10: Meter 17: User Interface 17a, 17a1-17a10: Knob 17b, 17b1-17b10: Feder 18:DSP 19: Bus 21a: Input patch 22a: Input channel 23a: Mixing bus 24a: Output channel 25a: Output patch
Claims
1. The sound signal processing device is The sound signal processing device accepts an operation to select at least one first channel from among a plurality of channels, The audio signal of the selected first channel is input, Based on the time-series volume data based on the input sound signal, or the data relating to a second channel different from the first channel among the multiple channels, setting data for setting in the sound signal processing device is identified. Output the identified setting data. A method for processing sound signals, The setting data is identified using a first trained model that has learned the relationship between the time-series volume data and the setting data, or using a second trained model that has learned the relationship between the data for the second channel and the setting data. Audio signal processing method.
2. The time-series volume data is obtained by reducing the sampling frequency of the sound signal or by reducing the number of quantization bits of the sound signal. The sound signal processing method according to claim 1.
3. The sound signal processing device includes a display that shows the level of the input sound signal based on the time-series volume data. The sound signal processing method according to claim 1 or claim 2.
4. The parameters for signal processing performed on each of the aforementioned multiple channels are stored in memory as scene data. The system accepts a scene recall request to read scene data stored in the aforementioned memory, When the aforementioned scene recall is received, the setting data is identified based on the retrieved scene data. The sound signal processing method according to claim 1 or claim 2.
5. Identify the type of sound source for the sound signal of the first channel, The setting data is identified according to the type of sound source identified above. The sound signal processing method according to claim 1 or claim 2.
6. The aforementioned sound signal processing device is equipped with a plurality of controls that accept user input, Multiple channels include the first channel group and the second channel group, The first group of channels is assigned to the plurality of operators. The second channel includes channels of the second channel group that are not assigned to the plurality of operators. The sound signal processing method according to claim 1 or claim 2.
7. The aforementioned sound signal processing device is It is equipped with a head amplifier that adjusts the gain of the audio signals input to the aforementioned multiple channels. The setting data includes the gain of the head amplifier. The gain of the head amplifier is determined based on the time-series volume data. The sound signal processing method according to claim 1 or claim 2.
8. The sound signal processing device is A fader that accepts the amount of level adjustment for the sound signals input to the aforementioned multiple channels, It is equipped with, The setting data includes the level adjustment amount received by the fader. The level adjustment amount is determined based on the data relating to the second channel. The sound signal processing method according to claim 1 or claim 2.
9. When the user requests an adjustment to the setting data, the output of the setting data is stopped, and the setting data received from the user is set in the sound signal processing device. The sound signal processing method according to claim 1 or claim 2.
10. If it is determined that the user has not requested any adjustments to the setting data, the output of the setting data will be resumed. The sound signal processing method according to claim 9.
11. Based on the output setting data, the contents of the setting data are displayed on the display unit. The sound signal processing method according to claim 1 or claim 2.
12. After displaying the contents of the setting data, the system accepts an operation to determine whether or not to set the setting data. When an operation to set the aforementioned setting data is received, the setting data is set in the sound signal processing device. The sound signal processing method according to claim 11.
13. The system accepts an operation to select at least one first channel from among multiple channels. The audio signal of the selected first channel is input, Based on the time-series volume data derived from the input sound signal, or data relating to a second channel different from the first channel among the multiple channels, setting data for setting the device is identified. Output the identified setting data. It has a processor that performs the processing. A sound signal processing device, The processor identifies the setting data using a first trained model that has learned the relationship between the time-series volume data and the setting data, or using a second trained model that has learned the relationship between the data relating to the second channel and the setting data. Sound signal processing device.
14. The processor acquires the time-series volume data by reducing the sampling frequency of the sound signal or by reducing the number of quantization bits of the sound signal. The sound signal processing device according to claim 13.
15. The system further includes a display that shows the level of the input sound signal based on the aforementioned time-series volume data. The sound signal processing device according to claim 13 or claim 14.
16. The system further includes a memory that stores the parameters of the signal processing performed on each of the audio signals of the aforementioned multiple channels as scene data. The aforementioned processor, The system accepts a scene recall request to read scene data stored in the aforementioned memory, When the aforementioned scene recall is received, the setting data is identified based on the retrieved scene data. The sound signal processing device according to claim 13 or claim 14.
17. The aforementioned processor, Identify the type of sound source for the sound signal of the first channel, The setting data is identified according to the type of sound source identified above. The sound signal processing device according to claim 13 or claim 14.
18. It further includes multiple controls that accept user input, Multiple channels include the first channel group and the second channel group, The processor assigns the first channel group to the plurality of operators, The second channel includes channels of the second channel group that are not assigned to the plurality of operators. The sound signal processing device according to claim 13 or claim 14.
19. The system further includes a head amplifier that adjusts the gain of the audio signals input to the aforementioned multiple channels. The setting data includes the gain of the head amplifier. The aforementioned processor, The gain of the head amplifier is determined based on the time-series volume data. The sound signal processing device according to claim 13 or claim 14.
20. The system further comprises faders that accept the amount of level adjustment for the sound signals input to the plurality of channels, The setting data includes the level adjustment amount received by the fader. The aforementioned processor, The level adjustment amount is determined based on the data relating to the second channel. The sound signal processing device according to claim 13 or claim 14.
21. When the processor receives a request from the user to adjust the setting data, it stops outputting the setting data and performs the settings based on the setting data received from the user. The sound signal processing device according to claim 13 or claim 14.
22. If the processor determines that it has not received any request from the user to adjust the setting data, it will resume outputting the setting data. The sound signal processing device according to claim 21.
23. It also has a display, The display unit displays the contents of the setting data based on the output setting data. The sound signal processing device according to claim 13 or claim 14.
24. The aforementioned processor, After displaying the contents of the setting data on the display unit, the system accepts an operation to determine whether or not to set the setting data. When an operation to set the aforementioned setting data is received, the setting data is set. The sound signal processing device according to claim 23.