Sound signal processing method and sound signal processing apparatus

WO2026168535A1PCT designated stage Publication Date: 2026-08-13YAMAHA CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-02-05
Publication Date
2026-08-13

Smart Images

  • Figure JP2026004176_13082026_PF_FP_ABST
    Figure JP2026004176_13082026_PF_FP_ABST
Patent Text Reader

Abstract

This sound signal processing method inputs sound signals of a plurality of sound sources, receives a designation of a specific sound source among the plurality of sound sources, and performs pseudo-jitter processing for adding fluctuation in a time axis direction to sound signals other than that of the designated specific sound source.
Need to check novelty before this filing date? Find Prior Art

Description

Sound signal processing method and sound signal processing apparatus

[0001] One embodiment of the present invention relates to a sound signal processing method and a sound signal processing apparatus.

[0002] Patent Document 1 discloses a configuration that performs signal processing to clarify a sound corresponding to a specific sound source.

[0003] Japanese Patent Application Laid-Open No. 2018-191127

[0004] Since the signal processing of Patent Document 1 changes the amplitude of the sound signal, the timbre and the sound image change.

[0005] An object of one embodiment of the present invention is to provide a sound signal processing method that can minimize changes in timbre and sound image and emphasize a specific sound source while maintaining the volume balance between sound sources.

[0006] The sound signal processing method according to one embodiment of the present invention inputs sound signals of a plurality of sound sources, accepts designation of a specific sound source among the plurality of sound sources, and applies a pseudo jitter process that adds fluctuations in the time axis direction to sound signals other than the designated specific sound source.

[0007] According to one embodiment of the present invention, it is possible to minimize changes in timbre and sound image and emphasize a specific sound source while maintaining the volume balance between sound sources.

[0008] This is a block diagram showing the configuration of the sound signal processing device 1. This is a block diagram showing the functional configuration of the processor 12. This is a flowchart showing the operation of the processor 12. This is a block diagram showing the configuration of the sound signal processing device 1 according to Modification 1. This is a block diagram showing the functional configuration of the processor 12 according to Modification 1. This is a block diagram showing the configuration of the sound signal processing device 1 according to Modification 2. This is a block diagram showing the functional configuration of the processor 12 according to Modification 2. This is a block diagram showing the functional configuration of the processor 12 according to Modification 3. This is a block diagram showing the functional configuration of the processor 12 according to Modification 4. This is a block diagram showing the functional configuration of the processor 12 according to Modification 5. This is a block diagram showing the functional configuration of the processor 12 according to Modification 6. This is a block diagram showing the functional configuration of the processor 12 according to Modification 8. This is a diagram showing an example of an application program screen (GUI). This is a GUI showing an example of displaying the first sound pressure distribution. This is a GUI showing an example of displaying the second sound pressure distribution. This is a GUI showing an example of displaying the third sound pressure distribution.

[0009] (First Embodiment) Figure 1 is a block diagram showing the configuration of the sound signal processing device 1. The sound signal processing device 1 includes a communication unit 11, a processor 12, a RAM 13, a flash memory 14, a display 15, a user interface 16, and an audio interface 17.

[0010] The sound signal processing device 1 may be, for example, an audio device such as an audio mixer, or an information processing device such as a personal computer, smartphone, or tablet computer.

[0011] The communication unit 11 communicates with other devices such as servers. The communication unit 11 has wireless communication functions such as Bluetooth® or Wi-Fi®, and wired communication functions such as USB or LAN. The communication unit 11 receives content data from other devices such as servers.

[0012] The display unit 15 consists of an LCD or OLED, etc. The display unit 15 displays the video output by the processor 12.

[0013] The user interface 16 is an example of an operating unit. The user interface 16 consists of a mouse, keyboard, or touch panel, etc. The user interface 16 receives user input. The touch panel may be stacked on the display unit 15.

[0014] The audio interface 17 has analog audio terminals or digital audio terminals, etc., and connects to audio equipment. In this embodiment, the audio interface 17 connects to a plurality of speakers 20, 21, 22, and 23 as an example of audio equipment, and outputs sound signals to these speakers.

[0015] The processor 12 consists of a CPU, DSP, or SoC (System on a Chip), etc. The processor 12 performs various operations by reading a program from the flash memory 14, which is a storage medium, and temporarily storing it in RAM 13. Note that the program does not need to be stored in the flash memory 14. The processor 12 may, for example, download a program from another device such as a server when necessary and temporarily store it in RAM 13.

[0016] Figure 2 is a block diagram showing the functional configuration of the processor 12. Functionally, the processor 12 includes a decoder 120, a pseudo-jitter processing unit 121, a localization processing unit 122, and a sound source designation unit 123.

[0017] Figure 3 is a flowchart showing the operation of the processor 12. The decoder 120 of the processor 12 decodes the content data received via the communication unit 11 and inputs the sound signals of multiple sound sources to the pseudo-jitter processing unit 121 (S11). The sound source designation unit 123 accepts the designation of a specific sound source from among the multiple sound sources (S12). The sound signal processing method of this embodiment corresponds to an object-based method. An object-based method is a method in which the sound signal for each object sound source, which has position information, is stored in the content data. In contrast, a channel-based method is a method in which the sound signals for each sound source are mixed in advance and stored in the sound signals of one or more channels.

[0018] The sound source designation unit 123 accepts the designation of a certain object sound source as a specific sound source. For example, the sound source designation unit 123 designates an object sound source specified by the user via the user interface 16 as a specific sound source. The user may, for example, designate a sound source that they are particularly interested in listening to as a specific sound source. For example, if the content data is audio data of performances by multiple artists, and the user is focusing on a particular artist, they may designate the performance sound of that artist as a specific sound source.

[0019] The pseudo-jitter processing unit 121 applies pseudo-jitter processing to sound signals other than the specified specific sound source, adding fluctuations in the time axis direction (S13). Pseudo-jitter processing is, for example, a process that changes the phase. The pseudo-jitter processing unit 121 rotates only the phase characteristics of the sound signal at a predetermined rotational speed without changing the amplitude characteristics of the frequency characteristics of the sound signal, for example, using a filter such as an all-pass filter. The predetermined rotational speed is, for example, 2 rotations per second (720° / second). When the phase rotational speed is set to 2 rotations per second (720° / second) or less, waveform distortion caused by the phase rotation is less likely to be perceived audibly. Furthermore, more preferably, the phase rotational speed is 1.5 rotations per second or less.

[0020] As a result, the pseudo-jitter processing unit 121 can reduce the sense of localization of sound sources by adding fluctuations in the time axis direction while suppressing the manifestation of waveform distortion caused by phase rotation. In other words, the pseudo-jitter processing unit 121 performs processing to emphasize a specific sound source by reducing the sense of localization of sound sources other than the specific sound source.

[0021] Furthermore, if the phase of the sound signal is rotated at an excessively slow speed, fluctuations in the time axis direction may be small, and it may not be possible to reduce the sense of localization of the sound source. Therefore, it is preferable to set the phase rotation speed to 0.01 rotations per second (3.6° / second) or more. That is, it is preferable to set the phase rotation speed to 0.01 rotations per second or more and 2 rotations per second or less. Alternatively, it is more preferable to set the rotation speed to 0.1 rotations per second or more and 1.5 rotations per second or less. It is even more preferable to set the rotation speed to 1 rotation per second (360° / second).

[0022] The localization processing unit 122 applies localization processing to the sound signals of all object sound sources, including the sound signals after pseudo-jitter processing, based on the position information of each object sound source. The localization processing unit 122 adjusts the amplitude of the sound signals distributed to the multiple speakers 20, 21, 22, and 23 so that the sound image is localized to the position information corresponding to the object sound source. For example, if the same sound signal is output to two speakers at the same level, the sound image will be localized at the center of the line segment connecting the two speakers. If the same sound signal is output to two speakers at different levels, the sound image will be localized closer to the position of the speaker with the higher level on the line segment connecting the two speakers. Therefore, the localization processing unit 122 adjusts the amplitude of the sound signals distributed to speakers 20, 21, 22, and 23 based on the position information of the sound source and the position information of the multiple speakers 20, 21, 22, and 23, thereby localizing the sound source.

[0023] The location information of the sound source and the speaker is represented by a two-dimensional or three-dimensional coordinate system with a specific location as the origin. This coordinate system may be a logical coordinate system (information normalized between 0 and 1), or it may be a real-world coordinate system corresponding to an actual venue such as a church, live music venue, or concert hall.

[0024] As described above, the sound signal processing device 1 performs pseudo-jitter processing on sound sources other than the specific sound source to be highlighted, thereby reducing the sense of localization. Pseudo-jitter processing is a process that introduces fluctuations in phase without changing the amplitude. As a result, the sound signal processing device 1 can highlight a specific sound source without changing the sound quality or volume balance of all sound sources. Therefore, users can obtain a new customer experience in which they can clearly hear a specific sound source while maintaining the sound quality and volume balance as intended by the content creator.

[0025] (Modification 1) Figure 4 is a block diagram showing the configuration of the sound signal processing device 1 according to Modification 1, and Figure 5 is a block diagram showing the functional configuration of the processor 12 according to Modification 1.

[0026] The audio I / F 17 of the sound signal processing device 1 according to Modification 1 connects headphones 40 as an example of audio equipment and outputs sound signals to the headphones 40.

[0027] The localization processing unit 122 according to Modification 1 performs binaural processing as localization processing. Binaural processing is a process that localizes an object sound source to a predetermined position by convolving a head-related transfer function (HRTF) into the sound signal. The HRTF corresponds to the transfer function between the predetermined position and the listener's ears. The HRTF is a transfer function that expresses the loudness, arrival time, and frequency characteristics of sound from a sound source at a certain position to the left and right ears, respectively. The localization processing unit 122 convolves the HRTF into the sound signal of the object sound source based on the position of the object sound source. As a result, the sound of the object sound source is localized to a position corresponding to the position information.

[0028] The sound signal processing device 1 of the modified example 1 also performs pseudo-jitter processing to reduce the sense of localization for sound sources other than the specific sound source that you want to highlight. Therefore, users can obtain a new customer experience in which they can clearly hear a specific sound source while maintaining the sound quality and volume balance as intended by the content creator.

[0029] (Modification 2) Figure 6 is a block diagram showing the configuration of the sound signal processing device 1 according to Modification 2, and Figure 7 is a block diagram showing the functional configuration of the processor 12 according to Modification 2. The audio I / F 17 of the sound signal processing device 1 according to Modification 2 connects a plurality of speakers 20C, 21FL, 22FR, 23SL, and 24SR as an example of audio equipment, and outputs sound signals to these speakers. Speakers 20C, 21FL, and 22FR correspond to front speakers installed in front of the listening position. More specifically, speaker 20C is the center channel speaker, speaker 21FL is the front left channel speaker, and speaker 22FR is the front right channel speaker. Speakers 23SL and 24SR correspond to rear speakers installed behind the listening position. More specifically, speaker 23SL is the left surround channel speaker, and speaker 24SR is the right surround channel speaker.

[0030] The content data related to Modification 2 corresponds to the channel-based method. Channel-based content data is multi-channel audio in which the sound signals of each sound source are mixed in advance and stored in the sound signals of multiple channels. In the channel-based method, there is no need to perform localization processing again in the sound signal processing device 1 on the playback side. The decoder 120 decodes the content data and extracts the multiple sound signals related to the multi-channel audio.

[0031] The sound source designation unit 123 in the modified example 2 designates one of the channels in the multi-channel audio as a specific sound source. For example, the sound source designation unit 123 accepts the designation of the center channel as a specific sound source. The center channel often contains sound sources of voices such as singing and conversation. Sound sources such as singing are often of particular interest to the user.

[0032] The sound signal processing device 1 of the modified example 2 also performs pseudo-jitter processing to reduce the sense of localization for sound sources other than the specific sound source that you want to highlight. Therefore, users can obtain a new customer experience in which they can clearly hear sound sources related to voices while maintaining the sound quality and volume balance as intended by the content creator.

[0033] Alternatively, the sound source designation unit 123 may accept designation of a sound source corresponding to the sound signal to be output to the front speakers as a specific sound source. For example, the sound source designation unit 123 may accept designation of the center channel, FL channel, and FR channel as specific sound sources. The front channels often contain the main sound source of the content.

[0034] In this case as well, users can enjoy a new customer experience where they can clearly hear the main sound sources of the content while maintaining the sound quality and volume balance as intended by the content creator.

[0035] (Modification 3) Figure 8 is a block diagram showing the functional configuration of the processor 12 according to Modification 3. The decoder 120 of the processor 12 according to Modification 3 decodes the content data and inputs video signals corresponding to the audio signals of multiple audio sources to the audio source designation unit 123. The audio source designation unit 123 accepts the designation of a specific audio source based on the video signal.

[0036] The sound source designation unit 123 extracts specific objects, such as singers and musicians, from the video signal. Extraction is performed using pattern matching or a machine-trained model such as a DNN (Deep Neural Network). The sound source designation unit 123 obtains positional information for the extracted specific objects within the video signal. In addition to the left-right and up-down positions within the video signal, the sound source designation unit 123 can also obtain distance information for the extracted specific objects based on their size. Then, the sound source designation unit 123 designates the object sound source corresponding to the obtained positional information as the specific sound source.

[0037] The sound signal processing device 1 according to Modification 3 also performs pseudo-jitter processing to reduce the sense of localization for sound sources other than the specific sound source that you want to highlight. Therefore, users can obtain a new customer experience in which they can clearly hear a specific sound source while maintaining the sound quality and volume balance as intended by the content creator.

[0038] (Modification 4) In Modification 4, the sound source designation unit 123 receives camera control information corresponding to the video signal. The camera information includes left-right position information (Pan), up-down position information (Tilt), and distance information (Zoom). Based on the camera control information, the sound source designation unit 123 designates the corresponding object sound source as a specific sound source.

[0039] The sound signal processing device 1 according to Modification 4 also performs pseudo-jitter processing to reduce the sense of localization for sound sources other than the specific sound source that you want to highlight. Therefore, users can obtain a new customer experience in which they can clearly hear a specific sound source while maintaining the sound quality and volume balance as intended by the content creator.

[0040] (Modification 5) Figure 9 is a block diagram showing the functional configuration of the processor 12 according to Modification 5. The decoder 120 of the processor 12 according to Modification 5 decodes the content data, obtains specification information that specifies a particular sound source, and inputs it to the sound source specification unit 123. The sound source specification unit 123 accepts the specification of a specific sound source based on the acquired specification information.

[0041] In this case, the content creator includes the specified information in the content data. The sound signal processing device 1 according to Modification 5 performs pseudo-jitter processing to reduce the sense of localization for sound sources other than the specific sound source that the content creator wants to highlight. In this case as well, the user can obtain a new customer experience in which they can clearly hear the specific sound source while maintaining the sound quality and volume balance as intended by the content creator.

[0042] (Modification Example 6) FIG. 10 is a block diagram showing the functional configuration of the processor 12 according to Modification Example 6. The sound source designating unit 123 acquires the orientation information of the listener's head and accepts the designation of a specific sound source based on the acquired orientation information. The orientation information of the listener's head is acquired, for example, from a head tracker built into the headphones.

[0043] The sound source designating unit 123 designates an object sound source corresponding to the direction in which the user's head is facing as the specific sound source. The sound signal processing apparatus 1 according to Modification Example 6 also performs a pseudo jitter process for reducing the sense of localization on sound sources other than the specific sound source to be emphasized. Therefore, the user can obtain a new customer experience of clearly hearing a specific sound source being focused on while maintaining the sound quality and volume balance as intended by the content creator.

[0044] (Modification Example 7) FIG. 11 is a block diagram showing the functional configuration of the processor 12 according to Modification Example 7. The sound source designating unit 123 acquires the gaze information of the listener and accepts the designation of a specific sound source based on the acquired gaze information. The gaze information of the listener can be obtained, for example, based on a video signal acquired by a camera that captures the listener.

[0045] The sound source designating unit 123 designates an object sound source corresponding to the direction in which the user's gaze is directed as the specific sound source. The sound signal processing apparatus 1 according to Modification Example 7 also performs a pseudo jitter process for reducing the sense of localization on sound sources other than the specific sound source to be emphasized. Therefore, the user can obtain a new customer experience of clearly hearing a specific sound source being focused on while maintaining the sound quality and volume balance as intended by the content creator.

[0046] (Modification Example 8) FIG. 12 is a block diagram showing the functional configuration of the processor 12 according to Modification Example 8. The sound source designating unit 123 according to Modification Example 8 inputs the sound signals of a plurality of sound sources decoded by the decoder 120. The sound source designating unit 123 designates a specific sound source based on the sound signals of the plurality of sound sources.

[0047] For example, the sound source specifying unit 123 identifies the type of content (genre, artist name, song title, etc.) based on the sound signals of a plurality of sound sources. The sound source specifying unit 123 designates a specific sound source according to the identified content type. For example, when the sound source specifying unit 123 identifies the genre of jazz, it may designate the saxophone sound source as the specific sound source. Alternatively, when the sound source specifying unit 123 identifies the artist name, it may designate the sound source corresponding to the artist as the specific sound source. Further, when the sound source specifying unit 123 identifies the song title, it may acquire content data and change the designation of the specific sound source according to the progress of the song.

[0048] Also, the sound source specifying unit 123 prepares a trained model obtained by training the relationship between the sound signals of a plurality of sound sources and a specific sound source using a DNN or the like, inputs the sound signals of the plurality of sound sources to the trained model, and may accept the designation of the specific sound source.

[0049] For example, as shown in Modification 5, the content creator includes designation information in the content data. The sound source specifying unit 123 trains the relationship between the sound signals of a plurality of sound sources and the designation information with respect to a predetermined model to create a trained model. The sound source specifying unit 123 inputs the sound signals of the plurality of sound sources to the trained model and accepts the designation of the corresponding specific sound source.

[0050] Thereby, even when the designation information is not included in the content data, the sound source specifying unit 123 can designate an appropriate sound source as the specific sound source.

[0051] (Other examples) The sound source specifying unit 123 inputs information regarding the preferences of the content creator or listener, prepares a trained model obtained by training the relationship between the information regarding the preferences and a specific sound source using a DNN or the like, inputs the information regarding the preferences to the trained model, and may accept the designation of the specific sound source.

[0052] Thereby, the user can obtain a new customer experience of clearly hearing a specific sound source that matches their preferences by simply inputting, for example, the name of a specific favorite artist.

[0053] Furthermore, the sound signal processing device 1 may perform enhancement processing on the sound signal of a specific sound source by adjusting the amplitude or using an equalizer.

[0054] (Second Embodiment) The sound signal processing device 1 of the second embodiment has the same configuration as the sound signal processing device 1 shown in Figure 1. In facilities such as restaurants and cafes, background music (BGM) is sometimes output from speakers placed on the ceiling or walls in order to provide customers with a more comfortable environment. On the other hand, it is preferable that customers do not become too aware of the BGM as they are enjoying eating, drinking, and conversation. However, there are differences in sound quality due to differences in volume and reverberation between people who are close to the speakers and those who are far away. As a result, customers who are close to the speakers have the problem of their attention being drawn to the speakers. One possible solution is to increase the number of speakers in the facility to make the sound output from a particular speaker less noticeable. In contrast, the sound signal processing device 1 of the second embodiment aims to make the sound output from a speaker less noticeable without increasing the number of speakers.

[0055] The sound signal processing device 1 of the second embodiment takes input the target space and the speaker arrangement in the space, determines a first sound pressure distribution in the space based on the input space and speaker arrangement, determines the optimal parameters for blurring the first sound pressure distribution, and determines a second sound pressure distribution when blurring is applied using the determined optimal parameters. The blurring process includes the pseudo-jitter processing shown in the first embodiment. The first and second sound pressure distributions are displayed on the display 15.

[0056] Figure 13 shows an example of the application program's GUI (Graphical User Interface). The user inputs space 101 by, for example, moving the mouse cursor on the GUI to a starting position, clicking, and performing a drag operation. The user also inputs the target sound pressure distribution in the input space. In the example in Figure 13, the user inputs the average sound pressure in space 101 as the target sound pressure distribution. Note that the sound pressure referred to in this embodiment is not the instantaneous sound pressure, but the average sound pressure (effective sound pressure) over a certain period of time.

[0057] In this embodiment, an example of setting a two-dimensional space viewed from a planar perspective is shown, but a one-dimensional space along a certain direction may also be set, or a three-dimensional space including the height direction may be set. Furthermore, the user may set a space formed not only by straight lines but also by curves.

[0058] Note that the setting of the acoustic space using drawing with a mouse cursor as shown in Figure 13 is just one example, and the application program may accept the acoustic space through any interface. For example, the user may input the name of a real hall, and the application program may accept the shape of the acoustic space based on 3D CAD data corresponding to the input hall.

[0059] The sound signal processing device 1 calculates a speaker arrangement corresponding to the target sound pressure distribution received in the received acoustic space, based on a predetermined model. The speaker arrangement includes the number of speakers to be placed, the output characteristics of each speaker, and the position information of each speaker.

[0060] The sound signal processing device 1 determines the speaker placement using, for example, a mathematical model or a pre-trained model. The pre-trained model is a model trained using a DNN to determine the relationship between sound pressure distribution and speaker placement.

[0061] A computer that generates a trained model (for example, a server for a speaker manufacturer) acquires a large number of datasets showing the relationship between speaker placement and sound pressure distribution during the training phase. The server uses these acquired datasets to train a given model using a predetermined algorithm to understand the relationship between speaker placement and sound pressure distribution.

[0062] The sound pressure distribution for a given speaker arrangement in a given space is uniquely determined. In other words, there is a correlation between the speaker arrangement in a given space and the sound pressure distribution in that space. Therefore, the server can train a given model on the relationship between speaker arrangement and sound pressure distribution, and generate a trained model.

[0063] The sound signal processing device 1 obtains the trained model, which has been trained as described above, from the server. In the execution phase, the sound signal processing device 1 receives the received space and the target sound pressure distribution in that space as input using the trained model, and determines the corresponding speaker placement.

[0064] The sound signal processing device 1 receives input information about the space and the speaker placement within that space, calculates the sound pressure at each position within the space, and obtains a first sound pressure distribution. In the example above, the user inputs the space and the target sound pressure distribution, but the user may manually input the space and the speaker placement within that space according to the actual environment of the facility.

[0065] The sound signal processing device 1 determines the optimal parameters for blurring the obtained first sound pressure distribution, calculates the sound pressure at each position when blurring is applied with the determined optimal parameters, and obtains the second sound pressure distribution.

[0066] As described above, the blurring process includes the pseudo-jitter process shown in the first embodiment. The pseudo-jitter process is a process that changes the phase. The parameters of the pseudo-jitter process include the phase rotation speed. The sound signal processing device 1 extracts the maximum and minimum values ​​of sound pressure in the first sound pressure distribution. The sound signal processing device 1 finds the optimal parameters so that the difference between the maximum and minimum values ​​of sound pressure is minimized. More specifically, the sound signal processing device 1 generates a number of second sound pressure distributions using predetermined random numbers, with the phase rotation speed of each speaker as a variable. The sound signal processing device 1 then calculates an average second sound pressure distribution by averaging the number of generated second sound pressure distributions over a predetermined time. The sound signal processing device 1 then extracts the best solution for the phase rotation speed of each speaker, using the variability of each second sound pressure distribution relative to the average second sound pressure distribution as an evaluation index. Furthermore, the sound signal processing device 1 uses the extracted best solution to further change each variable in smaller steps, repeatedly searching in the direction that minimizes the difference between the maximum and minimum sound pressure values, thereby finding the optimal parameters (optimal phase rotation speed) for each speaker. Alternatively, the sound signal processing device 1 may find the optimal parameters using a trained model that has been trained to output the optimal phase rotation speed for each speaker.

[0067] In this way, the sound signal processing device 1 of the second embodiment can make the sound output from the speakers less noticeable without increasing the number of speakers. The sound signal processing device 1 of the second embodiment does not need to accept the designation of a specific sound source. That is, the sound processing device 1 of the second embodiment only needs to apply blurring processing to the input sound signal. In this case, when the sound signal processing device 1 outputs background music from speakers placed on the ceiling or walls, it can prevent customers who are close to the speakers from becoming distracted by the speakers.

[0068] Furthermore, the blurring process is not limited to pseudo-jitter processing. The blurring process may also be sound pressure adjustment processing. In this case as well, the sound signal processing device 1 can make the sound output from the speakers less noticeable without increasing the number of speakers.

[0069] Figure 14 shows a GUI example of the display of the first sound pressure distribution. Figure 15 shows a GUI example of the display of the second sound pressure distribution. When the user selects the "blur" icon in the GUI of Figure 14, the sound signal processing device 1 displays the second sound pressure distribution.

[0070] Furthermore, the sound signal processing device 1 may determine the first and second sound pressure distributions by considering not only the direct sound from each speaker to each position, but also the reflected sound from the walls. In this case, the sound signal processing device 1 determines the sound pressure at each position in space by, for example, using the ray method to obtain a transfer function (impulse response) that represents the delay time and level change of the sound from each speaker, reflected from the walls, to each position. By considering not only the direct sound but also the reflected sound, the sound signal processing device 1 can obtain more accurate optimal parameters.

[0071] Furthermore, the sound signal processing device 1 may determine a third sound pressure distribution when blurring is applied in a space wider than the input space. Figure 16 is a GUI showing an example of the display of the third sound pressure distribution. When the user selects the "extension" icon in the GUI of Figure 14, the sound signal processing device 1 displays the third sound pressure distribution. That is, the sound signal processing device 1 displays the first sound pressure distribution, the second sound pressure distribution, or the third sound pressure distribution on the display 15. This allows the user to visually recognize the effects of blurring and optimization. When the input space is smaller than the space where the speaker is actually installed, the perceived volume may change abruptly inside and outside the input space. In response to this, the sound signal processing device 1 can suppress the abrupt change in perceived volume inside and outside the input space by determining a third sound pressure distribution when blurring is applied in a space wider than the input space.

[0072] Furthermore, the sound signal processing device 1 may determine a first sound pressure distribution, a second sound pressure distribution, and a third sound pressure distribution for a frequency band of several kHz that is perceived as noticeable by the ear. This allows the sound signal processing device 1 to reduce the computational load and make the sound output from the speakers less noticeable without increasing the number of speakers.

[0073] The sound signal processing device 1 may, by simulation, determine the speaker arrangement that optimizes the blurring effect, using the number of speakers as a constraint. For example, the sound signal processing device 1 generates a large number of second sound pressure distributions using predetermined random numbers, with the speaker arrangement as a variable. The sound signal processing device 1 then calculates an average second sound pressure distribution by further averaging the generated large number of second sound pressure distributions over time. The sound signal processing device 1 then extracts the best solution for the speaker arrangement using the variability of each second sound pressure distribution relative to the average second sound pressure distribution as an evaluation index. Alternatively, the sound signal processing device 1 may determine the optimal parameters using a trained model that has been trained to output the optimal speaker arrangement. This allows the user to know the optimal speaker arrangement that is less noticeable than the current speaker arrangement.

[0074] The concept of the second embodiment can be summarized as follows:

[0075] (1) The target space and the speaker arrangement in the space are input, a first sound pressure distribution in the space is determined based on the input space and speaker arrangement, and a second sound pressure distribution is determined when a blurring process including pseudo-jitter processing or speaker sound pressure adjustment processing is applied to the first sound pressure distribution.

[0076] (2) Input the target space and the speaker arrangement in the space, determine a first sound pressure distribution in the space based on the input space and speaker arrangement, determine the optimal parameters for blurring processing including pseudo-jitter processing or speaker sound pressure adjustment processing for the first sound pressure distribution, and determine a second sound pressure distribution when the blurring processing is applied with the determined optimal parameters.

[0077] (3) Input the target space and the speaker arrangement in the space; determine a first sound pressure distribution in the space based on the input space and speaker arrangement; determine the optimal parameters for blurring processing, including pseudo-jitter processing or speaker sound pressure adjustment processing, for the first sound pressure distribution; determine a second sound pressure distribution when the blurring processing is applied with the determined optimal parameters; and apply the blurring processing corresponding to the second sound pressure distribution to the sound signal.

[0078] (4) In any of (1) to (3) above, the optimal parameters are determined such that the difference between the maximum and minimum values ​​of sound pressure in the first sound pressure distribution is minimized.

[0079] (5) In any of (1) to (4) above, the first sound pressure distribution and the second sound pressure distribution are determined in the space, taking into account the sound reflected from the walls.

[0080] (6) In any of (1) to (5) above, determine a speaker arrangement that can optimize the effect of the blurring process, with the number of speakers as a constraint.

[0081] (7) In any of (1) to (6) above, a third sound pressure distribution is determined in a space wider than the input space.

[0082] (8) In any of (1) to (7) above, input the target sound pressure distribution in the space, determine the speaker arrangement corresponding to the input space and the target sound pressure distribution, and determine the first sound pressure distribution, the second sound pressure distribution, and the third sound pressure distribution based on the determined speaker arrangement and the space.

[0083] The first embodiment, the second embodiment, and each of their modifications can be combined as appropriate. For example, the sound signal processing device 1 may apply blurring to sound signals other than the specified specific sound source using optimized parameters corresponding to the second sound pressure distribution.

[0084] The description of this embodiment should be considered in all respects to be illustrative and not restrictive. The scope of the invention is indicated by the claims, rather than by the embodiments described above. Furthermore, the scope of the invention includes the scope equivalent to the claims.

[0085] 1: Sound signal processing unit, 11: Communication unit, 12: Processor, 13: RAM, 14: Flash memory, 15: Display unit, 16: User I / F, 17: Audio I / F, 20, 20C, 21, 21FL, 22, 22FR, 23, 23SL, 24SR: Speaker, 40: Headphones, 120: Decoder, 121: Pseudo-jitter processing unit, 122: Localization processing unit, 123: Sound source selection unit

Claims

1. An audio signal processing method that inputs audio signals from multiple sound sources, accepts the designation of a specific sound source among the multiple sound sources, and applies pseudo-jitter processing to the audio signals other than the designated specific sound source, thereby adding fluctuations in the time axis direction.

2. The sound signal processing method according to claim 1, wherein the plurality of sound sources include an object sound source having position information, the method accepts the designation of the object sound source as the specific sound source, and performs localization processing on the object sound source based on the position information.

3. The sound signal processing method according to claim 2, wherein the localization processing includes binaural processing.

4. The sound signal processing method according to claim 1, wherein the sound signals of the plurality of sound sources are multi-channel audio, and the method accepts the designation of a center channel as the specific sound source.

5. The sound signal processing method according to claim 1, wherein the sound signals of the plurality of sound sources are output to a plurality of speakers, the plurality of speakers comprising front speakers installed in front of the listening position and rear speakers installed behind the listening position, and the method accepts the designation of a sound source corresponding to a sound signal to be output to the front speakers as a specific sound source.

6. The sound signal processing method according to any one of claims 1 to 5, comprising inputting video signals corresponding to sound signals from the plurality of sound sources and accepting the designation of a specific sound source based on the video signals.

7. The sound signal processing method according to any one of claims 1 to 5, comprising inputting video signals corresponding to sound signals from the plurality of sound sources, inputting camera control information corresponding to the video signals, and accepting the designation of a specific sound source based on the camera control information.

8. The sound signal processing method according to any one of claims 1 to 5, comprising acquiring designation information that specifies a particular sound source, and accepting the designation of the particular sound source based on the acquired designation information.

9. The sound signal processing method according to any one of claims 1 to 5, comprising acquiring information on the orientation of the listener's head and accepting the designation of a specific sound source based on the acquired orientation information.

10. A sound signal processing method according to any one of claims 1 to 5, comprising acquiring gaze information of a listener and accepting the designation of a specific sound source based on the acquired gaze information.

11. The sound signal processing method according to any one of claims 1 to 5, comprising: preparing a trained model that has been trained on the relationship between the sound signals of the plurality of sound sources and the specific sound source; inputting the sound signals of the plurality of sound sources into the trained model and accepting the designation of the specific sound source.

12. A sound signal processing method according to any one of claims 1 to 5, comprising: inputting a target space and speaker arrangement in the space; determining a first sound pressure distribution in the space based on the input space and speaker arrangement; determining optimal parameters for the pseudo-jitter processing for the first sound pressure distribution; and determining a second sound pressure distribution when the jitter processing is applied using the optimal parameters.

13. The sound signal processing method according to claim 12, wherein the optimal parameters are determined such that the difference between the maximum and minimum values ​​of sound pressure in the first sound pressure distribution is minimized.

14. The sound signal processing method according to claim 12, wherein the first sound pressure distribution and the second sound pressure distribution are determined in the space, taking into account the sound reflected from the walls.

15. The sound signal processing method according to claim 12, wherein the first sound pressure distribution and the second sound pressure distribution are displayed on a display device.

16. The sound signal processing method according to claim 12, which determines a third sound pressure distribution when the jitter processing is performed in a space wider than the input space.

17. An audio signal processing device equipped with a processor that inputs audio signals from multiple sound sources, accepts the designation of a specific sound source among the multiple sound sources, and applies pseudo-jitter processing to the audio signals other than the designated specific sound source, thereby adding fluctuations in the time axis direction.