Sound processing method, sound processing device, and program
Patent Information
- Application Number
- JP2022107675
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-07-04
- Publication Date
- 2025-09-29
AI Technical Summary
Existing technologies fail to adjust the volume balance of multiple performers in a virtual space effectively.
A sound processing method that arranges performer objects and volume adjustment interfaces in a virtual space, receives user input for volume adjustments, and uses a trained model to determine and apply volume parameters based on the relationship between sound signals and volume adjustments.
Achieves appropriate volume balance for multiple performers in a virtual environment, allowing viewers to enjoy performances with optimal sound quality without manual adjustments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] An embodiment of the present invention relates to a sound processing method, a sound processing device, and a program. [Background technology]
[0002] Patent document 1 discloses an audio mixer that receives sound signals related to a performance from audio mixers B2, C3 via a network, measures the communication delay time between the audio mixers B2, C3, and mixes the sound signals of the audio mixers B2, C3 according to the measured communication delay time.
[0003] Patent Document 2 discloses a configuration for smoothing switching between pre-fader and post-fader by correcting the volume difference between pre-fader and post-fader.
[0004] Patent Document 3 discloses a configuration for measuring an impulse response from a speaker to a microphone and adjusting the volume taking into account indirect sound components, thereby adjusting the volume to an appropriate level taking into account the indirect sound components.
[0005] Patent Document 4 discloses a configuration for feeding back the result of measuring the volume of the direct sound on the near-end side and the distance between the speaker and the microphone to the far-end side, thereby allowing the far-end user to know that his or her voice is being amplified correctly.
[0006] Patent Document 5 discloses a configuration for adjusting the volume of a sound signal received from the far end based on an acoustic feature acquired by a microphone on the near end. This allows the invention of Patent Document 5 to adjust the volume in consideration of the listening environment.
[0007] Patent Document 6 discloses a configuration in which multiple amplifiers are placed for multiple performers positioned at different locations on the stage, and a mixer adjusts the volume of the sound signals from each amplifier before supplying them to the performers. This allows the mixer in Patent Document 6 to collectively adjust the volume balance of multiple monitor speakers. [Prior art documents] [Patent documents]
[0008] [Patent Document 1] JP 2005-128296 A [Patent Document 2] International Publication No. 2018 / 21402 [Patent Document 3] Patent Publication No. 2021-129145 [Patent Document 4] JP 2010-103853 A [Patent Document 5] JP 2020-202448 A [Patent Document 6] JP 2009-100185 A Summary of the Invention [Problem to be solved by the invention]
[0009] None of the configurations disclosed in the prior art documents adjusts the volume balance of multiple performers in a virtual space.
[0010] An object of one embodiment of the present invention is to provide a sound processing method capable of appropriately adjusting the volume balance between multiple performers in a virtual space. [Means for solving the problem]
[0011] A sound processing method according to one embodiment of the present invention includes arranging a plurality of performer objects and a plurality of volume adjustment interfaces corresponding to the plurality of performer objects within a virtual space, receiving a plurality of sound signals corresponding to each of the plurality of performers, accepting from a user volume adjustment parameters for each of the plurality of performers corresponding to the plurality of volume adjustment interfaces, determining the volume adjustment parameters for each of the plurality of performers using a trained model that has been trained on the relationship between the sound signals corresponding to each performer of the plurality of sound signals and the volume adjustment parameters corresponding to the sound signals, and adjusting and mixing the volume of the plurality of sound signals based on the volume adjustment parameters determined by the trained model. Effect of the Invention
[0012] According to one embodiment of the present invention, the volume balance of multiple performers in a virtual space can be appropriately adjusted. [Brief description of the drawings]
[0013] [Figure 1] 1 is a block diagram showing a configuration of a sound processing system 1. FIG. [Diagram 2] FIG. 2 is a block diagram showing the configuration of a PC12C. [Diagram 3] FIG. 2 is a perspective view showing an example of a virtual three-dimensional space R1. [Figure 4] 13 is a flowchart showing the operation of the PC 12C and the server 30 in the training stage. [Diagram 5] 13 is a flowchart showing the operation of the PC 12C in the execution stage. [Figure 6] 13 is a flowchart showing the operation of the PC 12A (or the PC 12B) according to the first modification. [Figure 7] FIG. 11 is a perspective view showing an example of a virtual three-dimensional space R1 according to Modification 2. [Figure 8] FIG. 11 is a block diagram showing a configuration of a sound processing system 1A according to a fourth modified example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] Fig. 1 is a block diagram showing the configuration of a sound processing system 1. The sound processing system 1 in Fig. 1 includes a PC (personal computer) 12A installed in a first venue 3, a PC 12B installed in a second venue 5, a PC 12C installed in a third venue 7, and a server 30. PC 12A, PC 12B, PC 12C, and server 30 are connected via a network 9. PC 12A, PC 12B, and PC 12C are examples of the sound processing device of the present invention.
[0015] The PC 12A in the first venue 3 is connected to the guitar amplifier 11A and the motion sensor 13A. The guitar amplifier 11A is connected to the electric guitar 10.
[0016] The electric guitar 10 is an example of an audio device. The guitar amplifier 11A is connected to the electric guitar 10 via an audio cable. The guitar amplifier 11A is also an example of an audio device. The guitar amplifier 11A is also connected to the PC 12A via, for example, a USB cable. Of course, the guitar amplifier 11A may also be connected to the PC 12A via wireless communication. The electric guitar 10 outputs an analog sound signal related to the sound being played to the guitar amplifier 11A.
[0017] The guitar amplifier 11A has an analog audio terminal. The guitar amplifier 11A receives an analog sound signal from the electric guitar 10 via an audio cable. The guitar amplifier 11A converts the received analog sound signal into a digital sound signal. The guitar amplifier 11A performs various signal processing such as effects on the digital sound signal. The guitar amplifier 11A converts the processed digital sound signal into an analog sound signal. The guitar amplifier 11A amplifies the analog sound signal. The guitar amplifier 11A outputs the performance sound of the electric guitar 10 based on the amplified analog sound signal via a built-in speaker. The guitar amplifier 11A also transmits the processed digital sound signal to the PC 12A.
[0018] The user of PC 12A is the player of electric guitar 10. Using PC 12A, the player of electric guitar 10 distributes the sound of his / her performance and moves a 3D model object that acts as his / her alter ego and performs virtually in a virtual space. However, the user who distributes the sound of the performance and the player do not have to be the same person. PC 12A controls motion data for controlling the movement of the object.
[0019] The motion sensor 13A is a sensor for capturing the motion of the performer, and may be, for example, an optical, inertial, or image sensor. The motion sensor 13A is connected to the PC 12A by, for example, a USB cable. The PC 12A controls the motion data based on the sensor information received from the motion sensor 13A. Of course, the motion sensor 13A may be connected to the PC 12A by wireless communication.
[0020] The PC 12A transmits to the server 30 a digital sound signal relating to the guitar performance sound received from the guitar amplifier 11A and motion data controlled based on sensor information from the motion sensor 13A.
[0021] The PC 12B in the second venue 5 is connected to a microphone 19 and a motion sensor 13B.
[0022] The microphone 19 is an example of an audio device. The microphone 19 is connected to the PC 12B via an audio cable, a USB cable, or the like. The PC 12B receives an analog audio signal from the microphone 19 via the audio cable. The PC 12B converts the received analog audio signal into a digital audio signal. Alternatively, the microphone 19 may output a digital audio signal to the PC 12B via a USB cable, or the like.
[0023] The user of PC 12B is a singer. Using PC 12B, the singer distributes his / her singing voice and moves an object that represents the singer and sings virtually in a virtual space. PC 12B controls motion data for controlling the movement of the object. However, the user who distributes the singing voice and the singer do not need to be the same person.
[0024] The motion sensor 13B is a sensor for capturing the motion of a singer, and may be, for example, an optical, inertial, or image sensor. The motion sensor 13B is connected to the PC 12B by, for example, a USB cable. The PC 12B controls the motion data based on the sensor information received from the motion sensor 13B. Of course, the motion sensor 13B may be connected to the PC 12B by wireless communication.
[0025] The PC 12B transmits to the server 30 a digital sound signal relating to the singing sound received from the microphone 19 and motion data controlled based on sensor information from the motion sensor 13B.
[0026] The PC 12C in the third venue 7 is connected to headphones 20. The headphones 20 are also an example of audio equipment. The user of the PC 12C is a viewer who watches a performance of a performer that is virtually performed in a virtual space.
[0027] Fig. 2 is a block diagram showing the configuration of PC 12C. PC 12C is a general-purpose information processing device. Fig. 2 shows the configuration of PC 12C, but the main configurations of PC 12A and PC 12B are also the same as the configuration shown in Fig. 2.
[0028] The PC 12C includes a communication unit 11, a processor 12, a RAM 13, a flash memory 14, a display unit 15, a user I / F 16, and an audio I / F 17.
[0029] The communication unit 11 has a wireless communication function such as Bluetooth (registered trademark) or Wi-Fi (registered trademark), and a wired communication function such as USB or LAN.
[0030] The display 15 is composed of an LCD, an OLED, etc. The display 15 displays the image output by the processor 12.
[0031] The user I / F 16 is an example of an operation unit. The user I / F 16 is composed of a mouse, a keyboard, a touch panel, or the like. The user I / F 16 accepts operations by a user. The touch panel may be laminated on the display 15.
[0032] The audio I / F 17 is an interface for connecting audio equipment, including an analog audio terminal or a digital audio terminal, etc. In this embodiment, the audio I / F 17 of the PC 12C connects to headphones 20 as an example of audio equipment, and outputs an audio signal to the headphones 20.
[0033] The processor 12 is composed of a CPU, a DSP, a SoC (System on a Chip), or the like. The processor 12 reads out a program from a flash memory 14, which is a storage medium, and temporarily stores the program in the RAM 13 to perform various operations. Note that the program does not have to be stored in the flash memory 14. The processor 12 may download the program from another device, such as a server, when necessary, and temporarily store the program in the RAM 13.
[0034] The processor 12 receives sound signals and motion data from the server 30 via the communication unit 11. The sound signals received from the server 30 include a first sound signal related to the performance sound of the performer in the first venue 3 and a second sound signal related to the singing sound of the singer in the second venue 5. The motion data received from the server 30 includes the motion of the performer in the first venue 3 and the motion of the singer in the second venue 5. The processor 12 also receives spatial information, model data, position information, and the like from the server 30 via the communication unit 11.
[0035] The spatial information is information indicating the shape of a three-dimensional space corresponding to a live venue such as a live music venue or a concert hall, and is expressed in three-dimensional coordinates with a certain position as the origin. The spatial information may be coordinate information based on 3D CAD data of a live venue such as an actual concert hall, or may be logical coordinate information (information normalized between 0 and 1) of a certain imaginary live venue.
[0036] The model data is three-dimensional CG image data for constructing a 3D model object, and is made up of multiple image parts. Model data is specified for each performer. For example, the performer in the first venue 3 specifies model data that will be his or her alter ego. The server 30 distributes the specified model data.
[0037] The position information is information that indicates the position of the model data in a three-dimensional space. The position information is expressed in three-dimensional coordinates in the virtual space. The position information may be position information corresponding to model data that does not change position, such as a device such as a speaker, or may be position information corresponding to model data that changes position, such as a performer.
[0038] Fig. 3 is a perspective view showing an example of a certain virtual three-dimensional space R1. The virtual three-dimensional space R1 in Fig. 3 shows a space having a rectangular parallelepiped shape as an example, but the shape of the space may be any shape.
[0039] The processor 12 arranges objects in a virtual three-dimensional space R1 as shown in FIG. 3 based on the spatial information and position information received from the server 30. The processor 12 also sets the position of the user of the PC 12C in the virtual three-dimensional space R1. The position of the user of the PC 12C corresponds to a viewpoint position 50 in the virtual three-dimensional space R1. In FIG. 3, the virtual three-dimensional space R1 is shown from above, but the processor 12 renders model data based on the spatial information, model data, position information, motion data of the object, and information on the set viewpoint position to generate an image of the virtual three-dimensional space R1 viewed from the set viewpoint position 50. The generated image is displayed via the display device 15. This allows the viewer of the PC 12C to view an image of the virtual three-dimensional space R1 viewed from the set viewpoint position 50. The user of the PC 12C can change the viewpoint position 50 in the virtual three-dimensional space R1 via the user I / F 16. The processor 12 generates an image of the virtual three-dimensional space R1 viewed from the changed viewpoint position 50. This allows the user of PC 12C to perceive himself / herself as moving within virtual three-dimensional space R1.
[0040] The processor 12 of the PC 12C adjusts the volume of the multiple sound signals received from the server 30 and mixes them to generate, for example, a stereo (L, R) channel sound signal. In this example, the processor 12 mixes a first sound signal of the first venue 3 and a second sound signal of the second venue 5. The processor 12 outputs the stereo channel sound signal to the headphones 20 via the audio I / F 17.
[0041] The processor 12 may perform effect processing such as equalizer or reverb processing on each of the first sound signal and the second sound signal. The processor 12 may also perform localization processing on the first sound signal and the second sound signal so that the sounds are localized at the positions of the corresponding objects.
[0042] PC12C uses a trained model that has been trained to determine the relationship between the sound signals corresponding to each performer and the volume adjustment parameters corresponding to the sound signals, determines the volume adjustment parameters for each of the multiple performers, and adjusts and mixes the volume of the multiple sound signals.
[0043] 4 is a flowchart showing the operation of the PC 12C and the server 30 in the training stage. The server 30 distributes the first sound signal and the second sound signal (S21). The processor 12 of the PC 12C receives the first sound signal and the second sound signal from the server 30 (S11).
[0044] The processor 12 places a plurality of performer objects and a plurality of volume control interfaces corresponding to the plurality of performer objects in the virtual space (S12).
[0045] Specifically, as shown in Fig. 3, the processor 12 places a first object 51 corresponding to a performer 31 present in a first venue 3, which is a remote location. The processor 12 places a second object 52 corresponding to a singer 32 present in a second venue 5, which is another remote location, in the virtual three-dimensional space R1. Furthermore, the processor 12 places a first volume adjustment interface 71 corresponding to the first object 51 and a second volume adjustment interface 72 corresponding to the second object 52. Note that in this embodiment, the processor 12 places objects and volume adjustment interfaces corresponding to performers in two venues, the first venue 3 and the second venue 5, but the number of venues is not limited to two. The processor 12 may place objects and volume adjustment interfaces for performers in a greater number of venues.
[0046] Next, the processor 12 receives volume adjustment parameters for the multiple performers corresponding to the multiple volume adjustment interfaces from the user of the PC 12C (S13). The user of the PC 12C operates the first volume adjustment interface 71 and the second volume adjustment interface 72 arranged in the virtual three-dimensional space R1 as shown in FIG. 3 to adjust the volume. For example, when the user of the PC 12C feels that the performance sound corresponding to the first object 51 is too loud, the user of the PC 12C operates the first volume adjustment interface 71 to lower the volume. In the example of FIG. 3, the first volume adjustment interface 71 and the second volume adjustment interface 72 are slider controls. Therefore, when the user of the PC 12C feels that the performance sound corresponding to the first object 51 is too loud, the user of the PC 12C moves the first volume adjustment interface 71 downward. Also, when the user of the PC 12C feels that the singing sound corresponding to the second object 52 is too quiet, the user of the PC 12C moves the second volume adjustment interface 72 upward to increase the volume.
[0047] The PC 12C transmits the received volume adjustment parameters to the server 30 (S14). The server 30 receives the volume adjustment parameters from the PC 12C (S22). In this example, the server 30 receives the volume adjustment parameters from the PC 12C, but also receives volume adjustment parameters from many other information processing devices. The server 30 uses the many received volume adjustment parameters to train a predetermined model on the relationship between the sound signals corresponding to the multiple performers distributed using a predetermined algorithm and the volume adjustment parameters (S23).
[0048] In this embodiment, the algorithm for training the model is not limited, and any machine training algorithm such as CNN (Convolutional Neural Network) or RNN (Recurrent Neural Network) can be used. The machine training algorithm may be supervised training, unsupervised training, semi-supervised training, reinforcement training, inverse reinforcement training, active training, transfer training, or the like. In addition, the server 30 may train the model using a machine training model such as HMM (Hidden Markov Model) or SVM (Support Vector Machine).
[0049] For example, if many viewers feel that the performance sound of a particular performer (e.g., performer 31 in first venue 3) is too loud, the majority of viewers will turn down the volume. In this case, the predetermined model is trained to output a volume adjustment parameter that turns down the volume for the sound signal of the performance sound in first venue 3. In this way, the sound signals of the multiple performers and the respective volume adjustment parameters have a correlation. Therefore, the server 30 can train the predetermined model on the relationship between the sound signals of the multiple performers and the respective volume adjustment parameters, and generate a trained model.
[0050] 5 is a flowchart showing the operation of the PC 12C in the execution stage. The processor 12 of the PC 12C receives the first sound signal, the second sound signal, and the trained model from the server 30 (S31). Note that the trained model may be received in advance separately from the first sound signal and the second sound signal.
[0051] The processor 12 places a plurality of performer objects in the virtual space (S32). Specifically, the processor 12 places a first object 51 and a second object 52 in the virtual three-dimensional space R1. In this example, the processor 12 does not place the first volume control interface 71 and the second volume control interface 72 in the execution stage.
[0052] The processor 12 uses the trained model to determine the volume adjustment parameters for each of the multiple performers (S33). As described above, the trained model is trained on the relationship between the sound signals of each of the multiple performers and the volume adjustment parameters. Therefore, the processor 12 uses the trained model to determine the first volume adjustment parameter and the second volume adjustment parameter corresponding to the first sound signal corresponding to the first object 51 and the second sound signal corresponding to the second object 52, respectively.
[0053] The processor 12 adjusts and mixes the volumes of the multiple sound signals based on the volume adjustment parameters obtained from the trained model (S34). Specifically, the processor 12 adjusts the volume of the first sound signal with the first volume adjustment parameter, adjusts the volume of the second sound signal with the second volume adjustment parameter, and mixes the first sound signal and the second sound signal after the volume adjustments.
[0054] In this way, the PC12C can appropriately adjust the volume balance of multiple performers virtually singing or playing in the virtual three-dimensional space R1 by adjusting and mixing the sound signals of multiple performers to an appropriate volume balance using a trained model trained with volume adjustment parameters received from multiple users. This allows viewers watching the performance of the performers virtually in the virtual three-dimensional space R1 to enjoy the customer experience of being able to easily watch the virtual performance in the virtual space with better volume balance without having to perform volume balance adjustment operations.
[0055] In this example, the processor 12 does not arrange the first volume adjustment interface 71 and the second volume adjustment interface 72 in the execution stage, but the first volume adjustment interface 71 and the second volume adjustment interface 72 may be arranged. In this case, the user of the PC 12C can further fine-tune the first volume adjustment parameter and the second volume adjustment parameter obtained by the processor 12 using the trained model. The PC 12C may also transmit the fine-tuned volume adjustment parameter to the server 30. The server 30 may also receive the fine-tuned volume adjustment parameter and further retrain the trained model. This updates the volume adjustment parameter as the performance progresses in the virtual three-dimensional space R1. Therefore, the viewer watching the performance can enjoy a customer experience in which the viewer can always watch with an appropriate volume balance in accordance with the progress of the performance in the virtual three-dimensional space R1.
[0056] (Variation 1) FIG. 6 is a flowchart showing the operation of the PC 12A (or the PC 12B) according to the first modification.
[0057] In the above embodiment, an example was shown in which the PC 12C receives a trained model from the server 30, adjusts the volume of the first sound signal with the first volume adjustment parameter, adjusts the volume of the second sound signal with the second volume adjustment parameter, and mixes the first sound signal and the second sound signal after the volume adjustment. In other words, the volume adjustment parameter obtained by the trained model is a volume adjustment parameter used by the receiving device that mixes the received multiple sound signals.
[0058] In the sound processing system 1 of variant example 1, the transmitting devices PC12A and PC12B each receive a trained model, adjust the volume of the first sound signal using a first volume adjustment parameter, and adjust the volume of the second sound signal using a second volume adjustment parameter.
[0059] Specifically, the PC 12A first receives a trained model from the server 30 (S41). The PC 12A uses the received trained model to determine a volume adjustment parameter for a sound signal to be transmitted (S42). As described above, the trained model is trained on the relationship between each sound signal of a plurality of performers and each volume adjustment parameter. Therefore, the PC 12A can use the trained model to determine a first volume adjustment parameter corresponding to the first sound signal. The PC 12A adjusts the volume of the first sound signal based on the first volume adjustment parameter determined by the trained model (S43). The PC 12A transmits the adjusted first sound signal to the server 30 (S44). Similarly, the PC 12B adjusts the volume of the second sound signal with the second volume adjustment parameter based on the trained model.
[0060] In other words, in variant 1, the volume adjustment parameters obtained by the trained model are volume adjustment parameters used by multiple devices each used by multiple performers, and each of the multiple devices adjusts the volume of the multiple sound signals based on the volume adjustment parameters, and the receiving device receives and mixes the multiple sound signals after the volume has been adjusted by the multiple devices.
[0061] As a result, even in the sound processing system 1 of variant example 1, viewers watching a performer's performance taking place virtually in the virtual three-dimensional space R1 can enjoy the customer experience of being able to easily watch the virtual performance in the virtual space with better volume balance without having to perform any volume balance adjustment operations.
[0062] In the first modification, an example has been shown in which PC 12A (or PC 12B) adjusts the volume of the sound signal, but for example, guitar amplifier 11A may adjust the volume of the sound signal based on a trained model, or electric guitar 10 may adjust the volume of the sound signal based on a trained model. Alternatively, PC 12A may determine a volume adjustment parameter for guitar amplifier 11A based on a trained model, input the volume adjustment parameter to guitar amplifier 11A, and guitar amplifier 11A may adjust the volume of the sound signal. Alternatively, PC 12A may determine a volume adjustment parameter for electric guitar 10 based on a trained model, input the volume adjustment parameter to electric guitar 10, and electric guitar 10 may adjust the volume of the sound signal.
[0063] (Variation 2) The trained model may be trained not only on the volume adjustment parameters, but also on the relationship between the sound signals corresponding to each performer of the multiple sound signals and the effect parameters of the effect processing applied to the sound signals.
[0064] Fig. 7 is a perspective view showing an example of a virtual three-dimensional space R1 according to Modification 2. The same components as those in Fig. 3 are given the same reference numerals and description thereof will be omitted.
[0065] The processor 12 of the PC 12C places a plurality of performer objects and a plurality of effect adjustment interfaces corresponding to the plurality of performer objects in the virtual space as a training stage. Specifically, the processor 12 places a first effect adjustment interface 71A corresponding to the first object 51 and a second effect adjustment interface 72A corresponding to the second object 52 as shown in FIG.
[0066] In this example, the first effect adjustment interface 71A and the second effect adjustment interface 72A are controls for adjusting the effect parameters of an equalizer. The first effect adjustment interface 71A and the second effect adjustment interface 72A each include controls for adjusting the levels of the high, mid, and low ranges.
[0067] A user of the PC 12C operates the first effect adjustment interface 71A and the second effect adjustment interface 72A to adjust the effect parameters.
[0068] The PC 12C transmits the received effect parameters to the server 30. The server 30 receives the effect parameters from a number of information processing devices including the PC 12C. The server 30 uses the received many effect parameters to train a predetermined model on the relationship between the sound signals corresponding to a number of performers delivered using a predetermined algorithm and the effect parameters.
[0069] In the execution stage, the processor 12 of the PC 12C receives the first sound signal, the second sound signal, and the trained model from the server 30. The processor 12 uses the trained model to determine effect parameters to be applied to the sound signals of the multiple performers. The processor 12 applies effect processing to the multiple sound signals based on the effect parameters determined by the trained model. The processor 12 also adjusts and mixes the volume of the multiple sound signals after the effect processing.
[0070] In this way, the PC 12C of the second modification can appropriately adjust the sound quality of multiple performers virtually singing or playing in the virtual three-dimensional space R1 by applying appropriate effect processing to the sound signals of multiple performers using a trained model trained with effect parameters received from multiple users and mixing them. This allows viewers watching the performance of the performers virtually in the virtual three-dimensional space R1 to enjoy the customer experience of being able to easily watch the virtual performance in the virtual space with better sound quality without having to adjust the effect parameters.
[0071] The effect is not limited to the equalizer shown in the above example. The effect may be a compressor, a reverb, or other effects. For example, if the user of the PC 12C feels that the performance sound of the first venue 3 does not reverberate, the user adjusts the effect parameters so that strong reverb processing is applied to the performance sound of the first venue 3. The server 30 receives effect parameters from a number of information processing devices including the PC 12C, and generates a trained model that applies strong reverb processing to the performance sound of the first venue 3. As a result, the performance sound of the first venue 3 is automatically subjected to strong reverb processing, so that the viewer can enjoy the customer experience of easily viewing the virtual performance in the virtual space with better sound quality without having to adjust the effect parameters to apply strong reverb processing to the performance sound of the first venue 3 again.
[0072] The effect processing may be performed by the PC 12A, PC 12B, guitar amplifier 11A, electric guitar 10, microphone 19, etc. on the transmitting side, instead of the PC 12C on the receiving side. In this case, the PC 12A and PC 12B on the transmitting side receive the trained model from the server 30, obtain effect parameters based on the trained model, and perform the effect processing. The guitar amplifier 11A may obtain effect parameters based on the trained model and perform the effect processing, or the electric guitar 10 may obtain effect parameters based on the trained model and perform the effect processing. Alternatively, the PC 12A may obtain effect parameters for the effect processing in the guitar amplifier 11A based on the trained model, input the effect parameters to the guitar amplifier 11A, and perform the effect processing based on the effect parameters input by the guitar amplifier 11A. Alternatively, the PC 12A may obtain effect parameters for the effect processing in the electric guitar 10 based on the trained model, input the effect parameters to the electric guitar 10, and the electric guitar 10 may perform the effect processing.
[0073] (Variation 3) The sound processing system 1 of variant example 3 acquires information on multiple audio devices used by multiple performers, and adjusts effect parameters of the effect processing applied to multiple sound signals based on the acquired information on the multiple audio devices.
[0074] For example, PC 12A transmits information about electric guitar 10 and guitar amplifier 11A to server 30. Information about electric guitar 10 and guitar amplifier 11A includes, for example, information such as the model names or serial numbers of electric guitar 10 and guitar amplifier 11A. Similarly, PC 12B transmits information about microphone 19 to server 30, and PC 12C transmits information about headphones 20 to server 30.
[0075] The server 30 stores information about a plurality of audio devices and appropriate effect parameters such as equalizers corresponding to each audio device as a table. The server 30 reads effect parameters corresponding to the information about the audio devices received from the PC 12A, PC 12B, or PC 12C from the table, and transmits the read effect parameters to the PC 12A, PC 12B, or PC 12C.
[0076] PC 12A, PC 12B, or PC 12C receives effect parameters from server 30 and adjusts effect parameters of effect processing to be applied to the sound signal of the corresponding audio device. For example, PC 12C adjusts equalizer parameters of the sound signal to be output to headphones 20 based on the effect parameters received from server 30.
[0077] This allows users at each venue to enjoy the customer experience of being able to easily adjust effect parameters to appropriate ones without having to manually adjust effect parameters such as the equalizer of the audio equipment they use. For example, the sound quality may differ between a case where a performer distributes singing sounds using a certain audio equipment (e.g., a certain microphone) and a case where a performer distributes singing sounds using a different audio equipment (a different microphone). In this way, the sound quality of the distributed singing sounds may differ due to differences in recording environments caused by differences in audio equipment. However, the sound processing system 1 of the third modification can correct such differences in recording environments caused by differences in audio equipment.
[0078] The server may determine the effect parameters of the corresponding audio device using a trained model that has been trained on the relationship between information on multiple audio devices and appropriate effect parameters corresponding to each audio device.
[0079] (Variation 4) Fig. 8 is a block diagram showing a configuration of a sound processing system 1A according to Modification 4. The same components as those in Fig. 1 are given the same reference numerals, and the description thereof will be omitted. In the sound processing system 1A, a performer in a first venue 3 and a performer in a second venue 5 transmit sound signals related to playing sounds or singing sounds to each other, and perform a remote session.
[0080] PC 12A receives an audio signal related to the singing sound of the performer in second venue 5, adjusts the volume and outputs it to headphones 20A. The performer in first venue 3 listens to the singing sound of the performer in second venue 5 through headphones 20A. The performer in first venue 3 also uses PC 12A to adjust the volume of the singing sound of the performer in second venue 5 and perform in accordance with the singing sound.
[0081] PC 12B receives an audio signal related to the performance sound of the performer in first venue 3, adjusts the volume and outputs it to headphones 20B. The performer in second venue 5 listens to the performance sound of the performer in first venue 3 through headphones 20B. The performer in second venue 5 also uses PC 12B to adjust the volume of the performance sound of the performer in first venue 3 and perform in accordance with the performance sound.
[0082] The server 30 receives the volume adjustment parameters adjusted by the PC 12A and the PC 12B, and trains a predetermined model. A trained model that has been trained for the band can be generated. When the band members next hold a remote session, they receive the trained model from the server 30 using the information processing device and adjust the volume using the trained model.
[0083] This allows the performers on PC12A and PC12B to enjoy a customer experience in which they can conduct remote sessions at a better volume than previously adjusted, without having to adjust the volume.
[0084] The above sound processing system 1A is an example of performing a remote session in the first venue 3 and the second venue 5. However, the sound processing system 1A can also transmit and receive sound signals related to musical performance sounds or singing sounds in a larger number of venues, and perform a remote ensemble in which each performer plays in sync with the other performers.
[0085] (Variation 5) The PC 12C of the fifth modification adjusts and mixes the volumes of a plurality of received sound signals based on the first position information of a plurality of performer objects and the second position information of the audience.
[0086] For example, the processor 12 of the PC 12C adjusts the volume of the first sound signal and the second sound signal based on the distance between the user's viewpoint position 50 and a first object 51 and the distance between the viewpoint position 50 and a second object 52 in the virtual three-dimensional space R1. The processor 12 of the PC 12C increases the volume of the sound signal corresponding to an object close to the viewpoint position 50 and decreases the volume of the sound signal corresponding to an object far from the viewpoint position 50.
[0087] This allows viewers watching a performer's performance taking place virtually in the virtual three-dimensional space R1 to enjoy the customer experience of being able to watch while being aware of the distance between themselves and the performer in the virtual three-dimensional space R1.
[0088] (Variation 6) The PC 12C of the sixth modified example performs sound processing with the viewpoint position 50 as the listening point, based on the viewpoint position 50, the position of the first object 51, and the position of the second object 52. The sound processing with the viewpoint position 50 as the listening point is, for example, a localization process.
[0089] The processor 12 of the PC 12C performs, for example, a localization process such that the sounds of the first object 51 and the second object 52 are localized at the positions of the first object 51 and the second object 52 as viewed from the viewpoint position 50.
[0090] The processor 12 performs localization processing based on, for example, HRTF (Head Related Transfer Function). HRTF represents a transfer function from a certain virtual sound source position to the right ear and left ear of the user. For example, as shown in FIG. 3, the position of the first object 51 is on the left side in front of the viewpoint position 50. The processor 12 performs binaural processing to convolve the HRTF on the sound signal corresponding to the first object 51 so as to localize the object at a position on the left side in front of the user. This allows the user of the PC 12C to perceive as if he or she is at the viewpoint position 50 in the virtual three-dimensional space R1 and is listening to the sound of the first object 51 on the left side in front of him or her.
[0091] (Variation 7) In the above embodiment, in the training stage, the server 30 receives volume adjustment parameters from a large number of information processing devices and trains a predetermined model. However, the server 30 may receive volume adjustment parameters from one information processing device and train a predetermined model. For example, when a skilled operator performs an adjustment operation of the volume balance, the server 30 generates a trained model that trains the volume adjustment operation of the skilled operator and distributes the trained model. As a result, the trained model is shared by a large number of information processing devices. The other information processing devices adjust the volume using the distributed trained model.
[0092] This allows viewers watching the performers' virtual performance in the virtual three-dimensional space R1 to enjoy the customer experience of being able to watch the virtual performance more easily with better volume balance adjusted by a skilled operator, without having to adjust the volume balance.
[0093] Furthermore, when a remote session is performed by a member of a certain band, the server 30 may receive volume adjustment parameters adjusted by the member of the band and train a predetermined model. This allows the server 30 to generate a trained model for the band. When the band member next performs a remote session, the band member receives the trained model from the server 30 using an information processing device and adjusts the volume using the trained model.
[0094] This allows band members participating in remote sessions to enjoy a customer experience in which they can perform remote sessions with better volume balance that has been adjusted in the past, without having to adjust the volume balance.
[0095] The server 30 may retrain a trained model that has been trained with a volume adjustment parameter received from one information processing device. The server 30 may receive the volume adjustment parameter again from the one information processing device and retrain the trained model, or may receive the volume adjustment parameter from another information processing device and retrain the trained model.
[0096] (More examples) The user may input a volume adjustment parameter by voice input such as "increase the volume of the first performer" rather than by using a control such as a slider.
[0097] The operator of the sound processing system including the server 30 may provide an environment for performance in the virtual three-dimensional space R1 and sell the trained model. For example, the server 30 is configured to perform a charging process for a specific trained model and download the trained model after confirming the charging. The charging process may be performed by a separate charging server instead of the server 30. After charging a certain amount to a certain user, the server 30 allows the user to download a trained model trained by, for example, a volume adjustment operation of a skilled operator. In this case, the server 30 may perform a payment process to pay a reward to the skilled operator each time a trained model is downloaded. In this way, the operator of the sound processing system including the server 30 may give an incentive to the operator. This allows the operator to sell his or her own volume adjustment technology. Therefore, the operator can increase the motivation of many users to use the sound processing system.
[0098] The server 30 may store a plurality of trained models trained with the volume adjustment parameters of a plurality of operators. The server 30 downloads an arbitrary trained model designated by a viewer from among the plurality of trained models to an information processing device used by the viewer. In this case, the server 30 may perform a process of paying a reward to the operator who trained the downloaded trained model. This allows the operator to increase the motivation of many operators to use the sound processing system.
[0099] The billing process may be a monthly or yearly subscription instead of a per-download billing.
[0100] The description of the present embodiment should be considered as illustrative in all respects and not restrictive. The scope of the present invention is indicated by the claims, not by the above-described embodiments. Furthermore, the scope of the present invention includes the scope equivalent to the claims. [Explanation of symbols]
[0101] 1,1A: Sound processing system 3: Hall 1 5: Venue 2 7: Venue 3 9: Network 10: Electric guitar 11: Communications Department 11A: Guitar amplifier 12: Processor 13: RAM 13A, 13B: Motion sensor 14: Flash memory 15:Display unit 16: User I / F 17: Audio I / F 19:Mike 20, 20A, 20B: Headphones 30: Server 31: Performer 32: Singer 50: Viewpoint position 51: First object 52: Second object 71: First volume control interface 71A: 1st effect adjustment interface 72: Second volume control interface 72A: 2nd effect adjustment interface
Claims
1. In the virtual space, multiple performer objects are placed to perform virtual performances. receiving a plurality of sound signals corresponding to each of the plurality of performers; obtaining a plurality of volume adjustment parameters corresponding to each of the plurality of performers; adjusting the volumes of the plurality of sound signals corresponding to each of the plurality of performers based on the plurality of volume adjustment parameters corresponding to each of the plurality of performers; mixing the plurality of sound signals after adjusting the volume thereof and outputting the resultant sound signal as a sound signal relating to the virtual performance performed by the plurality of performer objects in the virtual space; Sound processing methods.
2. Using a trained model that has been trained to understand the relationship between the plurality of sound signals corresponding to each of the plurality of performers and the plurality of volume adjustment parameters corresponding to each of the plurality of sound signals, the plurality of volume adjustment parameters corresponding to each of the plurality of performers are obtained. The sound processing method according to claim 1 .
3. In the training stage of the trained model, placing the plurality of performer objects and a plurality of volume control interfaces corresponding to the plurality of performer objects in the virtual space; receiving the plurality of sound signals corresponding to each of the plurality of performers; receiving the plurality of volume adjustment parameters corresponding to the plurality of volume adjustment interfaces, based on a user's volume adjustment operation on each of the plurality of volume adjustment interfaces; generating the trained model using the plurality of sound signals and the plurality of volume adjustment parameters received from the user; The sound processing method according to claim 2 .
4. A method for receiving the plurality of volume adjustment parameters from a plurality of information processing devices used by a plurality of users, generating the trained model using the plurality of volume adjustment parameters received from each of the plurality of information processing devices; The sound processing method according to claim 3 .
5. The trained model is further trained to determine a relationship between the plurality of sound signals corresponding to each of the plurality of performers and effect parameters of effect processing to be applied to the sound signals, Using the trained model, determine a plurality of effect parameters corresponding to each of the plurality of performers; applying effect processing to the plurality of sound signals based on the plurality of effect parameters determined by the trained model; The sound processing method according to any one of claims 2 to 4.
6. receiving information on a plurality of audio devices used by each of the plurality of performers; determining a plurality of effect parameters corresponding to each of the plurality of performers based on the received information on the plurality of audio devices; The sound processing method according to claim 5 .
7. Acquire first position information of the objects of the plurality of performers and second position information indicating a viewpoint position in a virtual space of a viewer who views the objects of the plurality of performers; adjusting the volume or localization of the plurality of sound signals based on the first position information and the second position information; The sound processing method according to any one of claims 1 to 4.
8. The plurality of volume adjustment parameters are used in a receiving-side information processing device that receives the plurality of sound signals, the receiving-side information processing device adjusts the volumes of the received sound signals based on the acquired volume adjustment parameters. The sound processing method according to any one of claims 1 to 4.
9. the plurality of volume adjustment parameters are used in a plurality of transmitting-side information processing devices that are used by the plurality of performers and that transmit the plurality of sound signals, respectively; each of the plurality of transmitting-side information processing devices adjusts the volume of each of the plurality of sound signals based on each of the plurality of acquired volume adjustment parameters; a receiving-side information processing device that receives and mixes the plurality of sound signals whose volumes have been adjusted by the plurality of transmitting-side information processing devices; The sound processing method according to any one of claims 1 to 4.
10. receiving a plurality of sound signals corresponding to each of the plurality of performers from a plurality of information processing devices used by each of the plurality of performers via a network; The sound processing method according to any one of claims 1 to 4.
11. In the execution stage of the trained model, further disposing a plurality of volume control interfaces corresponding to the plurality of performer objects in the virtual space; receiving the plurality of volume adjustment parameters corresponding to the plurality of volume adjustment interfaces, based on a user's volume adjustment operation on each of the plurality of volume adjustment interfaces; retraining the trained model using the plurality of volume adjustment parameters received from the user; The sound processing method according to any one of claims 2 to 4.
12. receiving the plurality of volume adjustment parameters from a first information processing device of a first user; generating the trained model using the received plurality of volume adjustment parameters; Transmitting the generated trained model to a second information processing device of a second user different from the first user; the second information processing device uses the trained model to obtain a plurality of volume adjustment parameters corresponding to each of the plurality of performers; The sound processing method according to claim 2 or 3.
13. the server performs a billing process for the second user and a remuneration payment process for the first user; The sound processing method according to claim 12.
14. In the virtual space, multiple performer objects are placed to perform virtual performances. receiving a plurality of sound signals corresponding to each of the plurality of performers; obtaining a plurality of volume adjustment parameters corresponding to each of the plurality of performers; adjusting the volumes of the plurality of sound signals corresponding to each of the plurality of performers based on the plurality of volume adjustment parameters corresponding to each of the plurality of performers; mixing the plurality of sound signals after adjusting the volume thereof and outputting the resultant sound signal as a sound signal relating to the virtual performance performed by the plurality of performer objects in the virtual space; With a processor, Sound processing device.
15. In the virtual space, multiple performer objects are placed to perform virtual performances. receiving a plurality of sound signals corresponding to each of the plurality of performers; obtaining a plurality of volume adjustment parameters corresponding to each of the plurality of performers; adjusting the volumes of the plurality of sound signals corresponding to each of the plurality of performers based on the plurality of volume adjustment parameters corresponding to each of the plurality of performers; mixing the plurality of sound signals after adjusting the volume thereof and outputting the resultant sound signal as a sound signal relating to the virtual performance performed by the plurality of performer objects in the virtual space; A program that causes an information processing device to perform sound processing.