Method for simultaneously playing multiple performance audio sources, simultaneous playback system, and simultaneous playback program for multiple performance audio sources.
The method for simultaneous playback of multiple performance sound sources addresses the lack of immersion in existing technologies by using multi-track recording and spatial audio processing to simulate a sound field, enabling effective practice accompaniment and virtual sessions with a sense of presence.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FRAME LUNCH CO LTD
- Filing Date
- 2023-10-14
- Publication Date
- 2026-07-22
AI Technical Summary
Existing technologies fail to provide a method for simultaneously playing multiple performance sound sources that simulate a sense of presence or immersion according to the listening position, particularly in big band performances, limiting their use as effective practice accompaniment sound sources.
A method involving multi-track recording, setting the space size and listening position, calculating three-dimensional coordinates, and applying object-based spatial audio processing to simulate a sound field within a specific space, allowing for the simultaneous playback of multiple performance sound sources with a sense of presence or immersion.
Enables virtual sessions with a sense of presence or immersion, making the sound sources usable as effective accompaniment for practice, and allowing for the selection and manipulation of listening positions and space size, enhancing the musical experience.
Smart Images

Figure 0007893581000001 
Figure 0007893581000002 
Figure 0007893581000003
Abstract
Description
Technical Field
[0001] The present invention relates to a method for simultaneous playback of a plurality of performance sound sources, a simultaneous playback system, and a program for simultaneous playback of a plurality of performance sound sources. More specifically, when playing a plurality of performance sound sources, by simulating a sense of presence or immersion according to the performance listening position, the present invention provides an excellent performance sound source for appreciation and can be utilized as a more effective practice accompaniment performance sound source, and relates to a method for simultaneous playback of a plurality of performance sound sources, a simultaneous playback system, and a program for simultaneous playback of a plurality of performance sound sources.
Background Art
[0002] In many music genres such as jazz, rock, and funk, a so-called N-1 accompaniment app is used regardless of whether the user is a professional or an amateur. The N-1 accompaniment app, for example, when a user practices a saxophone solo for a certain song, plays the rhythm accompaniment of the piano, bass, and drums, and plays the accompaniment performance sound source of parts other than one's own instrument. As such an N-1 accompaniment app, the software name 'IREAL PRO' is known. This app plays the rhythm accompaniment of the piano, bass, and drums while displaying the progress on a so-called chord sheet displayed on the screen for each of a large number of songs in various music genres such as pop, jazz, and rock. For example, instrument players such as saxophones or vocalists can practice accordingly. In addition, selection of rhythm accompaniment, selection of rhythms such as swing, funk, and Latin, adjustment of tempo, transposition, and selection of the number of repetitions for only a certain part of a song or the whole song are possible. However, this app is specifically dedicated to providing an accompaniment performance sound source and does not provide a performance sound source for appreciation. As an accompaniment performance sound source, it can be listened to as stereo, but it is not possible to perform while feeling a sense of immersion and presence like in a jazz session where instrument players gather on the spot and perform live.
[0003] In this regard, recently, accompaniment apps for multiple performance styles, especially for big bands, have been developed, such as the software called "WDR BIG BAND," The app "BOB Minzer BIG BAND" is well known. More specifically, a big band, unlike the so-called small combo style, is composed of a rhythm section consisting of piano, drums, bass, guitar, and vocals; a trumpet section consisting of multiple trumpets; a trombone section consisting of multiple trombones; and a saxophone section consisting of multiple saxophones, and primarily plays jazz. This accompaniment app features recordings of numerous songs with piano, bass, and drum rhythm accompaniment, along with front-line performances by trumpet, trombone, and saxophone. By selecting songs, it can be used not only as a source for listening to performances, but also to some extent as an N-1 accompaniment app, providing accompaniment for practice. However, it can only be used for a limited number of songs and cannot be used as a so-called virtual session. More specifically, a session, for example in the case of a small combo, involves gathering together musicians for the rhythm section, front section, and vocals to perform a chosen song live simultaneously. However, a so-called virtual session is not a simultaneous live performance with musicians gathered on the spot, nor is it even a simultaneous live performance at all. Instead, it involves adding one's own performance to a recorded track, ultimately completing it as a song and virtually simulating a session. While such virtual sessions have become possible to some extent with the recent advancements in communication networks, they are not yet implemented in the aforementioned bigband applications.
[0004] On the other hand, in recent years, acoustic and sound field software processing technologies have advanced, and as disclosed in Patent Documents 1, 2, and 3, for example, by applying so-called reverb software processing to the performance sound source, it has become possible to simulate a realistic performance through three-dimensional sound reproduction. Patent Document 1 discloses a 3D sound effect addition device, a program characterized by being executed by a computer, and a performance sound source device. The 3D sound effect addition device comprises a storage means for storing musical data, and a 3D sound effect addition means for extracting information that changes periodically according to the progress of the musical piece from the musical data stored in the storage means, generating 3D sound effect information that changes the localization of the 3D sound in synchronization with the periodically changing information when 3D sound is generated by playing the musical data, and outputting musical data with the 3D sound effect information added to the musical data. The program characterized by being executed by a computer comprises an analysis process for analyzing musical data and extracting information that changes periodically according to the progress of the musical piece from the musical data, and a 3D sound effect addition process for generating 3D sound that changes the localization of the 3D sound in synchronization with the periodically changing information extracted in the analysis process when 3D sound is generated by playing the musical data, and outputting musical data with the 3D sound effect information added to the musical data. The configuration involves a computer executing the following: The performance sound source device is a performance sound source device that reproduces three-dimensional sound according to musical data, and comprises: an analysis means that analyzes the musical data and extracts information that changes periodically according to the progress of the musical piece from the musical data; and a control means that generates three-dimensional sound effect information that changes the localization of the three-dimensional sound in synchronization with the periodically changing information extracted by the analysis means when three-dimensional sound is generated by the reproduction of the musical data, and controls the reproduction of three-dimensional sound based on the musical data based on the three-dimensional sound effect information. With such a 3D sound effect enhancement device, a program characterized by being executed by a computer, and a performance sound source device, it becomes possible to perform 3D sound playback using music data even if the music data was not created with 3D sound playback in mind or if the music data is for 3D sound playback with minimal 3D sound effects.
[0005] Patent Document 2 discloses a three-dimensional sound reproduction device and a program for making a computer function as a three-dimensional sound reproduction device. This three-dimensional sound reproduction device comprises an analysis unit that separates a multi-channel sound signal into a sound image component and an enveloping component, a sound image component processing unit that binauralizes the sound image component, an enveloping component enhancement processing unit that emphasizes and binauralizes the enveloping component, and an adder that adds the binauralized sound image component and the enveloping component. The analysis unit performs principal component analysis on the multi-channel sound signal using a variance-covariance matrix and separates the multi-channel sound signal into the sound image component and the enveloping component based on the magnitude of the obtained eigenvalues. The enveloping component enhancement processing unit changes the number of channels in the enveloping component so that the interaural correlation coefficient is smaller than that of the signal with binauralized enveloping component using the number of channels in the multi-channel sound signal, and emphasizes the enveloping component by binauralizing the enveloping component with the changed number of channels. With such a three-dimensional sound reproduction device and program, when reproducing a multi-channel sound signal in binaural format, it becomes possible to emphasize and reproduce the encapsulated component of the multi-channel sound signal.
[0006] Patent Document 3 discloses a stereophonic sound reproduction device, which includes a first speaker whose directional axis is directed toward one ear of the listener, a second speaker whose directional axis is directed toward the other ear of the listener, and a signal processing unit that performs crosstalk cancellation processing to cancel out the sound reaching one ear from the second speaker and the sound reaching the other ear from the first speaker, and generates a processed signal, wherein the first speaker and the second speaker emit sound based on the processed signal, and the first speaker and the second speaker are arranged on the midline plane or the same sagittal plane of the listener. With this type of 3D sound reproduction device, by arranging two speakers on the listener's midline or on the same sagittal plane, it becomes possible to reproduce 3D sound with greater robustness to the listener's movement than conventional methods.
[0007] However, none of the three patent documents (Patent Documents 1, 2, and 3) have applied such acoustic and sound field processing technologies to the playback of multiple sound sources, such as big band performances, and have failed to provide sound sources that offer a sense of presence or immersion according to the performance location. More specifically, it is not possible to select any of multiple recorded performance sound sources, or to select a listening position, for example, to use them as accompaniment sound sources for N-1 or as sound sources for listening to music, while simulating the performance space size of multiple recorded performances or simulating listening at an arbitrary position among multiple recorded performance sound sources. Furthermore, regarding the selected songs, a virtual session has not yet been realized in which musicians for the rhythm section, front section, and other instruments, as well as vocalists, are recruited not only from within Japan but from all over the world, and instead of having them perform simultaneously, the performance is gradually added based on recorded backing tracks, and once the song is complete, a virtual session simulating a live performance with a sense of presence is created using 3D sound reproduction.
[0008] In this regard, for example, Patent Document 4 discloses an online session server device. This online session server device is connected to terminal devices used by multiple users in a communicative manner, and has the following configuration for each terminal device: a setting unit that sets the part of the performance to be performed; a generation unit that generates part-specific audio data corresponding to each set part by processing karaoke performance data of a certain song used for karaoke performance; a transmission unit that transmits minus-one audio data, which is part-specific audio data corresponding to a part different from the one set, to the terminal device to which one part has been set; an acquisition unit that acquires live performance data based on the sound obtained by actually performing one part in sync with the performance based on the minus-one audio data from the terminal device to which one part has been set; and a storage processing unit that stores the live performance data of a certain song acquired from each of the multiple terminal devices. Such an online session server device allows for online session performances using karaoke performance data.
[0009] Furthermore, for example, Patent Document 5 discloses a session adjustment device. In this session adjustment device, the arithmetic unit acquires the composition information of each of the performance sounds of multiple songs, as well as the composition information of a model sound. Based on the composition information, the arithmetic unit calculates the difference between the composition information of the performance sound and the composition information of the model sound. The overall optimization processing unit then performs averaging based on this difference. The synthesis unit synthesizes each of the multiple performance sounds based on the average value obtained by the averaging process. The synthesis unit functions as an adjustment unit by adjusting the volume, rhythm, etc., of the multiple performance sounds to match the average value. Such session coordination devices make it possible to some extent to adjust sessions while eliminating network latency. However, neither Patent Document 4 nor Patent Document 5 describes a virtual session that simulates a live performance with a sense of presence through 3D sound reproduction, even if online session performances, that is, simultaneous performances, are possible.
[0010] [Patent Document 1] Patent No. 4983012 [Patent Document 2] Patent No. 6463955 [Patent Document 3] Japanese Patent Publication No. 2023-121744 [Patent Document 4] Japanese Patent Publication No. 2022-114309 [Patent Document 5] Japanese Patent Publication No. 2021-196556 [Disclosure of the Invention] [Problems that the invention aims to solve]
[0011] Therefore, the object of the present invention is to provide a method for simultaneously playing multiple performance sound sources, a simultaneous playback system, and a simultaneous playback program for multiple performance sound sources that can be used as more effective practice accompaniment sound sources, by simulating a sense of presence or immersion according to the listening position when playing multiple performance sound sources. Therefore, the object of the present invention is to provide a method for simultaneously playing multiple performance sound sources, a simultaneous playback system, and a simultaneous playback program for multiple performance sound sources that enable so-called virtual sessions when playing a song using multiple performance sound sources, and simulate a sense of presence or immersion, thereby providing excellent performance sound sources for listening, and that can be used as more effective accompaniment sound sources for practice. [Means for solving the problem]
[0012] To achieve the above objectives, the present invention provides a method for simultaneously playing multiple performance sound sources. A method for simultaneously playing multiple (N) performance sound sources, The process involves recording each of the multiple performance audio sources individually and independently using multi-track recording, On the operation screen, there is a step to set the space size when playing multiple audio sources, On the operation screen, there is a step to specify the listening position when playing the performance audio, Based on the identified listening location, the process involves calculating the three-dimensional coordinates of the listening location, The N-1 stage involves calculating the distance and orientation between the position of each performance sound source and the listening position, On the operation screen, there is a step to select one of the N-1 performance sound sources, ... N-1. Using object-based spatial audio processing, three-dimensional positional information is added to the performance sound source information, and based on the selected space size and listening position, the sound playback processing of each performance sound source is performed according to the positional relationship between the position of each performance sound source and the listening position, thereby simulating and constructing a sound field within a specific space size. The system has a configuration that includes a stage for simultaneously playing one or more selected performance recordings that have undergone spatial sound processing.
[0013] According to the simultaneous playback method of multiple performance sound sources having the above configuration, each of the multiple performance sound sources is recorded individually and independently using multi-track recording, the space size of the multiple performance sound sources is set on the operation screen, the listening position of the performance sound source is identified on the operation screen, the three-dimensional coordinates of the listening position are calculated based on the identified listening position, the distance and orientation between the position of each N-1 performance recording and the listening position are calculated, and then three-dimensional position information is added to the performance sound source using object-based spatial sound processing. At the same time, sound playback processing is performed for each performance sound source based on the selected space size and listening position, thereby simulating a sound field within a specific space size. By selecting one of the N-1 performance sound sources 1, ... N-1 on the operation screen and simultaneously playing the selected single or multiple performance recordings that have undergone spatial sound processing, a sense of presence or immersion corresponding to the listening position is simulated when playing multiple performance sound sources, thereby providing excellent performance sound sources for listening and making them usable as more effective accompaniment sound sources for practice. In summary, when simultaneously playing a piece of music using multiple performance sound sources, any of the multiple performance sound sources is selected, and any one of the multiple performance sound sources that is not selected is specified as the listening position of the music. The space size of the music performance by the multiple performance sound sources is set, and for each selected performance sound source information among the multiple performance sound sources, position information is associated according to the position vector relationship between each performance sound source position and the listening position. Thus, without being restricted by recording conditions such as the microphone position and the number of microphones, and playback conditions such as the speaker position and the number of speakers, object-based stereophonic processing is performed using the head-related transfer function (HRTF) to localize the sound image, and within the selected space size, a sound field is simulated and constructed. Thereby, it is possible to use a piece of music rich in immersion or a sense of presence for appreciation or practice accompaniment.
[0014] Also, including the direction of sound emission of the performances of the multiple performance sound sources, stereophonic processing is preferably performed according to the selected space size and the listening position, reflecting the volume difference, time difference, change in frequency characteristics, change in phase, change in reverberation, and Doppler effect. Furthermore, when playing the multiple performance sound sources, it is preferable to perform stereophonic processing after setting synchronization between the selected performance sound sources. Moreover, the multiple performance sound sources are composed of a rhythm section including a piano, drums, bass, and guitar, a trumpet section composed of multiple trumpets, a trombone section composed of multiple trombones, and a saxophone section composed of multiple saxophones. The simultaneous playback stage of the performance recording performance sound sources may be the performance playback of a big band piece of music. In addition, by selecting a listening position from multiple (N) performance sound sources and arbitrarily selecting any one of the remaining performance sound sources from 1 to N - 1 as a performance sound source, for the simultaneous playback of the multiple performance sound sources, it may also be used as a practice accompaniment performance sound source for the performance at the listening position.
[0015] Alternatively, apart from the plurality (N) of performance sound sources, an audience position may be set, and by selecting the audience position, a performance music piece for appreciation by a plurality of full performance sound sources may be obtained by simultaneously playing the plurality of performance sound sources. Furthermore, among the plurality (N) of performance sound sources, any performance sound source of the rhythm section may be selected and played, and another performance sound source of the rhythm section may be selected and played among the plurality (N) of performance sound sources, and this may be used for rhythm sense acquisition by comparing the individual performance sound sources of each part in the rhythm section. Furthermore, when any of the plurality of performance sound sources and / or the listening position moves, and the performance sound source selected as a performance sound source among the plurality of performance sound sources moves during performance, a BPM value may be selected when playing the plurality of performance sound sources. In addition, in the plurality of recorded performance sound sources, based on the selected recorded performance sound source and the selected listening position, after the simultaneous playback of the plurality of performance sound sources is started, stereophonic processing may be performed on the selected recorded performance sound source.
[0016] Also, when simultaneously playing the plurality of performance sound sources, by setting the playback start position and the playback end position in the selected music piece and / or setting the number of repetitions, the performance from the set playback start position to the set playback end position may be repeated according to the set number of repetitions. Furthermore, when individually and independently recording each of the plurality of performance sound sources by multi-track, the recording may be performed by binaural processing.
[0017] Furthermore, when playing a master performance sound source by the method for simultaneously playing a plurality of performance sound sources according to claim 1, a stage of selecting a performance sound source to be muted from the plurality of performance sound sources; A stage of recording the performance sound source while listening to the master performance sound source as accompaniment with stereo headphones; A stage of playing a music piece by an additional participation type virtual session by using the master performance sound source and the recorded performance sound source by the method for simultaneously playing a plurality of performance sound sources according to claim 1 may be included. In addition, when participating in an additional virtual session, it is also acceptable to decide whether to respond to invitations from existing participants and to accept or reject participation requests from those who wish to participate. Furthermore, in the aforementioned additional participation virtual session, each existing participant may be designated either as a Master, who has the authority to set up the session and reject participants, or as a Member, who only has the authority to reject the sharing of their own recorded performances. Furthermore, when participating in the aforementioned additional virtual session, the recorded performance audio can be shared by transferring it to an external storage device. In addition, the Master may decide which of the multiple performance recordings will be used to create the collaborative track by combining it with the master recording.
[0018] Alternatively, the Master may decide whether to limit the sharing of the collaborative track to participants only, or to make the simultaneous playback method of multiple performance sound sources described in claim 1 publicly available to users. Furthermore, if a participant has recorded a performance audio while listening to the master performance audio as accompaniment using stereo headphones, the next participant may use the simultaneous playback method of multiple performance audio sources described in claim 1 to perform stereophonic sound processing on the master performance audio source and the recorded performance audio source and use them as an accompaniment audio source. Furthermore, the master recording could include the rhythm section's performance, including piano, drums, and bass, with the front section (guitar, saxophone, trumpet, and trombone) participating. In addition, the evaluation results may be presented based on the number of plays and / or "likes" on the released collaborative song. Furthermore, in the aforementioned additional participation virtual session, the tempo, key, and spatial sound processing conditions of the accompanying music used as the basis for participating in the virtual session may be set in common among all participants.
[0019] To achieve the above objectives, the simultaneous playback program for multiple performance sound sources of the present invention is: A music playback program for a music playback device having a storage means for storing music data consisting of multiple performance sound sources, a playback means for playing the music data, a sound image localization processing means for processing the music sound signal output from the playback means to localize a sound image as stereophonic sound at an arbitrary position in a three-dimensional sound space based on a control signal, and a localization control table which defines a localization method for localizing the content of the music and the corresponding sound image as stereophonic sound in a three-dimensional sound space, and the contents of the localization control table can be arbitrarily changed by the user, The computer system is configured to function as follows: an analysis means for acquiring information about the content of a song from the song data when playback of the song data is instructed; a means for determining whether the information acquired by the analysis means is information about the content of a song as defined in the localization control table; a means for instructing the playback means to start playing the song data and sending a control signal to the sound image localization processing means to execute the corresponding sound image localization method in the localization control table when the means for determining whether the information acquired by the analysis means is information about the content of a song as defined in the localization control table; and a means for instructing the playback means to start playing the song data without performing sound image localization processing in the three-dimensional acoustic space by the sound image localization processing means when the means for determining whether the information acquired by the analysis means is not information about the content of a song as defined in the localization control table.
[0020] To achieve the above objectives, the simultaneous playback system of multiple performance sound sources of the present invention is A system for simultaneous playback of multiple sound sources, comprising: a control unit for controlling the operation of the entire music playback device when multiple sound sources are played simultaneously; a memory for storing various control programs and data; a music information storage unit for storing music data; a timer for outputting various timing signals; a music analysis unit for acquiring information about the content of music data stored in the music information storage unit; a display control unit for displaying various information on a display unit; a performance control unit for playing back the performance information contained in the music data stored in the music information storage unit and outputting it to a speaker; an operation unit for inputting operation signals; a communication unit for sending and receiving signals to and from the outside; a display unit for displaying information related to the performance information; at least two speakers capable of stereophonically reproducing the performance information from the performance control unit; an operation device for inputting various operation information; and a bus for transferring various data. The system has a localization control table that defines a localization method for each of the multiple performance sound sources, which in turn localizes the content of the music and the corresponding sound image as spatial sound within a three-dimensional sound space. Furthermore, multiple simultaneous playback systems for multiple performance sound sources are provided, each having a program for simultaneously playing multiple performance sound sources. These multiple systems are connected via network communication, and the multiple performance sound source data being played simultaneously in each system are mutually accessible. As a result, each simultaneous playback system can be used to record its own performance while utilizing the performance sound sources of other simultaneous playback systems, thus enabling additional participants to use it in virtual sessions. [Best Mode for Carrying Out the Invention]
[0021] The following describes embodiments of the method for simultaneous playback of multiple performance sound sources, the simultaneous playback system, and the simultaneous playback program for multiple performance sound sources of the present invention, using a big band having a rhythm section with multiple types of instruments and a front section with multiple types of instruments or vocals as an example of multiple performance sound sources, with reference to the attached diagrams.
[0022] The simultaneous playback program for multiple sound sources according to the present invention is an application program used in electronic devices such as smartphones and playback devices specialized for music playback (hereinafter collectively referred to as the simultaneous playback system 10 for multiple sound sources). The hardware configuration must include a CPU 12, memory 14, an operation unit 28 such as a touch panel or key switch, a display device such as an LCD, storage 32 for storing sound source data, and an audio output unit such as a speaker or earphone output. In other words, it is intended to be used in the simultaneous playback system 10 for multiple sound sources which has arithmetic processing capabilities equivalent to those of a computer.
[0023] For example, the simultaneous playback system 10 of multiple sound sources in one embodiment of the present disclosure may function as a computer that processes the session adjustment method of the present disclosure. Figure 1 is a diagram showing an example of the hardware configuration of the simultaneous playback system 10 of multiple sound sources according to one embodiment of the present disclosure. The simultaneous playback system 10 of multiple sound sources may be physically configured as a computer device including a CPU 12, memory 14, storage 32, communication unit 30, operation unit 28, output device, bus, etc. (Basic configuration)
[0024] More specifically, as shown in Figure 1, the simultaneous playback system 10 of multiple sound sources can be applied to various devices such as mobile terminals such as mobile phones, portable music players, and personal computers with AV functions, but Figure 1 shows the case of a mobile terminal such as a mobile phone. In Figure 1, 12 is a CPU that controls the operation of the entire music playback device 10, 14 is a memory consisting of ROM and RAM that stores various control programs and data, 18 is a music information holding unit that stores music data, 16 is a timer that outputs various timing signals, 20 is a display control unit 12 for displaying various information on a display unit 22, 24 is a performance control unit that plays back the performance information contained in the music data stored in the music information holding unit 18 and outputs it to a speaker 26, 28 is an operation unit that receives operation signals from an operating device, 30 is a communication unit, 22 is a display unit such as a liquid crystal display unit, 26 is a speaker (or headphones) of which at least two are provided that can reproduce the performance information from the performance control unit 24 in three dimensions, and 24 is a bus for transferring various data.
[0025] The performance control unit 24 includes a first interface circuit connected to a bus, a first FIFO (First in first out) buffer that stores performance information (sequence data) input via the first interface circuit, a sequencer that receives performance information via the first buffer and outputs event data to a performance sound source at timings specified by duration data in the performance information, a performance sound source that generates corresponding musical tones based on the performance information (event data) supplied from the sequencer, and a musical tone signal or MP3 decoder input from the performance sound source. The system includes a sound image localization processing circuit (3D circuit) that performs sound image localization processing based on a sound image localization control signal (3D control signal) input to the musical sound signal via a first interface circuit, a second interface circuit connected to the bus, a second FIFO buffer that stores audio compression format data (MP3 data in this example) supplied via the second interface circuit, an MP3 decoder that decodes the MP3 data supplied from the second buffer and outputs it to the sound image localization processing circuit, and an A / D converter (DAC) that converts the digital audio signal output from the sound image localization processing circuit into an analog signal, and the output signal of the DAC is amplified and output from stereo speakers 26.
[0026] The simultaneous playback system 10 for multiple sound sources comprises a CPU 12 equivalent to a computer's CPU, a memory 14 for temporary storage such as a playback buffer, a storage 32 for storing sound source data, an operation unit 28 used by the user, a display that shows the user interface of the simultaneous playback program for multiple sound sources, a playback unit that contributes to actual music playback, and a network connection unit equipped with functions for realizing network connectivity. Furthermore, Network Attached Storage (NAS) must comprise the CPU 12, memory 14, storage 32, and network connection unit of the simultaneous playback system 10 for multiple sound sources. Similarly, an external playback device must comprise the CPU 12, memory 14, operation unit 28, playback unit, and network connection unit of the simultaneous playback system 10 for multiple sound sources. Note that the external playback device may also have storage 32 and be capable of functioning like a NAS.
[0027] Memory 14 is a computer-readable recording medium and may consist of at least one of the following: ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. Memory 14 may also be called a register, cache, main memory 14 (main memory). Memory 14 can store executable programs (program code), software modules, etc., for implementing a method for simultaneously playing multiple sound sources according to one embodiment of the present disclosure.
[0028] The storage 32 is a computer-readable recording medium and may consist of at least one of the following: an optical disc such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disc, a digital multipurpose disc, a Blu-ray® disc), a smart card, flash memory 14 (e.g., a card, a stick, a key drive), a floppy® disk, a magnetic strip, etc. The storage 32 may also be called an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, server, or other suitable medium including at least one of memory 14 and storage 32.
[0029] The sequencer, upon receiving a playback command from the communication control process, reads the music data for playback pre-specified by the user from the music data in the music memory area and controls the playback sound source unit according to the music data. The communication control process is the process of causing the communication unit to perform the process of establishing the communication link described above, passing the information received by the communication unit to the audio processing unit, and controlling the communication unit to send terminal information supplied by the audio processing unit. In this process, when reading music data, event information is read and sent to the performance sound source unit. If duration information is read, the process is repeated, waiting for the time specified by that duration information to elapse before reading the next event information, thereby controlling the formation of the musical sound signal by the performance sound source unit. (Network configuration)
[0030] Figure 2 shows the overall configuration of a system to which multiple music playback systems, which are mobile terminals as described in Figure 1, are connected. In this figure, the mobile terminal 40 is configured to be able to connect to a large-scale network 34 such as the Internet via a base station 38. A distribution server 36 is connected to the large-scale network 34, and the mobile terminal 40 is configured to download music data from the distribution server 36, store it in the music information storage unit 18 (Figure 1), and play it back using the playback control unit 24.
[0031] The simultaneous playback program for multiple sound sources according to this invention makes it easy to play sound source data stored in a NAS or sound source data stored in a simultaneous playback system 10 on an external playback device 12, using the program as a remote control device, when external playback devices such as network-attached storage (NAS) and music playback components are connected to each other using compatible standards. In this embodiment, the term NAS includes not only literal network-attached storage but also all devices that have a storage 32 capable of storing sound sources and are network-connectable.
[0032] The communication unit 30 is hardware (transceiver / receiver device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, network controller, network card, communication module, etc. The communication unit 30 may be configured to include, for example, a high-frequency switch, duplexer, filter, frequency synthesizer, etc., in order to implement at least one of frequency division duplex (FDD) and time division duplex (TDD). (For example, it may be a touch panel.) (Data format)
[0033] The music data stored in the music information storage unit is basically data containing performance information (sequence data) such as SMF (Standard MIDI File) and SMAF (Synthetic music Mobile Application Format), but it may also be audio data (audio compression format data such as MP3 (MPEG Audio Layer-3) and AAC (MPEG-2 Audio AAC (Advanced Audio Coding))). Furthermore, this music data may be originally in stereo or 3D sound format, or it may be completely monaural data. In addition, the music analysis unit may be implemented by software processing by the CPU. Furthermore, the performance control unit 24 is a performance sound source unit that can play back the music data, and is equipped with a decoder such as an FM performance sound source, a WT (Wave Table) performance sound source, an MP3 decoder, or a performance sound source that combines these.
[0034] The simultaneous playback system 10 for multiple sound sources is equipped with a user interface that has a touch panel stacked on a display, and is equivalent to, for example, a personal computer, smartphone, digital still camera, digital video camera, or navigation device. Of course, the present invention can also be applied to other simultaneous playback systems for multiple sound sources.
[0035] Touch panels can detect not only contact from the user's fingers, etc., but also proximity (meaning approaching a predetermined distance without direct contact) by being transparent to the display screen. Here, "user's fingers, etc." refers to other body parts of the user, as well as styluses (pen-shaped pointers) compatible with touch panels. Hereafter, "user's fingers" will be used simply, but this will include other body parts and styluses.
[0036] For example, a capacitive touch panel can be used. Other methods are also acceptable, as long as they can detect not only the user's finger contact but also its proximity. The touch panel notifies the CPU 12 via the bus of detection information indicating the results of the user's finger contact or proximity detection.
[0037] The display consists of, for example, an LCD or an organic EL display, and displays screen images supplied from the CPU 12 via a bus. The memory 14 stores control programs executed by the CPU 12. The memory 14 may also store various types of content (image data, audio data, video data, application programs, etc.) to be processed by the simultaneous playback system 10 of multiple sound sources.
[0038] The CPU 12 controls the entire simultaneous playback system 10 of multiple sound sources by executing a control program stored in memory 14. In addition to the above configuration, hardware such as push buttons that serve as a user interface may be provided.
[0039] Next, the operation screen for using the simultaneous playback program for multiple sound sources of the present invention will be described. Regarding the graphical user interface, which is the operation screen for the simultaneous playback program for multiple sound sources, the display shown on the display of the simultaneous playback system 10 for multiple sound sources is an example of the graphical user interface (hereinafter referred to as GUI) for the simultaneous playback program for multiple sound sources. In order to enable the selection of a playback device and the selection of playback sound source data on the same screen, the simultaneous playback program for multiple sound sources according to the present invention may have a playback device selection area and a sound source data selection area fixedly arranged on the GUI screen. Furthermore, when making selections using the playback device selection area and the sound source data selection area, a hierarchical display area may be provided to display the selection targets hierarchically. (3D audio processing)
[0040] In the 3D sound effect addition process, music data specified by the operation of the control unit is read from the music data in the music memory area, converted into music data with appropriate 3D sound effect information added, and the converted music data is stored in the music memory area. In one preferred embodiment, the program for the 3D sound effect addition process is pre-written to the ROM of memory 14. In another embodiment, the program for the 3D sound effect addition process is downloaded to memory 14 by the user from a predetermined site on the internet.
[0041] There are three types of processes that can be performed to add spatial sound effect information in the spatial sound effect addition process. When the CPU 10 performs the spatial sound effect addition process, the user can specify the execution of one or more of these processes by operating the operation unit 28.
[0042] This process outputs the target localization position for sound image localization. This calculation determines the localization position of the sound image in a three-dimensional Cartesian coordinate system and outputs the localization position to the 3D sound system.
[0043] For example, if the positions of multiple sound sources and the listening position (point of interest) are calculated as distance R, azimuth angle θ, and elevation angle φ, the localized position (X,Y,Z) of the target in a three-dimensional Cartesian coordinate system is calculated using the following equations 1 to 3.
[0044] X = R·cosφ·cosθ Equation 1 Y = R·cosφ·sinθ Equation 2 Z = R·sinφ Equation 3 The spatial audio system instructs the audio system to set the transfer functions used in filters L and R. For example, the audio system has multiple transfer functions in reserve that correspond to the instructions from the spatial audio system, and it selectively sets the transfer function. As a result, the sound image formed by speakers L and R is localized to a localization position corresponding to the point of interest.
[0045] a. Perspective Effect Addition Process: In this perspective effect addition process, the sound intensity information (e.g., velocity, volume) contained in the original music data is converted into distance to the virtual performance sound source position. This generates spatial sound effect information that sets the virtual performance sound source position at that distance away from the user, and this information is added to the original music data.
[0046] b. Movement Effect Addition Process In this movement effect addition process, periodic changes such as note information contained in the original music data are extracted, and spatial sound effect information is generated that periodically moves the position of the virtual performance sound source in synchronization with the period of these changes, and this is added to the original music data. The trajectories for periodically moving the position of the virtual performance sound source are available as a straight trajectory extending left and right in front of the user, a circular trajectory surrounding the user, and an elliptical trajectory surrounding the user, and trajectory definition information for these trajectories (such as functions for calculating the coordinates of each point on the trajectory) is stored in the memory unit. The user can select a desired trajectory from among them by operating the control unit. In this movement effect addition process, the virtual performance sound source position to be assigned to each periodically changing note is calculated based on the trajectory definition information of the trajectory specified by the user, and spatial sound effect information is generated that localizes the sound image of the speaker playback sound to that virtual performance sound source position.
[0047] c. Doppler effect addition process: In this Doppler effect addition process, pitch control information such as pitch bend events contained in the original music data is converted into stereoscopic sound effect information that moves the virtual performance sound source from far away to closer to the user, or moves the virtual performance sound source from far away to further away from the user, and is added to the music data.
[0048] The 3D sound effect addition process is initiated when the user selects a song using the control unit and instructs the system to perform a conversion for the addition of 3D sound effects. First, the conversion conditions are set. Specifically, a screen is displayed on the display unit 22 asking the user which of the above-mentioned perspective effect addition process, motion effect addition process, and Doppler effect addition process they wish to execute, and the user's instructions are obtained via the control unit 28. At this time, the user may instruct the system to perform one of the perspective effect addition process, motion effect addition process, or Doppler effect addition process, or two or all of them. Also, if the user instructs the system to perform the motion effect addition process, a screen is displayed on the display unit 22 asking for the trajectory to move the virtual performance sound source position, and the user's instructions are obtained via the control unit 28.
[0049] Next, a perspective effect addition process, a motion effect addition process, or a Doppler effect addition process is executed to generate music data with stereoscopic sound effect information added to the original music data, and this data is stored in the music storage area of memory 14. The generated music data may overwrite the original music data in the music storage area, or it may be stored in the music storage area with a different file name than the original music data. The user instructs which storage method to use by operating the operation unit 28. The above describes the processing details of the stereoscopic sound effect addition process.
[0050] The subsystem includes a spatial sound system and an audio system. The spatial sound system instructs the audio system to set a transfer function in order to localize the sound image. Specifically, as will be described later, it instructs the transfer function to be set based on the target localization position output from the CPU 12. The transfer function is one used in well-known sound localization techniques, and more specifically, it is the head-related transfer function (HRTF) that represents the sound transmission characteristics from the virtual sound source to the listener's eardrum.
[0051] The audio system includes a sound source, L and R filters, and L and R speakers.
[0052] Speakers L and R are configured such that one is the left speaker (L) for outputting sound from the user's left side, and the other is the right speaker (R) for outputting sound from the user's right side.
[0053] The left speaker L corresponds to the left filter L, and the right speaker R corresponds to the right filter R. As described above, when localizing the sound image, the left filter L and the right filter R perform filtering based on the transfer function. The signal from the performance sound source, filtered in this way, is output to the left speaker L and the right speaker R, respectively.
[0054] The designation of the left speaker (L) and right speaker (R) indicates a 2-channel system, and each of these may consist of multiple speakers. Furthermore, under normal conditions where eye-tracking guidance is not performed, the audio signal is output from the sound source based on the audio environment set by the user. In this case, the audio system outputs the signal without passing through filters L and R. The left speaker (L) and right speaker (R) may also consist of stereo headphones.
[0055] The audio environment can be thought of as the equalizer settings that change the frequency characteristics of the output signal. It can also be thought of as the left / right balance setting. If the speakers are arranged in the front-to-back direction, the front-to-back balance setting is also included. The audio environment settings are made via a group of control switches and stored in the external storage unit 32. If no settings have been made by the user, the default audio environment is assumed to be stored in the external storage unit 32.
[0056] Figure 3 shows the touch panel operation screen 42. More specifically, the touch panel operation screen 42 is an operation screen for playing back digital recordings of a big band performance of a selected song under various conditions, when playing back a song by a big band consisting of a rhythm section 74 consisting of piano, drums, bass, guitar, and vocals, a trumpet section 76 consisting of multiple trumpets, a trombone section 78 consisting of multiple trombones, and a saxophone section 80 consisting of multiple saxophones. When playing a digital recording of a selected song, the top of the screen displays section diagrams 44 showing the instrument arrangements for the rhythm section 74, trumpet section 76, trombone section 78, and saxophone section 80. The user can then select one or more instruments from the rhythm section 74, trumpet section 76, trombone section 78, and saxophone section 80. More specifically, the rhythm section 74 displays drum icons 46, bass icons 48, piano icons 50, guitar icons 52, and vocal icons 54; the trumpet section 76 displays multiple trumpet icons 56; the trombone section 78 displays multiple trombone icons 58; and the saxophone section 80 displays multiple saxophone icons, each of which is clickable.
[0057] In the center of the screen, there is an image showing how to select sound, volume, and metronome by the length of a horizontal bar, and each of these is controlled by the length of the horizontal bar. More specifically, the following are displayed from top to bottom of the screen: a reset icon 64, a listening position setting icon 66, a space size selection bar 68, a volume selection bar 70, and a metronome volume adjustment bar 72. Clicking the reset icon 64 resets the state of the operation screen; clicking the listening position setting icon 66 sets the listening position; and clicking the space size selection bar 68 sets the space size for simultaneously playing multiple audio sources. The left vertical section of the screen displays the length of the song using a left vertical bar, while the right vertical bar shows an image for selecting the playback section of the song by its length and position, and the operation is performed using the length and position of the right vertical bar. More specifically, the left vertical bar is the progress bar 60, the right vertical bar is the playback section display bar 62, and at the bottom of the left vertical section of the screen, a repeat display icon 82 and a repeat time display icon 84 are displayed, showing an image for repeating the playback of the selected section of the song. At the bottom of the screen, separate from the section layout diagram 44, operation images are shown for each section: rhythm section 74, trumpet section 76, trombone section 78, and saxophone section 80, allowing for selection and deselection of each section. Instructions can be given using either these operation images or the section layout diagram 44, but the operation screen 42 is configured so that one instruction is reflected in the other. The display of the operation screen 42 and the commands issued by clicking on the touch panel screen are performed by the CPU 12, as in the conventional system.
[0058] Figure 4 shows a modified version of the operation screen 42. More specifically, in addition to the center position 90, left position 92 and right position 94 are added as audience positions. Furthermore, when selecting a space size and selecting instruments from multiple performance sound sources, the size of the section layout diagram 44 is enlarged or reduced in the space size selection bar 68 according to the selected space size (Figure 4(A) shows the case when the space size is large, and Figure 4(B) shows the case when the space size is small). This allows the user to visually grasp the image of the selected space size, and the direction of the selected performance instrument can be selected using the arrow display 96, reflecting the direction of sound production for each instrument and adding functionality to 3D sound processing. As described above, when creating a session, the selection of songs, playback conditions, and conditions for 3D audio processing can all be set visually.
[0059] As shown in Figure 6, in step 1, the system having a program for simultaneous playback of multiple sound sources is launched via an icon. Next, in step 2, if the sound source of the new song is stored in the external storage unit, storage 32, the new song information is imported into the internal storage unit, memory 14. Next, in step 3, if an existing session already exists, the user chooses whether or not to access the existing session. If they choose to access it, they proceed to step 5; otherwise, they proceed to step 4. Next, in step 4, the user selects songs, including the newly added song. Next, in step 5, the user chooses whether or not to play and preview the selected song. If they choose to preview it, they proceed to step 6; otherwise, they return to step 4.
[0060] Next, in step 7, you choose whether to create a session for the song you listened to. If you choose to create a session, proceed to step 8; otherwise, exit the app. Next, in step 8, set the tempo, volume, and space size for the song you want to create a session for on the operation screen. Next, in step 9, select the instruments to play and the listening position for the song you want to create a session for. More specifically, in a big band consisting of a rhythm section with piano, drums, bass, guitar, and vocals, a trumpet section with multiple trumpets, a trombone section with multiple trombones, and a saxophone section with multiple saxophones, for example, select only the drums in the rhythm section, the first trumpet in the trumpet section, the first trombone in the trombone section, and all the tenor, alto, and baritone saxophones in the saxophone section, and set the listening position to the piano position. Note that the instruments to play and the listening position can be selected by clicking directly on the instrument arrangement diagram on operation screen 42, as described above, or by selecting them in each section on operation screen 42.
[0061] Next, in step 10, spatial sound processing is performed based on the selected space size, instrument, and listening position. Then, in step 11, the start position, end position, and number of repetitions of the selected piece of music are selected. Next, in step 12, multiple performance sound sources are played simultaneously and can be used as accompaniment sound sources for N-1 for the user's instrument practice. In this case, the accompaniment sound source to be listened to is processed in spatial sound to create an immersive and realistic listening experience, as if listening from the piano's position, allowing for effective practice even when practicing alone, giving the feeling of performing live on the spot. Furthermore, if all instruments are selected and the listening position is set to the audience position, it is possible to listen to music with an immersive and realistic experience, as if listening to a completed piece of music in a large performance hall, depending on the selected space size. Next, in step 14, you choose whether to save the simultaneously played songs as a session. If you choose to save, proceed to step 14 and finish. If you choose not to save, return to the beginning of step 4 and repeat the same process for another song.
[0062] According to the simultaneous playback method of multiple performance sound sources having the above configuration, each of the multiple performance sound sources is recorded individually and independently using multi-track recording, the space size of the multiple performance sound sources is set on the operation screen, the listening position of the performance sound source is identified on the operation screen, the three-dimensional coordinates of the listening position are calculated based on the identified listening position, the distance and orientation between the position of each N-1 performance recording and the listening position are calculated, and then three-dimensional position information is added to the performance sound source using object-based spatial sound processing, while the positional relationship between the position of each performance sound source and the listening position is adjusted according to the selected space size and listening position. By processing the sound of each performance sound source to reflect differences in volume, time, frequency characteristics, phase, reverberation, and the Doppler effect, a sound field within a specific spatial size is simulated. By selecting one of the N-1 performance sound sources, ..., N-1 on the operation screen and simultaneously playing the selected single or multiple performance recordings that have undergone spatial sound processing, the system simulates a sense of presence or immersion according to the listening position when playing multiple performance sound sources, providing excellent performance sound sources for appreciation and making them usable as more effective accompaniment sound sources for practice.
[0063] A second embodiment of the present invention will be described below. In the following description, components similar to those in the first embodiment will be given the same reference numerals and their descriptions will be omitted. The characteristic features of this embodiment will be described in detail below with reference to Figures 6 to 28. The second embodiment of the present invention is characterized by its use of a method for simultaneously playing multiple performance sound sources. While the first embodiment is for N-1 practice accompaniment for big bands or for listening to music, this embodiment involves conducting a so-called additional participation type virtual session using multiple performance sound sources. Here, the "virtual" in "virtual session" means that, for example, in a normal session, each section of the combo, such as the rhythm section (piano, drums, bass, etc.) and the front section (saxophone, trumpet, trombone, etc.), gathers in one place and performs live simultaneously. However, in this embodiment, each section of the combo does not perform live, but rather creates recorded performance sound sources and participates sequentially. Therefore, it is not a simultaneous performance, but it means that the completed song is processed with stereophonic sound and played back simultaneously, thus virtually simulating a session.
[0064] To implement an additionally interactive virtual session, similar to the first embodiment, it is necessary to add functions for adding session members, recording performances, sharing performances, setting up performance collaborations, and joining publicly available sessions, all based on the premise of starting and listening to a session. The main functions are explained below using flowcharts.
[0065] Figure 6 shows the flowchart from launching the app to starting a session. As shown in Figure 6, if the app is launched in step 1, then in step 2, the data for the song, artist, and session (information only) is loaded. Next, in step 3, the user chooses whether to select a song; if selected, the user proceeds to step 4, otherwise, the user proceeds to step 6. Next, in step 4, the user enters session information including the title, description, and public / private setting. Alternatively, in step 6, an existing session is selected, and in step 7, the session is started. Next, a new session is created in step 5, and in step 7, the session is started.
[0066] Figure 7 shows a flowchart from the start of a session to saving it. As shown in Figure 7, in step 1, the session is started, and in step 2, data is loaded for the music audio data, spatial arrangement, tempo, session tonality, and user settings. Next, in step 3, the user chooses whether to set up the session; if so, they proceed to step 8, otherwise to step 4. Next, in step 4, it is determined whether the session is stopped; if so, they proceed to step 5, otherwise to step 6. Next, in step 5, the session is played. Next, in step 6, spatial audio processing is performed on the stream data, and multiple performance sound sources are played simultaneously in sync to put the session into a playing state. Next, in step 7, the user chooses whether to set up the session; if so, they proceed to step 8, otherwise to step 13. Next, in step 8, the session is set up, and in step 9, the performance parts are set by muting the specified part, muting other parts, adjusting the volume of the specified part, turning the metronome on and off, and adjusting the volume. Next, in step 10, the playback position is set by setting the playback start position, setting the playback range, and turning repeat playback on or off. Next, in step 11, the conditions for spatial sound processing are set by setting the listening position and setting the space size of the listening space. Next, in step 12, the music is set by setting the tempo and adjusting the settings. Next, in step 13, the session is stopped. Next, in step 14, you choose whether to save the settings. If you choose to save, proceed to step 15 to save the settings; otherwise, proceed to step 16 to finish.
[0067] Figure 11 shows the flowchart for adding a session member. As shown in Figure 11, in step 1, if a session member is to be added, in step 2, the invitee is notified using a link URL, email or LINE, a communication app, or a QR code as the means of invitation. Next, in step 3, the user chooses whether to cancel the invitation; if they choose to cancel, they return to step 4, or if they do not, they return to step 1. Next, in step 4, the user decides whether to resend the invitation notification; if they choose to resend, they return to step 5, or if they do not, they return to step 2. Next, in step 5, the person who received the invitation accepts it. On the other hand, if a participant request is to be handled in step 6, the session is made public in step 7, and then a participant request to join the session is received in step 8. Next, in step 9, a decision is made on whether to approve the request; if approved, the process proceeds to step 10, or if not approved, to step 11. Next, in step 10, if the invited person approves, or if the participant request is approved, participation in the session is completed, and the process ends in step 16. If the participant request is not approved, a notification of exclusion from the session is sent in step 11, and the process ends in step 16. Furthermore, if you manage the participating members in step 12, in step 13 you decide whether to change the permissions. If you do, proceed to step 14; otherwise, proceed to step 15. Note that changing permissions refers to a change between Member and Master, where Master can decide to change session settings and add session participants, and Member can only decide to share their own performances. Next, in step 14 you change the permissions and return to step 12. Alternatively, in step 15 you remove a participant and proceed to step 11. Figures 8 through 10 show the relevant operation screens. In the figures, 90 is the session invitation screen, 92 is the approval icon, 94 is the resend icon, 96 is the Master icon, 98 is the member icon, 102 is the role change screen, and 104 is the execute icon.
[0068] Figure 14 shows a flowchart for recording a performance. As shown in Figure 14, in step 1, when recording a performance, in step 2, the recording environment is set up, such as muting the part to be played. Next, in step 3, recording is started by pressing the record button, etc. Next, in step 4, the performance recording is converted into digital data from the microphone and saved to a temporary folder. Next, in step 5, it is decided whether to pause the recording; if paused, proceed to step 6; otherwise, proceed to step 7. Next, in step 6, it is decided whether to resume recording; if resumed, return to step 3; otherwise, proceed to step 7. Next, in step 7, recording is ended. Next, in step 8, the performance part is specified, the file name and comment information are added to the recording data and saved, and in step 9, the process ends. Figures 12 and 13 show the relevant operation screens. In the figures, 106 is the volume reset icon, 108 is the mute icon, and 110 is the recording entry screen.
[0069] Next, we will discuss how the session tracks and session members are chosen, and how the recordings of those performances are shared among the members. Figure 18 shows a flowchart for sharing a performance. As shown in Figure 18, in step 1, if you choose to share a performance, in step 2, you share the performance, and in step 3, you upload the file to a designated server via the network. Then, in step 4, the performance is shared with the session members, and in step 5, the process ends. On the other hand, if you delete the shared performance in step 1, in step 2, you check if the performance data exists on the server. If it does, you proceed to step 3; otherwise, you proceed to step 5. Then, in step 3, you unshare the performance. Next, in step 4, you delete the file on the server. Finally, in step 5, session members are no longer able to listen to the performance, and the process ends in step 6. Furthermore, if all performance-related data is to be deleted in step 1, then all performance-related data is deleted in step 2. Next, in step 3, it is determined whether performance data exists on the server; if it exists, the process proceeds to step 4, otherwise to step 5. Next, in step 4, the performance data is deleted from the server. Next, in step 5, the data on the terminal is deleted. Next, in step 6, the data attached to the recording data is deleted, and in step 7, the process ends. Figures 15 to 17 show the relevant operation screens. In the figures, 112 is the session progress screen, 114 is the session information screen, 116 is the cloud share icon, and 118 is the download icon.
[0070] Figure 23 shows a flowchart for setting up performance collaboration. Session collaboration refers to performance data that has been adjusted so that other users can listen to it as a virtual session. When it is made public, adjustments are made to determine whether not only the session members but also users of the program app that simultaneously plays multiple performance audio sources can view it. As shown in Figure 23, in step 1, when setting up a performance collaboration, in step 2, select the performance tracks of the members to be used for playback. Next, in step 3, set the sharing scope. Next, in step 4, decide whether to make it publicly available; if so, proceed to step 5, otherwise proceed to step 6. Next, in step 5, app users can listen to the session collaboration, and in step 7, it ends. On the other hand, in step 6, only session members can listen to the session collaboration, and in step 7, it ends. Figures 19 to 22 show the relevant operation screens. In the figures, 120 is the collaboration screen, 122 is the solo icon, 124 is the rhythm section icon, 126 is the trumpet icon, 128 is the pin icon, 130 is the song title screen, 132 is the session members screen, and 134 is the public selection screen.
[0071] Figure 28 shows a flowchart for joining a public session. As shown in Figure 28, in step 1, when joining a public session, in step 2, you find a public session by searching yourself, using system recommendations, or through shares from other users. Next, in step 3, you confirm the session by listening to the session information, session participants, or session collaborations. Next, in step 4, you send a join request to the confirmed session. Next, in step 5, a member of the session to which the request was sent who has Master privileges decides whether to approve the request; if approved, proceed to step 7, otherwise proceed to step 6. Next, in step 7, you start joining the session, and in step 8, you finish. Alternatively, in step 6, you are notified of the rejection of the request and return to step 2. Figures 24 to 26 show the relevant operation screens. In the figures, 136 is the session details screen, 138 is the listening icon, 140 is the participation request icon, 142 is the participation request icon, 144 is the recommendations screen, 146 is the popular sessions screen, and 148 is the popular artists screen.
[0072] As described above, by utilizing the simultaneous playback method, system, and program for multiple performance audio sources in a virtual session, it is possible to recreate an immersive and realistic performance, similar to a live performance, even without the combo members gathering in the same space and performing live simultaneously. This is achieved by simultaneously playing multiple performance audio sources with 3D sound processing. Furthermore, as an interactive virtual session, when each participant records their own performance while listening to the accompanying audio source, they can perform while listening to the 3D processed accompanying audio source, allowing for a more cohesive recording experience. As a variation, when each participant records their own performance while listening to the accompanying audio source, they can use the built-in microphone, such as in an iPhone. In this case, they can record while listening to the accompanying audio source without 3D sound processing, and when playing back the completed recording, listeners can choose the space size, listening position, and performance source before playing it back with 3D sound processing for enjoyment.
[0073] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Therefore, the descriptions in the present disclosure are for illustrative purposes only and are not intended to be restrictive in any way. For example, although this embodiment has been described as an N-1 practice accompaniment for a big band having a rhythm section of piano, bass, and drums, and a front section of saxophone, trumpet, and trombone, it is not limited to that. For example, by selecting the listening position as the audience position, it may also be used as a performance sound source for appreciation that provides a sense of presence and immersion. For example, although this embodiment has been described as an N-1 practice accompaniment for a big band having a rhythm section of piano, bass, and drums, and a front section of saxophone, trumpet, and trombone, it is not limited to that. For example, any of the rhythm section of piano, bass, and drums, and any of the front section of saxophone, trumpet, and trombone, for example, a saxophone, could be selected and used as an N-1 practice accompaniment for a so-called combo.
[0074] For example, in this embodiment, when describing the playback of multiple performance sound sources including a rhythm section and a front section, it was explained that any of the multiple performance sound sources may be used, for example, when the vocalist sings while moving. However, the invention is not limited to this, and the listening position may also move during the performance, or any of the multiple performance sound sources may move while the listening position moves during the performance. For example, in this embodiment, sound processing was described as being performed when playing back a selected song after its performance as a so-called virtual session is completed. However, the invention is not limited to this, and for example, when recording one's own performance based on the recorded performance sound source of N-1 before the virtual session is completed, the recorded performance sound source of N-1 may be individually processed. For example, in the second embodiment, in an additional-participation virtual session, the tempo, key, and spatial sound processing conditions of the accompanying music, which serves as the basis for participating in the virtual session, may be set in common among all participants.
[0075] The various modes / implementations described in this disclosure may be applied to at least one of the following systems: LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), FRA (Future Radio Access), NR (new Radio), W-CDMA (registered trademark), GSM (registered trademark), CDMA 2000, UMB (Ultra Mobile Broadband), IEEE802.11 (Wi-Fi (registered trademark)), IEEE802.16 (WiMAX (registered trademark)), IEEE802.20, UWB (Ultra-Wide Band), Bluetooth (registered trademark), and other appropriate systems, as well as next-generation systems extended based on these. Furthermore, multiple systems may be applied in combination (for example, a combination of at least one of LTE and LTE-A with 5G).
[0076] The processing procedures, sequences, flowcharts, etc., of each aspect / embodiment described herein may be reordered, provided they are consistent with each other. For example, the methods described herein present various step elements in an exemplary order and are not limited to that specific order.
[0077] Input and output information may be stored in a specific location (e.g., memory 14) or managed using a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be sent to other devices.
[0078] Each aspect / embodiment described herein may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of predetermined information (e.g., notification that "X is") is not limited to explicit notification, but may also be implicit (e.g., by not providing notification of predetermined information).
[0079] Software, whether called software, firmware, middleware, microcode, hardware description language, or by any other name, should be interpreted broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, and so on.
[0080] Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technology (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technology (such as infrared or microwave), then at least one of these wired and wireless technologies is included in the definition of a transmission medium.
[0081] The block diagrams used in describing the embodiments show functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired or wireless connections). A functional block may be realized by combining software with one or more devices.
[0082] Functions include, but are not limited to, judgment, decision, determination, calculation, processing, search, confirmation, reception, transmission, output, access, resolution, selection, comparison, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), and assigning. For example, a functional block (configuration unit) that enables transmission is called a transmitting unit or transmitter. As mentioned above, the method of implementation is not particularly limited.
[0083] The term "device" can be interpreted as a circuit, device, unit, etc. The hardware configuration of the simultaneous playback system 10 of multiple sound sources may be configured to include one or more of the devices shown in the figure, or it may be configured to omit some of the devices.
[0084] Furthermore, each device, such as the CPU 12 and memory 14, is connected by a bus for communicating information. The bus may be configured using a single bus, or different buses may be configured for each device.
[0085] Furthermore, the simultaneous playback system 10 for multiple sound sources may include hardware such as a microCPU 12, a digital signal CPU 12 (DSP: Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), and an FPGA (Field Programmable Gate Array), and some or all of each functional block may be realized by the hardware. For example, CPU 121001 may be implemented using at least one of these hardware components.
[0086] Information notification is not limited to the embodiments described herein and may be carried out by other means. For example, information notification may be carried out by physical layer signaling (e.g., DCI (Downlink Control Information), UCI (Uplink Control Information)), upper layer signaling (e.g., RRC (Radio Resource Control) signaling, MAC (Medium Access Control) signaling, broadcast information (MIB (Master Information Block), SIB (System Information Block))), other signals, or combinations thereof. RRC signaling may also be called RRC messages, and may be, for example, RRC Connection Setup messages, RRC Connection Reconfiguration messages, etc.
[0087] The system 10 for simultaneous playback of multiple sound sources, the NAS, and the external playback device recognize and link each other via a network connection. One possible method for this network connection is mutual recognition using a communication protocol called UPnP (Universal Plug and Play). For example, by adopting an industry standard (guideline) that enables network connections between electronic devices within the home by making products compatible with each other, such as DLNA (an abbreviation for Digital Living Network Alliance, and "DLNA" is a registered trademark of the said organization), it is possible to create a system where electronic devices that adopt this industry standard can easily connect to the network. Although DLNA is a standard for LAN connections, other communication standards may be used as long as they meet the requirements necessary for music playback, such as transfer speed, and are not limited to LAN connections. [Brief explanation of the drawing]
[0088] [Figure 1] This figure shows the overall configuration of a system for the simultaneous playback of multiple sound sources, which is a first embodiment of the present invention. [Figure 2] This diagram shows the network communication configuration in a system for the simultaneous playback of multiple sound sources, which is a first embodiment of the present invention. [Figure 3] This figure shows the operation screen in a system for simultaneous playback of multiple sound sources, which is a first embodiment of the present invention. [Figure 4] This figure shows a modified example of the operation screen in a system for simultaneous playback of multiple sound sources, which is the first embodiment of the present invention. [Figure 5] This is a flowchart showing the simultaneous playback of multiple performance sound sources in a system for simultaneous playback of multiple performance sound sources, which is the first embodiment of the present invention. [Figure 6] This is a flowchart showing the process from launching the application to starting a session in a system for simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 7] This is a flowchart showing the process from the start of a session to saving a session in a system for simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 8] This figure shows a screen for inviting someone to join a session in a system for simultaneously playing multiple performance sound sources, which is a second embodiment of the present invention. [Figure 9] This figure shows a screen for inviting someone to join a session in a system for simultaneously playing multiple performance sound sources, which is a second embodiment of the present invention. [Figure 10] This figure shows a screen for inviting someone to join a session in a system for simultaneously playing multiple performance sound sources, which is a second embodiment of the present invention. [Figure 11] This is a flowchart for adding session members in a system for the simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 12] This figure shows a screen for recording a performance in a system for simultaneously playing multiple performance sound sources, which is a second embodiment of the present invention. [Figure 13]This figure shows a screen for recording a performance in a system for simultaneously playing multiple performance sound sources, which is a second embodiment of the present invention. [Figure 14] This is a flowchart for recording a performance in a system for simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 15] This diagram shows a screen illustrating the management and sharing of recorded data in a system for the simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 16] This diagram shows a screen illustrating the management and sharing of recorded data in a system for the simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 17] This diagram shows a screen illustrating the management and sharing of recorded data in a system for the simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 18] This is a flowchart for sharing a performance in a system for the simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 19] This figure shows a collaboration settings screen for a performance in a system for simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 20] This figure shows a collaboration settings screen for a performance in a system for simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 21] This figure shows a collaboration settings screen for a performance in a system for simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 22] This figure shows a collaboration settings screen for a performance in a system for simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 23] This is a flowchart for setting up performance collaboration in a system for simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 24] This figure shows a screen displaying detailed session information in a system for the simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 25]This figure shows a screen displaying detailed session information in a system for the simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 26] This figure shows a screen displaying detailed session information in a system for the simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 27] This figure shows a screen for evaluating a publicly available session in a system for simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Figure 28] This is a flowchart for joining a publicly available session in a system for the simultaneous playback of multiple performance sound sources, which is a second embodiment of the present invention. [Explanation of symbols]
[0089] 10: System for simultaneous playback of multiple performance audio sources 12:CPU 14: Memory 16: Timer 18: Music Information Storage Unit 20: Display Control Unit 22: Display section 24: Performance Control Unit 26: Speaker 28:Operation unit 30: Communications Department 34 Network Communications Department 38 base station 40 devices 36 Distribution Servers 42 Operation screen 44 Section Layout Diagram 46 Drum Icons 48 Base Icon 50 piano icons 52 Guitar Icons 54 Vocal Icons 56 Trumpet icon 58 Trombones Icon 60 Progress bar 62 Performance position indicator bar 64 Reset Icon 66 Listening position setting icon 68 Space Size Selection Bar 70 Volume selection bar 72 Metronome Volume Adjustment Bar 74 Rhythm Section 76 Trumpet Section 78 Trombone Section 80 Saxophone Section 82 Repeat display icon 84 Repeat time display icon 90 Session Invitation Screen 92 Approval Icon 94 Resend icon 96 Master icon 98 member icon 102 Role Change Screen 104 Run icon 106 Volume Reset Icon 108 Mute Icons 110 Recording entry screen 112 Session progress screen 114 Session Information Screen 116 Cloud Share Icon 118 Download Icons 120 Collaboration screen 122 Solo Icons 124 Rhythm Section Icons 126 Trumpet Icon 128 Pin Icons 130 Song Title Screen 132 Session Member Screen 134 Publish Selection Screen 136 Session Details Screen 138 Listening icon 140 Participation Request Icon 142 Participation Request Icon 144 Recommended Screen 146 Popular Session Screens 148 Popular Artists Screen
Claims
1. A method for simultaneously playing multiple (N) performance sound sources using a computer having an information processing unit, an information storage unit, an information input unit, and an information output unit, The aforementioned computer performs the step of recording each of the multiple performance sound sources individually and independently using multi-track recording, On the operation screen, there is a step to set the space size when playing multiple audio sources, On the operation screen, there is a step in which the listening position when playing the performance audio source is identified from multiple performance audio source positions and other music listening positions, The process involves calculating the three-dimensional coordinates of the listening position during playback, The steps include calculating the distance and orientation between the position of each performance sound source and the listening position, When specifying the listening position for playback of a performance audio source from multiple performance audio source positions, the user must select one of the N-1 performance audio sources, ..., N-1, on the operation screen. Using object-based spatial audio processing, three-dimensional positional information is added to the performance sound source information, and based on the selected space size and listening position, the sound playback processing of each performance sound source is performed according to the positional relationship between the position of each performance sound source and the listening position, thereby simulating and constructing a sound field within a specific space size. A method for simultaneously playing multiple performance audio sources, characterized by performing the steps of simultaneously playing selected single or multiple performance recordings and utilizing them as accompaniment audio sources for performance practice or for musical appreciation.
2. A method for simultaneously playing multiple performance sound sources according to claim 1, wherein spatial sound processing is performed based on the selected space size and listening position, including the direction of sound production of the multiple performance sound sources, to reflect volume differences, time differences, changes in frequency characteristics, changes in phase, changes in reverberation, and the Doppler effect.
3. The method for simultaneously playing multiple performance sound sources according to claim 1, wherein, when playing the multiple performance sound sources, synchronization settings are made between the selected performance sound sources before performing stereophonic sound processing.
4. The method for simultaneously playing multiple performance sound sources according to claim 1, wherein the multiple performance sound sources consist of a rhythm section including piano, drums, bass, and guitar, a trumpet section consisting of multiple trumpets, a trombone section consisting of multiple trombones, and a saxophone section consisting of multiple saxophones, and the simultaneous playback stage of the performance recording sound sources is the playback of a song performed by a big band.
5. A method for simultaneously playing multiple performance sound sources according to claim 1, wherein a listening position is selected from multiple (N) performance sound sources, and any of the remaining 1, ..., N-1 performance sound sources are selected as a performance sound source, thereby simultaneously playing the multiple performance sound sources to use as a practice accompaniment performance sound source for the listening position.
6. The method for simultaneously playing multiple performance sound sources according to claim 1, wherein, in addition to multiple (N) performance sound sources, an audience position is set separately and the audience position is selected, thereby simultaneously playing multiple performance sound sources to create a performance piece for listening using multiple full performance sound sources.
7. A method for simultaneously playing multiple performance sound sources according to claim 1, wherein the method involves selecting and playing any performance sound source of the rhythm section from among multiple (N) performance sound sources, and then selecting and playing another performance sound source of the rhythm section from among multiple (N) performance sound sources, thereby being used to acquire a sense of rhythm by comparing individual performance sound sources of each part in the rhythm section.
8. A method for simultaneously playing multiple performance sound sources according to claim 1, wherein when playing multiple performance sound sources, a BPM value is selected in the case where one of the multiple performance sound sources and / or the listening position moves, and the selected performance sound source moves during playback.
9. A method for simultaneously playing multiple recorded performance sound sources according to claim 1, wherein, after simultaneous playback of the multiple performance sound sources is started based on the selected recorded performance sound source and the selected listening position, spatial sound processing is performed on the selected recorded performance sound source.
10. By setting the start and end positions of playback within the selected song, and / or the number of repetitions, when multiple performance audio sources are played simultaneously, the playback up to the set start and end positions is repeated according to the set number of repetitions, as described in claim 1. How to play multiple sound sources simultaneously.
11. A method for simultaneously playing multiple performance sound sources according to claim 1, wherein each of the multiple performance sound sources is recorded individually and independently using multi-track recording, and the recording is performed using binaural processing.
12. A method for simultaneously playing multiple (N) performance sound sources using a computer having an information processing unit, an information storage unit, an information input unit, and an information output unit, The aforementioned computer, when playing back a master performance audio source, includes the step of selecting a performance audio source to mute from among multiple performance audio sources, The stage of recording the performance audio while listening to the muted master performance audio, The process includes a stage in which the music is played back using the master performance audio and recorded performance audio through an additional interactive virtual session. The playback stage of the song is, On the operation screen, there is a step to set the space size when playing multiple audio sources, On the operation screen, there is a step to specify the listening position when playing the performance audio, The process involves calculating the three-dimensional coordinates of the listening position during music playback, The process involves calculating the distance and orientation between the master audio source and the recorded performance audio source, and the listening position. Using object-based spatial audio processing, three-dimensional positional information is added to the performance sound source information, and based on the selected space size and listening position, the sound playback processing of each performance sound source is performed according to the positional relationship between the position of each performance sound source and the listening position, thereby simulating and constructing a sound field within a specific space size. A method for simultaneously playing multiple performance sound sources, characterized by performing the steps of simultaneously playing a master performance sound source and a recorded performance sound source.
13. The method for simultaneously playing multiple performance sound sources according to claim 12, wherein when participating in the aforementioned additional-participation virtual session, the method determines whether to respond to invitations from existing participants and to accept or reject participation requests from those wishing to participate.
14. The method for simultaneously playing multiple performance sound sources according to claim 13, wherein, in the aforementioned additional-participation virtual session, each existing participant is determined to be either a Master who has the authority to set up the session and to reject participants, or a Member who has the authority only to reject the sharing of their own recorded performances.
15. The method for simultaneously playing multiple performance sound sources according to claim 13, wherein when participating as an additional virtual session, the recorded performance sound source is shared by transferring it to an external storage unit.
16. The method for simultaneously playing multiple performance audio sources according to claim 14, wherein the Master determines which combination of the performance audio source of one of the multiple performance audio sources will be used to create the collaborative song.
17. The method for simultaneously playing multiple performance sound sources according to claim 15, wherein the Master decides whether to limit the sharing of the collaborative song to the participants or to make it public to the users of the method for simultaneously playing multiple performance sound sources according to claim 1.
18. The method for simultaneously playing multiple performance sound sources according to claim 12, wherein if there is a recorded performance sound source that a participant has recorded while listening to the master performance sound source as accompaniment with stereo headphones, the next participant uses the method for simultaneously playing multiple performance sound sources according to claim 1 to perform stereophonic sound processing on the master performance sound source and the recorded performance sound source and use them as accompaniment performance sound sources.
19. A method for simultaneously playing multiple performance sound sources according to claim 12, wherein a performance sound source of the rhythm section, including piano, drums, and bass, is used as the master performance sound source, and the front section, consisting of guitar, saxophone, trumpet, and trombone, joins in.
20. A method for simultaneously playing multiple performance sound sources according to claim 16, which presents evaluation results based on the number of plays and / or "likes" on the released collaborative song.
21. A method for simultaneously playing multiple performance sound sources according to claim 12, wherein, in an additional-participation virtual session, the tempo, key, and spatial sound processing conditions of the accompanying performance sound source used as the basis for participating in the virtual session are set in common among all participants.
22. A system for simultaneous playback of multiple sound sources, comprising: a control unit for controlling the operation of the entire music playback device when multiple sound sources are played simultaneously; a memory for storing various control programs and data; a music information storage unit for storing music data; a timer for outputting various timing signals; a music analysis unit for acquiring information about the content of music data stored in the music information storage unit; a display control unit for displaying various information on a display unit; a performance control unit for playing back the performance information contained in the music data stored in the music information storage unit and outputting it to a speaker; an operation unit for inputting operation signals; a communication unit for sending and receiving signals to and from the outside; a display unit for displaying information related to the performance information; at least two speakers capable of stereophonically reproducing the performance information from the performance control unit; an operation device for inputting various operation information; and a bus for transferring various data. It has a localization control table that defines a localization method for localizing the content of the music and the corresponding sound image as spatial sound within a three-dimensional sound space for each of multiple performance sound sources. A system for simultaneously playing multiple performance sound sources, each having a program for simultaneously playing multiple performance sound sources, is provided with multiple simultaneous playback systems, the multiple simultaneous playback systems are connected by network communication, and the multiple performance sound source data played simultaneously in each of the multiple simultaneous playback systems are mutually accessible, thereby enabling each simultaneous playback system to record its own performance while utilizing the performance sound sources of other simultaneous playback systems, and is used for additional participation virtual sessions.