Method for generating reproduced data
The data output method enhances realism in automatic music performance by simulating a multi-musician environment based on user performance data, creating a more immersive experience for solo performers.
Patent Information
- Application Number
- JP2025131885
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-10-15
AI Technical Summary
Existing technologies for automatic music performance following a user's performance lack realism, failing to create an immersive experience for solo performers.
A data output method that includes acquiring performance data, identifying score performance positions, assigning position information to reproduced data, and outputting it with corresponding virtual positions and directions, simulating a performance environment with multiple musicians.
Enhances the realism of automatic performance by allowing solo performers to experience playing with others, providing a highly immersive musical experience.
Smart Images

Figure 2025157613000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for outputting data. [Background technology]
[0002] A technology has been proposed that identifies the playing position on the score of a given piece of music by analyzing sound data obtained when the user plays the piece. A technology has also been proposed that applies this technology to automatic performance, realizing automatic performance that follows the user's performance (for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-207615 Summary of the Invention [Problem to be solved by the invention]
[0004] By having the automatic performance follow the user's performance, it is possible to create the feeling that multiple people are playing together, even if only one person is playing. There is a demand for a more realistic experience for the user.
[0005] One of the objects of the present invention is to enhance the sense of realism given to the user in automatic processing that follows the user's performance. [Means for solving the problem]
[0006] According to one embodiment, a data output method is provided, which includes acquiring performance data generated by a performance operation, identifying a score performance position in a predetermined score based on the performance data, reproducing first data based on the score performance position, assigning first position information to the first data according to a first virtual position set corresponding to the first data, and outputting reproduced data including the first data to which the first position information has been assigned. [Effects of the Invention]
[0007] According to the present invention, it is possible to enhance the sense of realism given to the user in automatic processing that follows the user's performance. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating a system configuration according to a first embodiment. [Figure 2] 1 is a diagram illustrating the configuration of an electronic musical instrument according to a first embodiment. [Figure 3] FIG. 2 is a diagram illustrating the configuration of a data output device in the first embodiment. [Figure 4] FIG. 3 is a diagram illustrating position control data in the first embodiment. [Figure 5] 4A to 4C are diagrams illustrating position information and direction information in the first embodiment. [Figure 6] FIG. 2 is a diagram illustrating a configuration for realizing a performance follow-up function in the first embodiment. [Figure 7] FIG. 3 is a diagram illustrating a data output method in the first embodiment. [Figure 8] FIG. 10 is a diagram illustrating a configuration for realizing a performance follow-up function in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. The embodiment shown below is merely an example, and the present invention should not be construed as being limited to these embodiments. In the drawings referred to in the following embodiments, identical parts or parts having similar functions are designated by the same or similar symbols (symbols consisting of a number followed by A, B, etc.), and repeated explanations may be omitted. For clarity of explanation, the drawings may be illustrated schematically, with some components omitted from the drawings.
[0010] First Embodiment [overview] A data output device according to one embodiment of the present invention realizes automatic performance of a predetermined piece of music in response to a user's performance on an electronic musical instrument. Various instruments can be set as the target of automatic performance. If the electronic musical instrument played by the user is an electronic piano, the target of automatic performance may be an instrument other than the piano part, such as vocals, bass, drums, guitar, or a horn section. In this example, the data output device provides the user with the playback sound obtained by the automatic performance and an image of the instrument's performer (hereinafter sometimes referred to as a performer image). This data output device can give the user the feeling that they are playing together with other performers. A data output device and a system including the data output device will now be described.
[0011] [System Configuration] Fig. 1 is a diagram illustrating a system configuration in a first embodiment. The system shown in Fig. 1 includes a data output device 10 and a data management server 90 connected via a network NW such as the Internet. In this example, a head-mounted display 60 (hereinafter sometimes referred to as HMD 60) and an electronic musical instrument 80 are connected to the data output device 10. In this example, the data output device 10 is a computer such as a smartphone, tablet PC, laptop PC, or desktop PC. In this example, the electronic musical instrument 80 is an electronic keyboard device such as an electronic piano.
[0012] As described above, when a user plays a predetermined piece of music using the electronic musical instrument 80, the data output device 10 has a function (hereinafter referred to as a performance follow-up function) for executing an automatic performance that follows the performance and outputting data based on the automatic performance. A detailed description of the data output device 10 will be given later.
[0013] The data management server 90 includes a control unit 91, a storage unit 92, and a communication unit 98. The control unit 91 includes a processor such as a CPU and a storage unit such as RAM. The control unit 91 executes a program stored in the storage unit 92 using the CPU, thereby performing processing according to instructions written in the program. The storage unit 92 includes a storage unit such as a non-volatile memory and a hard disk drive. The communication unit 98 includes a communication module for connecting to a network NW and communicating with other devices. The data management server 90 provides music data to the data output device 10. The music data is data related to automatic performance, and details will be described later. If music data is provided to the data output device 10 by some other method, the data management server 90 need not exist.
[0014] In this example, the HMD 60 includes a control unit 61, a display unit 63, a behavior sensor 64, a sound emitting unit 67, an imaging unit 68, and an interface 69. The control unit 61 includes a CPU, RAM, and ROM, and controls each component in the HMD 60. The interface 69 includes a connection terminal for connecting to the data output device 10. The behavior sensor 64 includes, for example, an acceleration sensor, a gyro sensor, etc., and is a sensor that measures the behavior of the HMD 60, for example, changes in the orientation of the HMD 60. In this example, the measurement results by the behavior sensor 64 are provided to the data output device 10. This allows the data output device 10 to recognize the movement of the HMD 60. In other words, the data output device 10 can recognize the movement (head movement, etc.) of the user wearing the HMD 60. The user moves his / her head to receive information via the HMD 60. The user can input instructions to the data output device 10 using the operation unit. If the HMD 60 is provided with an operation unit, the user can also input instructions to the data output device 10 via the operation unit.
[0015] The imaging unit 68 includes an image sensor and captures an image in front of the HMD 60, i.e., in front of the user wearing the HMD 60, to generate imaging data. The display unit 63 includes a display that displays an image corresponding to video data. The video data is included, for example, in the playback data provided by the data output device 10. The display has a spectacle-like shape. The display may be semi-transparent, allowing the user wearing the HMD 60 to view the outside. If the display is non-transparent, the area captured by the imaging unit 68 may be displayed on the display superimposed on the video data. This allows the user to view the outside of the HMD 60 through the display. The sound emitting unit 67 is, for example, headphones or the like, and includes a vibrator that converts a sound signal corresponding to sound data into air vibrations and provides sound to the user wearing the HMD 60. The sound data is included, for example, in the playback data provided by the data output device 10.
[0016] [Electronic Instruments] 2 is a diagram illustrating the configuration of an electronic musical instrument according to the first embodiment. In this example, the electronic musical instrument 80 is an electronic keyboard device such as an electronic piano, and includes performance controls 84, a sound source unit 85, a speaker 87, and an interface 89. The performance controls 84 include multiple keys, and output signals to the sound source unit 85 in response to the operation of each key.
[0017] The sound source unit 85 includes a DSP (Digital Signal Processor) and generates sound data including sound waveform signals in response to operation signals. The operation signals correspond to signals output from the performance operators 84. The sound source unit 85 converts the operation signals into sequence data (hereinafter referred to as operation data) in a predetermined format for controlling the generation of sound (hereinafter referred to as sound generation), and outputs the converted data to the interface 89. In this example, the predetermined format is MIDI. This allows the electronic musical instrument 80 to transmit operation data corresponding to performance operations on the performance operators 84 to the data output device 10. The operation data is information that specifies the content of sound generation, and is output sequentially as sound generation control information such as note on, note off, and note number. The sound source unit 85 provides sound data to the interface 89, or can provide the sound data to the speaker 87 instead of providing the sound data to the interface 89.
[0018] The speaker 87 can convert a sound waveform signal corresponding to sound data provided from the sound source unit 85 into air vibrations and provide the sound to the user. The sound data may be provided to the speaker 87 from the data output device 10 via an interface 89. The interface 89 includes a module for transmitting and receiving data to and from an external device wirelessly or via a wired connection. In this example, the interface 89 is connected to the data output device 10 via a wired connection and transmits operation data and sound data generated in the sound source unit 85 to the data output device 10. These data may be received from the data output device 10.
[0019] [Data output device] 3 is a diagram illustrating the configuration of a data output device according to the first embodiment. The data output device 10 includes a control unit 11, a storage unit 12, a display unit 13, an operation unit 14, a speaker 17, a communication unit 18, and an interface 19. The control unit 11 is an example of a computer equipped with a processor such as a CPU and a storage unit such as a RAM. The control unit 11 executes a program 12a stored in the storage unit 12 using the CPU (processor), causing the data output device 10 to realize functions for executing various processes. The functions realized in the data output device 10 include a performance follow-up function, which will be described later.
[0020] The storage unit 12 is a storage device such as a non-volatile memory or a hard disk drive. The storage unit 12 stores various data, such as a program 12a executed by the control unit 11 and music data 12b required for executing the program 12a. The program 12a is downloaded from the data management server 90 or another server via the network NW and stored in the storage unit 12, thereby being installed in the data output device 10. The program 12a may be provided in a state recorded on a non-transitory computer-readable recording medium (e.g., a magnetic recording medium, an optical recording medium, a magneto-optical recording medium, a semiconductor memory, etc.). In this case, the data output device 10 only needs to be equipped with a device for reading the recording medium. The storage unit 12 can also be considered an example of a recording medium.
[0021] Similarly, the song data 12b may be downloaded from the data management server 90 or another server via the network NW and stored in the storage unit 12, or may be provided in a state recorded on a non-transitory computer-readable recording medium. The song data 12b is data stored in the storage unit 12 for each song, and includes setting data 120, background data 127, and musical score data 129. Details of the song data 12b will be described later.
[0022] The display unit 13 is a display having a display area that displays various screens under the control of the control unit 11. The operation unit 14 is an operation device that outputs signals to the control unit 11 in response to user operations. The speaker 17 generates sound by amplifying and outputting sound data supplied from the control unit 11. The communication unit 18 is a communication module that connects to the network NW under the control of the control unit 11 and communicates with other devices connected to the network NW, such as a data management server 90. The interface 19 includes a module for communicating with external devices via wireless communication, such as infrared communication or short-range wireless communication, or via wired communication. In this example, the external devices include an electronic musical instrument 80 and an HMD 60. The interface 19 is used for communication without going through the network NW.
[0023] [Song data] Next, the song data 12b will be described. The song data 12b is data stored in the storage unit 12 for each song, and includes setting data 120, background data 127, and musical score data 129. In this example, the song data 12b includes data for reproducing a predetermined live performance in accordance with the user's performance. The data for reproducing this live performance includes information regarding the form of the venue where the live performance was held, multiple instruments (performance parts), performers of each performance part, and the positions of the performers. One of the multiple performance parts is identified as the user's performance part. In this example, four performance parts (vocal part, piano part, bass part, and drum part) are defined. Of the four performance parts, the user's performance part is identified as the piano part.
[0024] The musical score data 129 corresponds to the musical score of the part played by the user. In this example, the musical score data 129 is data representing the musical score of the piano part of a piece of music, and is data written in a predetermined format such as MIDI. That is, the musical score data 129 includes time information and sound generation control information associated with the time information. The sound generation control information defines the content of sound generation at each time, and is represented by information including timing information such as note-on, note-off, and note number, as well as pitch information. The sound generation control information can also include character information, so that vocal sounds in the vocal part can also be included in the sound generation. The time information is information indicating playback timing relative to the start of the piece of music, and is represented by information such as delta time and tempo. The time information can also be considered information for identifying a position in the data. The musical score data 129 can also be considered data defining musical sound control information in a chronological order.
[0025] Background data 127 corresponds to the form of the venue where the live performance took place. The background data 127 includes data indicating the structure of the stage, the structure of the audience seats, the structure of the room, etc. For example, the background data 127 includes coordinate data specifying the position of each structure and image data for reproducing the space within the venue. The coordinate data is defined as coordinates in a predetermined virtual space. The background data 127 can also be said to include data for forming a background image that simulates a venue in the virtual space.
[0026] The setting data 120 corresponds to each performance part in the music piece. Therefore, the music piece data 12b may include multiple setting data 120. In this example, the music piece data 12b includes setting data 120 corresponding to three parts other than the piano part associated with the musical score data. Specifically, the three parts are a vocal part, a bass part, and a drum part. In other words, setting data 120 exists corresponding to the performers of each part. Setting data other than the performers may also exist; for example, the music piece data 12b may include setting data 120 corresponding to the audience. Even the audience, due to their movements, cheers, and the like during a live performance, can be treated as equivalent to one performance part.
[0027] The setting data 120 includes sound generation control data 121, video control data 123, and position control data 125. The sound generation control data 121 is data for reproducing sound data corresponding to a performance part, and is data written in a predetermined format such as MIDI. That is, the sound generation control data 121 includes time information and sound generation control information, just like the musical score data 129. In this example, the sound generation control data 121 and the musical score data 129 are similar data except for the performance part. The sound generation control data 121 can also be said to be data that defines musical sound control information in a time series.
[0028] The moving image control data 123 is data for playing moving image data, and includes time information and image control information associated with the time information. The image control information defines a performer image at each time. As described above, the performer image is an image that resembles a performer corresponding to a performance part. In this example, the moving image data to be played includes a performer image corresponding to a performer performing a performance related to the performance part. The moving image control data 123 can also be said to be data that defines image control information in chronological order.
[0029] FIG. 4 is a diagram illustrating position control data in the first embodiment. The position control data 125 includes information indicating the position of a performer corresponding to a performance part (hereinafter referred to as position information) and information indicating the direction of the performer (the front direction of the performer) (hereinafter referred to as direction information). In the position control data 125, the position information and direction information are associated with time information. The position information is defined as coordinates in a virtual space used in the background data 127. The direction information is defined as an angle based on a predetermined direction in this virtual space. As shown in FIG. 4, as the time information progresses from t1, t2, . . ., the position information changes from P1, P2, . . ., and the direction information changes from D1, D2, . . . The position control data 125 can also be said to be data that defines the position information and direction information in a time series.
[0030] FIG. 5 is a diagram illustrating position information and direction information in the first embodiment. As described above, in this example, there are four setting data 120 corresponding to three performance parts and the audience. Therefore, there are also position control data 125 corresponding to three performance parts and the audience. FIG. 5 is an example of a predetermined virtual space viewed from above. In FIG. 5, the wall RM of the venue and the stage ST are defined by background data 127. For each performer corresponding to each setting data, a virtual position corresponding to the position information defined in the position control data and a virtual direction corresponding to the direction information are set.
[0031] The virtual position and virtual direction of the vocal part performer are determined by the position information C1p and the direction information The virtual position and virtual direction of the bass part performer are set to position information C2p and direction information C2d. The virtual position and virtual direction of the drum part performer are set to position information C3p and direction information C3d. The virtual position and virtual direction of the audience are set to position information C4p and direction information C4d. Here, each performer is located on the stage ST. The audience is located in an area other than the stage ST (audience seats). The example shown in FIG. 5 is a situation at a specific time. Therefore, the virtual positions and virtual directions of the performers and audience may change over time.
[0032] In FIG. 5, the virtual position and virtual direction of the player of the piano part corresponding to the user are set as position information Pp and direction information Pd. This virtual position and virtual direction change according to the movement of the HMD 60 (measurement results of the behavior sensor 64) described above. For example, when the user wearing the HMD 60 changes the direction of their head, the direction information Pd changes corresponding to the direction of their head. When the user wearing the HMD 60 moves, the position information Pp changes corresponding to the movement of the user. The position information Pp and direction information Pd may be changed by operating the operation unit 14 or inputting an instruction from the user. The initial values of the position information Pp and direction information Pd may be set in advance in the music data 12b.
[0033] As will be described later, when a video is provided to the user via the HMD 60, the user can visually recognize other performers positioned in the virtual space at the positions and orientations (position information Pp, direction information Pd) shown in Fig. 5. For example, the performer of the drum part (position information C3p, direction information C3d) is visually recognized by the user as a performer image playing on the left side of the front direction (position indicated by vector V3) facing right.
[0034] [Performance tracking function] Next, the performance follow-up function realized by the control unit 11 executing the program 12a will be described.
[0035] 6 is a diagram illustrating the configuration for realizing the performance follow-up function in the first embodiment. The performance follow-up function 100 includes a performance data acquisition unit 110, a performance sound acquisition unit 119, a performance position identification unit 130, a signal processing unit 150, a reference value acquisition unit 164, and a data output unit 190. The configuration for realizing the performance follow-up function 100 is not limited to being realized by executing a program, and at least a part of the configuration may be realized by hardware.
[0036] The performance data acquisition unit 110 acquires performance data. In this example, the performance data corresponds to operation data provided from the electronic musical instrument 80. The performance sound acquisition unit 119 acquires sound data (performance sound data) corresponding to performance sounds provided from the electronic musical instrument 80. The reference value acquisition unit 164 acquires a reference value corresponding to the user's performance part. The reference value includes a reference position and a reference direction. The reference position corresponds to the above-mentioned position information Pp. The reference direction corresponds to the direction information Pd. As described above, the control unit 11 changes the position information Pp and the direction information Pd from their initial values set in advance in accordance with the movement of the HMD 60 (measurement results of the behavior sensor 64). The reference value may be set in advance. At least one of the reference position and the reference direction may be associated with time information, similar to the position control data 125. In this case, the reference value acquisition unit 164 may acquire a reference value associated with time information based on the correspondence between the musical score performance position and time information, which will be described later.
[0037] The performance position identification unit 130 refers to the musical score data 129 and identifies a musical score performance position corresponding to the performance data sequentially acquired by the performance data acquisition unit 110. The performance position identification unit 130 compares the history of the sound generation control information in the performance data (i.e., a set of time information and sound generation control information corresponding to the timing at which the operation data was acquired) with the set of time information and sound generation control information in the musical score data 129, and analyzes the correspondence between them by a predetermined matching process. The predetermined matching process may be a known matching process using a statistical estimation model, such as DP matching, a hidden Markov model, or matching using machine learning. The performance position on the score may be identified at a preset speed for a predetermined time after the start of performance.
[0038] From this correspondence, the performance position identification unit 130 identifies a score performance position that corresponds to the performance on the electronic musical instrument 80. The score performance position indicates a position currently being played in the score in the musical score data 129, and is identified, for example, as time information in the musical score data 129. The performance position identification unit 130 sequentially acquires performance data as the electronic musical instrument 80 is played, and sequentially identifies score performance positions that correspond to the acquired performance data. The performance position identification unit 130 provides the identified score performance positions to the signal processing unit 150.
[0039] The signal processing unit 150 includes data generating units 170-1, ..., 170-n (referred to as data generating units 170 when no particular distinction is made between them). The data generating units 170 are set in correspondence with the setting data 120. As in the example above, if the song data 12b includes three performance parts (vocal part, bass part, drum part) and four setting data 120 corresponding to the audience, the signal processing unit 150 includes four data generating units 170 (170-1 to 170-4). In this way, the data generating units 170 and the setting data 120 are associated with each other via the performance parts.
[0040] The data generation unit 170 includes a playback unit 171 and an assignment unit 173. The playback unit 171 acquires the sound generation control data 121 and the video control data 123 from the associated setting data 120. The assignment unit 173 acquires the position control data 125 from the associated setting data 120.
[0041] The playback unit 171 plays back sound data and video data based on the score performance position provided by the performance position identification unit 130. The playback unit 171 references the sound generation control data 121, reads out sound generation control information corresponding to the time information identified by the score performance position, and plays back the sound data. The playback unit 171 can also be said to have a sound source unit that plays back sound data based on the sound generation control data 121. This sound data is data corresponding to the performance sound of the associated performance part. In the case of a vocal part, the sound data may be data corresponding to a singing sound generated using at least character information and pitch information. The playback unit 171 references the video control data 123, reads out image control information corresponding to the time information identified by the score performance position, and plays back the video data. This video data is data corresponding to an image of the performer of the associated performance part, i.e., a performer image.
[0042] The assigning unit 173 assigns position information and direction information to the sound data and video data played by the playback unit 171. The assigning unit 173 references the position control data 125 and reads out the position information and direction information corresponding to the time information identified by the performance position on the musical score. The assigning unit 173 corrects the reference value acquired by the reference value acquisition unit 164, i.e., the read position information and direction information, using the position information Pp and direction information Pd. Specifically, the assigning unit 173 converts the read position information and direction information into relative information expressed in a coordinate system based on the position information Pp and direction information Pd. The assigning unit 173 assigns the corrected position information and direction information, i.e., the relative information, to the sound data and video data.
[0043] In the example shown in Fig. 5, the virtual position and virtual direction of the performer of the piano part corresponding to the user are the reference values. Therefore, among the relative information regarding the performers of each performance part and the audience, the portion related to position information includes information expressed by vectors V1 to V4. Among the relative information, the portion related to direction information includes direction information C1d, C2d, C3d, C4 with respect to direction information Pd. This corresponds to the direction of d (hereinafter referred to as the relative direction).
[0044] Adding relative information to sound data corresponds to performing signal processing on the sound signals of the left channel (Lch) and right channel (Rch) included in the sound data so that sound images are localized at predetermined positions in virtual space. The predetermined positions are positions defined by vectors included in the relative information. In the example shown in FIG. 5, for example, the performance sound of the drum part is localized at a position defined by vector V3. At this time, predetermined filter processing may be performed, such as using HRTF (Head Related Transfer Function) technology. The adding unit 173 may perform signal processing on the sound signal by referring to the background data 127 to add reverberation due to the structure of the room, etc. At this time, the adding unit 173 may add directionality so that sound is output from the sound image toward the relative direction included in the relative information.
[0045] Adding relative information to video data corresponds to performing image processing on the performer image included in the video data so that it is placed at a predetermined position in the virtual space and faces a predetermined direction. The predetermined position is the position where the above-mentioned sound image is localized. The predetermined direction corresponds to the relative direction included in the relative information. In the example shown in FIG. 5, the performer image of the drum part, for example, is visually recognized by the user wearing the HMD 60 as facing to the right (more precisely, forward and to the right) at the position defined by vector V3.
[0046] In this example, data generation unit 170-1 outputs moving image data and sound data with position information assigned for the vocal part. Data generation unit 170-2 outputs moving image data and sound data with position information assigned for the bass part. Data generation unit 170-3 outputs moving image data and sound data with position information assigned for the drum part. Data generation unit 170-4 outputs moving image data and sound data with position information assigned for the audience.
[0047] The data output unit 190 synthesizes the video data and audio data output from the data generation units 170-1, ..., 170-n and outputs the synthesized data as playback data. By supplying this playback data to the HMD 60, a user wearing the HMD 60 can view images of the vocal, bass, and drum performers at their respective positions and hear the corresponding performance sounds from those positions. This enhances the sense of realism imparted to the user. Furthermore, in this example, the user can also view the audience and hear their cheers. Since the video data and audio data included in the playback data follow the user's performance, the progression of the sound for each performance part and the movement of the performer images change according to the speed of the user's performance. In other words, musical performances and singing that follow the performance are realized in a virtual environment around the instrument played by the user. As a result, the user can experience the sensation of multiple people playing together, even when playing alone. This provides the user with a highly realistic customer experience.
[0048] The data output unit 190 may refer to the background data 127 and include in the video data a background image that resembles a venue in the virtual space. This allows the user to visually recognize performer images positioned in the positional relationship shown in FIG. 5 as they perform on the stage ST. The data output unit 190 may output playback data that further combines the performance sound data acquired by the performance sound acquisition unit 119. This allows the user to listen to the performance sounds made by the user through the HMD 60. This concludes the description of the performance follow-up function.
[0049] [Data output method] Next, a data output method executed in the performance follow-up function 100 will be described. The data output method described here begins when the program 12a is executed.
[0050] 7 is a diagram illustrating a data output method in the first embodiment. The control unit 11 acquires performance data provided sequentially (step S101) and identifies the performance position on the score (step S103). The control unit 11 plays back video data and sound data based on the performance position on the score (step S105), assigns position information to the played back video data and sound data (step S107), and outputs them as playback data (step S109). The control unit 11 repeats the processes from step S101 to step S109 until an instruction to end the process is input (step S111; No), and when an instruction to end the process is input (step S111; Yes), the control unit 11 ends the process.
[0051] Second Embodiment In the first embodiment, an example was described in which video data and sound data are reproduced in accordance with the performance of one user, but they may also be reproduced in accordance with the performances of multiple users. In the second embodiment, an example will be described in which video data and sound data are reproduced in accordance with the performances of two users.
[0052] 8 is a diagram illustrating a configuration for realizing the follow-performance function in the second embodiment. The follow-performance function 100A in the second embodiment has a configuration in which the two follow-performance functions 100 in the first embodiment run in parallel, with the background data 127, musical score data 129, and performance position identification unit 130 being shared by both. The two follow-performance functions 100 are provided corresponding to the first and second users.
[0053] The performance data acquisition unit 110A-1 acquires first performance data related to a first user. The first performance data is, for example, operation data output from the electronic musical instrument 80 played by the first user. The performance data acquisition unit 110A-2 acquires second performance data related to a second user. The second performance data is, for example, operation data output from the electronic musical instrument 80 played by the second user.
[0054] The performance position identification unit 130A identifies a score performance position by comparing the history of sound generation control information in either the first performance data or the second performance data with the sound generation control information in the score data 129. Whether the first performance data or the second performance data is selected is determined based on the first performance data and the second performance data. For example, the performance position identification unit 130A performs both a matching process on the first performance data and a matching process on the second performance data, and adopts the score performance position identified by the process with the higher calculation accuracy. The calculation accuracy may be determined, for example, by using an index indicating the matching error in the calculation results.
[0055] As another example, the performance position identification unit 130A determines whether to use the score performance position obtained from the first performance data or the score performance position obtained from the second performance data, depending on the position in the music piece identified by the score performance position. In this case, the performance period of the music piece may be divided into multiple periods in the score data 129, and priorities may be assigned to the performance parts for each period. The performance position identification unit 130A refers to the score data 129 and identifies the score performance position using the performance data corresponding to the performance part with the higher priority.
[0056] The signal processing units 150A-1 and 150A-2 have the same functions as the signal processing unit 150 in the first embodiment, and correspond to the first and second users, respectively. The signal processing unit 150A-1 plays back video data and sound data using the score performance position identified by the performance position identification unit 130A and the reference value related to the first user acquired by the reference value acquisition unit 164A-1. The data generating unit 170 may or may not be present. The data generating unit 170 relating to the performance part of the second user may not play sound data, but may play video data. When playing video data, the reference value relating to the second user acquired by the reference value acquiring unit 164A-2 may be used instead of using the position control data 125.
[0057] The signal processing unit 150A-2 plays back video data and sound data using the score performance position identified by the performance position identification unit 130A and the reference value related to the second user acquired by the reference value acquisition unit 164A-2. The signal processing unit 150A-2 may or may not include a data generation unit 170 related to the first user's performance part. The data generation unit 170 related to the first user's performance part may not play sound data, but may play video data. When playing back video data, the reference value related to the first user acquired by the reference value acquisition unit 164A-1 may be used instead of using the position control data 125.
[0058] The data output unit 190A-1 combines the video data and sound data output from the signal processing unit 150A-1 and outputs the combined data as playback data. This playback data is provided to the HMD 60 of the first user. The data output unit 190A-1 may reference background data 127 to include a background image simulating a venue in a virtual space in the video data. The data output unit 190A-1 may output playback data that further combines performance sound data acquired by the performance sound acquisition units 119A-1 and 119A-2. The sound data acquired by the performance sound acquisition unit 119A-1 is, for example, sound data output from the electronic musical instrument 80 played by the first user. The sound data acquired by the performance sound acquisition unit 119A-2 is, for example, sound data output from the electronic musical instrument 80 played by the second user. Relative information corresponding to the reference value of the second user relative to the reference value of the first user, or relative information attached to video data relating to the performance part of the second user, may be attached to the sound data acquired by the performance sound acquisition unit 119A-2 so that a sound image is localized at a predetermined position.
[0059] The data output unit 190A-2 combines the video data and sound data output from the signal processing unit 150A-2 and outputs the combined data as playback data. This playback data is provided to the HMD 60 of the first user. The data output unit 190A-2 may reference background data 127 to include a background image simulating a venue in the virtual space in the video data. The data output unit 190A-2 may output playback data that further combines the performance sound data acquired by the performance sound acquisition units 119A-1 and 119A-2. Relative information corresponding to the reference value of the first user relative to the reference value of the second user, or relative information assigned to the video data related to the first user's performance part, may be assigned to the sound data acquired by the performance sound acquisition unit 119A-1, so that a sound image is localized at a predetermined position.
[0060] In this way, the performance follow-up function 100A in the second embodiment can enhance the sense of realism given to the user even when two performance parts are played by the user.
[0061] <Modification> The present invention is not limited to the above-described embodiments, and includes various other modified examples. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and are not necessarily limited to those having all of the described configurations. Some modified examples will be described below. The first embodiment will be described as a modified example, but the modified examples can also be applied to other embodiments. Multiple modified examples can also be combined and applied to each embodiment.
[0062] (1) The video data and audio data included in the playback data are The video data and the audio data may be provided to different devices. For example, the video data may be provided to the HMD 60, and the audio data may be provided to a speaker device other than the HMD 60. This speaker device may be, for example, the speaker 87 of the electronic musical instrument 80. When the audio data is provided to a speaker device, the providing unit 173 may perform signal processing that takes into account the speaker device provided. For example, in the case of the speaker 87 of the electronic musical instrument 80, the left and right speaker units are fixed to the electronic musical instrument 80, and the position of the performer's ears is roughly estimated. In such a case, signal processing may be performed on the audio data to localize the sound image using crosstalk cancellation technology based on the two speaker units and the estimated positions of the performer's right and left ears. In this case, the shape of the room in which the electronic musical instrument 80 is installed may be acquired, and signal processing may be performed on the audio data to cancel the sound field caused by the shape of the room. The shape of the room may be acquired by a known method such as a method using sound reflection or a method using imaging.
[0063] (2) At least one of the video data and the sound data included in the playback data may not be present. That is, at least one of the video data and the sound data may follow the user's performance as an automatic process.
[0064] (3) The assigning unit 173 does not need to assign position information and direction information to at least one of the video data and the audio data.
[0065] (4) The functions of the data output device 10 and the functions of the electronic musical instrument 80 may be included in a single device. For example, the data output device 10 may be incorporated as a function of the electronic musical instrument 80. A portion of the configuration of the electronic musical instrument 80 may be included in the data output device 10, or a portion of the configuration of the data output device 10 may be included in the electronic musical instrument 80. For example, the configuration of the electronic musical instrument 80 other than the performance operators 84 may be included in the data output device 10. In this case, the data output device 10 may generate sound data from the acquired operation data using a sound source unit.
[0066] (5) The musical score data 129 may be included in the music data 12b in the same format as the setting data 120. In this case, by setting a performance part of the user in the data output device 10, the sound generation control data 121 included in the setting data 120 corresponding to that performance part may be used as the musical score data 129.
[0067] (6) The position information and direction information are determined as virtual positions and virtual directions in a virtual space, but may also be determined as virtual positions and virtual directions on a virtual plane, as long as they are determined as information defined in a virtual area.
[0068] (7) The position control data 125 in the setting data 120 does not have to include direction information or time information.
[0069] (8) The assigning unit 173 does not need to change the position information assigned to the sound data based on the position control data 125. For example, it may be fixed at an initial value. In this way, the performer image can be moved while assuming a situation in which sound is being emitted from a specific speaker (a situation in which the sound image is fixed). This may result in the sound image being localized at a position different from the position of the performer image. The information included in the position control data 125 may be provided separately for video data and for sound data, so that the performer image and the sound image can be controlled to different positions.
[0070] (9) The video data included in the playback data may be still image data.
[0071] (10) The performance data acquired by the performance data acquisition unit 110 may be sound data (performance sound data) instead of operation data. If the performance data is sound data, the performance position identification unit 130 performs a known matching process by comparing the sound data, which is the performance data, with sound data generated based on the musical score data 129. Through this process, the performance position identification unit 130 can identify the musical score performance position corresponding to the sequentially acquired performance data. The musical score data 129 may also be sound data. In this case, time information is associated with each part of the sound data.
[0072] (11) The sound generation control data 121 included in the music data 12b may be sound data. In this case, time information is associated with each part of the sound data. In the case of sound data for a vocal part, the sound data includes singing sounds. When the playback unit 171 reads out this sound data based on the score performance position, the sound data may be read out based on the relationship between the score performance position and time information, and the pitch may be adjusted according to the readout speed. The pitch may be adjusted, for example, so that it becomes the pitch when the sound data is read out at a predetermined readout speed.
[0073] (12) The control unit 11 may record the playback data output from the data output unit 190 on a recording medium or the like. The control unit 11 may generate recording data for outputting the playback data and record the data on the recording medium. The recording medium may be the storage unit 12 or a computer-readable storage medium connected as an external device. The recording data may be transmitted to a server device connected via the network NW. For example, the recording data may be transmitted to the data management server 90 and stored in the storage unit 92. The recording data may include video data and audio data, or may include the setting data 120 and time-series information on the performance position on the musical score. In the latter case, playback data may be generated from the recording data by functions corresponding to the signal processing unit 150 and the data output unit 190.
[0074] (13) The playing position identifying unit 130 may identify a score playing position during a portion of a musical piece, regardless of the performance data acquired by the performance data acquiring unit 110. In this case, the score data 129 may specify a progression speed of the score playing position to be identified during the portion of the musical piece. The playing position identifying unit 130 may identify the score playing position so that the score playing position is changed at the specified progression speed during this period.
[0075] (14) Of the setting data 120 included in the music piece data 12b, the setting data 120 usable in the follow-up performance function 100 may be restricted by the user. In this case, the data output device 10 may implement the follow-up performance function 100 on the premise of inputting a user ID. The restricted setting data 120 may be changed depending on the user ID. For example, if the user ID is a specific ID, the control unit 11 may perform control so that the setting data 120 related to the vocal part cannot be used in the follow-up performance function 100. The relationship between the ID and the restricted data may be registered in the data management server 90. In this case, when providing the music piece data 12b to the data output device 10, the data management server 90 may prevent the unusable setting data 120 from being included in the music piece data 12b.
[0076] The above is the explanation regarding the modified example.
[0077] As described above, according to one embodiment of the present invention, a method for playing a musical score includes acquiring performance data generated by a performance operation, identifying a musical score performance position in a predetermined musical score based on the performance data, reproducing first data based on the musical score performance position, and assigning first position information to the first data in accordance with a first virtual position set in correspondence with the first data. and outputting reproduced data including the first data to which the first position information is assigned.
[0078] The first virtual position may be set to correspond to the score playing position.
[0079] The first data may include sound data.
[0080] The sound data may include singing sounds.
[0081] The singing sound may be generated based on character information and pitch information.
[0082] Adding the first position information to the first data may include performing signal processing on the sound data to localize a sound image.
[0083] The first data may include video data.
[0084] The first position information corresponding to the first virtual position may include relative information of the first virtual position with respect to a set reference position and a set reference direction.
[0085] The method may include changing at least one of the reference position and the reference direction based on an instruction input by a user.
[0086] At least one of the reference position and the reference direction may be set in accordance with the musical score playing position.
[0087] The method may include adding, to the first data, first direction information corresponding to a first virtual direction set in correspondence with the first data.
[0088] The method may include reproducing second data based on the score playing position, and assigning second position information to the second data according to a second virtual position set corresponding to the second data. The reproduced data may include the first data assigned with the first position information and the second data assigned with the second position information.
[0089] The reproduction data may include performance sound data corresponding to the performance operation.
[0090] The method may include generating recording data for outputting the reproduced data.
[0091] The acquiring of the performance data may include acquiring first performance data generated by performing a performance operation on at least a first part and second performance data generated by performing a performance operation on a second part. The acquiring of the performance data may further include selecting one of the first performance data and the second performance data based on the first performance data and the second performance data. The score performance position may be identified based on the selected first performance data or the selected second performance data.
[0092] The performance data may include performance sound data corresponding to the performance operation.
[0093] The performance data may include operation data corresponding to the performance operation.
[0094] A program for causing a processor to execute any of the data output methods described above may be provided.
[0095] A data output device may be provided which includes a processor for executing the program described above.
[0096] The instrument may include a sound source unit that generates sound data in response to the performance operation.
[0097] An electronic musical instrument may be provided that includes the data output device described above and a performance operator for inputting the performance operation. [Explanation of symbols]
[0098] 10: Data output device, 11: Control unit, 12: Memory unit, 12a: Program, 12b: Music data, 13: Display unit, 14: Operation unit, 17: Speaker, 18: Communication unit, 19: Interface, 60: Head-mounted display, 61: Control unit, 63: Display unit, 64: Behavior sensor, 67: Sound emission unit, 68: Imaging unit, 69: Interface, 80: Electronic musical instrument, 84: Performance operator, 85: Sound source unit, 87: Speaker, 89: Interface, 90: Data management server, 91: Control unit, 92: Memory unit, 98: Communication unit, 100, 100A: Performance Tracking function, 110, 110A-1, 110A-2: performance data acquisition unit, 119, 119A-1, 119A-2: performance sound acquisition unit, 120: setting data, 121: sound generation control data, 123: video control data, 125: position control data, 127: background data, 129: musical score data, 130, 130A: performance position identification unit, 150, 150A-1, 150A-2: signal processing unit, 164, 164A-1, 164A-2: reference value acquisition unit, 170: data generation unit, 171: playback unit, 173: assignment unit, 190, 190A-1, 190A-2: data output unit
Claims
[Claim 1] Acquiring performance data generated by a performance operation; Identifying a musical score performance position in a predetermined musical score based on the performance data; playing first data including sound data corresponding to a performance sound of a preset first performance part or video data corresponding to an image of a performer of the first performance part based on the performance position of the score; assigning first position information corresponding to a first virtual position set in correspondence with the first data to the first data; outputting reproduction data including the first data to which the first position information is assigned; Including, the first virtual position is further set corresponding to the musical score playing position; Data output method.
Citation Information
Patent Citations
Visual display method of music play system and recording medium for recording visual display program of play system
JP1999352960A
Karaoke (orchestration without lyrics) device integrated with speaker
JP2000059880A
Information processor and method therefor, recording medium, program, and information processing system
JP2006085045A
Musical score position estimating device, musical score position estimating method and musical score position estimating robot
JP2011039511A
Sound field correction apparatus
JP2012014032A