Information processing system, moving-image editing method, and program

The information processing system automatically synchronizes video content with musical structure by analyzing music data to specify phrase periods, reducing user burden in generating edited videos for musical performances.

WO2025159007A1PCT designated stage Publication Date: 2025-07-31YAMAHA CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/001262
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-24
Filing Date
2025-01-17
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing techniques for generating videos of musical performances require excessive manual editing by users to synchronize video sources with music tempo, leading to a high user burden.

Method used

An information processing system that includes a music analysis unit to specify phrase periods in music data and a video editing unit to generate edited videos based on these periods, automatically synchronizing video content with musical structure.

Benefits of technology

Generates edited videos suitable for musical performances without requiring extensive user intervention, ensuring synchronization and focus on key performance aspects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025001262_31072025_PF_FP_ABST
    Figure JP2025001262_31072025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing system 100 includes: a music analysis unit 51 that identifies the phrase periods of a musical composition by analyzing performance data D representing the musical composition; and a moving-image editing unit 52 that generates an edited moving-image V by editing, according to the phrase periods, one or more recorded moving-images Ya showing a performance of the musical composition.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, video editing method and program

[0001] The present disclosure relates to a technique for generating a video that represents a performance by a musician.

[0002] Techniques for generating video images showing a performer performing a piece of music have been proposed. For example, Patent Document 1 discloses a configuration for switching between video sources capturing a performer at a timing that corresponds to the tempo at which the performer performs the piece of music.

[0003] Publication No. 2020-106753

[0004] However, in the technology of Patent Document 1, the timing of switching between multiple video sources is simply controlled according to the tempo of the performance. Therefore, in order to generate a video suitable for a musical performance, the user must manually edit the combined video, which places an excessive burden on the user. In consideration of the above circumstances, one aspect of the present disclosure aims to generate an edited video suitable for a musical performance without placing an excessive burden on the user.

[0005] In order to solve the above problems, an information processing system according to one aspect of the present disclosure includes a music analysis unit that identifies phrase periods of a song by analyzing music data representing the song, and a video editing unit that generates an edited video by editing one or more videos representing a performance of the song according to the phrase periods.

[0006] A video editing method according to one aspect of the present disclosure involves analyzing music data representing a song to identify music information about the song and phrase periods of the song, and generating an edited video by editing one or more videos representing a performance of the song according to the music information and the phrase periods.

[0007] A program according to one aspect of the present disclosure causes a computer system to function as a music analysis unit that identifies music information about a song and phrase periods of the song by analyzing music data representing the song, and a video editing unit that generates an edited video by editing one or more videos representing a performance of the song according to the music information and the phrase periods.

[0008] 1 is a block diagram illustrating the configuration of an information processing system in a first embodiment. FIG. 1 is a block diagram illustrating the configuration of a terminal device. FIG. 1 is a block diagram illustrating the configuration of a keyboard instrument. FIG. 2 is a schematic diagram of a keyboard video. FIG. 3 is a schematic diagram of a performance video. FIG. 4 is a schematic diagram of a pedal video. FIG. 5 is a schematic diagram of a setting screen. FIG. 6 is a schematic diagram of a setting screen when the performance scene is "recital". FIG. 7 is a schematic diagram of a setting screen when the performance scene is "street performance". FIG. 8 is a schematic diagram of a setting screen when the performance scene is "music class". FIG. 9 is a schematic diagram of performance data. FIG. 10 is an explanatory diagram of an edited video. FIG. 11 is a block diagram illustrating the functional configuration of a terminal device. FIG. 12 is a schematic diagram of a reference setting window. FIG. 13 is a schematic diagram of basic data. FIG. 14 is a schematic diagram of control data. FIG. 15 is an explanatory diagram of facial expression attention processing. FIG. 16 is an explanatory diagram of ascending follow-up processing and descending follow-up processing. FIG. 17 is an explanatory diagram of right hand attention processing and left hand attention processing. FIG. 18 is an explanatory diagram of pedal attention processing. FIG. 19 is a flowchart of video editing processing. FIG. 19 is an explanatory diagram of an edited video in a third embodiment. FIG. 19 is an explanatory diagram of an edited video in a fourth embodiment. Fig. 10 is a schematic diagram of a display image in a modified example. Fig. 11 is a schematic diagram of an edited moving image in a modified example. Fig. 12 is a schematic diagram of an edited moving image in a modified example. Fig. 13 is a schematic diagram of an edited moving image in a modified example. Fig. 14 is a schematic diagram of an edited moving image in a modified example.

[0009] A: First Embodiment Fig. 1 is a block diagram illustrating the configuration of an information processing system 100 according to the first embodiment. The information processing system 100 is a computer system (i.e., a performance recording system or a performance management system) for recording and managing a performance of a keyboard instrument 10 by a performer Ua. The information processing system 100 includes the keyboard instrument 10, a recording system 20, multiple terminal devices 30 (30a, 30b, 30c), and a measurement system 40.

[0010] Each terminal device 30 is an information device such as a smartphone, a tablet terminal, or a personal computer. Of the multiple terminal devices 30, terminal device 30a is used by a performer Ua who plays the keyboard instrument 10. The performer Ua is, for example, a person who practices playing the keyboard instrument 10. The performer Ua is, for example, a student belonging to a music school. The terminal device 30a is an example of a "control system."

[0011] Of the multiple terminal devices 30, terminal device 30b is used by instructor Ub. Instructor Ub instructs performer Ua in a music classroom. That is, instructor Ub uses terminal device 30 to manage the performer Ua's performance of the keyboard instrument 10. Terminal device 30c is used by guardian Uc (e.g., parent) of performer Ua. Like instructor Ub, guardian Uc uses terminal device 30c to manage the performer Ua's performance of the keyboard instrument 10. In the following description, when there is no need to particularly distinguish between performer Ua, instructor Ub, and guardian Uc, they will be collectively referred to as "user U."

[0012] 2 is a block diagram illustrating the configuration of a terminal device 30 (30a, 30b, 30c). The terminal device 30 includes a control device 31, a storage device 32, a communication device 33, a display device 34, an operation device 35, a sound emitting device 36, and a sound collecting device 37. The terminal device 30 may be realized as a single device, or may be realized as a plurality of devices configured separately from each other.

[0013] The control device 31 is configured with one or more processors that control each element of the terminal device 30. For example, the control device 31 is configured with one or more types of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a sound processing unit (SPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC).

[0014] The storage device 32 is one or more memories that store programs executed by the control device 31 and various data used by the control device 31. The storage device 32 is configured with a known storage medium such as a magnetic storage medium or a semiconductor storage medium. The storage device 32 may be configured with a combination of multiple types of storage media. Furthermore, a portable storage medium that is detachable from the terminal device 30, or a storage medium to which the control device 31 can write or read via a communication network (e.g., cloud storage) may be used as the storage device 32.

[0015] The communication device 33 communicates with peripheral devices (e.g., the keyboard instrument 10, the recording system 20, other terminal devices 30, and the measurement system 40). The communication device 33 communicates with the peripheral devices via wireless communication such as Wi-Fi (registered trademark) or Bluetooth (registered trademark). The communication between the communication device 33 and the peripheral devices may be wired communication. The communication between the communication device 33 and the peripheral devices may be communication via a communication network such as the Internet.

[0016] The display device 34 displays an image under the control of the control device 31. The display device 34 is configured with a display panel such as a liquid crystal panel or an organic EL panel. The operation device 35 is an input device that accepts operations by the users U (Ua, Ub, Uc). For example, an operator operated by the users U or a touch panel that detects contact by the users U is used as the operation device 35. Note that the display device 34 or the operation device 35, which are separate from the terminal device 30, may be connected to the terminal device 30 by wire or wirelessly.

[0017] The sound emitting device 36 emits sound waves under the control of the control device 31. The sound emitting device 36 is an output device such as a speaker or headphones. The sound emitting device 36 may include a D / A converter that converts the acoustic signal of the sound emitting target from digital to analog, and an amplifier that amplifies the acoustic signal. Note that the sound emitting device 36, which is separate from the terminal device 30, may be connected to the terminal device 30 by wire or wirelessly.

[0018] The sound collection device 37 collects ambient sounds to generate an acoustic signal (hereinafter referred to as a "collected sound signal"). Specifically, the sound collection device 37 collects the performance sound emitted from the keyboard instrument 10 by the performer Ua. In other words, the collected sound signal is a signal representing the waveform of the performance sound of the keyboard instrument 10. For example, the sound collection device 37 may be a single microphone or a microphone array in which multiple microphones are arranged. For convenience, an A / D converter that converts the collected sound signal from analog to digital and an amplifier that amplifies the collected sound signal are not shown in the figure. The sound collection device 37 may be separate from the terminal device 30 and connected to the terminal device 30 by wire or wirelessly. The control device 31 stores the collected sound signal generated by the sound collection device 37 in the storage device 32.

[0019] The terminal device 30 can establish communication with the keyboard instrument 10 by logging in to the keyboard instrument 10. Logging in to the keyboard instrument 10 requires authentication processing using, for example, the identification information of the keyboard instrument 10, the identification information (user ID) of the user U, and a password. The terminal device 30 can obtain the identification information by, for example, reading an image code attached to the keyboard instrument 10. The image code is a code represented by an optically readable image (for example, a QR code (registered trademark)). The identification information of the keyboard instrument 10 may be input by the user U, for example, by operating the operation device 35.

[0020] 3 is a block diagram illustrating the configuration of a keyboard instrument 10. The keyboard instrument 10 is an instrument that accepts a performance by a performer Ua, and is equipped with a control device 11, a storage device 12, a communication device 13, a display device 14, an operation device 15, a sound source device 16, a sound output device 17, and a performance mechanism 18. The keyboard instrument 10 is an example of a performance device used by the performer Ua for performance. Note that some or all of the functions of the terminal device 30 may be installed in the keyboard instrument 10.

[0021] The performance mechanism 18 includes a keyboard 181, a sound generation mechanism 183, pedals 184, and a detection device 185. The keyboard 181 is a mechanism with an array of multiple keys 182, including white keys and black keys. Each of the multiple keys 182 is a performance operator that accepts operation (e.g., key depression and key release) by the performer Ua. The sound generation mechanism 183 is an action mechanism that generates performance sounds by vibrating strings in response to operation of each key 182 on the keyboard 181. Specifically, the sound generation mechanism 183 includes strings, which are an example of a sound source, hammers that rotate to strike the strings, and a transmission mechanism (e.g., a wippen, jack, or repetition lever) that rotates the hammer in response to the displacement of each key 182. The sound generation mechanism 183 also includes a mute mechanism that prevents the strings from being struck in mute mode.

[0022] The pedal 184 is a performance operator operated by the performer Ua with the foot. For example, the pedal 184 may be a damper pedal (sustain pedal) for instructing the extension of a performance sound, or a soft pedal for reducing the volume of a performance sound.

[0023] The detection device 185 is a sensor that detects the operation of each key 182 and pedal 184. Specifically, the detection device 185 detects the amount of displacement of each key 182 and pedal 184. The detection device 185 is configured, for example, by an optical, mechanical, or magnetic sensor.

[0024] The control device 11 is composed of one or more processors that control each element of the keyboard instrument 10. For example, the control device 11 is composed of one or more types of processors such as a CPU, a GPU, an SPU, a DSP, an FPGA, or an ASIC.

[0025] The control device 11 generates music data in accordance with the detection result of the detection device 185. The music data is data representing a musical piece performed by the performer Ua. The musical piece to be performed includes a right-hand part that the performer Ua should play with his / her right hand HR and a left-hand part that the performer Ua should play with his / her left hand HL.

[0026] Specifically, the music data is a time series of event data conforming to the MIDI (Musical Instrument Digital Interface) standard, for example. That is, the music data specifies, in time series, the pitches (note numbers) corresponding to the keys 182 operated by the performer Ua and the intensity (velocity) of the operation. The music data also represents the operation of the pedal 184 by the performer Ua. As described above, the control device functions as a music data generator that generates music data in accordance with the results of detection by the detection device 185.

[0027] The display device 14 displays images under the control of the control device 31. The display device 14 is configured with a display panel such as a liquid crystal panel or an organic EL (Electroluminescence) panel. The operation device 15 is an input device that accepts operations by the performer Ua. For example, an operator operated by the performer Ua or a touch panel that detects contact by the performer Ua is used as the operation device 15. Note that the display device 14 or the operation device 15, which are separate from the keyboard instrument 10, may be connected to the keyboard instrument 10 by wire or wirelessly.

[0028] The sound source device 16 generates an audio signal representing a performance sound according to the result of detection by the detection device 185. Specifically, the sound source device 16 generates an audio signal according to the performance data generated by the control device 11. For example, an audio signal is generated having a pitch corresponding to the key 182 whose displacement has been detected by the detection device 185. Note that the control device 11 may realize the function of the sound source device 16 by executing a program stored in the storage device 12. As described above, the keyboard instrument 10 is an electronic musical instrument that can selectively generate natural sounds using the sound generation mechanism 183 and electronic sounds using the sound source device 16. Note that one of the sound generation mechanism 183 and the sound source device 16 may be omitted.

[0029] The sound emitting device 17 emits sound waves under the control of the control device 11. The sound emitting device 17 is an output device such as a speaker or headphones. Specifically, the sound emitting device 17 emits performance sounds represented by acoustic signals generated by the sound source device 16. As described above, the keyboard instrument 10 of the first embodiment can selectively generate performance sounds using the sound generation mechanism 183 and the sound source device 16. The sound emitting device 17 may include a D / A converter that converts the acoustic signals from digital to analog and an amplifier that amplifies the acoustic signals. The sound emitting device 17 may be separate from the keyboard instrument 10 and connected to the keyboard instrument 10 by wire or wirelessly.

[0030] The storage device 12 is one or more memories that store the programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is configured with a known storage medium, such as a magnetic storage medium or a semiconductor storage medium. The storage device 12 may also be configured with a combination of multiple types of storage medium. The storage device 12 may also be a portable storage medium that is detachable from the keyboard instrument 10, or a storage medium to which the control device 11 can write or read via a communication network (e.g., cloud storage). The storage device 12 stores the operating characteristic Xa, tuning information Xb, and history information Xc.

[0031] The operation characteristics Xa are device-specific characteristics (device profile) related to the operation of the keyboard instrument 10. The operation characteristics Xa include a dynamic range, a velocity range, a frequency range, a repeated key characteristic, and a key depression sensitivity.

[0032] The dynamic range is the range of volume (the ratio between maximum and minimum volume) of the performance sounds that the keyboard instrument 10 can generate via the sound generation mechanism 183 or the sound source device 16. The velocity range is the range of key pressing speed (key pressing strength). Because the volume of the performance sounds generated by the sound source device 16 corresponds to the key pressing speed, the velocity range is also expressed as the range of performance sound volume. The frequency range is the range of frequencies of the performance sounds that the keyboard instrument 10 can generate via the sound generation mechanism 183 or the sound source device 16. The repeated keying characteristic is the maximum number of key presses per unit time. The key pressing sensitivity is the relationship between the key pressing speed or displacement amount and the performance volume (touch sensitivity) for each key 182. Key pressing sensitivity includes, for example, key depth (mm), half-touch depth (mm), aftertouch depth (mm), resonance (Hz), downweight (g) for each 88th key, and return weight (g) for each key 182.

[0033] The information included in the performance characteristics Xa is not limited to the above examples. For example, the performance characteristics Xa may include pedal settings (differences in damper, sostenuto, muffler, shift, softness, etc.), location information of the installation location of the keyboard instrument 10, the type of building in which the installation location is located (recital hall, concert hall, salon, cathedral, club, plate, soundproof room, setting value), and reverberation measurement results (dB-Time) of the installation location.

[0034] The tuning information Xb is information relating to the tuning of the keyboard instrument 10. The tuning information Xb includes, for example, a reference pitch or temperament. The reference pitch is the pitch that serves as the reference for tuning (e.g., 440 Hz, 441 Hz, 442 Hz, etc.). The temperament is set to, for example, equal temperament, just intonation, Pythagorean temperament, or middle temperament. The tuning information Xb is input by a tuner when tuning the keyboard instrument 10. Note that the initial values ​​of each piece of information, including the tuning information Xb, are set, for example, before the keyboard instrument 10 is shipped.

[0035] The information included in the tuning information Xb is not limited to the above examples. For example, the tuning information Xb also includes information such as rapid-fire performance, key depth, aftertouch, velocity range, dynamic range, resonance, downweight, upweight, and comments. While some of the tuning information Xb overlaps with the operating characteristics Xa, the results of changes made during tuning or over time are saved. Even for acoustic pianos without a detection device 185, tuning information Xb is input and saved during tuning and used for comparison with other instruments. During tuning, the terminal device 30 is placed in a specific position, such as a music stand, and the volume or frequency measured by the terminal device 30 can be saved as tuning information Xb. The measured values ​​can also be converted using a fixed ratio and saved as tuning information Xb.

[0036] The historical information Xc is information related to the performance by the performer Ua, and is temporarily stored in the keyboard instrument 10 until it is transferred to the terminal device 30. In other words, after being transferred to the terminal device 30, the historical information Xc is deleted from the storage device 12. The historical information Xc is also expressed as information (performance log) that represents the history of the performance by the performer Ua.

[0037] Specifically, the history information Xc includes the identification information of the performer Ua, the sound source used, the mode used, the performance date and time (start time and end time), the performance location, the performance scene, the identification information of the keyboard instrument 10, the model, the training text used by the performer Ua, and the sheet music information of the music piece performed by the performer Ua.

[0038] The communication device 13 communicates with the terminal device 30. The communication device 13 communicates with the terminal device 30 via wireless communication such as Wi-Fi (registered trademark) or Bluetooth (registered trademark). For example, the communication device 13 transmits the operation characteristic Xa, tuning information Xb, and history information Xc stored in the storage device 12 to the terminal device 30. Note that the communication between the communication device 13 and the terminal device 30 may be wired communication. The communication device 13, which is separate from the keyboard instrument 10, may be connected to the keyboard instrument 10 via wire or wirelessly.

[0039] 1 records a performance by a performer Ua. Specifically, the recording system 20 includes multiple image capture devices 21 and one sound collection device 22. Note that some or all of the multiple image capture devices 21 may be mounted on the terminal device 30 or the keyboard instrument 10. Similarly, the sound collection device 22 may be mounted on the terminal device 30 or the keyboard instrument 10. The sound collection device 22 may also be mounted on any of the multiple image capture devices 21.

[0040] Each imaging device 21 is a video camera that generates a video of the performer Ua and the keyboard instrument 10 (hereinafter referred to as the "recorded video Ya"). That is, the recording system 20 generates multiple recorded videos Ya by capturing images of the performance by the performer Ua. Each imaging device 21 includes, for example, an optical system such as a photographing lens, an imaging element that receives incident light from the optical system, and a processing circuit that generates data for the recorded video Ya according to the amount of light received by the imaging element. Note that in addition to video equipment dedicated to imaging, a general-purpose information device (e.g., a smartphone or tablet terminal) equipped with an imaging function may also be used as the imaging device 21. Furthermore, the format of the data representing the recorded video Ya is arbitrary.

[0041] The multiple imaging devices 21 are installed at different positions and angles relative to the performer Ua and the keyboard instrument 10. Furthermore, the angle of view of each imaging device 21 is different. That is, each imaging device 21 captures a different part of the performer Ua or the keyboard instrument 10. The multiple imaging devices 21 generate recorded videos Ya in parallel with each other in parallel with the performance by the performer Ua. That is, the recording system 20 generates multiple recorded videos Ya that are recorded in parallel at different positions or angles.

[0042] 4 to 7 are schematic diagrams of the recorded videos Ya. The recorded videos Ya generated by the recording system 20 include a keyboard video Ya1, a performance video Ya2, a performance video Ya3, and a pedal video Ya4.

[0043] 4, the keyboard video Ya1 is a video of the keyboard 181 of the keyboard instrument 10 and the right hand HR and left hand HL of the performer Ua captured from above the keyboard instrument 10. In other words, the keyboard video Ya1 includes the keyboard 181 used to play the musical piece and both hands of the performer Ua. The keyboard video Ya1 does not include the face of the performer Ua.

[0044] 5, the performance video Ya2 is a video captured from the side of the performer Ua. The performance video Ya2 includes the performer Ua's face (e.g., a profile), the performer Ua's right hand HR and left hand HL, and the side of the keyboard instrument 10. In other words, the performance video Ya2 shows the performer Ua's facial expressions while performing the musical piece.

[0045] 6, the performance video Ya3 is a video captured from the right rear of the performer Ua. The performance video Ya3 includes the back and back of the head of the performer Ua, the right hand HR and both feet of the performer Ua, and the front of the keyboard instrument 10. The performance video Ya3 does not include the face of the performer Ua.

[0046] 7, the pedal video Ya4 is a video capturing images of the pedal 184 of the keyboard instrument 10 and the feet of the performer Ua. In other words, the pedal video Ya4 includes both feet of the performer Ua and the pedal 184 of the keyboard instrument 10. The pedal video Ya4 does not include the face of the performer Ua.

[0047] The sound collection device 22 in Fig. 1 is one or more microphones that collect sounds around the performer Ua (hereinafter referred to as "recorded sounds Yb"). Specifically, the sound collection device 22 is installed facing the keyboard instrument 10 and collects the performance sounds emitted by the keyboard instrument 10 as a result of the performance of the performer Ua as the collected sounds Yb. In addition to devices dedicated to collecting sounds, a general-purpose information device equipped with a sound collection function (e.g., a smartphone or tablet terminal) may also be used as the sound collection device 22. The format of the data representing the recorded sounds Yb is arbitrary.

[0048] Each imaging device 21 is connected to the terminal device 30 by wire or wirelessly. The recorded video Ya recorded by each imaging device 21 is transmitted to the terminal device 30. Similarly, the sound collection device 22 is connected to the terminal device 30 by wire or wirelessly. The recorded sound Yb by the sound collection device 22 is transmitted to the terminal device 30. As described above, the recording data Y representing the multiple recorded videos Ya and recorded sounds Yb is transmitted from the recording system 20 to the terminal device 30.

[0049] The terminal device 30 manages the temporal correspondence between each recorded video Ya and recorded sound Yb generated by the recording system 20 and the performance data generated by the keyboard instrument 10. Specifically, the terminal device 30 transmits synchronization data to the recording system 20 and the keyboard instrument 10. Each imaging device 21 adds synchronization data to the recorded video Ya, and each sound collection device 22 adds synchronization data to the recorded sound Yb. The keyboard instrument 10 also adds synchronization data to the performance data. Therefore, by referencing the synchronization data, the control device 31 can identify the corresponding points in time between each recorded video Ya, recorded sound Yb, and performance data. In other words, it is possible to synchronize each recorded video Ya, recorded sound Yb, and performance data D with each other. The synchronization data is, for example, LTC (Longitudinal Time Code) or MIDI time code.

[0050] 8 is a schematic diagram of a screen (hereinafter referred to as a "setting screen 60") displayed on the display device 34 for managing the connection between the recording system 20 and the terminal device 30. The setting screen 60 is displayed on the display device 34 in response to an operation by the user U on the operation device 35. After completing the settings for the recording system 20 using the setting screen 60, the performer Ua starts playing using the keyboard instrument 10.

[0051] The setting screen 60 includes a guide screen 61. The guide screen 61 is an image that guides the user U (Ua, Ub, Uc) to the position where each device of the recording system 20 should be installed. Specifically, the guide screen 61 displays a plurality of device icons 62 representing each device at the position where the device should be installed relative to the keyboard instrument 10.

[0052] The multiple device icons 62 include an imaging icon 62a and a sound collection icon 62b. The imaging icon 62a is an image representing the imaging device 21 connected to the terminal device 30. The sound collection icon 62b is an image representing the sound collection device 22 connected to the terminal device 30. By operating the operation device 35, the user U can move the imaging icon 62a to a position corresponding to the actual position of each imaging device 21 relative to the keyboard instrument 10. Similarly, the user U can move the sound collection icon 62b to a position corresponding to the actual position of the sound collection device 22 relative to the keyboard instrument 10.

[0053] While referring to the guide screen 61 , the user U sets up the recording system 20 around the keyboard instrument 10 and connects each imaging device 21 and sound collecting device 22 to the terminal device 30 .

[0054] As illustrated in Fig. 8, the setting screen 60 includes a scene selection window 63 and an equipment setting window 64. The scene selection window 63 is an image that allows the user U to select a performance scene by the performer Ua. Specifically, a plurality of performance scenes (normal performance, performance lesson, recital, street performance) with different performance purposes or situations are displayed in the scene selection window 63 as selection candidates for the user U. The user U can select one of the plurality of performance scenes by operating the operation device 35. The performance scene selected by the user U is shared between the terminal device 30 and the keyboard instrument 10.

[0055] The location where each device in the recording system 20 should be installed varies for each performance scene. Therefore, the guide screen 61 changes depending on the performance scene selected by the user U from the scene selection window 63. That is, the guide screen 61 guides the user U to the appropriate location of each device in each performance scene. Figure 8 shows the setting screen 60 when "performance training" is selected as the performance scene.

[0056] 9 shows a setting screen 60 when "Recital" is selected as the performance scene. The "Recital" guide screen 61 is arranged with an imaging icon 62a instructing the installation of an imaging device 21 facing the performer Ua from on stage, a terminal device icon 62a instructing the installation of a terminal device 30c that will image the performer Ua from the right side (rear) from the audience seats, a terminal device icon 62a' instructing the installation of a terminal device 30c' that will image the performer Ua from the rear right from the audience seats, a terminal device icon 62 installed on the music stand of the keyboard device 10 to image the face of the performer Ua, an icon of the keyboard device 10 instructing the installation of the keyboard device 10, and a sound pickup icon 62b instructing the installation of a sound pickup device 22 facing the keyboard instrument 10.

[0057] For example, specific connection terminals are terminal device 30c, which is a smartphone owned by the father who is the guardian Uc, and terminal device 30c', which is a smartphone owned by the mother, and all devices are connected to the smartphone owned by the father (terminal device 30c). Each device is connected to the same network, and a private network is formed using a password for login.

[0058] The terminal device 30c manages the temporal correspondence between the recorded videos Ya generated by the terminal devices 30c, 30c', and 30a, the imaging device 21, the recorded sounds Yb including the sounds recorded by the sound collection device 22, and the performance data generated by the keyboard instrument 10. Specifically, the terminal device 30 transmits synchronization data to the terminal devices 30a, 30c', the imaging device 21, the sound collection device 22, and the keyboard instrument 10. Each device adds synchronization data to the recorded videos Ya and the recorded sounds Yb. The keyboard instrument 10 also adds synchronization data to the performance data. Therefore, by referencing the synchronization data, the control device 30 can identify the corresponding points in time between the recorded videos Ya, the recorded sounds Yb, and the performance data. In other words, the recorded videos Ya, the recorded sounds Yb, and the performance data D can be synchronized with each other. The synchronization data is, for example, LTC (Longitudinal Time Code) or MIDI time code.

[0059] The device settings window 64 displays devices connected to the network and is a screen for specifying the installation location. It displays devices connected to the same private network as the terminal device 30, as well as public devices shared with other performers, such as an imaging device, keyboard device, and microphone. This screen displays the Piano (keyboard device 10), UI (30a) (performer Ua), UI (30c) (father), UI (30c') (mother), Mic (installed at the venue), and Camera #1 (installed at the venue) installed at the venue.

[0060] 10 shows a setting screen 60 when "Street Performance" is selected as the performance scene. The "Street Performance" guide screen 61 has an imaging icon 62a that instructs the installation of an imaging device 30a facing the performer Ua, an imaging icon 62b that instructs the installation of a terminal device 30c that images the performer Ua and the keyboard instrument 10 so as to include the surrounding spectators, and a sound pickup icon 62b that instructs the installation of a sound pickup device 22 facing the keyboard instrument 10.

[0061] For example, specific terminal devices include terminal device 30a, which is a smartphone owned by performer Ua, terminal device 30c, which is a smartphone owned by performer Ua's guardian Uc, and all devices are connected to the smartphone (terminal device 30c) owned by the father. Each device is connected to the same network, and a private network is formed using a password for login. As mentioned above, each device receives synchronization data transmitted from terminal device 30c, and the corresponding points in time of the recorded video and audio can be identified.

[0062] The device settings window 64 displays devices connected to the network and allows selection of installation locations. It displays devices connected to the same private network as the terminal device 30, as well as public devices shared with other performers, such as an imaging device, keyboard device, and microphone. This screen displays the piano (keyboard device 10), UI (30a) (performer Ua), UI (30c) (father), and microphone (installed at the venue) installed at the venue.

[0063] 11 is a schematic diagram of the setting screen 60 in a state where the actual keyboard device, imaging device, and sound collection device displayed in the device setting window 64 have been linked to the icons. For example, a user connects Camera #1 displayed in the device setting window 64 to the imaging icon 62a on the guide screen 61 by dragging and dropping. The imaging icon 62a has setting information for the imaging location, which may be preset or based on instructions (input) from the user to the management screen 60. The setting information may be linked to information such as above, face-NG (image not showing the face), above the keyboard, both hands, or the top of the keyboard device. The user uses a control device to connect the imaging device to the icon and installs the camera (imaging device) on the ceiling of the room or to the side of the piano.

[0064] Similarly, the control device 31 recognizes the video X from the imaging device 21 whose installation location is set to "right side" in the connection settings as a performance video Xb of the performer Ua's right side, right side profile, OK face (image showing face), right hand, left hand, right foot, and side of the keyboard device. The control device 31 recognizes the video X from the imaging device 21 whose installation location is set to "right rear" as a pedal video Xc of the performer Ua's right rear, right back head, NO face, right hand, right foot, and right rear side of the keyboard device. The control device 31 recognizes the video from the imaging device 21 whose installation location is set to "right foot" as a pedal right video Xd. For example, a user may indicate "above" as "ceiling" or "right arm" instead of "right hand." In addition to the displayed setting information and user input, the control device 31 searches for words similar to the displayed and input language and simultaneously saves them as detailed information. These detailed information can be displayed and edited at the user's request. The detailed information is saved as imaging device installation information for each imaging device 21 and used for comparison with the analysis information described below.

[0065] FIG. 12 shows the settings screen 60 when "Music Classroom" is selected as the performance scene. The classroom includes an instructor, student 1, student 2, and student 2's guardian Uc. The instructor, student 1, and student 2 use their own control terminals to connect to Piano 1, Piano 2, and Piano 3 by entering their user IDs and passwords. For example, the device settings window 64 of student 2's control terminal 30a displays devices connected to the network and allows users to select their installation locations. Piano 3, control terminal 3, control terminal 4, or camera 5 and camera 6, all connected to the same private network as terminal device 30a, are displayed, along with public devices shared with other performers, such as Piano 1, camera 1, camera 2, and control terminal 1. The user can link devices displayed in the device settings window 64 to icons. This linking associates preset image capture location settings and user information with each icon. Student 2 is also accompanied by guardian Uc, and guardian Uc's control terminal 30c is linked to control terminal 30a. The user can display the icon they wish to add by performing an add operation, and can add the device by entering the imaging position, etc. The set information can be saved, and the user can restore the saved settings by logging in to the keyboard device 10 the next time they visit.

[0066] The measurement system 40 in FIG. 1 measures biometric information Z related to the player Ua. Specifically, the measurement system 40 includes a biometric sensor 41 and a center-of-gravity sensor 42. The biometric sensor 41 is a portable sensor worn on the body of the player Ua. The biometric sensor 41 is, for example, a heart rate sensor that detects the heart rate of the player Ua. The center-of-gravity sensor 42 is a sensor installed on the chair on which the player Ua sits while playing the keyboard instrument 10, and detects movement of the center of gravity of the seat. The biometric information Z of the player Ua is data including, for example, the player Ua's playing posture, level of tension (strain), heart rate, and concentration.

[0067] The playing posture is an index relating to the posture of the player Ua playing the keyboard instrument 10. The playing posture is determined by analyzing the detection results of the center of gravity sensor 42. Note that the playing posture may also be determined by analyzing a recorded video Ya of the player Ua captured by the imaging device 21.

[0068] The heart rate is detected, for example, by a heart rate sensor. The tension level is an index of the player Ua's level of tension (strain) and is calculated, for example, based on the player Ua's heart rate and the stability of his / her playing posture. The concentration level is an index of the player Ua's level of concentration on playing and is calculated, for example, based on the player Ua's heart rate and the stability of his / her playing posture. Note that the concentration level may be set to a higher value the more times a key is pressed within a specified period of time. Note that the configuration and type of the measurement system 40 are not limited to the above examples.

[0069] In the configuration described above, the terminal device 30 acquires various information from the peripheral devices (the keyboard instrument 10, the recording system 20, and the measurement system 40). For example, the control device 31 transmits an information request to the peripheral devices from the communication device 33, and receives information transmitted from the peripheral devices in response to the information request via the communication device 33. Specifically, the terminal device 30 acquires the motion characteristics Xa, tuning information Xb, and history information Xc from the keyboard instrument 10, acquires recording data Y representing a plurality of recorded videos Ya and recorded sounds Yb from the recording system, and acquires biometric information Z of the performer Ua from the measurement system 40.

[0070] As illustrated in Fig. 13, performance data D is stored in the storage device 32 of the terminal device 30. The performance data D is comprehensive data related to the performance by the performer Ua. Specifically, the performance data D includes music data corresponding to the results of detection by the detection device 185, motion characteristics Xa, tuning information Xb, history information Xc, recorded data Y, and biometric information Z. The performance data D is stored in the storage device 32 for each performance by the performer Ua (each time the performer logs in to the keyboard instrument 10). The terminal device 30 executes various processes using the performance data D.

[0071] 1 is a computer system (cloud storage) that records various types of information. Some or all of the information stored in the storage device 32 of each terminal device 30 or the storage device 12 of the keyboard instrument 10 may be stored in the recording system 90. In other words, some or all of the storage device 32 or the storage device 12 may be replaced by the recording system 90. For example, some or all of the performance data D may be transferred from a peripheral device (the keyboard instrument 10, the recording system 20, the measurement system 40) to the recording system 90, and then transmitted from the recording system 90 to the terminal device 30 in response to a request from the terminal device 30.

[0072] The terminal device 30 (30a, 30b, 30c) generates a video (hereinafter referred to as an "edited video V") representing a performance by the performer Ua, using the performance data D stored in the storage device 32. Specifically, the edited video V is generated using a plurality of recorded videos Ya and recorded sounds Yb included in the recording data Y.

[0073] 14 is an explanatory diagram of edited video V. Edited video V is composed of video Va and audio Vb. Video Va and audio Vb are played in parallel with each other. Audio Vb is recorded sound Yb (i.e., the performance sound of the keyboard instrument 10) recorded by the sound collection device 22. Audio Vb may be generated by performing various types of acoustic processing on recorded sound Yb (for example, adding effects or removing noise).

[0074] The video Va is a video in which multiple partial videos Vq (Vq1, Vq2, ...) corresponding to different phrase periods Q (Q1, Q2, ...) on the time axis are arranged in chronological order. Each phrase period Q is a section (musical phrase) that constitutes a musical unit in a song. The partial video Vq corresponding to each phrase period Q is part or all of one of the multiple recorded videos Ya (Ya1 to Ya4) recorded by the recording system 20 in that phrase period Q. As can be understood from the above explanation, the terminal device 30 generates an edited video V in which the recorded videos Ya are switched sequentially for each phrase period Q of the song. In other words, in the edited video V, the imaging position and direction change at the boundaries of the phrase periods Q of the song.

[0075] The display device 34 displays a moving image Va of the edited moving image V. The sound emitting device 36 reproduces an audio Vb of the edited moving image V. That is, the display device 34 and the sound emitting device 36 function as a reproduction device that reproduces the edited moving image V.

[0076] 15 is a block diagram illustrating an example of the functional configuration of each terminal device 30. The control device 31 executes a program stored in the storage device 32 to realize multiple functions (music analysis unit 51, video editing unit 52, and playback control unit 53) for generating the edited video V.

[0077] The music analysis unit 51 identifies phrase periods Q of a piece of music by analyzing the performance data D. Specifically, the music analysis unit 51 identifies phrase periods Q using the performance data D and reference data. The reference data is data representing the musical score of a piece of music and is stored in the storage device 32. The reference data is, for example, a time series of event data conforming to the MIDI standard. Specifically, the reference data specifies the pitch (note number) and intensity (velocity) for each of the multiple notes that make up the piece of music. Specifically, the music analysis unit 51 of the first embodiment generates basic data B by analyzing the performance data D and the reference data, and generates control data C using the basic data B. The control data C is data representing each phrase period Q of the piece of music.

[0078] Note that the use of reference data may be omitted in the process of identifying the phrase period Q of the music piece. A reference setting window 65 of FIG. 16 regarding the presence or absence of reference data may be displayed on the display device 34, for example, together with the setting screen 60 described above. The user U can select one of multiple options displayed in the reference setting window 65 by operating the operation device 35.

[0079] Among the multiple options in the reference setting window 65, "Music Score" means the use of existing reference data stored in the storage device 32. "Past Performance" in the reference setting window 65 means reference data representing the performer Ua's past performances. In other words, music data generated when the performer Ua played a piece of music in the past is stored in the storage device 32 as reference data for "past performance." "Search" in the reference setting window 65 means searching for reference data for the piece of music, for example, from a sheet music distribution site. "None" in the reference setting window 65 means that no reference data is used. In other words, if "None" for reference data is selected, the music analysis unit 51 will not detect performance errors.

[0080] 17 is a schematic diagram of basic data B. Basic data B is a data table that specifies multiple pieces of performance information A (A1 to A11) for each of multiple unit periods P (P1, P2, ...) that divide a musical piece on the time axis. A unit period P is, for example, a period equivalent to one measure of a musical piece. However, unit period P may also be a period equivalent to multiple measures of a musical piece, or a period of a predetermined length that is unrelated to measures.

[0081] The performance information A is musical information related to the performance of the music piece by the performer Ua. Specifically, the performance information A includes difficulty level A1, ascending style information A2, descending style information A3, right hand attention level A4, left hand attention level A5, pedal information A6, slur information A7, staccato information A8, dynamic information A9, time information A10, and phrase information A11.

[0082] The difficulty level A1 is an index of the difficulty of playing a piece of music. The music analysis unit 51 determines the difficulty level A1 based on musical elements such as the number of notes or pitch difference within the unit period P, the complexity of the note sequence or rhythm, and the presence or absence of ornaments.

[0083] The ascending form information A2 is information indicating that the performance is an ascending form. An ascending form means that the pitch of the music piece increases over time. Specifically, the ascending form information A2 specifies, using a binary value, whether or not the performance within the unit period P corresponds to an ascending form. On the other hand, the descending form information A3 is information indicating that the performance is a descending form. A descending form means that the pitch of the music piece decreases over time. Specifically, the descending form information A3 specifies, using a binary value, whether or not the performance within the unit period P corresponds to a descending form.

[0084] The right-hand attention level A4 is the degree to which the right-hand part of a piece of music should be given attention. Specifically, the music analysis unit 51 sets the right-hand attention level A4 according to the number of notes, pitch difference, or volume of the right-hand part of the piece of music in the unit period P. Similarly, the left-hand attention level A5 is the degree to which the left-hand part of a piece of music should be given attention. Specifically, the music analysis unit 51 sets the left-hand attention level A5 according to the number of notes, pitch difference, or volume of the left-hand part of the piece of music in the unit period P.

[0085] The pedal information A6 is information that indicates whether or not the pedal 184 of the keyboard instrument 10 has been operated. The music analysis unit 51 determines whether or not the pedal 184 has been operated based on the performance data D within the unit period P, and generates the pedal information A6 based on the result of the determination.

[0086] The slur information A7 is information that represents a slur in a piece of music. The number of the slur is set as the slur information A7 over one or more unit periods P in which one slur is distributed. The staccato information A8 is information that represents a staccato in a piece of music. Specifically, the staccato information A8 represents the number of staccatos within the unit period P.

[0087] The dynamic information A9 represents the strength of the performance in the music piece. For example, dynamic symbols such as forte (f) or piano (p) are set as the dynamic information A9. The music analysis unit 51 generates the dynamic information A9 according to the strength of the performance represented by each piece of performance data D within the unit period P.

[0088] The time information A10 is information that indicates the start or end time of the unit period P. Specifically, the music analysis unit 51 generates, for example, the elapsed time from the start of the music piece to the start or end of the unit period P as the time information A10.

[0089] The phrase information A11 is information that represents a phrase period Q in a piece of music. Specifically, the phrase information A11 of each unit period P specifies the time of the start (or end) of the phrase period Q that exists in that unit period P. The phrase information A11 of a unit period P in which the start point of a phrase period Q does not exist is set to an invalid value (null).

[0090] The basic data B may be transmitted from a network server as data accompanying the musical score information and stored in the terminal device 30 or the keyboard instrument 10. The performance information A is information specific to the musical piece and is not limited to the items exemplified above. For example, various musical elements displayed on the musical score, such as the sound intensity of the highest note within the phrase period Q, the length of rests, modulation, pedal identification (left pedal una corda (uc): press down the left pedal, tre corde (tc): release the left pedal), and tempo, may be set as the performance information A. The basic data B may be created by analyzing the musical piece, but it may also be created by a skilled performer for each measure. In addition to information related to the performance, the basic data B may also include recommended camera selection and video processing for each phrase period Q.

[0091] The music analysis unit 51 uses the performance data D and the performance information A other than the phrase information A11 to identify the phrase period Q. To identify the phrase period Q of a piece of music, any known technology disclosed in, for example, Japanese Patent Laid-Open No. 2023-039332 or Japanese Patent No. 3812510 may be adopted. The music analysis unit 51 may identify multiple candidates for the phrase period Q from the performance data D and the performance information A, and select one of the multiple candidates as the phrase period Q.

[0092] The music analysis unit 51 compares the basic data B with the performance information A to create analysis information M, and then generates control data C. The control data C is information for managing each phrase period Q of a piece of music. FIG. 18 is a schematic diagram of the control data C. The control data C is a data table in which period information T and analysis information M are registered for each of the multiple phrase periods Q (Q1, Q2, Q3, ...) of a piece of music.

[0093] The period information T is information that specifies the start and end points of the phrase period Q. Specifically, the period information T represents the start and end times of the phrase period Q relative to the start point of the song. The analysis information M is information that represents the tendencies or characteristics of the performance within the phrase period Q of the song. Specific examples of the analysis information M are shown below.

[0094] When a difficulty level A1 is set within a phrase period Q of basic data B, the difficulty level A1 is established when the performance information A of performer Ua exceeds a predetermined threshold value determined by the music analysis unit 51. Furthermore, when ascending form information A2 within a phrase period Q of basic data B indicates an ascending form, the ascending form is established when the performance information A of performer Ua exceeds a predetermined threshold value determined by the music analysis unit 51. Similarly, when descending form information A3 within a phrase period Q of basic data B indicates a descending form, the descending form is established when the performance information A of performer Ua exceeds a predetermined threshold value determined by the music analysis unit 51.

[0095] When a right hand attention level A4 is set within a phrase period Q of basic data B, the right hand attention level A4 is established when the performance information A of performer Ua exceeds a predetermined threshold value determined by the music analysis unit 51. When a left hand attention level A5 is set within a phrase period Q of basic data B, the left hand attention level A5 is established when the performance information A of performer Ua exceeds a predetermined threshold value determined by the music analysis unit 51.

[0096] When pedal information A6 is set within phrase period Q of basic data B, if performance information A of performer Ua exceeds a predetermined threshold value determined by music analysis unit 51, pedal information A6 is established.

[0097] When the note information within a phrase period Q of basic data B indicates do-re-mi-fa-so-la, if the performance information A of performer Ua does not match the note information of basic data B as determined by the music analysis unit 51, a performance error is determined. Specifically, the music analysis unit 51 detects the presence or absence of a performance error by performer Ua by comparing performance data D with the reference data. A performance error is a portion of the performance represented by performance data D that differs from the performance represented by the reference data. For example, an incorrect operation related to pitch or rhythm is detected as a performance error. When a performance error is detected within a phrase period Q, the music analysis unit 51 includes information indicating the occurrence of a performance error in the analysis information M of the phrase period Q in the control data C.

[0098] The image analysis information obtained by analyzing the image of the performer is used by each imaging device 21 to analyze the movements of the performer Ua. For example, an imaging device 21 set up in the "upper" position can capture images of both hands of the performer Ua from above the keyboard 181, allowing for the determination of whether the performance expression is ascending or descending. Similarly, an imaging device 21 set up in the "right side" position can capture images of the movements of the right hand HR and left hand HL from the right side of the performer Ua, allowing for the determination of whether the performance expression is ascending or descending. The analysis of the performance expression from the image can be evaluated by determining the movements of the performer Ua's face, hands, feet, and posture using known image recognition technology, and identifying the direction of movement, facial expression, magnitude of movement, etc. relative to the keyboard instrument 10. The selection of the performance expression of the performer Ua to be selected through image analysis can be automatically or manually set on the setting screen 60 when configuring the imaging device 21.

[0099] The video editing unit 52 of FIG. 15 generates an edited video V by selecting multiple recorded videos Ya according to phrase periods Q. Specifically, as described above with reference to FIG. 14 , the video editing unit 52 generates a video Va in which a partial video Vq is sequentially switched to each of the multiple recorded videos Ya for each phrase period Q of the music. That is, in the video Va, the partial videos Vq are switched at the times specified by each period information T of the control data C (i.e., the start and end points of each phrase period Q). That is, editing by the video editing unit 52 includes a process of selecting different recorded videos Ya from the multiple recorded videos Ya as the partial videos Vq in one phrase period Q and the immediately following phrase period Q. For example, as illustrated in FIG. 14 , different recorded videos Ya are selected as the partial videos Vq in phrase periods Q1 and Q2. The video editing unit 52 generates an edited video V by adding recorded sound Yb as audio Vb to the video Va.

[0100] As explained above, in the first embodiment, an edited video V is generated in accordance with a phrase period Q identified by analyzing the performance data D, so that an edited video V suitable for a piece of music can be generated without placing an excessive burden on the user. In particular, in the first embodiment, different recorded videos Ya are selected as partial videos Vq in two consecutive phrase periods Q, so that a suitable edited video V can be generated in which the partial videos Vq switch in synchronization with musical divisions in the piece of music (i.e., the start or end point of a phrase period Q).

[0101] The video editing unit 52 of the first embodiment generates an edited video V by editing a plurality of recorded videos Ya in accordance with the phrase period Q of the music and the analysis information M of the control data C. Specifically, the video editing unit 52 selectively executes one of a plurality of different editing processes E in accordance with the analysis information M. As described above, since the analysis information M is reflected in the edited video V in addition to the phrase period Q of the music, there is a significant effect in that an edited video V suitable for the music can be generated.

[0102] The multiple editing processes E include facial expression attention process E1, ascending follow-up process E2, descending follow-up process E3, right hand attention process E4, left hand attention process E5, pedal attention process E6, normal process E7, and performance error process E8. The specific contents of each editing process E will be described in detail below.

[0103] [Facial Expression Attention Processing E1] The facial expression attention processing E1 is processing for selecting the performance video Ya2 of the performer Ua from among the multiple recorded videos Ya as the partial video Vq for the phrase period Q. Specifically, as illustrated in Fig. 19, the facial expression attention processing E1 includes processing for cutting out a partial area including the face of the performer Ua from the performance video Ya2 as the partial video Vq. For example, the face of the performer Ua is zoomed in on in the performance video Ya2.

[0104] When a difficult section of a musical piece is being played, there is a demand to focus on the facial expression of the performer Ua. Taking this demand into consideration, the video editing unit 52 executes facial expression attention processing E1 for a phrase period Q of the musical piece in which the analysis information M indicates that the performance is highly difficult. In other words, the performance video Ya2 is selected as the partial video Vq for the phrase period Q in which the difficulty level A1 exceeds a threshold. According to the above-described embodiment, it is possible to generate an edited video V that emphasizes the facial expression of the performer Ua playing a section of the musical piece with a high difficulty level A1.

[0105] [Ascending follow-up process E2] The ascending follow-up process E2 is a process of selecting the keyboard video Ya1 from the multiple recorded videos Ya as the partial video Vq for the phrase period Q. Specifically, the video editing unit 52 executes the ascending follow-up process E2 for the phrase period Q for which the analysis information M indicates that the performance is in an ascending form. That is, the video editing unit 52 executes the ascending follow-up process E2 when the analysis information M indicates an ascending form.

[0106] The ascending tracking process E2 is a process of panning the keyboard 181 in the keyboard video Ya1 from a low-pitched range to a high-pitched range. Specifically, as illustrated in FIG. 20 , the video editing unit 52 extracts a partial area of ​​the keyboard video Ya1 that includes the keyboard 181 and moves that area to the right at a predetermined speed. During phrase period Q in which the analysis information M indicates an ascending form, the performer Ua's right hand HR or left hand HL moves from a low-pitched range to a high-pitched range. Therefore, the partial video Vq extracted by the ascending tracking process E2 is a video whose imaging range moves to the right along with the performer Ua's right hand HR or left hand HL. This makes it possible to generate an edited video V that focuses on both hands of the performer Ua to follow the rising (ascending) pitch of the music piece.

[0107] [Descending Follow-Up Process E3] The descending follow-up process E3, like the ascending follow-up process E2, is a process of selecting the keyboard video Ya1 from the multiple recorded videos Ya as the partial video Vq for the phrase period Q. Specifically, the video editing unit 52 executes the descending follow-up process E3 for the phrase period Q for which the analysis information M indicates that the performance is in a descending form. That is, the video editing unit 52 executes the descending follow-up process E3 when the analysis information M indicates a descending form.

[0108] The descending tracking process E3 is a process of panning the keyboard 181 from a high-pitched range to a low-pitched range in the keyboard video Ya1. Specifically, as illustrated in FIG. 20 , the video editing unit 52 extracts an area of ​​the keyboard video Ya1 that includes the keyboard 181 and moves that area leftward at a predetermined speed. During phrase period Q in which the analysis information M indicates a descending pattern, the performer Ua's right hand HR or left hand HL moves from a high-pitched range to a low-pitched range. Therefore, the partial video Vq extracted by the descending tracking process E3 is a video whose imaging range moves leftward along with the performer Ua's right hand HR or left hand HL. This makes it possible to generate an edited video V that focuses on both hands of the performer Ua to follow the pitch drop (descending) in the music.

[0109] As explained above, the video editing unit 52 executes different processes (ascending follow-up process E2 and descending follow-up process E3) depending on whether the analysis information M indicates an ascending type or a descending type. Therefore, as mentioned above, it is possible to generate an edited video V that corresponds to an ascending (ascending) or descending (descending) pitch change in a piece of music.

[0110] [Right Hand Attention Processing E4] The right hand attention processing E4 is processing for selecting the keyboard video Ya1 from the multiple recorded videos Ya as the partial video Vq for the phrase period Q. Specifically, the video editing unit 52 executes the right hand attention processing E4 for the phrase period Q in the music piece where the right hand attention level A4 exceeds a threshold. That is, the video editing unit 52 executes the right hand attention processing E4 when the right hand attention level A4 exceeds the threshold.

[0111] The right hand attention process E4 is a process of enlarging the right hand HR of the player Ua in the keyboard video Ya1. Specifically, as shown in FIG. 21 , the video editing unit 52 cuts out an area of ​​the keyboard video Ya1 that includes the right hand HR of the player Ua as a partial video Vq. Therefore, in the partial video Vq of the edited video V, the right hand HR of the player Ua is enlarged compared to the keyboard video Ya1. For example, the right hand HR of the player Ua is zoomed in.

[0112] [Left hand attention processing E5] The left hand attention processing E5 is processing for selecting the keyboard animation Ya1 from the multiple recorded animations Ya as the partial animation Vq for the phrase period Q. Specifically, the animation editing unit 52 executes the left hand attention processing E5 for the phrase period Q in the music piece where the left hand attention level A5 exceeds a threshold. In other words, the animation editing unit 52 executes the left hand attention processing E5 when the left hand attention level A5 exceeds the threshold.

[0113] The left hand attention process E5 is a process of enlarging the left hand HL of the player Ua in the keyboard video Ya1. Specifically, as shown in FIG. 21 , the video editing unit 52 cuts out an area of ​​the keyboard video Ya1 that includes the left hand HL of the player Ua as a partial video Vq. Therefore, in the partial video Vq of the edited video V, the left hand HL of the player Ua is enlarged compared to the keyboard video Ya1. For example, the left hand HL of the player Ua is zoomed in.

[0114] As illustrated above, the editing process E in the first embodiment includes a right-hand attention process E4 and a left-hand attention process E5. Therefore, an edited video V can be generated that emphasizes the right hand HR or left hand HL of the performer Ua, whichever hand has a higher degree of attention in the music (right-hand attention degree A4, left-hand attention degree A5).

[0115] [Pedal Focusing Process E6] The pedal focus processing E6 is processing for selecting the pedal video Ya4 of the performer Ua from the multiple recorded videos Ya as the partial video Vq for the phrase period Q. Specifically, as illustrated in FIG. 22 , the pedal focus processing E6 includes processing for cutting out a partial area of ​​the pedal video Ya4 that includes the pedal 184 as the partial video Vq. The video editing unit 52 executes the pedal focus processing E6 for the phrase period Q in which the analysis information M indicates operation of the pedal 184. In other words, the video editing unit 52 executes the pedal focus processing E6 when the analysis information M indicates operation of the pedal 184. Therefore, an edited video V can be generated that focuses on the operation of the pedal 184 by the performer Ua.

[0116] As illustrated above, the editing process E by the video editing unit 52 includes a process of cutting out a specific region from the recorded video Ya in accordance with the analysis information M. Specifically, in the facial expression attention process E1, a region including the face of the performer Ua is cut out from the performance video Ya2. In the ascending tracking process E2 and the descending tracking process E3, a region including the keyboard 181 is cut out from the keyboard video Ya1. In the right hand attention process E4, a region including the right hand HR of the performer Ua is cut out from the keyboard video Ya1. In the left hand attention process E5, a region including the left hand HL of the performer Ua is cut out from the keyboard video Ya1. Furthermore, in the pedal attention process E6, a region including the pedal 184 is cut out from the pedal video Ya4. With the above configuration, as described above for each editing process E, an edited video V can be generated that focuses on a specific region from the recorded video Ya in accordance with the analysis information M. Note that the cutting out of some regions in each editing process E may be omitted.

[0117] Furthermore, as illustrated above, when any one of a plurality of conditions (hereinafter referred to as "execution conditions") related to the analysis information M is satisfied, one of the plurality of editing processes E (E1 to E6) corresponding to that execution condition is executed. Specifically, the execution condition for the facial expression attention process E1 is that the difficulty level A1 exceeds a threshold. The execution condition for the ascending following process E2 is that the analysis information M indicates an ascending form, and the execution condition for the descending following process E3 is that the analysis information M indicates a descending form. The execution condition for the right hand attention process E4 is that the right hand attention level A4 exceeds a threshold, and the execution condition for the left hand attention process E5 is that the left hand attention level A5 exceeds a threshold. Furthermore, the execution condition for the pedal attention process E6 is that the analysis information M indicates the operation of the pedal 184.

[0118] It should be noted that there may be cases where multiple execution conditions are met simultaneously. When multiple execution conditions are met simultaneously, the video editing unit 52 selectively executes one of the multiple editing processes E for which the execution conditions are met simultaneously. For example, a priority is set for each editing process E (E1 to E6). The priority of each editing process E is stored in the storage device 32. The video editing unit 52 executes the editing process E with the highest priority among the multiple editing processes E for which the execution conditions are met simultaneously.

[0119] [Normal Process E7] Normal process E7 is a process of selecting the performance video Ya2 or Ya3 of performer Ua from the multiple recorded videos Ya as the partial video Vq for phrase period Q. Specifically, normal process E7 is a process of selecting a wide area of ​​performer Ua's body (e.g., the entire body) from the performance video Ya2 or Ya3 as the partial video Vq. Normal process E7 includes a process of enlarging a portion of the performance video Ya2 or Ya3 over time (i.e., zooming in to enlarge the subject over time) and a process of shrinking the portion over time (i.e., zooming out to shrink the subject over time). The video editing unit 52 executes normal process E7, for example, when none of the multiple execution conditions is met.

[0120] [Performance Error Processing E8] The performance error processing E8 is processing for selecting the performance error processing Ya2 of the performer Ua from among the multiple recorded videos Ya as the partial video Vq for the phrase period Q. The execution condition for the performance error processing E8 is that the analysis information M indicates a performance error. That is, the video editing unit 52 executes the performance error processing E8 when a performance error occurs. In the performance error processing E8, the video editing unit 52 selects a partial area of ​​the performance video Ya2 as the partial video Vq and reduces this area over time. That is, the performer Ua or the keyboard instrument 10 is reduced (zoomed out) over time in the partial video Vq.

[0121] As described above, the partial moving image Vq is zoomed out during the phrase period Q in which the performer Ua made a performance error, so that an edited moving image V can be generated in which the performance error by the performer Ua is not noticeable.

[0122] 15 plays back the edited video V generated by the video editing unit 52. Specifically, the playback control unit 53 displays a video Va of the edited video V on the display device 34, and emits an audio Vb of the edited video V from the sound emission device 36.

[0123] 23 is a flowchart of the process (hereinafter referred to as "video editing process") executed by the control device 31 to generate the edited video V. For example, the video editing process is executed in response to a user's operation on the operation device 35. The video editing process may be executed in real time in parallel with the performance by the performer Ua, or may be executed after the performance by the performer Ua.

[0124] When the video editing process starts, the control device 31 (music analysis unit 51) analyzes (Sa1) the performance data D. Specifically, the control device 31 generates basic data B by analyzing the performance data D, and generates control data C using the basic data B. The analysis of the performance data D includes identifying each phrase period Q of the music piece.

[0125] The control device 31 (video editing unit 52) ​​generates an edited video V using the multiple recorded videos Ya and recorded sounds Yb included in the performance data D and the control data C that is the result of analyzing the performance data D (Sa2). For example, the control device 31 selectively executes one of the multiple editing processes E (E1 to E7) described above for each phrase period Q. The control device 31 (playback control unit 53) plays back the edited video V generated by the above procedure on the display device 34 and the sound output device 36 (Sa3).

[0126] B: Second Embodiment A second embodiment will be described. Note that, for elements in the following exemplary aspects that have the same functions as those in the first embodiment, the same reference numerals as those in the first embodiment will be used, and detailed descriptions of each element will be omitted as appropriate.

[0127] The video editing unit 52 of the second embodiment edits a plurality of recorded videos Ya under any of a plurality of different editing conditions. The editing conditions are conditions related to the editing of each recorded video Ya in the process of generating the edited video V. Specifically, the content of each editing process E, the priority of each editing process, and the execution conditions of each editing process are set as the editing conditions.

[0128] A plurality of editing conditions corresponding to different performance scenes are set in advance. The video editing unit 52 generates an edited video V from the plurality of recorded videos Ya under an editing condition corresponding to a performance scene selected by a user from the plurality of editing conditions.

[0129] For example, in a performance training session, editing conditions are set in advance to allow instructor Ub to easily create a video for enjoying a performance. In performance training, it is important for the student to view the video from a direction that allows for effective viewing of instructor Ub's fingering for each phrase, observe instructor Ub's posture during difficult phrases, and watch the pedal operation that affects the performance. It is also necessary for instructor Ub to be able to select camera switching and zoom in on fingering via voice commands. It is also necessary to enable the display of sheet music and a piano roll during performance, and to set a pedal operation screen to be overlaid on top of the screen for necessary phrases. Therefore, when performance training is selected as a performance scene, high priority is assigned to the right hand attention processing E4, the left hand attention processing E5, and the pedal attention processing E6. Furthermore, in performance training, the magnification rate of the right hand HR in the right hand attention processing E4 is set to a higher value than in normal performance, and the magnification rate of the left hand HL in the left hand attention processing E5 is set to a higher value than in normal performance.

[0130] Furthermore, the video editing unit 52 estimates the words spoken by the user U through speech recognition of the voice picked up by the sound collection device 22, and selects an imaging device that corresponds to a word that matches or is similar to the word. Therefore, for example, if the instructor Ub utters "Look at the hand from above," "Look from above," or "Upper camera," an imaging device installed above is selected. In a scene where it is effective to view the right hand from the side, if the instructor Ub utters "Right hand close-up" or "Zoom in from the side of the right hand," a side imaging device is selected and video processing for zooming in is performed. Furthermore, when the instructor Ub pronounces "pedal," an image of the pedal operation is superimposed on the screen during the corresponding phrase period. The instructor Ub can create a performance training video by playing while talking to the student.

[0131] The instructor Ub can create a performance training video using the basic data B. The basic data B includes control data for each phrase period that is used to control the camerawork optimal for performance training in accordance with the training piece.

[0132] The video editing unit 56 can generate an optimal framework using performance analysis, voice recognition, and basic data.

[0133] At a recital, it is important to check the facial expression and performance of the performer Ua. Therefore, when a recital is selected as a performance scene, high priority is assigned to the facial expression attention process E1, the upward tracking process E2, and the downward tracking process E3. In addition, by setting the enlargement rate in the right hand attention process E4 or the left hand attention process E5 to a lower value compared to normal performance, an edited video V that includes a wide range of the keyboard instrument 10 is generated.

[0134] In street performances, it is important to check the facial expressions of the performer Ua and the state of the performance. Therefore, when a recital is selected as a performance scene, high priority is set for the facial expression attention processing E1, the upward tracking processing E2, and the downward tracking processing E3. In addition, in street performances, it is also important to check the state of the audience. Therefore, an imaging device 21 is installed that captures a wide range of the environment in which the performer Ua is performing, and a priority is set so that the recorded video Ya captured by the imaging device 21 is preferentially selected as the partial video Vq.

[0135] In a music school, it is important to pay attention to the instructor's performance when the instructor is demonstrating a piece, and to pay attention to the student's performance when the student is performing. The video editing unit 56 detects the user ID and key input status of the keyboard device 10 and prioritizes editing the instructor's performance video for appropriate phrase periods during the instructor's performance. When editing the instructor's performance, as with performance training, it is important to view the instructor Ub's fingering from a direction that effectively visualizes each phrase period, as well as the instructor Ub's posture during difficult phrases and the pedal operation that affects the performance. It is also necessary to be able to select camera switching and zoom in on fingering via instructor Ub's voice command. It is also necessary to be able to display sheet music and a piano roll during performance, and to set up a pedal operation screen to be overlaid on top of the screen for necessary phrases. Therefore, when performance training is selected as a performance scene, high priority is assigned to the right hand attention processing E4, the left hand attention processing E5, and the pedal attention processing E6. In addition, in the performance training, the enlargement rate of the right hand HR in the right hand attention processing E4 is set to a higher value compared to normal performance, and the enlargement rate of the left hand HL in the left hand attention processing E5 is set to a higher value compared to normal performance.

[0136] Furthermore, the video editing unit 52 estimates the words spoken by the user U through speech recognition of the voice picked up by the sound collection device 22, and selects an imaging device that corresponds to a word that matches or is similar to the word. Therefore, for example, if the instructor Ub utters "Look at the hand from above," "Look from above," or "Upper camera," an imaging device installed above is selected. In a scene where it is effective to view the right hand from the side, if the instructor Ub utters "Right hand close-up" or "Zoom in from the side of the right hand," a side imaging device is selected and video processing for zooming in is performed. Furthermore, when the instructor Ub pronounces "pedal," an image of the pedal operation is superimposed on the screen during the corresponding phrase period. The instructor Ub can create a performance training video by playing while talking to the student.

[0137] The instructor Ub can create a performance training video using the basic data B. The basic data B includes control data for each phrase period that is used to control the camerawork optimal for performance training in accordance with the training piece.

[0138] The video editing unit 56 can generate an optimal framework using performance analysis, voice recognition, and basic data.

[0139] The video editing unit 56 edits the video of the students during the period when it is detecting the student's performance, without detecting the instructor's performance. It is important to be able to check the student's performance from multiple angles simultaneously. Hand movements are captured from above, below, left, and right, and edited so that they are displayed on the same screen. If a parent / guardian Uc is present in the classroom, video captured by the parent / guardian Uc's terminal device 30c can also be added. If the parent / guardian Uc has a scene in the classroom that they would like to prioritize recording, they can select which camera to prioritize from the connected devices using the terminal device 30c and specify the period. If no period is specified, recording is made for each phrase period of the performance.

[0140] The second embodiment also achieves the same effects as the first embodiment. Furthermore, in the second embodiment, an edited video V is generated by editing a plurality of recorded videos Ya under one of a plurality of editing conditions. Therefore, a variety of edited videos V can be generated using a plurality of recorded videos Ya as a common material with different editing conditions.

[0141] C: Third Embodiment As described above, performance data D is stored in the storage device 32 for each performance by the performer Ua. That is, multiple pieces of performance data D are accumulated in chronological order in the storage device 32 as the performer Ua practices playing the keyboard instrument 10. For example, multiple pieces of performance data D representing past performances by the performer Ua, such as practice at a music school, practice at the performer Ua's home, and performances at recitals, are stored in the storage device 32. The video editing unit 52 may generate an edited video V using the multiple pieces of performance data D representing past performances by the performer Ua.

[0142] The video editing unit 52 selects from the storage device 32 a plurality of pieces of performance data D corresponding to a performance of a specific piece of music (hereinafter referred to as a "target piece of music") from the plurality of pieces of performance data D stored in the storage device 32. The target piece of music is, for example, a piece of music that the performer Ua will play at a recital.

[0143] As illustrated in Fig. 24, the video editing unit 52 generates an edited video V by editing recorded videos Ya of multiple performance data D corresponding to the target song. Specifically, the video editing unit 52 generates the edited video V by selecting one of the multiple recorded videos Ya for each phrase period Q of the target song. For example, the multiple recorded videos Ya are selected sequentially in chronological order. Therefore, an edited video V is generated in which scenes of the performer Ua performing the target song are switched sequentially. Note that the conditions for switching between the recorded videos Ya are the same as those in the above-mentioned embodiments.

[0144] The third embodiment also achieves the same effects as the first embodiment. In the third embodiment, an edited video V is generated from a plurality of recorded videos Ya related to a target piece of music. Therefore, an edited video V suitable for recording, for example, a child's growth, can be automatically generated, showing the progress of the performer Ua's improvement in playing the target piece of music, or scenes of the performer Ua's past performances.

[0145] D: Fourth Embodiment The video editing unit 52 of the fourth embodiment edits a previously generated edited video V. For example, the video editing unit 52 edits the edited video V in response to an instruction from the user U via the operation device 35.

[0146] 25 is a schematic diagram of a screen (hereinafter referred to as an "editing screen 66") displayed on the display device 34 when editing an existing edited video V. The existing edited video V and a plurality of recorded videos Ya are displayed on the editing screen 66. For example, a plurality of recorded videos Ya corresponding to a target song are displayed on the display device 34 together with the edited video V on a common time axis.

[0147] The editing screen 66 can also display a musical score 67. The musical score 67 is displayed on the same time axis as the recorded video, and the performance sound can be played back simultaneously. A playback point 68 is placed at a specific point in the musical score 67. The playback point 68 is the point in the music piece that is the target of playback. By swiping the musical score 67 or the video, it is possible to fast forward or rewind and move to the desired playback point.

[0148] The user U can specify a specific range of the edited video V on the time axis and a specific range of a desired recorded video Ya among the multiple recorded videos Ya by operating the operation device 35. The video editing unit 52 replaces the range of the edited video V specified by the user U with the range of the recorded video Ya specified by the user U.

[0149] Specifically, the user U can adjust the phrase period specified by the video editing unit 52 by clicking on the time frame of the phrase period displayed overlaid on the recorded video or the musical score 67 and moving the time frame left or right.

[0150] Furthermore, the user U can specify a specific range of the edited video V and instruct the application of a desired visual effect by operating the operation device 35. For example, the user U can select one of a plurality of visual effect options (e.g., zoom in, zoom out, pan). The video editing unit 52 applies the visual effect selected by the user U to the range of the edited video V specified by the user U.

[0151] The fourth embodiment also achieves the same effects as the first embodiment. In the fourth embodiment, an existing edited video V is edited (re-edited), so that a suitable edited video V can be generated in accordance with the intentions or preferences of the user U.

[0152] E: Modifications Specific modifications that can be added to the above-mentioned embodiments are exemplified below. Two or more embodiments arbitrarily selected from the following examples may be combined as appropriate within the scope of not mutually contradicting each other.

[0153] (1) The method for generating the analysis information M is arbitrary, and for example, each recorded video Ya may be used to generate the analysis information M. Specifically, the music analysis unit 51 analyzes each of the multiple recorded videos Ya and generates the analysis information M according to the analysis results.

[0154] For example, if the right hand HR or left hand HL of the performer Ua moves to the right in the keyboard animation Ya1, the music analysis unit 51 determines that the performance is an ascending style, and includes information indicating the ascending style in the analysis information M. Therefore, the animation editing unit 52 executes the ascending follow-up process E2. That is, the ascending follow-up process E2 is executed according to the results of the analysis of the keyboard animation Ya1. Similarly, if the right hand HR or left hand HL of the performer Ua moves to the left in the keyboard animation Ya1, the descending follow-up process E3 is executed.

[0155] Furthermore, if the movement of the right hand HR is prominent in the keyboard animation Ya1, the music analysis unit 51 determines that attention should be paid to the right hand HR of the performer Ua, and includes information indicating attention to the right hand HR in the analysis information M. Therefore, the animation editing unit 52 executes right hand attention processing E4. That is, the right hand attention processing E4 is executed according to the results of the analysis of the keyboard animation Ya1. Similarly, if the movement of the left hand HL is prominent in the keyboard animation Ya1, the left hand attention processing E5 is executed.

[0156] In the above explanation, an example has been given in which each recorded video Ya is used in addition to the performance data D to generate the analysis information M, but in a form in which the analysis information M is generated based on the results of the analysis of each recorded video Ya, the use of the performance data D may be omitted.

[0157] (2) In a form in which reference data representing the musical score of a piece of music is stored in the storage device 32, the musical score 67 represented by the reference data may be displayed on the display device 34 together with the edited video V, as illustrated in FIG. 26 .

[0158] The score is displayed on a common timeline with the recorded video, and the performance sound can be played back simultaneously. A playback point 68 is placed at a specific point in the score 67. The playback point 68 is the point in the music piece that is the target for playback. The score scrolls from left to right in time with the performance. Swiping the score or video allows you to fast forward or rewind to move to the desired playback point.

[0159] The user can move the playback point 68 to any point in the musical score 67 by operating the operating device 35. When an instruction to move the playback point 68 is received from the user, the playback control unit 53 starts playing the portion of the edited video V that corresponds to the moved playback point 68. According to the above embodiment, the user can visually and intuitively grasp the correspondence between each portion of the edited video V and the musical score 67 of the music. Furthermore, the user can easily confirm the portion of the edited video V that corresponds to the desired point in the musical score 67.

[0160] (3) In the above-described embodiments, the video Va of the edited video V is composed only of the recorded video Ya. However, the edited video V may include elements other than the recorded video Ya recorded by the recording system 20. For example, as illustrated in FIG. 27 , the playback control unit 53 may display a guide image 69 that guides the performance of the keyboard instrument 10 together with the edited video V on the display device 34. FIG. 27 illustrates a state in which the position of the pedal video Ya4 in the edited video V is displayed as a partial video Vq. The playback control unit 53 displays the guide image 69 related to the partial video Vq on the display device 34. In FIG. 27 , the guide image 69 notifies the user of important points in operating the pedal 184.

[0161] Furthermore, in the above-described embodiments, one video from a plurality of recorded videos Ya is selected as the partial video Vq, but the partial video Vq may be composed of a combination of two or more videos from the plurality of recorded videos Ya (i.e., a multi-angle video). For example, as illustrated in Fig. 28, the partial video Vq of the edited video V may be composed of an arrangement of two recorded videos Ya (e.g., a keyboard video Ya1 and a performance video Ya2) selected for one phrase period Q. Furthermore, as illustrated in Fig. 29, the partial video Vq may be composed of an arrangement of one recorded video Ya (e.g., a performance video Ya2) selected from the plurality of recorded videos Ya and a recorded video Ya cut out from the recorded video Ya.

[0162] In the above-described embodiments, one of a plurality of recorded videos Ya is selected as the partial video Vq for each phrase period Q, but one of a plurality of videos cut out from one recorded video Ya may be selected as the partial video Vq for each phrase period Q. For example, as illustrated in FIG. 30 , the video editing unit 52 may sequentially select each of a plurality of videos cut out from different regions of one performance video Ya2 as the partial video Vq for each phrase period Q. As can be understood from the example of FIG. 30 , it is not necessary for there to be a plurality of recorded videos Ya that serve as material for the edited video V. In other words, it is not necessary to have a plurality of imaging devices 21 to generate the edited video V.

[0163] (4) The types of recorded videos Ya generated by the recording system 20 are not limited to the examples in the above-described embodiments (keyboard video Ya1, performance video Ya2, performance video Ya3, pedal video Ya4). For example, a recorded video Ya showing the operation of the sound generation mechanism 183 may be generated by the recording system 20. In a form in which the edited video V includes the recorded video Ya of the sound generation mechanism 183, it is possible to visually confirm how the sound generation mechanism 183 operates in response to the performance of the performer Ua.

[0164] (5) For example, in a scene such as a recital or street performance, the recorded video Ya recorded by the recording system 20 may include audience members unrelated to the performer Ua. From the perspective of protecting portrait rights, the playback control unit 53 may perform blurring processing, such as mosaic processing, on the audience members included in the recorded video Ya. Also, avatar images may be displayed superimposed on the audience members included in the recorded video Ya. Mosaic processing or avatar overlap display processing may be performed on only a portion of the face or other part that identifies the person.

[0165] (6) A recorded video Ya to be selected as a partial video Vq of the edited video V from among the multiple recorded videos Ya may be determined based on the results of voice recognition by the user U. For example, the user U may pronounce a phrase representing the installation location of the imaging device 21 (e.g., “above,” “to the right,” etc.) at any time while playing the keyboard instrument 10.

[0166] Each of the multiple recorded videos Ya is associated with a phrase representing the installation location of the recorded video Ya and stored in the storage device 32. For example, a phrase representing the installation location specified for each recorded video Ya on the setting screen 60 described with reference to Fig. 11 is stored in the storage device 32. Note that the phrase representing the installation location may be edited in response to an operation by the user U on the operation device 35, for example.

[0167] The video editing unit 52 estimates the phrase spoken by the user U by performing voice recognition on the sound collected by the sound collection device 22, and selects a recorded video Ya associated with a phrase that matches or is similar to the estimated phrase as a partial video Vq from the multiple recorded videos Ya. Therefore, for example, if the user U utters "above," the keyboard video Ya1 is selected as the partial video Vq of the edited video V, and if the user U utters "to the right," the performance video Ya2 is selected as the partial video Vq of the edited video V.

[0168] (7) In the above-described embodiments, the performer Ua plays the keyboard instrument 10, but the instrument played by the performer Ua is not limited to the keyboard instrument 10. In other words, the present disclosure can be applied to recording the performance of any type of instrument.

[0169] (8) In the above-described embodiments, the information processing system 100 includes a keyboard instrument 10 and a recording system 20. However, the keyboard instrument 10 and the recording system 20 may be omitted from the information processing system 100. For example, the keyboard instrument 10 and the recording system 20 may be installed in an acoustic space located remotely from the information processing system 100 (terminal device 30a). The performance data D generated by the keyboard instrument 10 and the recorded video Ya and recorded sound Yb data generated by the recording system 20 are transmitted to the information processing system 100 via a communication network such as the Internet. The operation of the terminal device 30a is the same as in the above-described embodiments. As can be understood from the above description, the "information processing system" in this disclosure may or may not include a musical instrument and a recording system. Therefore, the terminal device 30a alone in the above-described embodiments is also included in the "information processing system."

[0170] (9) The functions of the terminal device 30 exemplified in each of the above embodiments may be realized by a server device (e.g., a cloud server) that communicates with an information device such as a smartphone, a tablet terminal, or a personal computer.

[0171] (10) As mentioned above, the functions of the terminal device 30 exemplified above are realized through cooperation between one or more processors constituting the control device 31 and a program stored in the storage device 32. The program according to the present disclosure may be provided in a form stored on a computer-readable recording medium and installed on a computer. The recording medium may be, for example, a non-transitory recording medium, such as an optical recording medium (optical disk) such as a CD-ROM, but may also include any known form of recording medium, such as a semiconductor recording medium or a magnetic recording medium. Note that a non-transitory recording medium includes any recording medium other than a transient, propagating signal, and does not exclude volatile recording media. Furthermore, in a configuration in which a distribution device distributes a program via a communication network, the storage medium that stores the program in the distribution device corresponds to the non-transitory recording medium described above.

[0172] F: Supplementary Notes From the above-described exemplary embodiments, the following configurations can be understood, for example.

[0173] An information processing system according to one aspect (aspect 1) of the present disclosure includes a music analysis unit that identifies phrase periods of a piece of music by analyzing music data representing the piece of music, and a video editing unit that generates an edited video by editing one or more videos representing a performance of the piece of music according to the phrase periods.

[0174] According to the above aspect, an edited video is generated according to a phrase period identified by analyzing music data, so that an edited video suitable for a piece of music can be generated without placing an excessive burden on the user.

[0175] "Music data" is data in any format that represents a piece of music composed of a time sequence of multiple notes. For example, performance data that represents a performance of a piece of music by a user, or sheet music data that represents the musical score of a piece of music, are examples of "music data."

[0176] A "phrase period" is a time unit consisting of multiple notes in a piece of music. Specifically, a musical section (a phrase) is exemplified as a "phrase period." A phrase period is not necessarily uniquely determined, and multiple phrase period candidates may be identified for a single piece of music.

[0177] The "video" to be edited is, for example, a video recording of a musician playing a piece of music. "Editing" a video is an image process that uses the video as material to generate another video (edited video). Specifically, "editing" includes various processes such as cutting out a specific section of the video on the time axis, cutting out a specific area within the video, enlarging (zooming in) or reducing (zooming out) a specific area within the video, moving (panning, tilting) the cut-out area within the video, compositing multiple videos onto a single screen, and connecting multiple videos on the time axis.

[0178] Furthermore, "editing a video according to a phrase period" is, for example, a process of controlling conditions for editing a video when the conditions are synchronized with a phrase period. For example, when multiple videos are connected, it is assumed that the videos will be switched at the boundary (start point or end point) of the phrase period.

[0179] In a specific example (Aspect 2) of Aspect 1, the one or more moving images are multiple moving images recorded in parallel at different positions or angles, and the editing by the moving image editing unit includes processing for selecting different moving images from the multiple moving images during the phrase period and during another phrase period immediately following the phrase period. According to the above aspect, it is possible to generate a suitable edited moving image in which moving images are switched in synchronization with musical divisions of the music (i.e., the start or end point of a phrase period).

[0180] A specific example (aspect 3) of aspect 2 further includes an instrument that outputs performance data representing a musical piece performed by a performer as the music data, and a recording system that generates the plurality of videos by capturing images of the performance by the performer.

[0181] In a specific example (Aspect 4) of Aspect 3, the recording system includes a plurality of imaging devices that capture images of different parts of the performer, the music analysis unit generates further music information about the song by analyzing the music data, and the video editing unit assigns each of the plurality of imaging devices to a different part of the performer or the instrument. According to the above aspect, each part of the performer or the instrument can be appropriately captured by the plurality of imaging devices.

[0182] In a specific example (Aspect 5) of Aspect 3 or Aspect 4, the musical instrument includes a keyboard on which a plurality of keys are arranged, a sound generating mechanism that generates performance sounds by vibration of a sound source linked to operation of the keyboard, a detection device that detects operation of each of the plurality of keys, and a performance data generation unit that generates the performance data in accordance with the detection result by the detection device.

[0183] In a specific example (Aspect 6) of any of Aspects 1 to 5, the music analysis unit further identifies music information related to the song by analyzing the music data, and the video editing unit generates the edited video by editing the one or more videos in accordance with the phrase period and the music information. According to the above aspect, since the edited video reflects the music information of the song in addition to the phrase period of the song, it is possible to achieve a significant effect of generating an edited video that is suitable for the song.

[0184] "Music information" is musical information about a piece of music. Specifically, information that indicates the musical trends of a piece of music is exemplified as "music information." Examples of "music information" include (1) whether the piece of music is an ascending type, where the pitch tends to rise over time, or a descending type, where the pitch tends to fall over time; (2) the level of attention paid to the playing part of the piece of music (e.g., the right-hand part / left-hand part); (3) whether pedal operation is performed; (4) whether a particular playing technique, such as slur or staccato, is used; and (5) the dynamics of the performance (piano / forte, crescendo / decrescendo, etc.).

[0185] "Editing video according to music information" means changing video editing conditions according to music information. For example, during an ascending section of a song, a video of the performer's hands and keyboard is panned to the right, and during a descending section of a song, the video is panned to the left. Furthermore, during a section of a song with a large number of notes in the right-hand part, the video is zoomed in on the right-hand part, and during a section of a song with a large number of notes in the left-hand part, the video is zoomed in on the left-hand part.

[0186] In a specific example (aspect 7) of aspect 6, the editing by the video editing unit includes processing for cutting out a specific region from the one or more videos according to the music information. According to the above aspect, it is possible to generate an edited video that focuses on a specific region from the one or more videos according to the music information.

[0187] In a specific example (Aspect 8) of Aspect 7, the one or more videos include a keyboard video including a keyboard used to play the musical piece and both hands of a performer of the musical piece, the music information includes information indicating an ascending form or a descending form of the musical piece, and the video editing unit performs different processing when the music information indicates an ascending form and when the music information indicates a descending form. According to the above aspect, it is possible to generate edited videos corresponding to an increase (ascending) or decrease (descending) in pitch in the musical piece.

[0188] In a specific example (Aspect 9) of Aspect 8, the editing by the video editing unit includes a process of panning the keys in the keyboard video from a low range to a high range when the music information indicates an ascending pitch, and a process of panning the keys in the keyboard video from a high range to a low range when the music information indicates a descending pitch. According to the above aspect, an edited video can be generated that focuses on the performer's hands so as to follow the rise (ascending) or fall (descending) of the pitch in the music.

[0189] In a specific example (Aspect 10) of any of Aspects 7 to 9, the music information includes a performance difficulty level, the one or more videos include performance videos that include a face of a performer of the music piece, and the editing by the video editing unit includes a process of selecting the performance videos when the difficulty level exceeds a threshold. According to the above aspects, it is possible to generate an edited video that emphasizes the facial expression of a performer playing a section of the music piece that is difficult.

[0190] The "difficulty" of a performance is an index of the difficulty of playing a piece of music. For example, the "difficulty" is calculated based on various indicators such as the number of notes in a unit period such as a measure, the pitch difference between two adjacent notes, the number of notes that make up a chord, and the number of notes that correspond to black keys.

[0191] In a specific example (Aspect 11) of any of Aspects 7 to 10, the one or more videos include a keyboard video including a keyboard used to play the musical piece and both hands of a performer of the musical piece, the music information includes an attention level for a right-hand part and an attention level for a left-hand part of the musical piece, and the editing by the video editing unit includes a process of enlarging the performer's right hand in the keyboard video if the attention level for the right-hand part exceeds a threshold, and a process of enlarging the performer's left hand in the keyboard video if the attention level for the left-hand part exceeds a threshold. According to the above aspects, it is possible to generate an edited video that emphasizes the performer's right hand or left hand, whichever is more highly regarded in the musical piece.

[0192] The "attention level" of the right-hand (left-hand) part is an index that represents the degree to which attention should be paid to the performer's right hand (left hand) in a piece of music. For example, an index such as the number of notes in the right-hand part of a piece of music is used as the "attention level" of the right-hand part.

[0193] In a specific example (Aspect 12) of any of Aspects 4 to 11, the one or more videos include a pedal video including a pedal of a keyboard instrument, the music information includes whether or not the pedal is operated, and the editing by the video editing unit includes processing for selecting the pedal video when the music information indicates the pedal operation. According to the above aspect, it is possible to generate an edited video that focuses on the pedal operation by the performer.

[0194] In a specific example (Aspect 13) of any of Aspects 1 to 12, the video editing unit edits the one or more videos under any of a plurality of different editing conditions. According to the above aspect, an edited video is generated by editing one or more videos under any of a plurality of editing conditions. Therefore, a variety of edited videos can be generated using one or more videos as a common material, with different editing conditions.

[0195] The "editing conditions" are conditions for editing one or more videos. Specifically, examples of the "editing conditions" include various conditions that affect the editing by the video editing unit, such as conditions regarding the length of a phrase period in a song, conditions regarding the priority of each of multiple videos, etc.

[0196] The method of selecting the editing condition to be actually applied to the editing by the video editing unit from among the multiple editing conditions is arbitrary. For example, one of the multiple editing conditions may be selected in response to an instruction from a user, or one of the multiple editing conditions may be selected randomly.

[0197] A video editing method according to one aspect (aspect 14) of the present disclosure involves analyzing music data representing a song to identify music information about the song and phrase periods of the song, and generating an edited video by editing one or more videos representing a performance of the song according to the music information and the phrase periods.

[0198] A program according to one aspect (aspect 15) of the present disclosure causes a computer system to function as a music analysis unit that identifies music information about a song and phrase periods of the song by analyzing music data representing the song, and a video editing unit that generates an edited video by editing one or more videos representing a performance of the song according to the music information and the phrase periods.

[0199] 100...information processing system, 10...keyboard instrument, 11...control device, 12...storage device, 13...communication device, 14...display device, 15...operation device, 16...sound source device, 17...sound emission device, 18...performance mechanism, 181...keyboard, 182...key, 183...sound generation mechanism, 184...pedal, 185...detection device, 20...recording system, 21...imaging device, 22...sound collection device, 30 (30a, 30b, 30c)...terminal device, 31...control device, 32...storage Storage device, 33...communication device, 34...display device, 35...operation device, 36...sound emission device, 37...sound collection device, 40...measurement system, 41...biometric sensor, 42...center of gravity sensor, 51...music analysis unit, 52...video editing unit, 53...playback control unit, 60...setting screen, 61...guidance screen, 62...device icon, 62a...imaging icon, 62b...sound collection icon, 63...scene selection window, 64...device setting window, 90...recording system.

Claims

1. An information processing system comprising: a music analysis unit that identifies a phrase period of a piece of music by analyzing music data representing the piece of music; and a video editing unit that generates an edited video by editing one or more videos representing the performance of the piece of music according to the phrase period.

2. The information processing system according to claim 1, wherein the one or more videos are a plurality of videos recorded in parallel at different positions or angles, and the editing by the video editing unit includes a process of selecting different videos from the plurality of videos during the phrase period and another phrase period immediately after the phrase period.

3. The information processing system according to claim 2, further comprising: a musical instrument that outputs the music data representing the performance of the piece of music by a performer; and a recording system that generates the plurality of videos by imaging the performance by the performer.

4. The information processing system according to claim 3, wherein the recording system includes a plurality of imaging devices, the music analysis unit further generates music information regarding the piece of music by analyzing the music data, and the video editing unit assigns each of the plurality of imaging devices to each part of the performer or the musical instrument.

5. The information processing system according to claim 3, wherein the musical instrument includes: a keyboard having a plurality of keys arranged; a detection device that detects an operation on each of the plurality of keys; and a performance data generation unit that generates the performance data according to a result of the detection by the detection device.

6. The information processing system according to claim 1, wherein the music analysis unit further identifies music information regarding the piece of music by analyzing the music data, and the video editing unit generates the edited video by editing the one or more videos according to the phrase period and the music information.

7. The information processing system according to claim 6, wherein the editing by the video editing unit includes a process of cutting out a specific area according to the music information from the one or more videos.

8. The information processing system according to claim 7, wherein the one or more videos include a keyboard video including a keyboard used for the performance of the piece of music and both hands of the performer of the piece of music, the music information includes information indicating an ascending or descending form in the piece of music, and the video editing unit executes different processes when the music information indicates an ascending form and when the music information indicates a descending form.

9. The editing by the video editing unit includes, when the music information indicates an ascending form, a process of panning the keyboard in the keyboard video from the low pitch range to the high pitch range, and when the music information indicates a descending form, a process of panning the keyboard in the keyboard video from the high pitch range to the low pitch range. The information processing system according to claim 8.

10. The music information includes the difficulty level of performance. The one or more videos include a performance video including the face of the performer of the music piece. The editing by the video editing unit includes a process of selecting the performance video when the difficulty level exceeds a threshold value. The information processing system according to any one of claims 7 to 9.

11. The one or more videos include a keyboard video including the keyboard used in the performance of the music piece and both hands of the performer of the music piece. The music information includes the attention level of the right hand part and the attention level of the left hand part of the music piece. The editing by the video editing unit includes, when the attention level of the right hand part exceeds a threshold value, a process of enlarging the right hand of the performer in the keyboard video, and when the attention level of the left hand part exceeds a threshold value, a process of enlarging the left hand of the performer in the keyboard video. The information processing system according to claim 7.

12. The one or more videos include a pedal video including a pedal in a keyboard instrument. The music information includes the presence or absence of an operation on the pedal. The editing by the video editing unit includes a process of selecting the pedal video when the music information indicates an operation on the pedal. The information processing system according to claim 4.

13. The video editing unit edits the one or more videos according to any one of a plurality of different editing conditions. The information processing system according to claim 1. Video editing method realized by a computer system that identifies music information related to a music piece and the phrase period of the music piece by analyzing music data representing the music piece, and generates an edited video by editing one or more videos representing the performance of the music piece according to the music information and the phrase period. A program that causes a computer system to function as a music analysis unit that identifies music information related to a music piece and the phrase period of the music piece by analyzing music data representing the music piece, and a video editing unit that generates an edited video by editing one or more videos representing the performance of the music piece according to the music information and the phrase period.

Citation Information

Patent Citations

  • Music playing data processing method and musical sound signal synthesizing method

    JP2004070154A

  • Live video processing system, live video processing method, and program

    JP2018170678A

  • Musical performance lesson system

    JP2022039020A

  • Image output device, image output method, and program

    WO2022249555A1

  • Signal processing device and signal processing method

    WO2023139883A1