Performance analysis method, performance analysis system, and program

The performance analysis method and system address the challenge of synchronizing video and audio data by detecting changes in percussion instruments to generate precise temporal references, ensuring accurate alignment and playback.

JP7859038B2Active Publication Date: 2026-05-15YAMAHA CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
YAMAHA CORP
Filing Date
2021-11-08
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing techniques for synchronizing video and acoustic data in musical instrument performances face challenges in generating precise temporal references, making it difficult to accurately align video and audio data on the time axis.

Method used

A performance analysis method and system that includes acquiring video data of a percussion instrument, detecting changes in the instrument through video analysis, generating performance data representing the performance, and creating rhythmic data from the detected changes to synchronize video and audio data accurately.

Benefits of technology

The method and system enable high-precision synchronization of video and audio data by generating performance data that serves as a temporal reference, allowing for accurate alignment and playback of synchronized video and audio data, particularly for percussion instruments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007859038000001
    Figure 0007859038000001
  • Figure 0007859038000002
    Figure 0007859038000002
  • Figure 0007859038000003
    Figure 0007859038000003
Patent Text Reader

Abstract

To generate data to be used as a temporal reference for video data, from the video data.SOLUTION: A performance analysis system 40 includes: a video data acquisition unit 51 that acquires video data X generated by capturing an image of a percussion instrument; an analysis processing unit 53 that analyzes the video data X, thereby detecting vibrations caused by performance in the percussion instrument; a performance data generation unit 54 that generates performance data Q representing the performance in accordance with a result of detection; and a metrical data generation unit 56 that generates metrical data R representing a metrical structure from the performance data Q.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a technique for analyzing musical instrument performances.

Background Art

[0002] Various techniques for processing video data representing a video of a musical instrument performance have been conventionally proposed. For example, Patent Document 1 discloses a configuration for synchronizing video data with acoustic data representing the performance sound of a musical instrument. For the synchronization of video data and acoustic data, reference information such as a time code is used.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the technique of Patent Document 1, it is necessary to generate reference information independently of the video data. However, it is not practically easy to generate reference information that serves as a temporal reference for video data with high precision. In the above description, the case of synchronizing video data and acoustic data has been exemplified, but similar problems are assumed in various scenarios where video data is processed on the time axis. In consideration of the above circumstances, one aspect of this disclosure aims to generate data serving as a temporal reference for the performance of a percussion instrument from video data.

Means for Solving the Problems

[0005] To solve the above problems, a performance analysis method according to one aspect of the present disclosure includes acquiring video data generated by imaging a percussion instrument, detecting changes in the percussion instrument due to performance by analyzing the video data, generating performance data representing the performance according to the results of the detection, and generating rhythmic data representing the rhythmic structure from the performance data.

[0006] A performance analysis method according to another aspect of the present disclosure includes acquiring video data generated by imaging a percussion instrument, generating performance data representing the performance of the percussion instrument by processing the video data, and generating rhythmic data representing the rhythmic structure from the performance data.

[0007] A performance analysis system according to one aspect of the present disclosure comprises: a video data acquisition unit that acquires video data generated by imaging a percussion instrument; an analysis processing unit that detects changes in the percussion instrument due to performance by analyzing the video data; a performance data generation unit that generates performance data representing the performance according to the results of the detection; and a rhythm data generation unit that generates rhythm data representing the rhythmic structure from the performance data.

[0008] A program according to one aspect of this disclosure causes a computer system to function as follows: a video data acquisition unit that acquires video data generated by imaging a percussion instrument; an analysis processing unit that analyzes the video data to detect changes in the percussion instrument due to performance; a performance data generation unit that generates performance data representing the performance according to the results of the detection; and a rhythm data generation unit that generates rhythm data representing the rhythmic structure from the performance data. [Brief explanation of the drawing]

[0009] [Figure 1] This is a block diagram illustrating the configuration of the information processing system in the first embodiment. [Figure 2] This is a block diagram illustrating the configuration of a performance analysis system. [Figure 3] This is a block diagram illustrating the functional configuration of a performance analysis system. [Figure 4] This flowchart illustrates the detailed steps of the performance detection process. [Figure 5] This is an explanatory diagram of the processing performed by the analysis processing unit and the performance data generation unit. [Figure 6] This flowchart illustrates the detailed steps of the synchronous control process. [Figure 7] This flowchart illustrates the detailed steps involved in the performance analysis process. [Figure 8] This flowchart illustrates the detailed procedure for the performance detection process in the second embodiment. [Figure 9] This is a block diagram illustrating the functional configuration of the performance analysis system in the third embodiment. [Figure 10] This flowchart illustrates the detailed procedure of the synchronous control process in the third embodiment. [Figure 11] This flowchart illustrates the detailed procedure for the performance analysis process in the third embodiment. [Figure 12] This is a block diagram illustrating the functional configuration of the performance analysis system in the fourth embodiment. [Figure 13] This flowchart illustrates the detailed procedure for the performance analysis process in the fourth embodiment. [Figure 14] This is a block diagram illustrating the functional configuration of the performance analysis system in the fifth embodiment. [Figure 15] This flowchart illustrates the detailed procedure for the synchronization adjustment process in the fifth embodiment. [Figure 16] This flowchart illustrates the detailed procedure for the performance analysis process in the fifth embodiment. [Figure 17] This is an explanatory diagram of the trained model in the sixth embodiment. [Figure 18] This is a block diagram illustrating the functional configuration of the performance analysis system in a modified example. [Figure 19] It is a block diagram illustrating a functional configuration of a performance analysis system in a modified example.

Mode for Carrying Out the Invention

[0010] A: First Embodiment FIG. 1 is a block diagram illustrating a configuration of an information processing system 100 according to the first embodiment. The information processing system 100 is a computer system for recording and analyzing the performance of the percussion instrument 1 by the user U.

[0011] The percussion instrument 1 includes a drum set 10 and a foot pedal 12. The drum set 10 is composed of a plurality of drums including a bass drum 11. The bass drum 11 is a percussion instrument having a body portion 111 and a head 112. The body portion 111 is a cylindrical structure (shell). The head 112 is a plate-like elastic member that closes the opening of the body portion 111. Note that the opening of the body portion 111 on the side opposite to the head 112 is closed by a rear head, but the illustration of the rear head is omitted in FIG. 1. The user U plays the percussion part in a music piece by hitting the head 112 using the foot pedal 12. Note that the head 112 may be a mesh head for muffling. That is, it is not necessary for the opening of the body portion 111 to be completely sealed.

[0012] The foot pedal 12 includes a beater 121 and a pedal 122. The beater 121 is a striking body that hits the bass drum 11. The pedal 122 receives the depression by the user U. In conjunction with the depression of the pedal 122 by the user U, the beater 121 hits the head 112. The head 112 vibrates by the hit of the beater 121. That is, the head 112 is a vibrating body that vibrates by the performance of the user U. Also, the main performer of the drum set 10 is not limited to the user U. For example, a performance robot capable of performing automatic performance of a music piece may play the drum set 10.

[0013] The information processing system 100 comprises a recording device 20, a recording device 30, and a performance analysis system 40. The performance analysis system 40 is a computer system for analyzing the performance of percussion instrument 1 by user U. The performance analysis system 40 communicates with each of the recording devices 20 and 30. Communication between the performance analysis system 40 and the recording device 20 or 30 is by short-range wireless communication such as Wi-Fi (registered trademark) or Bluetooth (registered trademark). However, the performance analysis system 40 may communicate with the recording device 20 or 30 by wire. Alternatively, the performance analysis system 40 may be implemented by a server device that communicates with the recording devices 20 and 30 via a communication network such as the Internet.

[0014] Recording device 20 and recording device 30 each record user U's performance on drum set 10. Recording devices 20 and 30 are installed at different positions and angles relative to drum set 10.

[0015] The recording device 20 comprises an imaging device 21 and a communication device 22. The imaging device 21 generates video data X by imaging the user U playing the percussion instrument 1. That is, the video data X is generated by imaging the percussion instrument 1. The range imaged by the imaging device 21 includes the head 112 of the bass drum 11. Therefore, the image represented by the video data X includes the head 112. The imaging device 21 comprises, for example, an optical system such as a photographic lens, an image sensor that receives incident light from the optical system, and a processing circuit that generates video data X according to the amount of light received by the image sensor. The imaging device 21 starts and stops recording in response to instructions from the user U. That is, imaging by the imaging device 21 is started and stopped in response to instructions from the user U. The image represented by the video data X may include only a part of the bass drum 11, or it may include drums other than the bass drum 11 in the drum set 10, or it may include instruments other than the drum set 10. Furthermore, an operator other than user U may instruct the imaging device 21 to start or stop recording.

[0016] The communication device 22 transmits the video data X to the performance analysis system 40. For example, an information device such as a smartphone, tablet terminal, or personal computer may be used as the recording device 20. However, video equipment such as a video camera dedicated to recording may also be used as the recording device 20. Note that the imaging device 21 and the communication device 22 may be separate devices.

[0017] The recording device 30 comprises a sound pickup device 31 and a communication device 32. The sound pickup device 31 picks up ambient sounds. Specifically, the sound pickup device 31 generates sound data Y by picking up the sound of a percussion instrument 1 (drum set 10) being played. The sound is the musical sound produced by the percussion instrument 1 when played by user U. For example, the sound pickup device 31 comprises a microphone that generates an acoustic signal by picking up sounds, and a processing circuit that generates sound data Y from the acoustic signal. The sound pickup device 31 starts and stops recording in response to instructions from user U. However, an operator other than user U may also instruct the sound pickup device 31 to start or stop recording.

[0018] The communication device 32 transmits the acoustic data Y to the performance analysis system 40. For example, an information device such as a smartphone, tablet terminal, or personal computer may be used as the recording device 30. Alternatively, an acoustic device such as a standalone microphone may be used as the recording device 30. Furthermore, the sound pickup device 31 and the communication device 32 may be separate devices.

[0019] Image capture by the imaging device 21 and sound recording by the sound recording device 31 are performed in parallel with the user U's performance of the drum set 10. That is, video data X and audio data Y are generated in parallel for the same song. The performance analysis system 40 generates composite data Z by combining the video data X and audio data Y. Specifically, the composite data Z represents a video that includes the video represented by the video data X and the sound represented by the audio data Y.

[0020] Assuming the synthesis of video data X and audio data Y, it is desirable that the imaging device 21 and the sound recording device 31 simultaneously begin recording before the start of the percussion instrument 1 performance and simultaneously end recording after the performance ends. However, the start and end of recording are instructed individually to the imaging device 21 and the sound recording device 31. Therefore, the start and end times of recording may differ between the imaging device 21 and the sound recording device 31. In other words, the position on the time axis may differ between the video represented by video data X and the performance sound represented by audio data Y. Against this backdrop, the performance analysis system 40 synchronizes the video data X and the audio data Y with each other on the time axis.

[0021] Figure 2 is a block diagram illustrating the configuration of the performance analysis system 40. The performance analysis system 40 comprises a control device 41, a storage device 42, a communication device 43, an operating device 44, a display device 45, and a sound emission device 46. The performance analysis system 40 can be implemented as a single device or as multiple devices configured separately from each other. The recording device 20 or recording device 30 may be mounted on the performance analysis system 40.

[0022] The control device 41 consists of one or more processors that control each element of the performance analysis system 40. For example, the control device 41 consists of one or more types of processors such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), SPU (Sound Processing Unit), DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit).

[0023] The communication device 43 communicates with the recording device 20 and the recording device 30, respectively. Specifically, the communication device 43 receives video data X transmitted from the recording device 20 and audio data Y transmitted from the recording device 30.

[0024] The storage device 42 is one or more memories that store programs executed by the control device 41 and various data used by the control device 41. For example, video data X and audio data Y received by the communication device 43 are stored in the storage device 42. The storage device 42 is composed of a known recording medium such as a magnetic recording medium or a semiconductor recording medium, or a combination of multiple types of recording media. A portable recording medium that can be attached to and detached from the performance analysis system 40 may be used as the storage device 42. Alternatively, a recording medium that can be written to or read by the control device 41 via a communication network such as the Internet (e.g., cloud storage) may be used as the storage device 42.

[0025] The operating device 44 is an input device that receives instructions from user U. The operating device 44 is, for example, an operator operated by user U, or a touch panel that detects contact by user U. Alternatively, a separate operating device 44 (e.g., a mouse or keyboard) may be connected to the performance analysis system 40 by wire or wireless connection. Furthermore, an operator other than user U, who plays the percussion instrument 1, may operate the operating device 44.

[0026] The display device 45 displays various images under the control of the control device 41. For example, the display device 45 displays the image represented by the image data X of the composite data Z. Various display panels such as liquid crystal display panels or organic EL (electroluminescence) panels can be used as the display device 45. Alternatively, a separate display device 45 may be connected to the performance analysis system 40 by wire or wireless connection.

[0027] The sound emission device 46 reproduces the sound represented by the acoustic data Y in the synthesized data Z. The sound emission device 46 is, for example, a speaker or headphones. Alternatively, a separate sound emission device 46 may be connected to the performance analysis system 40 by wire or wireless connection. As can be understood from the above description, the display device 45 and the sound emission device 46 function as a playback device 47 that reproduces the synthesized data Z.

[0028] Figure 3 is a block diagram illustrating the functional configuration of the performance analysis system 40. The control device 41 executes a program stored in the storage device 42 to realize multiple functions (video data acquisition unit 51, sound data acquisition unit 52, analysis processing unit 53, performance data generation unit 54, synchronization control unit 55) for generating the synthesized data Z.

[0029] The video data acquisition unit 51 acquires video data X. Specifically, the video data acquisition unit 51 receives the video data X transmitted by the recording device 20 via the communication device 43. The audio data acquisition unit 52 acquires audio data Y. Specifically, the audio data acquisition unit 52 receives the audio data Y transmitted by the recording device 30 via the communication device 43.

[0030] The analysis processing unit 53 detects vibrations generated in the bass drum 11 by playing by analyzing the video data X. Specifically, the analysis processing unit 53 detects vibrations of the head 112 on the bass drum 11. Figure 4 is a flowchart illustrating the detailed procedure of the process by which the analysis processing unit 53 detects vibrations of the bass drum 11 (hereinafter referred to as the "play detection process").

[0031] When the performance detection process is started, the analysis processing unit 53 identifies the region where the bass drum 11 is located (hereinafter referred to as the "target region") from the video data X (Sa31). The target region is the region of the head 112 of the bass drum 11. The target region can also be described as the region that vibrates when the percussion instrument 1 is played. Any known object detection process can be arbitrarily used to identify the target region. For example, object detection processing using a deep neural network (DNN), such as a convolutional neural network (CNN), can be used to identify the target region.

[0032] The analysis processing unit 53 detects vibrations of the head 112 in accordance with changes in the image in the target region (Sa32). Specifically, as illustrated in Figure 5, the analysis processing unit 53 calculates a feature quantity F of the image in the target region and detects vibrations in accordance with the temporal change of the feature quantity F. Feature quantity F is an index that represents the characteristics of the image represented by the image data X. For example, feature quantity F is information that represents the optical characteristics of the image, such as the average value of the gradation (luminance) in the target region. The amount of reflected light that reaches the imaging device 21 from the head 112 of the bass drum 11 changes due to the vibration of the head 112. The analysis processing unit 53 detects the time point τ at which the amount of change in feature quantity F (e.g., increase or decrease) exceeds a predetermined threshold as the time point of vibration of the head 112. The head 112 of the bass drum 11 vibrates each time it is struck by the beater 121. Therefore, the time point τ of vibrations that the analysis processing unit 53 sequentially identifies corresponds to the time when the user U strikes the drum set 10 with the beater 121. Furthermore, the amplitude of vibrations generated in the head 112 depends on the force with which the user U strikes the bass drum 11 (hereinafter referred to as "strike intensity"). Therefore, the amount of change in the amount of reflected light reaching the imaging device 21 from the head 112 of the bass drum 11 depends on the strike intensity. Taking these relationships into consideration, the analysis processing unit 53 calculates the strike intensity according to the change in feature quantity F. For example, the analysis processing unit 53 sets the strike intensity to a larger value the larger the change in feature quantity F. As described above, the performance detection process includes a process to identify the target area of ​​the bass drum 11 from the image represented by the video data X (Sa31), and a process to detect vibrations of the head 112 according to changes in the image in the target area (Sa32). Note that the type of feature quantity F is not limited to the examples above. For example, the analysis processing unit 53 may extract feature points of the percussion instrument 1 by analyzing the video data X and calculate feature quantities F related to the movement of those feature points. For example, the speed or acceleration of the movement of the feature points may be calculated as feature quantity F. The calculation of the feature quantities F exemplified above utilizes known techniques such as optical flow. Furthermore, the feature points of percussion instrument 1 are characteristic points extracted from the video of percussion instrument 1 through predetermined image processing of the video data X.

[0033] The performance data generation unit 54 in Figure 3 generates performance data Q representing the performance of percussion instrument 1 by user U, according to the results of detection by the analysis processing unit 53. As illustrated in Figure 5, performance data Q is time-series data consisting of sound data q1 representing the sound of the drum set 10 and time point data q2 specifying the time of the sound. Sound data q1 is event data specifying the strike intensity detected by the analysis processing unit 53. Time point data q2 specifies the time of each sound of the drum set 10, for example, by the time interval between successive sounds, or by the elapsed time from the start of the performance of percussion instrument 1. The performance data generation unit 54 generates performance data Q that specifies the time point τ of vibration detected from video data X as the time of sound of the drum set 10 (hereinafter referred to as "sound point"). Performance data Q is time-series data in a format compliant with the MIDI standard, for example.

[0034] The synchronization control unit 55 in Figure 3 synchronizes the video data X and the audio data Y using the performance data Q. Figure 6 is a flowchart illustrating the detailed procedure of the process by which the synchronization control unit 55 synchronizes the video data X and the audio data Y (hereinafter referred to as the "synchronization control process").

[0035] When the synchronization control process is initiated, the synchronization control unit 55 identifies the point of sound production of the bass drum 11 by analyzing the acoustic data Y (Sa71). For example, the synchronization control unit 55 sequentially identifies the point of sound production where the increase in volume exceeds a predetermined value in the acoustic data Y. Note that known beat tracking techniques can be optionally used to identify the point of sound production using the acoustic data Y. The procedure for the synchronization control process is arbitrary, and processes such as beat tracking are not mandatory.

[0036] The synchronization control unit 55 synchronizes the video data X and the audio data Y using the performance data Q (Sa72). Specifically, the synchronization control unit 55 determines the position of the audio data Y on the time axis relative to the video data X so that each sound point specified by the performance data Q and each sound point identified from the audio data Y coincide on the time axis. As can be understood from the above explanation, synchronization of video data X and audio data Y means adjusting the position of one on the time axis relative to the other so that the sound represented by the audio data Y and the video represented by the video data X at any given point in the music correspond to each other on the time axis. Therefore, the processing by the synchronization control unit 55 can also be described as the processing of adjusting the temporal correspondence between video data X and audio data Y. As explained above, according to the first embodiment, it is possible to synchronize individually prepared video data X and audio data Y with each other.

[0037] The synchronization control unit 55 generates composite data Z which includes mutually synchronized video data X and audio data Y (Sa73). The composite data Z is played back by the playback device 47. As described above, in the composite data Z, the video data X and audio data Y are mutually synchronized. Therefore, when the video of a specific part of the music in the video data X is displayed by the display device 45, the sound of that part of the audio data Y is played back by the sound output device 46.

[0038] Figure 7 is a flowchart illustrating the detailed procedure of the process performed by the control device 41 (hereinafter referred to as the "performance analysis process"). For example, the performance analysis process is initiated by an instruction from the user U to the operating device 44. The performance analysis process in Figure 7 is an example of a "performance analysis method".

[0039] When the performance analysis process begins, the control device 41 functions as a video data acquisition unit 51 to acquire video data X (S1). The control device 41 also functions as an audio data acquisition unit 52 to acquire audio data Y (S2).

[0040] The control device 41 performs the aforementioned performance detection process (S3). Specifically, the control device 41 detects vibrations of the drum set 10 (head 112) by analyzing the video data X. In other words, the control device 41 functions as an analysis processing unit 53. The control device 41 generates performance data Q using the results of the performance detection process (S4). In other words, the control device 41 functions as a performance data generation unit 54.

[0041] The control device 41 performs the aforementioned synchronization control process (S7). Specifically, the control device 41 generates composite data Z by synchronizing the video data X and the audio data Y using the performance data Q. In other words, the control device 41 functions as a synchronization control unit 55. The control device 41 plays the composite data Z back using the playback device 47 (S9).

[0042] As described above, in the first embodiment, vibrations of the bass drum 11 (head 112) are detected by analyzing the video data X generated by imaging the percussion instrument 1, and performance data Q representing the performance of the bass drum 11 is generated according to the result of this detection. In other words, performance data Q, which serves as a temporal reference for the performance of the percussion instrument 1, can be generated from the video data X.

[0043] In general, the bass drum 11 is played in a fixed position. On the other hand, instruments other than percussion instruments, such as string instruments or wind instruments (hereinafter referred to as "non-percussion instruments"), move constantly in response to the movement or changes in posture of the performer. That is, the bass drum 11 tends to be less likely to move compared to, for example, non-percussion instruments. Therefore, according to the first embodiment of analyzing the video data X of the bass drum 11, there is also the advantage that the load required to generate performance data Q is reduced compared to the case where performance data is generated by analyzing the video data of non-percussion instruments.

[0044] Furthermore, in the first embodiment, the target area of ​​the bass drum 11 is identified from the video data X. Therefore, compared to a method of detecting vibrations without identifying a target area, the vibration of the bass drum 11 can be detected with high accuracy. As mentioned above, the bass drum 11 tends to be less prone to movement of the instrument itself compared to non-percussion instruments. Therefore, the target area where the bass drum 11 is located can be easily and accurately identified from the video data X. In other words, by targeting the bass drum 11, the processing load for detecting vibrations is reduced.

[0045] B: Second Embodiment A second embodiment will now be described. For elements whose function is the same as in the first embodiment in each of the embodiments described below, the same reference numerals as in the first embodiment will be used, and detailed descriptions of each will be omitted as appropriate.

[0046] In the first embodiment, an example was given in which the imaging device 21 images the bass drum 11. In the second embodiment, the imaging device 21 generates video data X by imaging the foot pedal 12 used to play the bass drum 11. In either the first or second embodiment, an embodiment in which the imaging device 21 images both the bass drum 11 and the foot pedal 12 is also conceivable.

[0047] The configuration of the performance analysis system 40 is the same as in the first embodiment (Figure 3). The control device 41 executes a program stored in the storage device 42 to realize multiple functions (video data acquisition unit 51, sound data acquisition unit 52, analysis processing unit 53, performance data generation unit 54, synchronization control unit 55) for generating synthesized data Z, similar to the first embodiment. The video data acquisition unit 51 acquires video data X, similar to the first embodiment. The sound data acquisition unit 52 acquires sound data Y, similar to the first embodiment.

[0048] As described above, the analysis processing unit 53 of the first embodiment detects vibrations generated in the bass drum 11 by playing. The analysis processing unit 53 of the second embodiment detects the impact of the beater 121 of the foot pedal 12 on the bass drum 11 by analyzing the video data X generated by imaging the foot pedal 12. Specifically, the analysis processing unit 53 detects the impact of the beater 121 on the bass drum 11 by the playing detection process illustrated in Figure 8. In other words, the playing detection process in Figure 3 of the first embodiment is replaced by the playing detection process in Figure 8 of the second embodiment.

[0049] When the performance detection process is started, the analysis processing unit 53 detects the beater 121 from the video data X (Sb31). Any known object detection process can be arbitrarily used to identify the beater 121. For example, object detection using a deep neural network such as a convolutional neural network can be used to identify the beater 121.

[0050] The analysis processing unit 53 detects the impact of the beater 121 on the drum set 10 in accordance with the change in the position of the beater 121 detected from the video data X (Sb32). Specifically, the analysis processing unit 53 detects the moment when the movement of the beater 121 reverses from a predetermined direction to the opposite direction as the moment of impact by the beater 121. Furthermore, the impact intensity by the user U depends on the movement speed of the beater 121. Taking these relationships into consideration, the analysis processing unit 53 calculates the impact intensity according to the movement speed of the beater 121 detected from the video data X. For example, the analysis processing unit 53 sets the impact intensity to a larger value the greater the movement speed of the beater 121. As described above, the performance detection process of the second embodiment includes a process for detecting the beater 121 from the video represented by the video data X (Sb31) and a process for detecting impact in accordance with the change in the position of the beater 121 (Sb32).

[0051] The performance data generation unit 54 of the second embodiment generates performance data Q representing the performance of the percussion instrument 1 by the user U, in accordance with the results of the detection by the analysis processing unit 53, similar to the first embodiment. Specifically, the performance data generation unit 54 generates performance data Q that designates the timing of the strike detected from the video data X as the sound-producing point of the bass drum 11. Similar to the first embodiment, the performance data Q consists of sound-producing data q1 that specifies the intensity of the strike and time-specific data q2 that specifies the timing of the sound.

[0052] The synchronization control unit 55 synchronizes the video data X and the audio data Y using the performance data Q. Specifically, the synchronization control unit 55 synchronizes the video data X and the audio data Y using the same synchronization control process as in the first embodiment (Figure 6).

[0053] The performance analysis process in the second embodiment is the same as the performance analysis process in the first embodiment illustrated in Figure 7. However, in the second embodiment, as described above, the performance detection process in Figure 3 in the performance analysis process is replaced with the performance detection process in Figure 8.

[0054] As described above, in the second embodiment, the impact of the beater 121 is detected by analyzing the video data X generated by imaging the beater 121, and performance data Q representing the performance of the bass drum 11 is generated according to the result of the detection. In other words, performance data Q, which serves as a temporal reference for the video data X, can be generated from the video data X.

[0055] C: Third Embodiment Figure 9 is a block diagram illustrating the functional configuration of the performance analysis system 40 in the third embodiment. The control device 41 in the third embodiment functions as a beat data generation unit 56 in addition to the same elements as in the first embodiment (video data acquisition unit 51, sound data acquisition unit 52, analysis processing unit 53, performance data generation unit 54, synchronization control unit 55) by executing a program stored in the storage device 42.

[0056] The rhythm data generation unit 56 generates rhythm data R from the performance data Q. Rhythm data R is data representing the rhythmic structure of a musical piece performed using the percussion instrument 1. Rhythm structure refers to the structure of the beats in a musical piece. Specifically, rhythmic structure is the structure of a rhythmic pattern (time signature) defined by a combination of multiple beats, such as strong beats or weak beats, and the timing at which each beat occurs. Typically, rhythmic structure is repeated periodically within a musical piece, such as every measure, but repetition is not mandatory. The rhythm data generation unit 56 generates rhythmic data R by analyzing the performance data Q. Specifically, the rhythm data generation unit 56 distinguishes the strikes specified in the time series by the performance data Q into strong beats and weak beats, and generates rhythmic data R by identifying the periodic pattern composed of strong beats and weak beats as the rhythmic structure. Note that known techniques may be arbitrarily employed to generate rhythmic data R using the performance data Q (i.e., to analyze the rhythmic structure). For example, techniques such as those described in Hamanaka et al., "Implementation of Musical Structure Analysis Based on GTTM: Acquisition of Grouping Structure and Rhythmic Structure," IPSJ Research Report MUS, [Music Information Science] 56, 1-8, 2004-08-02, or Goto et al., "Real-time Beat Tracking System for Acoustic Signals - Support for Music Without Percussion Sounds by Chord Change Detection -," IEICE Transactions on Electronics, Information and Communication Engineers D-2, Information & Systems 2-Information Processing 00081(00002), 227-237, 1998-02-25, are used for the analysis of rhythmic structure.

[0057] As described above, the synchronization control unit 55 of the first embodiment synchronizes the video data X and the audio data Y using the performance data Q. The synchronization control unit 55 of the second embodiment synchronizes the video data X and the audio data Y using the beat data R. Figure 10 is a flowchart illustrating the detailed procedure of the synchronization control process performed by the synchronization control unit 55 of the third embodiment. That is, the synchronization control process in Figure 6 in the first embodiment is replaced by the synchronization control process in Figure 10 in the third embodiment.

[0058] When the synchronization control process is started, the synchronization control unit 55 identifies the sound-producing point and sound-producing intensity of the bass drum 11 by analyzing the acoustic data Y (Sb71). The sound-producing intensity is the intensity of the sound (e.g., volume) identified from the acoustic data Y. For example, the synchronization control unit 55 sequentially identifies the point in time in the acoustic data Y where the increase in volume exceeds a predetermined value as the sound-producing point, and identifies the volume at that sound-producing point as the sound-producing intensity.

[0059] The synchronization control unit 55 synchronizes the video data X and the audio data Y using the rhythm data R (Sb72). For example, the synchronization control unit 55 identifies from the audio data Y a period during which the pattern of sound intensity at each sound point approximates the rhythmic structure specified by the rhythm data R. The synchronization control unit 55 then determines the position of the audio data Y on the time axis relative to the video data X such that the period identified from the audio data Y and the section of the video data X corresponding to that rhythmic structure coincide on the time axis. In other words, the synchronization of the video data X and the audio data Y is controlled by taking into account not only the simple time series of sound points, but also the rhythmic structure within the music.

[0060] The synchronization control unit 55 generates composite data Z, which includes mutually synchronized video data X and audio data Y, similar to the first embodiment (Sb73). The composite data Z is played back by the playback device 47. As described above, in the composite data Z, the video data X and audio data Y are mutually synchronized. Therefore, when the video of a specific part of the music in the video data X is displayed by the display device 45, the sound of that part of the audio data Y is played back by the sound output device 46.

[0061] Figure 11 is a flowchart illustrating the procedure for the performance analysis process in the third embodiment. When the performance analysis process is started, the control device 41 performs the following, similar to the first embodiment: acquisition of video data X (S1), acquisition of sound data Y (S2), performance detection processing (S3), and generation of performance data Q (S4). Once the performance data Q is generated, the control device 41 generates beat data R from the performance data Q (S5). In other words, the control device 41 functions as a beat data generation unit 56.

[0062] The control device 41 functions as a synchronization control unit 55 to execute the synchronization control process shown in Figure 10 (S7). Specifically, the control device 41 generates synthesized data Z by synchronizing the video data X and the audio data Y using the beat data R. The playback of the synthesized data Z (S9) is the same as in the first embodiment.

[0063] According to the third embodiment, similar to the first embodiment, performance data Q, which serves as a temporal reference for the video data X, can be generated by analyzing the video data X. Furthermore, in the third embodiment, rhythmic data R is used to synchronize the video data X and the audio data Y. That is, synchronization between the video data X and the audio data Y is achieved by taking into account the rhythmic structure of the music. Therefore, compared to the first embodiment, in which performance data Q specifying the timing of the bass drum 11's sound is used to synchronize the video data X and the audio data Y, it is possible to synchronize the video data X and the audio data Y with high accuracy.

[0064] In the above description, an example was given in which the generation of rhythmic data R is added to the first embodiment, in which the vibration of the drum set 10 (head 112) is detected by analyzing video data X representing percussion instrument 1. In the second embodiment, in which the striking of the drum set 10 is detected by analyzing video data X representing beater 121, the generation of rhythmic data R is also added, similar to the example in the third embodiment.

[0065] D: Fourth Embodiment Figure 12 is a block diagram illustrating the functional configuration of the performance analysis system 40 in the fourth embodiment. The control device 41 of the fourth embodiment functions as an acoustic processing unit 57 in addition to the same elements as in the third embodiment (video data acquisition unit 51, acoustic data acquisition unit 52, analysis processing unit 53, performance data generation unit 54, synchronization control unit 55, beat data generation unit 56) by executing a program stored in the storage device 42.

[0066] The sound represented by the acoustic data Y includes not only the sound of the bass drum 11, which is the original target of sound recording (hereinafter referred to as the "target sound"), but also the sounds of other instruments played by instruments other than the bass drum 11 (hereinafter referred to as "non-target sounds"). Non-target sounds include, for example, the sounds of drums other than the bass drum 11 in the drum set 10, or the sounds of various instruments played in the vicinity of the drum set 10. The acoustic processing unit 57 generates acoustic data Ya by performing acoustic processing on the acoustic data Y.

[0067] Acoustic processing is a process that relatively emphasizes the target sound compared to non-target sounds. For example, the target sound, which is the sound of the bass drum 11, is in the lower frequency range compared to the non-target sound. Therefore, the acoustic processing unit 57 performs low-pass filtering on the acoustic data Y, with the cutoff frequency set to the maximum value of the bass drum 11's frequency range. Since non-target sounds exceeding the cutoff frequency are reduced or removed by acoustic processing, the target sound is emphasized or extracted in the acoustic data Ya after acoustic processing. In addition, sound source separation processing, which emphasizes the target sound relative to the non-target sound by utilizing the difference in the direction from which the target sound arrives to the sound pickup device 31 and the direction from which the non-target sound arrives, is also used as acoustic processing on the acoustic data Y.

[0068] Furthermore, the synchronization control unit 55 of the fourth embodiment synchronizes the video data X with the audio data Ya after audio processing. The synchronization control process in the fourth embodiment is the same as the synchronization control process in the third embodiment, except that the processing target is changed from audio data Y to audio data Ya. That is, the synchronization control unit 55 synchronizes the video data X and the audio data Ya using rhythm data R.

[0069] Figure 13 is a flowchart illustrating the procedure for the performance analysis process in the fourth embodiment. In the fourth embodiment, acoustic processing (S6) on the acoustic data Y is added to the performance analysis process of the third embodiment. Specifically, the control device 41 generates acoustic data Ya by performing acoustic processing on the acoustic data Y. That is, the control device 41 functions as an acoustic processing unit 57. The control device 41 generates synthesized data Z by applying rhythm data R to a synchronous control process (S7). Other operations in the performance analysis process are the same as in the third embodiment.

[0070] According to the fourth embodiment, the same effects as in the third embodiment are achieved. Furthermore, in the fourth embodiment, since the sound of the bass drum 11 (target sound) is emphasized in the sound data Y, it is possible to synchronize the video data X and the sound data Y with high precision compared to a form in which the sound represented by the sound data Y sufficiently includes non-target sounds.

[0071] In the above description, an example was given in which sound processing of the sound data Y is added to the first embodiment. However, in the second embodiment as well, sound processing of the sound data Y may be applied in the same manner. Also, in the above description, an example was given in which the generation of rhythm data R as exemplified in the third embodiment is included. However, the generation of rhythm data R may be omitted from the fourth embodiment onward. That is, the synchronization control unit 55 may synchronize the video data X and the sound data Y after sound processing using the performance data Q.

[0072] The sound processing exemplified above applies to both the first and second embodiments. Furthermore, although the above description exemplified a form including the generation of rhythm data R in the third embodiment, the generation of rhythm data R (S5) may be omitted in the fourth embodiment. That is, in the fourth embodiment, the synchronization control unit 55 may synchronize the video data X and the sound data Y using the performance data Q, similar to the first or second embodiment.

[0073] E: Fifth Embodiment Figure 14 is a block diagram illustrating the functional configuration of the performance analysis system 40 in the fifth embodiment. The control device 41 of the fifth embodiment functions as a synchronization adjustment unit 58 in addition to the same elements as in the fourth embodiment (video data acquisition unit 51, sound data acquisition unit 52, analysis processing unit 53, performance data generation unit 54, synchronization control unit 55, beat data generation unit 56, sound processing unit 57) by executing a program stored in the storage device 42.

[0074] The synchronization control unit 55 of the fifth embodiment synchronizes the video data X and the audio data Ya, similar to the fourth embodiment. However, it is conceivable that the temporal relationship between the video data X and the audio data Ya after processing by the synchronization control unit 55 (hereinafter referred to as the "synchronization relationship") may not conform to the user U's intentions, or that the video data X and the audio data Ya may not be accurately synchronized. The synchronization adjustment unit 58 in Figure 14 changes the position of one of the video data X and audio data Ya on the time axis relative to the other (i.e., the synchronization relationship) after the synchronization control processing.

[0075] Figure 15 is a flowchart illustrating the detailed procedure of the synchronization adjustment unit 58's process of adjusting the temporal relationship between video data X and audio data Ya (hereinafter referred to as the "synchronization adjustment process"). When the synchronization adjustment process is started, the synchronization adjustment unit 58 sets the adjustment value α (S81).

[0076] User U, while viewing the video and audio of the composite data Z played back by the playback device 47, operates the control device 44 to instruct the adjustment of the synchronization relationship between video data X and audio data Ya. Specifically, User U instructs the adjustment of the synchronization relationship so that the temporal relationship between video data X and audio data Ya in the composite data Z becomes a desired relationship. For example, if it is determined that the audio data Ya is lagging behind the video data X, User U instructs that the audio data Ya be moved forward (in the opposite direction of the time axis) relative to the video data X by a predetermined amount. On the other hand, if it is determined that the audio data Ya is leading the video data X, User U instructs that the audio data Ya be moved backward (in the direction of the time axis) relative to the video data X by a predetermined amount. The synchronization adjustment unit 58 sets the adjustment value α according to the instruction from User U. For example, if it is instructed to move the audio data Ya forward relative to the video data X, the synchronization adjustment unit 58 sets the adjustment value α to a negative number according to the instruction from User U. Furthermore, if it is instructed to move the audio data Ya backward relative to the video data X, the synchronization adjustment unit 58 sets the adjustment value α to a positive number corresponding to the instruction from the user U.

[0077] The synchronization control unit 55 adjusts the position of the video data X and the audio data Ya on the time axis relative to one of them (i.e., the synchronization relationship) according to the adjustment value α (S82). Specifically, if the adjustment value α is negative, the synchronization control unit 55 moves the audio data Ya forward relative to the video data X by an amount of movement corresponding to the absolute value of the adjustment value α. If the adjustment value α is positive, the synchronization control unit 55 moves the audio data Ya backward relative to the video data X by an amount of movement corresponding to the absolute value of the adjustment value α. The synchronization control unit 55 generates composite data Z including the video data X and audio data Y whose synchronization relationship has been adjusted (S83).

[0078] Figure 16 is a flowchart illustrating the procedure for the performance analysis process in the fifth embodiment. In the fifth embodiment, the synchronization adjustment process illustrated in Figure 15 is added to the performance analysis process of the fourth embodiment. That is, the control device 41 functions as a synchronization adjustment unit 58 to adjust the synchronization relationship between the video data X and the audio data Ya according to the adjustment value α (S8). Other operations in the performance analysis process are the same as in the fourth embodiment. The synthesized data Z generated by the synchronization adjustment process is played back by the playback device 47 (S9).

[0079] According to the fifth embodiment, the same effects as the fourth embodiment are achieved. Furthermore, in the fifth embodiment, the position of one of the video data X and the audio data Ya on the time axis relative to the other can be adjusted after the synchronization control process. Moreover, in the fifth embodiment, since the adjustment value α is set according to the instruction from the user U, the position of one of the video data X and the audio data Ya relative to the other can be adjusted according to the user U's intention.

[0080] The synchronization relationship adjustments exemplified above are applicable to both the first and second embodiments. In addition, in the fifth embodiment, the generation of beat data R (S5) may be omitted. That is, in the fifth embodiment, the synchronization control unit 55 may synchronize the video data X and the audio data Y using the performance data Q, similar to the first or second embodiment. In addition, in the fifth embodiment, the audio processing (S6) for the audio data Y may also be omitted. That is, in the fifth embodiment, the synchronization control unit 55 may synchronize the video data X and the audio data Y, similar to the first or second embodiment.

[0081] F: Sixth Embodiment As described above, the synchronization adjustment unit 58 of the fifth embodiment sets the adjustment value α in response to instructions from the user U. The synchronization adjustment unit 58 of the sixth embodiment sets the adjustment value α using a trained model M. The configuration and operation other than the setting of the adjustment value α are the same as in the fifth embodiment.

[0082] Figure 17 is an explanatory diagram regarding the setting of the adjustment value α in the sixth embodiment. The synchronization adjustment unit 58 generates the adjustment value α by processing the input data C using the trained model M. In the sixth embodiment, video data X is supplied to the trained model M as input data C.

[0083] The temporal relationship (synchronization relationship) between the video data X and the audio data Ya synchronized by the synchronization control unit 55 tends to depend on conditions related to the bass drum 11. Conditions related to the bass drum 11 include, for example, the type (product model) or size of the bass drum 11. For example, it is assumed that electronic drums tend to have a greater delay in synchronized audio data Ya relative to the video data X than acoustic drums. Therefore, the adjustment value α for appropriately adjusting the synchronization relationship changes depending on the conditions of the bass drum 11 represented by the video data X. Taking the above correlation into consideration, the trained model M of the sixth embodiment is a statistical estimation model that has learned the relationship between the input data C (video data X) and the adjustment value α by machine learning. That is, the trained model M outputs a statistically valid adjustment value α for the input data C. Video data X is used as the input data C that indicates the conditions of the bass drum 11. Since the video data X reflects external conditions such as the type or model of the bass drum 11, the trained model M can generate a statistically valid adjustment value α for those conditions. Furthermore, information such as the type (model) or size of the bass drum 11 represented by the video data X may be supplied to the trained model M as input data C. Also, feature quantities F calculated from the video data X may be supplied to the trained model M as input data C.

[0084] Specifically, the trained model M is implemented by a combination of a program that causes the control device 41 to perform an operation to generate an adjustment value α from the input data C, and a number of variables (weights and biases) applied to the operation. The trained model M is composed of, for example, a deep neural network. For example, any form of deep neural network such as a recurrent neural network (RNN) or a convolutional neural network can be used as the trained model M. The trained model M may also be composed of a combination of multiple types of deep neural networks. In addition, additional elements such as long short-term memory (LSTM) or attention may be incorporated into the trained model M.

[0085] The trained model M described above is established through machine learning using multiple training datasets. Each of the training datasets includes training input data C (video data X) representing a bass drum 11 and an appropriate training adjustment value α (ground truth value) for that bass drum 11. In machine learning, multiple variables of the trained model M are iteratively updated so that the error between the adjustment value α generated by the provisional trained model M from the input data C of each training dataset and the adjustment value α of that training dataset is reduced. In other words, the trained model M learns the relationship between the training input data C and the training adjustment value α corresponding to the video of the percussion instrument.

[0086] In the synchronization adjustment process, the synchronization adjustment unit 58 obtains an adjustment value α by inputting the video data X as input data C to the trained model M (S81). The process of adjusting the synchronization relationship according to the adjustment value α (S82), and the process of generating composite data Z from the adjusted video data X and sound data Ya (S83) are the same as in the fifth embodiment.

[0087] According to the sixth embodiment, the same effects as in the fifth embodiment are achieved. Furthermore, in the sixth embodiment, since the adjustment value α is set using the trained model M, a statistically valid adjustment value α can be set for the input data C.

[0088] The synchronization adjustments exemplified above apply to both the first and second embodiments. In addition, in the sixth embodiment, the generation of beat data R (S5) may be omitted. That is, in the sixth embodiment, the synchronization control unit 55 may synchronize the video data X and the audio data Y using the performance data Q, similar to the first or second embodiment. In addition, in the sixth embodiment, the audio processing (S6) for the audio data Y may also be omitted. That is, in the sixth embodiment, the synchronization control unit 55 may synchronize the video data X and the audio data Y, similar to the first or second embodiment.

[0089] G: Variant The following are examples of specific modifications that may be added to each of the embodiments exemplified above. Multiple embodiments selected from the following examples may be merged as appropriate, provided they do not contradict each other.

[0090] (1) In each of the above-described embodiments, composite data Z was generated from one video data X and one audio data Y. However, multiple video data X generated by different recording devices 20 may be used to generate composite data Z. For each of the multiple video data X, the performance data generation unit 54 generates performance data Q and the rhythm data generation unit 56 generates rhythm data R. The synchronization control unit 55 generates composite data Z by synchronizing the multiple video data X and audio data Y. According to the above embodiments, a multi-angle video can be generated in which multiple videos taken at different locations and angles are arranged in parallel. Alternatively, the synchronization control unit 55 may generate composite data Z in which multiple video data X are switched sequentially in a time-division manner. For example, the synchronization control unit 55 generates composite data Z in which the video is switched at intervals corresponding to the rhythm structure represented by the rhythm data R. The interval corresponding to the rhythm structure is, for example, a period corresponding to n rhythm structures (where n is a natural number of 1 or more).

[0091] (2) In each of the above-described embodiments, composite data Z was generated from one video data X and one audio data Y. However, multiple audio data Y generated by different recording devices 30 may be used to generate composite data Z. The synchronization control unit 55 mixes the multiple audio data Y at a predetermined ratio and synchronizes the mixed audio data Y with the video data X. Alternatively, the synchronization control unit 55 may synchronize each of the multiple audio data Y with the video data X to generate composite data Z in which the multiple audio data Y are switched sequentially in a time-division manner.

[0092] (3) In the above-described embodiments, the recording device 20 generates video data X and the recording device 30 generates audio data Y. However, either or both of the recording devices 20 and 30 may generate both video data X and audio data Y. In addition, video data X or audio data Y may be transmitted to the performance analysis system 40 from each of the multiple recording devices. As described above, the number of recording devices is arbitrary, and the type of data transmitted by each recording device (either or both of video data X and audio data Y) is also arbitrary. Therefore, as illustrated in the above-described examples of modified form (1) or modified form (2), the total number of video data X or audio data Y acquired by the performance analysis system 40 is also arbitrary.

[0093] (4) In each of the above-described embodiments, the video data acquisition unit 51 acquires video data X from the recording device 20, but the video data X may also be data stored in the storage device 42. The video data acquisition unit 51 acquires video data X from the storage device 42. As can be understood from the above explanation, the video data acquisition unit 51 is any means of acquiring video data X and includes both the element of receiving video data X from an external device such as the recording device 20 and the element of acquiring video data X from the storage device 42.

[0094] (5) In each of the above-described embodiments, the sound data acquisition unit 52 acquires sound data Y from the recording device 30, but the sound data Y may also be data stored in the storage device 42. The sound data acquisition unit 52 acquires sound data Y from the storage device 42. As can be understood from the above explanation, the sound data acquisition unit 52 is an arbitrary means of acquiring sound data Y and includes both an element of receiving sound data Y from an external device such as the recording device 30 and an element of acquiring sound data Y from the storage device 42.

[0095] (6) In the above-described forms, examples were given in which video data X and audio data Y are recorded in parallel with each other, but it is not necessarily required that video data X and audio data Y be recorded in parallel. Even if video data X and audio data Y are recorded at different times or places, it is possible to synchronize them by using performance data Q or rhythm data R. In addition, there may be a difference in tempo between the performance represented by video data X and the performance represented by audio data Y. If there is a difference in tempo between video data X and audio data Y, the synchronization control unit 55 matches the tempo of audio data Y to the tempo of video data X by known time stretching, and then synchronizes video data X and audio data Y. The synchronization control unit 55 identifies the tempo of video data X from performance data Q or rhythm data R and performs time stretching on audio data Y to match that tempo. In other words, the performance data Q or rhythm data R used for synchronizing video data X and audio data Y is also used for time stretching of audio data Y.

[0096] (7) In the first embodiment, the vibration of the head 112 of the bass drum 11 was detected, but the object of detection using the video data X is not limited to the bass drum 11. For example, the vibration of other drums that make up the drum set 10 (e.g., tom-toms, floor toms, or snare drums) may be detected by analysis of the video data X. That is, the video represented by the video data X may include drums other than the bass drum 11 in the drum set 10.

[0097] Furthermore, while the aforementioned configurations focused on the bass drum 11 as an acoustic drum, a configuration in which the video data X represents the video of an electronic drum is also conceivable. The electronic drum is equipped with a pad (for example, a rubber pad) instead of the head 112 mentioned above. The analysis processing unit 53 detects the vibration of the pad in the electronic drum by analyzing the video data X. In addition, idiophones such as cymbals, or keyboard percussion instruments such as xylophones, may be included in the video data X. The analysis processing unit 53 detects the vibrations generated in the idiophones by analyzing the video data X. As can be understood from the above examples, the analysis processing unit 53 is comprehensively represented as an element for detecting vibrations generated in percussion instruments by performance, and the type of percussion instrument is arbitrary. It should be noted that idiophones such as cymbals tend to have a larger amplitude of vibration and a longer duration of vibration compared to the head 112 of membranophones such as the bass drum 11. Therefore, the processing load for the analysis processing unit 53 to detect the vibration of an idiophone exceeds the processing load for detecting the vibration of a membranophone. Considering the above trends, a method for detecting the vibrations of membranophones is preferable from the viewpoint of reducing the processing load required to detect the vibrations of percussion instruments. In percussion instruments, the elements that generate vibrations are comprehensively represented as vibrating bodies.

[0098] Furthermore, the support structures that hold up the bodies of various musical instruments, such as idiophones or membranophones, are also included in the concept of "percussion instruments." For example, a cymbal stand that supports a cymbal, or a hi-hat stand that supports a hi-hat, are vibrating bodies that vibrate when played, and are conceived as elements that constitute part of a percussion instrument. Also, the back head or body 111 that vibrates coupled with the impact of the head 112 is also included in the concept of vibrating bodies. As can be understood from the above examples, the vibrating bodies that the analysis processing unit 53 detects vibrations from include not only the elements that the user U directly strikes, but also other elements that vibrate in conjunction with those elements. In other words, vibrating bodies are comprehensively represented as elements that vibrate when played.

[0099] (8) In the second embodiment, the example given was that the video data X represents the video of the foot pedal 12, but the beater 121 included in the video data X is not limited to the foot pedal 12. For example, a stick used to play various percussion instruments such as a tom-tom, floor tom, or snare drum may be included in the video data X. The analysis processing unit 53 detects vibrations generated in the stick by analyzing the video data X. Also, a mallet used to play a keyboard percussion instrument such as a xylophone may be included in the video data X. The analysis processing unit 53 detects vibrations generated in the mallet by analyzing the video data X. As can be understood from the above examples, the analysis processing unit 53 is comprehensively represented as an element for detecting strikes by a striking body. The beater 121, stick, and mallet are examples of striking bodies. That is, a striking body is comprehensively represented as an element used for striking for performance.

[0100] As can be understood from the above-described modifications (7) and (8), the analysis processing unit 53 is comprehensively represented as an element for detecting changes in the percussion instrument due to performance. Changes in the percussion instrument due to performance include vibration of the vibrating body or strikes by the striking body. The striking body may be interpreted as the vibrating body of the percussion instrument.

[0101] (9) In each of the above-described forms, the bass drum 11 or foot pedal 12 may not be included in a portion of the video represented by the video data X. However, from the viewpoint of accurately synchronizing the video data X and the sound data Y at the start of the song, it is desirable that the video of the video data X includes the bass drum 11 or foot pedal 12 at that start point. However, the synchronization control unit 55 can also estimate the start point of the song by analyzing the performance data Q and the beat data R.

[0102] (10) In the first embodiment, an example was given in which the analysis processing unit 53 identifies a target region from the video data X, but the identification of the target region (Sa31) may be omitted. For example, if the video data X represents only the head 112 of the bass drum 11, the vibration of the head 112 can be detected by analyzing the video data X without identifying the target region. Therefore, the identification of the target region is omitted. In any form in which the analysis processing unit 53 detects vibration, the identification of the target region by the analysis processing unit 53 may be omitted.

[0103] (11) The trained model M in the sixth embodiment is not limited to a deep neural network. For example, a statistical estimation model such as an HMM (Hidden Markov Model) or an SVM (Support Vector Machine) may be used as the trained model M.

[0104] (12) In the sixth embodiment, video data X was used as input data C, but input data C is not limited to the above examples. As mentioned above, the synchronization relationship tends to depend on conditions relating to the bass drum 11. Considering this tendency, the synchronization control unit 55 may identify conditions relating to the bass drum 11 by analyzing the video data X and supply input data C representing these conditions to the trained model M. Conditions relating to the bass drum 11 include, for example, the size or type of the bass drum 11. The synchronization control unit 55 identifies conditions relating to the bass drum 11 by object detection processing on the video data X. As can be understood from the above description, input data C is comprehensively represented as data corresponding to the video data X, and includes not only the video data X itself but also data generated from the video data X.

[0105] (13) The order of each process in the performance analysis process may be changed as appropriate from the order exemplified in each of the above forms. For example, the order of acquiring video data X (S1) and acquiring sound data Y (S2) may be reversed. Also, the order of acquiring sound data Y (S2) and the performance detection process by the analysis processing unit 53 (S3) may be reversed.

[0106] (14) In the first and second embodiments, changes in the percussion instrument were detected by analyzing the video data X, and the results of the detection were used to generate performance data Q. However, as illustrated in Figure 18, a trained model (hereinafter referred to as the "first trained model") M1 may be used to generate the performance data Q. The first trained model M1 is a statistical estimation model that has learned the relationship between input data D and performance data Q by machine learning. The input data D supplied to the first trained model M1 is data corresponding to the video data X. Specifically, for example, the video data X itself, or the aforementioned feature quantity F calculated from the video data X, is used as input data D. The control device 41 (performance data generation unit 54) generates performance data Q by processing the input data D using the first trained model M1. Note that in the configuration of Figure 18, the analysis processing unit 53 exemplified in each of the above embodiments is omitted. Also, the video represented by the video data X includes at least one of the vibrating body and the striking body of the percussion instrument.

[0107] The first trained model M1 is implemented by a combination of a program that causes the control device 41 to perform an operation to generate performance data Q from input data D, and a number of variables (weights and biases) applied to the operation. The first trained model M1 is composed of a deep neural network, such as a convolutional neural network or a recurrent neural network.

[0108] The first trained model M1 is established through machine learning using multiple training datasets. Each of the training datasets includes training input data D and appropriate training performance data Q (ground truth value) for that input data D. In machine learning, multiple variables defining the first trained model M1 are iteratively updated so that the error between the performance data Q generated by the provisional first trained model M1 from the input data D of each training dataset and the performance data Q of that training dataset is reduced. In other words, the first trained model M1 learns the relationship between the training input data D and the training performance data Q corresponding to the video of the percussion instrument. The generation of rhythm data R using the performance data Q and the synchronization control processing using the rhythm data R are the same as in the forms described above.

[0109] In the configuration shown in Figure 18, performance data Q is generated by processing input data D corresponding to the video data X of percussion instrument 1 with the first trained model M1. That is, similar to the first or second embodiment, performance data Q, which serves as a temporal reference for the performance of percussion instrument 1, can be generated from the video data X.

[0110] Furthermore, the configuration of the fourth embodiment in which the acoustic processing unit 57 processes acoustic data Y, and the configuration of the fifth or sixth embodiment in which the synchronization adjustment unit 58 performs synchronization adjustment processing, are similarly applicable to the configuration in Figure 18.

[0111] (15) In Figure 18, performance data Q is generated by processing input data D with the first trained model M1, but as illustrated in Figure 19, rhythm data R may be generated by processing input data D with the second trained model M2. The second trained model M2 is a statistical estimation model that has learned the relationship between input data D and rhythm data R by machine learning. The input data D supplied to the second trained model is data corresponding to video data X. Specifically, for example, the video data X itself, or the aforementioned feature quantity F calculated from the video data X, is used as input data D. The control device 41 (rhythm data generation unit 56) generates rhythm data R by processing input data D using the second trained model M2. Note that in the configuration of Figure 19, the analysis processing unit 53 and performance data generation unit 54 exemplified in each of the above embodiments are omitted. Also, the video represented by video data X includes at least one of the vibrating body and striking body of a percussion instrument.

[0112] The second trained model M2 is implemented by a combination of a program that causes the control device 41 to perform an operation to generate rhythm data R from input data D, and multiple variables (weights and biases) applied to the operation. The second trained model M2 is composed of a deep neural network, such as a convolutional neural network or a recurrent neural network.

[0113] The second trained model M2 is established through machine learning using multiple training datasets. Each of the training datasets includes training input data D and appropriate training rhythm data R (ground truth value) for that input data D. In machine learning, several variables defining the second trained model are iteratively updated so that the error between the provisional rhythm data R generated by the second trained model M2 from the input data D of each training dataset and the rhythm data R of that training dataset is reduced. In other words, the second trained model M2 learns the relationship between the training input data D and the training rhythm data R corresponding to the video of the percussion instrument. The synchronization control processing using the rhythm data R is the same as in each of the forms described above.

[0114] In the configuration shown in Figure 19, rhythmic data R is generated by processing input data D corresponding to the video data X of percussion instrument 1 with the second trained model M2. That is, similar to the third embodiment, rhythmic data R, which serves as a temporal reference for the performance of percussion instrument 1, can be generated from the video data X. The configuration of the fourth embodiment, in which the sound processing unit 57 processes sound data Y, and the configuration of the fifth or sixth embodiment, in which the synchronization adjustment unit 58 performs synchronization adjustment processing, are also applied to the configuration shown in Figure 19.

[0115] (16) In each of the embodiments described above, an example was given in which the percussion instrument 1 includes a vibrating body (head 112) and a striking body (beater 121). In the embodiment that generates performance data Q or rhythm data R from video data X representing the striking body, performance data Q or rhythm data R can be generated even if the percussion instrument 1 does not include a vibrating body. Therefore, each of the embodiments described above also applies to an air drum in which, for example, a sound is reproduced by the user U shaking the striking body. As can be understood from the above explanation, the term "percussion instrument" in this disclosure also includes air drums. That is, for an embodiment that generates performance data Q or rhythm data R from video data X representing the image of the striking body, detection of the image and vibration of the percussion instrument is not essential.

[0116] (17) As described above, the functions of the performance analysis system 40 are realized through the cooperation of one or more processors constituting the control device 41 and the program stored in the storage device 42. The above program can be provided in a form stored on a computer-readable recording medium and installed on the computer. The recording medium is, for example, a non-transitory recording medium, such as an optical recording medium (optical disc) like a CD-ROM, but also includes any known form of recording medium such as a semiconductor recording medium or a magnetic recording medium. Note that a non-transitory recording medium includes any recording medium except for transient propagation signals (transitory, propagating signals), and volatile recording media are not excluded. Furthermore, in a configuration in which a distribution device distributes a program via a communication network, the recording medium that stores the program in the distribution device corresponds to the non-transitory recording medium described above.

[0117] H: Note From the forms exemplified above, the following configuration can be understood, for example.

[0118] A performance analysis method according to one embodiment (embodiment 1) includes acquiring video data generated by imaging a percussion instrument, detecting changes in the percussion instrument due to performance by analyzing the video data, generating performance data representing the performance according to the results of the detection, and generating rhythmic data representing the rhythmic structure from the performance data.

[0119] According to the above embodiment, changes in a percussion instrument are detected by analyzing the video data generated from imaging the instrument, and performance data Q representing the performance of the percussion instrument is generated according to the results of the detection. In other words, performance data that serves as a temporal reference for video data X can be generated from the video data. In addition, rhythmic data representing the rhythmic structure is generated from the performance data. Therefore, various processes utilizing the rhythmic structure can be realized.

[0120] "Changes in percussion instruments" refer to vibrations generated in the vibrating body of a percussion instrument, or strikes made by the striking body of a percussion instrument. The vibrating body is the part of a percussion instrument that vibrates when played. For example, in membranophones such as drums, the vibrating body includes not only the head (striking surface) that is struck during playing, but also the back head that vibrates coupled with the striking. In idiophones such as cymbals, the instrument body that is struck during playing is included in the vibrating body. It should be noted that "vibrations of percussion instruments" are not limited to vibrations of the vibrating body that the user directly strikes. For example, vibrations of the components that support the vibrating body of a percussion instrument are also included in "vibrations of percussion instruments."

[0121] Furthermore, a striking object is an element used for striking in the performance of a percussion instrument. For example, drumsticks or beaters used to strike drums, or mallets used to strike keyboard percussion instruments such as xylophones, are examples of striking objects. Also, if we consider percussion instruments that are struck by the performer's body (for example, hands), the performer's body can also be included in the concept of a "striking object."

[0122] "Performance data" refers to data in any format that represents the performance of a percussion instrument. For example, a time-series data set consisting of sound-producing data representing the strike of a percussion instrument and time-series data specifying the position of the strike on the time axis is an example of performance data. The sound-producing data may not only represent the occurrence of the strike but also specify the intensity of the strike.

[0123] "Meterial structure" refers to the structure (rhythm) of the beats in a musical piece. Specifically, a typical example of "meterial structure" is the structure of a rhythmic pattern (time signature) defined by a combination of multiple beats, such as strong or weak beats, and the timing of each beat.

[0124] In a specific example of Embodiment 1 (Embodiment 2), the percussion instrument includes a vibrating body that vibrates during performance, and detecting changes in the percussion instrument includes identifying a target region in the percussion instrument where the vibrating body exists from the image represented by the video data, and detecting the vibration of the vibrating body in accordance with changes in the image in the target region. According to the above embodiment, the target region of the vibrating body in the percussion instrument is identified from the image represented by the video data. Therefore, the vibration of the vibrating body can be detected with high accuracy by analyzing the video data.

[0125] In a specific example of Embodiment 1 or Embodiment 2 (Embodiment 3), the percussion instrument includes a striking body used for striking for the performance, and detecting changes in the percussion instrument includes identifying the striking body from the image represented by the video data and detecting strikes by the striking body in accordance with changes in the image of the striking body. In the above embodiments, strikes by the striking body are detected by analyzing the video data generated by imaging the striking body, and performance data representing the performance of the percussion instrument is generated according to the result of the detection. That is, performance data that serves as a temporal reference for the video data can be generated from the video data. In addition, rhythmic data representing the rhythmic structure is generated from the performance data. Therefore, various processes utilizing the rhythmic structure can be realized.

[0126] A performance analysis method according to any specific example (Aspect 4) of aspects 1 to 3 further includes acquiring acoustic data representing the sound of the performance and synchronizing the video data and the acoustic data using the rhythmic data. According to the above aspects, rhythmic data is used to synchronize the video data and the acoustic data. That is, synchronization of the video data and the acoustic data is achieved by taking into account the rhythmic structure of the musical piece. Therefore, it is possible to synchronize the video data and the acoustic data with higher accuracy compared to a form in which performance data is used to synchronize the video data and the acoustic data.

[0127] "Audio data" is any data that represents the sound of a performance. For example, audio data can represent the sound of the same song that is being performed in the video data. However, the song being performed in the video data and the song whose sound the audio data represents do not necessarily have to be exactly the same. The order in which video data and audio data are acquired is arbitrary.

[0128] "Synchronization" between video and audio data refers to the process of adjusting the temporal correspondence between the video and audio data. A typical example of "synchronization" is adjusting the temporal position of one of the video and audio data relative to the other so that the sound represented by the audio data and the image represented by the video data at any given point in time within a piece of music correspond to each other on the temporal axis (for example, they coincide on the temporal axis). It should be noted that the video and audio data do not necessarily need to be perfectly synchronized over the entire duration. For example, if the video and audio data correspond to each other at a specific point in time on the temporal axis, the relationship between the video and audio data can be interpreted as "synchronized" even if the temporal difference between the video and audio data expands over time from that point. Furthermore, "synchronization" is not limited to a temporally consistent relationship between video and audio data. That is, the process of adjusting the temporal correspondence between video and audio data so that the time difference between one of the video and audio data is a predetermined value is also included in the concept of "synchronization."

[0129] In a specific example of Embodiment 4 (Embodiment 5), the performance sound represented by the sound data includes the performance sound of the percussion instrument and the performance sound of instruments other than the percussion instrument, and further includes performing sound processing on the sound data to emphasize the performance sound of the percussion instrument compared to the performance sound of instruments other than the percussion instrument, and synchronizing the video data and the sound data includes synchronizing the video data and the sound data after sound processing. In the above embodiments, since the performance sound of the percussion instrument is emphasized in the sound data, it is possible to synchronize the video data and the sound data with higher precision compared to an embodiment in which the performance sound represented by the sound data sufficiently includes the performance sound of instruments other than the percussion instrument.

[0130] "Sound processing" refers to any processing that relatively emphasizes the sound of percussion instruments compared to the sounds of other instruments. For example, a low-pass filter with the cutoff frequency set to the maximum value of the percussion instrument's range is an example of "sound processing." Sound source separation processing that separates the sound of percussion instruments from the sound of other instruments is also an example of "sound processing." It should be noted that the sound of other instruments does not need to be completely eliminated. In other words, any processing that suppresses (ideally eliminates) the sound of other instruments relative to the sound of percussion instruments is included in "sound processing."

[0131] A performance analysis method according to an example of Embodiment 4 or Embodiment 5 (Embodiment 6) further includes setting an adjustment value and changing the position of the other on the time axis relative to one of the synchronized video data and the audio data according to the adjustment value. According to the above embodiments, the position of the other on the time axis relative to one of the video data and the audio data can be adjusted after synchronization using beat data.

[0132] In the specific example of Embodiment 6 (Embodiment 7), setting the adjustment value includes setting the adjustment value in response to instructions from the user. In the above embodiments, since the adjustment value is set in response to instructions from the user, the position of one of the video data and the other on the time axis can be adjusted according to the user's intentions.

[0133] In a specific example of Embodiment 6 (Embodiment 8), setting the adjustment value includes setting the adjustment value by processing the input data corresponding to the video data using a trained model that has learned the relationship between training input data corresponding to the video of the percussion instrument and the training adjustment value. In the above embodiment, since the adjustment value is generated using a trained model that has undergone machine learning, a statistically valid adjustment value can be generated for unknown input data based on the relationship between the input data and the adjustment value in multiple training data sets for machine learning. The input data includes, for example, the video data itself generated by imaging the percussion instrument, or features calculated from the video data. The features are video features that change in conjunction with the performance of the percussion instrument. The input data may also include conditions such as the type or size of the percussion instrument estimated from the video data.

[0134] A "trained model" is a model that has learned the relationship between input data and adjustment values ​​through machine learning. For example, various statistical estimation models such as deep neural networks (DNNs), hidden Markov models (HMMs), or support vector machines (SVMs) are used as "trained models."

[0135] "Input data" refers to any data corresponding to the video data. For example, the video data itself may be used as input data. Alternatively, features extracted from the video data may also be used as input data. For example, features such as the size or type of percussion instrument represented by the video data may be input to the trained model as input data. In addition, the distance between the imaging device and the percussion instrument (shooting distance) at the time of imaging may be input to the trained model as input data.

[0136] A performance analysis method according to another aspect of this disclosure (Aspect 9) includes acquiring video data generated by imaging a percussion instrument, processing the video data to generate performance data representing the rhythmic structure, and generating rhythmic data representing the rhythmic structure from the performance data. In the above aspect, performance data is generated by processing the video data, and rhythmic data is generated from the performance data. That is, performance data and rhythmic data that serve as a temporal reference for the performance of a percussion instrument can be generated from the video data.

[0137] In a specific example of Embodiment 9 (Embodiment 10), generating the performance data includes generating the performance data by processing the input data corresponding to the video data using a trained model that has learned the relationship between training input data corresponding to the video of the percussion instrument and training performance data. According to the above embodiment, statistically valid performance data can be generated for unknown input data, based on the relationship between input data and performance data in multiple training datasets for machine learning. The input data includes, for example, the video data itself generated by imaging the percussion instrument, or features calculated from the video data. The features are video features that change in conjunction with the performance of the percussion instrument.

[0138] Another aspect of the present disclosure (Aspect 11) of the performance analysis method includes acquiring video data generated by imaging a percussion instrument and generating rhythmic data representing the rhythmic structure by processing the video data. In this aspect, rhythmic data is generated by processing the video data. That is, rhythmic data that serves as a temporal reference for the performance of a percussion instrument can be generated from the video data.

[0139] In a specific example of Embodiment 11 (Embodiment 12), generating the rhythmic data includes generating the rhythmic data by processing the input data corresponding to the video data using a trained model that has learned the relationship between training input data corresponding to the video of a percussion instrument and training rhythmic data. According to the above embodiment, statistically valid rhythmic data can be generated for unknown input data based on the relationship between input data and rhythmic data in multiple training datasets for machine learning. The input data includes, for example, the video data itself generated by imaging a percussion instrument, or features calculated from the video data. The features are video features that change in conjunction with the performance of the percussion instrument.

[0140] In any specific example of embodiments 9 to 12 (embodiment 13), the input data processed by the trained model includes at least one of video data representing the video of the percussion instrument and feature quantities of the video calculated from the video data. In any specific example of embodiments 9 to 13 (embodiment 14), the feature quantities of the video are, for example, feature quantities relating to the movement of feature points of the percussion instrument.

[0141] In any specific example of embodiments 9 to 14 (embodiment 15), the percussion instrument includes a vibrating body that vibrates during performance and a striking body used for striking during performance, and the image represented by the video data includes the striking body. According to the above embodiments, performance data or rhythm data can be generated from the image of the striking body. Therefore, an image of the vibrating body in the percussion instrument is not necessary. Furthermore, performance data or rhythm data can be generated even in situations where the percussion instrument does not include a vibrating body (e.g., an air drum).

[0142] The performance analysis methods described above can also be implemented as performance analysis systems. Furthermore, the performance analysis methods described above can also be implemented as programs for causing a computer system to execute these performance analysis methods. [Explanation of Symbols]

[0143] 100...Information processing system, 1...Percussion instrument, 10...Drum set, 11...Bass drum, 111...Body, 112...Head, 12...Foot pedal, 121...Beater, 122...Pedal, 20...Recording device, 21...Imaging device, 22...Communication device, 30...Recording device, 31...Sound acquisition device, 32...Communication device, 40...Performance analysis system, 41...Control device, 42...Storage device, 43...Communication device, 44...Operation device, 45...Display device, 46...Sound emission device, 47...Playback device, 51...Video data acquisition unit, 52...Acoustic data acquisition unit, 53...Analysis processing unit, 54...Performance data generation unit, 55...Synchronization control unit, 56...Rhythm data generation unit, 57...Acoustic processing unit, 58...Synchronization adjustment unit, M...Trained model.

Claims

1. To acquire video data generated by imaging a percussion instrument that includes a vibrating body that vibrates when played, The process involves identifying the target region in the percussion instrument where the vibrating element is located from the image represented by the aforementioned video data, To detect the vibration of the vibrating body in accordance with the change in the image in the target region, To generate performance data representing the aforementioned performance according to the results of the detection, To generate rhythmic data representing the rhythmic structure from the aforementioned performance data, To obtain sound data representing the sound of the performance, To synchronize the video data and the audio data using the aforementioned rhythm data. A performance analysis method implemented by a computer system, including [the following].

2. To acquire video data generated by imaging a percussion instrument that is played by striking it using a striking body, Identifying the striking object from the video data represented by the aforementioned video data, The system detects the impact caused by the impacting body in accordance with changes in the image of the impacting body, To generate performance data representing the aforementioned performance according to the results of the detection, To generate rhythmic data representing the rhythmic structure from the aforementioned performance data, To obtain sound data representing the sound of the performance, To synchronize the video data and the audio data using the aforementioned rhythm data. A performance analysis method implemented by a computer system, including [the following].

3. The sound represented by the aforementioned acoustic data includes the sound of the percussion instrument and the sound of instruments other than the percussion instrument. The further includes performing sound processing on the sound data to emphasize the sound of the percussion instrument compared to the sound of instruments other than the percussion instrument, Synchronizing the video data and the audio data includes synchronizing the video data and the audio data after sound processing. A method for analyzing a performance according to claim 1 or claim 2.

4. Setting adjustment values, The position of the synchronized video data and the audio data relative to the other on the time axis is changed according to the adjustment value. A method for analyzing a performance according to any one of claims 1 to 3, further comprising the above.

5. Setting the aforementioned adjustment value means This includes setting the adjustment value in response to instructions from the user. The performance analysis method according to claim 4.

6. Setting the aforementioned adjustment value means This includes setting the adjustment values ​​by processing the input data corresponding to the video data using a trained model that has learned the relationship between training input data corresponding to the video footage of a percussion instrument and training adjustment values ​​for changing the position of one of the video data and the other on the time axis. The performance analysis method according to claim 4.

7. A video data acquisition unit that acquires video data generated by imaging a percussion instrument including a vibrating body that vibrates when played, An analysis processing unit identifies a target region in the percussion instrument where the vibrating element is located from the image represented by the aforementioned video data, and detects the vibration of the vibrating element in accordance with changes in the image in the target region. A performance data generation unit that generates performance data representing the performance according to the results of the detection, A rhythm data generation unit generates rhythm data representing the rhythmic structure from the performance data, An audio data acquisition unit that acquires audio data representing the sound of the performance, A synchronization control unit that synchronizes the video data and the audio data using the beat data. A performance analysis system equipped with the following features.

8. A video data acquisition unit that acquires video data generated by imaging a percussion instrument played by striking it using a striking body, An analysis processing unit identifies the striking body from the video data and detects the impact caused by the striking body in accordance with changes in the video of the striking body. A performance data generation unit that generates performance data representing the performance according to the results of the detection, A rhythm data generation unit generates rhythm data representing the rhythmic structure from the performance data, An audio data acquisition unit that acquires audio data representing the sound of the performance, A synchronization control unit that synchronizes the video data and the audio data using the beat data. A performance analysis system equipped with the following features.

9. A video data acquisition unit that acquires video data generated by imaging a percussion instrument including a vibrating body that vibrates when played, An analysis processing unit identifies a target region in the percussion instrument where the vibrating element is located from the video data, and detects the vibration of the vibrating element in accordance with changes in the video in the target region. A performance data generation unit generates performance data representing the performance according to the results of the detection. A rhythm data generation unit generates rhythm data representing the rhythmic structure from the performance data. An acoustic data acquisition unit that acquires acoustic data representing the sound of the performance, and A synchronization control unit that synchronizes the video data and the audio data using the aforementioned rhythm data. A program that makes a computer system function.

10. A video data acquisition unit that acquires video data generated by imaging a percussion instrument that is played by striking it using a striking body. An analysis processing unit identifies the striking body from the video data and detects the impact caused by the striking body in accordance with changes in the video of the striking body. A performance data generation unit generates performance data representing the performance according to the results of the detection. A rhythm data generation unit generates rhythm data representing the rhythmic structure from the performance data. An acoustic data acquisition unit that acquires acoustic data representing the sound of the performance, and A synchronization control unit that synchronizes the video data and the audio data using the aforementioned rhythm data. A program that makes a computer system function.