Performance recording method, performance recording system and program
The described method automates the selection and extraction of performance data from multiple performers using image recognition and machine learning, addressing the inefficiency of pre-recording in existing systems and improving the creation of musical works.
Patent Information
- Application Number
- JP2021147641
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-10
- Publication Date
- 2025-11-12
- Estimated Expiration
- 2041-09-10
AI Technical Summary
Existing music creation systems require time-consuming pre-recording of performance content for each performer, which is inefficient for creating musical works with multiple performers.
A performance recording method using a camera to capture multiple performers' performances, determine area candidates, select a target area, and extract relevant portions from recorded images or sounds, employing image recognition and machine learning to streamline the process.
Reduces the effort required to create a musical work by automating the selection and extraction of performance data, enhancing efficiency and reducing the time needed for content generation.
Smart Images

Figure 0007767788000001 
Figure 0007767788000002 
Figure 0007767788000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a performance recording method, a performance recording system, and a program. [Background technology]
[0002] The music creation system described in Patent Document 1 creates a musical work by multiple performers by combining performance content data for each performer. The performance content data for each performer is generated by recording the performance of each performer in advance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-31885 Summary of the Invention [Problem to be solved by the invention]
[0004] In the music creation system described in Patent Document 1, in order to create a musical work for a group with multiple performers, such as a band, it was necessary to create performance content data for each performer in advance by recording their performance, which was time-consuming.
[0005] One aspect of the present disclosure aims to provide a technology that can reduce the effort required to create a musical work by a group of multiple performers. [Means for solving the problem]
[0006] A performance recording method according to one aspect of the present disclosure is a performance recording method implemented by a computer system, which uses a first captured image generated by a camera capturing a scene in which multiple performers perform a first performance of a piece of music, determines multiple area candidates in the camera's capture area, selects a target area from the multiple area candidates, and extracts a portion corresponding to the target area from a performance record obtained by capturing an image of a scene in which the multiple performers perform a second performance of the piece of music or by recording the sound of the second performance.
[0007] A performance recording system according to another aspect of the present disclosure includes a determination unit that determines a plurality of area candidates in the camera's imaging area using a first captured image generated by the camera capturing a scene in which a plurality of performers perform a first performance of a piece of music; a selection unit that selects a target area from the plurality of area candidates; and an extraction unit that extracts a portion corresponding to the target area from a performance record obtained by capturing an image of a scene in which the plurality of performers perform a second performance of the piece of music or by collecting the sound of the second performance.
[0008] A program according to yet another aspect of the present disclosure causes a computer system to function as a determination unit that determines multiple area candidates in the camera's imaging area using a first captured image generated by the camera capturing a scene in which multiple performers perform a first performance of a piece of music, a selection unit that selects a target area from the multiple area candidates, and an extraction unit that extracts a portion corresponding to the target area from a performance record obtained by capturing an image of a scene in which the multiple performers perform a second performance of the piece of music or by collecting the sound of the second performance. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram showing a performance recording system 1 according to a first embodiment. [Figure 2] FIG. 1 is a diagram showing three axes virtually set for a camera 2. [Figure 3] FIG. 2 is a diagram showing an imaging area 2a of a camera 2 deployed on a plane. [Figure 4]FIG. 1 is a diagram showing a captured image K1 developed on a plane. [Figure 5] FIG. 10 is a diagram showing a captured image K2 developed on a plane. [Figure 6] FIG. 2 is a diagram showing an example of an object M and a plurality of region candidates 2d. [Figure 7] An example in which the region candidate 2d1 is selected as the target region 2e is shown. [Figure 8] FIG. 2 is a diagram showing an example of an output image P. [Figure 9] FIG. 1 is a diagram illustrating an example of a performance recording system 1. [Figure 10] FIG. 2 is a diagram showing an example of a processing device 1f. [Figure 11] 1 is a diagram illustrating a determination unit 11A, which is an example of the determination unit 11. FIG. [Figure 12] FIG. 1 is a diagram illustrating machine learning. [Figure 13] FIG. 10 is a diagram showing an example of learning data V2. [Figure 14] FIG. 10 is a diagram illustrating a selection unit 12A. [Figure 15] FIG. 10 is a diagram showing an example of an operation for determining a plurality of region candidates 2d. [Figure 16] 10 is a diagram showing an example of an operation for generating performance data Q. FIG. [Figure 17] FIG. 10 is a diagram showing examples of region candidates 2d3 and 2d4. [Figure 18] FIG. 10 is a diagram showing an example of weighting factors W1 and W2 according to the genre of music C. DETAILED DESCRIPTION OF THE INVENTION
[0010] A: First embodiment A1: Performance Recording System 1 1 is a diagram showing a performance recording system 1 according to the first embodiment. The performance recording system 1 is a computer system that records a performance of a piece of music C by a performing group B at a performance venue A.
[0011] Performance venue A is the place where the performance takes place. Performance venue A may be, for example, a music studio, a performance hall, an outdoor stage, or a classroom.
[0012] Performance group B is a musical band including multiple performers D. The multiple performers D are composed of two people: vocalist D1 and instrumentalist D2. Vocalist D1 and instrumentalist D2 are each an example of performer D. Vocalist D1 sings song C. Instrumentalist D2 plays song C using instrument E. Instrument E is a guitar. Instrument E is not limited to a guitar and may be, for example, a bass, drums, electronic piano, or synthesizer. The multiple performers D may include multiple vocalists D1. The multiple performers D may also include multiple instrumentalists D2. The multiple instruments E used by the multiple instrumentalists D2 may be different types of instruments or the same type of instruments.
[0013] Performance group B performs a rehearsal performance F1 and a live performance F2 of piece C at performance venue A. Rehearsal performance F1 is an example of a first performance. Live performance F2 is an example of a second performance.
[0014] Performance recording system 1 is connected to camera 2 and microphone 3. Performance recording system 1 may include at least one of camera 2 and microphone 3. Camera 2 and microphone 3 are each placed in the center of performance venue A. The positions of camera 2 and microphone 3 do not have to be the center of performance venue A, as long as they are placed at performance venue A. Performance recording system 1 uses camera 2 and microphone 3 to record a performance of musical piece C by performing group B at performance venue A.
[0015] Camera 2 is a 360-degree camera, which is also called a spherical camera or an omnidirectional camera.
[0016] FIG. 2 is a diagram showing three axes that are virtually set with respect to the camera 2. The three axes are a roll axis G1, a pitch axis G2, and a yaw axis G3. The roll axis G1 is an axis that is parallel to the front-to-back direction of the camera 2. The pitch axis G2 is an axis that is parallel to the left-to-right direction of the camera 2. The yaw axis G3 is an axis that is parallel to the up-and-down direction of the camera 2. The roll axis G1, the pitch axis G2, and the yaw axis G3 are perpendicular to each other.
[0017] The imaging area 2a of the camera 2 encompasses the entire periphery of the camera 2. FIG. 3 is a diagram showing the imaging area 2a of the camera 2 unfolded on a plane. The horizontal direction H1 in the imaging area 2a unfolded on a plane indicates a rotation angle θy around the yaw axis G3 as the rotation axis. The rotation angle θy is an angle ranging from 0 degrees to 360 degrees. The vertical direction H2 in the imaging area 2a unfolded on a plane indicates a rotation angle θp around the pitch axis G2. The rotation angle θp is an angle ranging from -90 degrees to 90 degrees.
[0018] The position of any point 2b in the imaging area 2a is determined by the rotation angles θy and θp. The direction 2c from the camera 2 to point 2b is also determined by the rotation angles θy and θp. The direction 2c is also called the angle.
[0019] Camera 2 generates rehearsal video data by capturing footage of performing group B performing rehearsal performance F1 of musical piece C at performance venue A. The rehearsal video data is video data showing performing group B performing rehearsal performance F1 in video. Camera 2 generates a series of captured image data J1 as the rehearsal video data. Each captured image data J1 in the series of captured image data J1 represents a still image that constitutes one frame of the video shown by the rehearsal video data. Each captured image data J1 is image data showing performing group B performing rehearsal performance F1 as a still image. The still image represented by the captured image data J1 is referred to as a "captured image K1." When camera 2 generates captured image data J1, it means that camera 2 has generated captured image K1. Captured image K1 is an example of a first captured image.
[0020] Camera 2 generates live video data by capturing footage of performing group B performing live performance F2 of musical piece C at performance venue A. The live video data is video data that represents performing group B performing live performance F2 in video. Camera 2 generates a series of captured image data J2 as the live video data. Each captured image data J2 in the series of captured image data J2 represents a still image that constitutes one frame of the video represented by the live video data. Each captured image data J2 is image data that represents performing group B performing live performance F2 in a still image. The still image represented by the captured image data J2 is referred to as a "captured image K2." The generation of captured image data J2 by camera 2 means that camera 2 has generated captured image K2. Captured image K2 is an example of a second captured image and an example of a performance record.
[0021] 4 is a diagram showing a captured image K1, which is an omnidirectional image showing a scene of a rehearsal performance F1, laid out on a plane. An omnidirectional image is also called, for example, a spherical image, a spherical panoramic image, or a 360-degree image. The process of laying out the captured image K1 on a plane is executed by the performance recording system 1 or the camera 2. An example of the process executed by the camera 2 to lay out the captured image K1 on a plane will be described below. The camera 2 generates captured image data J1 that shows the captured image K1 laid out on a plane.
[0022] The captured image K1 represents a vocalist D1, an instrument player D2, and an instrument E. The position of an arbitrary point K1a in the captured image K1 is determined by the rotation angles θy and θp.
[0023] 5 is a diagram showing a captured image K2, which is an omnidirectional image showing a scene of a live performance F2, laid out on a plane. The process of laying out the captured image K2 on a plane is executed by the performance recording system 1 or the camera 2. An example in which the camera 2 executes the process of laying out the captured image K2 on a plane will be described below. The camera 2 generates captured image data J2 that shows the captured image K2 laid out on a plane.
[0024] Like the captured image K1, the captured image K2 shows the vocalist D1, the instrument player D2, and the instrument E. The position of an arbitrary point K2a in the captured image K2 is determined by the rotation angles θy and θp. The position (coordinates) of any point in each of the captured images K1 and K2 may be determined by the x- and y-coordinates of the captured image developed on a plane, instead of by the rotation angles θy and θp. The x- and y-coordinates of the captured image developed on a plane are, for example, coordinates represented by the x-coordinate on the x-axis parallel to the lateral direction (horizontal direction) of the captured image developed on the plane, and the y-coordinate on the y-axis parallel to the longitudinal direction (vertical direction) of the captured image developed on the plane.
[0025] When there is no need to distinguish between a captured image K1 representing a scene of a rehearsal performance F1 and a captured image K2 representing a scene of a live performance F2, each of the captured images K1 and K2 will be referred to as "captured image K." When there is no need to distinguish between captured image data J1 representing the captured image K1 and captured image data J2 representing the captured image K2, each of the captured image data J1 and J2 will be referred to as "captured image data J."
[0026] The microphone 3 shown in FIG. 1 is a microphone set having multiple microphones. Each of the multiple microphones has directionality. The microphone 3 may be a single microphone that does not have directionality. The sound collection range of the microphone 3 encompasses the entire surroundings of the microphone 3. Note that the sound collection range of the microphone 3 only needs to cover the imaging range of the camera 2, and does not necessarily have to encompass the entire surroundings of the microphone 3.
[0027] Microphone 3 picks up the sound of the performance by performing group B at performance venue A. For example, microphone 3 picks up the sound of rehearsal performance F1 that performing group B gives of song C at performance venue A. Microphone 3 also picks up the sound of actual performance F2 that performing group B gives of song C at performance venue A.
[0028] The microphone 3 generates performance sound data L. The performance sound data L is data representing performance sounds obtained by the microphone 3 picking up the sounds of the live performance F2. The performance sounds represented by the performance sound data L are another example of a performance record.
[0029] The performance recording system 1 is, for example, a smartphone. The performance recording system 1 is not limited to a smartphone, but may be, for example, a personal computer or a tablet. A smartphone and a tablet are each an example of a portable information device. A personal computer is an example of a portable or stationary information device. The performance recording system 1 may be configured as a single device, or may be configured as multiple devices that are separate from each other.
[0030] The performance recording system 1 acquires captured image data J1 representing a still image of one frame of a moving image represented by the rehearsal moving image data and captured image data J2 representing a still image of one frame of a moving image represented by the actual moving image data. The captured image data J1 is image data representing a captured image K1 representing a scene of a rehearsal performance F1. The captured image data J2 is image data representing a captured image K2 representing a scene of an actual performance F2.
[0031] The performance recording system 1 determines a plurality of area candidates 2d in the imaging area 2a using the captured image K1 indicated by the captured image data J1. For example, the performance recording system 1 determines a plurality of area candidates 2d in the imaging area 2a based on the object M in the captured image K1.
[0032] The objects M are, for example, at least a portion of the bodies of multiple performers D and musical instruments E. The at least a portion of the bodies of multiple performers D is, for example, the upper body of a performer D. The at least a portion of the bodies of multiple performers D is not limited to the upper body of a performer D, but may be, for example, the hand of a performer D, the face of a performer D, or the entire body of a performer D.
[0033] FIG. 6 is a diagram showing an example of an object M in a captured image K1 representing a scene of a rehearsal performance F1, and a plurality of area candidates 2d in an image capture area 2a.
[0034] The object M includes a detection object M1 and a detection object M2. The detection object M1 is the upper body of a vocalist D1. The detection object M2 is composed of the whole body of a musical instrument player D2 and a musical instrument E.
[0035] The performance recording system 1 identifies an image area K11 and an image area K12 in a captured image K1 representing a scene of a rehearsal performance F1. The image area K11 and the image area K12 are examples of multiple image areas. The image area K11 is an area representing a detection object M1 in the scene of the rehearsal performance F1. The image area K12 is an area representing a detection object M2 in the scene of the rehearsal performance F1.
[0036] The performance recording system 1 automatically identifies the image area K11 and the image area K12 using, for example, image recognition technology.
[0037] Image region K11 and image region K12 are each rectangular. Image region K11 and image region K12 each have a common aspect ratio AP. Image region K11 and image region K12 may each have a different aspect ratio.
[0038] The performance recording system 1 generates image area data N1 indicating an image area K11. The image area K11 is an area representing the upper body of vocalist D1 in the rehearsal performance F1 scene. The image area data N1 includes position data N11 and size data N12. The position data N11 indicates the center position K11c of the image area K11 using rotation angles θy and θp. The center position K11c of the image area K11 is, for example, the position of the intersection of diagonal lines in the image area K11. The size data N12 indicates the size of the image area K11. The size data N12 indicates the ratio of the size of the image area K11 to the size of a reference rectangular area having an aspect ratio AP. The reference rectangular image is set in advance. The size data N12 is also referred to as zoom data.
[0039] The performance recording system 1 generates image area data N2 indicating the image area K12. The image area K12 is an area representing the entire body of the instrument player D2 and the instrument E in the scene of the rehearsal performance F1. The image area data N2 includes position data N21 and size data N22. The position data N21 indicates the center position K12c of the image area K12 by the rotation angles θy and θp. The center position K12c of the image area K12 is, for example, the position of the intersection of the diagonal lines in the image area K12. The size data N22 indicates the size of the image area K12. The size data N22 indicates the ratio of the size of the image area K12 to the size of a reference rectangular area. The size data N22 is also referred to as zoom data.
[0040] The multiple region candidates 2d in the imaging region 2a include region candidate 2d1 and region candidate 2d2. Region candidate 2d1 corresponds to image region K11, which represents the upper body of vocalist D1 in the scene of rehearsal performance F1. Region candidate 2d2 corresponds to image region K12, which represents the entire body of instrument player D2 and instrument E in the scene of rehearsal performance F1.
[0041] The performance recording system 1 determines area candidates 2d1 and 2d2 using image area data N1 indicating image area K11 and image area data N2 indicating image area K12. For example, the performance recording system 1 determines the range indicated by the image area data N1 in the imaging area 2a of the camera 2 as area candidate 2d1. The performance recording system 1 determines the range indicated by the image area data N2 in the imaging area 2a of the camera 2 as area candidate 2d2.
[0042] The performance recording system 1 selects a target region 2e from among the multiple region candidates 2d. Figure 7 shows an example in which the region candidate 2d1 is selected as the target region 2e from among the region candidates 2d1 and 2d2. The region candidate 2d2 may also be selected as the target region 2e.
[0043] The performance recording system 1 extracts an image of a target area 2e in the captured image K2 from the captured image K2 that represents the scene of the live performance F2 as an output image P. The image of the target area 2e in the captured image K2 refers to the image shown in the target area 2e in the captured image K2.
[0044] 8 is a diagram showing an example of the output image P. The output image P is an image extracted from the captured image K2 in a situation where the region candidate 2d1 is selected as the target region 2e.
[0045] A2: An example of performance recording system 1 9 is a diagram showing an example of a performance recording system 1. The performance recording system 1 includes an operation device 1a, a display device 1b, a speaker 1c, a communication device 1d, a storage device 1e, and a processing device 1f.
[0046] The operation device 1a is an input device that receives instructions from a user. The operation device 1a is, for example, a touch panel. The operation device 1a is not limited to a touch panel, and may be, for example, a controller operated by a user. The operation device 1a may be an input device (for example, a mouse or keyboard) that is connected to the performance recording system 1 by wire or wirelessly. The operation device 1a may be an external element of the performance recording system 1.
[0047] The display device 1b is a display panel. The display panel is, for example, a liquid crystal display panel or an organic EL (Electroluminescence) panel. The display device 1b may be a touch panel. A touch panel may be used as the operation device 1a and the display device 1b. The display device 1b may be connected to the performance recording system 1 by wire or wirelessly. The display device 1b may be an external element of the performance recording system 1. The display device 1b displays various images.
[0048] The speaker 1c is a speaker set having multiple speakers. The speaker 1c may be a single speaker. The speaker 1c may be connected to the performance recording system 1 by wire or wirelessly. The speaker 1c may be an external element of the performance recording system 1. The speaker 1c emits various sounds.
[0049] The communication device 1d communicates with an external device 5 via a communication network NW. For example, the communication device 1d transmits performance data Q representing a live performance F2 by a performing group B to the external device 5 via the communication network NW. The external device 5 is, for example, a distribution server or a terminal device. The distribution server is a server that distributes the performance data Q received from the performance recording system 1. The terminal device is, for example, a smartphone, tablet, or personal computer.
[0050] The storage device 1e is a computer-readable recording medium (e.g., a non-transitive recording medium readable by a computer). The storage device 1e includes one or more memories. The storage device 1e includes, for example, a non-volatile memory and a volatile memory. Examples of the non-volatile memory include a ROM (Read Only Memory), an EPROM (Erasable Programmable Read Only Memory), and an EEPROM (Electrically Erasable Programmable Read Only Memory). Examples of the volatile memory include a RAM (Random Access Memory).
[0051] The storage device 1e stores the program PG1 and various data. The program PG1 defines the operation of the performance recording system 1. The storage device 1e may store the program PG1 read from a storage device in a server that can communicate with the processing device 1f. In this case, the storage device in the server is another example of a computer-readable recording medium. The storage device 1e may be a portable recording medium that can be attached to and detached from the performance recording system 1. The storage device 1e may also be an external element of the performance recording system 1.
[0052] The processing device 1f includes one or more central processing units (CPUs). The one or more CPUs are an example of one or more processors. The processing device, processor, and CPU are each an example of a computer.
[0053] The processing unit 1f reads the program PG1 from the storage unit 1e and executes the program PG1.
[0054] A3: Processing equipment 1f 10 is a diagram showing an example of a processing device 1f. By executing a program PG1, the processing device 1f functions as a determination unit 11, a selection unit 12, an extraction unit 13, a generation unit 14, an output control unit 15, and a communication control unit 16. At least one of the determination unit 11, the selection unit 12, the extraction unit 13, the generation unit 14, the output control unit 15, and the communication control unit 16 may be configured by a circuit such as a DSP (Digital Signal Processor) or an ASIC (Application Specific Integrated Circuit).
[0055] The determination unit 11 uses the captured image K1 representing the scene of the rehearsal performance F1 to determine a plurality of area candidates 2d in the imaging area 2a of the camera 2. The determination unit 11 generates candidate data R indicating the plurality of area candidates 2d. The determination unit 11 stores the candidate data R in the storage device 1e.
[0056] The determination unit 11 may generate, as candidate data R, data indicating the plurality of area candidates 2d and the type of musical instrument represented by each image area in the captured image K1 (for example, each of image areas K11 and K12 in FIG. 6). The type of musical instrument may be input by the user to the operation device 1a, or may be identified by the determination unit 11 performing image recognition processing on the captured image K1. The determination unit 11 may generate, as candidate data R, data indicating the name of performing group B and an introduction to song C in addition to the plurality of area candidates 2d and the type of musical instrument. In this case, the candidate data R is easy to identify. Furthermore, the candidate data R can be used as both data indicating the name of performing group B and data indicating the introduction to song C. The determination unit 11 may determine a plurality of area candidates 2d for each of two or more different captured images K1. The two or more different captured images K1 are designated, for example, by a user. In this case, the determination unit 11 generates candidate data R for each of the two or more different captured images K1. For example, the determination unit 11 generates, as candidate data R, data indicating a plurality of area candidates 2d based on the captured image K1 and the elapsed rehearsal time from the start of the rehearsal performance F1 to the generation of the captured image K1 for each of the two or more different captured images K1.
[0057] The selection unit 12 selects a target area 2e from among a plurality of area candidate 2d. The selection unit 12 reads candidate data R from the storage device 1e. The selection unit 12 selects a target area 2e from among the plurality of area candidate 2d indicated by the candidate data R. For example, the selection unit 12 analyzes a captured image K2 representing a scene of a live performance F2 and selects the target area 2e. When multiple pieces of candidate data R exist, the selection unit 12 may switch from one piece of candidate data R to another piece of candidate data R during the actual performance F2. For example, the selection unit 12 first identifies, from among the multiple pieces of candidate data R, candidate data R indicating a rehearsal elapsed time that is shorter than the elapsed time of the actual performance F2 as provisional candidate data Ra. Next, the selection unit 12 identifies, from among the multiple pieces of candidate data Ra, the provisional candidate data Ra indicating the rehearsal elapsed time that is the smallest difference from the elapsed time of the actual performance F2 as target candidate data Rb. Note that, if no provisional candidate data Ra exists, the selection unit 12 identifies, from among the multiple pieces of candidate data R, the candidate data R indicating the rehearsal elapsed time that is the smallest difference from the elapsed time of the actual performance F2 as target candidate data Rb. Next, the selection unit 12 selects a target region 2e from among the multiple region candidates 2d indicated by the target candidate data Rb by analyzing the captured image K2 representing a scene at that elapsed time of the actual performance F2. The selection unit 12 may select the target area 2e in response to an instruction from the user. The selection unit 12 may also select the target area 2e randomly. When the selection unit 12 selects the target area 2e randomly, the captured image K2 (captured image data J2) representing the scene of the actual performance F2 can be eliminated in the process of selecting the target area 2e.
[0058] The extraction unit 13 extracts, as an output image P, an image of the target region 2e in the captured image K2, from the captured image K2 that represents the scene of the live performance F2.
[0059] For example, the extraction unit 13 extracts the output image P from the captured image K2 at a timing corresponding to the selection of the target region 2e. As an example, the extraction unit 13 extracts the output image P from the captured image K2 when the target region 2e is selected. In this case, the timing corresponding to the selection of the target region 2e is an example of the timing corresponding to the selection of the target region 2e. If the timing at which the captured image K2 is supplied to the extraction unit 13 is delayed from the timing at which the captured image K2 is supplied to the selection unit 12, the extraction unit 13 may extract the output image P from the captured image K2 at a timing when a certain time has elapsed since the target region 2e was selected. In this case, the timing at which a certain time has elapsed since the target region 2e was selected is an example of the timing corresponding to the selection of the target region 2e. Selection of the target region 2e is performed sequentially as the actual performance F2 progresses (time elapses). Therefore, the timing according to the selection of the target region 2e can be said to be timing according to the progress (time elapses) of the actual performance F2. If the target region 2e changes as the actual performance F2 progresses, the output image P extracted by the extraction unit 13 from the captured image K2 will change. Furthermore, if multiple candidate regions 2d change as the actual performance F2 progresses (time elapses), the output image P extracted by the extraction unit 13 from the captured image K2 may change. Therefore, the extraction unit 13 can extract a variety of output images P that change as the actual performance F2 progresses. The extraction unit 13 generates output image data T that represents the output image P.
[0060] The generation unit 14 receives output image data T and performance sound data L. The output image data T is image data showing an output image P4 extracted from a captured image K2 representing a scene of the live performance F2. The performance sound data L is sound data generated by a microphone 3 that collects the sound of the live performance F2. The generation unit 14 generates performance data Q including the output image data T and the performance sound data L. The performance data Q is data (video content) that represents the live performance F2 using images and sounds.
[0061] The output control unit 15 provides the display device 1b with the output image data T included in the performance data Q, thereby causing the display device 1b to display the output image P indicated by the output image data T. The output control unit 15 provides the speaker 1c with the performance sound data L included in the performance data Q, thereby causing the speaker 1c to emit the performance sound indicated by the performance sound data L.
[0062] The communication control unit 16 transmits the performance data Q from the communication device 1d to the external device 5 via the communication network NW.
[0063] A4: An example of the determination unit 11 11 is a diagram illustrating a determination unit 11A, which is an example of the determination unit 11. The determination unit 11A includes a detection unit 111, a candidate determination unit 112, an estimation model 41, and an estimation model 42. The estimation model 41 and the estimation model 42 may be external elements of the determination unit 11A.
[0064] The detection unit 111 detects, from a captured image K1 representing a scene of a rehearsal performance F1, objects M which are at least parts of the bodies of multiple performers D and musical instruments E. For example, the detection unit 111 detects, from the captured image K1, a detection object M1 which is the upper body of a vocalist D1, and a detection object M2 which is composed of the entire body of a musical instrument performer D2 and a musical instrument E (e.g., a guitar).
[0065] The detection unit 111 detects the detection object M1 using an estimation model 41. The estimation model 41 is a trained model that has learned the relationship between the captured image data J (captured image K) and the area representing the detection object M1 through machine learning. The estimation model 41 is configured by a deep neural network (DNN). The deep neural network is, for example, a convolutional neural network (CNN), a recurrent neural network (RNN), or a long short-term memory (LSTM). The estimation model 41 may include a combination of multiple types of deep neural networks.
[0066] The estimation model 41 has a plurality of coefficients U1. The plurality of coefficients U1 determine the behavior of the estimation model 41. The plurality of coefficients U1 have been adjusted by machine learning.
[0067] FIG. 12 is a diagram for explaining machine learning. The machine learning system 6 is a system separate from the performance recording system 1. The machine learning system 6 is, for example, a server system capable of communicating with the performance recording system 1 via a communication network NW. The machine learning system 6 generates an estimated model 41 from a tentative model 41a. The tentative model 41a is an estimated model (deep neural network) having multiple coefficients U1a. The machine learning system 6 generates multiple coefficients U1 and an estimated model 41 by updating the multiple coefficients U1a through machine learning. The multiple coefficients U1 are multiple coefficients U1a whose updates have been completed. The estimated model 41 is the tentative model 41a whose multiple coefficients U1a whose updates have been completed.
[0068] The machine learning system 6 updates a plurality of coefficients U1a using a plurality of pieces of training data V1. The plurality of pieces of training data V1 are different from one another. Each of the plurality of pieces of training data V1 includes a pair of image data V1a and region data V1b.
[0069] The image data V1a represents a known image including an image of the detection target M1. The image data V1a is generated by the camera 2. The image data V1a may be generated by a 360-degree camera different from the camera 2. The image data V1a may be generated by a known image synthesis technique.
[0070] The region data V1b indicates a region representing the detection object M1 in the image indicated by the image data V1a paired with the region data V1b. The region data V1b indicates a rectangular region that encompasses the detection object M1 as the region representing the detection object M1. The rectangular region that encompasses the detection object M1 has an aspect ratio AP. In other words, the aspect ratio of the rectangular region indicated by the region data V1b is the same as the aspect ratio of the multiple region candidates 2d.
[0071] The region data V1b includes position data V1b1 and size data V1b2. The position data V1b1 indicates the center position of a rectangular region containing the detection object M1 using rotation angles θy and θp. The center position of the rectangular region containing the detection object M1 is, for example, the position of the intersection of the diagonals of the rectangular region containing the detection object M1. The size data V1b2 indicates the size of the rectangular region containing the detection object M1. The size data V1b2 indicates the ratio of the size of the rectangular region containing the detection object M1 to the size of a reference rectangular region. The size data V1b2 is also referred to as zoom data.
[0072] The region data V1b means the correct answer that the provisional model 41a should output when the image data V1a is input to the provisional model 41a.
[0073] When the machine learning system 6 inputs the image data V1a to the provisional model 41a, the provisional model 41a outputs region data V1c. The region data V1c indicates a rectangular region in the image represented by the input image data V1a where the detection target M1 is estimated to exist.
[0074] The region data V1c includes position data V1c1 and size data V1c2. The position data V1c1 indicates the center position of a rectangular region where the detection object M1 is estimated to exist, using rotation angles θy and θp. The center position of the rectangular region where the detection object M1 is estimated to exist is, for example, the position of the intersection of the diagonals of the rectangular region where the detection object M1 is estimated to exist. The size data V1c2 indicates the size of the rectangular region where the detection object M1 is estimated to exist. The size data V1c2 indicates the ratio of the size of the rectangular region where the detection object M1 is estimated to exist to the size of the reference rectangular region. The size data V1c2 is also referred to as zoom data.
[0075] The machine learning system 6 calculates an error function using multiple pieces of training data V1 and the provisional model 41a. The error function represents the error between the region data V1c output by the provisional model 41a when the machine learning system 6 inputs image data V1a to the provisional model 41a, and the region data V1b paired with the input image data V1a. The machine learning system 6 updates the multiple coefficients U1a so as to reduce the error represented by the error function. The machine learning system 6 determines the provisional model 41a as the estimation model 41 when the process of updating the multiple coefficients U1a using each of the multiple pieces of training data V1 is completed.
[0076] The estimation model 41 outputs statistically valid region data V1c for unknown image data V1a in relation to image data V1a and region data V1b in a plurality of training data V1. The estimation model 41 is a trained model that has learned the relationship between image data V1a and region data V1b. When captured image data J1 showing a scene of a rehearsal performance F1 is used as unknown image data V1a, the estimation model 41 can accurately identify a rectangular region in which a detection target M1 exists in the image (captured image K1) indicated by the captured image data J1.
[0077] The detection unit 111 shown in FIG. 11 inputs captured image data J1 showing a scene of a rehearsal performance F1 to the estimation model 41. When the captured image data J1 is input to the estimation model 41, the detection unit 111 acquires region data V1c output from the estimation model 41 as image region data N1 showing an image region K11. As shown in FIG. 6, the image region K11 is a rectangular region that includes the detection object M1. Therefore, the detection unit 111 detects the detection object M1 by acquiring the image region data N1 from the estimation model 41.
[0078] The detection unit 111 detects the detection object M2 using an estimation model 42. The estimation model 42 is a trained model that has learned the relationship between the captured image data J (captured image K) and the area representing the detection object M2 through machine learning. The estimation model 42 is configured by a deep neural network. The estimation model 42 may include a combination of multiple types of deep neural networks.
[0079] The estimation model 42 has a plurality of coefficients U2. The plurality of coefficients U2 determine the operation of the estimation model 42. The plurality of coefficients U2 have been adjusted by machine learning.
[0080] The estimation model 42 is generated in the same manner as the estimation model 41. To generate the estimation model 42, a plurality of pieces of training data V2 are used instead of the plurality of pieces of training data V1. The plurality of pieces of training data V2 are different from each other.
[0081] 13 is a diagram showing an example of the training data V2. Each training data V2 includes a pair of image data V2a and region data V2b.
[0082] The image data V2a represents a known image including an image of the detection target M2. The image data V2a is generated by the camera 2. The image data V2a may be generated by a 360-degree camera different from the camera 2. The image data V2a may be generated by a known image synthesis technique.
[0083] The region data V2b indicates a region representing the detection object M2 in the image indicated by the image data V2a paired with the region data V2b. The region data V2b indicates a rectangular region that encompasses the detection object M2 as the region representing the detection object M2. The rectangular region that encompasses the region data V2b has an aspect ratio AP. In other words, the aspect ratio of the rectangular region indicated by the region data V2b is the same as the aspect ratio of the multiple region candidates 2d.
[0084] The region data V2b includes position data V2b1 and size data V2b2. The position data V2b1 indicates the center position of a rectangular region containing the detection object M2 using rotation angles θy and θp. The center position of the rectangular region containing the detection object M2 is, for example, the position of the intersection of the diagonals of the rectangular region containing the detection object M2. The size data V2b2 indicates the size of the rectangular region containing the detection object M2. The size data V2b2 indicates the ratio of the size of the rectangular region containing the detection object M2 to the size of a reference rectangular region. The size data V2b2 is also referred to as zoom data.
[0085] The region data V2b represents the correct answer that the estimation model 42 should output when the image data V2a is input to the estimation model 42. When the image data V2a is input to the estimation model 42, the estimation model 42 outputs region data V2c. The region data V2c indicates a rectangular region in the image represented by the input image data V2a where the detection object M2 is estimated to exist.
[0086] The region data V2c includes position data V2c1 and size data V2c2. The position data V2c1 indicates the center position of a rectangular region where the detection object M2 is estimated to exist, using rotation angles θy and θp. The center position of the rectangular region where the detection object M2 is estimated to exist is, for example, the position of the intersection of the diagonals of the rectangular region where the detection object M2 is estimated to exist. The size data V2c2 indicates the size of the rectangular region where the detection object M2 is estimated to exist. The size data V2c2 indicates the ratio of the size of the rectangular region where the detection object M2 is estimated to exist to the size of the reference rectangular region. The size data V2c2 is also referred to as zoom data.
[0087] 11 outputs statistically valid region data V2c for unknown image data V2a in relation to image data V2a and region data V2b in a plurality of training data V2. The estimation model 42 is a trained model that has learned the relationship between image data V2a and region data V2b. When captured image data J1 showing a scene of a rehearsal performance F1 is used as unknown image data V2a, the estimation model 42 can accurately identify a rectangular region in which a detection target M2 exists in the image (captured image K1) indicated by the captured image data J1.
[0088] The detection unit 111 inputs captured image data J1 showing a scene of a rehearsal performance F1 to the estimation model 42. When the captured image data J1 is input to the estimation model 42, the detection unit 111 acquires region data V2c output from the estimation model 42 as image region data N2 showing the image region K12. As shown in FIG. 6, the image region K12 is a rectangular region that includes the detection object M2. Therefore, the detection unit 111 detects the detection object M2 by acquiring the image region data N2 from the estimation model 42.
[0089] The candidate determination unit 112 determines at least one of the plurality of region candidates 2d based on the result of detection of the object M (detection objects M1 and M2) by the detection unit 111. For example, the candidate determination unit 112 determines all of the plurality of region candidates 2d based on the result of detection of the object M by the detection unit 111. The candidate determination unit 112 determines the region candidate 2d1 shown in FIG. 6 using the image region data N1 indicating the image region K11. The candidate determination unit 112 determines the region candidate 2d2 shown in FIG. 6 using the image region data N2 indicating the image region K12.
[0090] A5: An example of the selection unit 12 14 is a diagram showing a selection unit 12A, which is an example of the selection unit 12. The selection unit 12A selects a target region 2e based on the degree of change in the image of each region candidate 2d in a captured image K2 representing a scene of a live performance F2. The selection unit 12A includes a motion detection unit 121 and a region selection unit 122.
[0091] The motion detection unit 121 detects the degree of change in the image of each area candidate 2d in the captured image K2 representing the actual performance F2. The image of the area candidate 2d means the image shown in the area candidate 2d in the captured image K2.
[0092] For example, for each region candidate 2d, the motion detection unit 121 detects the degree of change in the image of the region candidate 2d based on the difference between the image shown in the region candidate 2d in the captured image K2 and the image shown in the region candidate 2d in the captured image K2 immediately before the captured image K2. The larger the difference, the greater the degree of change in the image of the region candidate 2d. The smaller the difference, the smaller the degree of change in the image of the region candidate 2d. The motion detection unit 121 generates a change index indicating the degree of change in the image for each region candidate 2d.
[0093] The region selection unit 122 selects a target region 2e from among the plurality of region candidates 2d based on the change index of each region candidate 2d.
[0094] A6: Operation of determining multiple region candidates 2d 15 is a diagram showing an example of an operation for determining a plurality of area candidates 2d. The operation for determining a plurality of area candidates 2d is executed before the actual performance F2. The operation for determining a plurality of area candidates 2d is started when the operation device 1a receives a determination instruction from the user. An example in which the determination unit 11A shown in FIG. 11 is used as the determination unit 11 will be described below.
[0095] In step S101, the detection unit 111 acquires captured image data J1 showing a scene of a rehearsal performance F1. For example, the detection unit 111 acquires the captured image data J1 from the camera 2. If the captured image data J1 is stored in the storage device 1e, the detection unit 111 may acquire the captured image data J1 from the storage device 1e.
[0096] Next, in step S102, the detection unit 111 detects the object M (detection objects M1 and M2) using the captured image data J1. For example, the detection unit 111 first inputs the captured image data J1 to each of the estimation models 41 and 42. Next, the detection unit 111 acquires the region data V1c output by the estimation model 41 as image region data N1 indicating the image region K11. As shown in FIG. 6, the image region K11 is a region representing the detection object M1. Next, the detection unit 111 acquires the region data V2c output by the estimation model 42 as image region data N2 indicating the image region K12. As shown in FIG. 6, the image region K12 is a region representing the detection object M2.
[0097] Next, in step S103, the candidate designator 112 shown in FIG. 11 determines a plurality of region candidates 2d in the imaging region 2a of the camera 2. For example, the candidate designator 112 determines a plurality of region candidates 2d as shown in FIG. 6. As an example, the candidate designator 112 determines a range indicated by image region data N1 indicating image region K11 in the imaging region 2a of the camera 2 as region candidate 2d1. In this case, the image region data N1 indicates region candidate 2d1 in addition to image region K11. The candidate designator 112 determines a range indicated by image region data N2 indicating image region K12 in the imaging region 2a of the camera 2 as region candidate 2d2. In this case, the image region data N2 indicates region candidate 2d2 in addition to image region K12.
[0098] Subsequently, in step S104, the candidate designator 112 generates candidate data R indicating the plurality of region candidates 2d. The candidate data R includes data indicating the region candidate 2d1 (image region data N1) and data indicating the region candidate 2d2 (image region data N2).
[0099] Subsequently, in step S105, the candidate designator 112 stores the candidate data R in the storage device 1e. When the candidate data R is stored in the storage device 1e, the operation of determining the plurality of region candidates 2d is completed.
[0100] A7: Operation to generate performance data Q (video content) FIG. 16 is a diagram showing an example of the operation of generating performance data Q. The operation of generating performance data Q is started when the operation device 1a receives a generation instruction from the user. Below, an example will be described in which the selection unit 12A shown in FIG. 12 is used as the selection unit 12. In response to receiving the generation instruction, the motion detection unit 121 temporarily resets the past captured image data J2. Furthermore, it is assumed that the operation of generating performance data Q is performed in parallel with the actual performance F2.
[0101] In step S201, the motion detection unit 121 reads candidate data R indicating a plurality of region candidates 2d from the storage device 1e.
[0102] Next, in step S202, the motion detection unit 121 acquires the oldest pair of two consecutive captured image data J2 from among the captured image data J2 that the motion detection unit 121 has not yet acquired.
[0103] Next, in step S203, the motion detection unit 121 generates a change index for each region candidate 2d using the two captured image data J2 acquired in the immediately preceding step S202. The change index indicates the degree of change in the image shown in the region candidate 2d within the captured image K2 representing the scene of the actual performance F2.
[0104] Hereinafter, of the two captured image data J2 acquired in the immediately preceding step S202, the captured image K2 indicated by the older captured image data J2 will be referred to as "captured image K21," and the captured image K2 indicated by the newer captured image data J2 will be referred to as "captured image K22."
[0105] In step S203, the motion detection unit 121 first detects, for each region candidate 2d, the difference between the image shown in the region candidate 2d in the captured image K21 and the image shown in the region candidate 2d in the captured image K22 as the degree of change in the image of the region candidate 2d. Next, the motion detection unit 121 generates a change index indicating the degree of change (difference) in the image for each region candidate 2d. The change index for region candidate 2d1 represents the change index for vocalist D1. The change index for region candidate 2d2 represents the change index for instrument player D2. The motion detection unit 121 increases the value of the change index as the degree of change (difference) in the image increases. The motion detection unit 121 may also decrease the value of the change index as the degree of change (difference) in the image increases. The motion detection unit 121 may generate a change index for each region candidate 2d at predetermined intervals (e.g., every second). The predetermined interval is not limited to one second and may be longer or shorter than one second. For example, the motion detection unit 121 acquires a plurality of consecutive captured image data J2 that are newly input at each predetermined interval. Next, the motion detection unit 121 generates a change index for each region candidate 2d using the newly input consecutive captured image data J2. For example, for each region candidate 2d, the motion detection unit 121 sums up the mutual differences between the images of the region candidate 2d in the captured images K2 indicated by the newly input consecutive captured image data J2. For each region candidate 2d, the motion detection unit 121 detects the sum of the mutual differences as the degree of change in the image of the region candidate 2d. Next, as shown in step S203, the motion detection unit 121 generates a change index for each region candidate 2d that indicates the degree of change in the image.
[0106] Next, in step S204, the region selection unit 122 selects a target region 2e from among the multiple region candidates 2d based on the change index of each region candidate 2d. Note that if the motion detection unit 121 generates a change index for each region candidate 2d at predetermined time intervals, the region selection unit 122 selects a target region 2e from among the multiple region candidates 2d based on the new change index for each region candidate 2d each time a new change index for each region candidate 2d is generated.
[0107] If the greater the degree of change in the image, the greater the value of the change index, the region selection unit 122 selects, from among the multiple region candidates 2d, the region candidate 2d having the largest change index value as the target region 2e.
[0108] When there are multiple region candidates 2d with the largest change index value, the region selection unit 122 selects a target region 2e from the multiple region candidates 2d with the largest change index value. For example, the region selection unit 122 randomly selects a target region 2e from the multiple region candidates 2d with the largest change index value. When priorities are set for the multiple region candidates 2d, the region selection unit 122 may select the region candidate 2d with the highest priority as the target region 2e from the multiple region candidates 2d with the largest change index value.
[0109] If the greater the degree of change in the image, the smaller the value of the change index, the region selection unit 122 selects, from among the multiple region candidates 2d, the region candidate 2d having the smallest change index value as the target region 2e.
[0110] When there are multiple region candidates 2d with the smallest change index value, the region selection unit 122 selects the target region 2e from the multiple region candidates 2d with the smallest change index value. For example, the region selection unit 122 randomly selects the target region 2e from the multiple region candidates 2d with the smallest change index value. When priorities are set for the multiple region candidates 2d, the region selection unit 122 may select the region candidate 2d with the highest priority from the multiple region candidates 2d with the smallest change index value as the target region 2e.
[0111] A large degree of change (difference) in the image means that the movement of the performer D shown in the image is large. Large movements of the performer D mean that the performer D is likely to be in a state of attracting attention. A state in which the performer D is attracting attention is, for example, when the performer is playing a solo part or when the performer is making a large action. For this reason, the region selection unit 122 selects the region candidate 2d that indicates the performer D in a state of attracting attention as the target region 2e.
[0112] 8, the extraction unit 13 extracts an image of the target region 2e from the captured image K2 representing the scene of the actual performance F2 as the output image P. For example, the extraction unit 13 extracts the latest image of the target region 2e from the captured image K2 as the output image P.
[0113] Subsequently, in step S206, the extraction unit 13 generates output image data T representing the output image P.
[0114] Next, in step S207, the generation unit 14 generates performance data Q including output image data T and performance sound data L. The performance sound data L is sound data generated by the microphone 3 during the actual performance F2. Therefore, the performance data Q represents the actual performance F2 in the form of images and sounds.
[0115] In step S207, the generation unit 14 adjusts the resolution (the number of pixels in the horizontal direction and the number of pixels in the vertical direction) of the output image data T to a resolution for transmission. The resolution for transmission is set in advance.
[0116] Subsequently, in step S208, the output control unit 15 provides the output image data T included in the performance data Q to the display device 1b, and causes the display device 1b to display the output image P indicated by the output image data T.
[0117] Subsequently, in step S209, the output control unit 15 provides the performance sound data L included in the performance data Q to the speaker 1c, causing the speaker 1c to emit the performance sound indicated by the performance sound data L.
[0118] Subsequently, in step S210, the communication control unit 16 transmits the performance data Q from the communication device 1d to the external device 5 via the communication network NW.
[0119] The processing order from step S208 to step S210 can be changed as appropriate.
[0120] Next, in step S211, the motion detection unit 121 determines whether or not there is any unacquired captured image data J2. If the motion detection unit 121 determines that there is any unacquired captured image data J2, the process returns to step S202, and the above-described operations are repeated. Therefore, the selection unit 12A sequentially selects the target areas 2e in parallel with, for example, the actual performance F2.
[0121] As the above-described operation is repeated, the selection unit 12A switches the target area 2e in accordance with the movements of each performer D during the performance F2 of the actual performance. As a result, performance data Q is generated that shows the performers D in a switching state, which are in a state of being noticed. If the motion detection unit 121 determines in step S211 that there is no unacquired captured image data J2, the operation shown in FIG. 16 ends. Note that this also ends the operation if, for example, processing of captured image data J2 older than the latest captured image data J2 showing a scene from the live performance F2 is completed before the latest captured image data J2, which represents a scene from the live performance F2, arrives at the performance recording system 1. Therefore, if the motion detection unit 121 determines in step S211 that there is no unacquired captured image data J2, it may wait a waiting time until at least consecutive captured image data J2 can be acquired. The waiting time is, for example, 0.5 seconds. The waiting time is not limited to 0.5 seconds and may be longer or shorter than 0.5 seconds. In this case, if the motion detection unit 121 acquires at least consecutive captured image data J2 during the waiting time, the process returns to step S202. If the motion detection unit 121 is unable to acquire at least consecutive captured image data J2 even after the waiting time has elapsed, the operation shown in FIG. 16 ends.
[0122] A8: Summary of the first embodiment The determination unit 11 determines multiple area candidates 2d using a captured image K1 representing a scene of a rehearsal performance F1. The selection unit 12 selects a target area 2e from the multiple area candidates 2d. The extraction unit 13 extracts an image (output image P) of a portion corresponding to the target area 2e from a captured image K2 representing a scene of the actual performance F2. This reduces the effort required to create a musical work for a group of multiple performers D. Furthermore, the output image P can be generated with a simple configuration consisting of a performance recording system 1, a camera 2, and a microphone 3.
[0123] The detection unit 111 detects objects M (at least parts of the bodies of multiple performers D and musical instruments E) from the captured image K1. The candidate determination unit 112 determines at least one of multiple region candidates 2d based on the detection results by the detection unit 111. Therefore, at least one of multiple region candidates 2d can be automatically determined based on the detection results of at least parts of the bodies of performers D and musical instruments E. This further reduces the user's effort.
[0124] The selection unit 12A selects the target region 2e based on the degree of change in the image of each candidate region 2d in the captured image K2 representing the scene of the live performance F2. This allows the target region 2e to be selected automatically, further reducing the user's effort.
[0125] The extraction unit 13 extracts the output image P from the captured image K2 that represents the scene of the live performance F2. This makes it easy to create an image work of the performance of a performing group B that includes multiple performers D.
[0126] The extraction unit 13 extracts the output image P from the captured image K2 at a timing based on the timing when the target area 2e is selected. Therefore, the output image P can be extracted at a timing based on the timing when the target area is selected (for example, at a timing according to the selection of the target area).
[0127] B: Modified example Modifications of the first embodiment are shown below. Two or more aspects arbitrarily selected from the following aspects may be combined as appropriate within the scope of not contradicting each other.
[0128] B1: First modified example In the first embodiment, the number of region candidates 2d may be greater than the number of performers D. For example, if the number of performers D is three, the number of region candidates 2d may be four or more. In addition to region candidates 2d1 and 2d2, the determination unit 11 may determine region candidate 2d3 corresponding to the position of the face of vocalist D1 and region candidate 2d4 corresponding to the position of the hands of instrument performer D2 as the multiple region candidates 2d.
[0129] FIG. 17 is a diagram showing examples of region candidates 2d3 and 2d4. If the multiple region candidates 2d include region candidate 2d3 corresponding to the position of the face of vocalist D1, it is possible to generate an output image P showing an action such as eye contact by vocalist D1. If the multiple region candidates 2d include region candidate 2d4 corresponding to the position of the hands of instrument player D2, it is possible to generate an output image P showing the operation of instrument E by instrument player D2. For example, it is possible to generate an output image P that focuses on the performance of instrument E by instrument player D2's hands. The determination unit 11 determines region candidates 2d3 and 2d4, similar to region candidates 2d1 and 2d2, by, for example, executing an image processing technique that uses an estimation model such as a trained model.
[0130] The determination unit 11 may identify image candidates that include all of the multiple performers D, and determine area candidates that correspond to the image candidates (image candidates that include all of the multiple performers D). In this case, the determination unit 11 identifies image candidates that include all of the multiple performers D by executing an image processing technique that uses an estimation model such as a trained model.
[0131] According to the first modified example, two or more region candidates 2d can be set for at least one performer D. This makes it possible to generate output images P from a variety of angles for at least one performer D.
[0132] The determination unit 11 may change the multiple area candidates 2d depending on the genre of musical piece C, the title of musical piece C, or the genre of performing group B. For example, if the genre of musical piece C is rock, the determination unit 11 selects area candidates 2d1 to 2d4 as the multiple area candidates 2d. If the genre of musical piece C is jazz, the determination unit 11 selects area candidates 2d1 to 2d2 as the multiple area candidates 2d. If the genre of performing group B is a rock band, the determination unit 11 selects area candidates 2d1 to 2d4 as the multiple area candidates 2d. If the genre of performing group B is a jazz band, the determination unit 11 selects area candidates 2d1 to 2d2 as the multiple area candidates 2d. In this case, the determination unit 11 receives section information indicating the genre of musical piece C, the title of musical piece C, or the genre of performing group B from the user via the operation device 1a. The determination unit 11 changes the multiple area candidates 2d based on the section information. Therefore, the user can change the multiple area candidates 2d by using the section information.
[0133] B2: Second variant In the first embodiment and the first variant, the determination unit 11 may determine at least one of the multiple region candidates 2d based on an image region selected by the user from multiple image regions (for example, image regions K11 and K12 in Figure 6).
[0134] For example, when the determination unit 11 identifies two or more image areas (e.g., image areas K11 and K12 in FIG. 6) in the captured image K1 representing the scene of the rehearsal performance F1, the determination unit 11 causes the display device 1b to display the two or more image areas. The determination unit 11 determines at least one of the multiple area candidates 2d based on an image area selected by the user from the two or more image areas displayed on the display device 1b.
[0135] For example, the determination unit 11 identifies an image area selected by a user from two or more image areas displayed on the display device 1b as a selected image area. The determination unit 11 generates image area data indicating the selected image area. The image area data indicating the selected image area indicates the center position of the selected image area by rotation angles θy and θp, and indicates the size of the selected image area by the ratio of the size of the selected image area to the size of a reference rectangular area. The determination unit 11 determines the range indicated by the image area data indicating the selected image area in the imaging area 2a as the area candidate 2d.
[0136] According to the second modification, the selection by the user is involved in determining the plurality of area candidates 2d, so that the area candidates 2d can be determined according to the user's preferences.
[0137] B3: Third variant In the first embodiment and the first to second variants, the determination unit 11 may determine at least one of the multiple area candidates 2d based on an image area set by the user in a captured image K1 representing a scene of a rehearsal performance F1.
[0138] For example, when the determination unit 11 identifies two or more image areas in the captured image K1 (for example, image areas K11 and K12 in FIG. 6), the determination unit 11 causes the display device 1b to display the two or more image areas. Based on image areas whose positions or sizes have been changed by the user from among the two or more image areas displayed on the display device 1b, the determination unit 11 determines the same number of area candidates 2d as the number of image areas (image areas whose positions or sizes have been changed by the user). The image areas changed by the user are an example of image areas set by the user.
[0139] The determination unit 11 may display the captured image K1 on the display device 1b. In this case, the determination unit 11 determines the same number of region candidates 2d as the number of image regions set by the user in the captured image K displayed on the display device 1b, based on the image regions set by the user. In this case, the determination unit 11, for example, limits the aspect ratio of the image regions set by the user to the aspect ratio AP. The determination unit 11 generates image region data indicating the image regions set by the user. The image region data indicating the image regions set by the user indicates the center position of the image region by rotation angles θy and θp, and indicates the size of the image region by the ratio of the size of the selected image region to the size of the reference rectangular region.
[0140] The method for determining the region candidate 2d based on the image region changed or set by the user is similar to the method for determining the region candidate 2d1 based on the image region K11.
[0141] According to the third modified example, the user's operation is involved in determining the multiple area candidates 2d, so the area candidates 2d can be determined according to the user's preferences.
[0142] B4: Fourth variant In the first embodiment and the first to third variants, the determination unit 11 may estimate the area in the imaging area 2a where the instrument E is located based on the sound obtained by the microphone 3 picking up the sound of the rehearsal performance F1.
[0143] In the fourth modification, the microphone 3 is, for example, a directional microphone. The directional microphone 3 is a microphone made up of multiple directional microphones. Each of the multiple microphones has a sound collection range according to its directivity. The sound collection ranges of the multiple microphones are different from each other. Note that as long as the sound collection ranges of the multiple microphones are different from each other, at least one of the multiple microphones may be an omnidirectional microphone.
[0144] The determination unit 11 identifies the microphone that picked up the loudest sound as the target microphone among the multiple microphones that make up the microphone 3. The determination unit 11 estimates the area in the imaging area 2a that overlaps with the sound pickup range of the target microphone as the area where the musical instrument E is present. The determination unit 11 may estimate the area where the musical instrument E exists based on multiple sound pickup results from multiple microphones constituting the microphone 3. For example, in a situation where the sound pickup ranges of the microphones partially overlap each other, the determination unit 11 first identifies the microphone that picked up sound at a volume equal to or higher than a reference level as the detection microphone. When the determination unit 11 identifies one microphone as the detection microphone, it estimates the area in the imaging area 2a that overlaps with the sound pickup range of the detection microphone as the area where the musical instrument E exists. When the determination unit 11 identifies multiple microphones as the detection microphones, it identifies the area where the sound pickup ranges of the detection microphones overlap as the overlap area. The determination unit 11 estimates the area in the imaging area 2a that overlaps with the overlap area as the area where the musical instrument E exists.
[0145] The determination unit 11 may determine at least one of the plurality of area candidates 2d based on the estimation result of the area where the instrument E exists. For example, the determination unit 11 determines the area where the instrument E is estimated to exist as the area candidate 2d.
[0146] According to the fourth modification, at least one of the plurality of region candidates 2d is determined based on the detection results of at least a portion of the body of the performer D and the instrument E, as well as the estimation results of the region where the instrument E is located. Therefore, the plurality of region candidates 2d can be made more diverse than in a configuration in which at least one of the plurality of region candidates 2d is determined based only on the detection results of at least a portion of the body of the performer D and the instrument E.
[0147] B5: Fifth variant In the first embodiment and the first to fourth variants, the selection unit 12 may select the target area 2e based on both the degree of change in the image of each area candidate 2d in the captured image K2 and the sound obtained when the directional microphone 3 picks up the sound of the actual performance F2.
[0148] For example, the selection unit 12 first identifies, as the detection microphone, a microphone that collects sound at or above a threshold level among the multiple microphones that make up the microphone 3. Next, the selection unit 12 identifies a region candidate 2d that overlaps with at least a portion of the sound collection range of the detection microphone. Next, the selection unit 12 changes the change index of the region candidate 2d that overlaps with at least a portion of the sound collection range of the detection microphone. If the value of the change index increases as the degree of change in the image represented by the region candidate 2d increases, the selection unit 12 increases the value of the change index of the region candidate 2d that overlaps with at least a portion of the sound collection range of the detection microphone by an adjustment value. The adjustment value is a preset value. If the value of the change index decreases as the degree of change in the image represented by the region candidate 2d increases, the selection unit 12 decreases the value of the change index of the region candidate 2d that overlaps with at least a portion of the sound collection range of the detection microphone by the adjustment value. Next, the selection unit 12 selects a target region 2e from the multiple region candidates 2d based on the change index of each region candidate 2d.
[0149] According to the fifth modification, the selection unit 12 selects the target area 2e based on both the degree of change in the image of each area candidate 2d and the sound obtained by recording the sound of the live performance F2. This allows for more diverse switching of the target areas 2e than in a configuration in which the target area 2e is selected based only on the degree of change in the image of each area candidate 2d.
[0150] The information used by the selection unit 12 to change the value of the change index is not limited to the sound of the live performance F2, but may also be one or more specific actions by the performer D. The specific action is, for example, a movement of raising the right hand, a movement of shaking the head, or a movement of moving the instrument E. The specific action is also called a special action. The multiple specific actions are different from each other. The selection unit 12 detects the specific action using, for example, image recognition technology.
[0151] The selection unit 12 changing the value of the change index based on the specific action by the performer D means that the selection unit 12 selects the target area 2e based on both the degree of change in the image of each area candidate 2d and the specific action by the performer D. This allows for more diverse switching of the target area 2e than a configuration in which the target area 2e is selected based only on the degree of change in the image of each area candidate 2d. For example, the performers D shown in the output image P can be switched by each of the multiple performers D raising their right hands in turn.
[0152] B6: 6th variant In the first embodiment and the first to fifth modified examples, the selection unit 12 may select the target region 2e based on an index obtained by weighting the degree of change in the image of each region candidate 2d in the captured image K2.
[0153] For example, in the first embodiment and the first to fifth modifications, if the instrumentalist D2 moves more than the vocalist D1 during the live performance F2, it is highly likely that many of the output images P will represent the instrumentalist D2. However, there may be a desire to have many of the output images P represent images of the vocalist D1, who moves less than the instrumentalist D2. The sixth modification is an example of a method for meeting such a desire. For example, the selection unit 12 assigns a greater weight to the change index of the vocalist D1 than to the change index of the instrumentalist D2.
[0154] The selection unit 12 calculates an index for vocalist D1 by multiplying the change index for vocalist D1 by a weighting coefficient W1. The selection unit 12 calculates an index for instrumentalist D2 by multiplying the change index for instrumentalist D2 by a weighting coefficient W2. The selection unit 12 selects a target region 2e based on both the index for vocalist D1 and the index for instrumentalist D2.
[0155] The weighting coefficients W1 and W2 are set, for example, by the user. The weighting coefficients W1 and W2 may be set in advance. The weighting coefficients W1 and W2 may be changed in response to a change instruction input by the user to the operation device 1a. The index of the vocalist D1 and the index of the instrument player D2 are each an example of an index obtained by weighting the degree of change in the image of each region candidate 2d in the captured image K2 representing the scene of the live performance F2.
[0156] If the greater the degree of change in the image shown in region candidate 2d, the greater the value of the change index, then weighting coefficient W1 is set to a value greater than weighting coefficient W2. In this situation, if the index of vocalist D1 is greater than the index of instrument player D2, the selection unit 12 selects region candidate 2d1 of vocalist D1 as target region 2e. If the index of instrument player D2 is greater than the index of vocalist D1, the selection unit 12 selects region candidate 2d2 of instrument player D2 as target region 2e.
[0157] If the greater the degree of change in the image shown in region candidate 2d, the smaller the value of the change index, then weighting coefficient W1 is set to a smaller value than weighting coefficient W2. In this situation, if the index of vocalist D1 is smaller than the index of instrument player D2, the selection unit 12 selects region candidate 2d1 of vocalist D1 as target region 2e. If the index of instrument player D2 is smaller than the index of vocalist D1, the selection unit 12 selects region candidate 2d2 of instrument player D2 as target region 2e.
[0158] If the index of vocalist D1 is equal to the index of instrumentalist D2, the selection unit 12 randomly selects the target region 2e from among the region candidate 2d1 of vocalist D1 and the region candidate 2d2 of instrumentalist D2. A situation may also be envisioned in which priority is set to the region candidate 2d1 of vocalist D1 and the region candidate 2d2 of instrumentalist D2. In this situation, if the index of vocalist D1 is equal to the index of instrumentalist D2, the selection unit 12 may select the region candidate 2d with the higher priority from the region candidates 2d1 and 2d2 as the target region 2e.
[0159] The weighting coefficients W1 and W2 are set according to the type (vocalist or instrumentalist) of the performer D. The type of performer D is not limited to vocalist or instrumentalist, but may be, for example, vocalist, guitarist, bassist, or drummer.
[0160] The weighting coefficients W1 and W2 may be set in accordance with information other than the type of performer D. For example, the weighting coefficients W1 and W2 may be set in accordance with the genre of the music piece C.
[0161] FIG. 18 is a diagram showing a genre table JT illustrating an example of weighting factors W1 and W2 corresponding to the genre of music C. The genre of music C is not limited to pop and jazz shown in FIG. 18, but may include, for example, rock and classical. The genre table JT is stored, for example, in a storage device 1e. The selection unit 12 selects weighting factors W1 and W2 corresponding to the genre input by the user to the operation device 1a from the weighting factors W1 and W2 shown in the genre table JT. Next, the selection unit 12 generates an index for vocalist D1 and an index for instrumentalist D2 by multiplying the change index by the weighting factors W1 and W2 corresponding to the input genre, respectively. The selection unit 12 selects a target region 2e based on the index for vocalist D1 and the index for instrumentalist D2.
[0162] The weighting coefficients W1 and W2 may be set according to the title of the music piece C. In this case, the selection unit 12 can change the target region 2e according to the title of the music piece C.
[0163] According to the sixth modification, the region candidate 2d selected as the target region 2e can be adjusted by weighting. Note that weighting may be performed not only on the image of the area candidate 2d but also on the volume of the sound picked up by the microphone 3, which is made up of multiple microphones. For example, to emphasize the singing voice of vocalist D1, a greater weight is applied to the gain of the microphone having the sound pickup area where vocalist D1 is present than to the gains of the other microphones. In this case, the singing voice of vocalist D1 is amplified with a higher amplification factor than other sounds, and the singing voice of vocalist D1 is emphasized. The weighting of the volume of the sound picked up by microphone 3 is not limited to a mode in which a greater weight is applied to the gain of the microphone having the sound pickup area where vocalist D1 is present than to the gains of the other microphones. For example, the weighting of the gain of each microphone may be changed according to the progress of the performance F2 during the performance.
[0164] B7: 7th variant In the first embodiment and the first to sixth variants, the selection unit 12 may select the target area 2e based on an image of the area candidate 2d selected by the user from the images of each area candidate 2d in the captured image K2 representing the scene of the actual performance F2.
[0165] For example, the selection unit 12 displays, on the display device 1b, images of each region candidate 2d in the captured image K2 representing the scene of the actual performance F2, in parallel with the actual performance F2. A user different from the multiple performers D uses the operation device 1a to select one image of the region candidate 2d from among the images of the region candidates 2d to be displayed on the display device 1b. The selection unit 12 selects the target region 2e based on the image of the region candidate 2d selected by the user. For example, the selection unit 12 selects the region candidate 2d representing the image selected by the user as the target region 2e. Note that the selection unit 12 may select the region candidate 2d representing the image selected by the user as the target region 2e by adjusting the change index of the region candidate 2d representing the image selected by the user. In this case, the target region 2e can be manually selected in parallel with the actual performance F2.
[0166] The selection unit 12 may select the target region 2e before the actual performance F2. For example, the selection unit 12 first displays, on the display device 1b, a video represented by a series of captured images K1 representing scenes from a rehearsal performance F1, indicated in each of the region candidates 2d. The user sequentially selects a video represented by each of the region candidates 2d in accordance with the progress of the video of the rehearsal performance F1. In this case, the user switches the selected video in accordance with the progress of the video of the rehearsal performance F1. The selection unit 12 sequentially selects, as the target region 2e, the region candidates 2d representing the video selected by the user in accordance with the progress of the video of the rehearsal performance F1. The selection unit 12 stores, in the storage device 1e, selection information indicating the sequential selection results of the target region 2e and the elapsed time in the rehearsal performance F1. When the actual performance F2 begins, the extraction unit 13 extracts an output image P based on the target region 2e, which changes in accordance with the elapsed time indicated by the selection information. In this case, before the actual performance F2, the output image P generated from the actual performance F2 can be predicted.
[0167] According to the seventh modification, the user is involved in the selection of the target area 2e, so that the target area 2e can be selected according to the user's preferences.
[0168] B8: Eighth Variation In the first embodiment and the first to seventh variants, the captured image K2 (an image representing the scene of the actual performance F1) input to the extraction unit 13 may be delayed from the captured image K2 (an image representing the scene of the actual performance F1) input to the selection unit 12.
[0169] The selection unit 12 selects the target region 2e based on a change in the image in the captured image K2. The extraction unit 13 extracts an output image P from the captured image K2 using the target region 2e selected based on the captured image K2. Therefore, when the captured image K2 input to the extraction unit 13 is synchronized with the captured image K2 input to the selection unit 12, an image representing a scene where performer D starts to move and an image representing a scene just before performer D starts to move are difficult to generate as the output image P. The eighth modified example is an example of a method for solving the problem that an image representing a scene where performer D starts to move and an image representing a scene just before performer D starts to move are difficult to generate as the output image P.
[0170] In the eighth modification, for example, the selection unit 12 uses the captured image K2 without delay, whereas the extraction unit 13 uses the captured image K2 with a delay of the adjustment time. The adjustment time is, for example, one second. The adjustment time may be shorter or longer than one second. The extraction unit 13 extracts the output image P from the captured image K2 that was generated the adjustment time before the captured image K2 from which the selection unit 12 extracted the target region 2e.
[0171] According to the eighth modification, an image showing the scene when the performer D starts to move and an image showing the scene just before the performer D starts to move can be easily generated as the output image P. In the eighth modification, the generation unit 14 uses the performance sound data L with a delay of the adjustment time. Therefore, synchronization between the image and the sound is maintained in the performance data Q.
[0172] B9: 9th variant In the first embodiment and the first to eighth modified examples, the extraction unit 13 may extract the output image P from the captured image K2 (switch the output image P) in time with the rhythm of the sound of the live performance F2.
[0173] For example, the extraction unit 13 estimates the rhythm (beat) of the music piece C based on the sound of an instrument E (e.g., a drum or bass) in the live performance F2. The extraction unit 13 extracts an output image P from the captured image K2 in time with the rhythm of the music piece C.
[0174] According to the ninth modification, the extraction unit 13 extracts the output image P from the captured image K2 in time with the rhythm of the sound obtained by recording the sound of the live performance F2. Therefore, the output image P2 can be extracted in time with the live performance F2.
[0175] B10: 10th variant In the first embodiment and the first to ninth modifications, the extraction unit 13 may switch the output image P in response to switching of the target area 2e, as if the camera were panned. In response to switching of the target area 2e, the extraction unit 13 may fade in the output image P before switching while fading out the output image P before switching.
[0176] According to the tenth modification, it is possible to smoothly switch the output image P. Furthermore, it is possible to visually present the switching of the output image P.
[0177] B11: 11th variant In the first embodiment and the first to tenth modifications, the selection unit 12 may select multiple target regions 2e. For example, if the greater the degree of change in the image, the greater the value of the change index, the selection unit 12 selects, from the multiple region candidates 2d, the region candidate 2d having the largest change index value and the region candidate 2d having the second largest change index value as target regions 2e.
[0178] When the selection unit 12 selects a plurality of target areas 2e, the extraction unit 13 extracts, as an output image P, an image of each target area 2e from the captured image K2 representing the scene of the live performance F2.
[0179] According to the eleventh modification, a plurality of output images P are extracted from one captured image K2, so that performance data Q showing a plurality of output images P at once can be generated.
[0180] B12: 12th variant In the first embodiment and the first to tenth modifications, the performance record is not limited to the captured image K2 generated by the camera 2 capturing an image of a scene in which multiple performers D perform a live performance F2. For example, the performance record may include performance sounds obtained by the microphone 3 picking up the sounds of the live performance F2. In this case, the portion of the performance record corresponding to the target area 2e includes, in addition to the image of the target area 2e in the captured image K2, the sound from the target area 2e among the performance sounds obtained by the microphone 3 picking up the sounds of the live performance F2. The sound from the target area 2e is identified by extracting sound pickup data indicative of the sound from the target area 2e from the sound pickup data of the directional microphone 3 (a set of multiple directional microphones).
[0181] The performance record may be either the captured image K2 or the performance sound obtained by the microphone 3 picking up the sound of the live performance F2. In this case, the portion of the performance record corresponding to the target area e2 is either the image of the target area 2e in the captured image K2 or the sound from the target area 2e among the performance sounds obtained by the microphone 3 picking up the sound of the live performance F2.
[0182] According to the twelfth modification, a musical performance piece based on the performance of a performing group B having a plurality of performers D can be easily created.
[0183] B13: 13th variant In the first embodiment and the first to twelfth modifications, the camera 2 is not limited to a 360-degree camera, and may be a camera having a field of view less than 360 degrees (for example, a 180-degree camera). If the camera 2 is not a 360-degree camera, there is no need to process the captured image generated by the camera 2 onto a plane. When the camera 2 is a 360-degree camera, the multiple performers D can perform without worrying about whether each of the multiple performers D is within the imaging area 2a.
[0184] B14: 14th variant In the first embodiment and the first to thirteenth modifications, the first and second performances are not limited to a rehearsal performance F1 and a performance F2 at a live performance. For example, if live performances are performed repeatedly, the first performance may be a performance at a past live performance, and the second performance may be a performance at a future live performance.
[0185] B15: 15th variant In the first embodiment and the first to thirteenth modifications, the performance recording system 1 may be configured by, for example, a server, rather than a smartphone, tablet, or personal computer.
[0186] C: Aspects understood from the above-mentioned embodiments and modifications The following aspects can be understood from at least one of the above-described embodiments and modifications.
[0187] C1: First mode A performance recording method according to one aspect (first aspect) of the present disclosure is a performance recording method implemented by a computer system. Using a first captured image generated by a camera capturing a scene in which multiple performers perform a first performance of a musical piece, the method determines multiple area candidates in the camera's capture area, selects a target area from the multiple area candidates, and extracts a portion corresponding to the target area from a performance recording obtained by capturing a scene in which the multiple performers perform a second performance of the musical piece or by capturing the sound of the second performance. This aspect eliminates the need to record each performer's performance. This reduces the effort required to create a musical work for a group of multiple performers.
[0188] C2: Second mode In an example of the first aspect (second aspect), determining the plurality of area candidates includes detecting at least a portion of the bodies of the plurality of performers and musical instruments from the first captured image, and determining at least one of the plurality of area candidates based on the detection results. According to this aspect, at least one of the plurality of area candidates can be automatically determined based on the detection results of at least a portion of the bodies of the performers and musical instruments. This further reduces the user's effort.
[0189] C3: Third mode In an example of the second aspect (third aspect), determining the plurality of area candidates further includes estimating an area in the imaging area where the musical instrument is located based on a sound obtained by capturing the sound of the first performance, and determining at least one of the plurality of area candidates based on the result of the estimation. According to this aspect, the plurality of area candidates can be more diverse than in a configuration in which at least one of the plurality of area candidates is determined based only on the detection result of at least a part of the player's body and the musical instrument.
[0190] C4: Fourth mode In an example of the first aspect (fourth aspect), determining the plurality of area candidates includes detecting at least parts of the bodies of the plurality of performers and musical instruments from the first captured image, identifying a plurality of image areas in the first captured image based on the detection results, and determining at least one of the plurality of area candidates based on an image area selected by a user from among the plurality of image areas. According to this aspect, the determination of the plurality of area candidates involves user selection, so that the area candidates can be determined according to the user's preferences.
[0191] C5: Fifth mode In an example of the first aspect (fifth aspect), determining the plurality of area candidates includes determining at least one of the plurality of area candidates based on an image area set by a user in the first captured image. According to this aspect, the area candidate can be determined according to the user's preference.
[0192] C6: Sixth mode In any one of the first to fifth aspects (sixth aspect), selecting the target area includes selecting the target area based on the degree of change in the image of each candidate area in the second captured image representing the second performance. According to this aspect, the target area can be automatically selected based on the degree of change in the image of each candidate area in the second captured image. This further reduces the user's effort.
[0193] C7: Seventh mode In an example of the sixth aspect (seventh aspect), selecting the target region includes selecting the target region based on an index obtained by weighting the degree of change in the image of each of the candidate regions. According to this aspect, the candidate regions selected as the target region can be adjusted by weighting.
[0194] C8: Eighth mode In any one of the first to fifth aspects (eighth aspect), selecting the target area includes selecting the target area based on an image of a candidate area selected by a user from images of candidate areas in the second captured image representing the second performance. According to this aspect, the target area can be selected according to the user's preference.
[0195] C9: 9th mode In any one of the first to eighth aspects (ninth aspect), extracting the portion corresponding to the target region from the performance record includes extracting the portion corresponding to the target region from the performance record at a timing corresponding to the selection of the target region. According to this aspect, the portion corresponding to the target region can be extracted at a timing corresponding to the selection of the target region.
[0196] C10: 10th mode In any one of the first to eighth aspects (tenth aspect), extracting the portion corresponding to the target region from the performance record includes extracting the portion corresponding to the target region from the performance record in time with the rhythm of the sound obtained by recording the sound of the second performance. According to this aspect, the portion corresponding to the target region can be extracted in time with the second performance.
[0197] C11: 11th mode In any one of the first to tenth aspects (eleventh aspect), the performance record is a second captured image generated by the camera capturing an image of the multiple performers performing the second performance, and the portion corresponding to the target area is an image of the target area in the second captured image. According to this aspect, a video work of a performance by a group of multiple performers can be easily created.
[0198] C12: 12th mode In any one of the first to tenth aspects (twelfth aspect), the performance record is a performance sound obtained by picking up the sound of the second performance with a microphone, and the part corresponding to the target area is a sound from the target area among the performance sounds. According to this aspect, a performance sound work of a performance by a group of multiple performers can be easily created.
[0199] C13: 13th mode A performance recording system according to a thirteenth aspect of the present disclosure includes: a determining unit that determines a plurality of area candidates in the camera's image capture area using a first captured image generated by the camera capturing a scene in which multiple performers perform a first performance of a musical piece; a selecting unit that selects a target area from the plurality of area candidates; and an extracting unit that extracts a portion corresponding to the target area from a performance record obtained by capturing an image of the scene in which the multiple performers perform a second performance of the musical piece or by capturing the sound of the second performance. This aspect eliminates the need to record each performer's performance, thereby reducing the effort required to create a musical work for a group of multiple performers.
[0200] C14: 14th mode A program according to a fourteenth aspect of the present disclosure causes a computer system to function as a determiner that uses a first captured image generated by a camera capturing a scene in which multiple performers perform a first performance of a musical piece to determine multiple area candidates in the camera's captured area, a selector that selects a target area from the multiple area candidates, and an extractor that extracts a portion corresponding to the target area from a performance record obtained by capturing an image of a scene in which the multiple performers perform a second performance of the musical piece or by recording the sound of the second performance. This aspect eliminates the need to record each performer's performance, thereby reducing the effort required to create a musical work for a group of multiple performers. [Explanation of symbols]
[0201] 1...performance recording system, 1a...operation device, 1b...display device, 1c...speaker, 1d...communication device, 1e...storage device, 1f...processing device, 2...camera, 3...microphone, 11...determination unit, 11A...determination unit, 12...selection unit, 12A...selection unit, 13...extraction unit, 14...generation unit, 15...output control unit, 16...communication control unit, 41...estimated model, 41a...tentative model, 42...estimated model, 111...detection unit, 112...candidate determination unit, 121...motion detection unit, 122...area selection unit.
Claims
1. A performance recording method realized by a computer system, determining a plurality of area candidates in an imaging area of the camera using a first captured image generated by the camera capturing an image of a scene in which a plurality of performers are performing a first performance of the musical piece; selecting a target region from the plurality of region candidates; extracting a portion corresponding to the target region from a performance record obtained by capturing an image of a scene in which the plurality of performers perform a second performance of the musical piece or by collecting sounds of the second performance, at a timing corresponding to the selection of the target region; How to record a performance.
2. A performance recording method realized by a computer system, determining a plurality of area candidates in an imaging area of the camera using a first captured image generated by the camera capturing an image of a scene in which a plurality of performers are performing a first performance of the musical piece; selecting a target region from the plurality of region candidates; extracting a portion corresponding to the target region from a performance record obtained by capturing an image of a scene in which the plurality of performers perform a second performance of the musical piece or by collecting sounds of the second performance in accordance with the rhythm of the sounds of the second performance; How to record a performance.
3. Determining the plurality of region candidates includes: detecting at least a portion of the bodies of the plurality of performers and musical instruments from the first captured image; determining at least one of the plurality of region candidates based on a result of the detection.
3. A performance recording method according to claim 1 or 2.
4. Determining the plurality of region candidates includes: estimating an area in the imaging area where a musical instrument is present based on a sound obtained by collecting the sound of the first performance; determining at least one of the plurality of region candidates based on a result of the estimation.
3. A performance recording method according to claim 1 or 2.
5. Determining the plurality of region candidates includes: determining at least one of the plurality of region candidates based on an image region selected by a user from among a plurality of image regions in the first captured image. The performance recording method according to any one of claims 1 to 4.
6. Determining the plurality of region candidates includes: determining at least one of the plurality of area candidates based on an image area set by a user in the first captured image. The performance recording method according to any one of claims 1 to 5.
7. selecting the region of interest comprises: selecting the target region based on a degree of change in the image of each candidate region in the second captured image representing the second performance; The performance recording method according to any one of claims 1 to 6.
8. selecting the region of interest comprises: selecting the target region based on an index obtained by weighting the degree of change in the image of each of the candidate regions; The performance recording method according to claim 7.
9. selecting the region of interest comprises: selecting the target area based on an image of a candidate area selected by a user from images of candidate areas in a second captured image representing the second performance; The performance recording method according to any one of claims 1 to 6.
10. the performance record is a second captured image generated by the camera capturing an image of the plurality of performers performing the second performance, the portion corresponding to the target area is an image of the target area in the second captured image. The performance recording method according to any one of claims 1 to 9.
11. the performance record is a performance sound obtained by collecting the sound of the second performance with a microphone, the portion corresponding to the target region is a sound from the target region among the performance sounds; The performance recording method according to any one of claims 1 to 9.
12. a determination unit that determines a plurality of area candidates in an image capture area of a camera using a first captured image generated by the camera capturing an image of a scene in which a plurality of performers are performing a first performance of a piece of music; a selection unit that selects a target region from the plurality of region candidates; an extracting unit that extracts a portion corresponding to the target region from a performance record obtained by capturing an image of a scene in which the plurality of performers perform a second performance of the musical piece or by collecting sounds of the second performance, at a timing corresponding to the selection of the target region; A performance recording system including
13. a determination unit that determines a plurality of area candidates in an image capture area of a camera using a first captured image generated by the camera capturing an image of a scene in which a plurality of performers are performing a first performance of a piece of music; a selection unit that selects a target region from the plurality of region candidates; an extracting unit that extracts a portion corresponding to the target region from a performance record obtained by capturing an image of a scene in which the plurality of performers perform a second performance of the musical piece or by capturing a sound of the second performance in accordance with the rhythm of the sound of the second performance; A performance recording system including
14. a determination unit that determines a plurality of area candidates in an image capture area of a camera using a first captured image generated by the camera capturing an image of a scene in which a plurality of performers are performing a first performance of a piece of music; a selection unit that selects a target region from the plurality of region candidates; and an extraction unit that extracts a portion corresponding to the target region from a performance record obtained by capturing an image of a scene in which the plurality of performers perform a second performance of the musical piece or by collecting sounds of the second performance, at a timing corresponding to the selection of the target region; A program that makes a computer system function as a
15. a determination unit that determines a plurality of area candidates in an image capture area of a camera using a first captured image generated by the camera capturing an image of a scene in which a plurality of performers are performing a first performance of a piece of music; a selection unit that selects a target region from the plurality of region candidates; and an extraction unit that extracts a portion corresponding to the target region from a performance record obtained by capturing an image of a scene in which the plurality of performers perform a second performance of the musical piece or by collecting sounds of the second performance in accordance with the rhythm of the sounds of the second performance; A program that makes a computer system function as a
Citation Information
Patent Citations
Music creating method, device, system and program
JP2015031885A
Network-based processing and delivery of multimedia content for live music performances
JP2019525571A
Video sound processing device, video sound processing method, and program
WO2017208820A1