Image processing method, image processing system, and storage medium

By generating a transformation matrix for projective transformation, the problem of image resolution accuracy caused by changes in the shooting conditions of keyboard instruments is solved, and high-precision performance image generation and score generation are achieved.

CN117083635BActive Publication Date: 2026-04-21YAMAHA CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YAMAHA CORP
Filing Date
2022-03-07
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In the existing technology, the shooting conditions of the shooting device for keyboard instruments may be different each time, making it difficult to generate high-precision scores.

Method used

By generating a transformation matrix to perform a projective transformation, the images of the instrument and the performer's fingers are made close to the reference image in the performance image. The projective transformation is then performed using the projective transformation unit to generate a high-precision performance image.

Benefits of technology

It enables the generation of high-precision performance images under specific shooting conditions and supports accurate generation of musical scores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117083635B_ABST
    Figure CN117083635B_ABST
Patent Text Reader

Abstract

The performance analysis system (100) has a matrix generating section (312) that generates a transformation matrix (W) that performs a projective transformation of a performance image in such a manner that an image of a musical instrument included in the performance image and an image of a plurality of fingers of a user playing the musical instrument are close to an image of a reference musical instrument included in a reference image, and a projective transformation section (314) that performs a projective transformation of the performance image using the transformation matrix (W).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a technique for analyzing a user's performance. Background Technology

[0002] Various techniques have been proposed for analyzing images of users playing instruments such as keyboard instruments captured by a camera. For example, Patent Document 1 discloses a technique for generating sheet music for a piece of music by analyzing images captured of a user playing a keyboard instrument.

[0003] Patent Document 1: US Patent No. 9,418,637 Summary of the Invention

[0004] However, the shooting conditions for a camera device used to photograph the keyboard of a keyboard instrument may differ each time it is shot. In view of the above, one objective of the present invention is to generate an image of an instrument played by a user, captured under specific shooting conditions.

[0005] To address the above issues, one aspect of the present invention relates to an image processing method in which a transformation matrix for projective transformation of the coordinates of the performance image is generated such that the image of the instrument in a performance image, including an image of the instrument and images of multiple fingers of a user playing the instrument, approximates a reference image representing a reference instrument, and the projective transformation of the performance image is performed using the transformation matrix.

[0006] One aspect of the present invention relates to an image processing system comprising: a matrix generation unit that generates a transformation matrix for projectively transforming the performance image in such a way that the image of the instrument in a performance image, including an image of an instrument and images of multiple fingers of a user playing the instrument, approximates an image of a reference instrument included in a reference image; and a projective transformation unit that performs a projective transformation of the performance image using the transformation matrix.

[0007] One aspect of the present invention relates to a program that enables a computer system to function as a matrix generation unit that generates a transformation matrix for projectively transforming the performance image in such a way that the image of the instrument in a performance image, which includes an image of the instrument and images of multiple fingers of a user playing the instrument, approximates an image of a reference instrument included in a reference image; and a projective transformation unit that performs a projective transformation of the performance image using the transformation matrix. Attached Figure Description

[0008] Figure 1 This is a block diagram illustrating the structure of the performance analysis system according to the first embodiment.

[0009] Figure 2This is a schematic diagram of the performance image.

[0010] Figure 3 This is a block diagram illustrating the functional structure of a performance analysis system.

[0011] Figure 4 This is a diagram illustrating the analysis of the screen.

[0012] Figure 5 This is a flowchart of the finger position estimation process.

[0013] Figure 6 This is a flowchart of the left and right decision-making process.

[0014] Figure 7 This is an illustration of image extraction and processing.

[0015] Figure 8 This is a flowchart of image extraction and processing.

[0016] Figure 9 This is an illustration of machine learning for creating an inference model.

[0017] Figure 10 This is a schematic diagram based on the reference image.

[0018] Figure 11 This is a flowchart of the matrix generation process.

[0019] Figure 12 This is a flowchart of the initial setup process.

[0020] Figure 13 This is a diagram of the settings screen.

[0021] Figure 14 This is a flowchart of the performance analysis and processing.

[0022] Figure 15 This is an explanatory diagram related to the topic of index estimation.

[0023] Figure 16 This is a block diagram illustrating the structure of the performance analysis system according to the second embodiment.

[0024] Figure 17 This is a schematic diagram of the control data in the second embodiment.

[0025] Figure 18 This is a flowchart of the performance analysis process in the second embodiment.

[0026] Figure 19 This is a flowchart of the performance analysis process in the third embodiment.

[0027] Figure 20 This is a flowchart of the initial setup process in the fourth embodiment.

[0028] Figure 21 This is a block diagram illustrating the structure of the performance analysis system according to the fifth embodiment.

[0029] Figure 22 This is a block diagram illustrating the functional structure of the image processing system according to the sixth embodiment.

[0030] Figure 23 This is a flowchart of the first image processing step in the sixth embodiment.

[0031] Figure 24 This is a block diagram illustrating the functional structure of the image processing system according to the seventh embodiment.

[0032] Figure 25 This is a flowchart of the second image processing step in the seventh embodiment. Detailed Implementation

[0033] 1: First Implementation Method

[0034] Figure 1 This is a block diagram illustrating the structure of the performance analysis system 100 according to the first embodiment. In the performance analysis system 100, a keyboard instrument 200 is connected via wired or wireless means. The keyboard instrument 200 is an electronic instrument having a keyboard 22 with a plurality of (N) keys 21 arranged thereon. Each of the plurality of keys 21 on the keyboard 22 corresponds to a different pitch n (n = 1 to N). The user (i.e., the performer) operates the desired keys 21 on the keyboard instrument 200 sequentially using their left and right hands. The keyboard instrument 200 supplies performance data P representing the user's performance to the performance analysis system 100. The performance data P is time-series data that specifies the pitch n of each of the plurality of notes played sequentially by the user. For example, the performance data P is data in the form of, for example, conforming to the MIDI (Musical Instrument Digital Interface) standard.

[0035] The performance analysis system 100 is a computer system that analyzes the performance of a keyboard instrument 200 by a user. Specifically, the performance analysis system 100 analyzes the user's finger movements. Finger movements are the methods by which the user uses the fingers of both hands (i.e., fingering techniques) in playing the keyboard instrument 200. That is, information such as which finger the user uses to operate each key 21 of the keyboard instrument 200 is analyzed as the user's finger movements.

[0036] The performance analysis system 100 includes a control device 11, a storage device 12, an operation device 13, a display device 14, and a camera device 15. The performance analysis system 100 can be implemented using a portable information device such as a smartphone or tablet terminal, or a portable or fixed information device such as a personal computer. Furthermore, the performance analysis system 100 can be implemented not only as a single device, but also as multiple devices configured separately from each other. Additionally, the performance analysis system 100 can be mounted on a keyboard instrument 200.

[0037] The control device 11 consists of one or more processors that control the various elements of the performance analysis system 100. For example, the control device 11 consists of one or more processors such as CPU (Central Processing Unit), SPU (Sound Processing Unit), DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit).

[0038] Storage device 12 is a single or multiple memory devices that store the program executed by control device 11 and various data used by control device 11. Storage device 12 may be composed of known recording media such as magnetic recording media or semiconductor recording media, or a combination of multiple recording media. In addition, a removable recording medium that can be detached from the performance analysis system 100, or a recording medium that can be written to or read from by control device 11 via a communication network such as the Internet (e.g., cloud storage) may also be used as storage device 12.

[0039] The operating device 13 is an input device that receives instructions from the user. The operating device 13 may be, for example, a control that is operated by the user, or a touch panel that detects the user's touch. Furthermore, the operating device 13 (e.g., a mouse or keyboard), which is separate from the performance analysis system 100, can be connected to the performance analysis system 100 in a wired or wireless manner.

[0040] The display device 14 displays images based on the control given by the control device 11. For example, various display panels such as liquid crystal display panels or organic EL (electroluminescence) panels can be used as the display device 14. In addition, the display device 14, which is separate from the performance analysis system 100, can be connected to the performance analysis system 100 in a wired or wireless manner.

[0041] The shooting device 15 is an image input device that generates a time sequence of image data D1 by shooting a subject. The time sequence of image data D1 is animation data representing animation. For example, the shooting device 15 includes an optical system such as a shooting lens, an imaging element that receives incident light from the optical system, and a processing circuit that generates image data D1 corresponding to the amount of light received by the imaging element. Furthermore, the shooting device 15, which is separate from the performance analysis system 100, can be connected to the performance analysis system 100 via wired or wireless means.

[0042] The user adjusts the position or angle of the shooting device 15 relative to the keyboard instrument 200 to achieve the shooting conditions recommended by the provider of the performance analysis system 100. Specifically, the shooting device 15 is positioned above the keyboard instrument 200 to capture images of the keyboard 22 of the keyboard instrument 200 and the user's left and right hands. Therefore, as Figure 2 As illustrated, the time sequence of image data D1, representing the performance image G1 including an image of the keyboard 22 of the keyboard instrument 200 (hereinafter referred to as "keyboard image") g1 and images of the user's left and right hands (hereinafter referred to as "finger images") g2, is generated by the shooting device 15. That is, the animation data representing the animation of the user playing the keyboard instrument 200 is generated in parallel with the performance. Furthermore, the shooting conditions of the shooting device 15 are, for example, the shooting range or the shooting direction. The shooting range is the range (field of view) captured by the shooting device 15. The shooting direction is the direction of the shooting device 15 relative to the keyboard instrument 200.

[0043] Figure 3 This is a block diagram illustrating the functional structure of the performance analysis system 100. The control device 11 functions as the performance analysis unit 30 and the display control unit 40 by executing programs stored in the storage device 12. The performance analysis unit 30 generates finger movement data Q representing the user's finger movements by analyzing performance data P and image data D1. The finger movement data Q specifies which of the user's fingers operated on each of the multiple keys 21 of the keyboard instrument 200. Specifically, the finger movement data Q specifies the pitch n corresponding to the key 21 operated by the user and the finger number (hereinafter referred to as "finger number") k used by the user in operating that key 21. The pitch n is, for example, a note number in the MIDI standard. The finger number k is a number assigned to each finger of the user's left and right hands.

[0044] The display control unit 40 displays various images on the display device 14. For example, the display control unit 40 displays an image (hereinafter referred to as "resolution screen") 61 representing the resolution result of the performance resolution unit 30 on the display device 14. Figure 4This is a schematic diagram of the analysis screen 61. The analysis screen 61 is an image in which multiple note images 611 are arranged on a coordinate plane with a horizontal time axis and a vertical pitch axis. The note images 611 are displayed for each note played by the user. The position of the note image 611 in the direction of the pitch axis is set corresponding to the pitch n of the note represented by that note image 611. The position and total length of the note image 611 in the direction of the time axis are set corresponding to the duration of the sounding of the note represented by that note image 611.

[0045] Each note image 611 is configured with a label (hereinafter referred to as "index number") corresponding to the finger number k assigned to that note by the finger movement data Q. The character "L" in index number 612 represents the left hand, and the character "R" represents the right hand. Furthermore, the numbers in index number 612 represent individual fingers. Specifically, the number "1" in index number 612 represents the thumb, "2" represents the index finger, "3" represents the middle finger, "4" represents the ring finger, and "5" represents the little finger. Therefore, for example, index number 612 "R2" represents the index finger of the right hand, and index number 612 "L4" represents the ring finger of the left hand. The note image 611 and index number 612 are displayed in different ways (e.g., hue or grayscale) for the right and left hands. The display control unit 40 uses the finger movement data Q to... Figure 4 The resolution screen 61 is displayed on the display device 14.

[0046] Furthermore, for notes in the multiple note images 611 within the analysis screen 61 where the estimation result of finger number k is of low reliability, the note image 611 is displayed in a manner different from the usual note image 611 (e.g., a dashed frame), and a specific label representing an invalid estimation result of finger number k is displayed, for example, "?".

[0047] like Figure 3 As illustrated, the performance analysis unit 30 includes a finger position data generation unit 31 and a finger movement data generation unit 32. The finger position data generation unit 31 generates finger position data F by analyzing the performance image G1. The finger position data F represents the positions of the fingers of the user's left hand and right hand. As described above, in the first embodiment, the positions of the user's fingers are distinguished as left and right, thus enabling estimation of the finger movements distinguishing the user's left and right hands. On the other hand, the finger movement data generation unit 32 generates finger movement data Q using the performance data P and the finger position data F. The finger position data F and the finger movement data Q are generated for each unit period on the time axis. Each unit period is a period of a predetermined length (time frame).

[0048] A: Finger position data generation unit 31

[0049] The finger position data generation unit 31 includes an image extraction unit 311, a matrix generation unit 312, a finger position estimation unit 313, and a projective transformation unit 314.

[0050] [Finger position estimation section 313]

[0051] The finger position estimation unit 313 estimates the positions c[h,f] of each finger of the user's left and right hands by analyzing the performance image G1 represented by image data D1. The position c[h,f] of each finger is the position of the fingertip in the xy coordinate system of the performance image G1. The position c[h,f] is represented by the combination (x[h,f], y[h,f]) of the coordinates x[h,f] on the x-axis and y[h,f] on the y-axis of the xy coordinate system of the performance image G1. The positive direction of the x-axis corresponds to the right direction of the keyboard 22 (from low to high notes), and the negative direction of the x-axis corresponds to the left direction of the keyboard 22 (from high to low notes). The notation h is a variable representing either the left or right hand (h = 1, 2). Specifically, the value "1" of variable h represents the left hand, and the value "2" of variable h represents the right hand. The variable f is the number of each finger of the left and right hands (f = 1 to 5). The variable f has the following values: 1 for thumb, 2 for index finger, 3 for middle finger, 4 for ring finger, and 5 for little finger. Therefore, for example, Figure 2 The illustrated position c[1,2] is the position of the fingertip of the index finger (f=2) of the left hand (h=1), and the position c[2,4] is the position of the fingertip of the ring finger (f=4) of the right hand (h=2).

[0052] Figure 5 This is a flowchart illustrating the specific process of the finger position estimation unit 313 estimating the position of each of the user's fingers (hereinafter referred to as "finger position estimation processing"). The finger position estimation processing includes image analysis processing Sa1, left-right determination processing Sa2, and interpolation processing Sa3.

[0053] Image analysis processing Sa1 estimates the positions c[h,f] of each finger of the user's left and right hands (hereinafter referred to as "first hand") and the positions c[h,f] of each finger of the user's left and right hands (hereinafter referred to as "second hand") by analyzing the performance image G1. Specifically, the finger position estimation unit 313 estimates the positions c[h,1] to c[h,5] of each finger of the first hand and the positions c[h,1] to c[h,5] of each finger of the second hand by image recognition processing, which estimates the user's skeletal structure or joints by analyzing the image. For example, known image recognition processing such as MediaPipe or OpenPose is used for image analysis processing Sa1. Furthermore, if no fingertip is detected from the performance image G1, the coordinate x[h,f] of that fingertip on the x-axis is set to an invalid value such as "0".

[0054] In image analysis processing Sa1, the positions of the fingers of the user's first hand (c[h,1] to c[h,5]) and the positions of the fingers of the second hand (c[h,1] to c[h,5]) are estimated, but it is not possible to determine whether the first or second hand belongs to the user's left or right hand. Furthermore, in playing the keyboard instrument 200, the user's right and left wrists sometimes cross, so it is inappropriate to determine the left or right hand solely based on the coordinates x[h,f] of the positions estimated by image analysis processing Sa1. Alternatively, if the imaging device 15 captures images of the user's wrists and torso, the user's left or right hand can be estimated based on the coordinates of the user's shoulders and wrists, according to the performance image G1. However, this presents problems such as the need for the imaging device 15 to capture a wider area and an increased processing load on image analysis processing Sa1.

[0055] Considering the above, the finger position estimation unit 313 of the first embodiment performs a determination of whether the first hand or the second hand belongs to the user's left or right hand. Figure 5 The left and right determination process Sa2. That is, the finger position estimation unit 313 determines the variable h of the finger position c[h,f] of the first hand and the second hand to be either the value "1" representing the left hand and the value "2" representing the right hand.

[0056] When playing the keyboard instrument 200, since the nails of both the left and right hands are positioned above the plumb line, the performance image G1 captured by the imaging device 15 includes images of the nails of both the user's left and right hands. Therefore, in the left hand within the performance image G1, the position of the thumb c[h,1] is to the right compared to the position of the little finger c[h,5], and in the right hand within the performance image G1, the position of the thumb c[h,1] is to the left compared to the position of the little finger c[h,5]. Considering the above, in the left / right determination processing Sa2, the finger position estimation unit 313 determines the hand in the first and second hands where the position of the thumb c[h,1] is to the right (positive x-axis direction) compared to the position of the little finger c[h,5] as the left hand (h=1). On the other hand, the finger position estimation unit 313 determines the hand in the first and second hands where the position of the thumb c[h,1] is to the left (negative x-axis direction) compared to the position of the little finger c[h,5] as the right hand.

[0057] Figure 6 This is a flowchart illustrating the specific process of the left / right determination process Sa2. The finger position estimation unit 313 calculates the determination index γ[h] for the first hand and the second hand respectively (Sa21). The determination index γ[h] is calculated, for example, by the following formula (1).

[0058] [Formula 1]

[0059]

[0060] The notation μ[h] in formula (1) is the average (for example, a simple average) of the coordinates x[h,1] to x[h,5] of the five fingers of the first and second hands respectively. As understood from formula (1), when the coordinate x[h,f] decreases from the thumb to the little finger (left hand), the determination index γ[h] is negative, and when the coordinate x[h,f] increases from the thumb to the little finger (right hand), the determination index γ[h] is positive. Therefore, the finger position estimation unit 313 determines the hand with a negative determination index γ[h] among the first and second hands as the left hand and sets the variable h to the value "1" (Sa22). In addition, the finger position estimation unit 313 determines the hand with a positive determination index γ[h] among the first and second hands as the right hand and sets the variable h to the value "2" (Sa23). Based on the left-right determination process Sa2 described above, the user's finger positions c[h,f] can be distinguished as right hand and left hand by using a simple processing of the relationship between the thumb and little finger positions.

[0061] Image analysis processing Sa1 and left / right determination processing Sa2 are used to estimate the position c[h,f] of each finger of the user for each unit period. However, sometimes the position c[h,f] cannot be properly estimated due to various factors such as noise present in the performance image G1. Therefore, when the position c[h,f] is missing in a specific unit period (hereinafter referred to as "missing period"), the finger position estimation unit 313 calculates the position c[h,f] of the missing period by using the interpolation processing Sa3 of the position c[h,f] of the unit period before and after the missing period. For example, when the position c[h,f] is missing in the central unit period (missing period) among three consecutive unit periods on the time axis, the average of the position c[h,f] of the unit period before and after the missing period is calculated as the position c[h,f] of the missing period.

[0062] [Image Extraction Unit 311]

[0063] As mentioned above, the performance image G1 includes a keyboard image g1 and a finger image g2. Figure 3 Image extraction unit 311, such as Figure 7 As illustrated, a specific region (hereinafter referred to as the "specific region") B within the performance image G1 is extracted. The specific region B is the region within the performance image G1 that contains the keyboard image g1 and the finger image g2. The finger image g2 corresponds to an image of at least a portion of the user's body.

[0064] Figure 8 This is a flowchart illustrating the specific process of the image extraction unit 311 extracting a specific region B from the performance image G1 (hereinafter referred to as "image extraction processing"). The image extraction processing includes region estimation processing Sb1 and region extraction processing Sb2.

[0065] Region estimation processing Sb1 is a process that estimates a specific region B for the performance image G1 represented by image data D1. Specifically, the image extraction unit 311 generates an image processing mask M representing the specific region B based on the image data D1 through region estimation processing Sb1. The image processing mask M is as follows: Figure 7 As illustrated, the mask is of the same size as the performance image G1 and consists of multiple elements corresponding to different pixels of the performance image G1. Specifically, the image processing mask M is a binary mask in which elements within a region corresponding to a specific region B of the performance image G1 are set to the value "1", and elements outside the specific region B are set to the value "0". The control device 11 performs region estimation processing Sb1, thereby estimating the elements (region estimation unit) for a specific region B of the performance image G1.

[0066] like Figure 3 As illustrated, the estimation model 51 is used in the generation of the image processing mask M by the image extraction unit 311. That is, the image extraction unit 311 generates the image processing mask M by inputting the image data D1 representing the performance image G1 into the estimation model 51. The estimation model 51 is a statistical model that has learned the relationship between the image data D1 and the image processing mask M through machine learning. The estimation model 51 is, for example, composed of a deep neural network (DNN). For example, any form of deep neural network such as a convolutional neural network (CNN) or a recurrent neural network (RNN) can be used as the estimation model 51. The estimation model 51 can also be composed of a combination of multiple deep neural networks. In addition, additional elements such as long short-term memory (LSTM) can be incorporated into the estimation model 51.

[0067] Figure 9 This is an illustration of machine learning used to create the inference model 51. For example, the inference model 51 is created through machine learning performed by a machine learning system 900 separate from the performance analysis system 100, and this inference model 51 is provided to the performance analysis system 100. The machine learning system 900 is a server system that can communicate with the performance analysis system 100 via a communication network such as the Internet. The inference model 51 is sent from the machine learning system 900 to the performance analysis system 100 via the communication network.

[0068] In the machine learning of the estimation model 51, multiple learning data T are utilized. Each of the multiple learning data T is composed of a combination of image data Dt and an image processing mask Mt used for learning. Image data Dt represents a known image including a keyboard image g1 of a keyboard instrument and an image of the surrounding area of ​​the keyboard instrument. The model of the keyboard instrument and the shooting conditions (e.g., shooting range and shooting direction) are different for each image data Dt. That is, multiple keyboard instruments are each photographed under different shooting conditions, thereby preparing image data Dt in advance. In addition, image data Dt can be prepared using known image compositing techniques. The image processing mask Mt of each learning data T is a mask representing a specific region B in the known image represented by the image data Dt of that learning data T. Specifically, elements in the region corresponding to the specific region B in the image processing mask Mt are set to the value "1", and elements in the region outside the specific region B are set to the value "0". That is, the image processing mask Mt represents the correct solution that the estimation model 51 should output based on the input of image data Dt.

[0069] The machine learning system 900 calculates an error function that represents the error between the image processing mask M output by the initial or temporary model (hereinafter referred to as the "temporary model") 51a and the image processing mask M of the learning data T when image data Dt is input for each learning data T. Furthermore, the machine learning system 900 updates multiple variables of the temporary model 51a to reduce the error function. The temporary model 51a, which has undergone the above processing repeatedly for each of the multiple learning data T, is determined as the inference model 51. Therefore, the inference model 51 outputs a statistically reasonable image processing mask M for the unknown image data D1 based on the potential relationship between the image data Dt and the image processing mask Mt of the multiple learning data T. That is, the inference model 51 is a trained model that has learned the relationship between the image data Dt and the image processing mask Mt.

[0070] As described above, in the first embodiment, an image processing mask M representing a specific region B is generated by inputting image data D1 of the performance image G1 into the inference model 51 that has completed machine learning. Therefore, it is possible to determine the specific region B with high accuracy for various unknown performance images G1.

[0071] Figure 8 The region extraction processing Sb2 is a process that extracts a specific region B from the performance image G1 represented by image data D1. Specifically, region extraction processing Sb2 is an image processing method that selectively removes regions other than the specific region from the performance image G1, thereby relatively emphasizing the specific region B. The image extraction unit 311 of the first embodiment generates image data D2 by applying an image processing mask M to image data D1 (performance image G1). Specifically, the image extraction unit 311 multiplies the pixel value of each pixel in the performance image G1 by the element corresponding to that pixel in the image processing mask M. Through region extraction processing Sb2, such as Figure 7 As illustrated, image data D2 is generated representing an image (hereinafter referred to as "performance image G2") of the performance image G1 after removing the area other than the specific region B. That is, the performance image G2 represented by image data D2 is an image in which the keyboard image g1 and the finger image g2 have been extracted from the performance image G1. The element (region extraction unit) for extracting the specific region B of the performance image G1 is realized by the region extraction process Sb2 performed by the control device 11.

[0072] [Projective Transformation Section 314]

[0073] The finger positions c[h,f] estimated by the finger position estimation process are coordinates set in the xy coordinate system of the performance image G1. The shooting conditions of the keyboard instrument 200 set by the shooting device 15 can vary depending on various situations such as the usage environment of the keyboard instrument 200. For example, imagine... Figure 2 The examples illustrate situations where the shooting range is too wide (or too narrow) compared to the ideal shooting conditions, or where the shooting direction is tilted relative to the plumb line. The values ​​of the coordinates x[h,f] and y[h,f] of each position c[h,f] depend on the shooting conditions of the performance image G1 set by the shooting device 15. Therefore, the image transformation unit 314 of the first embodiment transforms the position c[h,f] of each finger related to the performance image G1 into a position C[h,f] in the XY coordinate system that is substantially independent of the shooting conditions set by the shooting device 15. The finger position data F generated by the finger position data generation unit 31 is data representing the position C[h,f] transformed by the image transformation unit 314. That is, the finger position data F specifies the positions C[1,1] to C[1,5] of each finger of the user's left hand and the positions C[2,1] to C[2,5] of each finger of the user's right hand.

[0074] XY coordinate system as follows Figure 10 As illustrated, a specified image (hereinafter referred to as "reference image") Gref is defined. The reference image Gref is an image taken under standard shooting conditions of the keyboard of a standard keyboard instrument (hereinafter referred to as "reference instrument"). Furthermore, the reference image Gref is not limited to an image taken of an actual keyboard. For example, an image synthesized using known image compositing techniques can be used as the reference image Gref. Image data (hereinafter referred to as "reference data") Dref representing the reference image Gref and auxiliary data A associated with the reference image Gref are stored in storage device 12.

[0075] Auxiliary data A is data that specifies the combination of the area (hereinafter referred to as the "unit area") Rn where each key 21 of the reference instrument is located in the reference image Gref and the pitch n corresponding to that key 21. That is, auxiliary data A can also be referred to as data that defines the unit area Rn corresponding to each pitch n in the reference image Gref.

[0076] The transformation from position c[h,f] in the xy coordinate system to position C[h,f] in the XY coordinate system is performed using a projective transformation, as shown in the following equation (2), which utilizes a transformation matrix W. In equation (2), the notation X represents the coordinate on the X-axis of the XY coordinate system, and the notation Y represents the coordinate on the Y-axis. Additionally, the notation s is an adjustment value used to integrate the scale between the xy and XY coordinate systems.

[0077] [Formula 2]

[0078]

[0079] [Matrix Generation Unit 312]

[0080] Figure 3 The matrix generation unit 312 generates the transformation matrix W of the formula (2) applied to the projective transformation by the projective transformation unit 314. Figure 11 This is a flowchart illustrating the specific process of generating a transformation matrix W by the matrix generation unit 312 (hereinafter referred to as "matrix generation process"). The matrix generation process of the first embodiment is performed on the performance image G2 (image data D2) after being processed by the image extraction process. Based on the above structure, compared with the structure that performs matrix generation process on the entire performance image G1, which also includes the region other than the specific region B, it is possible to generate an appropriate transformation matrix W that approximates the keyboard image g1 with high accuracy to the reference image Gref.

[0081] The matrix generation process includes initialization process Sc1 and matrix update process Sc2. Initialization process Sc1 sets the initial value of the transformation matrix W, i.e., the initial matrix W0. Details of initialization process Sc1 will be described later.

[0082] The matrix update process Sc2 is a process of generating the transformation matrix W by repeatedly updating the initial matrix W0. That is, the projective transformation unit 314 repeatedly updates the initial matrix W0 in such a way that the keyboard image g1 of the playing image G2 becomes closer to the reference image Gref by utilizing the projective transformation of the transformation matrix W, thereby generating the transformation matrix W. For example, the transformation matrix W is generated in such a way that the X-axis coordinate X / s of a specific location in the reference image Gref is approximately or identical to the x-axis coordinate x of the corresponding location in the keyboard image g1, and the Y-axis coordinate Y / s of a specific location in the reference image Gref is approximately or identical to the y-axis coordinate y of the corresponding location in the keyboard image g1. That is, the transformation matrix W is generated in such a way that the coordinates of the key 21 in the keyboard image g1 corresponding to a specific pitch are transformed by applying the projective transformation of the transformation matrix W to the coordinates of the key 21 in the reference image Gref corresponding to that pitch. By executing the matrix update process Sc2 illustrated above by the control device 11, the elements (matrix generation unit 312) for generating the transformation matrix W are realized.

[0083] However, as a matrix update process Sc2, it is envisioned to update the transformation matrix W in a way that makes image features, such as SIFT (Scale-Invariant Feature Transform), close to those between the reference image Gref and the keyboard image g1. However, in the keyboard image g1, the pattern of multiple keys 21 arranged in the same way is repeated, so it may not be possible to properly estimate the transformation matrix W by utilizing image features.

[0084] Considering the above, the matrix generation unit 312 of the first embodiment repeatedly updates the initial matrix W0 in the matrix update processing Sc2 to increase (ideally maximize) the enhanced correlation coefficient (ECC) between the reference image Gref and the keyboard image g1. According to this method, compared to the aforementioned method utilizing image features, it is possible to generate an appropriate transformation matrix W that approximates the keyboard image g1 and the reference image Gref with high accuracy. The generation of a transformation matrix W utilizing the enhanced correlation coefficient is also disclosed in Georgios D. Evangelidis and Emmanouil Z. Psarakis, "Parametric Image Alignment Using Enhanced Correlation Coefficient Maximization", IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, VOL.30, NO.10, October 2008. Furthermore, as mentioned above, it is appropriate to enhance the correlation coefficient for generating the transformation matrix W used in the transformation of the keyboard image g1, but the transformation matrix W can also be generated in a way that approximates the reference image Gref and the keyboard image g1 using image features such as SIFT mentioned above.

[0085] Figure 3 The projective transformation unit 314 performs projective transformation processing. The projective transformation processing utilizes the projective transformation of the performance image G1, which is generated by the matrix generation process. Through the projective transformation processing, the performance image G1 is transformed into an image captured under the same shooting conditions as the reference image Gref (hereinafter referred to as the "transformed image"). For example, the area in the transformed image corresponding to the key 21 of pitch n is substantially the same as the unit area Rn of pitch n in the reference image Gref. In addition, the xy coordinate system of the transformed image is substantially the same as the xy coordinate system of the reference image Gref. In the projective transformation processing described above, the projective transformation unit 314 transforms the position c[h,f] of each finger into the position C[h,f] of the xy coordinate system, as expressed by the aforementioned formula (2). By executing the projective transformation processing exemplified above by the control device 11, the elements for performing the projective transformation of the performance image G1 (projective transformation unit 314) are realized.

[0086] The display control unit 40 causes the display device 14 to display the transformed image generated by the projective transformation process. For example, the display control unit 40 causes the display device 14 to display the transformed image and the reference image Gref in a state where they overlap. As described above, the area in the transformed image corresponding to the key 21 of each pitch n overlaps with the unit area Rn in the reference image Gref corresponding to that pitch n.

[0087] As described above, in the first embodiment, a transformation matrix W is generated such that the keyboard image g1 of the performance image G1 approximates the reference image Gref, and a projective transformation process utilizing the transformation matrix W is performed on the performance image G1. Therefore, the performance image G1 of the keyboard instrument 200 played by the user can be transformed into a transformed image corresponding to the shooting conditions of the reference instrument of the reference image Gref.

[0088] Figure 12 This is a flowchart illustrating the specific process of the initial setup process Sc1. If the initial setup process Sc1 begins, the projective transformation unit 314 will... Figure 13 The setup screen 62 shown is displayed on the display device 14 (Sc11). The setup screen 62 includes a performance image G1 captured by the imaging device 15 and instructions 622 for the user. Instructions 622 indicate that a region (hereinafter referred to as the "target region") 621 within the keyboard image g1 in the performance image G1 that corresponds to one or more specific pitches (hereinafter referred to as "target pitches") n is selected. The user visually confirms the setup screen 62 while operating the operation device 13, thereby selecting the target region 621 in the performance image G1 that corresponds to the target pitch n. The projection conversion unit 314 accepts the user's selection of the target region 621 (Sc12).

[0089] The projective transformation unit 314 determines one or more unit regions Rn specified for the target pitch n in the reference image Gref, represented by the reference data Dref (Sc13). Furthermore, the projective transformation unit 314 calculates a matrix for projectively transforming the target region 621 of the performance image G1 into one or more unit regions Rn determined from the reference image Gref, and uses this matrix as the initial matrix W0 (Sc14). As understood from the above description, the initial setting process Sc1 of the first embodiment sets the initial matrix W0 by using the projective transformation of the initial matrix W0 to make the target region 621 in the keyboard image g1, as indicated by the user, approximate the unit region Rn in the reference image Gref corresponding to the target pitch n.

[0090] Setting the initial matrix W0 is important for generating an appropriate transformation matrix W through matrix update processing Sc2. In particular, when using the enhanced correlation coefficient method for matrix update processing Sc2, there is a tendency for the appropriateness of the initial matrix W0 to easily affect the appropriateness of the final transformation matrix W. In the first embodiment, the initial matrix W0 is set such that the target region 621 in the performance image G1 corresponding to the user's instruction is close to the unit region Rn in the reference image Gref corresponding to the target pitch n. Therefore, an appropriate transformation matrix W can be generated that makes the keyboard image g1 and the reference image Gref highly approximate. Furthermore, in the first embodiment, the region in the performance image G1 specified by the user's operation on the operating device 13 is used as the target region 621 for setting the initial matrix W0. Therefore, compared to, for example, estimating the region in the performance image G1 corresponding to the target pitch n through computational processing, the processing load can be reduced and an appropriate initial matrix W0 can be generated. Furthermore, in the above description, the initial setting process Sc1 was performed on the performance image G1 as the object, but the initial setting process Sc1 can also be performed on the performance image G2.

[0091] B: Operations Data Generation Unit 32

[0092] Figure 3 As described above, the finger movement data generation unit 32 generates finger movement data Q using the performance data P generated by the keyboard instrument 200 and the finger position data F generated by the finger position data generation unit 31. The generation of finger movement data Q is performed for each unit period. The finger movement data generation unit 32 of the first embodiment has a probability calculation unit 321 and a finger movement estimation unit 322. Furthermore, in the above description, one finger of the user is represented by a combination of variables h and f, but in the following description, one finger of the user is represented by a finger number k (k = 1 to 10). Therefore, the position C[h,f] specified by the finger position data F for each finger is marked as position C[k] in the following description.

[0093] [Probability Calculation Section 321]

[0094] The probability calculation unit 321 calculates the probability p of playing the pitch n specified by the performance data P with each finger number k. The probability p is an index (likelihood) of the accuracy (likelihood) of the finger number k operating the key 21 of pitch n. The probability calculation unit 321 calculates the probability p in relation to whether the position C[k] of the finger number k exists within the unit area Rn of pitch n. The probability p is calculated for each unit period on the time axis. Specifically, when the performance data P specifies the pitch n, the probability calculation unit 321 calculates the probability p (C[k]|ηk=n) by performing the following equation (3).

[0095] [Formula 3]

[0096]

[0097] The condition "ηk = n" in the probability p(C[k]|ηk=n) represents the condition that the finger with finger number k is playing the pitch n. That is, the probability p(C[k]|ηk=n) represents the probability of observing position C[k] for the finger given that the finger with finger number k is playing the pitch n.

[0098] The notation I (C[k]∈Rn) in equation (3) is an indicator function that sets the value to "1" when the position C[k] exists within the unit region Rn, and sets the value to "0" when the position C[k] exists outside the unit region Rn. The notation |Rn| represents the area of ​​the unit region Rn. Additionally, the notation ν(0,σ) 2 E) represents the observation noise, consisting of the mean σ and the variance σ. 2 The normal distribution representation. The notation E is a 2x2 identity matrix. The notation * represents the observation noise ν(0,σ). 2 Convolution of E).

[0099] As understood from the above explanation, the probability p(C[k]|ηk=n) calculated by the probability calculation unit 321 is the accuracy of the position C[k] specified by the finger position data F for the finger, given that the finger with finger number k plays the pitch n specified by the performance data P. Therefore, the probability p(C[k]|ηk=n) is maximized when the position C[k] of the finger with finger number k is within the unit region Rn of the performance state, and decreases as the position C[k] moves further away from the unit region Rn.

[0100] On the other hand, when the performance data P does not specify any pitch n, that is, when the user does not operate any of the N keys 21, the probability calculation unit 321 calculates the probability p (C[k]|ηk=0) of each finger using the following formula (4).

[0101] [Formula 4]

[0102]

[0103] The notation |R| in equation (4) represents the total area of ​​the N unit regions R1 to RN of the reference image Gref. As understood from equation (4), when the user does not operate any key 21, the probability p(C[k]|ηk=0) is set to a common value (1 / |R|) for all finger numbers k.

[0104] As described above, during the period when the performance data P specifies pitch n, the multiple probabilities p(C[k]|ηk=n) corresponding to different fingers are calculated for each unit period on the time axis. On the other hand, during each unit period when the performance data P does not specify pitch n, the multiple probabilities p(C[k]|ηk=0) corresponding to different fingers are set to a sufficiently small fixed value (1 / |R|).

[0105] [Money Index Prediction Section 322]

[0106] The finger movement estimation unit 322 estimates the user's finger movements. Specifically, the finger movement estimation unit 322 estimates the finger (finger number k) that will play the pitch n specified by the performance data P based on the probability p (C[k]|ηk=n) of each finger. The estimation of the finger number k by the finger movement estimation unit 322 (generation of finger movement data Q) is performed for each calculation of the probability p (C[k]|ηk=n) of each finger (i.e., each unit period). Specifically, the finger movement estimation unit 322 determines the finger number k corresponding to the maximum value among the multiple probabilities p (C[k]|ηk=n) corresponding to different fingers. Moreover, the finger movement estimation unit 322 generates finger movement data Q that specifies the pitch n specified by the performance data P and the finger number k determined based on the probability p (C[k]|ηk=n).

[0107] Furthermore, during the period when the performance data P specifies the pitch n, if the maximum value among multiple probabilities p(C[k]|ηk=n) is less than a predetermined threshold, it indicates that the reliability of the estimated finger movement result is low. Therefore, the finger movement estimation unit 322 sets the finger number k to an invalid value representing an invalid estimation result during the unit period in which the maximum value of multiple probabilities p(C[k]|ηk=n) is less than the threshold. For notes where the finger number k is set to an invalid value, the display control unit 40 displays the following: Figure 4As illustrated, the note image 611 is displayed in a manner different from the usual note image 611, and the label “??” representing the invalid estimation result of the finger number k is displayed. The structure and operation of the fingering data generation unit 32 are as described above.

[0108] Figure 14 This is a flowchart illustrating the specific process of the processing (hereinafter referred to as "performance analysis processing") performed by the performance analysis unit 30. For example, the performance analysis processing is initiated when an instruction from the user to the operating device 13 is received.

[0109] If the performance analysis process begins, the control device 11 (image extraction unit 311) executes... Figure 8 Image extraction processing (S11). That is, the control device 11 generates a performance image G2 by extracting a specific region B in the performance image G1, which includes the keyboard image g1 and the finger image g2. The image extraction processing, as described above, includes region estimation processing Sb1 and region extraction processing Sb2.

[0110] If image extraction processing is performed, the control device 11 (matrix generation unit 312) executes... Figure 11 The matrix generation process (S12) involves the control device 11 repeatedly updating the initial matrix W0 to increase the enhanced correlation coefficient between the reference image Gref and the keyboard image g1, thereby generating the transformation matrix W. As described above, the matrix generation process includes an initial setting process Sc1 and a matrix update process Sc2.

[0111] If a transformation matrix W is generated, the control device 11 repeatedly performs the following illustrated processing (S13 to S18) for each unit period. First, the control device 11 (finger position estimation unit 313) executes... Figure 5 The finger position estimation process (S13) involves the control device 11 estimating the positions c[h,f] of each finger of the user's left and right hands by analyzing the performance image G1. As described above, the finger position estimation process includes image analysis processing Sa1, left / right determination processing Sa2, and interpolation processing Sa3.

[0112] The control device 11 (projective transformation unit 314) performs projective transformation processing (S14). That is, the control device 11 generates a transformed image by projective transformation of the performance image G1 using the transformation matrix W. In the projective transformation processing, the control device 11 transforms the position c[h,f] of each finger of the user into the position C[h,f] in the XY coordinate system, and generates finger position data F representing the position C[h,f] of each finger.

[0113] If finger position data F is generated through the above processing, the control device 11 (probability calculation unit 321) performs probability calculation processing (S15). That is, the control device 11 calculates the probability p(C[k]|ηk=n) that the pitch n specified by the playing data P is played by the finger with finger number k. Moreover, the control device 11 (finger movement estimation unit 322) performs finger movement estimation processing (S16). That is, the control device 11 estimates the finger number k of the finger that played the pitch n based on the probability p(C[k]|ηk=n) of each finger, and generates finger movement data Q that specifies the pitch n and the finger number k.

[0114] If finger movement data Q is generated through the above processing, the control device 11 (display control unit 40) updates the analysis screen 61 accordingly with the finger movement data Q (S17). Furthermore, the control device 11 determines whether a predetermined end condition is met (S18). For example, if the user indicates the end of the performance analysis process through operation of the operating device 13, the control device 11 determines that the end condition is met. If the end condition is not met (S18: NO), the control device 11 repeatedly performs the processing after the finger position estimation processing for the next unit period (S13-S18). On the other hand, if the end condition is met (S18: YES), the control device 11 ends the performance analysis process.

[0115] As explained above, in the first embodiment, finger movement data Q is generated using finger position data F generated by analyzing the performance image G1 and performance data P representing the user's performance. Therefore, finger movement can be estimated with higher accuracy compared to a structure that estimates finger movement solely based on performance data P.

[0116] Furthermore, in the first embodiment, the transformation matrix W, which is used to make the keyboard image g1 approximate the projective transformation of the reference image Gref, is used to transform the positions c[h,f] of each finger estimated by the finger position estimation process. That is, the positions C[h,f] of each finger are estimated with reference to the reference image Gref as a reference. Therefore, compared with the structure that does not transform the positions c[h,f] of each finger to positions with reference to the reference image Gref, finger movements can be estimated with high accuracy.

[0117] In the first embodiment, a specific region B containing the keyboard image g1 is extracted from the performance image G1. Therefore, as described above, a suitable transformation matrix W can be generated that allows the keyboard image g1 to approximate the reference image Gref with high accuracy. Furthermore, the convenience of the performance image G1 can be improved by extracting the specific region B. In particular, in the first embodiment, a specific region B containing the keyboard image g1 and the finger image g2 is extracted from the performance image G1. Therefore, a performance image G2 can be generated that allows for effective visual confirmation of the condition of the keyboard 22 of the keyboard instrument 200 and the condition of the user's fingers.

[0118] 2: Second Implementation Method

[0119] The second embodiment will be described. Furthermore, in the embodiments illustrated below, for elements that have the same function as those in the first embodiment, detailed descriptions are appropriately omitted using the same reference numerals as those used in the description of the first embodiment.

[0120] In the first embodiment, the probability p(C[k]|ηk=n) corresponding to whether the position C[k] of the finger with finger number k exists within a unit region Rn of pitch n is calculated. If it is assumed that only one finger exists within the unit region Rn, the finger movement can be estimated with high accuracy in the first embodiment. However, it is conceivable that in the actual performance of the keyboard instrument 200, there are multiple finger positions C[k] within a unit region Rn.

[0121] For example, such as Figure 15 As illustrated, when the user operates one key 21 with the middle finger of their left hand, and then moves the index finger of the left hand upwards in the direction of the plumb bob, the middle and index fingers of the left hand overlap in the performance image G1. That is, the positions C[k] of the middle finger and C[k] of the index finger exist within one unit area Rn. Furthermore, in performance methods where the user operates one key 21 with one finger and then passes other fingers above or below that finger (finger crossing), multiple fingers sometimes overlap. As described above, when multiple fingers overlap within one unit area Rn, the method of the first embodiment may not be able to accurately estimate finger movements. The second embodiment is a method to solve the above problems. Specifically, in the second embodiment, the positional relationships of multiple fingers and the temporal variations (fluctuations) of each finger's position are added to the estimation of finger movements.

[0122] Figure 16 This is a block diagram illustrating the functional structure of the performance analysis system 100 according to the second embodiment. The performance analysis system 100 of the second embodiment is a structure that adds a control data generation unit 323 to the same elements as the first embodiment.

[0123] The control data generation unit 323 generates N control data Z[1]~Z[N] corresponding to different pitches n. Figure 17 This is a schematic diagram of the control data Z[n] corresponding to any one pitch n. The control data Z[n] is a vector data representing the characteristics of the relative position (hereinafter referred to as "relative position") C'[k] of each finger relative to a unit area Rn of pitch n. The relative position C'[k] is information that transforms the position C[k] represented by the finger position data F into the relative position relative to the unit area Rn.

[0124] The control data Z[n] corresponding to a pitch n includes not only the pitch n, but also, for each of the multiple fingers, the positional average Za[n,k], positional variance Zb[n,k], velocity average Zc[n,k], and velocity variance Zd[n,k]. The positional average Za[n,k] is the average of the relative positions C'[k] over a specified length of time (hereinafter referred to as the "observation period") including the current unit period. The observation period is, for example, equivalent to a period of multiple unit periods arranged in front of the current unit period on a time axis. The positional variance Zb[n,k] is the variance of the relative positions C'[k] within the observation period. The velocity average Zc[n,k] is the average of the rate (i.e., the rate of change) at which the relative positions C'[k] change within the observation period. The velocity variance Zd[n,k] is the variance of the rate at which the relative positions C'[k] change within the observation period.

[0125] As described above, the control data Z[n] contains information related to the relative position C'[k] for each of the multiple fingers (Za[n,k], Zb[n,k], Zc[n,k], Zd[n,k]). Therefore, the control data Z[n] reflects the positional relationship of the user's multiple fingers. Furthermore, the control data Z[n] contains information related to the change in the relative position C'[k] for each of the multiple fingers (Zb[n,k], Zd[n,k]). Therefore, the control data Z[n] reflects the temporal change in the position of each finger.

[0126] In the probability calculation process performed by the probability calculation unit 321 in the second embodiment, multiple estimation models 52[k] (52[1] to 52

[10] ) prepared in advance for different fingers are used. The estimation model 52[k] for each finger is a trained model that has learned the relationship between control data Z[n] and the probability p[k] associated with that finger. The probability p[k] is an index (probability) of the accuracy with which the finger with finger number k plays the pitch n specified by the playing data P. The probability calculation unit 321 calculates the probability p[k] for each of the multiple fingers by inputting N control data Z[1] to Z[N] into the estimation model 52[k] of that finger.

[0127] The presumed model 52[k] corresponding to any one finger number k is a logit regression model expressed by the following formula (5).

[0128] [Formula 5]

[0129]

[0130] The variable β in equation (5) k and variable ω k,n The parameters are set through machine learning performed by the machine learning system 900. Specifically, each estimation model 52[k] is created through machine learning performed by the machine learning system 900, and each estimation model 52[k] is provided to the performance analysis system 100. For example, the variable β of each estimation model 52[k]... k and variable ω k,n It was sent from the machine learning system 900 to the performance analysis system 100.

[0131] Fingers positioned above, or moving above or below, a finger in a key-pressing state tend to move more easily than a finger in a key-pressing state. Taking this tendency into account, a estimation model 52[k] is designed to learn the relationship between control data Z[n] and probability p[k] in a way that the probability p[k] is small for fingers with a high rate of change in relative position C'[k]. The probability calculation unit 321 calculates multiple probabilities p[k] associated with different fingers for each unit period by inputting control data Z[n] into each of the multiple estimation models 52[k].

[0132] The finger movement estimation unit 322 estimates the user's finger movements by applying multiple probabilities p[k]. Specifically, the finger movement estimation unit 322 estimates the finger (finger number k) that will play the pitch n specified by the performance data P based on the probability p[k] of each finger. The estimation of the finger number k by the finger movement estimation unit 322 (generation of finger movement data Q) is performed for each calculation of the probability p[k] of each finger (i.e., each unit period). Specifically, the finger movement estimation unit 322 determines the finger number k corresponding to the maximum value among the multiple probabilities p[k] corresponding to different fingers. Moreover, the finger movement estimation unit 322 generates finger movement data Q, which specifies the pitch n specified by the performance data P and the finger number k determined according to the probability p[k].

[0133] Figure 18 This is a flowchart illustrating the specific process of the performance analysis processing in the second embodiment. In the performance analysis processing of the second embodiment, the generation of control data Z[n] is added to the same processing as in the first embodiment (S19). Specifically, the control device 11 (control data generation unit 323) generates N control data Z[1] to Z[N] corresponding to different pitches n based on the finger position data F (i.e., the position C[h,f] of each finger) generated by the finger position data generation unit 31.

[0134] The control device 11 (probability calculation unit 321) calculates the probability p[k] corresponding to the finger number k by performing probability calculation processing on N control data Z[1] to Z[N] input to each estimation model 52[k] (S15). In addition, the control device 11 (finger movement estimation unit 322) estimates the user's finger movement by applying multiple probabilities p[k] in finger movement estimation processing (S16). The operation of elements other than the finger movement data generation unit 32 (S11 to S14, S17 to S18) is the same as in the first embodiment.

[0135] In the second embodiment, the same effect as in the first embodiment is achieved. Furthermore, in the second embodiment, the control data Z[k] input to the estimation model 52[k] includes the average Za[n,k] and variance Zb[n,k] of the relative positions C'[k] of each finger, and the average Zc[n,k] and variance Zd[n,k] of the rate of change of the relative positions C'[k]. Therefore, even when multiple fingers overlap due to factors such as finger crossing, the user's finger movements can be estimated with high accuracy.

[0136] Furthermore, while the logit regression model was exemplified as the presumptive model 52[k] in the above description, the types of presumptive models 52[k] are not limited to those exemplified above. For example, statistical models such as multilayer perceptrons can also be used as presumptive models 52[k]. In addition, deep neural networks such as convolutional neural networks or recurrent neural networks can be used as presumptive models 52[k]. Combinations of multiple statistical models can also be used as presumptive models 52[k]. The various presumptive models 52[k] exemplified above can be generally represented as well-trained models that have learned the relationship between control data Z[n] and probability p[k].

[0137] 3: Third Implementation Method

[0138] Figure 19 This is a flowchart illustrating the specific process of the performance analysis processing in the third embodiment. If image extraction processing and matrix generation processing are performed, the control device 11 determines whether a user has played the keyboard instrument 200 by referring to the performance data P (S21). Specifically, the control device 11 determines whether any one of the multiple keys 21 of the keyboard instrument 200 has been operated.

[0139] When the keyboard instrument 200 is played (S21: YES), the control device 11 performs the generation of finger position data F (S13-S14), the generation of finger movement data Q (S15-S16), and the updating of the resolution screen 61 (S17), in the same manner as in the first embodiment. On the other hand, when the keyboard instrument 200 is not played (S21: NO), the control device 11 proceeds to step S18. That is, the generation of finger position data F (S13-S14), the generation of finger movement data Q (S15-S16), and the updating of the resolution screen 61 (S17) are not performed.

[0140] The third embodiment achieves the same effect as the first embodiment. Furthermore, in the third embodiment, the generation of finger position data F and finger movement data Q is stopped when the keyboard instrument 200 is not being played. Therefore, compared to a structure that continues to generate finger position data F regardless of whether the keyboard instrument 200 is being played, the processing load required to generate finger movement data Q can be reduced. Moreover, the third embodiment is also applicable to the second embodiment.

[0141] 4: Fourth Implementation Method

[0142] The fourth embodiment is a modification of the initial setting process Sc1 of the aforementioned methods. Figure 20 This is a flowchart illustrating the specific process of the initial setting process Sc1 performed by the control device 11 (matrix generation unit 312) of the fourth embodiment.

[0143] If the initial setup process Sc1 begins, the user operates the key 21 of the multiple keys 21 of the keyboard instrument 200 corresponding to the desired pitch (hereinafter referred to as "specific pitch") n using a specific finger (hereinafter referred to as "specific finger"). The specific finger is the finger (e.g., the index finger of the right hand) that is notified to the user, for example, through the display on the display device 14 or the instruction manual of the keyboard instrument 200. The result of the user's performance, the performance data P specifying the specific pitch n, is supplied from the keyboard instrument 200 to the performance analysis system 100. The control device 11 recognizes the user's performance of the specific pitch n by obtaining the performance data P from the keyboard instrument 200 (Sc15). The control device 11 determines the unit area Rn corresponding to the specific pitch n among the N unit areas R1 to RN of the reference image Gref (Sc16).

[0144] On the other hand, the finger position data generation unit 31 generates finger position data F through finger position estimation processing. The finger position data F includes the position C[h,f] of a specific finger used by the user in playing a specific pitch n. The control device 11 determines the position C[h,f] of the specific finger by acquiring the finger position data F (Sc17).

[0145] The control device 11 sets the initial matrix W0 using the unit region Rn corresponding to a specific pitch n and the position C[h,f] of a specific finger represented by finger position data F (Sc18). That is, the control device 11 sets the initial matrix W0 in such a way that the position C[h,f] of the specific finger represented by finger position data F is close to the unit region Rn of the specific pitch n in the reference image Gref. Specifically, the matrix used to project the position C[h,f] of the specific finger onto the center of the unit region Rn is set as the initial matrix W0.

[0146] The fourth embodiment achieves the same effect as the first embodiment. Furthermore, in the fourth embodiment, if the user plays a desired pitch n with a specific finger, the initial matrix W0 is set such that the position c[h,f] of the specific finger in the playing image G1 is close to the portion (unit area Rn) corresponding to the specific pitch n in the reference image Gref. Since the user can simply play the desired pitch n, the workload of setting the initial matrix W0 is reduced compared to the first embodiment, which requires the user to select the target area 621 through the operation of the operating device 13. On the other hand, according to the first embodiment where the user specifies the target area 621, the estimation of the user's finger position C[h,f] is not required; therefore, compared to the second embodiment, the influence of estimation errors can be reduced to set an appropriate initial matrix W0. Moreover, the fourth embodiment is also applicable to the second or third embodiments.

[0147] Furthermore, while the fourth embodiment envisions a scenario where the user plays one specific pitch n, it is also possible for the user to play multiple specific pitches n using specific fingers. The control device 11 sets an initial matrix W0 for each of the multiple specific pitches n in such a way that the position C[h,f] of the specific finger when playing that specific pitch n is close to the unit region Rn of that specific pitch n.

[0148] 5: Fifth Implementation Method

[0149] Figure 21 This is a block diagram illustrating the functional structure of the performance analysis system 100 according to the fifth embodiment. The performance analysis system 100 of the fifth embodiment includes a pickup device 16. The pickup device 16 picks up sounds played from the keyboard instrument 200 by the user's playing, thereby generating an audio signal V. The audio signal V is an audio signal representing the time region of the waveform of the sound played by the keyboard instrument 200. Furthermore, the pickup device 16, which is separate from the performance analysis system 100, can be connected to the performance analysis system 100 in a wired or wireless manner. Furthermore, the time sequence of samples constituting the audio signal V can be interpreted as "performance data P".

[0150] The control device 11 of the performance analysis system 100 functions as the performance analysis unit 30 by executing a program stored in the storage device 12. The performance analysis unit 30 generates finger movement data Q using an acoustic signal V supplied from the pickup device 16 and image data D1 supplied from the imaging device 15. Similar to the first embodiment, the finger movement data Q specifies the pitch n corresponding to the key 21 operated by the user and the finger number k of the finger used by the user when operating the key 21. In the first embodiment, the pitch n is specified by performance data P, but the acoustic signal V in the fifth embodiment is not a signal that directly specifies the pitch n. Therefore, the performance analysis unit 30 simultaneously estimates the pitch n and the finger number k using the acoustic signal V and the image data D1.

[0151] To estimate the pitch n and the finger number k, we hypothesize a latent variable w. t,n,k The notation t is a variable representing time. One unit period on the time axis can be indicated by the variable t. In addition, the finger number k in the fifth embodiment is set to any one of 11 values, including 10 values ​​(k = 1 to 10) corresponding to different fingers and a specified invalid value (k = 0).

[0152] Prepare latent variables w for each combination of pitch n and finger number k t,n,k Latent variable w t,n,k This is a variable used to represent the one-hot expression of a variable, set to either "0" or "1". Latent variable w t,n,k The value "1" represents playing pitch n using the finger numbered k, and the latent variable w t,n,k The value "0" indicates that no finger was used for playing.

[0153] In addition, let's assume the post-hoc probability U t,n And probability π t,n,k Post-hoc probability U t,n Given the observed acoustic signal V, the probability of producing a pitch n at time t is post-hoc. Therefore, the probability (1 - U) is... t,n This is equivalent to the latent variable w given the observed acoustic signal V. t,n,0 The probability of being 1 (the probability that no note n is played). Post-hoc probability U. t,n By analyzing the audio signal V and the subsequent probability U t,n The relationship between them is estimated using a well-known inference model. The inference model is a trained model used for automatic spectral acquisition. For example, deep neural networks such as convolutional neural networks or recurrent neural networks are used for estimating the posterior probability U. t,n The estimated model is used for estimation. Probability π t,n,kIt is the probability that the pitch n is played by the finger numbered k while the pitch n is played.

[0154] The acoustic signal V and probability π were observed. t,n,k The latent variable w at time t,n,k The probability p(w|V,π) is expressed by the following formula (6).

[0155] [Formula 6]

[0156]

[0157] The first term on the right side of equation (6) represents the probability that any pitch n is not pronounced, and the second term represents the probability that pitch n is pronounced and played by the finger numbered k.

[0158] In addition, after observing the latent variable w t,n,k The probability p(C[k]|w) of observing position C[k] from the performance image G1 is expressed by the following formula (7).

[0159] [Formula 7]

[0160]

[0161] The probability p(C[k]|σ) of equation (7) 2 Rn) is the probability expressed by the aforementioned formula (3) or formula (4).

[0162] In addition, as probability π t,n,k The prior distribution is conceived as a symmetrical Dirichlet distribution (Dir) represented by the following formula (8).

[0163] [Formula 8]

[0164]

[0165] The notation α in equation (8) is a variable that defines the shape of the symmetrical Dirichlet distribution.

[0166] Under the above premise, the latent variable w will be executed. t,n,k The maximum after-the-fact probability estimation (MAP) that maximizes the after-the-fact probability p(z|V,π,C[k]) can simultaneously estimate the presence or absence of pitch n and finger number k. However, estimating the probability distribution of the after-the-fact probability p(z|V,π,C[k]) is difficult. Therefore, in the fifth embodiment, the mean-field approximation (variational Bayesian estimation) is studied.

[0167] Specifically, determine the distribution among the distributions factored as shown in the following equation (9) that best approximates the probability distribution of the posterior probability p(z|V,π,C[k]). For example, determine the distribution whose KL (Kullback-Leibler) distance to the posterior probability p(z|V,π,C[k]) is minimized.

[0168] [Formula 9]

[0169]

[0170] Specifically, the performance analysis section 30 repeatedly performs the following calculations of formulas (10) and (11).

[0171] [Formula 10]

[0172]

[0173] [Formula 11]

[0174] q(π t,n,k ) = Dir(π t,n,k |α+ρ t,n,k (11)

[0175] The notation c of equation (10) is a probability distribution ρ such that the range of multiple finger numbers k is given by the expression. t,n,k The sum of these sums becomes "1" in this way for the probability distribution ρ t,n,k The coefficients are standardized. Additionally, the notation <> represents the expected value.

[0176] Specifically, the performance analysis unit 30 repeatedly performs calculations of formulas (10) and (11) on all combinations of pitch n and finger number k for a single time t on the time axis. The performance analysis unit 30 determines the result of the calculation of formula (10) at the time point on which formulas (10) and (11) have been repeatedly performed a predetermined number of times as the latent variable w. t,n,k probability distribution ρ t,n,k The probability distribution ρ is calculated for each time t on the time axis. t,n,k .

[0177] However, based on the probability distribution ρ calculated independently for each time t on the time axis... t,n,k However, in the method of calculating pitch n and finger number k for each time t, sometimes the finger number k changes before and after the user plays a note, or the duration of pitch n is too short. Therefore, the performance analysis unit 30 of the fifth embodiment utilizes a probability distribution ρ t,n,kThe time series of the combination of pitch n and finger number k (i.e., finger movement data Q) is generated by the Hidden Markov Model (HMM).

[0178] Specifically, the HMM used for finger deduction consists of potential states corresponding to the pronunciation (key press) and silencing of pitch n, and multiple potential states corresponding to different finger numbers k. As for state transitions, only three are allowed: (1) self-transition, (2) no sound → any finger number k, and (3) any finger number k → no sound. The transition probability for other state transitions is set to "0". The above conditions are constraints used to ensure that the finger number k does not change during the pronunciation of one note. In addition, the probability distribution ρ calculated by the operations of formulas (10) and (11) is... t,n,k The expected value is set as the observation probability associated with each potential state of the HMM. The performance analysis unit 30 uses the HMM described above to estimate the state sequence using a dynamic programming method such as the Viterbi algorithm. The performance analysis unit 30 generates a time series of the fingering data Q corresponding to the estimated state sequence.

[0179] According to the fifth embodiment, finger movement data Q is generated using the audio signal V and image data D1. That is, finger movement data Q can be generated even when performance data P is unavailable. Furthermore, in the fifth embodiment, pitch n and finger number k are estimated simultaneously using the audio signal V and image data D1. Therefore, compared to estimating pitch n and finger number k separately, the processing load can be reduced while achieving high-precision finger movement estimation. Moreover, the fifth embodiment is also applicable to embodiments 2 through 4.

[0180] 6: Sixth Implementation Method

[0181] As illustrated in the aforementioned embodiments, the projective transformation unit 314 generates a transformed image based on the performance image G1. That is, the projective transformation unit 314 changes the shooting conditions of the performance image G1. The sixth embodiment is an image processing system 700 that utilizes the above-mentioned function of changing the shooting conditions of the performance image G1. Furthermore, the performance analysis system 100 of the first to fifth embodiments, if considering the processing of the performance image G1 by the projective transformation unit 314, also manifests as an image processing system 700. Furthermore, in the sixth embodiment, the estimation of the user's finger movements is not necessary.

[0182] Figure 22This is a block diagram illustrating the functional structure of the image processing system 700 according to the sixth embodiment. Similar to the performance analysis system 100 of the first embodiment, the image processing system 700 includes a control device 11, a storage device 12, an operation device 13, a display device 14, and a shooting device 15. Similar to the first embodiment, the shooting device 15 generates a time series of image data D1 representing the performance image G1 by shooting the keyboard instrument 200 under specific shooting conditions.

[0183] Storage device 12 stores multiple reference data Drefs. Each reference data Dref represents a reference image Gref that has been photographed of the keyboard of a standard keyboard instrument, i.e., the reference instrument. The shooting conditions for the reference instrument are different for each reference image Gref (each reference data Dref). Specifically, for example, one or more conditions, such as the shooting range or shooting direction, are different for each reference image Gref. In addition, storage device 12 stores auxiliary data A for each reference data Dref.

[0184] The control device 11 implements the matrix generation unit 312, the projective transformation unit 314, and the display control unit 40 by executing a program stored in the storage device 12. The matrix generation unit 312 selectively generates a transformation matrix W using any one of a plurality of reference data Dref. The projective transformation unit 314 generates image data D3 of the transformed image G3 based on the image data D1 of the performance image G1 by utilizing the projective transformation of the transformation matrix W. The display control unit 40 displays the transformed image G3 represented by the image data D3 on the display device 14.

[0185] Figure 23 This is a flowchart illustrating the specific flow of the process (hereinafter referred to as "first image processing") performed by the control device 11 of the sixth embodiment. For example, the first image processing is started based on an instruction from the user to the operating device 13.

[0186] The user selects any one of multiple shooting conditions corresponding to different reference images Gref by operating the operation device 13. The control device 11 (matrix generation unit 312) determines whether it has received a shooting condition selection from the user (S31). If a shooting condition selection is received (S31: YES), the control device 11 (matrix generation unit 312) obtains the reference data Dref (hereinafter referred to as "selection reference data Dref") corresponding to the shooting condition selected by the user from the multiple reference data Dref stored in the storage device 12 (S32). The user's selection of shooting conditions is equivalent to the action of selecting any one of the multiple reference images Gref (reference data Dref) corresponding to different shooting conditions.

[0187] The control device 11 (matrix generation unit 312) performs the same matrix generation process as in the first embodiment (S33) using the selected reference data Dref. Specifically, the control device 11 sets an initial matrix W0 using the initial setting process Sc1 of the selected reference data Dref. In addition, the control device 11 generates a transformation matrix W by repeatedly updating the initial matrix W0 in a matrix update process Sc2, which makes the keyboard image g1 of the performance image G1 close to the reference image Gref of the selected reference data Dref. On the other hand, if no shooting condition selection is received (S31: NO), the selection of reference data Dref (S32) and the matrix generation process (S33) are not performed.

[0188] The control device 11 (projective transformation unit 314) generates a transformed image G3 (S34) by performing a projective transformation process using a transformation matrix W on the performance image G1. The projective transformation process is the same as in the first embodiment. Image data D3 representing the result of the projective transformation process, the transformed image G3, is generated. Specifically, the transformed image G3 is generated based on the performance image G1, and this transformed image G3 corresponds to the shooting conditions of the reference image Gref selected by the reference data Dref. That is, the transformed image G3 is an image that transforms the shooting conditions of the performance image G1 to the same shooting conditions as the reference image Gref. As understood from the above description, according to the sixth embodiment, a transformed image G3 corresponding to the shooting conditions selected by the user is generated.

[0189] The control device 11 (display control unit 40) displays the transformed image G3 generated by the projective transformation process on the display device 14 (S35). The control device 11 determines whether the end condition is met (S36). If, for example, the user indicates the end of the first image processing through operation of the operation device 13, the control device 11 determines that the end condition is met. If the end condition is not met (S36: NO), the control device 11 proceeds to step S31. That is, it performs the generation of the transformation matrix W with the acceptance of the shooting condition selection (S31: YES) as a condition (S32-S33) and the generation and display of the transformed image G3 (S34-S35). On the other hand, if the end condition is met (S36: YES), the control device 11 ends the first image processing.

[0190] As described above, in the sixth embodiment, a transformation matrix W is generated such that the keyboard image g1 of the performance image G1 approximates the reference image Gref, and a projective transformation process utilizing this transformation matrix W is performed on the performance image G1. Therefore, the performance image G1 of the keyboard instrument 200 played by the user can be transformed into a transformed image G3 corresponding to the shooting conditions of the reference instrument of the reference image Gref.

[0191] Furthermore, in the sixth embodiment, any one of the multiple reference data Drefs with different shooting conditions is selectively used in the matrix generation process. Therefore, it is possible to generate a transformed image G3 corresponding to multiple shooting conditions based on a performance image G1 captured under specific shooting conditions. In the sixth embodiment, the reference data Dref among the multiple reference data Drefs corresponding to the shooting conditions selected by the user is specifically used in the matrix generation process, thus enabling the generation of a transformed image G3 corresponding to the shooting conditions desired by the user. As described above, by changing the shooting conditions of the performance image G1, a transformed image G3 applicable to various purposes can be generated. For example, by performing the first image processing of the sixth embodiment on each of the multiple performance images G1 captured by a music instructor, multiple transformed images G3 with unified shooting conditions can be generated as teaching materials for music instruction.

[0192] 7: Implementation Method 7

[0193] As illustrated in the aforementioned embodiments, the image extraction unit 311 extracts a specific region B within the performance image G1, which includes the keyboard image g1 and the finger image g2. The seventh embodiment is an image processing system 700 that utilizes the above-described function of extracting the specific region B of the performance image G1. Furthermore, the performance analysis system 100 of the first to fifth embodiments, if considering the processing of the performance image G1 by the image extraction unit 311, is also manifested as the image processing system 700. Moreover, in the seventh embodiment, the estimation of the user's finger movements is not necessary.

[0194] Figure 24 This is a block diagram illustrating the functional structure of the image processing system 700 according to the seventh embodiment. Similar to the performance analysis system 100 of the first embodiment, the image processing system 700 includes a control device 11, a storage device 12, an operation device 13, a display device 14, and a shooting device 15. The shooting device 15 generates a time series of image data D1 representing a performance image G1 by capturing images of the keyboard instrument 200 under specific shooting conditions. As in the aforementioned embodiments, the performance image G1 includes a keyboard image g1 and a finger image g2.

[0195] The control device 11 functions as an image extraction unit 311 and a display control unit 40 by executing a program stored in the storage device 12. The image extraction unit 311 generates image data D2 representing a region of a performance image G2 from which a portion of the performance image G1 has been extracted. Specifically, similar to the first embodiment, the image extraction unit 311 performs a region estimation process Sb1 that generates an image processing mask M and a region extraction process Sb2 that applies the image processing mask M to the performance image G1. The display control unit 40 displays the performance image G2 represented by the image data D2 on the display device 14.

[0196] In the first embodiment, an inference model 51 for a single entity is illustrated. In the seventh embodiment, the inference model 51 used by the region inference process Sb1 includes a first model 511 and a second model 512. Both the first model 511 and the second model 512 are composed of deep neural networks such as convolutional neural networks or recurrent neural networks.

[0197] The first model 511 is a statistical model used to generate a first mask representing a first region in the performance image G1. The first region is the region in the performance image G1 that includes the keyboard image g1. The finger image g2 is not included in the first region. The first mask is, for example, a binary mask in which each element in the first region is set to the value "1", and each element in the region outside the first region is set to the value "0". The image extraction unit 311 generates the first mask by inputting the image data D1 representing the performance image G1 into the first model 511. That is, the first model 511 is a trained model that has learned the relationship between the image data D1 and the first mask (first region) through machine learning.

[0198] The second model 512 is a statistical model used to generate a second mask representing a second region within the performance image G1. The second region is the area within the performance image G1 that includes the finger image g2. The keyboard image g1 is not included in the second region. The second mask is, for example, a binary mask where each element within the second region is set to the value "1", and each element outside the second region is set to the value "0". The image extraction unit 311 generates the second mask by inputting the image data D1 representing the performance image G1 into the second model 512. That is, the second model 512 is a trained model that has learned the relationship between the image data D1 and the second mask (second region) through machine learning.

[0199] Figure 25 This is a flowchart illustrating the specific flow of the process (hereinafter referred to as "second image processing") performed by the control device 11 in the seventh embodiment. For example, the second image processing is started when the user gives an instruction to the operating device 13.

[0200] If the second image processing begins, the control device 11 (image extraction unit 311) performs region estimation processing Sb1 (S41 to S43). The region estimation processing Sb1 in the seventh embodiment includes a first estimation process (S41), a second estimation process (S42), and a region compositing process (S43).

[0201] The first estimation process is a process of estimating a first region of the performance image G1. Specifically, the control device 11 generates a first mask representing the first region by inputting image data D1 representing the performance image G1 into the first model 511 (S41). The second estimation process is a process of estimating a second region of the performance image G2. Specifically, the control device 11 generates a second mask representing the second region by inputting image data D1 representing the performance image G1 into the second model 512 (S42).

[0202] The region compositing process is the process of generating an image processing mask M representing a specific region B containing the first region and the second region. Specifically, the specific region B represented by the image processing mask M is equivalent to the sum of the first region and the second region. That is, the control device 11 generates the image processing mask M by compositing the first mask and the second mask (S43). As understood from the above description, the image processing mask M, like in the first embodiment, is a binary mask used to extract the specific region B containing the keyboard image g1 and the finger image g2 in the performance image G1.

[0203] The control device 11 (image extraction unit 311) performs the same region extraction process Sb2 (S44) as in the first embodiment using the image processing mask M generated in the region estimation process Sb1. That is, the control device 11 extracts a specific region B in the performance image G1 represented by the image data D1 using the image processing mask M, thereby generating image data D2 representing the performance image G2.

[0204] The control device 11 (display control unit 40) displays the performance image G2 generated by the region extraction process Sb2 on the display device 14 (S45). The control device 11 determines whether the end condition is met (S46). If, for example, the user indicates the end of the second image processing through operation of the operation device 13, the control device 11 determines that the end condition is met. If the end condition is not met (S46: NO), the control device 11 causes the processing to proceed to step S41. That is, the region estimation process Sb1 (S41-S43), the region extraction process Sb2 (S44), and the display of the performance image G2 are performed (S45). On the other hand, if the end condition is met (S46: YES), the control device 11 ends the second image processing.

[0205] In the seventh embodiment, similar to the first embodiment, a specific region B containing the keyboard image g1 is extracted from the performance image G1. Therefore, the convenience of the performance image G1 can be improved. Specifically, in the seventh embodiment, a specific region B containing both the keyboard image g1 and the finger image g2 is extracted from the performance image G1. Therefore, a performance image G2 that effectively provides visual confirmation of the condition of the keyboard 22 of the keyboard instrument 200 and the condition of the user's fingers can be generated.

[0206] Furthermore, according to the seventh embodiment, the first region containing the keyboard image g1 in the performance image G1 is estimated by the first model 511, and the second region containing the finger image g2 in the performance image G1 is estimated by the second model 512. Therefore, compared with the structure of the estimation model 51 that extracts both the keyboard image g1 and the finger image g2 in a concentrated manner, the specific region B containing the keyboard image g1 and the finger image g2 can be extracted with high accuracy. In addition, the first model 511 and the second model 512 are each created through independent machine learning, thus reducing the processing load associated with the machine learning of the first model 511 and the second model 512.

[0207] Furthermore, it is envisioned that the image extraction unit 311 can switch between a first mode and a second mode. The first mode is the operation mode for extracting both the keyboard image g1 and the finger image g2 from the performance image G1. That is, in the first mode, the image extraction unit 311 performs both the first estimation process and the second estimation process. Therefore, similar to the seventh embodiment, an image processing mask M representing a specific region B is generated. That is, in the first mode, a specific region B containing both the keyboard image g1 and the finger image g2 is extracted from the performance image G1.

[0208] The second mode is the action mode for extracting the keyboard image g1 from the performance image G1. That is, in the second mode, the image extraction unit 311 performs the first estimation process but does not perform the second estimation process. In other words, the first mask generated by the first estimation process is determined as the image processing mask M applied to the region extraction process Sb2. Therefore, in the second mode, the keyboard image g1 is extracted from the performance image G1.

[0209] As described above, by switching between the first mode and the second mode, the extraction target from the performance image G1 can be easily switched. Furthermore, in the above description, the image extraction unit 311 performed the first estimation process in the second mode, but it is also envisioned that in the second mode, the image extraction unit 311 performs the second estimation process but not the first estimation process. In the above manner, the finger image g2 is extracted from the performance image G1. As understood from the above example, the second mode is an action mode that performs either the first estimation process or the second estimation process.

[0210] 8: Variation Example

[0211] The following examples illustrate specific variations of the methods illustrated above. Two or more methods arbitrarily selected from the examples below can be appropriately combined without contradiction.

[0212] (1) In the aforementioned methods, image extraction processing will be used ( Figure 8 The matrix generation process is performed on the performance image G2 after processing, but the matrix generation process can also be performed on the performance image G1 captured by the imaging device 15. That is, the image extraction process (image extraction unit 311) for generating the performance image G2 based on the performance image G1 can be omitted.

[0213] In the aforementioned methods, finger position estimation processing using the performance image G1 is illustrated. However, finger position estimation processing can also be performed using the performance image G2 processed by image extraction processing. That is, the positions C[h,f] of each finger of the user can be estimated by analyzing the performance image G2. Furthermore, in the aforementioned methods, projective transformation processing is performed on the performance image G1 as the object. However, projective transformation processing can also be performed on the performance image G2 processed by image extraction processing. That is, a transformed image can be generated by performing a projective transformation on the performance image G2.

[0214] (2) In the aforementioned methods, the position c[h,f] of each user's finger is transformed into the position C[h,f] in the XY coordinate system through projective transformation. However, finger position data F representing the position c[h,f] of each finger can also be generated. That is, the projective transformation process (projective transformation unit 314) that transforms the position c[h,f] into the position C[h,f] can be omitted.

[0215] (3) In embodiments 1 to 5, an example is given of a method in which the transformation matrix W generated after the start of the performance analysis process is continued to be used in subsequent processes. However, the transformation matrix W can also be updated at appropriate points in time during the execution of the performance analysis process. For example, it is envisioned that the transformation matrix W is updated when the position of the shooting device 15 relative to the keyboard instrument 200 changes. Specifically, the transformation matrix W is updated when a change in the position of the shooting device 15 is detected by the analysis of the performance image G1 (hereinafter referred to as "position change"), or when the user indicates a change in the position of the shooting device 15.

[0216] Specifically, the matrix generation unit 312 generates a transformation matrix δ representing the position change (offset) of the shooting device 15. For example, the following formula (12) is envisioned to represent the coordinates (x,y) within the performance image G(G1, G2) after the position change.

[0217] [Formula 12]

[0218]

[0219] The matrix generation unit 312 generates a transformation matrix δ in such a way that the coordinate x' / ε calculated by formula (12) based on the x-coordinate of a specific location after the position change is approximately or consistent with the x-coordinate of the location corresponding to that location in the performance image G before the position change, and the coordinate y' / ε calculated by formula (12) based on the y-coordinate of the specific location after the position change is approximately or consistent with the y-coordinate of the location corresponding to that location in the performance image G before the position change, generates a transformation matrix δ. Furthermore, the matrix generation unit 312 generates an initial matrix W0 by using the product Wδ of the transformation matrix W before the position change and the transformation matrix δ representing the position change as the initial matrix W0, and updates the initial matrix W0 by matrix update processing Sc2, thereby generating the transformation matrix W.

[0220] In the above structure, the transformation matrix W after the position change is generated using the transformation matrix W calculated before the position change and the transformation matrix δ representing the position change. Therefore, the load of matrix generation processing can be reduced to generate a transformation matrix W that can accurately determine the position C[h,f] of each finger. Furthermore, while the first to fifth embodiments were envisioned in the above description, the transformation matrix W can also be updated at appropriate time points during the execution of the first image processing in the sixth embodiment.

[0221] (4) In the foregoing embodiments, a keyboard instrument 200 with a keyboard 22 is shown as an example, but the type of instrument to which the present invention is applied is arbitrary. For example, the foregoing embodiments are equally applicable to any instrument that can be manually operated by the user, such as string instruments, wind instruments, or percussion instruments. Typical examples of instruments are those that can be played by the user with the fingers of one or both hands.

[0222] (5) The performance analysis system 100 can be implemented by a server device that communicates with an information device, such as a smartphone or tablet. For example, performance data P generated by a keyboard instrument 200 connected to the information device and image data D1 generated by a camera 15 mounted or connected to the information device are sent from the information device to the performance analysis system 100. The performance analysis system 100 generates fingering data Q by performing performance analysis processing on the performance data P and image data D1 received from the information device, and sends the fingering data Q to the information device. Similarly, the image processing system 700 illustrated in the sixth or seventh embodiment can also be implemented by a server device that communicates with the information device.

[0223] (6) The functions of the performance analysis system 100 according to embodiments 1 to 5, or the image processing system 700 according to embodiments 6 to 7, as described above, are achieved through the coordinated operation of one or more processors constituting the control device 11 and the program stored in the storage device 12. The program involved in this invention is provided and installed on a computer in a manner stored on a computer-readable recording medium. The recording medium is, for example, a non-transitory recording medium, preferably an optical recording medium (optical disc) such as a CD-ROM, and also includes any known form of recording medium such as a semiconductor recording medium or a magnetic recording medium. Furthermore, as a non-transitory recording medium, it includes any recording medium other than a transient propagating signal, and may also include volatile recording media. In the configuration where the transmission device transmits the program via a communication network, the recording medium 12 in the transmission device that stores the program is equivalent to the aforementioned non-transitory recording medium.

[0224] 9: Appendix

[0225] Based on the examples above, for instance, understand the following structure.

[0226] One aspect (Aspect 1) of the present invention relates to an image processing method in which a transformation matrix for projective transformation of the coordinates of the performance image is generated such that the image of the instrument in the performance image, which includes an image of the instrument and images of multiple fingers of a user playing the instrument, approximates a reference image representing a reference instrument. The projective transformation of the performance image is then performed using the transformation matrix. In this aspect, the transformation matrix is ​​generated such that the image of the instrument in the performance image approximates a reference image, and a projective transformation using this transformation matrix is ​​performed on the performance image. Therefore, it is possible to transform a performance image of an instrument played by a user into an image corresponding to the shooting conditions of a reference instrument in the reference image.

[0227] One aspect (Aspect 2) of the present invention relates to an image processing method in which a transformation matrix is ​​generated such that the image of the keyboard in a performance image, including an image of the keyboard of a keyboard instrument and images of multiple fingers of a user playing the keyboard instrument, approximates a reference image representing a reference instrument. The projective transformation of the performance image is performed using the transformation matrix. In the above aspect, the transformation matrix is ​​generated such that the image of the keyboard in the performance image approximates a reference image, and a projective transformation using this transformation matrix is ​​performed on the performance image. Therefore, it is possible to transform a performance image of a keyboard instrument played by a user into an image corresponding to the shooting conditions of a reference instrument in the reference image.

[0228] In a specific example of Method 2 (Method 3), during the generation of the transformation matrix, an initial value, i.e., an initial matrix, is set. During the generation of the transformation matrix, the initial matrix is ​​repeatedly updated to increase the enhanced correlation coefficient between the reference image and the keyboard image in the performance image. Since the keyboard image shows repeated patterns of multiple keys, it may be impossible to properly estimate the transformation matrix using image feature quantities such as SIFT (Scale-Invariant Feature Transform). By repeatedly updating the initial matrix to increase the enhanced correlation coefficient (ECC), the transformation matrix can be properly estimated even when dealing with images containing repeated patterns.

[0229] In a specific example of Method 3 (Method 4), the initial matrix is ​​set as the matrix used to projectively transform the target region in the keyboard image of the playing image corresponding to the user's instruction into the region in the reference image corresponding to a specific pitch. In the process of repeatedly updating the initial matrix to increase the correlation coefficient, there is a tendency for the appropriateness of the initial matrix to easily affect the appropriateness of the final transformation matrix. The structure of setting the initial matrix according to the target region corresponding to the user's instruction enables the generation of an appropriate transformation matrix that makes the keyboard image and the reference image highly approximate.

[0230] In a specific example of Method 4 (Method 5), the initial matrix is ​​set by using the region in the performance image specified by the user through operation of the operating device as the target region. In the above methods, the user uses the region in the performance image specified by the user through operation of the operating device as the target region for setting the initial matrix. Therefore, compared to methods that estimate, for example, the region corresponding to a specific pitch in the performance image through computational processing, an appropriate initial matrix can be set while reducing the processing load.

[0231] In a specific example of Method 4 (Method 6), the initial matrix setting involves obtaining performance data specifying the pitch played by the user using the keyboard instrument, and finger position data indicating the position of the finger playing the pitch in the performance image. The initial matrix is ​​set such that the finger position indicated by the finger position data is close to the portion of the reference image corresponding to the pitch specified by the performance data. In the above method, if the user plays a desired pitch with a specific finger, the initial matrix is ​​set such that the position of that finger in the performance image is close to the portion of the reference image corresponding to that pitch. Based on this structure, the user can simply play the desired pitch, thus reducing the workload of setting the initial matrix.

[0232] One aspect (aspect 7) of the present invention relates to an image processing system comprising: a matrix generation unit that generates a transformation matrix for projectively transforming the performance image in such a way that the image of the instrument in a performance image, including an image of an instrument and images of multiple fingers of a user playing the instrument, approximates an image of a reference instrument included in a reference image; and a projective transformation unit that performs a projective transformation of the performance image using the transformation matrix.

[0233] One aspect (aspect 8) of the present invention relates to a program that causes a computer system to function as a matrix generation unit that generates a transformation matrix for projectively transforming the performance image in such a way that the image of the instrument in a performance image, which includes an image of the instrument and images of multiple fingers of a user playing the instrument, approximates an image of a reference instrument included in a reference image; and a projective transformation unit that performs a projective transformation of the performance image using the transformation matrix.

[0234] Furthermore, the contents of Japanese Patent Application No. 2021-051180, filed on March 25, 2021, are incorporated herein by reference.

[0235] Explanation of the label

[0236] 100… Performance Analysis System

[0237] 11…Control device

[0238] 12… Storage devices

[0239] 13…operating device

[0240] 14… Display device

[0241] 15…Filming device

[0242] 200… Keyboard Instruments

[0243] 21…key

[0244] 22…keyboard

[0245] 30…Performance Analysis Section

[0246] 31…Finger position data generation unit

[0247] 311…Image Extraction Department

[0248] 312…Matrix Generation Department

[0249] 313… Finger position estimation section

[0250] 314…Projective Transformation Department

[0251] 32… Operations Data Generation Department

[0252] 321…Probability Calculation Department

[0253] 322…Traffic Index Prediction Department

[0254] 323…Control Data Generation Department

[0255] 40… Display Control Unit

[0256] 51…Presumed Model

[0257] 51a… Temporary Model

[0258] 52[k]...Estimated Model

[0259] 700… Image Processing System

Claims

1. An image processing method implemented by a computer system, A transformation matrix is ​​generated to projectively transform the coordinates of the performance image in such a way that the image of the instrument in the performance image, which includes an image of the instrument and images of multiple fingers of the user playing the instrument, approximates a reference image representing a reference instrument. The projective transformation of the performance image is performed using the transformation matrix.

2. An image processing method implemented by a computer system. A transformation matrix for projective transformation of the performance image is generated in such a way that the image of the keyboard in the performance image, which includes an image of the keyboard of a keyboard instrument and images of multiple fingers of the user playing the keyboard instrument, approximates a reference image representing a reference instrument. The projective transformation of the performance image is performed using the transformation matrix.

3. The image processing method according to claim 2, wherein, In the generation of the transformation matrix, The initial value of the transformation matrix, i.e., the initial matrix, is set. In the generation of the transformation matrix, the initial matrix is ​​repeatedly updated to increase the enhanced correlation coefficient between the reference image and the keyboard image in the performance image.

4. The image processing method according to claim 3, wherein, In the setting of the initial matrix, The matrix used to project and transform the target region in the keyboard image of the performance image corresponding to the instructions from the user into the region in the reference image corresponding to a specific pitch is set as the initial matrix.

5. The image processing method according to claim 4, wherein, In the setting of the initial matrix, The initial matrix is ​​set by using the region in the performance image specified by the user through the operation of the operating device as the target region.

6. The image processing method according to claim 4, wherein, In the setting of the initial matrix, Acquire performance data specifying the pitch played by the user using the keyboard instrument, and finger position data indicating the position of the finger played by the user in the performance image. The initial matrix is ​​set such that the finger position represented by the finger position data is close to the portion in the reference image corresponding to the pitch specified by the performance data.

7. An image processing system, comprising: A matrix generation unit generates a transformation matrix for projectively transforming the performance image in a manner that makes the image of the instrument in the performance image, including an image of the instrument and images of multiple fingers of the user playing the instrument, approximate an image of a reference instrument included in the reference image; and The projective transformation unit performs the projective transformation of the performance image using the transformation matrix.

8. The image processing system according to claim 7, wherein, The matrix generation unit sets the initial value of the transformation matrix, i.e., the initial matrix. In the generation of the transformation matrix, the initial matrix is ​​repeatedly updated to increase the enhanced correlation coefficient between the reference image and the image of the instrument in the performance image.

9. The image processing system according to claim 8, wherein, In the setting of the initial matrix, The matrix used to project and transform the target region in the image of the instrument in the performance image corresponding to the instruction from the user into the region in the reference image corresponding to a specific pitch is set as the initial matrix.

10. The image processing system according to claim 9, wherein, In the setting of the initial matrix, The initial matrix is ​​set by using the region in the performance image specified by the user through the operation of the operating device as the target region.

11. The image processing system according to claim 9, wherein, In the setting of the initial matrix, Acquire performance data specifying the pitch played by the user using the instrument, and finger position data indicating the position of the finger played by the user in the performance image. The initial matrix is ​​set such that the finger position represented by the finger position data is close to the portion in the reference image corresponding to the pitch specified by the performance data.

12. A storage medium storing a program that causes a computer system to function as the following functional unit: A matrix generation unit generates a transformation matrix for projectively transforming the performance image in a manner that makes the image of the instrument in the performance image, including an image of the instrument and images of multiple fingers of the user playing the instrument, approximate an image of a reference instrument included in the reference image; and The projective transformation unit performs the projective transformation of the performance image using the transformation matrix.

Citation Information

Patent Citations

  • Display device, display system, method for adjusting display, and display adjustment program

    JP2021051180A

  • Methods and systems for visual music transcription

    US9418637B1

  • Adaptive video image splicing method and device based on video frame matching information

    CN109035145A

  • Image processing device, image capture system, image processing method, and image processing program

    CN109791685A