Signal processing device and signal processing method

The signal processing device automatically detects and estimates attention levels in musical performances using a learning model, addressing the challenge of identifying attention-grabbing scenes beyond the song's climax, thereby enhancing video editing efficiency.

JP7771772B2Active Publication Date: 2025-11-18YAMAHA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022007337
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-20
Publication Date
2025-11-18
Estimated Expiration
2042-01-20

AI Technical Summary

Technical Problem

Existing video editing technologies fail to accurately detect scenes in a musical performance that generate significant attention beyond the song's climax, requiring user expertise in musical performance to identify attention-grabbing moments.

Method used

A signal processing device and method that utilizes a trained learning model to estimate the attention level of drum performances in images based on features such as rhythm similarity, movement analysis, and sound characteristics, enabling automatic detection of attention-grabbing scenes.

Benefits of technology

Enables users to recognize and select images or videos that attract significant attention, facilitating enhanced video editing by highlighting high-attention moments without requiring musical knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007771772000001
    Figure 0007771772000001
  • Figure 0007771772000002
    Figure 0007771772000002
  • Figure 0007771772000003
    Figure 0007771772000003
Patent Text Reader

Abstract

To estimate a degree of attention in a performance shown in an image.SOLUTION: A signal processor includes: an image acquisition unit for acquiring a performance image captured so as to include a drum performance; an estimation unit for estimating, based on a feature quantity related to the drum performance obtained from the performance image, a degree of attention by inputting the performance image to a learned learning model having carried out machine learning for estimating the degree of attention which is a degree of attention received by the drum performance in the performance image; and an output unit for outputting the degree of attention estimated by the estimation unit.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a signal processing device and a signal processing method. [Background technology]

[0002] Video posting services allow users to freely post videos, and the posted videos can be viewed by users and other viewers. Some users attract the attention of many viewers by posting videos, and there is a need to create videos that will attract viewers' attention. Patent Document 1 discloses a technology for detecting the climax of a song from music information. Using this technology, for example, when posting a video of a band playing a song, it is possible to create a highly attention-grabbing video that will attract viewers' attention by editing the video to include images that correspond to the climax of the song. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2004-127019 Summary of the Invention [Problem to be solved by the invention]

[0004] However, when focusing on a particular instrument that makes up a band, the phrases that generate excitement in a song may not coincide with the phrases that generate significant attention. For example, a particular instrument may generate significant attention in a phrase that is different from the climax of a song, such as a drum fill-in played in a phrase before the climax of a song. To determine whether a particular instrument is generating significant attention, a user editing the video is required to have knowledge and experience regarding musical performance. It is desirable to be able to detect scenes in which significant attention is generated, even in phrases that are different from the climax of a song, as images that generate significant attention.

[0005] The present invention has been made in view of the above circumstances, and its purpose is to provide a signal processing device and a signal processing method that can estimate the level of attention to a performance shown in an image. [Means for solving the problem]

[0006] In order to solve the above-mentioned problems, one aspect of the present invention includes an image acquisition unit that acquires a performance image captured so that the performance of a drum is included; an estimation unit that estimates an attention level by inputting the performance image into a trained learning model that has undergone machine learning to estimate an attention level, which is the degree to which the drum performance in the performance image is attracting attention, based on a feature amount related to the drum performance obtained from the performance image; and an output unit that outputs the attention level estimated by the estimation unit. The learning data to be machine-learned by the learning model is associated with the attention level based on a feature amount corresponding to a degree of similarity between the rhythm of a drum tone included in a performance sound corresponding to a learning image and the rhythm of a performance sound output from a tone of an instrument other than a drum. It is a signal processing device.

[0007] In order to solve the above-mentioned problems, one aspect of the present invention is a method for generating a performance image including a drum performance, the method comprising: an image acquisition unit that acquires a performance image captured so that the drum performance is included; an estimation unit that estimates an attention level, which is the degree to which the drum performance in the performance image is attracting attention, based on a feature amount related to the drum performance obtained from the performance image; and an output unit that outputs the attention level estimated by the estimation unit. The estimation unit estimates the attention level based on a feature amount relating to the drum performance, the feature amount corresponding to a degree of similarity between the rhythm of a drum tone included in the performance sound corresponding to the performance image and the rhythm of a performance sound output from a tone of an instrument other than the drum. It is a signal processing device.

[0008] In one aspect of the present invention, a performance image captured so as to include a drum performance is acquired, and an attention level, which is the degree to which the drum performance in the performance image is attracting attention, is estimated based on a feature amount related to the drum performance obtained from the performance image, and the estimated attention level is output. In the estimating step, the attention level is estimated based on a feature amount relating to the drum performance, the feature amount corresponding to the degree of similarity between the rhythm of the drum tone included in the performance sound corresponding to the performance image and the rhythm of the performance sound output from the tone of an instrument other than the drum. This is a signal processing method. [Effects of the Invention]

[0009] According to the present invention, it is possible to estimate the degree of attention of a performance shown in an image, thereby enabling a user to recognize the degree of attention of the image and to select an image showing a performance that is attracting a great deal of attention, for example. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a block diagram showing an example of the configuration of a signal processing system 1 according to an embodiment. [Figure 2] 1 is a block diagram showing an example of the configuration of a user terminal 10 according to an embodiment. [Figure 3] 1 is a block diagram showing an example of the configuration of a signal processing device 20 according to an embodiment. [Figure 4] FIG. 2 is a diagram showing an example of image information 220 in the embodiment. [Figure 5] FIG. 2 is a diagram showing an example of musical score information 221 in the embodiment. [Figure 6] FIG. 10 is a diagram illustrating a process performed by an editing unit 232 in the embodiment. [Figure 7] 10A and 10B are diagrams illustrating an example of a method for determining an attention level from an image according to an embodiment. [Figure 8] 10A and 10B are diagrams illustrating an example of a method for determining an attention level from a performance sound in an embodiment. [Figure 9] 3 is a sequence diagram showing the flow of processing performed by a signal processing system 1 in the embodiment. FIG. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0012] The signal processing system 1 of this embodiment is a system that detects images that attract a lot of attention from moving images of a musical performance (hereinafter referred to as a performance video). Fig. 1 is a block diagram showing an example of the configuration of the signal processing system 1. The signal processing system 1 includes a plurality of user terminals 10 (user terminals 10-1, 10-2, ..., 10-N, N is any natural number) and a signal processing device 20. Each of the plurality of user terminals 10 and the signal processing device 20 are communicatively connected via a communication network NW.

[0013] The user terminal 10 is a computer, such as a PC (Personal Computer), a tablet terminal, or a smartphone. The user terminal 10 acquires a performance video captured by a user. The user terminal 10 transmits the acquired performance video to the signal processing device 20.

[0014] The signal processing device 20 is a computer, such as a server device or a PC. The signal processing device 20 receives a performance video from the user terminal 10, estimates the attention level of an image (hereinafter referred to as a performance image) included in the received performance video, and transmits the estimation result to the user terminal 10. The signal processing device 20 may edit the performance video based on the estimation result and transmit the edited performance video to the user terminal 10.

[0015] In this embodiment, the degree of attention is the degree to which an image attracts the interest of a viewer. For example, if a viewer is very interested in the way a performer plays, that image is an image with a high degree of attention. The degree of attention may also be the degree to which the performance sound associated with the performance video attracts the interest of a listener. For example, if the way a performer plays is monotonous and does not attract much interest from a viewer, but the listener is very interested in the performance sound produced by that performer, that image is an image with a high degree of attention.

[0016] In the following, an example will be described in which a performance video includes footage of drums being played, and a highly popular drum performance is detected from the performance video. The footage of drums being played here does not necessarily include both the drum set and the drummer; it is sufficient to include at least a portion of the drum set or the drummer. For example, it is sufficient if at least some of the images in the image group include a portion of the drum set or a portion of the drummer's body (such as the face or arm). Alternatively, it is sufficient if the sound associated with the performance video includes the sound of the drums being played.

[0017] 2 is a block diagram showing an example configuration of the user terminal 10. The user terminal 10 includes, for example, a communication unit 11, a storage unit 12, a control unit 13, a display unit 14, and an imaging unit 15. The communication unit 11 communicates with the signal processing device 20. The communication unit 11 transmits a performance video captured by the user to the signal processing device 20.

[0018] The storage unit 12 is configured by a storage medium such as a HDD, flash memory, EEPROM (Electrically Erasable Programmable Read Only Memory), RAM (Random Access Read / Write Memory), ROM (Read Only Memory), or a combination of these. The storage unit 12 stores programs for executing various processes of the user terminal 10 and temporary data used when performing various processes.

[0019] The control unit 13 realizes its functions by causing a CPU (Central Processing Unit) provided as hardware in the user terminal 10 to execute a program stored in the storage unit 12. The control unit 13 comprehensively controls the user terminal 10. The control unit 13 controls each of the communication unit 11, the storage unit 12, the display unit 14, and the imaging unit 15.

[0020] The display unit 14 includes a display device such as a liquid crystal display, and displays still images and moving images according to the control of the control unit 13. The imaging unit 15 includes an imaging device, and captures a moving image of a performance according to the control of the control unit 13.

[0021] 3 is a block diagram showing an example configuration of the signal processing device 20. The signal processing device 20 includes, for example, a communication unit 21, a storage unit 22, and a control unit 23. The communication unit 21 communicates with the user terminal 10. The communication unit 21 receives a performance video from the user terminal 10.

[0022] The storage unit 22 is configured by a storage medium such as a HDD, a flash memory, an EEPROM, a RAM, a ROM, or a combination of these. The storage unit 22 stores programs for executing various processes of the signal processing device 20 and temporary data used when performing various processes.

[0023] The storage unit 22 stores, for example, image information 220, musical score information 221, and a trained model 222. The image information 220 is information indicating a performance video. The musical score information 221 is information indicating the musical score of the song performed in the performance video. The trained model 222 is information indicating a trained model used for estimation by the estimation unit 231, which will be described later. The trained model 222 stores information used to construct the model. For example, if the trained model is a model based on a DNN (Deep Neural Network), the information stored therein indicates the number of units in each of the input layer, intermediate layer, and output layer, the number of layers in the intermediate layer, the coupling coefficients and bias values ​​between units, and the activation function, etc.

[0024] FIG. 4 is a diagram showing an example of image information 220. The image information 220 stores information corresponding to each of the following items: time information, image information, and sound information. The time information is information indicating the elapsed time based on a predetermined time, such as the time when imaging of a performance video started. The image information is information indicating a performance image captured at a time specified by the time information. The sound information is information indicating a performance sound performed at a time specified by the time information. The performance sound includes the sound of an instrument played by a performer, vocal sounds, special sound effects, pre-sampled sounds, etc.

[0025] FIG. 5 is a diagram showing an example of the musical score information 221. The musical score information 221 stores, for example, information corresponding to each item of time information and event information. The time information is information indicating the elapsed time based on a predetermined time such as the start of performance. The event information is information indicating the tone to be output at a time specified by the time information, as well as the intensity and duration of that tone. In addition to the time information and event information, the musical score information 221 may also include information such as the song title, tempo, beat, and lyrics. For example, a MIDI (Musical Instrument Digital Interface) file can be used as the musical score information 221.

[0026] 3 , the control unit 23 is realized by causing a CPU provided as hardware in the signal processing device 20 to execute a program. The control unit 23 comprehensively controls the signal processing device 20. The control unit 23 controls each of the communication unit 21 and the storage unit 22.

[0027] The control unit 23 includes, for example, an image acquisition unit 230, an estimation unit 231, an editing unit 232, an output unit 233, and a learning unit 234.

[0028] The image acquisition section 230 acquires a performance video captured by a user via the communication section 21. The image acquisition section 230 stores the acquired performance video in the storage section 22 as image information 220.

[0029] The estimation unit 231 estimates the attention level of an image included in a performance video using a trained model. The estimation unit 231 constructs a trained model by referring to the trained model 222 in the storage unit 22. The estimation unit 231 inputs an image into the constructed trained model. The trained model estimates the attention level of the input image and outputs the estimation result. The estimation unit 231 regards the estimation result output from the trained model as the attention level of the image.

[0030] The editing unit 232 edits the performance video. For example, the editing unit 232 edits the performance images according to the attention level. Specifically, the editing unit 232 zooms in and enlarges the performance images whose attention level is greater than a threshold, and generates a video using the enlarged images.

[0031] Alternatively, the editing unit 232 may generate a video by shortening the performance video. For example, if there is a limit on the file size of videos that can be posted to a video sharing site, it is necessary to shorten the performance video to generate a video with a file size that can be posted. For example, to accommodate such a limit, the editing unit 232 generates a video by shortening the performance video. The editing unit 232 selects performance images with an attention level greater than a threshold based on the attention level of the performance images included in the performance video. The editing unit 232 generates a video using the selected images. For example, the editing unit 232 generates a video in which the selected performance images are arranged in chronological order. In this case, the editing unit 232 may enlarge performance images with particularly high attention levels from the performance video used for editing, and generate a video using the enlarged performance images. When selecting a performance image with a level of attention greater than the threshold, the editing unit 232 may select a group of images that includes a target image with a level of attention greater than the threshold and images before and after the target image. This allows the target image and the images before and after it to be displayed in chronological order, and makes it possible to display images with low to high levels of attention. Therefore, the target image can attract more attention from viewers than when only the target image is displayed.

[0032] Furthermore, the editing unit 232 may generate one moving image using a plurality of performance videos in which the drum performance is captured from different directions.

[0033] Here, a method by which the editing unit 232 generates one moving image using a plurality of performance videos will be described with reference to Fig. 6. Fig. 6 is a diagram illustrating the processing performed by the editing unit 232. Fig. 6 shows performance images included in a plurality of performance videos G (performance videos G1 to G3) in chronological order.

[0034] 6 is based on the assumption that the estimation unit 231 has estimated the attention level for each performance image included in the performance video G. In the example shown in this figure, in performance video G1, the image group shown between times T1 and T2 (symbol A) and the image group shown between times T5 and T6 are estimated to have an attention level higher than the threshold. In performance video G2, the image group shown between times T3 and T4 (symbol C) are estimated to have an attention level higher than the threshold. In performance video G3, the image group shown between times T7 and T8 (symbol D) are estimated to have an attention level higher than the threshold. Furthermore, in the example shown in this figure, it is determined that no performance is taking place in the images shown before time T0 and after time T8 (symbol X), and an attention level indicating that they are receiving little attention (for example, the lowest value) is associated with them.

[0035] The editing unit 232 identifies captured images captured at the same time from each of the performance videos G. For example, the editing unit 232 identifies performance images captured at the same time based on commonalities in performance sounds associated with each of the performance videos G. Alternatively, the editing unit 232 may identify performance images captured at the same time based on time codes set in each of the performance videos G.

[0036] The editing unit 232 selects performance images with a level of attention greater than a threshold based on the level of attention of the performance images included in each performance video. Specifically, the editing unit 232 selects a group of images corresponding to the symbols A to D. The editing unit 232 generates a moving image in which the selected group of images is arranged in chronological order. For example, the editing unit 232 generates a moving image in which the group of images is arranged in the order of symbol A, symbol C, symbol B, and symbol D. In this case, the editing unit 232 may enlarge some of the performance images that make up the moving image, especially performance images with a high level of attention, and generate a moving image using the enlarged performance images.

[0037] The output unit 233 outputs the estimation result estimated by the estimation unit 231, i.e., the attention level estimated in the performance image. Alternatively, the output unit 233 may output a moving image edited by the editing unit 232. The information output by the output unit 233 is transmitted to the user terminal 10 via the communication unit 21.

[0038] What the output unit 233 outputs is determined based on a user request. For example, if the user edits the performance video based on the attention level estimated in the performance image, the output unit 233 outputs the attention level estimated in the performance image. On the other hand, if the user requests the signal processing device 20 to edit the performance image, the output unit 233 outputs the performance video edited by the editing unit 232. Furthermore, the output unit 233 may output information that enables the user to edit the video using an image with a high level of attention. For example, the output unit 233 outputs information to the user terminal 10 for sorting and displaying performance videos captured by multiple cameras according to their levels of attention. As a result, on the display screen of the user terminal 10, for example, of multiple performance videos, those with high levels of attention are displayed at the top and those with low levels of attention are displayed at the bottom. Therefore, the user can view the images from top to bottom and can select an image with high levels of attention as an image to use for editing without having to view all of the images. Alternatively, the output unit 233 may extract a video portion that has a length corresponding to the time length specified by the user and that is relatively attracting attention, and output information about the extracted video portion to the user terminal 10. This makes it possible to suggest a video portion that is attracting a lot of attention from the user and that has an appropriate length that matches the time period specified by the user. Alternatively, the output unit 233 may output, as an estimation result, information showing an image with a particularly large attention point as a thumbnail to the user terminal 10. This allows the output unit 233 to propose an image with a large attention point in a display format that is easy for the user to understand. Alternatively, the output unit 233 may generate a thumbnail using an image with a large focus point, and allow the generated thumbnail to be posted to an SNS using the user's account. In this case, for example, the output unit 233 generates a thumbnail using an image with a large focus point, and when transmitting the generated thumbnail to the user terminal 10, transmits information indicating a button labeled "Post" or the like along with the thumbnail. As a result, a button labeled "Post" or the like is displayed on the display screen of the user terminal 10 along with the thumbnail. The user views the thumbnail and touches the button to post to the SNS. When the touch operation is performed, the user terminal 10 acquires operation information indicating the touch operation and transmits the acquired operation information to the signal processing device 20. The signal processing device 20 posts the thumbnail to the SNS using the user's pre-registered account based on the operation information received from the user terminal 10.

[0039] The learning unit 234 generates a trained model. The trained model is a model that has been trained to output the attention level of an input image by having a machine learning model learn a training dataset. The model here is, for example, a DNN. However, the model is not limited to a DNN, and any learning model such as a CNN (Convolutional Neural Network), an RNN (Recurrent Neural Network), a combination of a CNN and an RNN, an HMM (Hidden Markov Model), or an SVM (Support Vector Machine) may be used.

[0040] The training dataset in this embodiment is information that pairs (sets) of training images and the attention levels for the training images. The training images are images included in unspecified performance videos that capture scenes such as past band performances. The training images include scenes of drum performances, and include images that are highly attention-grabbing because the drum performance attracts the viewer's attention, and images that are not so attention-grabbing. In this way, a correlation exists between the training images and the attention levels. By creating a trained model that has learned this correlation, the trained model can estimate the attention levels for images. For example, a training dataset is generated by having an expert or the like who views the training images assign attention levels to the images. The expert here is someone with extensive experience in playing drums, editing videos related to drum performances, or watching videos, i.e., someone who is well-versed in attention-grabbing scenes in drum performances. Such an expert can generate a training dataset in which appropriate attention levels are associated with the training images. Therefore, even a user who does not have knowledge or experience regarding musical performance can understand that a performance that attracts a lot of attention is being performed in a phrase that is different from the exciting part of the song by utilizing a trained model that has learned from a training dataset to which appropriate attention levels are associated.

[0041] The learning unit 234 causes the model to learn the correlation between training images in the training dataset and the attention level. For example, if the model is a model constructed using a DNN, the learning unit 234 sets model parameters (e.g., coupling coefficients between units and bias values) so that, when a training image in the training dataset is input, the attention level associated with the image is output. When parameters that can accurately output the attention level for all training images in the training dataset can be set, the learning unit 234 designates the model as a trained model. In this way, by determining appropriate parameters based on the correlation in the training dataset, the trained model can accurately estimate the attention level for an image. The learning unit 234 stores information indicating the generated trained model in the storage unit 22 as image information 220.

[0042] Here, a method for determining the level of attention in a learning image will be described with reference to Figs. 7 and 8. Fig. 7 is a diagram illustrating an example of a method for determining the level of attention from the movement of a performer in an image. Fig. 8 is a diagram illustrating an example of a method for determining the level of attention from a performance sound associated with an image. The process for determining the level of attention shown in Figs. 7 and 8 may be performed by a person such as an expert, or may be executed by the learning unit 234 using image processing technology or the like. A case where the learning unit 234 performs the process for determining the level of attention will be described below.

[0043] For example, the training data set is associated with an attention level based on a feature amount corresponding to the movement of the drummer shown in the training image. That is, a feature amount corresponding to the movement of the drummer shown in the training image is calculated, and an attention level is determined based on the calculated feature amount. The feature amount corresponding to the movement of the drummer is an example of a "feature amount related to drum performance."

[0044] For example, when a drummer is playing a monotonous rhythm, the drummer's arms will repeatedly move in a fixed manner. In contrast, when a performance is a highlight, such as when the drummer hits the drumsticks with a full stroke, the drummer's arms will move more widely than when the drummer is playing a monotonous rhythm. By determining the degree of movement of the drummer's arms as a feature, scenes in which the drummer is performing with large arm movements can be detected as images with a high degree of attention.

[0045] For example, when drums are playing a monotonous rhythm, the drummer's gaze is fixed toward the drum set, e.g., toward the drums or cymbals, which are the target instruments being played (hit), and barely moves. In contrast, when performing a so-called "kime" (a rhythmic climax), in which the drummer plays in time with other performers, the drummer moves his or her gaze from the direction of the target instrument being played, makes eye contact with the other performers, and synchronizes with the other performers. Similarly, when performing a "break" (a rhythmic climax), in which the drummer stops playing in time with the other performers, the drummer moves his or her gaze from the direction of the drums or cymbals being played, makes eye contact with the other performers, and synchronizes with the other performers. Thus, when performing a performance that requires a high level of attention, such as a "kime" (a rhythmic climax) or a "break," the drummer's gaze tends to be directed in a direction different from the target instrument being played. In light of this tendency in drum performances, the drummer's gaze direction is used as a feature to determine the level of attention. For example, a feature value is calculated that reduces the level of attention when the drummer's line of sight is directed in the direction of the object he or she is actually playing, and a feature value is calculated that increases the level of attention when the drummer's line of sight is not directed in the direction of the object he or she is actually playing. By determining the level of attention using the direction of the drummer's line of sight as a feature value, scenes in which the drummer performs a "finish" or "break" can be detected as images with a high level of attention.

[0046] For example, while the drummer is playing a monotonous rhythm, the drummer moves less than the other players. In contrast, when a drummer is playing a solo, it is thought that only the drummer moves, and that the other players do not move. By determining the attention level using the difference between the drummer's movement and the movement of the other players as a feature, it is possible to detect a scene in which the drummer is playing a solo as an image with high attention level.

[0047] For example, when a drummer plays a monotonous rhythm, the drummer's arms move, but the drummer maintains a seated posture and the upper half of the body, other than the arms, barely moves. In contrast, when adding a fill-in to liven up a song or playing a specific instrument such as a wind chime or tambourine, the drummer's upper body posture changes. For example, the drummer may change the direction of their upper body to play a specific instrument or hit multiple cymbals. Or, when the drummer joins the chorus, the microphone moves in a set direction. By comparing the degree of movement of the drummer's upper body, scenes in which the drummer performs a special performance, a flashy performance such as spinning sticks, or a solo performance can be detected as images that are likely to attract attention.

[0048] In this way, by focusing on the movement of the performer, it is possible to determine an appropriate attention level. Figure 7 shows the process of calculating a feature amount according to the presence or absence of the performer's movement, and associating an attention level based on the calculated feature amount.

[0049] As shown in FIG. 7 , the learning unit 234 acquires a learning image (step S10). The learning unit 234 determines whether the acquired image indicates that a performance is in progress (step S11). The learning unit 234 determines whether a performance is in progress based on, for example, the presence or absence of musical sounds associated with the image. Alternatively, the learning unit 234 may determine whether a performance is in progress based on whether a performer depicted in the image is performing a musical instrument. Note that the learning unit 234 may determine whether a performance is in progress not only based on the target image, which is the object of the determination of whether a performance is in progress, but also on images before and after the target image in chronological order. For example, there may be a case where a long break occurs during a performance, causing all the band members to stop moving and for the sound to become silent. If the performers' movements and the output of musical sounds resume after such a long break, the target image is determined to be "in progress."

[0050] If the image indicates that a drummer is playing, the learning unit 234 determines whether or not a drummer is captured in the image (step S12). For example, the learning unit 234 uses image recognition technology to identify a person captured in the image and determines whether or not a drummer is captured in the image.

[0051] If a drummer is captured in the image, the learning unit 234 calculates the degree of arm movement of the drummer (step S13). The learning unit 234 calculates the degree of arm movement of the drummer based on the amount of change in each of the consecutive frames. For example, the learning unit 234 calculates the degree of arm movement of the drummer based on the difference between the position of the drummer's arm in the previous frame image and the position of the drummer's arm in the current frame image. The learning unit 234 determines that the degree of arm movement is large when the difference is large. The learning unit 234 determines that the degree of arm movement is small when the difference is small.

[0052] Next, the learning unit 234 calculates the degree of eye movement of the drummer (step S14). For example, if the direction of the drummer's eye movement is toward the drum set, the learning unit 234 determines that the degree of eye movement is small. If the direction of the drummer's eye movement is different from the direction of the drum set, the learning unit 234 determines that the degree of eye movement is large.

[0053] Next, the learning unit 234 calculates the degree of movement of the upper body of the drum player (step S15). For example, the learning unit 234 calculates the degree of movement of the upper body of the drum player using a method similar to that used to calculate the degree of movement of the arms.

[0054] Next, the learning unit 234 calculates the difference between the degree of movement of the drummer and the degree of movement of the other players (step S16). For example, the learning unit 234 calculates the degree of movement of the drummer and the other players using a method similar to that used to calculate the degree of arm movement. The learning unit 234 calculates the difference between the calculated degrees of movement. In this case, if the degree of movement of the drummer is greater than the degree of movement of the other players, the learning unit 234 increases the difference.

[0055] The learning unit 234 then determines the level of attention according to the sum of the feature amounts calculated in each of steps S13 to S16. As a result, for example, a high level of attention can be associated with a scene in the learning image in which the drummer's arms move vigorously. Also, a high level of attention can be associated with a scene in which the drummer's line of sight is directed in a direction different from the drum set. A high level of attention can be associated with a scene in which the drummer's upper body moves vigorously. Also, a high level of attention can be associated with a scene in which there is no movement of any other performers other than the drummer, i.e., a solo performance. Furthermore, when these are combined, for example, a high level of attention can be associated with a scene in which the drummer moves his arms or upper body vigorously to perform a special or flashy performance, or a solo performance, etc.

[0056] An image determined in step S11 not to be in performance, or an image determined in step S12 not to include a drummer, is associated with a level of attention indicating that it receives little attention (for example, the lowest value).

[0057] Although the above description has been given taking an example in which steps S13 to S16 are performed in order, the order in which steps S13 to S16 are performed may be reversed.Furthermore, it is sufficient that at least one of steps S13 to S16 is performed.

[0058] The training data set is also associated with an attention level based on a feature amount obtained from the performance sound corresponding to the training image.

[0059] The feature quantity obtained from the sound being played is, for example, a feature quantity according to the rhythm. For example, the rhythm differs when the drums are playing a monotonous rhythm, when a fill-in is played, and when a "finishing" part is played. By determining the attention level based on the feature quantity according to the rhythm, it is possible to detect a scene in which a performance is being performed with a rhythm different from a monotonous rhythm as an image with a high attention level.

[0060] The feature obtained from the performance sound is, for example, a feature corresponding to the number of tones. For example, when drums are playing a monotonous rhythm, specific instruments such as a snare drum, a bass drum, and a hi-hat cymbal are often played. In this case, at least one of the tones of these specific instruments is output. In contrast, to add excitement to a song, a different tonality is output than when a monotonous rhythm is played. For example, the sound of a crash cymbal may be added, or a wind chime or tambourine may be added, or the sound may change so that it flows from the snare drum to the toms. Furthermore, the hi-hat cymbal may be played in an open position, or a ride cymbal may be used instead of striking the hi-hat cymbal. Alternatively, special sound effects may be output. By determining the attention level based on the feature corresponding to the number of tones, a scene in which a performance that adds the sound of a crash cymbal is being performed can be detected as an image with a high attention level.

[0061] The feature obtained from the sound of the performance is, for example, a feature corresponding to the loudness of the sound. For example, when a drumstick is struck with a full stroke, a louder sound is output compared to when a monotonous rhythm is played. By determining the attention level based on the feature corresponding to the loudness of the sound, a scene in which a performance is being performed that outputs a louder sound compared to when a monotonous rhythm is played can be detected as an image with a high level of attention.

[0062] The feature amount obtained from the sound being played is, for example, a feature amount according to the musical score. For example, fill-ins are often performed in measures before a change in the melody, such as in the A melody, B melody, or chorus. Some musical scores indicate the measures where fill-ins are to be placed. Furthermore, based on the musical score, it is possible to determine whether a measure has a monotonous rhythm, a fast rhythm, or a slow rhythm. By determining the attention level based on the feature amount according to the musical score, it is possible to detect scenes where a fill-in is likely to be played or scenes where a performance is performed with a rhythm different from a monotonous rhythm as images with a high attention level.

[0063] In this way, by calculating the feature quantities corresponding to the played sounds, it is possible to determine an appropriate attention level. Fig. 8 shows the flow of a process for determining the attention level based on the feature quantities corresponding to the played sounds.

[0064] 8, the learning unit 234 acquires sound information and musical score information of the performance sounds produced in the learning performance video (step S20). The sound information is, for example, information on sounds picked up by a microphone when the performance video is captured. The musical score information is information on the musical score corresponding to the performance sounds.

[0065] The learning unit 234 calculates a feature quantity corresponding to the rhythm being played based on the acquired sound information (step S21). For example, the learning unit 234 determines whether or not a drum tone is included in the sound information for a predetermined time, for example, for each time corresponding to a measure in the musical score. Whether or not a drum tone is included can be determined, for example, based on the frequency characteristics of the sound included in the sound information. The frequency characteristics of the sound can be calculated by frequency converting the sound information. The learning unit 234 determines the rhythm based on the number of times the drum tone is output within a predetermined time. Alternatively, the learning unit 234 may determine the rhythm based on the number of notes per measure shown in the musical score. The learning unit 234 uses the rhythm that is most frequently included in the entire song as a standard and calculates a feature quantity that will attract more attention when a rhythm that differs from the standard is played.

[0066] The learning unit 234 calculates a feature amount depending on whether a specific drum tone is being output (step S22). For example, the learning unit 234 determines the output tone based on the frequency characteristics of the sound included in the sound information. The learning unit 234 calculates a feature amount that will attract less attention when a tone used to keep a monotonous rhythm, such as a snare drum, bass drum, or hi-hat cymbal, is being output. On the other hand, the learning unit 234 calculates a feature amount that will attract more attention when a tone used to add excitement to a song, such as a crash cymbal, ride cymbal, open hi-hat cymbal, tambourine, or wind chime, is being output.

[0067] Here, the type of instrument used to create a monotonous rhythm and the type of instrument used to add excitement to a song may differ depending on the drummer. For this reason, the learning unit 234 may determine the timbre used to create a monotonous rhythm and the timbre used to add excitement to a song individually for each played note. For example, the learning unit 234 calculates a feature value that reduces the level of attention for a timbre that is output frequently throughout the song. The learning unit 234 calculates a feature value that increases the level of attention for a timbre that is used only a few times throughout the song.

[0068] The learning unit 234 calculates a feature value according to the number of tones used in the drum performance (step S23). For example, the learning unit 234 determines whether or not a drum tone is included in the sound information for a predetermined time, for example, for each time corresponding to a measure in a musical score, using a method similar to that of step S21. The learning unit 234 determines the number of drum tones output within the predetermined time. The learning unit 234 calculates a feature value that attracts more attention as the number of drum tones increases.

[0069] The learning unit 234 calculates a feature value according to the degree of rhythmic similarity between the drum timbre and the timbre of a non-drum instrument, such as a guitar, bass, or keyboard (step S24). For example, the learning unit 234 determines, using a method similar to step S21, whether or not the sound information for a predetermined time, for example, for each time corresponding to a measure in a musical score, includes a drum timbre and whether or not a non-drum timbre is included. The learning unit 234 calculates the rhythm of the drum timbre and the rhythm of the non-drum timbre for a time section that includes both the drum timbre and the non-drum timbre. The method for calculating the rhythm can be the same as step S21. When the rhythm of the drum timbre matches the rhythm of the non-drum timbre, the learning unit 234 calculates a feature value that attracts attention. This makes it possible to calculate a feature value that attracts attention when a "key moment" is played in which the drummer and another performer play the same rhythm at the same time. On the other hand, when the rhythm of the drum tone and the rhythm of the tone other than the drum tone do not match, the learning unit 234 calculates a feature amount that reduces the degree of attention.

[0070] The learning unit 234 calculates a feature amount according to whether or not the performance is for a measure before a change in musical style based on the musical score information (step S25). The learning unit 234 determines the musical style for each measure based on the musical score information. For example, if the musical score information describes musical styles such as the A melody, B melody, and chorus, the learning unit 234 determines the musical style based on this description. Alternatively, the learning unit 234 may determine the musical style using conventional techniques such as those described in prior art documents. The learning unit 234 extracts the measure before the change in musical style and calculates a feature amount that will draw more attention to the part where the performance shown in the extracted measure is performed.

[0071] The learning unit 234 then determines the attention level according to the total value of the values ​​calculated in steps S21 to S25. The learning unit 234 uses the determined attention level as a corresponding attention level. As a result, for example, a high attention level can be associated with a scene in which the drum performance sound in the learning image is played with a rhythm different from a monotonous rhythm, such as a faster rhythm or an irregular rhythm. A high attention level can be associated with a performance image showing a scene in which a specific sound, such as a wind chime, is output. A high attention level can be associated with a performance image showing a scene in which a luxurious sound is output by adding more cymbal tones or a tambourine. A high attention level can be associated with a performance image showing a scene in which a "finishing" part is played. Furthermore, a high attention level can be associated with a performance image showing a scene in which a measure before a melody change is played, which suggests that a fill-in will be inserted. Furthermore, when these are combined, a high attention level can be associated with, for example, a scene in which a fill-in is played with a rhythm different from a monotonous rhythm in a measure before a melody change.

[0072] Although the above description has been given taking an example in which steps S21 to S25 are performed in order, the order in which steps S21 to S25 are performed may be reversed.Furthermore, it is sufficient that at least one of steps S21 to S25 is performed.

[0073] Here, the flow of processing performed by the signal processing system 1 will be described with reference to Fig. 9. Fig. 9 is a sequence diagram showing the flow of processing performed by the signal processing system 1.

[0074] The user terminal 10 captures a performance video (step S30). The user terminal 10 transmits the captured performance video to the signal processing device 20.

[0075] The signal processing device 20 acquires a performance video by receiving the performance video from the user terminal 10 (step S31). The signal processing device 20 estimates the attention level of each performance image included in the acquired performance video (step S32). The signal processing device 20 selects performance images to be used for editing according to the estimated attention levels (step S33). The signal processing device 20 generates a moving image using the selected performance images (step S34). The signal processing device 20 transmits the generated moving image to the user terminal 10.

[0076] As described above, the signal processing device 20 in the embodiment includes the image acquisition unit 230, the estimation unit 231, and the output unit 233. The image acquisition unit 230 acquires a performance image captured so as to include a drum performance. The estimation unit 231 estimates an attention level, which is the degree to which the drum performance in the performance image is attracting attention, based on feature amounts obtained from the performance image. The output unit 233 outputs the attention level estimated by the estimation unit 231. This allows the signal processing device 20 in the embodiment to estimate the attention level for the performance shown in the image.

[0077] Furthermore, in the signal processing device 20 of the embodiment, the estimation unit 231 estimates the attention level using a trained model. The trained model is a model generated by machine learning a training dataset in which training images including drum performances are associated with the attention levels of the training images. The trained model is a model trained to output the attention level of an input image. As a result, the signal processing device 20 of the embodiment can easily estimate the attention level using the trained model.

[0078] Furthermore, in the signal processing device 20 of the embodiment, the training data set is associated with an attention level based on feature amounts corresponding to the movements of the drummer shown in the training image. As a result, the signal processing device 20 of the embodiment can associate a greater attention level with, for example, a scene in which the drummer is performing a solo with large movements of his arms and upper body.

[0079] Furthermore, in the signal processing device 20 of the embodiment, the training data set is associated with an attention level based on a feature depending on whether or not a specific drum tone is included in the performance sound corresponding to the training image. As a result, the signal processing device 20 of the embodiment can associate a high attention level with a performance image showing a scene in which a specific sound, such as a wind chime, is output to liven up the song.

[0080] Furthermore, in the signal processing device 20 of the embodiment, the training data set is associated with an attention level based on a feature amount corresponding to the number of drum tones included in the performance sound corresponding to the training image. As a result, the signal processing device 20 of the embodiment can associate a high attention level with a performance image showing a scene in which a gorgeous sound is output by increasing the number of cymbal tones or adding a tambourine, for example.

[0081] Furthermore, in the signal processing device 20 of the embodiment, the training data set is associated with an attention level based on a feature amount corresponding to the degree of similarity between the performance sounds output from the drum-related timbres included in the performance sounds corresponding to the training images and the timbres unrelated to the drums. This allows the signal processing device 20 of the embodiment to associate a high attention level with a performance image showing a scene in which a "finishing" part is being played, for example.

[0082] Furthermore, in the signal processing device 20 of the embodiment, the training data set is associated with an attention level based on a feature obtained from the drum score information corresponding to the training image. As a result, the signal processing device 20 of the embodiment can associate a high attention level with a performance image showing a scene in which a measure before a change in melody is being played, which suggests that a fill-in will be inserted, for example.

[0083] Moreover, the signal processing device 20 in the embodiment further includes an editing unit 232. The editing unit 232 generates a moving image using a plurality of performance images. The editing unit 232 selects images having a score equal to or greater than a threshold from the plurality of performance images based on a score according to the attention level estimated by the estimation unit 231. The estimation unit 231 generates a moving image using the selected images. In this way, the signal processing device 20 in the embodiment can generate a moving image including performance images with high attention levels.

[0084] Furthermore, in the signal processing device 20 of the embodiment, the image acquisition unit 230 acquires a plurality of images of a drum performance captured from different directions. The editing unit 232 identifies a plurality of performance images captured at the same time. The editing unit 232 selects, from the plurality of performance images, images whose scores are equal to or greater than a threshold. The editing unit 232 generates a moving image using the selected images. In this way, the signal processing device 20 of the embodiment can generate a moving image that includes, from the plurality of images captured from different directions, performance images that are attracting a lot of attention.

[0085] (Modification of the embodiment) A modification of the embodiment will now be described. In this modification, the estimation unit 231 estimates the attention level using a rule-based model (rather than a trained model). This model derives the attention level from an image based on a set of rules predetermined by experts or the like. The set of rules is associated with an attention level corresponding to the scene depicted in the image. For example, in a drum performance, a relatively high attention level is associated with a scene in which the performer moves significantly, a scene in which a fill-in is played, a scene in which the drumsticks are struck with a full stroke, a scene in which the drums are played solo, etc. On the other hand, a relatively low attention level is associated with a scene in which the rhythm is played monotonously, a scene in which the performance has not yet started, or a scene in which the performance has ended, etc.

[0086] The estimation unit 231 estimates the level of attention in the performance image, for example, using a method similar to the method used by the learning unit 234 to determine the level of attention. For example, the estimation unit 231 estimates the level of attention based on the movement of the drum player shown in the performance image. The estimation unit 231 estimates the level of attention based on whether a specific drum tone is included in the performance sound played in the performance video. The estimation unit 231 estimates the level of attention based on the number of drum tones included in the performance sound. The estimation unit 231 estimates the level of attention based on the degree of similarity between the performance sounds output from the drum-related tones included in the performance sound and the tones unrelated to drums. As a result, in the modified embodiment, the level of attention can be quantitatively estimated based on predetermined rules.

[0087] The estimation unit 231 may also estimate the level of attention based on features obtained from musical score information. In this case, the signal processing device 20 acquires, for example, from the user terminal 10, musical score information of the song being played in the performance video together with the performance video. The estimation unit 231 estimates the level of attention using the musical score information acquired by the signal processing device 20. As a result, in the modified example of the embodiment, the level of attention can be estimated using the musical score information.

[0088] In the above-described embodiment, the signal processing device 20 may be configured to execute all of the functions performed by the signal processing system 1, and to display the processing results executed by the signal processing device 20, i.e., the estimation results of the attention level. In this case, for example, the user terminal 10 captures a performance video and transmits the captured video to the signal processing device 20. The signal processing device 20 estimates the attention level of a performance image constituting the performance video received from the user terminal 10 and transmits the estimation results to the user terminal 10. The user terminal 10 receives the estimation results from the signal processing device 20 and displays the received estimation results. In such a configuration, the user terminal 10 does not need to store a program related to the processing of estimating the attention level in the storage unit 12. In other words, the program related to the processing of estimating the attention level is stored in the storage unit 22 of the signal processing device 20. In this case, the user terminal 10 may omit the storage unit 12.

[0089] Furthermore, the functions performed by the signal processing system 1 may be realized by the signal processing device 20 and a computer other than the signal processing device 20. That is, the process of estimating the attention level in an image, which is a function performed by the signal processing system 1, may be executed by one or more computers.

[0090] Furthermore, the method for generating a trained model according to the embodiment is a generation method performed by a computer, in which the training unit 234 generates the trained model by having the training model perform machine learning on a training dataset. The training dataset is information in which training images including drum performances are associated with the attention levels of the training images. By having the trained model perform machine learning on the training dataset, when an image is input to the trained model, the trained model can output the attention level of the input image. Therefore, the signal processing device 20 in the embodiment can generate a trained model that can estimate the attention level of an image.

[0091] In the above-described embodiment, the "learning stage" and the "execution stage" are both executed by one computer (e.g., the signal processing device 20). Here, the "learning stage" refers to a stage in which a learning model is trained, specifically, a stage in which the learning unit 234 generates a trained model. The "execution stage" refers to a stage in which estimation is performed using the trained model, specifically, a stage in which the estimation unit 231 estimates the attention level in an image using the trained model. However, this is not limited to this. The "learning stage" and the "execution stage" may each be executed by different computers. For example, the "learning stage" may be configured to be executed by a learning server that is a computer different from the signal processing device 20. In this case, for example, information on the trained model generated by the learning server is transmitted to the signal processing device 20 and stored as the trained model 222 in the storage unit 22 of the signal processing device 20. The signal processing device 20 then executes the "execution stage" by performing estimation using the trained model based on the trained model 222 stored in the storage unit 22.

[0092] The signal processing system 1 and the signal processing device 20 in the above-described embodiment may be implemented in whole or in part by a computer. In this case, a program for implementing the functions may be recorded on a computer-readable recording medium, and the program may be loaded into a computer system and executed. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into a computer system. Furthermore, the term "computer-readable recording medium" may also include devices that dynamically store programs for a short period of time, such as communication lines used when transmitting programs via networks such as the Internet or telephone lines, or devices that store programs for a fixed period of time, such as volatile memory within a computer system that serves as a server or client. The program may be designed to implement some of the functions described above, or may be capable of implementing the functions in combination with a program already stored in the computer system, or may be implemented using a programmable logic device such as an FPGA.

[0093] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention described in the claims and their equivalents. [Explanation of symbols]

[0094] 1... signal processing system, 10... user terminal, 20... signal processing device, 230... image acquisition unit, 231... estimation unit, 232... editing unit, 233... output unit, 234... learning unit

Claims

1. an image acquisition unit that acquires a performance image captured so as to include a drum performance; an estimation unit that estimates an attention level by inputting the performance image into a trained learning model that has undergone machine learning to estimate an attention level, which is a degree to which the drum performance in the performance image is attracting attention, based on a feature amount related to the drum performance obtained from the performance image; and an output unit that outputs the attention level estimated by the estimation unit; Equipped with The learning data to be machine-learned by the learning model is associated with the attention level based on a feature amount corresponding to a degree of similarity between the rhythm of a drum tone included in the performance sound corresponding to the learning image and the rhythm of a performance sound output from a tone of an instrument other than a drum. Signal processing device.

2. an image acquisition unit that acquires a performance image captured so as to include a drum performance; an estimation unit that estimates an attention level by inputting the performance image into a trained learning model that has undergone machine learning to estimate an attention level, which is a degree to which the drum performance in the performance image is attracting attention, based on a feature amount related to the drum performance obtained from the performance image; and an output unit that outputs the attention level estimated by the estimation unit; Equipped with The learning data to be machine-learned by the learning model includes a feature quantity, which is determined using the musical score information corresponding to the learning image, according to whether or not the performance corresponds to the measure before the melody changes. The attention level based on Signal processing device.

3. further comprising an editing unit that generates a moving image using the plurality of performance images; the editing unit selects images having a score equal to or greater than a threshold from the plurality of performance images based on the score according to the attention level estimated by the estimation unit, and generates the moving image using the selected images.

3. The signal processing device according to claim 1 or 2.

4. the image acquisition unit acquires a plurality of performance images capturing images of drum performances, the editing unit identifies images captured at the same time from the plurality of performance images acquired by the image acquisition unit, selects images from the identified images having scores equal to or greater than a threshold, and generates a moving image using the selected images. The signal processing device according to claim 3 .

5. an image acquisition unit that acquires a performance image captured so as to include a drum performance; an estimation unit that estimates an attention level, which is a level of attention that the drum performance in the performance image is receiving, based on a feature amount related to the drum performance obtained from the performance image; an output unit that outputs the attention level estimated by the estimation unit; Equipped with the estimation unit estimates the attention level based on a feature quantity relating to the drum performance, the feature quantity corresponding to a degree of similarity in rhythm between a drum timbre included in the performance sound corresponding to the performance image and a timbre of a musical instrument other than the drum; Signal processing device.

6. an image acquisition unit that acquires a performance image captured so as to include a drum performance; an estimation unit that estimates an attention level, which is a level of attention that the drum performance in the performance image is receiving, based on a feature amount related to the drum performance obtained from the performance image; an output unit that outputs the attention level estimated by the estimation unit; Equipped with the estimation unit estimates the attention level based on a feature amount relating to the drum performance, the feature amount being determined using musical score information corresponding to the performance image, and indicating whether the performance corresponds to a measure before a change in melody. Signal processing device.

7. Acquire a performance image captured so as to include a drum performance; an attention level, which is a level of attention that the drum performance in the performance image is drawing, based on a feature amount related to the drum performance obtained from the performance image; outputting the estimated attention level; In the estimating step, the attention level is estimated based on a feature quantity relating to the drum performance, the feature quantity corresponding to the degree of similarity in rhythm between the timbre of the drum included in the performance sound corresponding to the performance image and the timbre of a musical instrument other than the drum. Signal processing methods.

8. Acquire a performance image captured so as to include a drum performance; an attention level, which is a level of attention that the drum performance in the performance image is drawing, based on a feature amount related to the drum performance obtained from the performance image; outputting the estimated attention level; In the estimating step, the attention level is estimated based on a feature amount relating to the drum performance, which is determined using musical score information corresponding to the performance image and indicates whether the performance corresponds to a measure before a change in melody. Signal processing methods.

Citation Information

Patent Citations

  • Information processing system, method and program for controlling graphic display

    JP2004127019A

  • Communication karaoke system

    JP2011175170A

  • Moving image combining system, moving image combining method, moving image combining program and storage medium of the same

    JP2012182724A