Image capture area determination device, image capture area learning device, and their programs
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NIPPON HOSO KYOKAI
- Filing Date
- 2022-10-27
- Publication Date
- 2026-08-05
AI Technical Summary
【0018】 本発明によれば、装置の処理時間によるカメラワークの遅延を抑制することができる。
Smart Images

Figure 0007901000000007 
Figure 0007901000000008 
Figure 0007901000000009
Abstract
Description
Technical Field
[0001] The present invention relates to a shooting area determination device, a shooting area learning device, and programs thereof for determining a shooting area of a sports video.
Background Art
[0002] To acquire know-how regarding camera operation in sports program production, rich experience in relay programs is required. However, the human resources of experienced cameramen are limited. Also, at the relay site, low-cost, efficient, and high-quality program production is required. For these reasons, automation of camera work by an AI robot camera equipped with know-how of camera operation is required.
[0003] Non-Patent Document 1 describes a technique for automatically generating camera work by collecting and learning operation information of a cameraman involved in a relay program and information of a shooting target (for example, a player or a ball). In the technique described in this Non-Patent Document 1, a formation map of a measurement target is used for automatically generating camera work. However, it is considered that simply using the formation map of the measurement target does not reach the camera work by an experienced cameraman.
[0004] Therefore, the technique described in Non-Patent Document 2 has been proposed. In the technique described in this Non-Patent Document 2, when automatically generating camera work, in addition to the formation map of the measurement target, by using an event during the competition and past camera work, it is possible to automatically generate more natural and suitable camera work.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, in the technology described in Non-Patent Document 2, when automatically generating camera work, there is a problem that the camera work is delayed due to the processing time of the device.
[0007] The present invention has been made in view of the above problems, and an object thereof is to provide a shooting area determination device, a shooting area learning device, and programs thereof that can suppress the delay of camera work due to the processing time of the device.[[ID=1??]]
Means for Solving the Problems
[0008] In order to solve the above problems, a shooting area determination device according to the present invention is a shooting area determination device that determines a shooting area of a camera that shoots a sports video using a formation map representing at least the position or speed of each subject in the field of the sports video, and includes a convolutional part, a recursive processing part, a combining part, and a determination part.
[0009] According to such a configuration, the convolutional part generates a column vector of the formation map by convolutional the formation map with a pre-trained convolutional neural network. The recursive processing part generates a vector of shooting area candidates continuous in the time direction by inputting the prediction result of the shooting area at a time earlier than the formation map into a pre-trained recurrent neural network.
[0010] The combining part generates a combined vector by combining the column vector of the formation map and the vector of the shooting area candidates. The decision unit determines the imaging area by inputting connection vectors into a pre-trained neural network for determining the imaging area.
[0011] Thus, since the imaging area determination device utilizes a recurrent neural network that has learned the prediction results of future imaging areas, it can determine the imaging area in a way that takes into account the delay caused by the device's processing time.
[0012] Furthermore, in order to solve the above problems, the shooting area learning device according to the present invention is a shooting area learning device that learns to determine the shooting area of a camera shooting ball game video using a formation map that represents at least the position or velocity of each subject within the field of the ball game video, and comprises an operation information conversion unit, a learning data storage unit, an input shooting area output unit, a convolution unit, a recursive processing unit, a coupling unit, a determination unit, an error calculation unit, and a learning unit.
[0013] In this configuration, the operation information conversion unit receives camera operation information and converts the input operation information into an input shooting area according to predetermined conversion rules. The learning data storage unit stores, as learning data, the formation map and the input shooting area for each time period, obtained when actual ball game footage is filmed. The input shooting area output unit outputs the input shooting area from the learning data storage unit at a time point prior to the formation map.
[0014] The convolutional section generates column vectors of the formation map by convolving the formation map of the learning data storage section with a convolutional neural network. The recursive processing unit inputs the input imaging region output by the input imaging region output unit into a recursive neural network, thereby generating a vector of continuous imaging region candidates in the time direction. The joining section generates a combined vector by combining the column vectors of the formation map and the vectors of the candidate shooting areas.
[0015] The decision unit determines the imaging area by inputting connection vectors into a neural network for determining the imaging area. The error calculation unit calculates the error of the shooting area determined by the determination unit. The learning unit trains the convolutional neural network, recurrent neural network, and neural network for determining the imaging area so that the error calculated by the error calculation unit is minimized.
[0016] In this way, the imaging region learning device learns predictions of future imaging regions, allowing it to incorporate the processing time delay of the device into the candidate imaging regions determined by the recurrent neural network.
[0017] Furthermore, the present invention can also be realized by a program that causes a computer to function as the aforementioned imaging area determination device or imaging area learning device. [Effects of the Invention]
[0018] According to the present invention, delays in camera work caused by the processing time of the device can be suppressed. [Brief explanation of the drawing]
[0019] [Figure 1] This is an explanatory diagram illustrating the overview of the camera work prediction and generation device according to the first embodiment. [Figure 2] This is a block diagram showing the configuration of the camera work prediction and generation device according to the first embodiment. [Figure 3] This image shows an example of a formation map in the first embodiment. [Figure 4] In the first embodiment, this is an image showing the normalization of the formation map. [Figure 5] This is an explanatory diagram illustrating an example of a convolutional neural network in the first embodiment. [Figure 6A] This is an explanatory diagram illustrating the camera's shooting area in the first embodiment. [Figure 6B]This is an explanatory diagram illustrating the coordinate system of the imaging area in the first embodiment. [Figure 7] This is an explanatory diagram illustrating an example of a neural network for determining the imaging area in the first embodiment. [Figure 8] This is a block diagram showing the configuration of a camera work prediction learning device according to the first embodiment. [Figure 9] This is an explanatory diagram illustrating the processing of the camera work prediction learning device in the first embodiment. [Figure 10] This is a flowchart showing the operation of the camera work prediction and generation device according to the first embodiment. [Figure 11A] This is a flowchart showing the operation of the camera work prediction learning device according to the first embodiment. [Figure 11B] This is a flowchart showing the operation of the camera work prediction learning device according to the first embodiment. [Figure 12] This is a block diagram showing the configuration of the camera work prediction and generation device according to the second embodiment. [Figure 13] This is a block diagram showing the configuration of a camera work prediction learning device according to the second embodiment. [Modes for carrying out the invention]
[0020] The embodiments of the present invention will be described below with reference to the drawings. However, the embodiments described below are intended to embody the technical concept of the present invention, and unless otherwise specified, the present invention is not limited to these embodiments. In addition, the same reference numerals are used for the same means, and their descriptions may be omitted.
[0021] (First Embodiment) [Overview of the Camera Work Prediction and Generation System] Referring to Figure 1, the camera work prediction and generation device (shooting area determination device) 1 according to the first embodiment will be described. The camera work prediction and generation device 1 determines the shooting area of the robot camera 4 that captures the ball game footage, using a formation map that represents at least the position or velocity of each subject within the field of the ball game footage.
[0022] From now on, "ball game footage" will be explained as referring to soccer footage, which is footage of a soccer match. Furthermore, each neural network used in the camera work prediction and generation device 1 is assumed to have been pre-trained in the camera work prediction learning device 4 (Figure 8), which will be described later.
[0023] Figure 1 shows the camera work prediction and generation device 1, as well as the camera 2, the formation map generation unit 3, and the robot camera 4. Camera 2 is a camera that captures an overhead view of the entire soccer field 20. For example, Camera 2 could be a sensor camera. Camera 2 then outputs the captured soccer video to the formation map generation unit 3.
[0024] The formation map generation unit 3 generates a formation map from soccer footage captured by camera 2. Here, the formation map generation unit 3 generates the formation map using known methods (see Non-Patent Literature 1). Specifically, when generating the formation map, background subtraction and chroma keying are used for subject extraction from the soccer footage, an extended Kalman filter is used for player tracking, a particle filter is used for ball tracking, and an 11-layer convolutional neural network is used for face orientation estimation.
[0025] The robot camera 4 is a camera that captures soccer footage based on camera control information input from the camera work prediction and generation device 1. For example, the robot camera 4 comprises a camera body having a shooting lens and an image sensor, and an automatic pan / tilt head on which the camera body is mounted. The robot camera 4 then changes its position and orientation according to the camera control information from the camera work prediction and generation device 1.
[0026] <Formation Map> The following is a concrete example of a formation map. A formation map represents at least the position or velocity of each subject within the field of a soccer video. For example, in a soccer video, subjects could be the players of each team or the ball. In other words, a formation map is a map that represents the characteristics of multiple subjects (e.g., players, ball) and is generated sequentially from the start to the end of a soccer match. Here, the formation map may represent the positional distribution and velocity distribution of each subject within the field of the soccer video. Furthermore, the formation map may represent the gaze area and acceleration of each subject (each player) within the field of the soccer video.
[0027] For example, the method described in Japanese Patent Publication No. 6596804 can be used to obtain the information necessary for generating a formation map. When generating a formation map corresponding to soccer footage, the following data can be used for each frame of the soccer footage.
[0028] Player's Team: The team to which each player belongs. This refers to the distinction between teams that attack to the right of the field or teams that attack to the left of the field. Player position: The position of a player on the field. For example, a player's position is expressed as a number of meters along the long and short sides of the field, with the center as the center. Ball position: The position of the ball on the field. The ball's position can be described in the same way as the players' positions. Athlete's speed: The direction and speed at which the athlete is running. Player's face direction: The direction the player's face is pointing.
[0029] In the example shown in Figure 3, the formation map is divided into two halves: the upper half shows the positional distribution and face direction of each player, and the lower half shows the velocity distribution of each player and the positional distribution of the ball. Each pixel in the formation map is represented by an RGB value between 0 and 255 (256 levels).
[0030] The upper part of Figure 3 illustrates the positional distribution of each subject included in the formation map. In the upper part of Figure 3, the positional distribution of players on the team attacking to the right of the field is shown in red, the positional distribution of players on the team attacking to the left of the field is shown in green, and the direction of the players' faces is shown in blue. Note that the positional distribution of each player is represented by a circle of the same radius, but this may not be the case in reality. For example, when moving in the same direction for a long period of time, the formation map may not show a player at the center of the circle representing the player's positional distribution, or the area representing the player's positional distribution may not be perfectly circular.
[0031] The lower part of Figure 3 illustrates the velocity distribution of each subject included in the formation map. In the lower part of Figure 3, the velocity distribution of players of the team attacking to the right of the field is shown in red, the velocity distribution of players of the team attacking to the left of the field is shown in green, and the ball's position information is shown in blue. Note that there are many false detections of the ball's position, so it is shown in blue in multiple places. In addition, each pixel in the player's velocity distribution represents 1 meter / second.
[0032] [Configuration of the camera work prediction and generation device] Referring to Figure 2, the configuration of the camera work prediction and generation device 1 will be explained in detail. As shown in Figure 2, the camera work prediction generation device 1 comprises a convolutional unit 10, a play status determination unit 11, an LSTM unit (recurrent processing unit) 12, a coupling unit 13, a decision unit 14, and a camera control unit 15. It should be assumed that each neural network used by the camera work prediction generation device 1 has been trained by the camera work prediction learning device 4 described later.
[0033] The convolutional unit 10 generates column vectors of the formation map by convolving the formation map with a pre-trained convolutional neural network. Here, the formation map is input to the convolutional unit 10 from the formation map generation unit 3. The convolutional unit 10 then normalizes the formation map and then convolves it with the convolutional neural network.
[0034] As shown in Figure 4, the convolution unit 10 normalizes the size of the formation map and the RGB values of each pixel. Generally, a soccer field is 105 x 68 meters. Therefore, the convolution unit 10 normalizes the size of the position distribution and velocity distribution of the formation map to 106 x 69 pixels, with one pixel of the formation map representing one meter, leaving a margin of one pixel. The convolution unit 10 also normalizes the RGB values of each pixel of the formation map to a value between 0 and 1.
[0035] For example, the convolutional section 10 can utilize a residual neural network (ResNet) as the convolutional neural network (see, for example, Reference 1). Reference 1: He etc, “Deep Residual Learning for Image Recognition”, Computer Vision and Pattern Recognition, 2016, pp.770-pp.778
[0036] <Convolutional Neural Network> Figure 5 illustrates a specific example of a convolutional neural network 100. As shown in Figure 5, the convolutional neural network 100 comprises, in order from the input side to the output side, a first convolutional layer 110, a first pooling layer 120, a second convolutional layer 130, a third convolutional layer 140, and a second pooling layer 150.
[0037] The first convolutional layer 110 convolves a formation map 111 of size 128×128×3 (3D array) into data of size 56×56×64. At this time, the first convolutional layer 110 applies an activation function (e.g., ReLU) every 2×2 stride to the numerical values calculated using an equation with the weight coefficients of the kernel 112. This kernel 112 has a range of 7×7 vertically and horizontally. The first convolutional layer 110 then outputs the convolution result to the first pooling layer 120.
[0038] Here, the three-dimensional array "128×128×3" of the formation map 111 represents "vertical size (number of pixels) of the formation map" × "horizontal size (number of pixels) of the formation map" × "RGB value of each pixel". In other words, the formation map 111 input to the first convolutional layer 110 is the input formation map resized to 128×128. Also, the stride 2×2 indicates that the 7×7 kernel 112 is convolved by shifting it horizontally or by 2 horizontally. Furthermore, the activation function ReLU is a function that outputs values greater than or equal to 0 as they are, and outputs values less than 0 by replacing them with 0.
[0039] The first pooling layer 120 converts the 64x64x64 size data 121 into 32x32x64 size data by pooling (maximum pooling). At this time, the kernel 122 is in the range of 3x3 and the stride is in the range of 2x2. The first pooling layer 120 then outputs the pooling result to the second convolutional layer 130.
[0040] The second convolutional layer 130 convolves the data 131, which is 32 × 32 × 64 in size, into data 16 × 16 × 128 in size. At this time, the second convolutional layer 130 applies an activation function (e.g., ReLU) every 2 × 2 stride to the numerical value calculated using the weight coefficients of the kernel 132. This kernel 132 is in a 3 × 3 area vertically and horizontally. The second convolutional layer 130 then outputs the convolution result to the third convolutional layer 140.
[0041] The third convolutional layer 140 convolves the 16×16×128 size data 141 into 8×8×25 size data. At this time, the third convolutional layer 140 applies an activation function (e.g., ReLU) every 2×2 stride to the numerical value calculated using the weight coefficients of the kernel 142. This kernel 142 has a 3×3 area vertically and horizontally. The third convolutional layer 140 then outputs the convolution result to the second pooling layer 150.
[0042] The second pooling layer 150 transforms the 8x8x256 size data 151 into a 1x1x256 size (i.e., a column vector with 256 elements) by performing pooling (global average pooling). The second pooling layer 150 then outputs the column vector 152, which is the convolution result of the formation map, to the play status determination unit 11 and the merging unit 13.
[0043] It goes without saying that the convolutional neural network 100 in the convolutional section of Figure 5 is just one example and is not limited to it.
[0044] Returning to Figure 2, we continue the explanation of the camera work prediction and generation device 1. The play status determination unit 11 generates play status vectors representing events that occurred in the soccer video by inputting column vectors of the formation map into a pre-trained neural network for play status determination. Here, the play status determination unit 11 receives column vectors of the formation map from the convolution unit 10.
[0045] In other words, the play status determination unit 11 determines the events (play status) that occur in a soccer match and generates a play status vector representing the determined events. For example, the play status determination unit 11 can use a neural network capable of multi-class classification as the neural network for play status determination (for example, a multilayer perceptron).
[0046] The connection structure of each node in the neural network for determining the play status and the weight coefficients set for each node are calculated by the camera work prediction learning device 4 described later. In other words, the play status determination unit 11 determines the play status by inputting the column vector of the formation map into the trained neural network for determining the play status. After that, the play status determination unit 11 outputs the generated play status vector to the connection unit 13.
[0047] In this embodiment, the play status vector is a column vector in one-hot representation. For example, if the play status vector represents events that occur in a soccer match, such as in play, goal kick, corner kick, free kick, and throw-in, it will be a column vector in one-hot representation corresponding to each event, as shown below.
[0048] Inplay = [1,0,0,0,0] Goal kick = [0,1,0,0,0] Corner kick = [0,0,1,0,0] Free kick = [0,0,0,1,0] Throw-in = [0,0,0,0,1]
[0049] In addition, the play status vector may represent whether play is ongoing or stopped in a soccer match, or it may be detailed information that divides a corner kick into four sections according to its position on the field.
[0050] The LSTM unit 12 generates a vector of continuous candidate shooting areas in the time direction by inputting the prediction results of the shooting areas at time points ahead of the formation map into a pre-trained recurrent neural network.
[0051] Here, the LSTM unit 12 receives the prediction result (vector) of the shooting area from the decision unit 14, which will be described later. In other words, the prediction result of the shooting area input to the LSTM unit 12 represents a vector of the shooting area from the current time onward, estimated from the formation map prior to the current time. For example, consider the case where the decision unit 14, which will be described later, estimates the prediction result of the shooting area 1 second later. In this case, the LSTM unit 12 finds candidate shooting areas from the current time to 1 second ahead, which were estimated in the past 1 second.
[0052] Specifically, the LSTM unit 12 uses vectors of imaging regions at regular intervals as input data. For example, the input data includes a vector of the imaging region at the current time estimated using a formation map from 1 second ago, a vector of the imaging region 0.1 seconds ahead estimated using a formation map from 0.9 seconds ago, a vector of the imaging region 0.2 seconds ahead estimated using a formation map from 0.8 seconds ago, ..., and a vector of the imaging region 0.9 seconds ahead estimated using a formation map from 0.1 seconds ago.
[0053] For example, the LSTM unit 12 can utilize LSTM (Long Short-Term Memory) as a recurrent neural network (see, for example, Reference 2). LSTM is a type of recurrent neural network (RNN) that handles time series and overcomes the vanishing gradient problem inherent in RNNs. By using this LSTM, it is possible to predict future imaging area candidates from imaging areas that change over time. Reference 2: S. Hochreiter etc, “Long Short-Term Memory”, Neural Computation, 1997, pp. 1735-1780
[0054] The LSTM unit 12 has a recursive structure in which it outputs the generated vector of candidate imaging areas to the coupling unit 13, and the prediction result of the imaging area is input from the determination unit 14.
[0055] The merging unit 13 generates a combined vector by combining the column vectors of the formation map, the vectors of the candidate shooting area, and the play status vectors of the soccer video. Here, the merging unit 13 receives the column vectors of the formation map from the convolution unit 10, the play status vectors from the play status determination unit 11, and the vectors of the candidate shooting area from the LSTM unit 12.
[0056] In other words, the coupling unit 13 combines the predicted shooting area considering the time direction, the column vector of the convolved formation map, and the play situation of the soccer video to create an element for predicting the shooting area. As a result, the coupling unit 13 suppresses large discrepancies in the prediction results of the shooting area in the time direction, and obtains results with temporal continuity. Subsequently, the coupling unit 13 outputs the generated combined vector to the determination unit 14.
[0057] The determination unit 14 determines the imaging area by inputting a connected vector into a pre-trained neural network for determining the imaging area. Here, the connected vector is input to the determination unit 14 from the connected unit 13.
[0058] The shooting area determined by the determination unit 14 represents a predicted shooting area that is a predetermined time later than the time on the formation map. Figure 6A illustrates a soccer field 20. As shown in Figure 6A, the shooting area of the robot camera 4 is described in a coordinate system where the center of the field 20 (center of the center circle) is the origin K, the longitudinal direction of the field 20 is the X axis, and the transverse direction of the field 20 is the Y axis. In this case, as shown in Figure 6B, the shooting area of the robot camera 4 is a column vector [C] with 3 columns. x ,C y ,C s It is represented by ]. C x This is the horizontal coordinate of the intersection point between the camera optical axis and field 20. y This is the perpendicular coordinate of the intersection point between the camera optical axis and field 20. s This represents the size of the shooting area, and is half the horizontal length of the shooting area.
[0059] Here, the determination unit 14 can use a known neural network (e.g., a multi-layer perceptron) as the neural network for determining the shooting area. FIG. 7 illustrates an example of the neural network 200 for determining the shooting area. As shown in FIG. 7, the neural network 200 for determining the shooting area is a three-layer neural network and includes an input layer 201, an intermediate layer 202, and an output layer 203. Each of the input layer 201, the intermediate layer 202, and the output layer 203 is composed of a plurality of nodes 204. For example, the input layer 201 is composed of nodes 204 corresponding to each element of the combined vector. Also, the output layer 203 is composed of three nodes 204 corresponding to each element C x , C y , C s in the shooting area. The connection structure of the nodes 204 and the parameters (weights, biases) set for each node 204 are calculated by the camera operation prediction learning device described later from the operation results of an experienced cameraman.
[0060] After that, the determination unit 14 outputs the column vector [C x , C y , C s of the shooting area obtained by the output layer 203 to the camera control unit 15. Further, the determination unit 14 outputs the column vector [C<000**********]] of the shooting area to the LSTM unit 12 as the prediction result of the shooting area.
[0061] In this way, the determination unit 14 can suitably determine the future shooting area by using, in addition to the column vector of the formation map, the play situation vector from the play situation determination unit 11 and the vector of the shooting area candidates from the LSTM unit 12. Even if the formation map changes significantly due to noise, the determination unit 14 determines the shooting area according to the past estimation results, so temporal variations can be suppressed.
[0062] Returning to FIG. 2, the description of the camera operation prediction generation device 1 will be continued. The camera control unit 15 converts the shooting area determined by the determination unit 14 into a camera movement based on a predetermined conversion rule, and controls the robot camera 4 to shoot soccer footage according to the converted camera movement.
[0063] Here, the camera control unit 15 controls the robot camera 4 based on the shooting area input from the determination unit 14. For example, the camera control unit 15 determines the column vector of the shooting area [C x ,C y ,C s The camera control unit 15 converts the position of the robot camera 4 (pan angle, tilt angle, and field of view) into the posture of the robot camera 4. Furthermore, the camera control unit 15 refers to a correspondence table between the field of view and zoom demand and converts the field of view into the rotation amount of the zoom demand. In addition, the camera control unit 15 converts the installation position of the robot camera 4 and the camera's gaze point [C] in field coordinates into the position of the robot camera 4. x ,C y ] from robot camera 4 and camera focus point [C x ,C y The camera control unit 15 calculates the distance to the point of focus. Then, it converts this point of focus distance into the rotation amount of the focus demand. After that, the camera control unit 15 outputs the pan angle, tilt angle, rotation amount of the zoom demand, and rotation amount of the focus demand as camera control information to the robot camera 4.
[0064] The transformation rule is as follows: [C] x ,C y ,C s The correspondence between the robot camera 4 and the robot camera 4, the correspondence table between the field of view and zoom demand, and the correspondence between the gaze point distance and the rotation amount of the focus demand are set in advance.
[0065] [Configuration of the camera work prediction learning device] Referring to Figure 8, the configuration of the camera work prediction learning device (shooting area learning device) 4 according to the first embodiment will be described in detail. The camera work prediction learning device 4 uses a formation map to learn how to determine the shooting area of the robot camera 4 that captures soccer footage.
[0066] As shown in Figure 8, the camera work prediction learning device 4 comprises a convolutional unit 10, a play status determination unit 11, an LSTM unit (recursive processing unit) 12B, a coupling unit 13, a decision unit 14, an operation information conversion unit 40, a learning data expansion unit 41, a learning data storage unit 42, an input shooting area output unit 43, a learning decision unit 44, an evaluation unit (error calculation unit) 45, and a weight coefficient calculation unit (learning unit) 46. Note that the convolutional unit 10, play status determination unit 11, coupling unit 13, and decision unit 14 are the same as those in the camera work prediction generation device 1, so their explanation is omitted.
[0067] Here, the camera work prediction learning device 4 receives, as learning data, formation maps and play situations corresponding to past soccer videos, as well as operation information from when the soccer video was filmed. In other words, the camera work prediction learning device 4 assumes a case where formation maps and play situations corresponding to the actual filming results are generated in advance, and the actual camera operations by the cameraman can be acquired.
[0068] The operation information conversion unit 40 receives camera operation information and converts the input operation information into an input shooting area according to predetermined conversion rules. Here, the operation information conversion unit 40 receives camera operation information from the cameraman (for example, pan angle, tilt angle, and zoom demand rotation amount). For example, the operation information conversion unit 40 can acquire camera operation information using the technology described in Japanese Patent Publication No. 5771117. Then, the operation information conversion unit 40 converts the acquired operation information into a column vector [C] of the input shooting area according to the reverse conversion rules of the camera control unit 15. x ,C y ,C s Convert to ]. Then the operation information conversion unit 40 converts the input shooting area column vector [C x ,C y ,C s The output is sent to the learning data expansion unit 41.
[0069] The learning data augmentation unit 41 augments (inflates) the learning data to prevent overfitting. Here, the learning data augmentation unit 41 receives past formation maps and play status, as well as the input shooting area from the operation information conversion unit 40, as learning data. The learning data augmentation unit 41 then augments (data augmentes) the input learning data.
[0070] Specifically, the learning data augmentation unit 41 augments the formation map by horizontally flipping it. Furthermore, the learning data augmentation unit 41 adjusts the input shooting area [C] to correspond to the horizontally flipped formation map. x ,C y [-C] is a horizontally flipped shooting area. x ,C y This generates the data. Furthermore, regarding the play situation, if the learning data expansion unit 41 distinguishes between events such as goal kicks and corner kicks on the left and right sides of the field 20, it simply swaps the left and right sides. After that, the learning data expansion unit 41 writes the expanded learning data to the learning data storage unit 42.
[0071] The learning data storage unit 42 is a storage device such as a memory or hard disk that stores the learning data from the learning data expansion unit 41. Specifically, the learning data storage unit 42 stores the formation map, input shooting area, and play situation for each time period as learning data when actual soccer technique videos are filmed.
[0072] The input shooting area output unit 43 outputs the input shooting area at a time point later than the formation map from the learning data storage unit 42. Specifically, the input shooting area output unit 43 outputs future shooting areas for multiple frames (for example, 10 frames every 6 frames) to the LSTM unit 12B for the formation map and play status that are to be combined in the combination unit 13.
[0073] The LSTM unit 12B generates a vector of continuous shooting region candidates in the time direction by inputting the input shooting region output by the input shooting region output unit 43 into a recurrent neural network. For example, the LSTM unit 12B uses an LSTM to predict future shooting region candidates from the shooting regions of multiple frames by a predetermined prediction time (e.g., 1 second ahead). The processing of the LSTM unit 12B itself is the same as that of the LSTM unit 12 in Figure 2, so the explanation is omitted. The output of the LSTM unit 12B is a column vector of shooting region candidates [C x ,C y ,C s The format does not have to be ], and it can also be a column vector with four or more columns.
[0074] The learning decision unit 44 inputs the column vectors of the formation map to the second imaging region determination neural network, thereby determining the learning imaging region at a time point later than that of the formation map in the learning data storage unit 42.
[0075] Here, the learning decision unit 44 receives the column vectors of the formation map from the convolution unit 10. The learning shooting area determined by the learning decision unit 44 represents the predicted value of the shooting area a predetermined time interval (for example, 60 frames) in the future relative to the time of the input column vectors of the formation map. The learning shooting area is a 3-column column vector [C x ,C y ,C s It can be represented as ].
[0076] The learning decision unit 44 can utilize a known neural network as the neural network for determining the second shooting area (for example, a multilayer perceptron). The parameters (weights, biases) of the neural network for determining the second shooting area are calculated from the results of operations performed by experienced camera operators, similar to the neural network for determining the shooting area in the decision unit 14. Furthermore, the parameters of the neural network for determining the second shooting area are not used in the camera work prediction generation device 1. Subsequently, the learning determination unit 44 outputs the determined learning imaging area to the evaluation unit 45.
[0077] When the formation map changes significantly, the shooting area also changes significantly. In this case, if the learning decision unit 44 is not provided, a problem arises in that the predicted change in the shooting area is delayed. This problem stems from the fact that while the LSTM unit 12B can acquire a smooth shooting area over time, it becomes slow to respond to large changes in the formation map. Therefore, the learning decision unit 44 can calculate the parameters of the neural network for determining the shooting area in the decision unit 14 and the recurrent neural network in the LSTM unit 12B in order to improve the slow response of the LSTM unit 12B to large changes in the formation map and to predict an appropriate shooting area.
[0078] The evaluation unit 45 calculates the error of the shooting area determined by the decision unit 14, as well as the error of the learning shooting area determined by the learning decision unit 44. In other words, the evaluation unit 45 evaluates the accuracy of the shooting area and the learning shooting area relative to the actual shooting area (correct data) stored in the learning data storage unit 42. Furthermore, the evaluation unit 45 evaluates the accuracy of the play situation determined by the play situation determination unit 11 relative to the actual play situation (correct data) stored in the learning data storage unit 42.
[0079] Here, the evaluation unit 45 receives the shooting area from the decision unit 14, the learning shooting area from the learning decision unit 44, and the play status from the play status determination unit 11. The evaluation unit 45 also reads the formation map, input shooting area, and play status from the learning data storage unit 42 as correct answer data. The input shooting area (actual shooting area) read from the learning data storage unit 42 is data that is a predetermined time (for example, 1 second) ahead of the formation map. The play status read from the learning data storage unit 42 is data from the same time as the formation map.
[0080] As shown in Figure 9, the evaluation unit 45 evaluates each element C of the column vector of the imaging area. x ,C y ,Cs For each instance, the accuracy of the shooting area is evaluated using a loss function, based on the error between the correct data of the shooting area and play status and the prediction results of the play status determination unit 11 and the determination unit 14 for the shooting area and play status.
[0081] For example, as loss functions, the mean squared error can be used to evaluate the shooting area, and the cross-entropy error can be used to evaluate the gameplay situation. Mean squared error L mse The actual shooting area (input shooting area) is the true value y i The imaging area determined by the neural network 200 for determining the imaging area in the determination unit 14 is used as the predicted value y i ^ is represented by the following equation (1).
[0082]
number
[0083] The learning imaging area determined by the learning determination unit 44 can also be evaluated in the same way as in equation (1) above. The actual imaging area is true value y i The second neural network 230 for determining the imaging region of the learning decision unit 44 determines the imaging region, and the predicted value y i ^ can be expressed in the same way as in equation (1) above (mean squared error L mse2 ).
[0084] Cross-entropy error L ce This is the true value y of the elements of the column vector corresponding to the actual gameplay situation. i The elements (probabilities) of the column vector corresponding to the play status determined by the play status determination neural network 210 of the play status determination unit 11 are used as predicted values y i ^ is represented by the following equation (2).
[0085]
number
[0086] Subsequently, the evaluation unit 45 outputs the evaluation results (mean squared error, cross-entropy error) to the weight coefficient calculation unit 46. For evaluation purposes, the results of the loss function calculations performed by the play status determination unit 11, the decision unit 14, and the learning decision unit 44 are used, but pre-set weights may be reflected in the calculation results of each unit.
[0087] The weight coefficient calculation unit 46 trains the convolutional neural network, recurrent neural network, imaging region determination neural network, and second imaging region determination neural network so as to minimize the error calculated by the evaluation unit 45.
[0088] Here, the weight coefficient calculation unit 46 calculates the parameters (weights, biases) of the convolutional neural network, recurrent neural network, neural network for determining the shooting area, and second neural network for determining the shooting area, based on the evaluation results (loss function calculation results) from the evaluation unit 45, in order to improve the accuracy of determining the shooting area and play status, that is, to reduce the error between the actual shooting area and play status and the determined shooting area and play status.
[0089] For example, the weight coefficient calculation unit 46 can use Adam (Adaptive Moment Estimation) as the optimization algorithm for the loss function. The parameters calculated by the weight coefficient calculation unit 46 are applied to each neural network of the camera work prediction learning device 4 during learning, and to each neural network of the camera work prediction generation device 1 after learning. If learning is started with the parameters unlearned, the initial values of the parameters can be set arbitrarily (for example, randomly).
[0090] The weight coefficient at this point is w (t) The evaluation function (loss function) is E (w) In this case, the gradient g of the evaluation function (t) This is expressed by the following equation (3).
[0091]
number
[0092] Here, t is a parameter related to the number of iterations, which is the number of training steps, and it represents which iteration it is. The maximum value of t is the number of batches per epoch multiplied by the set number of epochs.
[0093] Also, m t ,v t Let each of these be set to satisfy the following equations (4) and (5).
[0094]
number
[0095] In this case, the unbiased estimator m of the gradient of the evaluation function is m. t ^, and the unbiased estimator of the squared gradient of the evaluation function v t ^ is represented by equations (6) and (7) below. Note that β1 and β2 are hyperparameters. Also, m0 = 0 and v0 = 0.
[0096]
number
[0097] The weight coefficient calculation unit 46 calculates the next weight coefficient w as shown in equations (8) and (9) below. (t+1) The following is calculated. Note that η and ε are hyperparameters, and η represents the learning rate. Also, w (t) ,g (t) ,m t ,v t It is a vector.
[0098]
number
[0099] Furthermore, the camera work prediction learning device 4 does not necessarily have to include a learning decision unit 44. In this case, the evaluation unit 45 does not need to calculate the error of the learning shooting area. In addition, the weight coefficient calculation unit 46 does not need to train the neural network for determining the second shooting area.
[0100] [Operation of the camera work prediction and generation device] Referring to Figure 10, the operation of the camera work prediction and generation device 1 will be explained. As shown in Figure 10, in step S1, the camera work prediction generation device 1 sets the initial value of the shooting area to be input to the LSTM unit 12.
[0101] In step S2, the formation map is input to the convolution unit 10 from the formation map generation unit 3. In step S3, the convolution unit 10 normalizes the formation map. In step S4, the convolution unit 10 convolves the normalized formation map.
[0102] In step S5, the play status determination unit 11 generates a play status vector of the soccer video by inputting the column vector of the formation map into the neural network for play status determination. In step S6, the LSTM unit 12 inputs the current imaging region into the recurrent neural network and generates a vector of continuous imaging region candidates in the time direction. In step S7, the coupling unit 13 generates a combined vector of the column vectors of the formation map, the vector of the candidate shooting area, and the play status vector of the soccer video.
[0103] In step S8, the determination unit 14 determines the imaging area by inputting a connection vector to the neural network for determining the imaging area. In step S9, the camera work prediction and generation device 1 determines whether the soccer video has ended. For example, the camera work prediction and generation device 1 determines that the soccer video has ended when it reaches the final frame of the soccer video.
[0104] If the soccer video ends (Yes in step S9), the camera work prediction and generation device 1 terminates processing. If the soccer video has not finished (No in step S9), the camera work prediction and generation device 1 proceeds to the process in step S10. In step S10, the camera work prediction and generation device 1 updates the shooting area input to the LSTM unit 12 to the next frame group and returns to the processing in step S2.
[0105] [Operation of the camera work predictive learning device] The operation of the camera work prediction learning device 4 will be explained with reference to Figures 11A and 11B. As shown in Figure 11A, in step S20, the camera work prediction learning device 4 sets the number of learning iterations. In step S21, the camera work prediction learning device 4 sets the counter to 1.
[0106] In step S22, the camera work prediction learning device 4 receives input for all frames of the actual soccer video training data, including formation maps, play situations, and camera operation information.
[0107] In step S23, the operation information conversion unit 40 converts the camera's operation information into an input shooting area. In step S24, the training data expansion unit 41 expands the training data. In step S25, the learning data expansion unit 41 stores the expanded learning data in the learning data storage unit 42. In step S26, the convolution unit 10 reads out formation maps for all frames from the learning data storage unit 42.
[0108] In step S27, the convolution unit 10 normalizes the read formation map. In step S28, the convolution unit 10 reads out the input shooting area for all frames from the learning data storage unit 42.
[0109] As shown in Figure 11B, in step S29, the convolution unit 10 convolves the formation maps for all normalized frames. In step S30, the play status determination unit 11 generates a play status vector of the soccer video by inputting column vectors of formation maps for all frames into the play status determination neural network.
[0110] In step S31, the LSTM unit 12B inputs the captured regions for all frames into the recurrent neural network and generates a vector of continuous captured region candidates in the time direction. In step S32, the coupling unit 13 generates a combined vector for all frames, consisting of the column vector of the formation map, the vector of the candidate shooting area, and the play status vector of the soccer video.
[0111] In step S33, the determination unit 14 determines the shooting area for all frames by inputting connection vectors to the neural network for determining the shooting area. In step S34, the learning decision unit 44 determines the learning shooting region by inputting the column vectors of the formation map for all frames into the second shooting region determination neural network. Note that the processes in steps S30 to S33 and the process in step S34 may be executed in parallel.
[0112] In step S35, the evaluation unit 45 calculates the error of the shooting area determined by the determination unit 14, the error of the learning shooting area determined by the learning determination unit 44, and the error of the play status determined by the play status determination unit 11. In step S36, the weight coefficient calculation unit 46 trains the convolutional neural network, the recurrent neural network, the neural network for determining the imaging region, and the second neural network for determining the imaging region so as to minimize the error calculated by the evaluation unit 45.
[0113] In step S37, the camera work prediction learning device 4 determines whether the counter is equal to the number of learning iterations. If the counter is equal to the number of learning iterations (Yes in step S37), the camera work prediction learning device 4 terminates the process. If the counter is not equal to the number of learning iterations (No in step S37), the camera work prediction learning device 4 proceeds to the process in step S38. In step S38, the camera work prediction learning device 4 increments the counter and returns to the process in step S29.
[0114] [Effects / Effects] As described above, the camera work prediction learning device 4 according to the first embodiment learns the prediction results of future shooting areas, so it can incorporate the delay due to the processing time of the device into the candidate shooting areas determined by the recurrent neural network. Furthermore, the camera work prediction and generation device 1 according to the first embodiment utilizes a recurrent neural network that has learned the prediction results of future shooting areas, so it can determine the shooting area in a way that incorporates the delay caused by the processing time of the device. In this way, delays in camera work can be suppressed.
[0115] Furthermore, the camera work prediction and generation device 1 and the camera work prediction and learning device 4 can reflect the gameplay of a soccer match in the shooting area, thereby enabling more optimal camera work.
[0116] (Second Embodiment) [Configuration of the camera work prediction and generation device] Referring to Figure 12, the configuration of the camera work prediction and generation device (shooting area determination device) 1B according to the second embodiment will be explained in terms of differences from the first embodiment. As shown in Figure 12, the camera work prediction generation device 1B comprises a convolution unit 10, an LSTM unit (recursive processing unit) 12, a coupling unit 13B, a determination unit 14, and a camera control unit 15.
[0117] In other words, the camera work prediction and generation device 1B does not include the play status determination unit 11 shown in Figure 2. The coupling section 13B is the same as the coupling section 13 in Figure 2, except that it does not combine the play status vectors. In other respects, it is the same as the first embodiment, so no further explanation is necessary.
[0118] [Configuration of the camera work prediction learning device] Referring to Figure 13, the configuration of the camera work prediction learning device (shooting area learning device) 4B according to the second embodiment will be explained in terms of differences from the first embodiment. As shown in Figure 13, the camera work prediction learning device 4B comprises a convolutional unit 10, an LSTM unit 12B, a coupling unit 13B, a decision unit 14B, an operation information conversion unit 40, a learning data expansion unit 41, a learning data storage unit 42B, an input shooting area output unit 43, a learning decision unit 44, an evaluation unit 45B, and a weight coefficient calculation unit 46.
[0119] In other words, the camera work prediction learning device 4B does not include the play status determination unit 11 shown in Figure 8. The coupling unit 13B, the determination unit 14B, and the learning data storage unit 42B are the same as the means in Figure 2, except that they do not utilize the play status vector. In other respects, it is the same as the first embodiment, so no further explanation is necessary.
[0120] [Effects / Effects] As described above, the camera work prediction generation device 1B and camera work prediction learning device 4B according to the second embodiment can suppress camera work delays due to the processing time of the device, similar to the first embodiment. Furthermore, the camera work prediction generation device 1B and the camera work prediction learning device 4B allow for the omission of the play status determination unit 11, thus enabling a simpler configuration.
[0121] Although embodiments have been described in detail above, the present invention is not limited to the embodiments described above, and includes design changes and the like that that do not depart from the spirit of the present invention.
[0122] In the embodiments described above, the ball game video was assumed to be soccer video, but the invention is not limited to this. For example, ball game videos could be basketball videos of a basketball game, handball videos of a handball game, or rugby videos of a rugby game.
[0123] In the embodiments described above, the camera work prediction generation device and the camera work prediction learning device were described as independent hardware, but the present invention is not limited thereto. For example, the present invention can also be realized by a program that causes hardware resources such as the CPU, memory, and hard disk of a computer to function as the camera work prediction generation device or camera work prediction learning device described above. This program may be distributed via a communication line, or it may be written to a recording medium such as a CD-ROM or flash memory and distributed. [Explanation of Symbols]
[0124] 1.1B Camera Work Prediction and Generation Device (Shooting Area Determination Device) 2 cameras 3. Formation Map Generation Unit 4,4B Camera work prediction learning device (shooting area learning device) 10 Folding section 11. Play Status Determination Unit (Play Status Determination Unit) 12,12B LSTM section (recursive processing section) 13,13B Joint part 14. Decision-making section 15 Camera Control Unit 40 Operation Information Conversion Unit 41. Training Data Expansion Unit 42 Learning Data Storage Unit 43 Input shooting area output section 44 Learning Decision Unit 45. Evaluation Unit (Error Calculation Unit) 46. Weight coefficient calculation unit (learning unit)
Claims
1. A shooting area determination device for determining the shooting area of a camera that shoots ball game footage, using a formation map that represents at least the position or speed of each subject within the field of ball game footage, A convolution unit generates column vectors of the formation map by convolving the formation map with a pre-trained convolutional neural network, A recurrent processing unit generates a vector of continuous shooting area candidates in the time direction by inputting the prediction results of the shooting area at a time earlier than the formation map into a pre-trained recurrent neural network, A coupling unit that generates a combined vector by combining the column vectors of the formation map and the vectors of the candidate shooting area, A determination unit determines the imaging area by inputting the connection vector into a pre-trained neural network for determining the imaging area, A device for determining the imaging area, characterized by comprising the following features.
2. The system further includes a play situation determination unit that generates play situation vectors representing events that occurred in the ball game video by inputting the column vectors of the formation map into a pre-trained neural network for determining play situations, The shooting area determination device according to claim 1, characterized in that the combining unit generates a combined vector by combining the column vector of the formation map, the vector of the candidate shooting area, and the play status vector of the ball game video.
3. The shooting area determination device according to claim 1, further comprising a camera control unit that converts the shooting area determined by the determination unit into a camera work based on a predetermined conversion rule, and controls the camera according to the converted camera work.
4. A shooting area learning device that learns how to determine the shooting area of a camera shooting ball game footage, using a formation map that represents at least the position or speed of each subject within the field of ball game footage, An operation information conversion unit receives camera operation information and converts the input operation information into an input shooting area according to a predetermined conversion rule, The learning data storage unit stores the formation map and the input shooting area for each time period as learning data, when the actual ball game video was filmed. An input shooting area output unit outputs an input shooting area at a time earlier than the formation map from the learning data storage unit, A convolutional unit generates column vectors of the formation map by convolving the formation map of the learning data storage unit using a convolutional neural network, The input imaging region output by the aforementioned input imaging region output unit is input to a recurrent neural network to generate a vector of continuous imaging region candidates in the time direction, and the recurrent processing unit provides this to the network. A coupling unit that generates a combined vector by combining the column vectors of the formation map and the vectors of the candidate shooting area, By inputting the connection vector into the neural network for determining the imaging area, a determination unit determines the imaging area, An error calculation unit calculates the error of the imaging area determined by the determination unit, A learning unit that trains the convolutional neural network, the recurrent neural network, and the neural network for determining the imaging area so that the error calculated by the error calculation unit is minimized, A camera area learning device characterized by comprising the following features.
5. The system further comprises a learning determination unit that determines a learning shooting area at a time point earlier than the formation map of the learning data storage unit by inputting a coupling vector of the column vector of the formation map into a second neural network for determining the shooting area, The error calculation unit calculates the error of the shooting area determined by the determination unit, and also calculates the error of the learning shooting area determined by the learning determination unit. The imaging region learning device according to claim 4, characterized in that the learning unit learns the convolutional neural network, the recurrent neural network, the imaging region determination neural network, and the second imaging region determination neural network so as to minimize the error calculated by the error calculation unit.
6. The system further includes a play situation determination unit that generates play situation vectors representing events that occurred in the ball game video by inputting the column vectors of the formation map into a pre-trained neural network for determining play situations, The shooting area learning device according to claim 4, characterized in that the coupling unit generates a combined vector by combining the column vector of the formation map, the vector of the candidate shooting area, and the play status vector of the ball game video.
7. A program for causing a computer to function as the imaging area determination device described in claim 1.
8. A program for causing a computer to function as the imaging area learning device described in claim 4.