Course information generation device and program
The course information generating device addresses the challenges of costly and non-scalable ball trajectory judgment systems by using a neural network-based feature extraction unit to determine the ball's course from input images, achieving accurate and cost-effective results without dedicated equipment.
Patent Information
- Application Number
- JP2023199597
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-06-06
AI Technical Summary
Existing systems for judging the course and landing point of a ball in sports competitions are costly, require dedicated equipment, and lack scalability, leading to variability in human judgment and the need for expensive infrastructure.
A course information generating device and program that uses a neural network-based feature extraction unit to calculate features from input images, allowing for the determination of the ball's course without the need for expensive dedicated filming equipment, and can be implemented using existing equipment.
The system enables accurate and cost-effective judgment of the ball's course using existing equipment, reducing variability in human judgment and eliminating the need for expensive infrastructure, thereby enhancing scalability and efficiency.
Smart Images

Figure 2025085900000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a course information generating device and a program. [Background technology]
[0002] In sports competitions, technology for determining the ball's trajectory and landing point is being researched and some is being put to practical use.
[0003] Patent Document 1 describes a device that determines whether a ball thrown by a pitcher in baseball has passed through the strike zone. The technology in Patent Document 1 involves installing multiple cameras in a baseball field to capture images of the ball being thrown, and then comparing the images with data to determine whether the ball has passed through the strike zone.
[0004] Patent Document 2 describes a technology that displays a strike zone on a head-mounted display worn by a baseball umpire (home plate umpire). The strike zone display system of Patent Document 2 assists the umpire in judging whether a pitch is a strike or a ball.
[0005] Patent Document 3 describes a "Hawk-Eye" identification method and system for determining the trajectory of a tennis ball. In the technique of Patent Document 3, a tennis ball is periodically photographed using multiple cameras. The image of the tennis ball is then positioned within each image. The photographed tennis ball is then projected onto the X-axis, Y-axis, and Z-axis to obtain the coordinates of the tennis ball on these three axes. This process is performed on a series of images to obtain the coordinates of the tennis ball over time and determine the trajectory of the tennis ball. This makes it possible to determine whether the tennis ball is in or out. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Japanese Patent Application Publication No. 09-290037 [Patent Document 2] JP 2011-161111 A [Patent Document 3] International Publication No. 2017 / 008218 Summary of the Invention [Problem to be solved by the invention]
[0007] When people (umpires, etc.) judge the course of the ball, there is a problem that the judgment varies depending on the individual, and support from machines, etc. Existing systems for judging the course and landing point of the ball have a problem that they require expensive dedicated equipment and systems.
[0008] Furthermore, conventional technology has problems with scalability, such as the need to install judging equipment in the stadium (baseball field, etc.).
[0009] It is therefore desirable to be able to judge the course of the ball using equipment that is already in use, thereby enabling the system to be realized at low cost.
[0010] The present invention has been made based on the recognition of the above problems, and aims to provide a course information generating device and program that can generate information regarding the ball's course, etc., without the need to install expensive dedicated filming equipment, etc. [Means for solving the problem]
[0011] [1] In order to solve the above problems, a course information generating device according to one aspect of the present invention includes a feature extraction unit that calculates, based on an input image, a feature for identifying a label corresponding to the image and representing the ball's course.
[0012] [2] Also, one aspect of the present invention is the course information generating device according to [1] above, wherein the feature extraction unit is configured using a neural network, and based on learning data provided as a set of a first image, a second image, and flag information indicating whether a label corresponding to the first image and a label corresponding to the second image are the same, (A) when the flag information indicates that the label corresponding to the first image and the label corresponding to the second image are the same, a first feature amount calculated by the feature extraction unit based on the first image and a second feature amount calculated by the feature extraction unit based on the second image are and (B) when the flag information indicates that the label corresponding to the first image and the label corresponding to the second image are different from each other, calculates a loss value that monotonically increases with an increase in distance between a first feature amount calculated by the feature extraction unit based on the first image and a second feature amount calculated by the feature extraction unit based on the second image; and a control unit that controls learning of the feature extraction unit to adjust values of internal parameters of the feature extraction unit in a direction that reduces the loss calculated by the loss calculation unit.
[0013] [3] In one aspect of the present invention, in the course information generation device of [1] or [2] above, the device further includes a feature distribution storage unit that stores a relationship between feature samples and known labels for each of the samples, and a determination unit that determines to which label a target feature corresponds, based on the relationship between the feature samples and the known labels stored in the feature distribution storage unit and a target feature calculated by the feature extraction unit based on an unknown image.
[0014] [4] Furthermore, according to one aspect of the present invention, in the course information generating device of [3] above, the determination unit determines which label the feature to be determined corresponds to using a K-nearest neighbor algorithm, based on a set of relationships between the feature samples stored in the feature distribution storage unit and the known labels.
[0015] [5] Moreover, one aspect of the present invention is that in the course information generating device of [3] or [4] above, M systems of processing units are provided, from a first system to an Mth system (where M is an integer and M≧2), and each of the M systems of processing units is provided with the feature extraction unit, the feature distribution memory unit, and the judgment unit, thereby making a unique label judgment for each system for the input image.
[0016] [6] Furthermore, one aspect of the present invention is that in any of the course information generating devices described above in [1] to [5], the image is an image cut out from a video of a baseball broadcast, and is an image of the moment when the ball thrown by the pitcher is caught by the catcher.
[0017] [7] Another aspect of the present invention is a program for causing a computer to function as the course information generation device described in any one of [1] to [6] above. Effect of the Invention
[0018] According to the present invention, the course information generating device can use a trained model to calculate, based on an input image, a feature amount for identifying a label corresponding to the image and representing the course of the ball. The label may be determined based on the feature amount. [Brief description of the drawings]
[0019] [Figure 1] 1 is a block diagram showing a schematic functional configuration of a course information generating device according to a first embodiment of the present invention. [Diagram 2] FIG. 2 is a block diagram showing a schematic internal functional configuration of a feature extraction unit 12 in the first embodiment. [Diagram 3] FIG. 2 is a schematic diagram showing a basic configuration of teacher data used in the first embodiment. [Figure 4]4 is a schematic diagram showing a method for creating the teacher data shown in FIG. 3 based on video of a baseball broadcast or the like in the first embodiment. FIG. [Diagram 5] FIG. 2 is a schematic diagram showing a method for selecting two images from N images (N is an integer equal to or greater than two) and generating data for learning by the course information generating device in the first embodiment. [Figure 6] FIG. 2 is a schematic diagram showing the configuration of data (teacher data) that includes two selected images and is used for learning in the first embodiment. [Figure 7] 7 is a schematic diagram for explaining a process for learning a feature extraction model held by a feature extractor using the learning data shown in FIG. 6 in the first embodiment. FIG. [Figure 8] 5 is a flowchart showing a procedure of a process for learning a feature extraction model in the course information generation device according to the first embodiment. [Figure 9] 4 is a graph showing a distribution of samples of feature amounts associated with each label in a feature amount space in the first embodiment. [Figure 10] 4 is a schematic diagram showing an example of the configuration of data stored in a feature distribution storage unit in the first embodiment. FIG. [Figure 11] 5 is a flowchart showing a procedure of a determination process in a determination section in the course information generating device according to the first embodiment. [Figure 12] FIG. 11 is a block diagram showing a schematic functional configuration of a course information generating device according to a second embodiment. [Figure 13] FIG. 11 is a schematic diagram showing two ways of attaching labels in the second embodiment. [Figure 14] FIG. 11 is a block diagram showing a schematic functional configuration of a course information generating device according to a third embodiment. [Figure 15] 1 is a block diagram showing an example of the internal configuration of a course information generating device according to a first embodiment, a second embodiment, and a third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0020] Next, a number of embodiments of the present invention will be described with reference to the drawings.
[0021] [First embodiment] The course information generating device of this embodiment is realized with a model capable of machine learning inside. In this embodiment, an image of the moment when the catcher catches the ball pitched by the pitcher in a baseball broadcast video is used as learning data for learning the model. In addition, a label (correct answer) representing the course of the ball is assigned to this image, for example, manually. In this embodiment, a Siamese network is used as the model. The model is trained so that the features of each image can be extracted well. Then, by using the trained model, the course information generating device calculates the feature amount of an unknown image (however, the image of the moment when the catcher catches the ball) based on the image. In addition, the course information generating device can determine the label corresponding to the image based on the feature amount of the image. In this embodiment, the course information generating device uses a K-nearest neighbor algorithm (k-NN algorithm) for the determination.
[0022] FIG. 1 is a block diagram showing a schematic functional configuration of a course information generating device according to this embodiment. As shown in the figure, the course information generating device 1 includes an image acquiring unit 11, a feature extracting unit 12, a determining unit 13, a determining result output unit 14, a feature distribution storage unit 16, and a loss calculating unit 17. Each of these functional units can be realized, for example, by a computer and a program. Each functional unit also has a storage unit as necessary. The storage unit is, for example, a variable in a program or a memory allocated by the execution of a program. Also, non-volatile storage units such as a magnetic hard disk drive or a solid state drive (SSD) may be used as necessary. Also, at least a part of the functions of each functional unit may be realized as a dedicated electronic circuit rather than a program.
[0023] The image acquisition unit 11 acquires an image from an external source. The image acquisition unit 11 passes the acquired image as an input to the feature extraction unit 12. The image acquired by the image acquisition unit 11 may be, for example, an image cut out from a video of a baseball broadcast, as described later. In other words, the image acquired by the image acquisition unit 11 does not have to be an image taken by a camera installed in a baseball field or the like exclusively for course judgment. The accuracy of the installation position of the camera required in this embodiment may be lower than the accuracy of the position required by a dedicated device for course judgment. The image acquisition unit 11 does not necessarily have to acquire an image from a device on the side of a content production company (including a broadcasting company), for example. The image acquisition unit 11 may receive a video of a baseball broadcast from a television receiver or an Internet terminal (such as a PC) owned by a general content viewer, and acquire the image.
[0024] The image acquired by the image acquisition unit 11 is, for example, an image cut out from a video of a baseball broadcast, and is an image of the moment when the catcher catches the ball pitched by the pitcher. Also, the image may be an image cut out from only a predetermined range (i.e., the vicinity of home base) from the video of the baseball broadcast so as to clearly show the situation near home base, the catcher's catching posture, etc.
[0025] The feature extraction unit 12 (feature extraction model) calculates a feature amount for identifying a label corresponding to an image and indicating the course of a ball (for example, a ball thrown by a pitcher). That is, the feature extraction unit 12 calculates a feature amount for an image based on the image acquired by the image acquisition unit 11. The image is information in which pixel values are appropriately arranged, and the feature extraction unit 12 calculates a feature amount based on these pixel values. The feature amount calculated by the feature extraction unit 12 is expressed as a multidimensional (for example, 768 dimensions, but is not limited to this number of dimensions) vector. The feature extraction unit 12 is configured to be able to adjust internal parameters by machine learning. Specifically, the feature extraction unit 12 is configured using a neural network. The relationship between the feature amount and the label will be further explained later.
[0026] The determination unit 13 determines which label the feature to be determined corresponds to. The determination unit 13 refers to information stored in the feature distribution storage unit 16 for this determination process. That is, what is stored in the feature distribution storage unit 16 is information representing the relationship between a feature sample and a known label corresponding to the feature. Note that the feature distribution storage unit 16 stores the relationship between the feature and the known label for a large number of feature samples. The determination unit 13 determines the label that should correspond to the feature to be determined based on the determination target feature, which is a feature calculated by the feature extraction unit 12 based on an unknown image, and the set of relationships between the feature sample and the known label corresponding to the feature stored in the feature distribution storage unit 16. The determination unit 13 may perform processing using, for example, a K-nearest neighbor algorithm (k-NN algorithm) for this label determination. Note that, here, the label is, for example, information representing the course of the ball.
[0027] That is, the determination unit 13 determines which label the unknown feature corresponds to based on the distance between the known feature and the unknown feature. That is, when the unknown feature and the known feature are close to each other, the determination unit 13 determines that the label corresponding to the unknown feature is equal to the label corresponding to the known feature. The specific processing procedure of the K-nearest neighbor algorithm will be described in more detail later.
[0028] The judgment result output unit 14 outputs to the outside the result of the judgment made by the judgment unit 13. That is, the judgment result output unit 14 outputs a label (this label indicates the course of the ball) which is the judgment result for the unknown image acquired by the image acquisition unit 11.
[0029] The feature distribution storage unit 16 stores information representing the relationship between the feature samples and the known labels corresponding to the feature samples. An example of the configuration of the data stored in the feature distribution storage unit 16 will be described later with reference to another drawing.
[0030] The loss calculation unit 17 calculates a loss used for machine learning of the feature extraction unit 12 (feature extraction model). As described later, the learning data used for learning by the feature extraction unit 12 is a set of a first image, a second image, and flag information indicating whether a label corresponding to the first image and a label corresponding to the second image are the same. (A) When the flag information indicates that a label corresponding to the first image and a label corresponding to the second image are the same, the loss calculation unit 17 calculates a loss value that monotonically decreases with an increase in the distance between a first feature amount calculated by the feature extraction unit 12 based on the first image and a second feature amount calculated by the feature extraction unit 12 based on the second image. Furthermore, (B) when the flag information indicates that the label corresponding to the first image and the label corresponding to the second image are different from each other, the loss calculation unit calculates a loss value that monotonically increases with an increase in the distance between a first feature amount calculated by the feature extraction unit 12 based on the first image and a second feature amount calculated by the feature extraction unit 12 based on the second image. The formula for calculating the loss as described above will be described in further detail later.
[0031] When the feature extraction unit 12 is trained, the feature extraction unit 12 is trained so as to adjust the values of the internal parameters of the feature extraction unit 12 in a direction that reduces the loss calculated by the loss calculation unit 17. The backpropagation method can be used for the training. The backpropagation method is an existing method.
[0032] The distance between the features may be, for example, the Euclidean distance (L2 norm) in the feature space. Alternatively, other norms or values of the powers of these norms (non-negative values) may be used as the distance between the features.
[0033] The course information generating device 1 includes a control unit (not shown). The control unit controls the overall operation of the course information generating device 1. The control unit controls at least whether the course information generating device 1 operates in a mode in which a feature extraction model is learned (which may be called a learning mode), or in a mode in which the learned feature extraction model is used to extract image features and determine (predict) labels (which may be called a prediction mode). Each unit constituting the course information generating device 1 performs processing according to these modes under the control of the control unit. The control unit executes control for the course information generating device 1 to perform the above-mentioned learning processing so that the internal parameters of the feature extraction unit 12 can be adjusted (optimized).
[0034] 2 is a block diagram showing a schematic internal functional configuration of the feature extraction unit 12. As shown in the figure, the feature extraction unit 12 includes an encoder 122 and a predictor 123. The feature extraction unit 12 can be realized using a neural network. That is, the internal configuration of the neural network 121 includes the encoder 122 and the predictor 123. The feature extraction unit 12 is also called a feature extraction model.
[0035] The encoder 122 receives an image, performs a calculation based on the input image and the values of the internal parameters, and outputs a predetermined vector. The feature amount output from the encoder 122 is passed to the predictor 123.
[0036] The predictor 123 receives a vector output from the encoder 122, performs a calculation based on the vector and the values of internal parameters, and outputs a predetermined feature amount (vector). The feature amount output from the predictor 123 represents the features of the original image. In this embodiment, the number of dimensions of the feature amount (vector) output from the predictor 123 is 768, but a vector with another number of dimensions may be output.
[0037] The neural network 121 is a hierarchical network consisting of many nodes. Each node in the neural network 121 accepts values input to the neural network 121 or values output from multiple other nodes, weights and adds these values, and outputs the calculation result. Here, the weight value in each node is an internal parameter of the neural network 121. The value of the internal parameter can be updated. When the neural network 121 is learning, the value of the internal parameter can be adjusted based on the learning data using a method such as backpropagation. In other words, by using a large amount of learning data, learning can be performed so that the neural network 121 realizes a desired input / output relationship. The neural network itself is an existing technology.
[0038] In other words, by appropriately characterizing the neural network 121 using the learning data, the feature extraction unit 12 is expected to output a feature amount that satisfactorily represents the feature of the input image. In other words, the feature extraction unit 12 is expected to calculate a feature amount that is suitable for a label that matches the input image.
[0039] As shown in the figure, the feature extraction unit 12 outputs two types of feature amounts. The first feature amount is a feature amount that is output as a result of processing by the predictor 123. The second feature amount is a feature amount that is not processed by the predictor 123, but is output as a result of processing by the encoder 122. These two types of feature amounts are used differently as follows. That is, when learning a feature extraction model, the loss calculation unit 17 calculates loss using the feature amount (first feature amount) output from the predictor 123. On the other hand, when determining a course using a learned feature extraction model, the feature amount (second feature amount) output from the encoder 122 is used.
[0040] In this embodiment, the feature extraction model is configured using a neural network composed of an encoder 122 and a predictor 123. However, the feature extraction model may be realized using a neural network different from this configuration. In this embodiment, learning is performed using a Siamese network technique so that a desired feature amount (or an approximate value thereof) can be obtained for an input image. That is, the feature amount extracted by the feature extraction model has information about a label associated with an image. In other words, the feature extraction model is learned so that the distance between feature amounts belonging to the same label is relatively close, and the distance between feature amounts belonging to different labels is relatively far. At this time, the form of the neural network, which is an element constituting the Siamese network, is not necessarily limited to the form in this embodiment.
[0041] Next, the teacher data used for learning by the course information generating device 1 will be described. The teacher data is basically a set of an image and a label (correct answer) associated with the image. This label indicates the course of the pitch by the pitcher. In addition, when the course information generating device 1 learns, the teacher data is a set of two images and a flag indicating whether the two images have the same label.
[0042] FIG. 3 is a schematic diagram showing a basic configuration of training data. As shown in the figure, one piece of training data is configured as a pair of one image and a label assigned to the image. As shown in the example, the image is an image near home base in baseball. This image is an image taken from approximately the direction of the back screen of the baseball stadium (the direction of the center field of the outfield). As an image of training data, a still image of the moment when the catcher catches the ball thrown by the pitcher is used. That is, the image usually includes the catcher, the batter, and the plate umpire, and the background includes the spectator seats behind the backstop. The label associated with the image represents the course of the ball. For example, the course of the ball is classified into three types, such as left, center, and right, as viewed from the direction of the back screen. For example, "left" may be associated with label "0", "center" with label "1", and "right" with label "2". However, the association between the course and the label is not limited to this example, and other association methods may be used. Regardless of whether the batter in the image is standing in the right-handed or left-handed batter's box, the correct label will be assigned according to his position (left, center, or right) when viewed from the back screen.
[0043] FIG. 4 is a schematic diagram showing a method for creating the training data shown in FIG.
[0044] FIG. 4(A) is an image of one frame included in a video of a live baseball game broadcast on television or distributed over the Internet. A frame at the moment when the catcher catches the ball thrown by the pitcher is selected as the training data. The frame selection at this timing may be performed by human judgment or by some automated judgment. The image in FIG. 4(A) is an image of a size suitable for live broadcast such as television broadcast, and contains information other than that required for determining the course of the ball. Therefore, an image from which as much information as possible other than that related to the course of the ball has been removed by cutting out only an appropriate range from the image in FIG. 4(A) is used as the training data.
[0045] FIG. 4(B) is an image cut out from the image in FIG. 4(A). In other words, FIG. 4(B) is an image in which the part near the catcher is cut out from an image of a baseball broadcast (an image showing the direction of home base from the direction near the back screen in the outfield). In the illustrated example, the image in FIG. 4(B) is an image shaped into a square of 512 pixels vertically and horizontally. The areas indicated by 0, 1, and 2 in the image shown in FIG. 4(B) correspond to labels representing the course of the ball. In other words, in this example, the left corresponds to label "0", the center corresponds to label "1", and the right corresponds to label "2". In the example image in FIG. 4(B), the course of the ball corresponds to label "0". In other words, in the case of this example image, "0" is assigned as the correct label.
[0046] In Fig. 4(B), plane 200 is an imaginary plane that is perpendicular to the line connecting the center of the pitcher's plate and the center of home base when the baseball field is viewed from above. In this embodiment, the ball's trajectory is information that indicates which area of 0 (left), 1 (center), or 2 (right) the center of the ball passed through when the ball thrown by the pitcher passes through plane 200. In other words, the ball's trajectory in this embodiment may be unrelated to the strike zone regulations in the rules of baseball.
[0047] FIG. 5 is a schematic diagram showing a method of selecting two images from N images (N is an integer equal to or greater than 2) to generate data for learning by the course information generating device 1. As shown in the figure, in this embodiment, one learning image is generated by selecting two images from N images to make a pair. Here, as also shown in FIG. 3, a label (correct answer) corresponding to that image is assigned to each of the N images. Any pair (two images) selected from the N images can be used as learning data. In other words, when N images and their correct answer labels are given, N C 2 (=N(N+1) / 2) pieces of training data can be created. In other words, the amount of training data can be increased.
[0048] FIG. 6 is a schematic diagram showing the configuration of learning data (teacher data) including two selected images. As shown in the figure, one learning data is configured as a set of a flag, an image 1, and an image 2. The flag is information indicating whether the label (correct answer) associated with the image 1 and the label (correct answer) associated with the image 2 are the same or not. In other words, if the label (course) of the image 1 and the label (course) of the image 2 are different, the value of the flag is "0". Also, if the label (course) of the image 1 and the label (course) of the image 2 are the same, the value of the flag is "1". In this way, the model held by the course information generating device 1 is learned using a flag indicating whether the two images correspond to the same label or not.
[0049] FIG. 7 is a schematic diagram for explaining a process when the model held by the feature extraction unit 12 is learned using the learning data shown in FIG. 6. As shown in the figure, the feature extraction unit 12 adjusts internal parameters using a pair of images (image 1 and image 2). That is, the feature extraction unit 12 reads image 1, and calculates feature amount 1 (vector) based on the image 1 using the internal parameter values held by the feature extraction unit 12 at that time. The feature extraction unit 12 also reads image 2, and calculates feature amount 2 (vector) based on the image 2 using the internal parameter values held by the feature extraction unit 12 at that time. The feature amount 1 and feature amount 2 calculated by the feature extraction unit 12 are passed to the loss calculation unit 17. The loss calculation unit 17 calculates a loss based on these two feature amounts and the value of a flag (see FIG. 6) held by the learning data. The feature extraction unit 12 updates the internal parameter values using the backpropagation method based on the loss calculated by the loss calculation unit 17.
[0050] At a given point in time, the feature extraction unit 12 that calculates feature amount 1 based on image 1 and the feature extraction unit 12 that calculates feature amount 2 based on image 2 share the internal parameters of the model. In other words, for one piece of learning data (a pair of images), feature amounts are extracted using common internal parameters.
[0051] The loss calculation unit 17 calculates the loss L by the following formula (1).
[0052]
number
[0053] In formula (1), Y is a flag value of the learning data (see FIG. 6). The value of Y is 1 or 0. That is, when the label of the input image 1 is the same as the label of the input image 2, Y=1. When the label of the input image 1 is different from the label of the input image 2, Y=0. D is a distance between the feature amount of the image 1 and the feature amount of the image 2 output from the feature extraction unit 12. The feature amount used in calculating the loss L is the feature amount output from the predictor 123 (see FIG. 2). Furthermore, m is a hyperparameter that is appropriately determined. For example, m may be 1.0. However, the value of m is not limited to this.
[0054] The feature quantity of image 1 and the feature quantity of image 2 calculated by the feature extraction unit 12 are respectively z 1 and z 2 Then, the distance between these two features is expressed, for example, by the following formula (2). In other words, the distance D is the L2 norm. However, if the feature (vector) z 1 and z 2 Regarding, it is as shown in equation (3).
[0055]
number
[0056]
number
[0057] In other words, by training the feature extraction unit 12 using the loss function expressed by equation (1), when Y=1 (image 1 and image 2 have the same label), the loss L is D2 In other words, the feature value z of image 1 is proportional to 1 and feature value z of image 2 2 Learning is performed in the direction in which the difference between the feature value of image 1 and the feature value of image 2 becomes smaller. In addition, when Y=0 (image 1 and image 2 have different labels), the loss L is proportional to the square of (mD) (however, if (mD) is negative, it is 0). In other words, the direction in which D>m (the feature value z of image 1) 1 and feature value z of image 2 2 Learning is performed in the direction that increases the difference between the
[0058] As described above, by performing learning based on the loss L, the feature amount (vector) extracted from an image by the feature extraction unit 12 becomes information that can identify the label of the image.
[0059] 8 is a flowchart showing the procedure for learning a feature extraction model in the course information generating device 1. The learning processing procedure will be described below with reference to this flowchart.
[0060] First, in step S1, the course information generating device 1 inputs one piece of teacher data. The configuration of this teacher data is as described with reference to Fig. 6. That is, the course information generating device 1 reads two images (image 1 and image 2) and a flag.
[0061] Next, in step S2, the feature extraction unit 12 calculates the feature amounts of each of the images 1 and 2 input in step S1 based on the internal parameters at that time. Note that the feature amounts calculated in this step are the feature amounts output from the predictor 123, as described above.
[0062] Next, in step S3, the loss calculation unit 17 calculates the loss L by the above formula (1) based on the feature amounts of each of images 1 and 2 calculated in step S2 and the given flag value (Y). In addition, the course information generation device 1 adjusts (updates) the values of the internal parameters of the feature extraction unit 12 by the backpropagation method based on the calculated loss L.
[0063] Next, in step S4, the course information generating device 1 determines whether or not to end the learning. The determination of whether or not to end the learning may be made, for example, based on whether or not processing using a predetermined amount of learning data has been completed, or whether or not the value of the internal parameter set has converged. If the learning is to be ended (step S4: YES), the process proceeds to the next step S5. If the learning is not to be ended (step S4: NO), the process returns to step S1.
[0064] When the process proceeds to step S5, the course information generating device 1 saves the feature extraction model obtained as a result of the learning. In other words, the course information generating device 1 writes the values of the internal parameters of the feature extraction unit 12 after learning to a recording medium or the like. After the process of step S5 ends, the entire process of this flowchart ends.
[0065] Next, a processing method in which the determination unit 13 determines the course based on the feature amount will be described.
[0066] FIG. 9 is a graph showing the distribution of samples of features associated with each label in the feature space. Each of these samples corresponds to each of the original teacher data (see FIG. 3). In reality, the feature is a multidimensional vector (768 dimensions in this embodiment as an example), but in this graph, for convenience, such a multidimensional vector is projected onto a two-dimensional space. The features shown here are the features output from the encoder 122 (see FIG. 2). As shown in the figure, in the feature space, features associated with the same label generally exist relatively close to each other, and features having different labels exist relatively far from each other. In the example shown in the figure, the feature associated with label "0" is located in the lower left area of the graph. The feature associated with label "1" is located in the area slightly above the center of the graph. The feature associated with label "2" is located in the lower right area of the graph. In the graph, the areas in which the features of each label exist are surrounded by dashed lines. Note that within the regions corresponding to each label, there are multiple regions (clusters) with small clusters of multiple features, but this is a bias in the features due to factors other than the course of the pitch. The reason for this bias is that each live baseball broadcast video has its own characteristics (for example, the characteristics of the images shown at each baseball stadium).
[0067] The determination unit 13 determines with which label (0, 1, or 2) an unknown image should be associated, based on the distribution of feature amounts exemplified in FIG.
[0068] That is, since the learning of the feature extraction model is completed using the teacher data, it is already known which label (course: 0, 1, or 2) the feature value corresponding to each teacher data corresponds to. That is, the judgment unit 13 calculates the feature value of an image based on an image included in each teacher data using the trained feature extraction model. The judgment unit 13 also associates the calculated feature value with the label value of the teacher data and stores them in the feature distribution storage unit 16 (see also FIG. 10). Then, the feature extraction unit 12 calculates the feature value for an input unknown image using the trained feature extraction model. The judgment unit 13 judges which label the unknown image corresponds to based on the relationship between the feature value calculated by the feature extraction unit 12 for the unknown image and the feature values of a large number of known samples.
[0069] Specifically, the determination unit 13 determines a label (ball course) to be assigned to the image to be determined based on the distance of the feature to be determined (feature of an unknown image) to the feature of the teacher data (a set of feature of images of the bases). Note that the determination unit 13 can use a K-nearest neighbor algorithm (k-NN algorithm) for this determination.
[0070] The procedure of the K-nearest neighbor algorithm is as follows. That is, the judgment unit 13 receives a feature quantity to be judged from the feature extraction unit 12. This feature quantity is conveniently called a "feature quantity to be judged." Then, the judgment unit 13 acquires labels of the feature quantities of k samples (k is a positive integer determined appropriately) that are closest to the feature quantity to be judged from a set of samples of known feature quantities (for example, a large number of samples shown in FIG. 9). That is, the judgment unit 13 acquires k labels that are close to the feature quantity to be judged. Here, "close" means that the distance between the feature quantities is small. For example, the Euclidean distance (L2 norm) may be used as the distance between the feature quantities. However, other distances may be used as the distance between the feature quantities. Then, the judgment unit 13 determines the label of the feature quantity to be judged by majority vote of the k acquired labels. However, if there are multiple labels with the largest number of labels, the judgment unit 13 may determine the label with the smallest total distance from among the multiple labels, for example, by using the total sum of the above distances, or may determine the label according to a predetermined priority order. The value of k is arbitrary. For example, k may be set to 1. When k=1, the label of the feature closest to the target feature is determined as the label of the target feature. Alternatively, k=2, 3, etc. may be used. The K-nearest neighbor algorithm itself is an existing method.
[0071] FIG. 10 is a schematic diagram showing an example of the configuration of data stored in the feature distribution storage unit 16. This feature distribution storage unit 16 can store a set of pairs of feature amounts and labels. That is, the judgment unit 13 can store feature amounts (feature amounts whose labels are known) and labels corresponding to the feature amounts in the feature distribution storage unit 16 in association with each other. As shown in the figure, the feature distribution storage unit 16 can store data having a tabular structure. This table has data items of feature amounts and labels. One row in this table holds information about one feature amount (one sample). The feature distribution storage unit 16 can be realized by using, for example, a magnetic hard disk device, a semiconductor memory, or the like.
[0072] 11 is a flowchart showing the procedure of the determination process in the determination unit 13 in the course information generating device 1. The procedure of the determination process will be described below with reference to this flowchart.
[0073] First, in step S11, the course information generating device 1 inputs each of a plurality of training data (see FIG. 3) into the trained feature extraction model. The training data includes images and labels.
[0074] Next, in step S12, the feature extraction unit 12 calculates a feature amount (vector) for each image included in the training data input in step S11, using the trained model. Note that the feature amount calculated in this step is the feature amount output from the encoder 122, as described above. The feature extraction unit 12 passes to the determination unit 13 a set of pairs of the calculated feature amount and the label associated with the image on which the calculated feature amount is based.
[0075] Next, in step S13, the determination unit 13 writes the features and labels passed from the feature extraction unit 12 in step S12 in the feature distribution storage unit 16 (see FIG. 10) in association with each other. By the processing up to this step, information indicating the distribution state of the features for the known image is stored in the feature distribution storage unit 16, and each feature sample is associated with a label (a label indicating the pitching trajectory).
[0076] Next, in step S14, the course information generating device 1 inputs one image to be determined (an image in which the label representing the course is unknown) into the feature extraction model.
[0077] Next, in step S15, the feature extraction unit 12 calculates a feature amount for the image input in step S14 using the trained model. The feature extraction unit 12 passes the feature amount to be judged (judgment target feature amount) to the judgment unit 13.
[0078] Next, in step S16, the determination unit 13 determines the label for the feature to be determined based on a set of pairs of feature and label for the teacher data (written in the feature distribution storage unit 16 in step S13 above). In the determination process of this step, the determination unit 13 can use a K-nearest neighbor algorithm.
[0079] Next, in step S17, the determination unit 13 outputs the determination result. That is, the determination unit 13 outputs a label to which the above-mentioned determination target feature quantity corresponds. That is, the determination unit 13 outputs a label corresponding to the image input in step S14. The label of the determination result indicates the pitch course (for example, distinction between left / center / right) in the image.
[0080] Next, in step S18, it is determined whether there are any more images to be determined. If there are more images to be determined (step S18: YES), the process returns to step S14. If there are no more images to be determined (step S18: NO), the entire process of this flowchart ends.
[0081] Through the above procedure, the course information generating device 1 can use the trained feature extraction model to extract features based on an image, and determine a label (corresponding to the course of the ball) based on the features.
[0082] According to this embodiment, the course information generating device 1 can calculate feature amounts representing information about the course of the ball based on an image. The course information generating device 1 can also learn a feature extraction model for calculating the feature amounts. The course information generating device 1 can also determine a label (a label corresponding to the course of the ball) based on the calculated feature amounts.
[0083] [Evaluation Experiment (First Embodiment)] A demonstration experiment was carried out to confirm the performance of this embodiment, and the results of the experiment will be described below.
[0084] The learning data consisted of 575,128 pairs of images (image 1 and image 2) and flag information (0 or 1) for the pairs (see Figure 6). The number of learning data was the number of all combinations of selecting two images from 1073 images ( 1073 C 2 ) Of these 575,128 pieces of data, 372,644 pieces of data had a flag value of "0," and 202,484 pieces of data had a flag value of "1." 184 images not included in the training data (images of the moment the catcher catches the ball) were used as evaluation data. Additionally, experiments were conducted using k=1 as the hyperparameter value in the K-nearest neighbors algorithm.
[0085] The experimental results are shown in Table 1 below.
[0086] [Table 1]
[0087] As a result of the experiment, the accuracy was calculated as (15+44+55) / 184, and the value was 0.62. In particular, there were zero instances where the determination unit 13 determined that the correct answer was "right" for an image whose correct answer was "left". There were also zero instances where the determination unit 13 determined that the correct answer was "left" for an image whose correct answer was "right". The proportion of instances where the determination unit 13 determined that the correct answer was "middle" for an image whose correct answer was "middle" was high at 0.94 (=44 / 47). It can be said that the effectiveness of the course information generating device 1 according to this embodiment was confirmed by this experiment.
[0088] [Second embodiment] Next, a second embodiment of the present invention will be described. Note that the matters already described in the previous embodiment may not be described below. Here, the matters unique to this embodiment will be mainly described.
[0089] 12 is a block diagram showing a schematic functional configuration of a course information generating device according to this embodiment. As shown in the figure, the course information generating device 2 is equipped with two systems of judgment means. That is, the course information generating device 2 is configured to include an image acquisition unit 11, a feature extraction unit 12, a judgment unit 13, a judgment result output unit 14, a feature distribution storage unit 16, a feature extraction unit 22, a judgment unit 23, a judgment result output unit 24, and a feature distribution storage unit 26. The course information generating device 2 of this embodiment can also be realized using a computer, an electronic circuit, or the like.
[0090] In this configuration, the feature extraction unit 12, the judgment unit 13, the judgment result output unit 14, and the feature distribution storage unit 16 form a first system of judgment means. Also, the feature extraction unit 22, the judgment unit 23, the judgment result output unit 24, and the feature distribution storage unit 26 form a second system of judgment means. In this embodiment, the first system and the second system can be assigned labels in different ways. The image acquisition unit 11 can supply acquired images to both the first system and the second system.
[0091] The feature extraction unit 22 has the same function as the feature extraction unit 12. However, the second system performs processing independent of the first system.
[0092] The determination unit 23 has the same function as the determination unit 13. However, the second system performs processing independent of the first system.
[0093] The decision result output unit 24 has the same function as the decision result output unit 14. However, the second system performs processing independent of the first system.
[0094] The feature distribution storage unit 26 has the same function as the feature distribution storage unit 16. However, the feature distribution storage unit 26 belonging to the second system stores information on a feature distribution different from that stored in the feature distribution storage unit 16 belonging to the first system.
[0095] FIG. 13 is a schematic diagram showing two ways of labeling in this embodiment. FIG. 13(A) shows a labeling method in the first system. FIG. 13(B) shows a labeling method in the second system. As shown in the figure, in the labeling method in the first system, labels 0, 1, and 2 are assigned to the left, center, and right, respectively, as in the first embodiment. In addition, in the labeling method in the second system, labels 0, 1, and 2 are assigned to the top, center, and bottom, respectively, unlike the first embodiment. In other words, even if the image is the same, labels are assigned differently in the first system and the second system.
[0096] That is, the feature extraction model in the first system is trained based on the labels assigned in Fig. 13(A). On the other hand, the feature extraction model in the second system is trained based on the labels assigned in Fig. 13(B). That is, the feature extraction unit 12 (first system) and the feature extraction unit 22 (second system) in this embodiment are trained using different learning data (labels). However, both may be trained using a common image.
[0097] The judgment unit 13 (first system) and the judgment unit 23 (second system) judge the label (pitching course) of the unknown image based on the feature amount calculated by the feature extraction model of each system. Each of the judgment unit 13 (first system) and the judgment unit 23 (second system) may perform judgment using the K-nearest neighbor algorithm, as in the first embodiment. Furthermore, the value of the hyperparameter may be k=1, or another value of k may be used. As a result, each of the judgment unit 13 (first system) and the judgment unit 23 (second system) determines (assigns) a unique label to the unknown image as a judgment result.
[0098] That is, the labels 0, 1, and 2 of the judgment result given by the judgment unit 13 (first system) respectively mean that the course is left, middle, and right. Also, the labels 0, 1, and 2 of the judgment result given by the judgment unit 23 (second system) respectively mean that the course is up, middle, and down.
[0099] For a single unknown input image, the first and second systems may extract independent features and assign independent labels to the images. In this case, the combination of the two systems can provide nine different course determination results for a single image, as shown in Table 2 below.
[0100] [Table 2]
[0101] [Modification of the second embodiment] As the second embodiment, an example has been described in which image features are calculated and labels (ball courses) are determined using two systems, a first system and a second system. As a modification of the second embodiment, the number of systems may be three or more. That is, the course information generating device 2 may generally include processing systems from a first system to an Mth system (M≧2). That is, in each of these systems, features for different types of labels can be calculated. Also, a unique label can be assigned to each system based on the calculated features.
[0102] That is, the course information generating device 2 of the second embodiment (including modified examples) includes M processing systems from a first system to an M-th system (where M is an integer and M≧2). Each of the M processing systems includes a feature extraction unit, a feature distribution storage unit, and a determination unit, so that a unique label determination is performed for each system for an input image. In the above example, the unique labels are labels for distinguishing left / middle / right and labels for distinguishing top / middle / bottom. These unique labels for each system are determined independently from the label determinations in the processing units of the other systems.
[0103] As described above, according to this embodiment (including the modified examples), feature amounts relating to a plurality of different labels can be generated and the labels can be determined.
[0104] [Third embodiment] Next, a third embodiment of the present invention will be described. Note that the description of the matters already described in the previous embodiments may be omitted below. Here, the description will focus on matters unique to this embodiment.
[0105] 14 is a block diagram showing a schematic functional configuration of a course information generating device according to this embodiment. As shown in the figure, the course information generating device 3 includes an image acquiring unit 11, a feature extracting unit 12, and a loss calculating unit 17. The course information generating device 3 of this embodiment can also be realized using a computer, an electronic circuit, or the like.
[0106] The image acquisition unit 11 in this embodiment has the same functions as those in the first embodiment, etc. The feature extraction unit 12 in this embodiment has the same functions as those in the first embodiment, etc. The loss calculation unit 17 in this embodiment has the same functions as those in the first embodiment, etc.
[0107] A feature of this embodiment is that the course information generating device 3 does not have a judgment unit 13 or a feature distribution storage unit 16. In other words, the course information generating device 3 extracts and outputs feature amounts related to the course of the pitch based on an image, but does not judge the labels of the feature amounts. The feature amounts output by the course information generating device 3 are information about the course of the ball. In other words, the course information generating device 3 generates course information.
[0108] The course information generating device 3 having the trained feature extracting unit 12 can extract the feature amount of an unknown image based on the image.
[0109] The course information generating device 3 has a loss calculating unit 17. That is, the course information generating device 3 can learn a feature extraction model based on the loss calculated by the loss calculating unit 17, similarly to the first embodiment.
[0110] [Effects of the embodiment] The first to third embodiments (including the modified examples) have been described above. According to any of these modified examples, it is possible to generate and output information about the trajectory of the ball based only on the video of a live baseball broadcast, without using expensive dedicated equipment.
[0111] In the first to third embodiments (including the modified examples), when the course information generating device includes the loss calculation unit 17, it is possible to learn a feature extraction model. Based on the images for learning (the image at the time when the catcher catches the ball), two of these images can be combined to create learning data. In other words, it is possible to increase the amount of learning data. That is, even if there is not a large amount of images at the time when the catcher catches the ball, it is possible to learn a feature extraction model.
[0112] The course information can be generated by any one of the course information generating devices according to the first to third embodiments. As the course information, a label representing the course can be generated and output, or the feature amount extracted by the feature extracting unit can be output as it is. The course information generating devices according to these embodiments can be used as a course determining device, and can also be used for various other services. As an example, the course information generating device according to any one of the first to third embodiments can be used for a commentary audio service that provides information about the pitching course to visually impaired people who watch content. In other words, the course information generating device can be used for automatically adding commentary audio.
[0113] [Realization using computers and programs] FIG. 15 is a block diagram showing an example of the internal configuration of the course information generating device in each of the first embodiment, the second embodiment, and the third embodiment. The course information generating device in each embodiment can be realized using a computer. As shown in the figure, the computer is configured to include a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, and a bus 906. The computer itself can be realized using existing technology. The central processing unit 901 executes instructions included in a program read from the RAM 902 or the like. The central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic and logical operations according to each instruction. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. Note that RAM is an abbreviation for "random access memory." The input / output port 903 is a port through which the central processing unit 901 exchanges data with an external input / output device or the like. The input / output devices 904 and 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used inside the computer. For example, the central processing unit 901 reads and writes data from and to the RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port 903 via the bus 906.
[0114] At least some of the functions of the course information generating device in each of the above-mentioned embodiments can be realized by a computer and a program. In that case, the program for realizing the functions may be recorded in a computer-readable recording medium, and the program recorded in the recording medium may be read into a computer system and executed to realize the functions. The term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, CD-ROMs, DVD-ROMs, and USB memories, and storage devices such as hard disks built into a computer system. In other words, the term "computer-readable recording medium" may be a non-transitory computer-readable recording medium. Furthermore, the term "computer-readable recording medium" may include a medium that temporarily and dynamically holds a program, such as a communication line when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, and a medium that holds a program for a certain period of time, such as a volatile memory inside a computer system that is a server or client in that case. The above-mentioned program may be a program for realizing some of the above-mentioned functions, and may further be a program that can be realized in combination with a program already recorded in the computer system.
[0115] Although a number of embodiments have been described above, the present invention can also be embodied in the following modified examples.
[0116] [Variation 1] In the above embodiment, a course information generating device for classifying the ball's course into three classes (for example, left / middle / right or top / middle / bottom) has been described. As a modification, the number of classes for classification is arbitrary. That is, the number of classes for classification may be two, or any number equal to or greater than three. This modification 1 can be realized by assigning an appropriate label according to the number of classes.
[0117] [Variation 2] In the above embodiment, a configuration has been described in which the course information generating device includes the loss calculation unit 17, and learns a feature extraction model using the backpropagation method based on the loss calculated by the loss calculation unit 17. As a modified example, the course information generating device may not have a function for learning a model. In that case, the course information generating device does not need to have the loss calculation unit 17. In this modified example, the course information generating device can calculate the feature amount of an input image using a trained feature extraction model (i.e., by reading in the internal parameter values of the trained feature extraction model). In addition, the course information generating device can determine a label (ball course) based on the calculated feature amount.
[0118] The above describes in detail an embodiment of the present invention (including modified examples) with reference to the drawings. However, the specific configuration is not limited to this embodiment, and also includes designs that do not deviate from the gist of the present invention. [Industrial Applicability]
[0119] The present invention can be used, for example, to analyze sports images (baseball broadcast images, etc.). Or, it can be used to analyze the course (trajectory, etc.) of an object shown in the image. However, the scope of use of the present invention is not limited to the examples given here. [Explanation of symbols]
[0120] 1,2,3 Course information generator 11 Image acquisition section 12 Feature Extraction Unit 13 Judgment section 14. Judgment result output section 16 Feature distribution memory unit 17 Loss Calculation Section 22 Feature Extraction Unit 23 Judgment section 24 Judgment result output section 26 Feature distribution memory unit 121 Neural Networks 122 Encoder 123 Predictor 200 planes 901 Central Processing Unit 902 RAM 903 Input / Output Ports 904,905 Input / Output Devices 906 Bus
Claims
1. a feature extraction unit that calculates, based on an input image, a feature amount for identifying a label corresponding to the image and indicating a ball's course; A course information generating device comprising:
2. The feature extraction unit is configured using a neural network, Based on learning data given as a set of a first image, a second image, and flag information indicating whether a label corresponding to the first image and a label corresponding to the second image are the same, (A) when the flag information indicates that a label corresponding to the first image and a label corresponding to the second image are the same, a loss value is calculated that monotonically decreases with an increase in a distance between a first feature amount calculated by the feature extraction unit based on the first image and a second feature amount calculated by the feature extraction unit based on the second image; (B) when the flag information indicates that a label corresponding to the first image and a label corresponding to the second image are different from each other, a loss value is calculated that monotonically increases with an increase in a distance between a first feature amount calculated by the feature extraction unit based on the first image and a second feature amount calculated by the feature extraction unit based on the second image. A loss calculation unit; a control unit that controls learning of the feature extraction unit so as to adjust values of internal parameters of the feature extraction unit in a direction that reduces the loss calculated by the loss calculation unit; The course information generating device according to claim 1 , further comprising:
3. a feature distribution storage unit that stores a relationship between a sample of feature amounts and a known label for each of the samples; a determination unit that determines to which label a feature to be determined corresponds based on a relationship between the sample of the feature stored in the feature distribution storage unit and the known label and a feature to be determined, the feature being a feature calculated by the feature extraction unit based on an unknown image; and The course information generating device according to claim 1 , further comprising:
4. the determining unit determines to which label the feature to be determined corresponds by using a K-nearest neighbor algorithm, based on a set of relationships between the feature samples and the known labels stored in the feature distribution storage unit; The course information generating device according to claim 3.
5. The present invention includes a processing unit having M systems, from a first system to an M-th system (where M is an integer and M≧2), Each of the M processing units includes the feature extraction unit, the feature distribution storage unit, and the determination unit, and performs a unique label determination for each of the M systems for the input image. The course information generating device according to claim 3.
6. The image is an image cut out from a video of a baseball broadcast, and is an image of the moment when the ball pitched by the pitcher is caught by the catcher. The course information generating device according to any one of claims 1 to 5.
7. A course information generating device according to any one of claims 1 to 5, A program that makes a computer function as a
Citation Information
Patent Citations
Automatic strike judging device
JP1997290037A
Strike zone presentation system
JP2011161111A
Hawk-eye identification method and system in tennis match
WO2017008218A1