Data Structure

The data structure addresses the challenge of selecting promising racehorses in POG by using weighted coefficients to rank candidate horses based on historical data, providing an objective and efficient method for predicting performance.

JP7691606B2Active Publication Date: 2025-06-12杉村 紀夫
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2020211903
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-12-22
Publication Date
2025-06-12
Estimated Expiration
2040-12-22

AI Technical Summary

Technical Problem

In POG (Paper Owner Game), selecting promising racehorses with a high possibility of winning large prize money is challenging due to the vast amount of subjective information and the time-consuming process of comparing multiple factors for each horse.

Method used

A data structure that uses a combination of numerical data from past generations for factors such as the father, mother, trainer, owner, breeder, and transaction price, with a scoring system based on weighted coefficients to efficiently rank and select candidate horses.

Benefits of technology

This data structure enables users to predict the ranking of racehorse performance objectively, reducing the reliance on experience and subjective judgments, and facilitating the efficient selection of high-quality horses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007691606000004
    Figure 0007691606000004
  • Figure 0007691606000005
    Figure 0007691606000005
  • Figure 0007691606000006
    Figure 0007691606000006
Patent Text Reader

Abstract

To provide a data structure which is used for ranking prediction of future scores of race horses so as to support a game participant to be able to efficiently select a candidate horse in POG.SOLUTION: The data structure includes at least three or more factors selected from a father, mother's father, sex, trainer, owner, producer, transaction price, mother's racing score, birthday, and mother's age at a birthday of each of race horses, and each piece of data comprises numerical data based on scores of race horses in past generations collected for each of the factors, and a computer is caused to perform computation based on weight coefficients of individual pieces of data to output score values based on which the race horses can be compared to one another.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data structure used for predicting the performance of racehorses.

Background Art

[0002] POG (Paper Owner Game) is a type of game in which participants select an actual racehorse as a virtual horse owner and compete for points converted from the prize money obtained based on the race results of the selected racehorse. Since the participants in the game do not become actual horse owners but become "paper horse owners," it is called Paper Owner. In POGs hosted by racing magazines and Internet media, there are large-scale ones with tens of thousands of participants, and prizes are awarded to the top prize winners in terms of the obtained points, so it is widely spread. As a general rule for POG targeting central horse racing, a period of one year from the start of the 2-year-old race to the final race on the construction day of the Tokyo Yushun the following year is set. Around 10 unraced 2-year-old horses are selected. Points are converted from the first to fifth place prize money by about one ten-thousandth, and the outcome is often determined by the total points obtained by the selected horses.

[0003] From April to June every year, books, so-called POG books, that publish information on 2-year-old horses as selection candidates are published. For example, there are "POG Master" by Kobunsha, "Genius POG Blue Book" by Media Boy, "POG Book for Horse Racing King" by Guideworks, "Complete POG" by Sankei Sports, etc. Also, during this period, many racing magazines and racing specialized newspapers, for example, "Yushun" by Central Racing P.R. Center, "Sarabure" by KADOKAWA, etc., feature special articles on POG.

[0004] In POG, the ability to successfully select 2-year-old horses that are likely to obtain a large amount of prize money in the future depends greatly on the technology of efficiently sorting through a large amount of information such as the Internet in addition to the information in books and magazines and selecting important information. However, the information on 2-year-old horses provided by various media mainly consists of comments by ranch, stable-related persons, horse owners, etc., and subjective evaluations by reporters, critics, etc., and lacks objectivity.

[0005] Compiled by the Editorial Department of the Horse Racing King, "Introduction to Unbeatable POG" published by Hakuya Shobo in 2009 describes a method of narrowing down candidate horses by analyzing objective data such as the father, mother, mother's father, gender, trainer, owner, breeder, date of birth, transaction price, and asking price of racehorses. In the book, the index of prize money per race is used to compare each factor such as the father, gender, trainer, owner, breeder, and date of birth of racehorses.

[0006] The prize money per race is, for example, when organizing data by the factor of trainer, the value obtained by dividing the total prize money during the POG period of racehorses belonging to each trainer's stable in the past several generations by the total number of starts during the POG period. For example, when evaluating 2-year-old horses debuting in 2009, if calculating the prize money per race for the past 5 generations, it targets the 2-year-old races from 2004 to the 3-year-old races including the Tokyo Yushun in 2005 for horses born in 2002, the 2-year-old races from 2005 to the 3-year-old races including the Tokyo Yushun in 2006 for horses born in 2003, the 2-year-old races from 2006 to the 3-year-old races including the Tokyo Yushun in 2007 for horses born in 2004, the 2-year-old races from 2007 to the 3-year-old races including the Tokyo Yushun in 2008 for horses born in 2005, and the 2-year-old races from 2008 to the 3-year-old races including the Tokyo Yushun in 2009 for horses born in 2006.

[0007] The prize money per race for each factor can be obtained by importing data into a computer via the Internet from the horse racing information service JRA-VAN Data Lab. (registered trademark), which uses JRA official data provided by JRA System Services Co., Ltd., and using the corresponding application program, such as "TARGET frontier JV".

[0008] The author emphasizes three factors in particular among each factor, namely the father, breeder, and trainer of racehorses, and recommends a method of narrowing down by excluding from the nomination candidates if there is even one factor with a low prize money per race.

Prior Art Documents

Non-Patent Documents

[0009] [Non-Patent Document 1] Compiled by the editors of "Unbeatable POG Introduction", published by Byakuya Shobo in 2009 [Summary of the Invention] [Problems to be Solved by the Invention]

[0010] In order to select a promising horse with a high possibility of winning a large amount of prize money in POG, as a preliminary preparation, it is efficient if candidate horses can be narrowed down or ranked using some objective indicators.

[0011] For example, by comparing the performance of several generations in the past using factors such as the father, grandfather, gender, trainer, owner, breeder, transaction price, mother's racing performance, birth date, and mother's age at birth of racehorses, and ranking all horses, candidate horses can be selected efficiently.

[0012] However, there are about 7,000 racehorses registered with pedigrees in Japan every year, and among them, about 4,000 racehorses debut in central horse racing every year. There was a problem that it took an enormous amount of time to compare each factor of each horse in detail and comprehensively.

[0013] In addition, among the multiple factors that a single racehorse has, which factor to emphasize when comparing each horse depends on the judgment of the game participants in the existing method. Therefore, even if objective data such as prize money per race is used, it has been difficult to accurately predict the ranking of the future performance of racehorses.

[0014] The problem to be solved by the present invention is to provide a data structure that can obtain certain prediction results without relying on the experience and sense of game participants in POG, and by using it for predicting the ranking of the future performance of racehorses, it supports game participants to select candidate horses more efficiently. [Means for Solving the Problems]

[0015] The data structure used for predicting the performance of racehorses according to the present invention is a data structure including at least three or more factors selected from the father of the racehorse, the father of the mother, gender, trainer, owner, breeder, transaction price, the racing performance of the mother, date of birth, and the age of the mother at birth. Each data consists of numerical data based on the performance of racehorses in past generations bracketed for each of the factors, and a score value capable of relatively comparing each racehorse is output by causing a computer to perform an operation based on the weighting coefficient of each data.

[0016] Further, the data structure according to the present invention uses the data structure of racehorses in past generations from the target generation for which performance prediction is to be performed as learning data, sets each data as an explanatory variable, and causes a computer to learn using information representing the performance of racehorses in past generations as an objective variable, thereby obtaining the weighting coefficient.

[0017] Further, as a method for determining the weighting coefficient, the data structure according to the present invention is such that the weighting coefficient is determined so that the integrated value from i = 1 to n of the difference between the total prize money from the first to the i-th (where i = 1, 2, 3, ···, n (n ≦ m)) horses arranged in descending order of prize money among m racehorses in past generations and the total prize money from the first to the i-th horses arranged in descending order of the score value is minimized.

Effect of the Invention

[0018] According to the present invention, it is possible to provide a data structure used for predicting the ranking of the performance of racehorses, which can obtain a certain prediction result without relying on the experience and sense of game participants in POG. As a result, it is expected that users can efficiently select high-quality horses with a high possibility of winning a large amount of prize money in POG using objective indicators.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Embodiment for Carrying Out the Invention

[0020] The data structure according to the present invention causes a computer that has taken in the data structure to perform an operation and output a score value. The data structure of the present invention is not a program that directly commands a computer, but defines the processing of the computer by the mutual relationship between data elements and built-in arithmetic expressions, and functions the computer in combination with other programs.

[0021] First, the configuration of the computer A that the data structure according to the embodiment of the present invention functions will be described. FIG. 1 is a schematic block diagram showing the functional configuration of the computer A that is functioned by the data structure. The computer A includes an input / output interface unit 10, a control unit 20, and a storage unit 30.

[0022] The input / output interface unit 10 is composed of an interface, an input device, and a display device (none of which are shown). The interface transmits and receives information to and from an external database via a wired or wireless communication network. The input device is a device including a keyboard, a mouse, etc. for inputting various data input by the user to the control unit 20. The display device is a display device such as a liquid crystal display for presenting the display data output from the control unit 20 to the user.

[0023] The control unit 20 is composed of a hardware processor, a CPU (Central Processing Unit), and a main storage device that is a program memory (none of which are shown).

[0024] The memory unit 30 is composed of an auxiliary storage device (not shown). The auxiliary storage device uses a non-volatile memory such as an HDD (Hard Disc Drive) or an SSD (Solid State Drive) that can be written to and read from at any time as a storage medium, and also includes those using a magneto-optical disk, a CD-ROM, a DVD-ROM, etc.

[0025] The data structure according to the embodiment of the present invention may be stored in the auxiliary storage device. The CPU reads the data structure from the auxiliary storage device, expands it in the main storage device, and performs predetermined processing defined by the data structure.

[0026] Also, when the data structure according to the embodiment of the present invention is distributed to computer A via a communication line, the CPU of computer A that has received the distribution via the interface may expand the data structure in the main storage device and execute predetermined processing.

[0027] Functionally, computer A is configured to include a supplementary data acquisition unit 21, a data structure supplement unit 22, a score calculation unit 23, and an output control unit 24 in the control unit 20, and a data structure storage unit 31 in the memory unit 30. The memory unit 30 may include a factor data storage unit 32.

[0028] The data structure storage unit 31 stores a data structure that is a performance prediction model for racehorses of the target generation for which performance prediction is to be performed. The data structure stored in the data structure storage unit 31 includes profile information regarding racehorses of the target generation that is minimally required to calculate a score value representing the likelihood of occurrence of a racehorse with a high prize money, numerical data corresponding to each factor, a weight coefficient used for score value calculation, and a score value calculation formula.

[0029] The factor data storage unit 32 stores numerical data based on the performance of racehorses in past generations grouped by factor, such as prize money data per race.

[0030] The supplementary data acquisition unit 21 acquires data for supplementing the profile of the racehorse of the target generation for which performance prediction is to be made from an external database, for example, the JRA-VAN data laboratory of the horse racing information service, via the Internet through the input / output interface unit 10. Alternatively, the same type of data may be read from a computer-readable recording medium such as a magnetic disk or an optical disk via the storage unit 30.

[0031] The data structure supplementing unit 22 reads out the data structure stored in the data structure storage unit 31 of the storage unit 30, combines the data acquired by the supplementary data acquisition unit 21 with the data structure, and performs a process of supplementing the profile. Alternatively, the data structure supplementing unit 22 also functions when the profile information of the data structure is corrected or supplemented by manual operation of an input device such as a keyboard or a mouse by the user, and also functions when the weighting factor used to calculate the score value is directly input. Further, the data structure supplementing unit 22 reads out the numerical data stored in the factor data storage unit 32 of the storage unit 30, extracts the numerical data corresponding to the profile supplemented by the above procedure, and performs a process of incorporating it into the data structure.

[0032] The score calculation unit 23 calculates a score value representing the ease of occurrence of a racehorse with a high prize money by performing an operation based on the weighting factor of each data in the data structure. Further, the score calculation unit 23 performs a process of incorporating the calculated score value into the data structure.

[0033] The output control unit 24 creates output data based on the score value calculated by the score calculation unit 23, and performs a process of outputting it to a display device or an external terminal via the input / output interface unit 10. For example, the output control unit 24 can create a priority list with priorities assigned to the nominated candidate horses in descending order of the score value as output data.

[0034] Hereinafter, an embodiment in which prize money data per race is used as numerical data corresponding to each factor will be described as an example, but the present invention is not limited thereto. Various numerical data based on the performance of racehorses can be used without departing from the gist of the present invention.

[0035] A computer B used for creating a data structure according to an embodiment of the present invention will be described. FIG. 2 is an example of a schematic block diagram and a system configuration diagram of the computer B. The computer B includes, as hardware, an input / output interface unit 10, a control unit 20, and a storage unit 30.

[0036] The detailed configuration of each unit is the same as that of the computer A that enables the data structure to function. The computer B used for creating the data structure may be the same as the computer A that enables the data structure to function, or may be another computer.

[0037] The storage unit 30 includes, as a storage area necessary for realizing the present embodiment, a performance data storage unit 301, a prize money data storage unit 302 per race, a learning data storage unit 303, a learned data storage unit 304, a score calculation data storage unit 305, and a data structure storage unit 306.

[0038] The performance data storage unit 301 stores performance data D1 in which profile information on past racehorses whose POG period has already ended is associated with information representing performance such as the number of starts and the prize money obtained during the period.

[0039] The prize money data storage unit 302 per race stores a prize money data set D2 per race including prize money information per race for each divided group of each factor, which is calculated based on the performance data D1.

[0040] The learning data storage unit 303 stores learning data D3, which includes the profile information of past racehorses extracted from the performance data D1 and the per-race prize money information corresponding to each factor extracted from the per-race prize money data set D2.

[0041] The learned data storage unit 304 stores the weight coefficient W calculated as a result of generating a prediction model using the learning data D3, and the same score value calculation formula as the prediction model.

[0042] The score calculation data storage unit 305 stores score calculation data D5, which is used to create a prediction model for the racehorses of the target generation for which performance prediction is to be performed.

[0043] The data structure storage unit 306 stores the data structure that is the performance prediction model for the racehorses of the target generation for which performance prediction is to be performed.

[0044] The control unit 20 includes a performance data acquisition unit 201, a per-race prize money calculation unit 202, a learning data creation unit 203, a learning data acquisition unit 204, a learning unit 205, a prediction data acquisition unit 206, a score calculation data creation unit 207, a data structure creation unit 208, and a score calculation unit 209 to execute the processing functions in this embodiment. The processing functions in each of these units are all realized by causing the hardware processor to execute a program stored in the program memory. Note that these processing functions may be realized not by using a program stored in the program memory but by using a program provided through a network.

[0045] The performance data acquisition unit 201 acquires, via the input / output interface unit 10, from an input device, an external database, etc., the profile information regarding past racehorses whose POG period has already ended, and information representing performance such as the number of starts during a period and the main prize money obtained during the period, creates performance data D1 by associating them, and stores it in the performance data storage unit 301.

[0046] The per-race prize calculation unit 202 reads out the performance data D1 stored in the performance data storage unit 301 of the storage unit 30, and executes a process of generating a data set D2 representing the per-race prize for each group of respective factors. The per-race prize calculation unit 202 may calculate the per-race prize from all the acquired past data, or may calculate the per-race prize from data of an arbitrary number of generations. For each factor, the number of generations of data used for calculating the per-race prize may be different.

[0047] The learning data creation unit 203 reads out the performance data D1 stored in the performance data storage unit 301 of the storage unit 30 and the per-race prize data set D2 stored in the per-race prize data storage unit 302, and performs a process of creating learning data D3 used for generating a prediction model for predicting the performance of racehorses. The learning data D3 uses the performance data D1 of racehorses in past generations from the target generation for which performance prediction is to be performed. Further, in the learning data D3, the per-race prize data set D2 calculated using the performance data of generations one or more generations further in the past than the generation targeted by the performance data D1 is combined with the performance data D1 and used. The learning data creation unit 203 stores the created learning data D3 in the learning data storage unit 303.

[0048] The learning data acquisition unit 204 reads out the data stored in the learning data storage unit 303 of the storage unit 30, and performs a process of generating one learning data D3 obtained by combining a plurality of learning data D3 used for generating a prediction model for predicting the performance of racehorses.

[0049] The learning unit 205 executes a process of performing statistical analysis using the learning data D3. For example, the learning unit 205 sets the prize money per race in the learning data D3 as an explanatory variable, and further, using the prize money, which is information representing the results of past generations of racehorses included in the data set, as a target variable, executes a process of optimizing the weight coefficient corresponding to each factor for calculating a score value representing the likelihood of occurrence of racehorses with a higher target variable from the explanatory variable. The obtained weight coefficient W is stored in the learned data storage unit 304. Also, the obtained weight coefficient W can be used for prediction processing by being incorporated into the prediction model. Further, the same score value calculation formula used for the process of optimizing the weight coefficient W is stored in the learned data storage unit 304.

[0050] The prediction data acquisition unit 206 acquires data representing profile information regarding a racehorse of the target generation for which performance prediction is to be performed from an input device, an external database, etc. via the input / output interface unit 10, and creates prediction data D4. The prediction data acquisition unit 206 reads out the prize money per race data set D2 stored in the prize money per race data storage unit 302, which is calculated using the performance data of generations one or more generations past the target generation for which performance prediction is to be performed corresponding to the prediction data D4.

[0051] The score calculation data creation unit 207 performs a process of creating score calculation data D5 for generating a prediction model for predicting the performance of a racehorse, using the prediction data D4 created by the prediction data acquisition unit 206 and the prize money per race data set D2. The score calculation data creation unit 207 stores the created score calculation data D5 in the score calculation data storage unit 305.

[0052] The data structure creation unit 208 generates a data structure, which is a prediction model, by incorporating the weight coefficient W stored in the learned data storage unit 304 and the score value calculation formula into the score calculation data D5 stored in the score calculation data storage unit 305.

[0053] The score calculation unit 209 calculates a score value representing the likelihood of a racehorse with a high prize money by performing an operation based on the weight coefficient of each data in the data structure. Further, the score calculation unit 209 performs a process of incorporating the calculated score value into the data structure. The score calculation unit 209 stores the created data structure in the data structure storage unit 306.

[0054] The data structure according to the embodiment of the present invention will be described. The data structure is created by using a spreadsheet software such as EXCEL (registered trademark) of Microsoft Corporation.

[0055] An example of the data structure is shown in FIG. 3. In one row, so-called profiles such as the name of a racehorse, date of birth, gender, father's name, mother's name, mother's age at birth, mother's father's name, breeder, trainer, owner, and transaction price are represented. The data structure of the present invention includes at least three or more factors representing such profiles.

[0056] Among these, information such as the name of the racehorse, trainer, and owner is often undetermined for horses that have not yet raced and may be blank. Even if the name of the racehorse is undetermined, since usually only one horse is born to a mare in a year, the racehorse can be identified by the mother's name.

[0057] Also, since the transaction price often does not become public except when there is a market transaction, it may be blank.

[0058] Each data constituting the data structure is numerical data corresponding to each factor representing these profiles. As shown in FIG. 3, it is preferable to write each data side by side closest to the factor.

[0059] Further, the data structure may include a score value that is the result of an operation based on the weight coefficient of each data. The operation is based on an arbitrarily determined combination of expressions or functions, and the purpose is to relatively compare each racehorse by comparing the magnitudes of the score values.

[0060] Next, each factor representing the profile according to the embodiment of the present invention will be described. Each data consists of numerical data based on the performance of racehorses in past generations grouped by each factor. Hereinafter, the calculation method of the numerical data corresponding to each individual factor will be described.

[0061] The father of a racehorse, that is, a factor meaning a sire. The racehorses in past generations grouped by the father of a racehorse represent a group of foals having the same sire in generations of 3 years old and above corresponding to past generations from the 2-year-old generation. Here, a generation means a group of racehorses born in the same year.

[0062] To calculate the numerical data based on the performance of racehorses in past generations grouped by the father of a racehorse, performance data for a sufficient number of generations is used. As a preferable number of generations, performance data of racehorses in 3 to 10 generations including the immediately previous generation with respect to the non-started generation for which the ranking prediction of performance is to be made is used. If the number of generations is too small, a sufficient amount of data with sufficient reliability cannot be obtained. Conversely, if the number of generations is too large, past performance is strongly reflected, and there is a tendency to overestimate even if the performance of recent foals is poor.

[0063] The format of the numerical data is not limited to the prize money per race. For example, to calculate the prize money per race as numerical data, the total sum of the main prize money during the POG period of racehorses in a predetermined past generation grouped by the father of a racehorse is divided by the total number of starts during the POG period. For example, if the total main prize money of all foals having sire A is 1 billion yen and the total number of starts is 1000 times, the numerical data will be 1 million yen.

[0064] When obtaining the numerical data, if the total number of starts is too small, the calculation result from individual prize money data becomes unstable, and reliable numerical data cannot be obtained. If it does not reach the predetermined total number of starts, the numerical data of the father factor of the racehorse may be regarded as non-existent. As the predetermined total number of starts, preferably, it is 30 starts or more to 100 starts or more as the boundary. Even in the case of a new sire where the target generation for which the ranking prediction of performance is to be made is the first generation in which the racehorse debuts, since there are no past generations, there is no numerical data.

[0065] The sire of a racehorse means a factor that corresponds to the grandsire on the dam side of the racehorse. The racehorses of past generations enclosed by the sire of the racehorse represent a group of foals born in generations of three years old or older, which correspond to past generations from the two-year-old generation, and have the same stallion as the grandsire on the dam side. Here, a generation means a group of racehorses born in the same year.

[0066] To calculate numerical data based on the performance of racehorses of past generations enclosed by the sire of the racehorse, performance data for a sufficient number of generations is used. However, since mares are not only produced domestically but may also be imported from foreign countries, the sires of racehorses are more diverse and numerous than the sires of racehorses. Therefore, the total number of starts of racehorses with the same sire per generation tends to be less than the total number of starts of racehorses with the same sire. Therefore, it is preferable to use performance data from more past generations when calculating numerical data than when the sire of the racehorse is a factor. As a preferable number of generations, performance data of racehorses in 5 to 20 generations, including the immediately preceding generation, is used for the unstarted generation for which the ranking prediction of performance is to be made. If the number of generations is too small, a sufficient amount of data for reliability cannot be obtained. Conversely, if the number of generations is too large, past performance is strongly reflected, and there is a tendency to overestimate even if the performance of recent foals is poor.

[0067] The form of the numerical data is not limited to the prize money per race. For example, to calculate the prize money per race as numerical data, it is obtained by dividing the total sum of the main prize money during the POG period of racehorses of a predetermined past generation enclosed by the sire of the racehorse by the total number of starts during the POG period. For example, if the total main prize money of all foals with stallion B as the sire is 1 billion yen and the total number of starts is 1000 times, the numerical data will be 1 million yen.

[0068] When obtaining numerical data, if the total number of starts is too small, the calculation results from individual prize money data will become unstable, and reliable numerical data cannot be obtained. If the total number of starts does not reach the predetermined number, the numerical data of the sire and dam factors of the racehorse may be regarded as non-existent. As the predetermined total number of starts, preferably, it is bounded by 30 starts to 200 starts or more. Even in the case of the sire and dam of the target generation for which the prediction of the order of results is to be made and which is the first generation in which the racehorse makes its debut, since there is no past generation, the numerical data is regarded as non-existent.

[0069] Numerical data may be calculated using the gender of the racehorse as a factor. Although there are three types of genders for racehorses: stallions, mares, and geldings, since it is rare for a stallion to become a gelding among 2-year-old racehorses before debut, stallions and geldings may be treated as a single factor.

[0070] The past-generation racehorses grouped by the gender of the racehorse represent the foal groups of the group grouped by stallions / geldings and the group grouped by mares in generations 3 years old and above, which correspond to past generations from the 2-year-old generation. Here, a generation refers to racehorses born in the same year as one generation.

[0071] To calculate the numerical data based on the results of past-generation racehorses grouped by the gender of the racehorse, performance data for a sufficient number of generations is used. As the preferred number of generations, performance data for racehorses in 5 to 20 generations, including the immediately preceding generation, is used for the non-started generation for which the prediction of the order of results is to be made.

[0072] The form of the numerical data is not limited to the prize money per start. For example, to calculate the prize money per start as numerical data, the total sum of the total main prize money during the POG period of past-generation racehorses grouped by the gender of the racehorse is divided by the total number of starts during the POG period. For example, if the total main prize money of all foals of stallions / geldings is 100 billion yen and the total number of starts is 100,000 times, the numerical data will be 1 million yen.

[0073] Alternatively, numerical data may be obtained by combining gender with other factors. For example, it is known that depending on the sire, the performance of the foals is biased by gender. Therefore, fillies or colts with the same sire as the racehorse, or fillies with the same sire as the racehorse, may be grouped and numerical data may be calculated separately for each group. In this case, except for grouping by fillies or colts, numerical data is calculated in the same way as in the case where the sire of the racehorse is a factor.

[0074] The trainer of a racehorse is a factor that means the stable to which the racehorse belongs. The racehorses of past generations grouped by the trainer of the racehorse represent a group of foals born in the 2-year-old generation and belonging to the same stable in generations 3 years old and above, which are past generations. Here, a generation means a group of racehorses born in the same year.

[0075] To calculate numerical data based on the performance of racehorses of past generations grouped by the trainer of the racehorse, performance data for a sufficient number of generations is used. As a preferable number of generations, performance data for 3 to 10 generations of racehorses including the immediately previous generation is used for the unraced generation for which the order prediction of performance is to be made. If the number of generations is too small, a sufficient amount of data for reliability cannot be obtained. Conversely, if the number of generations is too large, past performance is strongly reflected, and there is a tendency to overestimate even if recent performance is poor.

[0076] The form of the numerical data is not limited to the prize money per race. For example, to calculate the prize money per race as numerical data, it is obtained by dividing the total sum of the main prize money during the POG period of racehorses of a predetermined past generation grouped by the trainer of the racehorse by the total number of starts during the POG period. For example, if the total main prize money of all foals managed by trainer C is 1 billion yen and the total number of starts is 1000, the numerical data will be 1 million yen.

[0077] When obtaining numerical data, if the total number of starts is too small, the calculation results from individual prize money data will become unstable, and reliable numerical data cannot be obtained. If the total number of starts does not reach the predetermined number, the numerical data of the trainer factor of the racehorse may be regarded as non-existent. As the predetermined total number of starts, preferably, it is bounded by 30 starts to 100 starts or more. In the case of a novice trainer whose target generation for which the order prediction of results is to be made is the first generation in which the racehorse makes its debut, since there is no past generation, there is no numerical data. Also, for 2-year-old horses before debut, in many cases, the trainer has not been determined or, even if determined, has not been announced. Even when the trainer is undetermined in this way, it may be regarded as having no numerical data.

[0078] Numerical data may be calculated using the owner of the racehorse as a factor. The past-generation racehorses grouped by the owner of the racehorse represent a group of foals owned by the same owner in generations of 3 years old or more corresponding to past generations from the 2-year-old generation. Here, a generation means a group of racehorses born in the same year.

[0079] To calculate the numerical data based on the results of past-generation racehorses grouped by the owner of the racehorse, performance data of a sufficient number of generations are used. As the preferred number of generations, performance data of 3 to 10 generations of racehorses including the immediately preceding generation to the non-started generation for which the order prediction of results is to be made are used. If the number of generations is too small, a sufficient amount of reliable data cannot be obtained. Conversely, if the number of generations is too large, past results are strongly reflected, and there is a tendency to overestimate even if recent results are poor.

[0080] The form of the numerical data is not limited to the prize money per start. For example, to calculate the prize money per start as numerical data, the total sum of the total main prize money during the POG period of past-generation racehorses grouped by the owner of the racehorse is divided by the total number of starts during the POG period. For example, if the total main prize money of all the foals owned by Owner D is 1 billion yen and the total number of starts is 1000 times, the numerical data will be 1 million yen.

[0081] When obtaining numerical data, if the total number of starts is too small, the calculation results from individual prize money data will become unstable, and reliable numerical data cannot be obtained. If the total number of starts does not reach a predetermined number, the numerical data of the horse owner factor of the racehorse may be regarded as non-existent. As the predetermined total number of starts, preferably, it is bounded by 30 starts to 100 starts or more. In the case of a new horse owner whose target generation for which the order prediction of results is to be made is the first generation in which the racehorse makes its debut, since there is no past generation, there is no numerical data. For 2-year-old horses before debut, in many cases, the horse owner has not been determined or, even if determined, has not been announced. Thus, even when the horse owner is undetermined, it may be regarded as having no numerical data.

[0082] The producer of a racehorse is a factor that means the production farm of the racehorse. The racehorses of past generations covered by the producer of the racehorse represent a group of foals produced at the same farm in generations of 3 years old or more corresponding to past generations from the 2-year-old generation. Here, a generation means a group of racehorses born in the same year.

[0083] To calculate numerical data based on the results of racehorses of past generations covered by the producer of the racehorse, performance data of a sufficient number of generations is used. As the preferable number of generations, performance data of 3 to 10 generations of racehorses including the immediately previous generation to the non-started generation for which the order prediction of results is to be made is used. If the number of generations is too small, a sufficient amount of reliable data cannot be obtained. Conversely, if the number of generations is too large, past results are strongly reflected, and even if recent results are poor, there is a tendency to be overestimated.

[0084] The form of the numerical data is not limited to the prize money per start. For example, to calculate the prize money per start as numerical data, it is obtained by dividing the total sum of the total prize money during the POG period of racehorses of a predetermined past generation covered by the producer of the racehorse by the total number of starts during the POG period. For example, if the total prize money of all foals produced by producer E is 1 billion yen and the total number of starts is 1000 times, the numerical data will be 1 million yen.

[0085] When obtaining numerical data, if the total number of races is too small, the calculation results from individual prize money data will become unstable, and reliable numerical data cannot be obtained. If the total number of races does not meet the predetermined number, the numerical data of the producer factor of the racehorse may be regarded as non-existent. As the predetermined total number of races, preferably, it is bounded by 30 races to 100 races or more. In the case of a new producer whose target generation for which the order prediction of results is to be made is the first generation in which the racehorse debuts, since there is no past generation, there is no numerical data.

[0086] Numerical data may be calculated using the transaction price of racehorses as a factor. The data provided by JRA-VAN includes the transaction price data of the racehorse transaction market held in Japan, and the transaction price of racehorses can be referred to. The transaction price of racehorses is divided into several price ranges, and numerical data based on the results of past-generation racehorses enclosed by the price range of racehorses is calculated. Examples of price ranges are divided as less than 1 million yen, less than 1 million to 5 million yen, less than 5 million to 10 million yen, less than 10 million to 20 million yen, less than 20 million to 30 million yen, less than 30 million to 50 million yen, less than 50 million to 70 million yen, less than 70 million to 100 million yen, 100 million yen or more.

[0087] The past-generation racehorses enclosed by the price range of racehorses represent the foal groups of the same price range in the 3-year-old and older generations corresponding to the past generations from the 2-year-old generation. Here, a generation refers to racehorses born in the same year as one generation.

[0088] To calculate the numerical data based on the results of past-generation racehorses enclosed by the price range of racehorses, performance data of a sufficient number of generations is used. As the preferable number of generations, performance data of 5 to 20 generations of racehorses including the immediately previous generation for the non-raced generation for which the order prediction of results is to be made is used. If the number of generations is too small, a sufficient amount of reliable data cannot be obtained. On the contrary, if the number of generations is too large, since the transaction price of racehorses is affected by the economic situation, the correlation between the transaction price and the results will deviate greatly between recent data and past data, resulting in problems.

[0089] Although not limited to the prize money per race as the form of numerical data, for example, to calculate the prize money per race as numerical data, the total sum of the main prize money during the POG period of racehorses in a predetermined past generation bracketed by the price range of racehorses is obtained by dividing it by the total number of starts during the POG period. For example, if the total main prize money of all foals with a transaction price of 10 million to 20 million yen is 10 billion yen and the total number of starts is 10,000 times, the numerical data will be 1 million yen.

[0090] When obtaining numerical data, if the total number of starts is too small, the calculation results from individual prize money data will become unstable, and reliable numerical data cannot be obtained. Avoid divisions where a price range with less than 100 starts occurs.

[0091] The numerical data may be calculated using the racing results of the mother of the racehorse as a factor. The racing results of the mother of the racehorse are divided into several result groups, and numerical data based on the results of racehorses in past generations bracketed by the racing result group of the mother of the racehorse is calculated. As a method of dividing into racing result groups, for example, there is a method of dividing by the total main prize money amount or the prize money per race during the POG period.

[0092] The racehorses in past generations bracketed by the racing result group of the mother of the racehorse represent a group of foals born from the 2-year-old generation to past generations of 3-year-old and older generations, whose mothers belonged to the same result group. Here, a generation refers to a group of racehorses born in the same year.

[0093] To calculate numerical data based on the results of racehorses in past generations bracketed by the racing results of the mother of the racehorse, performance data for a sufficient number of generations is used. As a preferable number of generations, performance data for 5 to 20 generations of racehorses including the immediately preceding generation is used for the unstarted generation for which the ranking prediction of the results is to be made. If the number of generations is too small, a sufficient amount of reliable data cannot be obtained.

[0094] Although not limited to the prize money per race as the format of numerical data, for example, to calculate the prize money per race as numerical data, the total sum of the prize money during the POG period of racehorses of a predetermined past generation grouped by the race performance group of the mother of the racehorse is obtained by dividing it by the total number of starts during the POG period. For example, if the total prize money of all the foals born to the mother of racehorses with a total prize money of 5 million to 10 million yen during the POG period is 10 billion yen and the total number of starts is 10,000 times, the numerical data will be 1 million yen.

[0095] When obtaining numerical data, if the total number of starts is too small, the calculation results from individual prize money data will become unstable, and reliable numerical data cannot be obtained. Avoid divisions where a race performance group of the mother of racehorses with less than 100 starts occurs. If the mother of the racehorse has no domestic race performance due to reasons such as being imported from overseas, it may be regarded as having no numerical data.

[0096] The numerical data may be calculated using the birth date of the racehorse as a factor. The birth date is divided into several periods, and the numerical data based on the performance of racehorses of past generations grouped within the same birth period is calculated. The simplest method is to divide by birth month. The racehorses of past generations grouped within the birth period represent a group of foals born from the 2-year-old generation to past generations of 3-year-old and older generations, whose birth dates were divided and belonged to the same period. Here, a generation refers to racehorses born in the same year as one generation.

[0097] To calculate the numerical data based on the performance of racehorses of past generations grouped within the birth period, performance data of a sufficient number of generations is used. As a preferable number of generations, performance data of 5 to 20 generations of racehorses including the immediately previous generation to the non-starting generation for which the ranking prediction of performance is to be made is used. If the number of generations is too small, a sufficient amount of reliable data cannot be obtained.

[0098] Although it is not limited to the prize money per race as the format of numerical data, for example, to calculate the prize money per race as numerical data, it is obtained by dividing the total sum of the main prizes during the POG period of racehorses of a predetermined past generation bracketed by the birth period by the total number of starts during the POG period. For example, if the total main prize money of all foals born in April is 10 billion yen and the total number of starts is 10,000 times, the numerical data will be 100,000 yen.

[0099] When obtaining numerical data, if the total number of starts is too small, the calculation results from individual prize money data will become unstable and reliable numerical data cannot be obtained. Avoid divisions where the birth period results in a total number of starts of less than 100 races. Racehorses born from July to December, for example, correspond to horses produced in the southern hemisphere, but the number of horses registered as racehorses in Japan is relatively small, and if divided by individual birth months, the total number of starts may be less than 100 races. To avoid the occurrence of divisions with a small total number of starts, for example, racehorses born from July to December can be grouped as one group.

[0100] Numerical data may be calculated using the age of the mother at the time of foaling of the racehorse as a factor. Numerical data is calculated based on the performance of racehorses of past generations bracketed by the age of the mother at the time of foaling of the racehorse. The racehorses of past generations bracketed by the age of the mother at the time of birth of the racehorse represent a group of foals born to mothers of the same age at the time of foaling in generations of 3 years old and above corresponding to past generations from the 2-year-old generation. Here, a generation refers to racehorses born in the same year as one generation.

[0101] To calculate numerical data based on the performance of racehorses of past generations bracketed by the age of the mother at the time of birth of the racehorse, performance data of a sufficient number of generations is used. As a preferable number of generations, performance data of racehorses of 5 to 20 generations including the immediately preceding generation to the unstarted generation for which the order prediction of performance is to be made is used. If the number of generations is too small, a sufficient amount of reliable data cannot be obtained.

[0102] Although not limited to the prize money per race as the format of numerical data, for example, to calculate the prize money per race as numerical data, the total sum of the main prize money during the POG period of racehorses in a predetermined past generation bracketed by the mother's age at the time of the birth of the racehorse is obtained by dividing it by the total number of starts during the POG period. For example, if the total main prize money of all foals born to a mother aged 10 at the time of birth is 10 billion yen and the total number of starts is 10,000, the numerical data will be 1 million yen.

[0103] When obtaining numerical data, if the total number of starts is too small, the calculation results from individual prize money data will become unstable and reliable numerical data cannot be obtained. Racehorses born to mothers aged 20 or older are relatively few, and when classified by individual mother's age, the total number of starts may be less than 100. To avoid the occurrence of a category with a small total number of starts, for example, foals born to mothers aged 20 or older can be grouped as one group.

[0104] Next, the weight coefficients of each data used in the calculation according to the embodiment of the present invention will be described. By causing a computer to perform an operation based on the weight coefficient on the numerical data for each factor, a score value that can relatively compare each racehorse is output.

[0105] As an operation method, for example, there is a method of calculating the weighted average of the numerical data of each factor. For the numerical data x1, x2, ···, xf of f factors, when the respective weight coefficients are w1, w2, ···, wf, the operation method is represented by Equation 1. Equation 1 ··· (w1x1 + w2x2 + ··· + wfxf) / (w1 + w2 + ··· wf)

[0106] As an operation method, for example, there is a method of calculating the weighted harmonic average of the numerical data of each factor. For the numerical data x1, x2, ···, xf of f factors, when the respective weight coefficients are w1, w2, ···, wf, the operation method is represented by Equation 2. Equation 2 ··· (w1 + w2 + ··· + wf) / (w1 / x1 + w2 / x2 + ··· wf / xf)

[0107] When adopting weighted average values such as weighted average and weighted harmonic average, the greater the factor with a larger weight coefficient, the higher the importance. The higher the score value of the calculation result, the higher the ranking of the predicted sequence result of the performance, which means that the probability of a racehorse with a higher performance ranking occurring is high.

[0108] For the weight coefficient, for past generations whose performance during the POG period has already been determined, optimized numerical values are used so that the result of predicting the ranking of the performance is closest to the best result.

[0109] As a method for optimizing the weight coefficient, the sum of the prizes from the first to the i-th (where i = 1, 2, 3, ···, n (n ≤ m)) horses arranged in descending order of the prize money of m racehorses in past generations is compared with the sum of the prizes from the first to the i-th horses arranged in descending order of the score value based on the weight coefficient of each data. The weight coefficient of each data is determined so that the integrated value from i = 1 to n of the difference is minimized.

[0110] If the prize money of the k-th horse arranged in descending order of the prize money is represented by a k then the sum of the prizes from the first to the i-th horses is expressed by Equation A. TIFF0007691606000001.tif1779

[0111] On the other hand, if the prize money of the k-th horse arranged in descending order of the score value of the calculation result based on the weight coefficient of each data is represented by b k then the sum of the prizes from the first to the i-th horses is expressed by Equation B. TIFF0007691606000002.tif1779

[0112] The difference between the sum of the prizes from the first to the i-th horses arranged in descending order of the prize money and the sum of the prizes from the first to the i-th horses arranged in descending order of the score value of the calculation result based on the weight coefficient of each data is (Equation A - Equation B).

[0113] Therefore, the integrated value from i = 1 to n of the difference between the total prize money from the 1st to the i-th horse arranged in descending order of the amount of this prize money and the total prize money from the 1st to the i-th horse arranged in descending order of the score value of the calculation result based on the weight coefficient of each data is expressed by Equation C. TIFF0007691606000003.tif1779

[0114] Here, arranging the m racehorses of past generations in descending order of the amount of this prize money means the arrangement order in which the prediction sequence result of the performance is the best. That is, the total prize money from the 1st to the i-th horse in the arrangement represented by Equation A is the largest among all possible arrangement orders, which is i!.

[0115] The closer the prediction sequence result of the performance of past generations is to the best, the larger the total prize money from the 1st to the i-th horse arranged in descending order of the score value of the calculation result based on the weight coefficient of each data represented by Equation B will be, but it will not exceed the total prize money from the 1st to the i-th horse actually arranged in descending order of the amount of this prize money represented by Equation A.

[0116] For example, by assuming the weight coefficients of each factor step by step, calculating the integrated value represented by Equation C, and having the computer execute a program to compare them, the optimal weight coefficient that minimizes the integrated value can be obtained.

[0117] In this case, as the weight coefficients w1, w2, ···, wf, natural numbers from 1 to 10 or decimals from 0.1 to 1.0 can be adopted. For example, when calculating numerical values from 1 to 10 as candidates for the weight coefficients, the f weight coefficients (w1, w2, ···, wf) are changed step by step from (1, 1, ···, 1) to (10, 10, ···, 10), and 10 f combinations of the integrated values represented by Equation C are calculated. The combination with the minimum calculation result is adopted as the weight coefficient of each factor.

[0118] As a method for selecting m racehorses of past generations, it is preferable to select racehorses of 5 to 20 generations including the immediately previous generation for the generation for which the order prediction of performance is to be performed. It is also possible to select all the horses of the generation, or it is also possible to extract and select several hundred top-ranked horses in terms of performance in each generation so that the data does not become too large.

[0119] By using the data structure incorporating the weight coefficients thus obtained, certain prediction results that are convenient for predicting the order of performance of unraced generations can be obtained.

Example

[0120] Next, the processing procedure for creating the data structure of the foals born in 2017 will be described as an example. To calculate the weight coefficients, learning data of foals born from 2012 to 2016 corresponding to the past five generations was used. To calculate the prize money per race as numerical data for each factor, performance data corresponding to the past ten generations from that generation was used. That is, the prize money per race for the data structure of the foals born in 2017 was the performance data of the foals born from 2007 to 2016, the prize money per race for the learning data of the foals born in 2012 was the performance data of the foals born from 2002 to 2011, the prize money per race for the learning data of the foals born in 2013 was the performance data of the foals born from 2003 to 2012, the prize money per race for the learning data of the foals born in 2014 was the performance data of the foals born from 2004 to 2013, the prize money per race for the learning data of the foals born in 2015 was the performance data of the foals born from 2005 to 2014, and the prize money per race for the learning data of the foals born in 2016 was the performance data of the foals born from 2006 to 2015 were respectively used.

[0121] (1) Calculation of prize money per race FIG. 4 is a flowchart showing an example of the processing procedure and processing content of the prize money per race calculation used for creating the score calculation data of the foals born in 2017 by the control unit 20 shown in FIG. 2.

[0122] In step S101, under the control of the performance data acquisition unit 201, the control unit 20 acquires profile data and performance data of foals born from 2007 to 2016 from an input device, an external database, etc. via the input / output interface unit 10, and stores them in the performance data storage unit 301 as performance data D1(2007) to D1(2016).

[0123] Next, in step S102, the performance data acquisition unit 201 combines the performance data D1(2007) to D1(2016), creates performance data D1(2007 - 2016), and stores it in the performance data storage unit 301. FIG. 9 shows an example of the created performance data D1. The performance data D1 includes at least the name of the racehorse, the total prize money obtained during a certain period as information representing the performance of the racehorse, and the number of starts during the said period. The performance data D1 further includes factors representing the profile of the racehorse, such as date of birth, gender, father's name, mother's name, mother's age at birth, mother's father's name, producer, trainer, owner, transaction price, etc.

[0124] In step S103, under the control of the prize money per race calculation unit 202, the control unit 20 reads the performance data D1(2007 - 2016) from the performance data storage unit 301, refers to the columns of the factors of the performance data D1(2007 - 2016), and performs a process of extracting total prize money data and number of starts data for each group of factors of the performance data D1(2007 - 2016). For a plurality of factors, processing is performed for each group classified according to the characteristics of the profile.

[0125] Subsequently, in step S104, based on the data extracted for each group of factors, the prize money per race is calculated. The calculation of the prize money per race is performed by dividing the total prize money of the group by the total number of starts of the group.

[0126] In step S105, the per-race prize calculation unit 202 combines the calculated per-race prize for each group of factors and sets this as per-race prize data d2 (2007 - 2016). A plurality of per-race prize data d2 are created for each factor. FIG. 10 shows an example of per-race prize data d2 for some groups with respect to the factors. The per-race prize data includes notations indicating the classified groups and the per-race prize for the respective groups.

[0127] In step S106, the per-race prize calculation unit 202 consolidates the plurality of per-race prize data d2 (2007 - 2016) created for each factor to create per-race prize data set D2 (2007 - 2016). Specifically, the per-race prize data d2 (2007 - 2016), which consists of separate files for each factor such as birth month and day, the father of the racehorse, the mother's age at birth, the mother's father of the racehorse, the producer, the trainer, the horse owner, and the transaction price, is compiled into one file. For example, when using EXCEL (registered trademark) to create data set D2, it is preferable for data set D2 to take the form of one workbook having a plurality of sheets that hold data d2 for each factor.

[0128] In step S107, the created per-race prize data set D2 (2007 - 2016) is stored in the per-race prize data storage unit 302.

[0129] The prize money per race data set D2 (2002 - 2011) used for creating learning data of foals born in 2012, the prize money per race data set D2 (2003 - 2012) used for creating learning data of foals born in 2013, the prize money per race data set D2 (2004 - 2013) used for creating learning data of foals born in 2014, the prize money per race data set D2 (2005 - 2014) used for creating learning data of foals born in 2015, the prize money per race data set D2 (2006 - 2015) used for creating learning data of foals born in 2016, and the prize money per race data set D2 (2007 - 2016) used for creating score calculation data of foals born in 2017 are all created using the same processing procedure and stored in the prize money per race data storage unit 302.

[0130] (2) Generation of Prediction Model (2 - 1) Creation of Learning Data Figure 5 is a flowchart showing an example of the processing procedure and processing content of the learning data creation process for foals born in 2012, which is part of the learning data for calculating the weight coefficients used in the data structure of foals born in 2017 by the control unit 20 shown in Figure 2.

[0131] In step S201, the control unit 20 reads out the performance data D1(2012) stored in the performance data storage unit 301 under the control of the learning data creation unit 203.

[0132] In step S202, the learning data creation unit 203 reads out the prize money per race data set D2 (2002 - 2011) stored in the prize money per race data storage unit 302. Step S202 may be executed after step S201, executed in parallel with step S201, or executed before step S201.

[0133] In step S203, the learning data creation unit 203 refers to each factor of the performance data D1 (2012), extracts the prize money data per race corresponding to those conditions from the prize money data set D2 (2002 - 2011) per race, and combines them to create the learning data D3 (2012). Specifically, the learning data creation unit 203 refers to the birth date, gender, father's name, mother's name, mother's age at birth, mother's father's name, producer, trainer, owner, transaction price, etc. from the performance data D1 (2012), extracts the prize money data per race corresponding to those conditions from the prize money data set D2 (2002 - 2011) per race, combines it with the performance data D1 (2012), and creates the learning data D3 (2012).

[0134] Figure 11 shows an example of the learning data D3. The learning data D3 includes, for example, the horse name, main prize money, birth date, gender, father's name, mother's name, mother's age at birth, mother's father's name, producer, trainer, owner, transaction price, etc. extracted from the performance data D1, and the prize money per race corresponding to each factor extracted from the prize money data set D2 per race.

[0135] In step S204, the control unit 20 stores the created learning data D3 (2012) in the learning data storage unit 303.

[0136] The learning data D3 (2013), D3 (2014), D3 (2015), and D3 (2016) for the foals born in 2013 - 2016 are also created by the same processing procedure as the learning data D3 (2012) for the foals born in 2012. However, for the creation of the learning data D3 (2013), the performance data D1 (2013) and the prize money data set D2 (2003 - 2012) per race are used, for the creation of the learning data D3 (2014), the performance data D1 (2014) and the prize money data set D2 (2004 - 2013) per race are used, for the creation of the learning data D3 (2015), the performance data D1 (2015) and the prize money data set D2 (2005 - 2014) per race are used, and for the creation of the learning data D3 (2016), the performance data D1 (2016) and the prize money data set D2 (2006 - 2015) per race are used.

[0137] (2-2) Optimization of Weight Coefficient FIG. 6 is a flowchart showing an example of a learning data acquisition process and a process procedure and process content for calculating a weight coefficient used for the data structure of the 2017-born foal by the control unit 20 shown in FIG. 2.

[0138] In step S301a, the control unit 20 reads the learning data D3(2012) from the learning data storage unit 303 under the control of the learning data acquisition unit 204.

[0139] Similarly, in steps S301b to S301e, the learning data acquisition unit 204 reads the learning data D3(2013) to learning data D3(2016) stored in the learning data storage unit 303. Steps S301a to S301e may be executed sequentially or simultaneously in parallel.

[0140] In step S302, the learning data acquisition unit 204 combines the learning data D3(2012) to learning data D3(2016) to obtain learning data D3(2012 - 2016).

[0141] In step S303, the learning data acquisition unit 204 incorporates a score value calculation formula composed of an arbitrarily determined calculation formula or a combination of functions and a temporary weight coefficient W into the learning data D3(2012 - 2016). At this point, the learning data D3(2012 - 2016) includes profile information, prize money data per race as numerical data, a weight coefficient, and a score value calculation formula, and satisfies the requirements of the data structure. The learning data D3(2012 - 2016) may include a temporary score value calculated based on the score value calculation formula based on the temporary weight coefficient.

[0142] In step S304, the learning unit 205 obtains the learning data D3 (2012 - 2016) from the learning data acquisition unit 204 and generates a prediction model by executing a process of performing statistical analysis. In this embodiment, the learning unit 205 uses the prize money in the learning data D3 (2012 - 2016) as the target variable, and conducts learning with the prize money per race for each factor as the explanatory variable, and optimizes the weight coefficient W for calculating a score value representing the likelihood of the occurrence of top - performing racehorses.

[0143] In step S305, the learning unit 205 stores the calculated final weight coefficient W in the learned data storage unit 304. Also, the learning unit 205 stores the same score value calculation formula as the prediction model used for calculating the weight coefficient W in the learned data storage unit 304.

[0144] (3) Creation of data structure (3 - 1) Acquisition of prediction data FIG. 7 is a flowchart showing an example of the processing procedure and processing content of the prediction data creation process for the 2017 foals by the control unit 20 shown in FIG. 2.

[0145] In step S401, under the control of the prediction data acquisition unit 206, the control unit 20 obtains the profile data of the 2017 foals from an input device or an external database etc. via the input / output interface unit 10, and creates the prediction data D4 (2017). FIG. 12 shows an example of the prediction data D4. For example, the prediction data D4 at least includes information such as the name of the racehorse, or in the case where the name is undetermined, the combination of the dam name representing the identity of the racehorse and the birth year. The prediction data D4 further includes factors representing the profile of the racehorse such as the birth date, gender, sire name, dam name, dam's age at birth, dam's sire name, breeder, trainer, owner, transaction price, etc.

[0146] In step S402, the prediction data acquisition unit 206 reads out the per-race prize money data set D2 (2007 - 2016) stored in the per-race prize money data storage unit 302. Step S402 may be executed after step S401, may be executed in parallel with step S401, or may be executed before step S401.

[0147] (3-2) Creation of data for score calculation In step S403 of FIG. 7, under the control of the control unit 20, the control unit 20 refers to each factor of the prediction data D4 (2017) generated by the prediction data acquisition unit 206, extracts the per-race prize money data corresponding to those conditions from the per-race prize money data set D2 (2007 - 2016), combines them, and creates the score calculation data D5 (2017). Specifically, the score calculation data creation unit 207 refers to the prediction data D4 (2017) to obtain the birth date, gender, father's name, mother's name, mother's age at birth, mother's father's name, producer, trainer, owner, transaction price, etc., extracts the per-race prize money data corresponding to those conditions from the per-race prize money data set D2 (2007 - 2016), combines it with the prediction data D4 (2017), and creates the score calculation data D5 (2017).

[0148] FIG. 13 shows an example of the score calculation data D5. The score calculation data D5 includes, for example, information such as the horse name or the combination of the mother's name and the birth year extracted from the prediction data D4, the birth date, gender, father's name, mother's name, mother's age at birth, mother's father's name, producer, trainer, owner, transaction price, etc., and the per-race prize money corresponding to each factor extracted from the per-race prize money data set D2.

[0149] In step S404, the control unit 20 stores the created score calculation data D5 (2017) in the score calculation data storage unit 305.

[0150] (3-3) Incorporation of weight coefficients and score calculation formula FIG. 8 is a flowchart showing an example of a creation procedure and processing content of a data structure which is a performance prediction model for foals born in 2017 by the control unit 20 shown in FIG. 2.

[0151] In step S501, the control unit 20 reads out the score calculation data D5(2017) from the score calculation data storage unit 305 under the control of the data structure creation unit 208.

[0152] Next, in step S502, the data structure creation unit 208 acquires the weight coefficient W stored in the learned data storage unit 304 and the score value calculation formula.

[0153] In step S503, the data structure creation unit 208 generates a data structure (2017) which is a performance prediction model for foals born in 2017 by incorporating the weight coefficient W and the score value calculation formula into the score calculation data D5.

[0154] Next, in step S504, the score calculation unit 209 may calculate a score value by performing an operation based on the score calculation formula with the prize money per race of each factor of the data structure (2017) as a variable.

[0155] Finally, in step S505, the control unit 20 stores the created data structure (2017) in the data structure storage unit 306.

[0156] A data structure is created by the above processing procedure. However, a series of processing such as reading data via the Internet, extracting information, processing, creating a prediction model, constructing a data structure, and incorporating a combination of formulas or functions into the data structure may be automatically performed in whole or in part by an application program.

[0157] Alternatively, all or part of the above series of processing may be performed by an operation by the user of the application program.

[0158] Next, using the data structure (2017) created by the above series of processes, a processing procedure for predicting the performance of foals born in 2017 will be described as an example.

[0159] (1) Supplement of profile data FIG. 14 is a flowchart showing an example of the processing procedure and processing content of supplementing profile data of the data structure of foals born in 2017 by the control unit 20 shown in FIG. 1.

[0160] In step S1001, the control unit 20 acquires supplementary data D4+(2017) of foals born in 2017 from an external database or the like via the input / output interface unit 10 under the control of the supplementary data acquisition unit 21. Alternatively, the supplementary data D4+(2017) may be acquired from a magnetic disk, an optical disk, or the like via the storage unit 30. Similar to the prediction data D4, D4+ includes information combining the name or dam name of the racehorse and the birth year, and also includes the latest information at the time of acquisition such as the birth date, gender, sire name, dam name, dam's age at birth, dam's sire name, producer, trainer, owner, and transaction price as factors representing the profile of the racehorse.

[0161] In step S1002, the control unit 20 reads out the data structure (2017) from the data structure storage unit 31 under the control of the data structure supplementing unit 22. Step S1002 may be executed after step S1001, may be executed in parallel with step S1001, or may be executed before step S1001.

[0162] Next, in step S1003, the data structure supplementing unit 22 refers to each factor of the data structure (2017) and the supplementary data D4+(2017) acquired by the supplementary data acquisition unit 21, and if there is a missing profile in the data structure, it supplements by combining with the supplementary data. Specifically, information such as the name of the horse, trainer, owner, and transaction price, which was undetermined at the time of creating the data structure, is supplemented. Also, information on foals that were not registered at the time of creating the data structure, such as overseas-produced horses imported after the creation of the data structure, is supplemented.

[0163] In step S1004, the data structure supplementing unit 22 reads out the prize money data set D2 (2007 - 2016) per race stored in the factor data storage unit 32.

[0164] In step S1005, the data structure supplementing unit 22 extracts the prize money data per race of the factor corresponding to the profile supplemented above from the prize money data set D2 (2007 - 2016) per race and incorporates it into the data structure (2017). (2) Calculation of score value

[0165] In step S1006, the score calculation unit 23 uses the prize money per race of each factor in the data structure (2017) as a variable and performs an operation based on the score calculation formula to calculate the score value. The score value obtained as a result of the operation is incorporated into the data structure by the score calculation unit 23.

[0166] In step S1007, the output control unit 24 outputs the score value via the input / output interface unit 10. It is preferable that the score value is output near the beginning of the row representing the profile of each racehorse as shown in FIG. 3, which makes it easier for comparison. Also, the output control unit 24 may perform sorting of the data based on the score value. Depending on the content of an arbitrarily defined formula or function, for example, if the higher the score value, the higher the predicted ranking order, it is efficient and preferable to sort in descending order of the score value for the selection of candidate horses.

[0167] The performance prediction of the foals born in 2017 is performed using the data structure (2017) according to the above processing procedure. However, a series of processes such as reading data via the Internet, supplementing data, extracting information, and performing calculations may be automatically performed in whole or in part by an application program.

[0168] In step S1003, instead of supplementing the data structure by referring to the supplementary data D4+(2017), the data structure supplementary unit 22 may correct or supplement the profile information using the data input by manual operation of the input / output interface unit 10 such as a keyboard or a mouse by the user.

[0169] The weighting factor of each data used for score calculation in the data structure may be changed by manual operation of the input / output interface unit 10 at an arbitrary timing by the user.

[0170] When the correction or supplementation of the profile data and the change of the weighting factor are performed starting from the manual operation of the user as described above, the score calculation unit 23 performs the calculation and calculates the score value. Next, the updated score value is output by the output control unit 24. It is preferable that the operation of the calculation and the update of the score value are automatically performed by the application program each time each data and the weighting factor are updated.

[0171] (Verification) In order to evaluate the usefulness of the score value calculated according to the embodiment, verification was performed using the performance data D1(2017) created by acquiring the performance data of the foals born in 2017 from June 2019 to May 2020. Specifically, the prize money obtained by each competing horse during the period was extracted from the performance data D1(2017), combined with the data structure (2017), and the score value was evaluated.

[0172] FIG. 15 is a histogram showing the distribution of the score values of racehorses in the data structure (2017) and a graph showing the occurrence rate within the top 100 and the occurrence rate within the top 200 of the prize money in each section of the score value.

[0173] In the histogram, the interval between intervals was set to every 20 in terms of score value. Racehorses with a score value greater than the minimum boundary value of the interval and less than or equal to the maximum boundary value were counted, and the number of horses in each interval was determined. For example, in the data structure (2017), there were 74 racehorses corresponding to the 160 - 180 interval where the score value was greater than 160 and less than or equal to 180.

[0174] The winning rate within the top 100 in prize money is the value obtained by dividing the number of horses in each interval of racehorses that have won a prize money of 24.9 million yen or more in 2017 foals within the top 100 arranged in the above order of prize money by the total number of horses in each interval counted above. For example, in the data structure (2017), since there were 13 racehorses within the top 100 in prize money in the 160 - 180 interval, the winning rate within the top 100 in prize money was calculated to be 17.6%.

[0175] The winning rate within the top 200 in prize money is the value obtained by dividing the number of horses in each interval of racehorses that have won a prize money of 17.1 million yen or more in 2017 foals within the top 200 arranged in the order of prize money by the total number of horses in each interval counted above. For example, in the data structure (2017), since there were 17 racehorses within the top 200 in prize money in the 160 - 180 interval, the winning rate within the top 200 in prize money was calculated to be 23.0%.

[0176] For example, in the data structure (2017), out of 97 racehorses with a score value greater than 160, 17 were within the top 100 in prize money and 23 were within the top 200 in prize money. The winning rate within the top 100 in prize money was 17.5% and the winning rate within the top 200 in prize money was 23.7% respectively. By considering candidate horses from racehorses with a high score value before debut, a result was shown that it was expected to efficiently select racehorses that win a lot of prize money in POG.

Explanation of symbols

[0177] 1 ··· Computer A 2 ··· Computer B 10 ··· Input / output interface unit 20 ··· Control unit 21... Supplementary data acquisition unit 22... Data structure supplementary unit 23... Score calculation unit 24... Output control unit 30... Memory unit 31... Data structure storage unit 32... Factor data storage unit 201... Performance data acquisition unit 202... Prize money calculation unit per race 203... Learning data creation unit 204... Learning data acquisition unit 205... Learning unit 206... Prediction data acquisition unit 207... Data creation unit for score calculation 208... Data structure creation unit 209... Score calculation unit 301... Performance data storage unit 302... Prize money data storage unit per race 303... Learning data storage unit 304... Learned data storage unit 305... Data storage unit for score calculation 306... Data structure storage unit birth... Column indicating the prize money per race with the birth date as a factor sire... Column indicating the prize money per race with the father of the racehorse as a factor broodmare... Column indicating the prize money per race with the mother's age at birth as a factor BMS... Column indicating the prize money per race with the mother's father of the racehorse as a factor farm... Column indicating the prize money per race with the producer as a factor trainer... Column indicating the prize money per race with the trainer as a factor owner... Column indicating the prize money per race with the horse owner as a factor value... Column indicating the prize money per race with the transaction price as a factor

Claims

**Claim 1**: A program for predicting the performance of racehorses, which causes a computer to perform the following steps for a data structure including profile information containing at least three or more factors selected from the father, father's father, gender, trainer, owner, breeder, transaction price, mother's racing performance, date of birth, and mother's age at birth of racehorses in the target generation for which performance prediction is to be made, numerical data calculated for each factor based on the prize money of racehorses in a generation or more past than the target generation for which performance prediction is to be made and having the same corresponding factors, a weight coefficient used for calculating the score value, and a score value calculation formula including the following formula (a) or formula (b) in the formula: calculating a score value that can relatively compare each racehorse in the target generation for which performance prediction is to be made by causing the computer to perform the calculation of the score value calculation formula; creating output data based on the calculated score value; and outputting the data to a display device or an external terminal. Formula (a) ··· w1x1 + w2x2 + ··· + wfxf Formula (b) ··· w1 / x1 + w2 / x2 + ··· + wf / xf However, x1, x2, ···, xf represent the numerical data of f factors, and w1, w2, ···, wf represent the respective weight coefficients. **Claim 2** The program according to claim 1, wherein the weight coefficient is obtained by causing a computer to learn using the data structure of racehorses in a generation or more past than the target generation for which performance prediction is to be made as learning data, using the numerical data as explanatory variables, and using the information representing the performance of racehorses in the generation or more past than the target generation as objective variables. **Claim 3** In the program according to claim 2, as a method for determining the weight coefficient, the cumulative value from i = 1 to n of the difference between the total prize money from the 1st to the i-th horse (where i = 1, 2, 3, ···, n (n ≤ m)) arranged in descending order of prize money of m racehorses in a generation or more past than the target generation for which performance prediction is to be made and the total prize money from the 1st to the i-th horse arranged in descending order of the score value is minimized, and the weight coefficient is determined.

Citation Information

Patent Citations

  • Prediction performance curve estimation program, prediction performance curve estimation device and prediction performance curve estimation method

    JP2017049674A

  • Information processing system, information processing method, information processing program, and information processing device

    JP2020087334A

  • Program, method, and system for predicting potential ability of race horse, seed horse candidate presentation program, method for presenting seed horse candidate, seed horse candidate presentation system, and prediction career earnings database preparation program used for the same

    JP2020149583A