Medal prediction optimization method

By constructing a multi-level and multi-dimensional medal prediction model, using the Tobit model, random effect correction and RCA index, the problem of ignoring key factors and dynamic information in the existing technology is solved, and more accurate and reliable medal prediction is achieved.

CN120218353APending Publication Date: 2025-06-27SHAANXI UNIV OF SCI & TECH

Patent Information

Application Number
CN202510372047.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing medal prediction methods ignore the influence of key factors such as competition venues, athletes, and events, and it is difficult to fully integrate dynamic information, resulting in inaccurate predictions.

Method used

The Tobit model is used to combine random effect correction and RCA index to construct a multi-level and multi-dimensional medal prediction model. This model extracts key indicators from the perspectives of athletes, competition events, sponsor countries, etc., processes restricted dependent variables, enhances the explanatory power and robustness of the model, and dynamically captures the competitive advantages of each country in different projects.

Benefits of technology

It significantly improves the accuracy and explanatory power of medal predictions. It not only predicts the number of awards won by various countries, but also captures the advantages of countries in this event and provides scientific and reliable international event medal prediction solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218353A_ABST
    Figure CN120218353A_ABST
Patent Text Reader

Abstract

The invention relates to a medal prediction optimization method, which comprises the following steps of: 1, acquiring basic information, competition categories and items of competitors in previous competitions, and prize winning condition data of the dominant countries of the previous competitions and each country; 2, data preprocessing and key index extraction are carried out, the data obtained in the step 1 are cleaned, and then key indexes are extracted from three core perspectives of athletes, competition items and a host country; 3, constructing a Tobit model, introducing a virtual variable to capture heterogeneity in data, and meanwhile, correcting the influence of a time invariant feature on the model by introducing a random effect obeying normal distribution with the mean value being 0; and step 4, integrating the display comparative advantage index into the Tobit model established in the step 3, optimizing the prediction precision of the model, predicting the number of meals obtained by each country, obtaining a match medal list, and capturing advantage items of each country in the current match. Existing data resources can be utilized to the maximum extent, and resource allocation is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medal prediction, and particularly relates to an optimization method for medal prediction. Background Art

[0002] Existing methods for predicting medals mainly rely on the historical medal-winning situations of various countries, especially the numbers of gold, silver, and bronze medals. However, these predictions often neglect the impacts of key factors such as competition venues, participating athletes, and participating events, and also overlook the problem that the number of medals won is a non-negative dependent variable subject to constraints. In addition, the specific lists of participating athletes in many international events are usually announced only within one month before the start of the competition, and it is difficult for existing technologies to comprehensively integrate such dynamic information. This uncertainty makes it difficult for national Olympic committees to accurately predict the competitive performances of their own countries in sports events with a relatively long preparation period.

[0003] For example, the patent with the publication number: CN111695680A discloses a performance prediction method, a performance prediction model training method, a device, and an electronic device. The specific implementation solution is as follows: determining the team information and past performance information of the team to be predicted; inputting the team information and past performance information of the team to be predicted into a pre-trained performance prediction model; obtaining the predicted performance of the team to be predicted output by the performance prediction model. Through the above solution, on the one hand, combining the team information and past performance information of the team as a reference for performance prediction can improve the accuracy of prediction. On the other hand, since the performance prediction model is an overall model and is trained as an overall model during training, the finally trained output performance prediction result is closer to the actual situation. Its existing defect is that it only considers the historical medal-winning situations of the team and the capabilities of the athletes, so it can only predict the performance of a single event. In addition, there is a lack of comprehensive consideration of competition venues and competition events, and it is impossible to predict the number of medals won by each country.

[0004] Therefore, in view of the above problems, an optimization method for medal prediction is needed. Summary of the Invention

[0005] In order to overcome the above problems existing in the prior art, the purpose of the present invention is to provide an optimization method for medal prediction. The established prediction model can not only help coaches of various countries scientifically formulate training plans during the intense preparation period of the competition, assist athletes in improving their competitive levels, but also maximize the utilization of existing data resources, optimize resource allocation, and provide strong support for each country to strive for more medals in the competition.

[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0007] An optimization method for medal prediction, comprising the following steps:

[0008] Step 1: Collect data, obtaining the basic information of athletes participating in previous events, the categories and items of the competitions, as well as data on the host countries of previous events and the medal-winning situations of each country;

[0009] Step 2: Conduct data preprocessing and extract key indicators. Clean the data obtained in Step 1 to ensure the integrity and accuracy of the data. Subsequently, extract key indicators from three core perspectives: athletes, competition events, and host countries, providing high-quality data support for the training and prediction of subsequent models;

[0010] Step 3: Establish a Tobit model and perform random effect correction

[0011] The Tobit model is particularly suitable for dealing with limited dependent variables (such as the number of medals being non-negative integers) and can provide more robust prediction results in the case of unbalanced data distribution.

[0012] Based on the three sets of key indicators extracted in Step 2, construct a Tobit model. To further enhance the interpretability of the model, introduce dummy variables to capture the heterogeneity in the data. At the same time, by introducing a random effect that follows a normal distribution with a mean of 0, correct the impact of time-invariant characteristics on the model, thereby significantly improving the robustness and prediction accuracy of the model.

[0013] Step 4: Update the model of the RAC index

[0014] Integrate the Revealed Comparative Advantage (RCA) index into the Tobit model established in Step 3. The RCA index can further optimize the prediction accuracy of the model by measuring the intensity of production factors of each country in specific items. Furthermore, the updated model can dynamically capture the competitive advantages of each country in different items;

[0015] Finally, not only can the number of medals obtained by each country be predicted to obtain the medal standings of the event, but also the advantageous events of each country in this event can be captured;

[0016] In the above Step 1, the basic information of athletes participating in previous events of the competition includes gender, nationality, participating year and event, and whether they won awards;

[0017] The competition categories include major events such as swimming, track and field, equestrian, gymnastics, badminton, etc. These competition categories are further divided into different events, such as: synchronized swimming, diving, relay swimming, etc.;

[0018] For the medal-winning situations of each country in previous events, it is necessary to obtain the number of gold, silver, and bronze medals obtained by each participating country in each event.

[0019] In the above Step 2, cleaning the data collected in Step 1 includes checking for missing values and unreasonable values;

[0020] The extracted metrics include three perspectives: athletes, competition events, and host countries, which are as follows:

[0021] (1) Perspective of athletes

[0022] Average number of athletes (ANA ijk ): The average number of athletes participating in the two previous events before the i-th event in the j-th country, specifically:

[0023]

[0024] Gender of athletes (SEX ijk ):

[0025]

[0026] Whether retired (RET ijk ): Based on the average participation times of athletes in each competition event, it is calculated whether the athlete will participate in the upcoming event;

[0027]

[0028] (2) Perspective of competition events:

[0029] Proportion of gold medals (PG ijk ):

[0030]

[0031] Among them, represents the total number of gold medals in the previous two years of events, is the total number of medals in the previous two years of events;

[0032] Event category (TYPE ijk ): One-hot encoding is used for distinction;

[0033] Number of event participations (NPP ij ): The total number of times the j-th country participates in the k-th competition event before the i-th event;

[0034] Number of event awards (NPW ij ): Refers to the total number of awards the j-th country receives for participating in the k-th competition event before the i-th event;

[0035] Event participation rate (NPT ijk ): The ratio of the number of awards for the j-th country's participation in the k-th competition event to the total number of participations before the i-th event. Specifically:

[0036]

[0037] (3) Perspective of the host country

[0038] Is it the host country (H ij ): A 0-1 variable. If the j-th country is the host country of the i-th event, then H ij = 1; otherwise, H ij = 0;

[0039] Number of new events (NE i ): The number of competition events added by the host country;

[0040] Reduction in the number of medals (RN i ): The number of competition events cancelled by the host country.

[0041] In Step 3, the core idea of the Tobit model is to divide the dependent variable Y into two parts, the latent variable y* and the observed value y. This mechanism enables the Tobit model to effectively handle the situation where the dependent variable is restricted, such as the scenario where the number of medals is a non-negative integer. Specifically:

[0042] y* = Xβ + ε, ε ∈ -N(0, σ 2 )

[0043]

[0044] The construction process of the Tobit model is as follows:

[0045] Step1) Using the three core indicator groups extracted in Step 2 as the basis for the medal production function of each country, the athlete perspective group (A ijk ), the competition event perspective group (P ijk ), and the host country perspective group (HC ij );

[0046] A ijk = ANA ijk + SEX ijk + RET ijk

[0047] P ijk = PG ijk + TYPE ijk + NPP ijk + NPW ijk + NPT ijk

[0048] HC ijk = H ij + NE i + RN i

[0049] Step2) Based on the above variables, establish a double-sided truncated regression model:

[0050] Y *ijk = β1lnA ijk + β2lnP ijk + β3lnHC ijk + ε + μ

[0051] ε ~ N(0, σ 2 )

[0052] Where Y ijk represents the number of medals obtained by the j-th country in the k-th event of the i-th competition. Y ijk * is the latent dependent variable, A ijk , P ijk , HC ij are the explanatory variables, and β is the regression coefficient. When the latent variable Y ijk * is greater than 0, the dependent variable Y ijk is equal to Y ijk * itself, and the result of Y ijk is rounded; otherwise, the dependent variable Y ijk is equal to 0. Assume that the disturbance term u ijk obeys a normal distribution with a mean of 0 and a variance of σ 2 .

[0053] Step3) When fitting the equation to the data of different competitions

[0054] , consider the influence of dummy variables on the change in the number of medals and further introduce random effects to correct the time-invariant characteristics. The dummy variables include the type of event (TYPE k ), the gender of the athletes required for the event (SEX k ), whether the country is the host country (H ij ), and whether there are retired athletes (RET ijk );

[0055] The random effects model treats the effects of some variables as randomly drawn from the population rather than fixed values, so as to better capture the heterogeneity in the data. The specific form is:

[0056]

[0057] Where v ij is the random effect, and it is assumed to obey a normal distribution with a mean of 0 and a variance of σ v 2 ;

[0058] In step four, introduce the Revealed Comparative Advantage (RCA) index. The calculation formula of the RCA index is:

[0059]

[0060] Among them, M (i-l)jk represents the number of medals obtained by country i in the k-th event, represents the total number of medals obtained by country i in all events, represents the number of medals allocated by the host to the k-th event, represents the total number of medals in this event;

[0061] Subsequently, the RCA model is integrated into the Tobit model obtained in Step 3, and the updated model is:

[0062] Y * ijk =β1lnA ijk +β2lnP ijk +β3lnHC ijk +β4RCA ijk +ε.

[0063] Advantages of the present invention:

[0064] 1. Through the Tobit model, random effect correction, and RCA index integration, the present invention constructs a multi-level and multi-dimensional medal prediction model. This model not only solves the problem of limited dependent variables where the number of medals is a non-negative integer but also effectively captures the data heterogeneity between different countries and different sessions through random effect correction, significantly improving the interpretability and prediction accuracy of the model. In addition, the RCA index can accurately reflect the competitive advantages of countries in different events by dynamically measuring the intensity of production factors in specific events for each country. This enables the model to not only handle static historical data but also dynamically capture the competitive situation of countries in different events, providing a more scientific and reliable solution for international event medal prediction.

[0065] 2. Considering the randomness of international competitive sports, the present invention distinguishes the characteristics of different events by introducing dummy variables (such as whether it is the host country, event category, etc.), significantly improving the interpretability and robustness of the model, and providing strong support for countries to optimize resource allocation and enhance competitive strength during the preparation period. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a schematic flow chart of a medal prediction optimization method provided by an embodiment of the present invention.

[0067] Figure 2 It is a schematic diagram of three core index groups extracted in Step 2 of an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0068] The present invention will be further described in detail below with reference to the accompanying drawings.

[0069] AsFigure 1 As shown in the figure, an optimized method for medal prediction includes the following steps:

[0070] Step 1: Data collection

[0071] Obtain data such as the basic information of athletes participating in previous events, the competition categories and events, as well as the host countries of previous events and the medal-winning situations of each country; the basic information of athletes participating in previous events includes gender, nationality, participation year and event, and whether they won medals; for the medal-winning situations of each country in previous events, it is necessary to obtain the number of gold, silver, and bronze medals obtained by each participating country in each event.

[0072] Step 2: Data preprocessing and key index extraction

[0073] As Figure 2 shown, clean the data obtained in Step 1 to ensure the integrity and accuracy of the data. Subsequently, extract key indicators from three core perspectives: athletes, competition events, and host countries, providing high-quality data support for the subsequent training and prediction of the model. These indicators include:

[0074] (1) From the perspective of athletes

[0075] Average number of athletes (ANA ijk ): The average number of athletes participating in the two previous events of the jth country in the ith event, specifically:

[0076]

[0077] Athlete gender (SEX ijk ):

[0078]

[0079] Whether retired (RET ijk ): Calculate whether an athlete will participate in the upcoming event based on the average number of participations of athletes in each competition event;

[0080]

[0081] (2) From the perspective of competition events:

[0082] Proportion of gold medals (PG ijk ):

[0083]

[0084] Among them, represents the total number of gold medals in the previous two events, is the total number of medals in the previous two events;

[0085] Event category (TYPE ijk): Use one-hot encoding for distinction;

[0086] Number of times a country participated in events (NPP ij ): The total number of times the j-th country participated in the k-th event before the i-th competition;

[0087] Number of awards for an event (NPW ij ): Refers to the total number of awards the j-th country won in the k-th event before the i-th competition;

[0088] Participation rate of an event (NPT ijk ): The ratio of the number of awards won by the j-th country in the k-th event to the total number of participations before the i-th competition. Specifically:

[0089]

[0090] (3) Perspective of the host country

[0091] Whether it is the host country (H ij ): A 0-1 variable. If the j-th country is the host country of the i-th competition, then H ij = 1; otherwise, H ij = 0;

[0092] Number of new events (NE i ): The number of new events added by the host country;

[0093] Number of medals reduced (RN i ): The number of events cancelled by the host country;

[0094] Step 3: Establishment of the Tobit model and random effect correction

[0095] The Tobit model is a regression model specifically designed to handle truncated or censored dependent variables, suitable for analyzing problems where some observed values are truncated, especially here where the number of medals won is greater than or equal to 0. The Tobit model divides the dependent variable Y into two parts, the latent variable y* and the observed value y. This mechanism enables the Tobit model to effectively handle the situation where the dependent variable is restricted, such as the scenario where the number of medals is a non-negative integer. Specifically:

[0096] y* = Xβ + ε, ε ∈ -N(0,σ 2 )

[0097]

[0098] In model construction, the three core indicator groups extracted in Step 2 are used as the basis for the medal production function of each country, the athlete perspective group (A ijk ), the event perspective group (P ijk ), and the host country perspective group (HCij );

[0099] A ijk = ANA ijk + SEX ijk + RET ijk

[0100] P ijk = PG ijk + TYPE ijk + NPP ijk + NPW ijk + NPT ijk

[0101] HC ijk = H ij + NE i + RN i

[0102] Based on the above variables, a two-sided truncated regression model is established:

[0103] Y * ijk = β1lnA ijk + β2lnP ijk + β3lnHC ijk + ε + μ

[0104] ε ~ N(0, σ 2 )

[0105] where Y ijk represents the number of medals obtained by the j-th country in the k-th event of the i-th competition. Y ijk * is the latent dependent variable, and A ijk , P ijk , HC ij are the explanatory variables, and β is the regression coefficient. When the latent variable Y ijk * is greater than 0, the dependent variable Y ijk is equal to Y ijk * itself, and the result of Y ijk is rounded; otherwise, the dependent variable Y ijk is equal to 0. Assume that the disturbance term u ijk follows a normal distribution with a mean of 0 and a variance of σ 2 .

[0106] When fitting the equation to the data of different competitions, to consider the impact of dummy variables on the change in the number of medals, a random effect is further introduced to correct the time-invariant characteristics. The dummy variables include the type of event (TYPE k ), the gender of the athletes required for the event (SEX k), whether the country is the host country (H ij ), and whether there are retired athletes (RET ijk );

[0107] The random effects model can treat the effects of certain variables as randomly drawn from the population rather than fixed values, thus better capturing the heterogeneity in the data. The specific form is:

[0108]

[0109] where v ij is the random effect, assumed to follow a normal distribution with a mean of 0 and a variance of σ v 2 ;

[0110] Step 4: Update the RAC index integration model

[0111] To more comprehensively reflect the competitive advantages of countries in specific events, the Revealed Comparative Advantage (RCA) index is introduced and integrated into the Tobit model to further optimize the prediction accuracy;

[0112] The calculation formula of the RCA index is:

[0113]

[0114] where M (i-l)jk represents the number of medals obtained by country i in the kth event, represents the total number of medals obtained by country i in all events, represents the number of medals allocated by the host to the kth event, represents the total number of medals in this event;

[0115] Integrate the RCA model into the Tobit model obtained in Step 3 to update the model:

[0116] Y * ijk = β1lnA ijk + β2lnP ijk + β3lnHC ijk + β4RCA ijk + ε. According to the above steps, the comparison between the predicted results and the actual results of the 2024 Olympic Games is as follows in the table:

[0117]

[0118]

[0119] The accuracy rate of the predicted results in 2024 is as follows in the table:

[0120]

[0121] The mean squared error (MSE) measures the average of the squares of the differences between the predicted values and the true values in a prediction. The smaller the MSE value, the closer the predicted value is to the true value, and the higher the prediction accuracy. In the examples of the present invention, the MSE results of the gold medal prediction and the total medal prediction are both less than 100, indicating good prediction effects.

[0122] The correlation coefficient is an index that measures the degree of linear correlation between two variables, and its value range is between -1 and 1. When the correlation coefficient is closer to 1, it indicates a stronger positive correlation between the two variables; when the correlation coefficient is closer to -1, it indicates a stronger negative correlation; when the correlation coefficient is close to 0, it indicates that there is almost no linear correlation between the two variables. In the examples of the present invention, the correlation coefficients between the predicted values and the true values of the total medal numbers are all greater than 0.75, which indicates a strong positive correlation between the prediction results and the true values, further indicating that the error between the predicted values and the true values is small and the prediction effect is relatively accurate.

[0123] It can be found that the accuracy rates of the total medal numbers and the gold medal numbers predicted by the present invention are both at a relatively high level. Further, the prediction results for 2028 are as follows in the table:

[0124]

Claims

1. An optimization method for medal prediction, characterized in that: The following steps are involved: Step 1: Collect data to obtain basic information of athletes participating in previous competitions, competition categories and events, as well as data on the host countries and awards of various countries; Step 2: Perform data preprocessing and extract key indicators. Clean the data obtained in step 1 to ensure the integrity and accuracy of the data. Then, extract key indicators from three core perspectives: athletes, events, and host countries, to provide high-quality data support for subsequent model training and prediction. Step 3: Based on the three sets of key indicators extracted in step 2, a Tobit model is constructed, and dummy variables are introduced to capture the heterogeneity in the data. At the same time, random effects with a normal distribution with a mean of 0 are introduced to correct the impact of time-invariant characteristics on the model. Step 4: Integrate the display comparative advantage index into the Tobit model established in step 3 to further optimize the prediction accuracy of the model. The updated model can dynamically capture the competitive advantages of countries in different events. Finally, predict the number of medals won by each country, derive the medal table of the event, and capture the advantageous events of each country in this event.

2. The optimization method for medal prediction according to claim 1, characterized in that: In step 1, the basic information of athletes participating in the competition over the years includes gender, nationality, year of participation and event, and whether they won an award; The winning results of each country in previous competitions require the number of gold, silver and bronze medals won by each participating country in each competition.

3. The optimization method for medal prediction according to claim 1, characterized in that: In the step 2, the data collected in the step 1 is cleaned, including checking for missing values ​​and unreasonable values; The extracted indicators include three perspectives: athletes, events, and host countries, as follows: (1) Athlete perspective Athlete Average (ANA ijk ): The average number of athletes from the jth country in the two previous events before the i-th event, specifically: Athlete gender (SEX ijk ): Retired (RET ijk ): Calculate whether an athlete will participate in the upcoming event based on the average number of times the athlete participates in each event; (2) Competition perspective: Gold Medal Percentage (PG ijk ): in, represents the total number of gold medals in the previous two years of competition, is the total number of medals won in the previous two years of the competition; Project Category ijk ): Use one-hot encoding to distinguish; Number of project participations (NPP ij ): the total number of times the jth country participated in the kth event before the i-th event; Number of project awards (NPW ij ): refers to the total number of awards won by the jth country in the kth event before the ith event; Project participation rate (NPT ijk ): The ratio of the number of awards won by the jth country in the kth event to the total number of times it participated before the i-th event. Specifically: (3) Host country perspective Host country? ij ): 0-1 variable. If the jth country is the host country of the i-th event, then H ij =1, otherwise, H ij =0; Number of new projects (NE i ): the number of events added by the host country; Reduced Medal Count (RN i ): Number of events cancelled by the host country.

4. The optimization method for medal prediction according to claim 1, characterized in that: In the step three, specifically: y*=Xβ+ε,ε∈-N(0,σ 2 ) The Tobit model construction process is as follows: Step 1) The three core indicator groups extracted in step 2 are used as the basis of each country's medal production function, and the athlete perspective group (A ijk ), Competition Perspective Group (P ijk ), and the host country perspective group (HC ij ); A ijk =ANA ijk +SEX ijk +RET ijk P ijk =PG ijk +TYPE ijk +NPP ijk +NPW ijk +NPT ijk HC ijk =H ij +NE i +RN i Step 2) Based on the above variables, a bilateral truncated regression model is established: Y * ijk =β1lnA ijk +β2lnP ijk +β3lnHC ijk +e+m ε~N(0,σ 2 ) Among them, Y ijk Y represents the number of medals won by the jth country in the kth event in the i-th competition, ijk * is the potential dependent variable, A ijk , P ijk HC ij is the explanatory variable, and β is the regression coefficient. ijk * When it is greater than 0, the dependent variable Y ijk Equal to Y ijk * itself, and Y ijk The result of rounding; otherwise, the dependent variable Y ijk is equal to 0. Assume that the disturbance term u ijk The mean is 0 and the variance is σ 2 Normal distribution of Step 3) When fitting the equations for the data of different events, we further introduce random effects to correct the time-invariant characteristics by considering the impact of dummy variables on the change in the number of medals. The dummy variables include the event type (TYPE k ), the gender of the athlete required for the event (SEX k ), whether the country is the host country (H ij ), and whether there are retired athletes (RET ijk ); The random effects model treats the effects of certain variables as randomly drawn from the population rather than fixed values, and the specific form is: Among them, v ij is a random effect, which is assumed to have a mean of 0 and a variance of σ v 2 The normal distribution of .

5. The optimization method for medal prediction according to claim 1, characterized in that: The fourth step introduces the revealed comparative advantage index, and the calculation formula of the RCA index is: Among them, M (i-l)jk represents the number of medals won by country i in the kth event, represents the number of medals won by country i in all events, represents the number of medals allocated by the host country to the kth event, Indicates the total number of medals in the event; Then, the RCA model is integrated into the Tobit model obtained in step 3, and the updated model is: Y * ijk =β1 ln A ijk +β2ln P ijk +β3 ln HC ijk +β4RCA ijk +e.

Citation Information

Patent Citations

  • Score prediction method and device, score prediction model training method and device, and electronic equipment

    CN111695680A

Cited By

  • Sports meeting medal prediction and excellent coach effect analysis method and system

    CN121562876A