Competition popularity analysis method and system based on big data
By using big data analytics to structure event information and calculate the propagation factor of popularity bands, the problem of insufficient data structuring in traditional event popularity analysis is solved. This enables hierarchical attribution and path matching of event popularity, improving the accuracy of popularity analysis and decision support capabilities.
Patent Information
- Application Number
- CN202511612909.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-01-02
AI Technical Summary
Traditional methods for analyzing the popularity of sporting events rely on manual or semi-automatic data collection, resulting in insufficient data structuring. This makes it difficult to identify the continuity of popularity peaks and the gaps between peaks, accurately distinguish the focus of attention for different events, and lack a hierarchical assessment of the dissemination relationship. Consequently, it is difficult to focus resource allocation and monitoring efforts, especially in scenarios with multiple events running concurrently, where trend analysis becomes ambiguous.
By employing big data-based methods for analyzing the popularity of events, including structured processing of event information, identification of popularity timestamps, calculation of popularity band propagation factors, hierarchical analysis, and popularity path matching, the continuity, aggregation, and dissemination of event popularity are structured and enhanced, ensuring the accuracy and traceability of popularity data.
It achieves completeness and relevance in event popularity analysis, ensures the accuracy of comparisons between multiple types of events, avoids trend distortion, can identify key nodes and dissemination paths, and improves the accuracy and decision support capabilities of popularity analysis.
Smart Images

Figure CN121256199A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of big data analysis, in particular to a big data-based event heat analysis method and system. BACKGROUND
[0002] The technical field of big data analysis involves collecting, storing, managing and computing processing of data with extensive sources and large scale, extracting the correlation and characteristic information hidden in the data through the integration and analysis of structured, semi-structured and unstructured data, and providing decision support and trend analysis for specific application scenarios. This technical field covers core links such as data collection, cleaning, conversion, storage, query, statistical analysis and pattern recognition, and uses distributed computing architecture and multi-dimensional data modeling methods to process large-scale data sets.
[0003] Among them, the traditional event heat analysis method refers to the means of statistics and analysis of the attention degree of sports events, e-competitive games or competitive activities. The technical matter targeted by this method is to obtain event-related information and measure and compare the attention degree. The traditional event heat analysis collects data from news reports, forum posts, social media comments and forwarding sources through manual or semi-automatic means, and then summarizes and analyzes according to specific event counting rules, keyword frequency statistics or media report quantity statistics to reflect the attention trend of the event at different time nodes.
[0004] The traditional event heat analysis relies on manual or semi-automatic data collection, which is prone to insufficient structure due to scattered sources and different data formats. In heat trend statistics, only keyword frequency or report quantity is used for summary, lacking recognition of heat peak continuity and band gap, resulting in insufficient trend refinement. When the type of event is complex or simultaneous, the heat of different events is easily interfered with each other, making it difficult to distinguish the real focus of attention. The path analysis is mostly limited to the comparison of time node quantity, lacking hierarchical evaluation of transmission relationship, and unable to reveal the transmission rule of heat between time and event. In the selection of key events, there is a lack of transmission priority ranking, making it difficult to focus on resource allocation and monitoring, and the trend analysis is prone to be fuzzy and the decision is delayed in the scene of multiple events in parallel. SUMMARY
[0005] In order to solve the technical problems existing in the prior art, the embodiments of the present application provide a big data-based event heat analysis method, which comprises the following steps: In order to achieve the above-mentioned purpose, the present application adopts the following technical scheme: a big data-based event heat analysis method, comprising the following steps: S1: based on the event association data, extract the event name, participating team, competition time, identify and classify the event type, regional distribution, and organize the data in a structured manner to obtain a standardized event information dataset; S2: based on the standardized event information dataset, extract the event heat timestamp, arrange in chronological order, identify the adjacent heat peak time interval, judge whether the heat fluctuation is continuous, filter the associated heat into the same heat waveband, calculate the heat waveband propagation factor, divide the heat propagation path, and obtain the event heat time series; S3: call the event heat time series, count the frequency of event types in different heat wavebands, evaluate the heat attribution weight, judge whether the heat is merged or split, and obtain the hierarchical event heat analysis result; S4: based on the hierarchical event heat analysis result, extract the key heat node, analyze the time window association relationship, filter the correlation degree heat into the same heat path, and obtain the event heat path matching dataset.
[0006] As a further scheme of the present application, the standardized event information dataset includes event name, participating team, and subject-predicate-object structure analysis result, the event heat time series includes heat timestamp, heat fluctuation judgment result, and heat waveband propagation factor, the hierarchical event heat analysis result includes heat attribution weight, event co-occurrence probability, and heat level classification, and the event heat path matching dataset includes heat time offset, association relationship analysis result, and correlation degree heat classification.
[0007] As a further scheme of the present application, the step of the standardized event information dataset is specifically: S101: based on the event association data, including news reports, social media comments, and forum posts, extract the event name, participating team, and competition time, remove duplicate records, and obtain the event information matching data volume; S102: based on the event information matching data volume, analyze the semantic relationship of event name, participating team, and competition time, extract the subject-predicate-object structure, combine the event field mapping corresponding to the event type, merge the event names under the same event type, count the number of participating teams for multiple event types, and classify according to the event type quantity threshold to obtain the event type distribution interval dataset; S103: call the event type distribution interval dataset, classify and analyze the regional distribution according to the event name, group the region, industry category, and event level, cross-match the event type and summarize the key event type, and generate the standardized event information dataset.
[0008] As a further scheme of the present application, the step of the event heat time series is specifically: S201: arranging in time sequence based on the heat time stamp of the standardized event information dataset, identifying the time interval of adjacent heat peaks, screening the heat peaks not exceeding the threshold value, and generating an event heat time interval sequence; S202: based on the event heat time interval sequence, for the heat peaks exceeding the time threshold value, analyzing the event regional consistency, extracting the region, event level and event type, screening the heat peaks meeting the conditions and belonging to the same heat band, and counting the number of heat peaks, time span and number of regions in the band to obtain heat band statistical data; S203: calling the heat band statistical data, calculating the heat band propagation factor, dividing the heat propagation path according to the number of heat peaks, time span and number of regions, and generating an event heat time sequence.
[0009] As a further scheme of the application, the step of hierarchical event heat analysis result is specifically: S301: calling the event heat time sequence, counting the frequency of event type in the differential heat band, classifying according to the event name, event level and event type, and obtaining the event type frequency distribution data; S302: based on the event type frequency distribution data, identifying the event co-occurrence probability, extracting the event type combination in the differential heat band, counting the number of combinations, calculating the heat attribution weight according to the probability threshold value according to the co-occurrence number and heat total amount, and generating a heat attribution weight matrix; S303: calling the heat attribution weight matrix, dividing the heat according to the attribution weight, merging the heat higher than the threshold value, splitting the heat lower than the threshold value, and obtaining the hierarchical event heat analysis result.
[0010] As a further scheme of the application, the heat attribution weight adopts the formula: ; Among them, represents the heat attribution weight, represents the heat value of event i in the kth frequency, represents the heat value of event j in the kth frequency, represents the total number of frequencies.
[0011] As a further scheme of the application, the step of event heat path matching data set is specifically: S401: based on the hierarchical event heat analysis result, extracting key heat nodes, arranging the heat nodes according to time sequence, screening the nodes with time axis offset characteristics, calculating the time difference between adjacent nodes, and obtaining the key heat time offset amount; S402: Based on the key heat time offset, analyze the correlation in the time window, identify the temporal relationship between adjacent heat nodes, identify the correlation strength between nodes according to the time interval and the triggering order between nodes, filter the heat node combination with a correlation greater than a preset threshold, and obtain the time correlation dataset. S403: Based on the time-related relationship dataset, filter nodes with high relevance and popularity, cluster nodes belonging to the same event type, merge nodes of the same type and divide the belonging paths to generate an event popularity path matching dataset.
[0012] As a further aspect of the present invention, the method further includes step S5: S5: Call the event popularity path matching dataset, calculate the node propagation value, sort the popularity priority according to the propagation value, filter the propagation paths and classify them into key monitoring, and obtain the event popularity propagation collection information table; The event popularity dissemination information table includes popularity priority, node dissemination value, and key monitoring paths.
[0013] As a further aspect of the present invention, the steps of the event popularity dissemination information table are specifically as follows: S501: Call the event popularity path matching dataset, extract the frequency of popularity node triggering, count the number of times the node appears in the differentiated path, and obtain the probability of popularity node triggering. S502: Based on the trigger probability of the hot nodes, calculate the node propagation value, combine the trigger frequency and path distribution to statistically analyze the node contribution, sort the hot priority in descending order of propagation value, filter the paths with high propagation values and classify them into key monitoring, and obtain the key hot paths. S503: Based on the key heat paths, collect the heat events under the key monitoring paths, count the scope of influence and the number of event types, and generate a heat dissipation collection information table for events.
[0014] A big data-based sports event popularity analysis system includes: The event popularity identification module extracts event name, participating teams, and match time based on event-related data, calculates the corresponding offset of popularity time, compares the degree of overlap of event regions, filters the popularity of target events, summarizes the feature combination of popularity patterns, and obtains the event popularity pattern. The heat band construction module, based on the aforementioned event heat pattern, compares the cross-frequency of event regions, filters event heat that fits the time window, analyzes the influence relationship of differentiated nodes, summarizes the attribution of multiple event heats, adjusts the transmission path of heat bands, filters bands with large propagation factors, and obtains heat band structure information. Based on the information about the trending band structure, the trending event attribution module calculates the crossover frequency of events in the band, analyzes the co-occurrence relationship of events, identifies the attribution boundaries of trending events, filters out trending events with conflicting attributions, adjusts the band attribution of events with uneven distribution, and obtains the trending event attribution mapping. Based on the hot event attribution mapping, the heat path reconstruction module extracts the core heat nodes of the path, adjusts the order of the heat path, compares the heat correlation, summarizes logically similar paths, filters key path derivative relationships, and obtains the event heat correlation path. The heat impact assessment module extracts the trigger frequency of heat nodes based on the heat association path of the event, calculates the impact diffusion value, filters heat nodes with a large impact range, summarizes the high-heat event paths, and obtains the event heat propagation aggregation information table.
[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, by structuring and organizing event information, standardized expression of event types and geographical distribution is achieved, avoiding statistical inconsistencies caused by data clutter. By identifying peak intervals of popularity and merging related bands, a continuous popularity sequence is formed, avoiding trend distortion caused by fragmented data. By statistically analyzing frequencies and evaluating weights in different popularity bands, hierarchical classification of popularity can be achieved, ensuring the accuracy of comparisons between multiple types of events. By extracting key nodes and constructing popularity path matching data, it is ensured that related events can be identified within the time window, avoiding the impact of scattered popularity on overall judgment. By calculating and sorting propagation values, propagation aggregation information with clear priorities is formed, making it easier to track the popularity trends of key events, thereby improving the completeness and relevance of event popularity analysis. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a detailed schematic diagram of S3 of the present invention; Figure 5 This is a detailed schematic diagram of S4 of the present invention; Figure 6This is a detailed schematic diagram of S5 of the present invention; Figure 7 This is a system module diagram of the present invention. Detailed Implementation
[0018] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0019] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0020] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0021] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0022] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0023] Please see Figure 1 This invention provides a method for analyzing the popularity of sports events based on big data, including the following steps: S1: Based on the event association data, extract the event name, participating teams, and match time, identify and classify the event type and geographical distribution, and organize the data in a structured manner to obtain a standardized event information dataset; S2: Based on the standardized event information dataset, extract the event popularity timestamps, arrange them in chronological order, identify the time interval between adjacent popularity peaks, determine whether the popularity fluctuations are continuous, filter related popularity into the same popularity band, calculate the popularity band propagation factor, divide the popularity propagation path, and obtain the event popularity time series. S3: Call the event popularity time series, count the frequency of event types in the differentiated popularity bands, evaluate the popularity attribution weight, determine whether the popularity is merged or split, and obtain the hierarchical event popularity analysis results. S4: Based on the hierarchical event popularity analysis results, key popularity nodes are extracted, the correlation between time windows is analyzed, and the popularity of the correlation is filtered and classified into the same type of popularity path to obtain the event popularity path matching dataset. The direct relationship between the "Event Popularity Path Matching Dataset" and popularity lies in its further structuring and strengthening of the continuity, aggregation, and dissemination of popularity. Its core function is to compare key popularity nodes obtained through hierarchical analysis according to time sequence and offset characteristics, identify the time intervals and trigger sequences between nodes, and then filter out highly correlated combinations of popularity nodes. It also clusters similar event types to generate attribution paths, thus avoiding trend distortion caused by the dispersion of single-point popularity. This allows popularity originally scattered across different times or nodes to be uniformly aggregated along the path dimension. Through this dataset, popularity is reorganized into traceable propagation paths, strengthening the temporal correlation between events and ensuring the integration of similar popularity, thus providing a solid foundation for subsequent node propagation value calculation, priority ranking, and key monitoring. Therefore, it directly determines the attribution method and propagation structure of popularity, transforming event popularity from "scattered fluctuations" to "path-based continuous propagation," achieving accurate characterization and in-depth analysis of overall popularity. S5: Call the event popularity path matching dataset, calculate the node propagation value, sort the popularity priority according to the propagation value, filter the propagation paths and classify them into key monitoring, and obtain the event popularity propagation collection information table.
[0024] The standardized event information dataset includes event name, participating teams, and subject-verb-object structure parsing results; the event popularity time series includes popularity timestamps, popularity fluctuation judgment results, and popularity band propagation factors; the hierarchical event popularity analysis results include popularity attribution weights, event co-occurrence probabilities, and popularity hierarchy classifications; the event popularity path matching dataset includes popularity time offsets, correlation analysis results, and correlation popularity classifications; and the event popularity propagation aggregation information table includes popularity priority, node propagation values, and key monitoring paths.
[0025] Please see Figure 2 The specific steps for standardizing the event information dataset are as follows: S101: Based on event-related data, including news reports, social media comments, and forum posts, extract the event name, participating teams, and match time. After removing duplicate records, obtain the event information matching data volume. Based on event-related data, the system accurately identifies and extracts event names, participating teams, and match times from multiple online information sources, including news reports, social media comments, and forum posts. For example, from an online news report about sports, it identifies "Global Esports Challenge" as the event name, "TeamAlpha" and "TeamBeta" as participating teams, and "August 10th" as the match time. Simultaneously, it captures information such as "World Football League," "Real Madrid vs. Barcelona," and "August 12th" from big data streams on social media platforms. Using a rule-based text parsing engine and a specially trained named entity recognition model, it performs character-by-character scanning and structured extraction of the data content from each information source. After extraction, a duplicate record removal operation is performed on all captured event information. This process involves calculating unique hash values for three key data fields: event name, core participating team list (e.g., retaining only the two teams mentioned most frequently), and match time. Subsequently, the hash values are strictly compared. For example, if two independent records generate completely identical hash signatures after hash calculation, the second record is determined to be redundant data and removed. This operation ensures the uniqueness of each event record in the dataset. For instance, in the initial data collection phase, a total of 1250 original records were obtained. After precise deduplication, 1187 unique and non-repeating event information records were ultimately retained, resulting in the event information matching data volume.
[0026] S102: Based on the amount of data matching the event information, analyze the semantic relationship between the event name, participating teams and the match time, extract the subject-verb-object structure, combine the event field to map the corresponding event type, merge the event names under the same event type, count the number of participating teams for multiple event types, and classify them according to the threshold of the number of event types to obtain the event type distribution interval dataset. Based on the amount of event information matching data, a dataset containing 1187 unique event records is received. Each record includes the event name, participating teams, and match time. The semantic relationship between the event name, participating teams, and match time is analyzed. By constructing a detailed dependency parsing tree and semantic role labeling model, the core constituent elements and their logical relationships are identified from the event name. For example, for the event name "Honor of Kings Professional League Summer Season," "Honor of Kings" is identified as the game product category, "Professional League" is determined as the event organization level, and "Summer Season" is identified as a specific event period feature. The subject-verb-object structure is extracted, and standardized "subject-verb-object" structures are further extracted from the event-related descriptive text. For example, from the text "Team A defeated Team B and advanced," "Team A" is accurately identified as the subject, "defeated" as the verb, and "Team B" as the object. Combining the event field mapping with the corresponding event type, the parsed semantic features, such as the game category "Honor of Kings," the sports event "football," and the action verb "defeated," are precisely mapped one-to-one with the preset event type classification system. For example, "Honor of Kings" is mapped to "...". MOBA esports events, "football" mapped to "traditional football events," merging event names under the same event type. All specific event names mapped to the same event type are logically aggregated. For example, event names such as "Chinese Super League" and "English Premier League" are all merged into the unified type "football events." The number of participating teams across multiple event types is counted. For each merged event type, such as "football events," the total number of unique participating teams involved in all event names under that type is iterated. For example, the "football events" type has 385 different professional football clubs counted. These are then categorized based on a threshold for the number of participating teams. A preset threshold is used for each event type category. For example, event types with fewer than 60 participating teams are categorized as "regional events," those between 60 and 250 as "national events," and those with more than 250 as "international events." For instance, a basketball event type with 180 participating teams is categorized as a "national event," and a football event type with 385 participating teams is categorized as an "international event," resulting in a dataset of event type distribution intervals.
[0027] S103: Call the event type distribution interval dataset, classify and analyze the regional distribution by event name, group by region, industry category and event level, cross-match event types and summarize key event types to generate a standardized event information dataset; The dataset receives a distribution range of competition types, clearly indicating the various competition types and their classification in terms of the number of participating teams. For example, "football" is precisely classified as "international competition," while "basketball" is classified as "national competition." The dataset is then analyzed by competition name to determine geographical distribution. For each competition name in the dataset, such as "Premier League," specific geographical attributes, such as "United Kingdom," are extracted from its associated competition information sources. The frequency of media mentions and user attention for the competition in different geographical regions (e.g., countries, provinces, specific cities) are also analyzed. For example, by analyzing global media data, the frequency of mentions of "Premier League" in "United Kingdom," "China," and "United States" is identified as 9000, 7500, and 4000 times, respectively. The dataset is then grouped by region, industry category, and competition level. All competition data are categorized according to their geographical attributes (e.g., "Asia," "Europe"), industry category (e.g., "sports," "e-sports"), and competition level (e.g., [missing information - likely a specific category]). The "International Leagues" and "National Leagues" groups are cross-grouped in multiple dimensions. For example, the "English Premier League" is precisely classified under the "European-Sports-National Leagues" group. The event types are cross-matched and key event types are summarized. The results of the multi-dimensional grouping are cross-compared with the previously obtained event type distribution interval dataset to identify the event types that appear most frequently or have a significant impact on user attention under specific regional, industry and level combinations. For example, in the "European-Sports-International Leagues" group, "Football (UEFA Champions League)" is identified and summarized as the key event type. At the same time, its core attributes, such as "top event" and "cross-border influence", are extracted to generate a standardized event information dataset. This dataset stores all cleaned, classified and summarized event information in a unified structure. Each record contains a standard event name, precise event type, determined event level, clear regional attributes, a summarized list of key attributes and the corresponding popularity timestamp sequence.
[0028] Please see Figure 3 The specific steps for determining the time series of event popularity are as follows: S201: Based on the popularity timestamps of the standardized event information dataset, arrange them in chronological order, identify the time intervals between adjacent popularity peaks, filter out popularity peaks that do not exceed the threshold, and generate a sequence of event popularity time intervals; Based on the popularity timestamps of a standardized event information dataset, a popularity timestamp sequence associated with each event item is extracted. This sequence accurately records the changes in the event's popularity or search index at different time points. For example, for the event "World Cup Football Tournament," its popularity timestamp data sequence is [{Time: 11-2018:00:00, Popularity Value: 75}, {Time: 11-21 09:00:00, Popularity Value: 92}, ..., {Time: 12-18 23:00:00, Popularity Value: 80}]. Arranged in chronological order, the popularity timestamp data points for each event are strictly sorted in ascending order to ensure that all data points are accurately organized and processed according to the chronological order of their occurrence. The time interval between adjacent popularity peaks is identified. First, a peak identification algorithm based on relative gain is applied to determine whether the popularity value at a certain time point is higher than the popularity value within a specific time window before and after it. For example, if a certain popularity value is higher than... If the average of all popularity values within 12 hours before and after a point is at least 15% higher, then that point is identified as a popularity peak. Subsequently, the time difference between two adjacent identified popularity peaks is calculated. For example, if the first peak occurs at 09:00:00 on November 21, and the next peak occurs at 14:00:00 on November 26, then the time interval is calculated as 5 days and 5 hours. Popular peaks that do not exceed the threshold are filtered out. A time interval threshold is preset. For example, the time interval between adjacent popularity peaks must be less than or equal to 8 days. This threshold is an empirical value set based on statistical analysis of the periodic characteristics of the popularity spread of various events. If the time interval between a pair of adjacent peaks, for example, 10 days, exceeds this preset threshold, then the pair of peaks is considered to lack continuous spread. Only those continuous popularity peaks whose time interval does not exceed the threshold are retained for subsequent analysis. For example, the above 5-day and 5-hour interval does not exceed the 8-day threshold and is retained, thus obtaining the event popularity time interval sequence.
[0029] S202: Based on the time interval sequence of event popularity, analyze the regional consistency of the popularity peaks that exceed the time threshold, extract the region, event level and event type, filter the popularity peaks that meet the conditions and classify them into the same popularity band, and count the number of popularity peaks, time span and number of regions within the band to obtain the popularity band statistics. Based on the event popularity time interval sequence, which contains a series of popularity peak records with adjacent time intervals not exceeding an 8-day threshold, for popularity peaks exceeding the time threshold, those popularity peaks with time intervals exceeding a preset threshold (e.g., 8 days) are precisely screened from the original popularity timestamp data. The screened peaks usually mark significant interruptions in the popularity propagation chain. For example, if a peak of an event occurs on April 10th, and the next related peak occurs on April 25th, the interval is 15 days, exceeding the 8-day threshold, then this peak will be marked as a peak exceeding the time threshold. The regional consistency of the event is analyzed, and detailed regional information tracing and cross-analysis are performed on the popularity peaks identified as exceeding the time threshold. For example, by analyzing the amount of regional news releases and social media discussion heat related to the event during the peak occurrence period, it is determined whether there is significant inconsistency in its main regional distribution. For example, if a peak is mainly contributed by audience attention in China, while another distant peak is mainly driven by attention in North America, then it is determined that there is regional inconsistency, and regional and event-level data are extracted. Ignoring the event type, from the metadata associated with each identified popularity peak, precisely extract its regional attributes (e.g., "North America"), event level (e.g., "International"), and event type (e.g., "Basketball"), and filter popularity peaks that meet the criteria and group them into the same popularity band. Based on the results of the regional consistency judgment and the consistency of metadata such as event level and event type, cluster those popularity peaks that have large intervals on the timeline but maintain consistent regional attributes, event level, and event type to form a unified "popularity band." For example, if... If two peaks 15 days apart both belong to the "North America - International - Basketball" category, they are grouped into the same trending band. The number of trending peaks, time span, and geographical scope within the band are counted. For each established trending band, the total number of trending peaks contained within it is precisely calculated. For example, if a band contains 7 trending peaks, the complete time span covered by the band from the first peak to the last peak is calculated, for example, a total of 35 days. The total number of different geographical scopes involved in the band is also counted, for example, the band involves 5 different states. The trending band statistics are then obtained.
[0030] S203: Call the heat band statistics, calculate the heat band propagation factor, divide the heat propagation path according to the number of heat peaks, time span and number of regions, and generate the event heat time series; The system retrieves statistical data for the popularity bands, for example, {Band A: {Number of peaks: 5, Time span: 20, Number of regions: 3}}. Based on the number of peaks, time span, and number of regions, it calculates the popularity band propagation factor. For example, the propagation factor is calculated as (Number of peaks / Time span) × Number of regions. If the propagation factor of Band A is (5 / 20) × 3 = 0.75, a threshold is set based on the propagation factor value. For example, a value above 0.6 indicates strong propagation, 0.3 to 0.6 indicates medium propagation, and a value below 0.3 indicates weak propagation. For example, if the propagation factor of Band A is 0.75, which is higher than 0.6, it is classified as a strong propagation path. This generates a time series of event popularity data, which contains the propagation characteristics of each event's popularity at different time points. For example, {Event Name: World Cup, Popularity Band: [{Band ID: A, Propagation Factor: 0.75, Propagation Path Type: Strong Propagation Path}]}.
[0031] Please see Figure 4 The specific steps for analyzing the tiered event popularity results are as follows: S301: Call the event popularity time series, count the frequency of event types in the differentiated popularity bands, classify by event name, event level and event type, and obtain event type frequency distribution data; The system retrieves the time series data on the popularity of various events, which records the popularity bands and propagation characteristics of each event. For example, it tracks the strong propagation path of the "World Cup" event. It iterates through all popularity bands and, based on the propagation path type (e.g., "strong propagation path"), counts the total number of times each event type (e.g., "football" appears 50 times in the "strong propagation path" band and "basketball" appears 30 times in the "medium propagation path" band. The statistical results are then aggregated and categorized at multiple levels by event name, event level (e.g., "international"), and event type. For example, it generates frequency data for "World Cup - International - Football," which is further subdivided into "World Cup - International - Football - Strong Propagation Path Occurrence Frequency" (50 times) to obtain the frequency distribution data for each event type.
[0032] S302: Based on the frequency distribution data of event types, identify the co-occurrence probability of events, extract the combination of event types in the differentiated popularity bands, count the number of times the combination occurs, and calculate the popularity attribution weight according to the probability threshold based on the number of co-occurrences and the total popularity, and generate the popularity attribution weight matrix. The weighting of popularity is determined by the following formula: ; in, Represents the weighting of popularity. The popularity value of event i at the k-th frequency. The popularity value of event j at the k-th frequency. The total number of frequencies; Based on the frequency distribution data of event types, this data records in detail the frequency of different event types appearing in various popularity bands, identifies the co-occurrence probability of events, and identifies the frequency of different event types appearing simultaneously by analyzing the combinations of event types in differentiated popularity bands. For example, in the "strong propagation path" band, the event types "football" and "basketball" appear simultaneously, indicating a co-occurrence relationship. Event type combinations are extracted from differentiated popularity bands. All popularity bands are traversed to extract all event type combinations contained therein, such as "football-e-sports". The frequency of occurrence of this combination is counted. All extracted event type combinations are counted, and the total number of times each combination appears in different popularity bands is counted. For example, the "football-e-sports" combination appears 10 times in all "strong propagation path" bands. Based on the co-occurrence frequency and the total popularity, the popularity attribution weight is calculated. ,in Represents the weighting of popularity, quantifying the competition. With the event The degree of difference in heat at different frequencies, Representative competition In the The normalized heat value in the frequency range reflects the event's... The level of attention or influence in a specific communication context ranges from 0 to 1. For example, it can be obtained by normalizing the maximum peak popularity within a band to 1 and scaling the peak value proportionally. Representative competition In the Normalized heat value in the frequency range, The total number of samples representing frequencies refers to the total number of samples from different frequencies or intensity bands considered when calculating the co-occurrence probability. For example, if considering two frequencies, "strong propagation path" and "medium propagation path," In the formula Indicates all frequencies Perform accumulation operations. Computational competition and events At a single frequency The absolute difference in thermal values indicates the degree of thermal dissimilarity at that frequency. The summation of popularity dissimilarity across all frequencies reflects the overall difference in popularity curves between the two events. The denominator... Normalization using the square root of the sum of squares aims to eliminate the influence of differences in overall popularity levels between different events on the weighting calculation, making the results more comparable. This formula reflects the degree of similarity in popularity dissemination behavior by measuring the differences in popularity values of two events at various frequencies. The lower the value, the more similar the popularity curves of the two events are, that is, the higher the weight of popularity attribution; Example calculation: Assumption The event (Football) The popularity values at high and medium frequencies are respectively ; Event (Esports) are respectively molecule is ; The denominator is ; Therefore, the weighting of popularity The results show that the popularity dissemination patterns of football and e-sports are highly similar, and the degree of popularity attribution is high. The advantage of the formula is that by calculating the multi-dimensional popularity differences, it can accurately quantify the similarity of the popularity dissemination patterns of the events, identify the event combinations with synergistic effects, and generate a popularity attribution weight matrix.
[0033] S303: Call the popularity attribution weight matrix, divide the popularity according to the attribution weight, merge the popularity above the threshold, split the popularity below the threshold, and obtain the hierarchical event popularity analysis results. The popularity attribution weight matrix is invoked, which quantifies the degree of popularity difference between different event types at different frequencies. For example, the elements in the matrix... A threshold for assigning popularity weights is set, for example, a threshold of 0.25. This threshold is determined by clustering historical event popularity data to identify the critical points for different similarity levels. Simultaneously, it is calibrated by incorporating the experience of domain experts regarding the correlation between events, ensuring that the threshold effectively distinguishes between strong and weak correlations. All popularity weights in the matrix corresponding to popularity weights below this threshold are considered "strongly correlated popularity." For example, if... A score less than 0.25 is considered a strong correlation in popularity. Scores above or equal to the threshold are considered "weakly correlated popularity." Popularity scores below the threshold are merged. The system merges all event combinations classified as "strongly correlated popularity" and their popularity values to form a higher-level popularity aggregation. For example, the co-occurrence popularity of the "football-e-sports" combination is merged into a unified "comprehensive competitive popularity" event. Popularity scores above or equal to the threshold are broken down. For those with affiliation weights above or equal to the threshold, for example... For event combinations with a correlation coefficient greater than 0.25, the correlation coefficient is determined to be "weakly correlated". The correlation coefficient is then split and no longer treated as a related event. Instead, the correlation coefficient of each event is calculated and analyzed separately. For example, the correlation coefficient of "basketball" is analyzed independently of that of "volleyball", resulting in a hierarchical analysis of event correlation coefficients.
[0034] Please see Figure 5 The specific steps for matching the event popularity path dataset are as follows: S401: Based on the hierarchical event popularity analysis results, extract key popularity nodes, arrange the popularity nodes according to time order, filter the nodes with offset features on the time axis, calculate the time difference between adjacent nodes, and obtain the key popularity time offset. Based on the hierarchical event popularity analysis results, the analysis results are received, and the event popularity has been merged and split according to the attribution weight. Aggregated popularity events with high popularity values or significant attribution weights, such as weights higher than 0.004, are identified as key popularity nodes. Each node includes event type, popularity intensity, and occurrence time. For example, {Node ID: N1, Event Type: Comprehensive Competition, Popularity Intensity: 0.8, Time: 07-25 10:00:00}. The key popularity nodes are arranged in ascending order of occurrence time. By comparing the expected and actual occurrence times of adjacent popularity nodes, for example, if a node is expected to occur on 07-26 but actually occurs on 07-28, it is determined to have a 2-day time offset feature. The time difference between adjacent nodes is calculated. For example, if node N1 occurs on 07-25 10:00:00 and node N2 occurs on 07-27 15:00:00, the time difference between the two is 2 days and 5 hours, and the key popularity time offset is obtained.
[0035] S402: Based on the key popularity time offset, analyze the correlation within the time window, identify the temporal relationship between adjacent popularity nodes, identify the correlation strength between nodes based on the time interval and the triggering order between nodes, filter the combination of popularity nodes with a correlation greater than a preset threshold, and obtain the time correlation dataset. Based on the key heat time offset, offset data is received. For example, the time difference between N1 and N2 is 2 days and 5 hours. A time window threshold is set, for example, from 1 hour to 72 hours. If the time difference between nodes falls within this window, a correlation is considered to exist. By sliding the time window, potential correlations between nodes are analyzed. The order of adjacent nodes is directly obtained from the offset. For example, N1 always appears before N2. Based on the time interval and the triggering order between nodes, the correlation strength between nodes is identified. For example, the correlation strength is 1 / (time interval in hours) × (order matching coefficient). The order matching coefficient is set to 1. If the time interval between N1 and N2 is 53 hours, then the correlation strength is 1 / 53≈0.0189. A correlation threshold is set, for example, 0.015, to filter node combinations with significant time correlations. The calculated correlation strength is compared with this threshold. For example, if 0.0189 is greater than 0.015, then a significant correlation is considered to exist between N1 and N2, resulting in a time correlation dataset.
[0036] S403: Based on the time-related relationship dataset, filter highly correlated hot nodes, cluster nodes belonging to the same event type, merge similar nodes and divide the belonging path to generate an event hot path matching dataset. Based on a time-related dataset, the system receives a dataset, for example, the association degree of combination (N1, N2) is 0.0189. It then filters out nodes with significantly higher association degrees than the average, such as the top 20% of nodes (those with association degrees above 0.015), as core associated nodes. For these highly associated nodes, it performs clustering based on the types of events they are associated with. For example, it groups all key nodes related to the "football" event type into the "football" cluster. It then logically merges nodes of the same type within each cluster, for example, aggregating multiple highly associated nodes belonging to the "football" event type into a "football popularity path." This path reflects the continuous propagation trajectory of football popularity over time, and each path is assigned a unique path ID, such as "Path_Football," generating an event popularity path matching dataset.
[0037] Please see Figure 6 The specific steps for compiling the event popularity dissemination information table are as follows: S501: Call the event popularity path matching dataset, extract the trigger frequency of popularity nodes, count the number of times the node appears in the differentiated path, and obtain the trigger probability of popularity nodes; The system calls the event popularity path matching dataset and receives the dataset, for example, {"Path_Football": [node F1, node F2]}. It iterates through each path and identifies the occurrence count of each popular node in the path. For example, in "Path_Football", node F1 appears 3 times and node F2 appears 2 times. For each popular node, such as node F1, it counts the total number of times it appears in all different event popularity paths. For example, node F1 appears 3 times in "Path_Football" and 1 time in "Path_SportsNews", for a total of 4 times. It compares the total number of times each popular node appears in all paths with the number of times it appears in a specific path to calculate its trigger frequency in the specific path. Then it divides the trigger frequency by the total number of node occurrences to obtain a standardized ratio. For example, if node F1 appears 3 times in "Path_Football" and appears a total of 4 times, then its trigger ratio in "Path_Football" is 3 / 4 = 0.75, thus obtaining the trigger probability of the popular node.
[0038] S502: Calculate the node propagation value based on the trigger probability of hot nodes, combine the trigger frequency and path distribution to count the node contribution, sort the hot priority in descending order of propagation value, filter the paths with high propagation value and classify them into key monitoring, and obtain the key hot paths. Based on the trigger probability of hot nodes, trigger ratio data is received. For example, node F1 has a trigger ratio of 0.75 in "Path_Football". Combining the trigger ratio of each hot node with the path distribution characteristics of that node, such as path length and the number of regions involved, the comprehensive propagation value of that node is calculated. For example, propagation value = trigger ratio × (path length coefficient + number of regions coefficient), where the path length coefficient is 0.6 and the number of regions coefficient is 0.4. If node F1 is in a path with a length of 10 nodes and involves 3 regions, its propagation value is... The propagation value is 0.75×(0.6×10+0.4×3)=5.4. The propagation value is regarded as the node contribution and the heat priority is sorted in descending order. For example, if the propagation value of node F1 is 5.4 and the propagation value of node E1 is 4.8, then F1 has a higher priority than E1. A propagation value threshold is set, for example, 5.0. The paths of nodes with propagation values higher than this threshold are identified as "critical heat paths". For example, if the propagation value of node F1 is 5.4, which is higher than 5.0, its "Path_Football" path is marked as a key monitoring path, thus obtaining the critical heat path.
[0039] S503: Based on key heat paths, collect heat events under key monitoring paths, count the scope of impact and the number of event types, and generate a heat dissipation collection information table for events. Based on key trending paths, such as "Path_Football" being marked as a key monitoring path, all paths identified as "key trending paths" are traversed, and all related trending events under the path are centrally aggregated. For example, the peak trending data, related news mentions, and social media discussions under the "Path_Football" path are aggregated into a "Football Event Communication Event Set". For each aggregated trending event set, its geographical influence is statistically analyzed, such as the number of countries or regions involved (e.g., affecting 20 countries globally), and the number of event types included (e.g., even though it is a "football" path, if it includes "e-sports football", it is counted as 1). An event trending dissemination aggregation information table is generated, which details the aggregated events, influence range, number of event types involved, and other core indicators for each key trending path.
[0040] Please see Figure 7 A big data-based sports event popularity analysis system includes: The event popularity identification module extracts event name, participating teams, and match time based on event-related data, calculates the corresponding offset of popularity time, compares the degree of overlap of event regions, filters the popularity of target events, summarizes the feature combination of popularity patterns, and obtains the event popularity pattern. The heat band construction module is based on the event heat pattern, compares the cross frequency of event regions, filters the event heat that fits the time window, analyzes the influence relationship of differentiated nodes, summarizes the attribution of multiple event heat, adjusts the transmission path of heat bands, filters bands with large propagation factors, and obtains heat band structure information. The hot event attribution module calculates the crossover frequency of events in the bands based on the hot band structure information, analyzes the co-occurrence relationship of events, identifies the hot event attribution boundary, filters hot events with conflicting attributions, adjusts the band attribution of events with uneven distribution, and obtains the hot event attribution mapping. The heat path reconstruction module extracts the core heat nodes of the path based on the heat event attribution mapping, adjusts the heat path order, compares the heat correlation, summarizes logically similar paths, filters key path derivative relationships, and obtains the event heat correlation path. The heat impact assessment module extracts the trigger frequency of heat nodes based on the heat association path of the event, calculates the impact diffusion value, filters heat nodes with a large impact range, summarizes the high heat event paths, and obtains the event heat propagation information table.
[0041] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for analyzing the popularity of sports events based on big data, characterized in that, Includes the following steps: S1: Based on the event association data, extract the event name, participating teams, and match time, identify and classify the event type and geographical distribution, and organize the data in a structured manner to obtain a standardized event information dataset; S2: Based on the standardized event information dataset, extract the event popularity timestamps, arrange them in chronological order, identify the time interval between adjacent popularity peaks, determine whether the popularity fluctuations are continuous, filter related popularity into the same popularity band, calculate the popularity band propagation factor, divide the popularity propagation path, and obtain the event popularity time series. S3: Call the event popularity time series, count the frequency of event types in the differentiated popularity bands, evaluate the popularity attribution weight, determine whether the popularity is merged or split, and obtain the hierarchical event popularity analysis results. S4: Based on the hierarchical event popularity analysis results, extract key popularity nodes, analyze the correlation between time windows, filter the popularity of the correlation and classify them into the same type of popularity path, and obtain the event popularity path matching dataset.
2. The method for analyzing the popularity of sports events based on big data according to claim 1, characterized in that, The standardized event information dataset includes event name, participating teams, and subject-verb-object structure parsing results. The event popularity time series includes popularity timestamps, popularity fluctuation judgment results, and popularity band propagation factors. The hierarchical event popularity analysis results include popularity attribution weights, event co-occurrence probability, and popularity hierarchy classification. The event popularity path matching dataset includes popularity time offset, correlation analysis results, and correlation popularity classification.
3. The method for analyzing the popularity of sports events based on big data according to claim 1, characterized in that, The specific steps for creating the standardized event information dataset are as follows: S101: Based on event-related data, including news reports, social media comments, and forum posts, extract the event name, participating teams, and match time. After removing duplicate records, obtain the event information matching data volume. S102: Based on the amount of matching data for the event information, analyze the semantic relationship between the event name, participating teams and the match time, extract the subject-verb-object structure, combine the event field to map the corresponding event type, merge the event names under the same event type, count the number of participating teams for multiple event types, and classify them according to the event type number threshold to obtain the event type distribution interval dataset. S103: Call the event type distribution interval dataset, classify and analyze the regional distribution by event name, group by region, industry category and event level, cross-match event types and summarize key event types to generate a standardized event information dataset.
4. The method for analyzing the popularity of sports events based on big data according to claim 3, characterized in that, The specific steps for the time series analysis of event popularity are as follows: S201: Based on the heat timestamps of the standardized event information dataset, arrange them in chronological order, identify the time intervals between adjacent heat peaks, filter out heat peaks that do not exceed the threshold, and generate a sequence of event heat time intervals; S202: Based on the time interval sequence of the event popularity, for the popularity peaks that exceed the time threshold, analyze the regional consistency of the event, extract the region, event level and event type, filter the popularity peaks that meet the conditions and classify them into the same popularity band, count the number of popularity peaks, time span and number of regions within the band, and obtain the popularity band statistics data. S203: Call the aforementioned heat band statistics, calculate the heat band propagation factor, divide the heat propagation path according to the number of heat peaks, time span and number of regions, and generate the event heat time series.
5. The method for analyzing the popularity of sports events based on big data according to claim 4, characterized in that, The specific steps for analyzing the hierarchical event popularity results are as follows: S301: Call the event popularity time series, count the frequency of event types in the differentiated popularity bands, classify them by event name, event level and event type, and obtain event type frequency distribution data; S302: Based on the frequency distribution data of the event types, identify the co-occurrence probability of events, extract the combination of event types in the differentiated popularity bands, count the number of times the combination occurs, and calculate the popularity attribution weight according to the probability threshold based on the number of co-occurrences and the total popularity, and generate the popularity attribution weight matrix. S303: Call the aforementioned popularity attribution weight matrix, divide popularity according to attribution weight, merge popularity above the threshold, split popularity below the threshold, and obtain hierarchical event popularity analysis results.
6. The method for analyzing the popularity of sports events based on big data according to claim 5, characterized in that, The weighting of popularity is determined using the following formula: ; in, Represents the weighting of popularity. The popularity value of event i at the k-th frequency. The popularity value of event j at the k-th frequency. The total number of frequencies.
7. The method for analyzing the popularity of sports events based on big data according to claim 5, characterized in that, The specific steps for matching the event popularity path dataset are as follows: S401: Based on the hierarchical event popularity analysis results, extract key popularity nodes, arrange the popularity nodes according to time order, filter nodes with offset features on the time axis, calculate the time difference between adjacent nodes, and obtain the key popularity time offset. S402: Based on the key heat time offset, analyze the correlation in the time window, identify the temporal relationship between adjacent heat nodes, identify the correlation strength between nodes according to the time interval and the triggering order between nodes, filter the heat node combination with a correlation greater than a preset threshold, and obtain the time correlation dataset. S403: Based on the time-related relationship dataset, filter nodes with high relevance and popularity, cluster nodes belonging to the same event type, merge nodes of the same type and divide the belonging paths to generate an event popularity path matching dataset.
8. The method for analyzing the popularity of sports events based on big data according to claim 1, characterized in that, The method also includes step S5: S5: Call the event popularity path matching dataset, calculate the node propagation value, sort the popularity priority according to the propagation value, filter the propagation paths and classify them into key monitoring, and obtain the event popularity propagation collection information table; The event popularity dissemination information table includes popularity priority, node dissemination value, and key monitoring paths.
9. The method for analyzing the popularity of sports events based on big data according to claim 8, characterized in that, The specific steps for compiling the event popularity dissemination information table are as follows: S501: Call the event popularity path matching dataset, extract the frequency of popularity node triggering, count the number of times the node appears in the differentiated path, and obtain the probability of popularity node triggering. S502: Based on the trigger probability of the hot nodes, calculate the node propagation value, combine the trigger frequency and path distribution to statistically analyze the node contribution, sort the hot priority in descending order of propagation value, filter the paths with high propagation values and classify them into key monitoring, and obtain the key hot paths. S503: Based on the key heat paths, collect the heat events under the key monitoring paths, count the scope of influence and the number of event types, and generate a heat dissipation collection information table for events.
10. A sports event popularity analysis system based on big data, characterized in that, The system is used to implement the big data-based sports event popularity analysis method according to any one of claims 1-9, and the system includes: The event popularity identification module extracts event name, participating teams, and match time based on event-related data, calculates the corresponding offset of popularity time, compares the degree of overlap of event regions, filters the popularity of target events, summarizes the feature combination of popularity patterns, and obtains the event popularity pattern. The heat band construction module, based on the aforementioned event heat pattern, compares the cross-frequency of event regions, filters event heat that fits the time window, analyzes the influence relationship of differentiated nodes, summarizes the attribution of multiple event heats, adjusts the transmission path of heat bands, filters bands with large propagation factors, and obtains heat band structure information. Based on the information about the trending band structure, the trending event attribution module calculates the crossover frequency of events in the band, analyzes the co-occurrence relationship of events, identifies the attribution boundaries of trending events, filters out trending events with conflicting attributions, adjusts the band attribution of events with uneven distribution, and obtains the trending event attribution mapping. Based on the hot event attribution mapping, the heat path reconstruction module extracts the core heat nodes of the path, adjusts the order of the heat path, compares the heat correlation, summarizes logically similar paths, filters key path derivative relationships, and obtains the event heat correlation path. The heat impact assessment module extracts the trigger frequency of heat nodes based on the heat association path of the event, calculates the impact diffusion value, filters heat nodes with a large impact range, summarizes the high-heat event paths, and obtains the event heat propagation aggregation information table.