Sports event content comprehensive processing method based on artificial intelligence
By improving the BERTopic model and combining multiple analytical methods, the problem of comprehensive processing of sports events is solved, and the accurate judgment and communication effect of sports events is achieved.
Patent Information
- Application Number
- CN202510267304.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing technology cannot effectively use artificial intelligence-based models to comprehensively process sports events, and thus cannot accurately analyze and judge the hot spots of sports events.
By improving the BERTopic model, combining hierarchical analysis, gray correlation analysis and fuzzy evaluation method, a heat index judgment matrix is constructed, and the preheat event weight and heat index weight are generated, and the comprehensive sports event attention event is determined.
It has achieved timely and accurate determination of hot events in sports events, improved the dissemination effect of sports events-related content, made it more in line with the audience's interests, and increased the audience's stickiness.
Smart Images

Figure CN120196752A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sports event content processing, and particularly to a comprehensive processing method for sports event content based on artificial intelligence. Background Art
[0002] Sports event content is an important channel for people to participate in and watch sports. In today's mobile Internet era, people have a broad vision and diverse spiritual pursuits. As a result, they have more diverse concerns about major sports events than in the past, and the measurement dimension indicators for these concerns are also diverse.
[0003] Therefore, today's sports event communicators need to use an artificial intelligence-based model to theoretically analyze and speculate on concerns through measurement dimension indicators, and then quickly and objectively determine the relevant events of the sports events that people are most concerned about. Thus, communicators can timely and effectively promote the dissemination of sports events based on the events that people are most concerned about, making the dissemination process of sports events more in line with the interests of the audience, narrowing the distance from the audience, increasing the stickiness of the audience's attention to sports events, and creating a good in-depth communication environment for sports events.
[0004] Currently, for example, the patent document with the application number 202311620846.0 discloses a method for identifying the technical themes of scientific and technological projects based on the BERTopic theme recognition model. However, it does not disclose how to comprehensively process the themes output by the artificial intelligence model BERTopic, and thus cannot be applied to the analysis of sports event hotspots.
[0005] Based on this, how to use an artificial intelligence-based model to comprehensively process sports event content and generate sports events with the highest comprehensive attention is a technical problem to be solved at present. Summary of the Invention
[0006] For this purpose, the present invention provides a comprehensive processing method for sports event content based on artificial intelligence. By improving the BERTopic model, it realizes the mining of the sports event heat event network based on artificial intelligence. Through the analytic hierarchy process, grey relational analysis, and fuzzy evaluation method, it determines the sports event with high attention that synthesizes the weights of multiple heat indicators and the time-series heat trend, and realizes the timely and accurate determination of the hot events of sports events, which is conducive to the dissemination and promotion of sports event-related content that is more in line with the interests of the audience.
[0007] To achieve the above object, the present invention proposes a comprehensive processing method for sports event content based on artificial intelligence, including:
[0008] Generating multiple heat events from the sports event content data transmitted over the network through a heat word mining model based on the improved BERTopic;
[0009] Construct a heat index judgment matrix for multiple heat indexes corresponding to multiple of the heat events through the expert evaluation method, and generate a pre-heat event weight and a heat index weight for the heat index judgment matrix through a heat analysis model based on the analytic hierarchy process;
[0010] Generate an index correlation degree for the time series data of multiple of the heat indexes through a correlation analysis model based on grey relational analysis, and adjust the pre-heat event weight according to the index correlation degree to generate a time series heat event weight;
[0011] Generate an evaluation matrix for the heat index through grade evaluation, and determine a comprehensive sports event attention event through a heat event judgment model based on the fuzzy evaluation method for the evaluation matrix, the time series heat event weight, and the heat index weight.
[0012] Furthermore, the heat word mining model includes a BERT text embedder, a UMAP dimensionality reducer, a self-attention mechanism dimensionality selector, and an HDBSCAN text clusterer. The process of generating multiple heat events for the sports event content data propagated on the network through the heat word mining model based on improved BERTopic includes:
[0013] Generate text embedding vectors for the sports event content data through the BERT text embedder;
[0014] Perform dimensionality reduction processing on the text embedding vectors through the UMAP dimensionality reducer to generate multiple dimensionality-reduced text vectors with reduced dimensions;
[0015] Generate optimal dimensionality-reduced text vectors with optimal reduced dimensions for the dimensionality-reduced text vectors of multiple dimensions respectively through the self-attention mechanism dimensionality selector;
[0016] Perform text clustering calculation on the optimal dimensionality-reduced text vectors through the HDBSCAN text clusterer to generate representations of the heat events.
[0017] Furthermore, the self-attention mechanism dimensionality selector includes a self-attention mechanism, an output layer, and a dimensionality selection layer. The process of generating optimal dimensionality-reduced text vectors with optimal reduced dimensions for the dimensionality-reduced text vectors of multiple dimensions respectively through the self-attention mechanism dimensionality selector includes:
[0018] Perform context awareness on the dimensionality-reduced text vectors through the self-attention mechanism to generate an embedding matrix;
[0019] Classify the embedding matrix through the output layer to determine the class probability of the dimensionality-reduced text;
[0020] Calculate the similarity of the category probabilities through the dimension selection layer to determine the optimal dimensionality-reduced text vector.
[0021] Further, the dimension selection layer includes a channel similarity calculation module, a time window similarity calculation module, and a comprehensive selection module. The process of calculating the similarity of the category probabilities through the dimension selection layer to determine the optimal dimensionality-reduced text vector includes:
[0022] Generate a channel similarity for the publishing channels of the category probabilities through the channel similarity calculation module;
[0023] Generate a time distribution similarity for the time window sequence of the category probabilities through the time window similarity calculation module;
[0024] Calculate and determine the optimal dimensionality-reduced text vector through the comprehensive selection module for the channel similarity and the time distribution similarity.
[0025] Further, the channel similarity calculation module is constructed based on the cosine similarity of the same publishing channels and different publishing channels;
[0026] The time window similarity is constructed based on the cosine similarity of adjacent time window sequences.
[0027] In the above solution, the problem that the BERTopic model performs static dimensionality reduction, reducing the performance of the clustering algorithm and resulting in a mismatch with diverse sports event texts is overcome, and the model's learning and mining of context texts are strengthened through the self-attention mechanism, achieving a better clustering and mining effect of the BERTopic model on texts.
[0028] Further, the hot word mining model also includes a tokenizer and a keyword determiner based on TF-IDF;
[0029] Generate multiple candidate keywords for the hot event through the tokenizer;
[0030] Generate keywords associated with the hot event through the keyword determiner for the candidate keywords.
[0031] Further, the process of generating an evaluation matrix of hotness indicators through rank evaluation includes:
[0032] Obtain multiple hotness indicators of the hot event containing the keywords;
[0033] Generate the evaluation matrix through normalization and hierarchical judgment for the hotness indicators;
[0034] Among them, the hotness indicators include the number of event videos, the number of event clicks, the number of event comments, the number of event follows, and the number of event reports.
[0035] In the above solution, keywords determined based on the improved BERTopic model are realized, and an evaluation matrix accurate and conforming to the network data situation is generated based on the keywords, so as to further accurately judge hot events of sports competitions.
[0036] Further, the process of constructing a heat index judgment matrix for multiple heat indexes through the expert evaluation method includes:
[0037] Obtaining the scores of the heat indexes in the questionnaire published on the network;
[0038] Constructing the heat index judgment matrix according to the scores.
[0039] Further, the process of generating index correlation degrees for time series data of multiple heat indexes through a correlation analysis model based on grey correlation analysis and calculating and adjusting the weights of pre-heat events according to the index correlation degrees to generate time series heat event weights includes:
[0040] Calculating the correlation coefficients between multiple time series data;
[0041] Performing arithmetic mean on the correlation coefficients to obtain the index correlation degrees;
[0042] Multiplying the weights of the pre-heat events by the index correlation degrees to obtain the time series heat event weights.
[0043] Further, the process of determining the comprehensive sports competition attention events through a heat event judgment model based on the fuzzy evaluation method for the evaluation matrix, time series heat event weights, and heat index weights includes:
[0044] Calculating the event heat score by multiplying the evaluation matrix by the heat index weights;
[0045] Calculating the total heat score by multiplying the index score by the time series heat event weights;
[0046] Determining the heat events with the total heat score exceeding the total score threshold and the event heat score exceeding the event score threshold as the comprehensive sports competition attention events.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows.
[0048] 1. Through the improved BERTopic model, network mining of sports competition heat events based on artificial intelligence is realized. By using the analytic hierarchy process, grey correlation analysis, and fuzzy evaluation method, high-concern sports competition events integrating multiple heat index weights and time series heat trends are determined, realizing timely and accurate determination of hot events of sports competitions, which is beneficial to the dissemination and promotion of sports competition-related content that is more in line with the interests of the audience.
[0049] 2. It overcomes the problem that the BERTopic model performs static dimensionality reduction, which reduces the performance of the clustering algorithm and leads to a mismatch with diverse sports event texts. The model strengthens the learning and mining of context texts through the self-attention mechanism, achieving a better clustering and mining effect of the BERTopic model on texts.
[0050] 3. It realizes the generation of an accurate evaluation matrix that conforms to the network data situation based on the keywords determined by the improved BERTopic model, and thus can make a more accurate judgment of hot events in sports competitions. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic diagram of the general process of the comprehensive processing method for sports event content based on artificial intelligence according to an embodiment of the present invention;
[0052] Figure 2 It is a schematic diagram of the detailed process of determining heat events in the comprehensive processing method for sports event content based on artificial intelligence according to an embodiment of the present invention;
[0053] Figure 3 It is a schematic diagram of the structure of the heat word mining model based on the improved BERTopic in the comprehensive processing method for sports event content based on artificial intelligence according to an embodiment of the present invention;
[0054] Figure 4 It is a schematic diagram of the detailed process of determining comprehensive sports event attention events in the comprehensive processing method for sports event content based on artificial intelligence according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0056] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and do not limit the protection scope of the present invention.
[0057] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0058] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0059] As Figures 1 to 4 shown, the present invention provides a comprehensive processing method for sports event content based on artificial intelligence. By improving the BERTopic model, it realizes the mining of the sports event heat event network based on artificial intelligence. Through the analytic hierarchy process, grey relational analysis, and fuzzy evaluation method, it determines the highly concerned sports event with comprehensive weights of multiple heat indexes and time-series heat trends, and realizes the timely and accurate determination of the hot events of sports events, which is beneficial to the dissemination and promotion of sports event-related content that is more in line with the interests of the audience.
[0060] As Figures 1 to 4 shown, this embodiment proposes a comprehensive processing method for sports event content based on artificial intelligence, including:
[0061] Generating multiple heat events from the sports event content data spread on the network through a heat word mining model based on the improved BERTopic;
[0062] Constructing a heat index judgment matrix for multiple heat indexes corresponding to the multiple heat events through the expert evaluation method, and generating a pre-heat event weight and a heat index weight for the heat index judgment matrix through a heat analysis model based on the analytic hierarchy process;
[0063] Generating an index correlation degree for the time-series data of the multiple heat indexes through a correlation analysis model based on grey relational analysis, and adjusting the pre-heat event weight according to the index correlation degree to generate a time-series heat event weight;
[0064] Generating an evaluation matrix for the heat index through grade evaluation, and determining the comprehensive sports event attention event through a heat event judgment model based on the fuzzy evaluation method for the evaluation matrix, the time-series heat event weight, and the heat index weight.
[0065] It is understandable that BERTopic, as a pre-trained topic model, discovers the potential semantic structure information of documents by mapping the high-dimensional word vector space to a low-dimensional topic space. It is a mainstream topic model based on artificial intelligence / deep learning. It uses a pre-trained model to represent documents through embedding and then uses a clustering method to identify topics. BERTopic performs well in topic tasks. However, the clusters it outputs are heat events of multiple topic categories, which are classified only according to text semantics and do not consider the heat metrics attached to the heat events on the network. Therefore, it is impossible to determine the heat events that people are most concerned about.
[0066] It should be noted that the heat events are preferably reports on the heat events of online sports events, such as a certain athlete winning the championship, a certain athlete competing against a certain athlete, etc., and can also be content related to the organization of sports events, such as various promotional activities of local characteristic events. Thus, the disseminators and reporters of sports events can capture the latest heat events in real time and conduct more timely reports. The heat events and the heat metrics can be obtained by exporting through the API interface of the sports event website.
[0067] Furthermore, as Figure 2 and 3 shown, the heat word mining model includes a BERT text embedder, a UMAP dimensionality reducer, a self-attention mechanism dimensionality selector, and an HDBSCAN text clustering algorithm. The process of generating multiple heat events from the sports event content data spread on the network through the heat word mining model based on the improved BERTopic includes:
[0068] Generating text embedding vectors from the sports event content data through the BERT text embedder;
[0069] Reducing the dimensionality of the text embedding vectors through the UMAP dimensionality reducer to generate multiple dimensionality-reduced text vectors;
[0070] Generating optimal dimensionality-reduced text vectors with the optimal dimensionality from the dimensionality-reduced text vectors of multiple dimensions through the self-attention mechanism dimensionality selector respectively;
[0071] Performing text clustering calculation on the optimal dimensionality-reduced text vectors through the HDBSCAN text clustering algorithm to generate the representation of the heat events.
[0072] Specifically, the BERT text embedder (Bidirectional Encoder Representations from Transformers) is a context-based pre-trained model capable of generating dynamic semantic vector representations. It is preferably bert-base-chinese, with a default learning rate of e-4, a batch size of 256, and uses the Adam optimizer. The UMAP (Uniform Manifold Approximation and Projection) reducer is an efficient non-linear dimensionality reduction algorithm for static dimensionality reduction tasks of datasets, suitable for processing text embedding vectors of high-dimensional embeddings generated by BERT. It preferably uses WordPiece tokenization. HDBSCAN (Hierarchical Density-Based Clustering) belongs to the density clustering algorithm and is applicable to the text vectors after dimensionality reduction. Its preferred minimum cluster size is 5, and the noise label is -1.
[0073] It can be understood that the popularity event is preferably a report on the popularity event of a network sports event. The popularity word mining model is preferably used to process the report on the popularity event of a network sports event. The report has the text semantic characteristics of clear logic, standardized word usage, and prominent key points. Therefore, by inserting a self-attention mechanism dimension selector for further semantic extraction, a better clustering and classification effect can be obtained, which is more in line with the characteristics of the report on the popularity event of a sports event.
[0074] Furthermore, as Figure 2 and 3 shown, the self-attention mechanism dimension selector includes a self-attention mechanism, an output layer, and a dimension selection layer. The process of generating the optimal dimensionality reduction text vector with the optimal dimensionality reduction from the dimensionality reduction text vectors of multiple dimensions through the self-attention mechanism dimension selector includes:
[0075] Performing context awareness on the dimensionality reduction text vector through the self-attention mechanism to generate an embedding matrix;
[0076] Classifying the embedding matrix through the output layer to determine the class probability of the dimensionality reduction text;
[0077] Calculating the similarity of the class probability through the dimension selection layer to determine the optimal dimensionality reduction text vector.
[0078] It can be understood that by embedding the self-attention mechanism into BERTopic and setting dynamic dimensionality reduction, the semantic focusing ability of the clustering process of the HDBSCAN text clustering algorithm is enhanced, and the model's learning ability for texts of different lengths and different emphases is enhanced. As a pre-trained model, BERTopic itself has problems in the topic classification of specific sports event texts, such as information dilution of long texts and insufficient capture of local dependencies through the default pooling dimensionality reduction strategy. For example, key "tactical adjustments" details are missed in long sports event reports. The self-attention mechanism (Self-Attention) trained and learned through the dataset of sports event texts can significantly improve the topic classification accuracy and semantic focusing ability of sports event texts. For example, key details such as "tactical adjustments" and "physical fitness advantages" can be captured as secondary topics in long sports event reports.
[0079] Specifically, the output layer uses a fully connected layer and a Softmax activation function to achieve output class probabilities. The specific processes of the self-attention mechanism and the output layer can be expressed as:
[0080] q, k, v = linear(x)
[0081]
[0082] Output = softmax(MLP(X))
[0083] In the formula, q, k, v represent the query vector, key vector, and value vector obtained by linearly transforming the input dimensionality-reduced text vector x through linear, Q, K, V represent the matrices of all query vectors, all key vectors, and all value vectors respectively, d is the dimension of the vector, Attention(Q, K, V) represents the self-attention mechanism, Output represents the class probability output by the output layer, MLP represents the multi-layer perceptron, and X represents the embedding matrix output by the self-attention mechanism.
[0084] Furthermore, as Figure 2 and 3 shown, the dimension selection layer includes a channel similarity calculation module, a time window similarity calculation module, and a comprehensive selection module. The process of determining the optimal dimensionality-reduced text vector by calculating the similarity of the class probability through the dimension selection layer includes:
[0085] Generating a channel similarity for the publishing channel of the class probability through the channel similarity calculation module;
[0086] Generating a time distribution similarity for the time window sequence of the class probability through the time window similarity calculation module;
[0087] The optimal dimensionality-reduced text vector is determined by calculating the channel similarity and the time distribution similarity through the comprehensive selection module.
[0088] It can be understood that the dimensionality of UMAP dimensionality reduction directly affects the effect of downstream text clustering tasks. Traditional methods rely on the clustering silhouette coefficient of downstream tasks to select dimensions, but ignore the data channel source and time dynamics. Therefore, the dimensionality selection layer is used to measure the effect of dimensionality reduction to achieve scientific dimensionality selection.
[0089] Preferably, the dimensionality selection layer determines the optimal dimensionality reduction among dimensionality reductions of 10D, 20D, 30D, and 50D.
[0090] Furthermore, the channel similarity calculation module is constructed based on the cosine similarity of the same publishing channels and different publishing channels;
[0091] The time window similarity is constructed based on the cosine similarity of adjacent time window sequences.
[0092] Specifically, the operation process of the channel similarity calculation module can be expressed as:
[0093]
[0094] In the formula, S channel represents the channel similarity, and E S , E D respectively represent the cosine similarity of document embedding vectors of the same publishing channels classified into the same category by the output layer and the cosine similarity of document embedding vectors of different publishing channels, both of which are obtained from the embedding matrix. It can be understood that the higher the cosine similarity E S of the document embedding vectors of the same publishing channels, the better, and the lower the cosine similarity E D of the document embedding vectors of different publishing channels, the better.
[0095] Specifically, the channel preferably refers to the publishing institutions of sports event news reports, such as domestic news agencies, international news agencies, new media platforms, and event official agencies. Domestic news agencies such as Xinhua News Agency and CCTV Sports, international news agencies such as the Associated Press, new media platforms such as major sports websites, and event official agencies such as the International Olympic Committee.
[0096] Specifically, the operation process of the time window similarity calculation module can be expressed as:
[0097]
[0098] In the formula, S time represents the time window similarity, and E n , E n-1respectively represent the mean vectors of the nth time window and the (n - 1)th time window classified into the same category by the output layer, both of which are obtained from the embedding matrix.
[0099] Specifically, the data is divided into multiple time windows at fixed intervals, and the mean vector of the embedding vectors classified into the same category by the output layer in each window is calculated. The fixed interval is preferably 1 hour or 3 hours, and thus the time window similarity can measure the temporal stability.
[0100] Specifically, the operation process of the comprehensive selection module can be expressed as:
[0101] S(d) = λS time +(1 - λ)S channel
[0102]
[0103] In the formula, S is the comprehensive similarity, λ represents the weight coefficient, preferably 0.3, S channel and S time respectively represent the channel similarity and the time window similarity, d * represents the optimal dimensionality reduction, d represents the current dimensionality reduction, D is the set of candidate dimensions, and D preferably includes 10D, 20D, 30D, 50D.
[0104] In the above solution, the problem that the BERTopic model performs static dimensionality reduction, reducing the performance of the clustering algorithm and resulting in a mismatch with diverse sports event texts is overcome, and the model's learning and mining of context texts are strengthened through the self-attention mechanism, achieving a better clustering and mining effect of the BERTopic model on texts.
[0105] Furthermore, as Figure 2 and 3 shown, the hot word mining model further includes a tokenizer and a TF-IDF-based keyword determiner;
[0106] Generate multiple candidate keywords for the hot event through the tokenizer;
[0107] Generate keywords associated with the hot event from the candidate keywords through the keyword determiner.
[0108] Specifically, the tokenizer is Jieba or CountVectorizer, and the keyword determiner determines the keywords by evaluating the importance of each candidate keyword in the clustering through the TF-IDF method and MNR.
[0109] Furthermore, as Figure 4 shown, the process of generating the evaluation matrix of the hotness index through rank evaluation includes:
[0110] Obtain multiple heat metrics of heat events containing the keyword;
[0111] Generate the evaluation matrix by normalizing and grading the heat metrics;
[0112] Among them, the heat metrics include the number of event videos, the number of event clicks, the number of event comments, the number of event followers, and the number of event reports.
[0113] In the above solution, based on the keyword determined by the improved BERTopic model, an accurate evaluation matrix that conforms to the network data situation is generated according to the keyword, and then a more accurate judgment of hot events in sports competitions can be made.
[0114] Further, as Figure 4 shown, the process of constructing a heat metric judgment matrix for multiple heat metrics corresponding to multiple heat events through the expert evaluation method includes:
[0115] Obtain the scores for the heat metrics in the questionnaire published on the network;
[0116] Construct the heat metric judgment matrix according to the scores.
[0117] Specifically, the questionnaire includes a comparative score for the importance between two heat metrics. For example, if metric 1 is significantly more important than metric 2, the score is 7; conversely, if metric 2 is more important than metric 1, the score is 1 / 7. Generate the index judgment matrix corresponding to the scores. After geometric mean normalization of the index judgment matrix, obtain the weight vector and perform a consistency test. If the consistency test fails, remove the questionnaire with the largest extreme difference and regenerate the heat metric judgment matrix.
[0118] Further, as Figure 4 shown, the process of generating the index correlation degree for the time series data of multiple heat metrics through a correlation analysis model based on grey relational analysis and calculating and adjusting the weights of pre-heat events according to the index correlation degree to generate the time series heat event weights includes:
[0119] Calculate the correlation coefficients between multiple time series data, specifically:
[0120]
[0121] In the formula, γ i,k represents the correlation coefficient of the i-th heat event at the k-th moment, N represents the total number of moments, x i (j) represents the value of the i-th heat event at the j-th moment, and x0(j) represents the value of the reference sequence at the j-th moment.
[0122] Arithmetically average the correlation coefficients to obtain the index correlation degree, specifically:
[0123]
[0124] In the formula, γ i,k represents the correlation coefficient of the i-th heat event at the k-th moment, N represents the total number of moments, and r i represents the index correlation degree.
[0125] Multiply the pre-heat event weight by the index correlation degree to obtain the time-series heat event weight, specifically:
[0126] w i ′ = w i × r i
[0127] In the formula, r i represents the index correlation degree, and w i ′ and w i represent the time-series heat event weight and the pre-heat event weight respectively.
[0128] Furthermore, as Figure 4 shown, the process of determining the comprehensive sports event attention events through the heat event judgment model based on the fuzzy evaluation method for the evaluation matrix, the time-series heat event weight, and the heat index weight includes:
[0129] Calculate the event heat score by multiplying the evaluation matrix by the heat index weight;
[0130] Calculate the total heat score by multiplying the index score by the time-series heat event weight;
[0131] Determine the heat events with the total heat score exceeding the total score threshold and the event heat score exceeding the event score threshold as the comprehensive sports event attention events.
[0132] Specifically, perform maximum-minimum normalization on the heat index with a unified time window by day to generate a dimensionless index to eliminate the dimension difference. Map the dimensionless index to a numerical interval, and each numerical interval corresponds to a fuzzy number representation. Integrate all the fuzzy number representations to generate an evaluation matrix. Multiply the evaluation matrix by the heat index weight to obtain the event heat score, and multiply the index score by the time-series heat event weight to obtain the total heat score.
[0133] Preferably, the total score threshold is 60, and the event score threshold is 80.
[0134] In this embodiment, the improved BERTopic model is used to realize the network mining of sports event popularity events based on artificial intelligence. The analytic hierarchy process, grey relational analysis and fuzzy evaluation method are used to determine the high-concern sports event events that synthesize the weights of multiple popularity indicators and the time-series popularity trend, realizing the timely and accurate determination of the hot events of sports events, which is beneficial to the dissemination and promotion of sports event-related content that is more in line with the interests of the audience. It overcomes the problem that the BERTopic model performs static dimensionality reduction, reducing the performance of the clustering algorithm and resulting in a mismatch with diverse sports event texts, and strengthens the model's learning and mining of context texts through the self-attention mechanism, achieving a better clustering and mining effect of the BERTopic model on texts. An evaluation matrix that is accurate and conforms to the network data situation is generated based on the keywords determined by the improved BERTopic model, and then a more accurate judgment of the hot events of sports events can be made.
[0135] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
[0136] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A comprehensive processing method for sports event content based on artificial intelligence, characterized in that: include: The sports event content data disseminated on the Internet is used to generate multiple hot events through the hot word mining model based on the improved BERTopic; The multiple heat indices corresponding to the multiple heat events are used to construct a heat index judgment matrix through an expert evaluation method, and the heat index judgment matrix is used to generate preheat event weights and heat index weights through a heat analysis model based on the hierarchical analysis method; Generate an index correlation degree through a correlation analysis model based on grey correlation analysis for the time series data of the plurality of heat indices, and adjust the preheating event weight according to the index correlation degree to generate a time series heat event weight; An evaluation matrix of the heat index is generated through grade evaluation, and the evaluation matrix, the temporal heat event weights and the heat index weights are used to determine the comprehensive sports event focus events through a heat event judgment model based on a fuzzy evaluation method.
2. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 1, characterized in that: The hot word mining model includes a BERT text embedder, a UMAP dimension reducer, a self-attention mechanism dimension selector, and a HDBSCAN text clusterer. The process of generating multiple hot events from the sports event content data disseminated on the Internet through the hot word mining model based on the improved BERTopic includes: Generate a text embedding vector for the sports event content data through the BERT text embedder; The text embedding vector is subjected to dimensionality reduction processing by the UMAP dimensionality reducer to generate a plurality of reduced dimensionality text vectors; The reduced-dimensionality text vectors of multiple dimensions are respectively generated by the self-attention mechanism dimension selector to generate the optimal reduced-dimensionality text vector with the optimal reduced dimension; The optimal dimension-reduced text vector is subjected to text clustering calculation by the HDBSCAN text clusterer to generate the hot event.
3. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 2 is characterized in that: The self-attention mechanism dimension selector includes a self-attention mechanism, an output layer and a dimension selection layer. The process of generating an optimal dimension reduction text vector with optimal dimension reduction by respectively passing the dimension reduction text vectors of multiple dimensions through the self-attention mechanism dimension selector includes: The reduced-dimensional text vector is context-aware through the self-attention mechanism to generate an embedding matrix; Classifying the embedding matrix through the output layer to determine the category probability of the reduced-dimensional text; The category probability is subjected to similarity calculation through the dimension selection layer to determine the optimal dimension-reduced text vector.
4. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 3 is characterized in that: The dimension selection layer includes a channel similarity calculation module, a time window similarity calculation module and a comprehensive selection module. The process of performing similarity calculation on the category probability through the dimension selection layer to determine the optimal dimension reduction text vector includes: The channel similarity calculation module is used to calculate the channel similarity of the publication channel of the category probability; The time window sequence of the category probability is used to generate time distribution similarity through the time window similarity calculation module; The channel similarity and the time distribution similarity are calculated by the comprehensive selection module to determine the optimal dimension-reduced text vector.
5. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 4 is characterized in that: The channel similarity calculation module is constructed based on the cosine similarity of the same posting channel and different posting channels; The time window similarity is constructed based on the cosine similarity of adjacent time window sequences.
6. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 2, characterized in that: The hot word mining model also includes a word segmenter and a keyword determiner based on TF-IDF; Generate multiple candidate keywords from the hot event through the word segmenter; The candidate keywords are used by the keyword determiner to generate keywords associated with the hot event.
7. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 6 is characterized in that: The process of generating an evaluation matrix of heat indexes through grade evaluation includes: Obtain multiple heat indexes of heat events containing the keyword; Generate the evaluation matrix by normalizing and grading the heat index; The popularity index includes the number of event videos, the number of event clicks, the number of event comments, the number of event attentions and the number of event reports.
8. The method for comprehensive processing of sports event content based on artificial intelligence according to any one of claims 1 to 7, characterized in that: The process of constructing a heat index judgment matrix by expert evaluation method for multiple heat indexes corresponding to multiple heat events includes: Obtaining scores of the heat index in a questionnaire published on the Internet; The heat index judgment matrix is constructed according to the scores.
9. The method for comprehensive processing of sports event content based on artificial intelligence according to any one of claims 1 to 7, characterized in that: The process of generating the index correlation degree by using the correlation analysis model based on grey correlation analysis to generate the index correlation degree of the time series data of multiple heat indexes, and calculating and adjusting the preheating event weight according to the index correlation degree to generate the time series heat event weight includes: Calculate the correlation coefficient between multiple time series data; Taking the arithmetic mean of the correlation coefficients to obtain the index correlation degree; The preheating event weight is multiplied by the indicator correlation to obtain the temporal heat event weight.
10. The method for comprehensive processing of sports event content based on artificial intelligence according to any one of claims 1 to 7, characterized in that: The process of determining the events of interest in comprehensive sports events by using the evaluation matrix, the time series heat event weights and the heat index weights through the heat event judgment model based on the fuzzy evaluation method includes: Calculate the evaluation matrix and the heat index weight to obtain an event heat score; Calculate the index score and the temporal heat event weight to obtain a total heat score; The hot events whose total heat score exceeds the total score threshold and whose event heat score exceeds the event score threshold are determined as the comprehensive sports event attention events.
Citation Information
Patent Citations
Science and technology project technology theme identification method based on BERTopic theme identification model
CN117725212A
Prediction method for network topic popular degree
CN106557552A
Online public opinion popularity value quantitative identification method based on grey correlation analysis
CN111414550A
Interactive hierarchical topic modeling visual analysis method and device based on BERTopic
CN117407520A
Cited By
Online sports event management system based on reliable transmission of event information
CN120768955A
Online sports event activity management system based on reliable transmission of event information
CN120768955B