A comprehensive processing method for sports event content based on artificial intelligence
By improving the BERTopic model and combining it with multiple analysis methods and clusterers, the accuracy problem of comprehensive processing of sports event content was solved, timely and accurate judgment of hot events in sports events was achieved, and the relevance of communication was improved.
Patent Information
- Application Number
- CN202510267304.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-03-07
AI Technical Summary
Existing technologies are unable to effectively utilize AI-based models to comprehensively process sports event content, and are unable to accurately determine the sports events that people are most concerned about, resulting in dissemination that does not match the audience's interests.
By improving the BERTopic model and combining the hierarchical analysis method, grey relational analysis and fuzzy evaluation method, the weights and temporal popularity trends of multiple hot events are generated. The BERT text embedder, UMAP dimension reducer, self-attention mechanism dimension selector and HDBSCAN text clusterer are used to mine the hot event network and generate accurate sports event attention events.
It has achieved timely and accurate identification of hot events in sports events, improved the dissemination of content related to sports events, and enhanced the matching with audience interests.
Smart Images

Figure CN120196752B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sports event content processing, and in particular to a sports event content comprehensive processing method based on artificial intelligence. Background Art
[0002] Sports content is an important channel for people to participate in and watch sports. In today's mobile internet era, people have a broader perspective and diverse spiritual pursuits, and therefore have a more diverse focus on major sports events than before, and the measurement dimensions of these focus are also diverse.
[0003] Therefore, today's communicators of sports events need to use measurement dimension indicators to conduct theoretical analysis and speculation on the focus points through artificial intelligence-based models, and then quickly and objectively determine the related events of sports events that people are most concerned about, so that communicators can timely and effectively promote sports events based on the events that people are most concerned about, make the dissemination process of sports events more in line with the audience's interests, shorten the distance with the audience, increase the audience's attention to sports events, and create a good in-depth communication environment for sports events.
[0004] At present, for example, the patent document with application number 202311620846.0 discloses a method for identifying technical topics of scientific and technological projects based on the BERTopic topic recognition model. However, it does not disclose how to comprehensively process the topics output by the artificial intelligence model BERTopic, and thus cannot be applied to the analysis of hot spots of sports events.
[0005] Based on this, how to use artificial intelligence-based models to comprehensively process sports event content and generate the sports events with the highest overall attention is a technical problem that needs to be solved. Summary of the Invention
[0006] To this end, the present invention provides an artificial intelligence-based comprehensive processing method for sports event content, which realizes artificial intelligence-based sports event hot event network mining by improving the BERTopic model, and determines high-attention sports event events based on the weights of multiple hot indexes and temporal hot trends through hierarchical analysis method, grey correlation analysis and fuzzy evaluation method, thereby realizing timely and accurate determination of hot events in sports events, which is conducive to the dissemination and promotion of sports event-related content that is more in line with the audience's interests.
[0007] To achieve the above objectives, the present invention proposes a method for comprehensive processing of sports event content based on artificial intelligence, comprising:
[0008] The sports event content data disseminated on the Internet is used to generate multiple hot events through the hot word mining model based on the improved BERTopic;
[0009] A heat index judgment matrix is constructed by using an expert evaluation method for multiple heat indices corresponding to the multiple heat events, and the heat index judgment matrix is used to generate preheat event weights and heat index weights through a heat analysis model based on the hierarchical analysis method;
[0010] The time series data of the plurality of heat indices are subjected to a correlation analysis model based on grey correlation analysis to generate an index correlation, and the preheating event weight is adjusted according to the index correlation to generate a time series heat event weight;
[0011] An evaluation matrix of the heat index is generated through grade evaluation, and the evaluation matrix, the temporal heat event weights and the heat index weights are used to determine comprehensive sports event attention events through a heat event judgment model based on a fuzzy evaluation method.
[0012] Furthermore, the hot word mining model includes a BERT text embedder, a UMAP dimension reducer, a self-attention mechanism dimension selector, and an HDBSCAN text clusterer. The process of generating multiple hot events from the sports event content data disseminated on the Internet through the hot word mining model based on the improved BERTopic includes:
[0013] Generate a text embedding vector for the sports event content data through the BERT text embedder;
[0014] Performing dimensionality reduction processing on the text embedding vector through the UMAP dimensionality reducer to generate a plurality of reduced-dimensionality text vectors;
[0015] The reduced-dimensionality text vectors of multiple dimensions are respectively passed through the self-attention mechanism dimension selector to generate the optimal reduced-dimensionality text vector with the optimal reduced-dimensionality;
[0016] The optimal dimension-reduced text vector is subjected to text clustering calculation by the HDBSCAN text clusterer to generate a representation of the hot event.
[0017] Furthermore, the self-attention mechanism dimension selector includes a self-attention mechanism, an output layer, and a dimension selection layer. The process of generating an optimal reduced dimension text vector with optimal reduced dimension by respectively passing the reduced dimension text vectors of multiple dimensions through the self-attention mechanism dimension selector includes:
[0018] The reduced-dimensional text vector is context-aware through the self-attention mechanism to generate an embedding matrix;
[0019] Classifying the embedding matrix through the output layer to determine the category probability of the reduced-dimensionality text;
[0020] The category probabilities are subjected to similarity calculation through the dimension selection layer to determine the optimal dimensionality reduction text vector.
[0021] Furthermore, the dimension selection layer includes a channel similarity calculation module, a time window similarity calculation module, and a comprehensive selection module. The process of performing similarity calculation on the category probabilities through the dimension selection layer to determine the optimal dimension reduction text vector includes:
[0022] The channel similarity calculation module is used to calculate the channel similarity of the publication channel of the category probability;
[0023] The time window sequence of the class probability is passed through the time window similarity calculation module to generate time distribution similarity;
[0024] The channel similarity and the time distribution similarity are calculated by the comprehensive selection module to determine the optimal dimension-reduced text vector.
[0025] Furthermore, the channel similarity calculation module is constructed based on the cosine similarity between the same posting channel and different posting channels;
[0026] The time window similarity is constructed based on the cosine similarity of adjacent time window sequences.
[0027] The above solution overcomes the problem that the BERTopic model performs static dimensionality reduction, which reduces the performance of the clustering algorithm and leads to a mismatch with diverse sports event texts. The self-attention mechanism strengthens the model's learning and mining of contextual text, achieving a better clustering and mining effect for text using the BERTopic model.
[0028] Furthermore, the hot word mining model also includes a word segmenter and a TF-IDF-based keyword determiner;
[0029] Generate multiple candidate keywords from the hot event through the word segmenter;
[0030] The candidate keywords are used by the keyword determiner to generate keywords associated with the hot event.
[0031] Furthermore, the process of generating an evaluation matrix of heat indexes through grade evaluation includes:
[0032] Obtaining multiple heat indexes of heat events containing the keyword;
[0033] The heat index is normalized and graded to generate the evaluation matrix;
[0034] The popularity index includes the number of event videos, the number of event clicks, the number of event comments, the number of event attentions and the number of event reports.
[0035] In the above scheme, keywords are determined based on the improved BERTopic model, and an evaluation matrix that is accurate and fits the network data situation is generated based on the keywords, thereby making it possible to make more accurate judgments on hot events in sports events.
[0036] Furthermore, the process of constructing a heat index judgment matrix by using the expert evaluation method for the multiple heat indexes corresponding to the multiple heat events includes:
[0037] Obtaining scores for the popularity index from a questionnaire published online;
[0038] The popularity index judgment matrix is constructed according to the ratings.
[0039] Furthermore, the time series data of multiple heat indices are subjected to a correlation analysis model based on grey correlation analysis to generate an indicator correlation, and the preheating event weight is calculated and adjusted according to the indicator correlation to generate the time series heat event weight. The process includes:
[0040] Calculate the correlation coefficient between multiple time series data;
[0041] Performing arithmetic averaging on the correlation coefficients to obtain the indicator correlation degree;
[0042] The preheating event weight is multiplied by the indicator correlation to obtain the temporal heat event weight.
[0043] Furthermore, the process of determining the focus events of comprehensive sports events by using the evaluation matrix, the time series hot event weights, and the hot index weights through the hot event judgment model based on the fuzzy evaluation method includes:
[0044] Calculate the event heat score by combining the evaluation matrix and the heat index weight;
[0045] Calculate the index score and the temporal hot event weight to obtain a total hot score;
[0046] A hot event whose total heat score exceeds a total score threshold and whose event heat score exceeds an event score threshold is determined as the comprehensive sports event of interest event.
[0047] Compared with the prior art, the present invention has the following advantages:
[0048] 1. By improving the BERTopic model, we realized the network mining of hot events of sports events based on artificial intelligence. Through the hierarchical analysis method, grey correlation analysis and fuzzy evaluation method, we determined the high-profile sports events that integrated the weights of multiple hot indicators and the temporal hot trends. This made it possible to determine the hot events of sports events in a timely and accurate manner, which is conducive to the dissemination and promotion of sports event-related content that is more in line with the audience's interests.
[0049] 2. This overcomes the problem of static dimensionality reduction performed by the BERTopic model, which reduces the performance of the clustering algorithm and leads to a mismatch with diverse sports event texts. Furthermore, the self-attention mechanism is used to enhance the model's learning and mining of contextual texts, achieving a better clustering and mining effect for texts using the BERTopic model.
[0050] 3. We have realized the use of keywords determined based on the improved BERTopic model, and generated an accurate evaluation matrix that fits the network data situation based on the keywords, thereby enabling more accurate judgment of hot events in sports events. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a schematic diagram of the general flow of a method for comprehensive processing of sports event content based on artificial intelligence according to an embodiment of the present invention;
[0052] Figure 2 Detailed flowchart of determining hot events in the method for comprehensive processing of sports event content based on artificial intelligence according to an embodiment of the present invention;
[0053] Figure 3 This is a structural diagram of a hot word mining model based on an improved BERTopic in a method for comprehensive processing of sports event content based on artificial intelligence according to an embodiment of the present invention;
[0054] Figure 4 This is a detailed flowchart of determining comprehensive sports event focus events in the method for comprehensive sports event content processing based on artificial intelligence according to an embodiment of the present invention. DETAILED DESCRIPTION
[0055] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0056] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0057] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0058] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0059] like Figures 1 to 4 As shown, the present invention provides an artificial intelligence-based comprehensive processing method for sports event content, realizes artificial intelligence-based sports event hot event network mining by improving the BERTopic model, and determines high-attention sports event events based on the weights of multiple heat indexes and time series heat trends through the hierarchical analysis method, grey correlation analysis and fuzzy evaluation method, realizes timely and accurate determination of hot events in sports events, and is conducive to the dissemination and promotion of sports event-related content that is more in line with the audience's interests.
[0060] like Figures 1 to 4 As shown, this embodiment proposes a comprehensive processing method for sports event content based on artificial intelligence, including:
[0061] The sports event content data disseminated on the Internet is used to generate multiple hot events through the hot word mining model based on the improved BERTopic;
[0062] A heat index judgment matrix is constructed by using an expert evaluation method for multiple heat indices corresponding to the multiple heat events, and the heat index judgment matrix is used to generate preheat event weights and heat index weights through a heat analysis model based on the hierarchical analysis method;
[0063] The time series data of the plurality of heat indices are subjected to a correlation analysis model based on grey correlation analysis to generate an index correlation, and the preheating event weight is adjusted according to the index correlation to generate a time series heat event weight;
[0064] An evaluation matrix of the heat index is generated through grade evaluation, and the evaluation matrix, the temporal heat event weights and the heat index weights are used to determine comprehensive sports event attention events through a heat event judgment model based on a fuzzy evaluation method.
[0065] It's understandable that BERTopic, a pre-trained topic model, discovers the underlying semantic structure of documents by mapping high-dimensional word vectors into a low-dimensional topic space. It's a mainstream topic model based on artificial intelligence and deep learning. It uses a pre-trained model to embed documents and then uses clustering to identify topics. While BERTopic performs well on topic tasks, its output clusters into multiple topic categories. These events are categorized solely based on text semantics and don't consider the associated online attention metrics. Consequently, they can't identify the most engaging topics.
[0066] It should be noted that the hot events are preferably reports on online sports events, such as athlete X winning a championship, athlete X competing against athlete X, etc. Related content can also be organized for sports events, such as various local event promotional activities. This allows sports event communicators and reporters to capture the latest hot events in real time and provide more timely reporting. The hot events and hot indicators can be exported and obtained through the API interface of the sports event website.
[0067] Furthermore, if Figure 2 and 3 As shown, the hot word mining model includes a BERT text embedder, a UMAP dimension reducer, a self-attention mechanism dimension selector, and an HDBSCAN text clusterer. The process of generating multiple hot events from the sports event content data disseminated on the Internet through the hot word mining model based on the improved BERTopic includes:
[0068] Generate a text embedding vector for the sports event content data through the BERT text embedder;
[0069] Performing dimensionality reduction processing on the text embedding vector through the UMAP dimensionality reducer to generate a plurality of reduced-dimensionality text vectors;
[0070] The reduced-dimensionality text vectors of multiple dimensions are respectively passed through the self-attention mechanism dimension selector to generate the optimal reduced-dimensionality text vector with the optimal reduced-dimensionality;
[0071] The optimal dimension-reduced text vector is subjected to text clustering calculation by the HDBSCAN text clusterer to generate a representation of the hot event.
[0072] Specifically, the BERT text embedder (Bidirectional Encoder Representations from Transformers) is a context-based pre-trained model that can generate dynamic semantic vector representations, preferably bert-base-chinese, with a default learning rate of e-4, a batch size of 256, and an Adam optimizer. The UMAP (Uniform Manifold Approximation and Projection) dimensionality reducer is an efficient nonlinear dimensionality reduction algorithm used for static dimensionality reduction tasks of datasets. It is suitable for processing high-dimensional embedded text embedding vectors generated by BERT, and preferably uses WordPiece word segmentation. HDBSCAN (Hierarchical Density-Based Clustering) is a density clustering algorithm applicable to text vectors after dimensionality reduction, with a preferred minimum cluster size of 5 and a noise label of -1.
[0073] It can be understood that the hot events are preferably reports on hot events of online sports events, and the hot word mining model is preferably used to process reports on hot events of online sports events. The reports have the text semantic characteristics of clear logic, standardized wording and prominent emphasis. Therefore, by inserting the self-attention mechanism dimension selector for further semantic extraction, better clustering and classification effects can be obtained, which is more in line with the characteristics of hot event reports of sports events.
[0074] Further, if Figure 2 and 3 As shown, the self-attention mechanism dimension selector includes a self-attention mechanism, an output layer, and a dimension selection layer. The process of generating the optimal reduced dimension text vector with the optimal reduced dimension by respectively passing the reduced dimension text vectors of multiple dimensions through the self-attention mechanism dimension selector includes:
[0075] The reduced-dimensional text vector is context-aware through the self-attention mechanism to generate an embedding matrix;
[0076] Classifying the embedding matrix through the output layer to determine the category probability of the reduced-dimensionality text;
[0077] The category probabilities are subjected to similarity calculation through the dimension selection layer to determine the optimal dimensionality reduction text vector.
[0078] It is understandable that by embedding the self-attention mechanism into BERTopic and performing dynamic dimensionality reduction settings, the semantic focus capability of the clustering process of the HDBSCAN text clusterer is enhanced, and the model's learning ability for texts of different lengths and different focuses is enhanced. BERTopic itself, as a pre-trained model, has problems in the topic classification of specific sports event texts, such as the default pooling dimensionality reduction strategy, which leads to the dilution of long text information and insufficient capture of local dependencies. For example, in long sports event reports, key details such as "tactical adjustments" are omitted. The self-attention mechanism (Self-Attention), which is trained and learned through a dataset of sports event texts, can significantly improve the topic classification accuracy and semantic focus capability of sports event texts. For example, in long sports event reports, key details such as "tactical adjustments" and "physical advantages" can be captured as secondary topics.
[0079] Specifically, the output layer uses a fully connected layer and a Softmax activation function to output the class probability. The specific process of the self-attention mechanism and the output layer can be expressed as:
[0080] q,k,v=linear(x)
[0081]
[0082] Output = softmax(MLP(X))
[0083] Where q, k, v represent the query vector, key vector, and value vector obtained by linearly transforming the input reduced-dimensional text vector x. Q, K, V represent the matrix of all query vectors, the matrix of all key vectors, and the matrix of all value vectors, respectively. d is the dimension of the vector. Attention(Q, K, V) represents the self-attention mechanism, Output represents the category probability output by the output layer, MLP represents the multi-layer perceptron, and X represents the embedding matrix output by the self-attention mechanism.
[0084] Further, if Figure 2 and 3 As shown, the dimension selection layer includes a channel similarity calculation module, a time window similarity calculation module and a comprehensive selection module. The process of performing similarity calculation on the category probability through the dimension selection layer to determine the optimal dimensionality reduction text vector includes:
[0085] The channel similarity calculation module is used to calculate the channel similarity of the publication channel of the category probability;
[0086] The time window sequence of the class probability is passed through the time window similarity calculation module to generate time distribution similarity;
[0087] The channel similarity and the time distribution similarity are calculated by the comprehensive selection module to determine the optimal dimension-reduced text vector.
[0088] It's understandable that the dimensions reduced by UMAP directly impact the performance of downstream text clustering tasks. Traditional methods rely on the clustering silhouette coefficient of downstream tasks to select dimensions, but ignore the source of data channels and temporal dynamics. Therefore, a dimension selection layer is used to measure the effectiveness of dimensionality reduction and achieve scientific dimension selection.
[0089] Preferably, the dimension selection layer determines the optimal dimension reduction among the dimension reductions of 10D, 20D, 30D, and 50D.
[0090] Furthermore, the channel similarity calculation module is constructed based on the cosine similarity between the same posting channel and different posting channels;
[0091] The time window similarity is constructed based on the cosine similarity of adjacent time window sequences.
[0092] Specifically, the operation process of the channel similarity calculation module can be expressed as:
[0093]
[0094] Where S channel represents the channel similarity, E S 、E D They represent the cosine similarity of embedding vectors of documents with the same publishing channel and documents with different publishing channels that are classified into the same category by the output layer, respectively, both of which are obtained from the embedding matrix. It can be understood that the cosine similarity of embedding vectors of documents with the same publishing channel E S The higher the better, the cosine similarity E of document embedding vectors of different publishing channels D The lower the better.
[0095] Specifically, the channels preferably refer to publishers of sports event news reports, such as domestic news organizations, international news organizations, new media platforms, and official event organizations. Domestic news organizations include Xinhua News Agency and CCTV Sports, international news organizations include the Associated Press, new media platforms include major sports websites, and official event organizations include the International Olympic Committee.
[0096] Specifically, the operation process of the time window similarity calculation module can be expressed as:
[0097]
[0098] Where S time represents the time window similarity, E n 、E n-1They respectively represent the mean vectors of the nth time window and the n-1th time window classified into the same class by the output layer, both of which are obtained from the embedding matrix.
[0099] Specifically, the data is divided into multiple time windows at fixed intervals, and the mean vector is calculated for the embedding vectors classified into the same class in the output layer in each window. The fixed interval is preferably 1 hour or 3 hours, and the time window similarity can measure the temporal stability.
[0100] Specifically, the operation process of the comprehensive selection module can be expressed as:
[0101] S(d)=λS time +(1-λ)S channel
[0102]
[0103] In the formula, S is the comprehensive similarity, λ represents the weight coefficient, preferably 0.3, S channel 、S time Represent channel similarity and time window similarity respectively, d * represents the optimal dimension reduction, d represents the current dimension reduction, D represents the candidate dimension set, and D preferably includes 10D, 20D, 30D, and 50D.
[0104] The above solution overcomes the problem that the BERTopic model performs static dimensionality reduction, which reduces the performance of the clustering algorithm and leads to a mismatch with diverse sports event texts. The self-attention mechanism strengthens the model's learning and mining of contextual text, achieving a better clustering and mining effect for text using the BERTopic model.
[0105] Further, if Figure 2 and 3 As shown, the hot word mining model also includes a word segmenter and a TF-IDF-based keyword determiner;
[0106] Generate multiple candidate keywords from the hot event through the word segmenter;
[0107] The candidate keywords are used by the keyword determiner to generate keywords associated with the hot event.
[0108] Specifically, the word segmenter is Jieba or CountVectorizer, and the keyword determiner evaluates the importance of each candidate keyword in the cluster through the TF-IDF method and MNR to determine the keyword.
[0109] Further, if Figure 4 As shown in FIG, the process of generating an evaluation matrix of heat indexes through grade evaluation includes:
[0110] Obtaining multiple heat indexes of heat events containing the keyword;
[0111] The heat index is normalized and graded to generate the evaluation matrix;
[0112] The popularity index includes the number of event videos, the number of event clicks, the number of event comments, the number of event attentions and the number of event reports.
[0113] In the above scheme, keywords are determined based on the improved BERTopic model, and an evaluation matrix that is accurate and fits the network data situation is generated based on the keywords, thereby making it possible to make more accurate judgments on hot events in sports events.
[0114] Furthermore, if Figure 4 As shown in FIG, the process of constructing a heat index judgment matrix by using the expert evaluation method for multiple heat indexes corresponding to multiple heat events includes:
[0115] Obtaining scores for the popularity index from a questionnaire published online;
[0116] The popularity index judgment matrix is constructed according to the ratings.
[0117] Specifically, the questionnaire includes a comparative score of the importance between two heat indicators. For example, if indicator 1 is obviously more important than indicator 2, the score is 7. Conversely, if indicator 2 is scored 1 / 7 of indicator 1, the score is generated corresponding to the indicator judgment matrix. After the indicator judgment matrix is normalized by the geometric mean, the weight vector is obtained and a consistency test is performed. If the consistency test fails, the questionnaire with the largest extreme value difference is removed and the heat indicator judgment matrix is regenerated.
[0118] Furthermore, if Figure 4 As shown, the process of generating indicator correlation by using a correlation analysis model based on grey correlation analysis for the time series data of multiple heat indicators and calculating and adjusting the preheating event weight according to the indicator correlation to generate the time series heat event weight includes:
[0119] Calculate the correlation coefficient between multiple time series data, specifically:
[0120]
[0121] Where, γ i,k represents the correlation coefficient of the i-th hot event at the k-th moment, N represents the total number of moments, x i (j) represents the value of the i-th heat event at the j-th moment, and x0(j) represents the value of the reference sequence at the j-th moment.
[0122] The correlation coefficients are arithmetic averaged to obtain the index correlation degree, specifically:
[0123]
[0124] Where, γ i,k represents the correlation coefficient of the i-th hot event at the k-th moment, N represents the total number of moments, r i Indicates the correlation degree of indicators.
[0125] The preheating event weight is multiplied by the indicator correlation to obtain the temporal heat event weight, specifically:
[0126] w i ′=w i ×r i
[0127] Where r i Indicates the index correlation, w i ′、w i They represent the weight of temporal heat events and the weight of preheat events respectively.
[0128] Further, if Figure 4 As shown in FIG, the process of determining the focus events of comprehensive sports events by using the evaluation matrix, the time series hot event weights, and the hot index weights through the hot event judgment model based on the fuzzy evaluation method includes:
[0129] Calculate the event heat score by combining the evaluation matrix and the heat index weight;
[0130] Calculate the index score and the temporal hot event weight to obtain a total hot score;
[0131] A hot event whose total heat score exceeds a total score threshold and whose event heat score exceeds an event score threshold is determined as the comprehensive sports event of interest event.
[0132] Specifically, the maximum and minimum values of the heat index in the unified time window of each day are normalized to generate a dimensionless index to eliminate dimensional differences. The dimensionless index is corresponded to a numerical interval, and each numerical interval corresponds to a fuzzy number representation. All fuzzy number representations are integrated to generate an evaluation matrix. The evaluation matrix is multiplied by the heat index weight to obtain the event heat score, and the index score is multiplied by the time series heat event weight to obtain the total heat score.
[0133] Preferably, the total score threshold is 60, and the event score threshold is 80.
[0134] In this embodiment, an artificial intelligence-based network mining of hot events in sports events is achieved by improving the BERTopic model. The hierarchical analysis method, grey correlation analysis and fuzzy evaluation method are used to determine the high-profile sports events that integrate the weights of multiple hot indicators and the temporal hot trends. This allows for timely and accurate determination of hot events in sports events, which is conducive to the dissemination and promotion of sports event-related content that is more in line with the audience's interests. The problem of the BERTopic model performing static dimensionality reduction, which reduces the performance of the clustering algorithm and leads to mismatching of diverse sports event texts, is overcome. The model's learning and mining of contextual texts is enhanced through the self-attention mechanism, achieving a better clustering mining effect of the BERTopic model on text. Based on the keywords determined by the improved BERTopic model, an evaluation matrix that is accurate and fits the network data situation is generated based on the keywords, thereby making it possible to make more accurate judgments on hot events in sports events.
[0135] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
[0136] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A comprehensive processing method for sports event content based on artificial intelligence, characterized in that: include: The sports event content data disseminated on the Internet is used to generate multiple hot events through the hot word mining model based on the improved BERTopic; A heat index judgment matrix is constructed by using an expert evaluation method for multiple heat indices corresponding to the multiple heat events, and the heat index judgment matrix is used to generate preheat event weights and heat index weights through a heat analysis model based on the hierarchical analysis method; The time series data of the plurality of heat indices are subjected to a correlation analysis model based on grey correlation analysis to generate an index correlation, and the preheating event weight is adjusted according to the index correlation to generate a time series heat event weight; An evaluation matrix of the heat index is generated through grade evaluation, and the evaluation matrix, the temporal heat event weights and the heat index weights are used to determine comprehensive sports event attention events through a heat event judgment model based on a fuzzy evaluation method.
2. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 1, characterized in that: The hot word mining model includes a BERT text embedder, a UMAP dimension reducer, a self-attention mechanism dimension selector, and an HDBSCAN text clusterer. The process of generating multiple hot events from the sports event content data disseminated on the Internet through the hot word mining model based on the improved BERTopic includes: Generate a text embedding vector for the sports event content data through the BERT text embedder; Performing dimensionality reduction processing on the text embedding vector through the UMAP dimensionality reducer to generate a plurality of reduced-dimensionality text vectors; The reduced-dimensionality text vectors of multiple dimensions are respectively passed through the self-attention mechanism dimension selector to generate the optimal reduced-dimensionality text vector with the optimal reduced-dimensionality; The optimal dimension-reduced text vector is subjected to text clustering calculation by the HDBSCAN text clusterer to generate the hot event.
3. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 2, characterized in that: The self-attention mechanism dimension selector includes a self-attention mechanism, an output layer, and a dimension selection layer. The process of generating an optimal reduced dimension text vector with optimal reduced dimension by respectively passing the reduced dimension text vectors of multiple dimensions through the self-attention mechanism dimension selector includes: The reduced-dimensional text vector is context-aware through the self-attention mechanism to generate an embedding matrix; Classifying the embedding matrix through the output layer to determine the category probability of the reduced-dimensionality text; The category probabilities are subjected to similarity calculation through the dimension selection layer to determine the optimal dimensionality reduction text vector.
4. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 3, characterized in that: The dimension selection layer includes a channel similarity calculation module, a time window similarity calculation module, and a comprehensive selection module. The process of performing similarity calculation on the category probability through the dimension selection layer to determine the optimal dimensionality reduction text vector includes: The channel similarity calculation module is used to calculate the channel similarity of the publication channel of the category probability; The time window sequence of the class probability is passed through the time window similarity calculation module to generate time distribution similarity; The channel similarity and the time distribution similarity are calculated by the comprehensive selection module to determine the optimal dimension-reduced text vector.
5. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 4, characterized in that: The channel similarity calculation module is constructed based on the cosine similarity of the same posting channel and different posting channels; The time window similarity is constructed based on the cosine similarity of adjacent time window sequences.
6. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 2, characterized in that: The hot word mining model also includes a word segmenter and a keyword determiner based on TF-IDF; Generate multiple candidate keywords from the hot event through the word segmenter; The candidate keywords are used by the keyword determiner to generate keywords associated with the hot event.
7. The method for comprehensive processing of sports event content based on artificial intelligence according to claim 6, characterized in that: The process of generating an evaluation matrix of heat indexes through grade evaluation includes: Obtaining multiple heat indexes of heat events containing the keyword; The heat index is normalized and graded to generate the evaluation matrix; The popularity index includes the number of event videos, the number of event clicks, the number of event comments, the number of event attentions and the number of event reports.
8. The method for comprehensive processing of sports event content based on artificial intelligence according to any one of claims 1 to 7, characterized in that: The process of constructing a heat index judgment matrix by expert evaluation method based on multiple heat indexes corresponding to multiple heat events includes: Obtaining scores for the popularity index from a questionnaire published online; The popularity index judgment matrix is constructed according to the ratings.
9. The method for comprehensive processing of sports event content based on artificial intelligence according to any one of claims 1 to 7, characterized in that: The process of generating indicator correlation by using a correlation analysis model based on grey correlation analysis for the time series data of multiple heat indicators and calculating and adjusting the preheating event weights based on the indicator correlation to generate the time series heat event weights includes: Calculate the correlation coefficient between multiple time series data; Performing arithmetic averaging on the correlation coefficients to obtain the indicator correlation degree; The preheating event weight is multiplied by the indicator correlation to obtain the temporal heat event weight.
10. The method for comprehensive processing of sports event content based on artificial intelligence according to any one of claims 1 to 7, characterized in that: The process of determining the focus events of comprehensive sports events by using the evaluation matrix, time series hot event weights, and hot index weights through the hot event judgment model based on the fuzzy evaluation method includes: Calculate the event heat score by combining the evaluation matrix and the heat index weight; Calculate the index score and the temporal hot event weight to obtain a total hot score; A hot event whose total heat score exceeds a total score threshold and whose event heat score exceeds an event score threshold is determined as the comprehensive sports event of interest event.
Citation Information
Patent Citations
Science and technology project technology theme identification method based on BERTopic theme identification model
CN117725212A
Prediction method for network topic popular degree
CN106557552A
Online public opinion popularity value quantitative identification method based on grey correlation analysis
CN111414550A