AI-based large sports event comprehensive management method

By constructing a multimodal data fusion system, utilizing video, audio, crowd flow sensor, and text-based public opinion data, and combining transfer learning and spatiotemporal risk evolution prediction models, the problem of single data in large-scale sports events has been solved, enabling accurate prediction and graded early warning of future regional unit risks.

CN122089060APending Publication Date: 2026-05-26CHONGQING DONGYE SPORTS DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING DONGYE SPORTS DEVELOPMENT CO LTD
Filing Date
2026-01-30
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies for large-scale sporting events rely on single-mode data collection, failing to achieve multimodal data fusion and thus hindering effective risk prediction for future regional units.

Method used

By constructing a multimodal data fusion system, utilizing video, audio, pedestrian flow sensor, and text-based public opinion data, and combining transfer learning and spatiotemporal risk evolution prediction models, cross-modal alignment and adaptive multimodal fusion are achieved to predict risks in regional units.

Benefits of technology

It enables accurate and efficient prediction of future risk types and levels for various regional units of large-scale sports events, supporting tiered early warning and command and dispatch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089060A_ABST
    Figure CN122089060A_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of sports event risk prediction, and specifically relates to a comprehensive control method for large-scale sports events based on AI, including the following steps: S1: Define regional units and time window samples; S2: Construct a risk label system; S3: Collect and organize multimodal data; S4: Perform preprocessing and alignment; S5: Construct a transfer learning modal encoder; S6: Decomposed cross-modal alignment; S7: Adaptive multimodal fusion, fuse shared risk semantic representations and modality-specific semantic representations, construct a gated weighting module, and allocate dynamic weights to each modality in combination with quality indicators and risk correlations; S8: Establish a spatio-temporal risk evolution prediction model; S9: Risk output and closed-loop update, list and display the prediction results according to regional units and time windows, trigger early warnings according to risk levels and give disposal suggestions.The present invention realizes multimodal data fusion based on each regional unit, conducts spatio-temporal risk evolution prediction, facilitates quick positioning of risk locations and improves the efficiency of manual review.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of risk prediction technology for sports events, specifically to a comprehensive management and control method for large-scale sports events based on AI. Background Technology

[0002] Risk management for large-scale sporting events is an important and complex technology, involving personnel safety, event operation, public safety, and fire safety. Although some existing technologies are related to the safety management of large-scale sporting events, and some early warning and command platforms are publicly available, they often have shortcomings: the collected data is singular, multimodal data fusion is not achieved, and comprehensive predictions of the risks of each regional unit over a future period are not made. Summary of the Invention

[0003] The present invention aims to provide an AI-based comprehensive management and control method for large-scale sports events, which can achieve multimodal data fusion based on various regional units, and perform spatiotemporal risk evolution prediction with accurate and efficient prediction results.

[0004] The AI-based comprehensive management and control method for large-scale sports events includes the following steps:

[0005] S1: Define regional units and time window samples, divide regional units, and construct a venue topology map according to connectivity; discretize the event process with sampling step size Δτ, set observation window P and prediction window H, and form a regional time window sample index set Ω;

[0006] S2: Construct a risk labeling system, classify risks, and refine risk subcategories; define risks as the probability and level of risk for a regional unit within a future window, and complete mixed labeling of strong and weak labels by combining disposal records and verification results;

[0007] S3: Collect and organize multimodal data, acquire video, audio, time-series data from people flow sensors, and text-based public opinion data, covering the audience seating area, corridor, entrances and exits, turnstiles and escalators, and key areas respectively; uniformly bind regional units and timestamps to form a multimodal input set oriented towards regional time window samples;

[0008] S4: Perform preprocessing and alignment, unify the video frame rate and sample keyframes and complete the viewpoint to region mapping, denoise and segment the audio and align it with the time axis, clean and smooth the pedestrian flow data and resample it according to Δτ, deduplicate and merge the text, and extract risk-related words; aggregate multimodal segments according to the observation window and generate quality indicators of occlusion, noise, missing rate and confidence.

[0009] S5: Construct a transfer learning modality encoder, and establish video, audio, pedestrian flow time-series and text encoders to extract risk cues; initialize with a general pre-trained model and transfer-adapt it on venue data, and map it to the same-dimensional semantic vector space through a unified projection layer to provide a consistent interface for cross-modal alignment and fusion;

[0010] S6: Decompositional Cross-Modal Alignment: Each modal representation is decomposed into a shared risk semantic representation and a modality-specific representation; the shared layer applies consistency and distribution alignment constraints to suppress modality shift, and the specific layer introduces risk type prototype anchors to achieve comparable alignment; positive and negative sample pairs and modality mask constraints are used to support robust alignment under weak supervision and missing modalities.

[0011] S7: Adaptive multimodal fusion, which integrates shared risk semantic representation and modality-specific representation, constructs a gated weighting module, and assigns dynamic weights to each modality by combining quality indicators and risk correlation; outputs a fused risk representation vector, and retains modality weights and key contribution fragment indices for interpretation and verification;

[0012] S8: Establish a spatiotemporal risk evolution prediction model, with regional units as nodes, inheriting the venue topology adjacency matrix constructed in step S1 as the connection relationship between nodes, used to model the spatial diffusion and bottleneck convergence relationship, and TCN is used in the time dimension to characterize short-term outbreaks and long-term trends; set up a multi-step prediction head to output risk type, risk level and risk probability for multiple future time windows, forming a unified prediction interface.

[0013] S9: Risk output and closed-loop update. The prediction results are displayed in a list format by regional unit and time window. Early warnings are triggered according to risk level and disposal suggestions are given. The threshold is adaptively adjusted according to capacity, baseline and situation. The verification and disposal results are written back to the sample. Periodic retraining and threshold updates realize closed-loop optimization.

[0014] This invention organizes multimodal data through a unified index of regional units and prediction time windows, enabling cross-modal alignment, adaptive multimodal fusion, and spatiotemporal risk evolution prediction. It outputs risk probability and risk level for the risk type within the future prediction time window of each regional unit, and presents it in the form of a regional unit-level risk list, supporting hierarchical early warning triggering and command and dispatch. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the framework of the AI-based comprehensive management and control method for large-scale sports events of this invention. Detailed Implementation

[0016] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described below are only for explaining the present invention and do not limit the scope of protection of the present invention.

[0017] The present invention will be further described in detail below through preferred embodiments:

[0018] As attached Figure 1 As shown: This embodiment discloses an AI-based risk management method for large-scale sporting events, which predicts the risks of large-scale sporting events based on collected multimodal data. It includes the following steps:

[0019] S1: Define regional units and time window samples. Divide the region into units and construct a venue topology map based on connectivity; discretize the event process with a sampling step size Δτ, set the observation window P and the prediction window H, and form the regional time window sample index set Ω.

[0020] S2: Construct a risk labeling system. Classify risks and refine risk subcategories; define risks as the probability and level of risk for a regional unit within a future window, and complete mixed labeling of strong and weak labels by combining disposal records and verification results.

[0021] S3: Collect and organize multimodal data. Acquire video, audio, time-series data from pedestrian flow sensors, and text-based public opinion data, covering the audience seating area, corridors, entrances and exits, turnstiles and escalators, and key areas; uniformly bind regional units and timestamps to form a multimodal input set oriented towards regional time window samples.

[0022] S4: Perform preprocessing and alignment, unify the video frame rate and sample keyframes and complete the viewpoint to region mapping, denoise and segment the audio and align it with the time axis, clean and smooth the pedestrian flow data and resample it according to Δτ, deduplicate and merge the text and extract risk-related words; aggregate multimodal segments according to the observation window and generate quality indicators of occlusion, noise, missing rate and confidence.

[0023] S5: Construct transfer learning modal encoders. Video, audio, pedestrian flow time-series, and text encoders are established to extract risk cues; these are initialized with a general pre-trained model and transferred and adapted onto venue data, then mapped to a semantic vector space of the same dimension via a unified projection layer, providing a consistent interface for cross-modal alignment and fusion.

[0024] S6: Decompositional cross-modal alignment: Each modal representation is decomposed into a shared risk semantic representation and a modality-specific representation; the shared layer applies consistency and distribution alignment constraints to suppress modality shift, and the specific layer introduces risk type prototype anchors to achieve comparable alignment; positive and negative sample pairs and modality mask constraints are used to support robust alignment under weak supervision and missing modalities.

[0025] S7: Adaptive multimodal fusion, which integrates shared risk semantic representation and modality-specific representation, constructs a gated weighting module, and assigns dynamic weights to each modality by combining quality indicators and risk correlation; outputs a fused risk representation vector, and retains modality weights and key contribution fragment indices for interpretation and verification.

[0026] S8: Establish the spatiotemporal risk evolution prediction model GTCN-RiskNet, with regional units as nodes, inheriting the venue topology adjacency matrix constructed in step S1 as the connection relationship between nodes, which is used to model the spatial diffusion and bottleneck convergence relationship. The time dimension uses TCN to characterize short-term outbreaks and long-term trends; set up a multi-step prediction head to output risk type, risk level and risk probability for multiple future time windows, forming a unified prediction interface.

[0027] S9: Risk output and closed-loop update. The prediction results are displayed in a list format by regional unit and time window. Early warnings are triggered according to risk level and disposal suggestions are given. The threshold is adaptively adjusted according to capacity, baseline and situation. The verification and disposal results are written back to the sample. Periodic retraining and threshold updates realize closed-loop optimization.

[0028] Step S1 specifically includes the following steps:

[0029] S11: Divide the stadium space into a set of non-overlapping and fully covered regional units. Each regional unit includes the core area of ​​the football field, spectator seating areas, corridor segments, entrance / exit turnstile areas, escalator areas, and the restricted area boundary zone. The set of regional units is defined as follows:

[0030]

[0031] in, A set of regional units; For the first Each regional unit; This represents the total number of regional units.

[0032] To facilitate subsequent spatial propagation and capacity constraint modeling, each region unit is assigned a region attribute vector:

[0033]

[0034] in, This is a vector of region unit attributes; It is a regional capacity indicator used to characterize the population size or traffic capacity that a regional unit can support; The area type is coded to represent whether the area unit belongs to the core area of ​​the venue, the audience seating area, the corridor, the turnstile location, the escalator location, or the boundary zone of the restricted area; This is the regional location vector, used to represent the position of the regional unit in the venue's planar coordinate system.

[0035] S12: Based on the connectivity of venue passageways, the connection between entrances and exits and passageways, the connection between escalators going up and down, and the contact relationship between restricted areas and adjacent areas, construct a regional unit topology diagram as shown below:

[0036]

[0037] in, A regional topology map; Let it be the set of edges; edges Representing regional units With regional units There are reachable connections between people or events that affect each other.

[0038] Define the topological weight adjacency matrix As shown below, it is used to characterize connectivity strength and support subsequent spatiotemporal modeling.

[0039]

[0040] in, It is an adjacency matrix; To the regional unit To regional unit The connectivity weights; For regional units With regional units The connectivity strength index between passages is determined based on factors such as passage width, direction of passage, turnstile capacity, and escalator capacity. ;

[0041] Then, according to the uniform sampling step size For the entire event being discrete, a past observation window length is set. With future predicted length The regional time window sample is composed of regional units and observation time windows; only samples in which both observation and prediction fall within the monitoring interval are selected to form the regional time window sample index set. .

[0042] S2 specifically includes the following steps:

[0043] S21: Define the set of risk subclasses As shown below, the risk subclasses are limited to specific subclasses among population risk, public security risk, fire risk, and facility and equipment risk.

[0044]

[0045] in, A collection of risk subclasses; For congestion and stampedes; Running in panic; For conflict and fighting; To trespass into the restricted area; For smoke; For fire; The gate is malfunctioning. The escalator is malfunctioning.

[0046] S22: Based on historical handling records and verification results, calculate the occurrence and overall severity of events within the future prediction window to form a time-slot-level event count. Weighted severity and weighted cumulative severity within the future forecast window .

[0047] S23: Using regional time window samples Define risk subclasses based on units. The binary label indicating whether something will happen within the future prediction window is shown below: This binary label can be used to train the risk probability output.

[0048]

[0049] in, Risk subclass In regional units Time Index The corresponding risk occurrence label; Representing regional units In the future At least one risk subclass occurs within each time slot. event; Representing regional units In the future No risk subclass occurred within the time slot. event; For regional units In time index Risk subclasses occurring within the corresponding time slot Event count; This is an indicator function that takes the value 1 when the condition within the curly braces is true, and 0 otherwise. For the prediction window length, it represents the time index. The number of time slots to be covered in the future; For the time slot offset index within the prediction window, satisfying ; For regional cell indexes, For discrete-time indexing, Index for risk subclasses; (·) indicates that in The maximum value of the indicator function is taken within the range to implement the label definition of "whether the event has occurred at least once in the future window".

[0050] S24: To unify the severity of events within the future prediction window to discrete risk levels, let the total number of risk levels be... And for each risk subclass Set a monotonically increasing sequence of level thresholds. The risk level labels are defined as follows.

[0051]

[0052] in, This is a risk level label; the value range is... ; Risk subclass The Level threshold, satisfying

[0053]

[0054] S25: Samples for each regional time window The following set of monitoring labels is constructed to simultaneously characterize risk type, risk occurrence, risk level, occurrence area, and occurrence time window.

[0055]

[0056] in, The set of supervised labels for regional time window samples; Risk type label; Label the occurrence of risk; Risk level label; Label the area where the event occurred; This is a label for the time window in which the event occurs.

[0057] S3 specifically includes the following steps:

[0058] S31: Divide the venue into several regional units and construct a set of regional units:

[0059]

[0060] In the formula, For a set of regional units, For the first Each regional unit This represents the number of regional units.

[0061] S32: Establish a unified timestamp sequence to obtain the target sampling time.

[0062]

[0063] In the formula, For the first Each target sampling time, At the starting time, To standardize the sampling step size, This is the aligned sampling sequence number.

[0064] S33: Acquire video data and bind it to the region unit and timestamp.

[0065] Camera clusters were deployed in the audience seating area, the corridor area, the entrance / exit turnstile area, and the escalator area. Camera Output video frame sequence Construct the spatial mapping matrix from camera to region unit:

[0066]

[0067] And form regional unit-level video observation records:

[0068]

[0069] in, For a moment Regional Unit A collection of video observations; For camera At any moment The acquired video observation data is represented as image frames or video segments composed of consecutive frames; For the collection of cameras in the venue; Index the camera to satisfy ; A camera coverage relationship indicator used to characterize the camera. With regional units The correspondence between them; Indicates camera Coverage area unit , Indicates camera Uncovered area unit ; For the first Each regional unit; For region cell indexing; This is a discrete-time index.

[0070] S34: Acquire audio data and bind it to the region unit and timestamp.

[0071] A collection of microphones was deployed in key areas both inside and outside the venue. ,pickup Output audio waveform Spatial mapping matrix from pickup to regional unit. This forms regional unit-level audio observation records:

[0072]

[0073] In the formula, For regional units At any moment The audio observation set.

[0074] S35: Collect time-series data from the turnstile and escalator pedestrian flow sensors and bind it to the area unit and timestamp.

[0075] A set of pedestrian flow sensors is installed at the turnstiles and escalators. Each sensor The output count and speed-related timing are denoted as:

[0076]

[0077] In the formula, For a moment Inward throughput For a moment Outbound throughput For a moment Through rate. Through the mapping from sensor to regional unit. Bind it to the corresponding region unit .

[0078] S36: Collect text-based public opinion data and bind it to regional units and timestamps.

[0079] Collect text sets related to the event A single text record is denoted as:

[0080]

[0081] In the formula, For the publication timestamp, For text content, To assess source credibility or source weight, text records are grouped into regional units based on geographic tags, key location terms, and event terms. .

[0082] S37: Forming a multimodal input set for regional time window samples

[0083] Let the length of the observation time window be... Then The observation time window for the final moment is defined as:

[0084]

[0085] And construct a multimodal input set of regional time window samples:

[0086]

[0087] In the formula, For regional units During the observation time window A collection of video clips within. A collection of audio clips, A collection of time-series segments of human flow. A collection of text fragments These serve as multimodal input samples for subsequent models.

[0088] Step S4 specifically includes the following steps:

[0089] S41: Unified timestamps and resampling step size

[0090] Set a uniform sampling step size Align each modal input set to the target sampling time sequence. .

[0091] S42: Video Preprocessing and Spatial Mapping

[0092] S421: Frame Rate Unification and Keyframe Sampling

[0093] For the camera The video frame sequence at the target time Extract from the specified location to obtain the aligned frame:

[0094]

[0095] In the formula, For the aligned frame, This represents the moment when the distance is minimized.

[0096] S422: Spatial mapping from camera viewpoint to area unit

[0097] For each camera With regional units Constructing a region mask And obtain regional unit-level video frames:

[0098]

[0099] in, For regional units At any moment The region frame representation is used to characterize the fused video observation of the region unit at that moment; For the collection of cameras in the venue; Index the camera to satisfy ; For camera coverage relationship indication, Indicates camera Coverage area unit , Indicates camera Uncovered area unit ; For camera Corresponding regional unit The spatial mask matrix is ​​used to select regional units from the camera's viewpoint. The pixel or feature location; For camera At any moment The obtained preprocessed image frame representation; This is an element-wise multiplication operation; This indicates that the masked observations from all cameras are accumulated and fused.

[0100] S43: Audio Preprocessing and Timing Alignment

[0101] S431: Noise Reduction Processing

[0102] Taking short-time spectral representation as an example, the original spectrum... Noise reduction achieved :

[0103]

[0104] In the formula, For frequency index, For noise spectrum estimation, The noise reduction coefficient is... Used to ensure non-negativity.

[0105] S432: Segment and align with the target time.

[0106] The aligned audio segment is defined as:

[0107]

[0108] In the formula, For regional units exist Audio segments at the location.

[0109] S44: Data cleaning, smoothing, and resampling from pedestrian flow sensors

[0110] S441: Outlier Removal and Smoothing

[0111] For the original sequence Obtained by exponential smoothing :

[0112]

[0113] In the formula, For smoothing coefficients, These are the smoothed sequence values.

[0114] S442: Resample according to uniform sampling step size

[0115] When the target sampling time Falling on the original adjacent sampling time and When the alignment values ​​are between these points, linear interpolation is used to obtain the alignment values.

[0116]

[0117] In the formula, For the aligned observations, To and Adjacent original sampling times, These are the corresponding observations.

[0118] S45: Textual sentiment analysis deduplication and filtering, time merging, and risk-related keyword extraction.

[0119] S451: Deduplication Filtering and Time Merging

[0120] By regional unit With observation time window For indexing, merge the text collection. It also retains the latest or highest credibility record for texts from the same source and with high similarity.

[0121] S452: Extraction of Risk-Related Keywords

[0122] Building a risk vocabulary For terms The risk-related score within the observation time window is defined as follows:

[0123]

[0124] In the formula, For indicator functions, For word frequency, Inverse document frequency, This refers to the credibility or weight of the text source.

[0125] S46: After alignment, aggregate multimodal fragments according to the observation time window.

[0126] Each mode in The alignment results within the region are aggregated into multimodal fragments of the regional time window samples:

[0127]

[0128] In the formula, This is a sequence of video clips at the regional unit level. It is an audio segment sequence. For human flow time sequence segments, It is a collection of text fragments or a vectorized sequence of text.

[0129] S47: Generate a quality indicator for each mode

[0130] S471: Quality Indicator Corresponding to Video Occlusion Degree

[0131]

[0132] In the formula, For video quality indicators, The attenuation coefficient is... This refers to the occlusion ratio.

[0133] S472: Audio noise intensity corresponding to quality indicator

[0134]

[0135] In the formula, This is an indicator of audio quality. This is a noise intensity index.

[0136] S473: Quality Indicator Corresponding to Missing Rate of People Flow Data

[0137]

[0138] In the formula, As an indicator of the quality of human flow data, The missing rate.

[0139] S474: Quality Indicator for Text Credibility

[0140]

[0141] In the formula, This is a text quality indicator. To observe the number of text entries within the time window. To assess the credibility or weight of the text source. This is a single text entry.

[0142] This step outputs the aligned regional time window sample multimodal fragment set and the quality indicators of each modality, which are then received by step S5 as encoding input and quality weighting basis.

[0143] Step S5 specifically includes the following steps:

[0144] S51: Determine the input object and modality set. Based on the multimodal segments of the regional time window samples obtained in step S3, denot the video segment as... The audio clip is The time sequence of the flow of people is The text fragment is Simultaneously, based on the modal quality indicator obtained in step S4, the video quality indicator is denoted as... The audio quality indicator is The quality indicator of pedestrian flow is Text quality indicator is ;

[0145] S52: Establish a video encoder and extract high-level video representation;

[0146] Building a video encoder ,right Characterization is performed to obtain the high-level representation of the video:

[0147]

[0148] In the formula, For regional units In the time window The high-level representation vector of the video; These are the parameters for the video encoder; the video encoder is used to extract cues of crowd density changes, motion patterns, and abnormal behavior.

[0149] S53: Establish an audio encoder and extract high-level audio representation; construct the audio encoder. ,right Characterization is performed to obtain a high-level audio representation:

[0150]

[0151] In the formula, This is the high-level audio representation vector; These are the parameters for the audio encoder; the audio encoder is used to extract clues such as screams, sudden bursts of sound, and abnormal changes in overall sound pressure.

[0152] S54: Establish a pedestrian flow temporal encoder and extract the high-level representation of pedestrian flow; construct the pedestrian flow temporal encoder. ,right Characterization was performed to obtain the representation of the upper levels of pedestrian flow:

[0153]

[0154] In the formula, The vector representing the upper layer of human flow; These are the parameters of the pedestrian flow time sequence encoder; the pedestrian flow time sequence encoder is used to extract dynamic clues such as turnstile and escalator throughput, queue accumulation and congestion reduction;

[0155] S55: Establish a text encoder and extract high-level text representations; construct the text encoder. ,right Characterization is performed to obtain a high-level representation of the text;

[0156]

[0157] In the formula, A high-level representation vector of the text; These are the parameters for the text encoder; the text encoder is used to extract public sentiment polarity, topic clustering, and risk trigger signals.

[0158] S56: Output semantic vectors of the same dimension through a unified projection layer;

[0159] To ensure a consistent representation interface for the outputs of each modality, a projection layer is constructed for each modality to obtain semantic vectors of the same dimension:

[0160]

[0161] In the formula, For modal identifiers, the set of modal identifiers is: ; For modality semantic vector; For modality Projection matrix; For modality Projection bias vector; It is a 2-norm; the dimension of the semantic vectors of each modality is unified as . ;

[0162] S57: Initialize with a general pre-trained model and perform transfer learning and domain adaptation;

[0163] Each modal encoder is initialized using common pre-trained parameters, denoted as . After transfer learning, the parameters are Optimize and update the venue data:

[0164]

[0165] In the formula, For modality Learning rate; For about The gradient operator; For modality Transfer learning objective function;

[0166] To enhance the controllability of domain adaptation, a risk-label-based auxiliary supervision head is introduced. Output the probability vector of risk occurrence:

[0167]

[0168] in, For modality For regional units In the time window The probability prediction vector of risk occurrence; For modality The auxiliary supervisory head function is used to map the modal representation to the risk occurrence probability output; For regional units In the time window modality The fusion representation serves as the input features for the auxiliary supervision head; For modality Parameters of the auxiliary monitoring head; Modal identifier;

[0169] The auxiliary supervision loss adopts the class-by-class binary cross-entropy form, as shown below:

[0170]

[0171] in, For modality The loss of auxiliary supervision; For modal identification, used to refer to one of the following modalities: video modality, audio modality, crowd flow time sequence modality, or text modality; This is a sample index pair consisting of a regional cell index and a time window index; A collection of risk subclasses; For risk subclass indexes, satisfying ; For regional units In the time window Regarding risk subcategories The risk occurrence label takes a value of 0 or 1; Modal identifier In regional units With time window Addressing risk subclasses The predicted probability of occurrence; symbol Indicates the predicted quantity; For items corresponding to the "risk not occurred" label; This is the logarithm of the probability that the risk has not occurred;

[0172] To suppress overfitting caused by venue noise and occlusion, a feature regularization term weighted by the quality indicator is introduced.

[0173]

[0174] In the formula, Modal identifier Quality indicators; the lower the quality, the greater the penalty.

[0175] Final mode The transfer learning objective function is defined as

[0176]

[0177] In the formula, The regularization weight coefficients, through the above transfer learning and domain adaptation, enable each modal encoder to output semantic vectors of the same dimension in the venue scenario, providing a consistent representation interface for cross-modal alignment and subsequent fusion in step S6.

[0178] S6 specifically includes the following steps:

[0179] S61: Introduce a modality mask and determine the set of available modes;

[0180] For scenarios with missing modalities, a modal mask is defined for each region's time window samples:

[0181]

[0182] In the formula, Indicates sample modality Available Representing modes Missing or unavailable; the modal mask is used to limit the set of valid modalities that participate in the computation during subsequent cross-modal alignment and aggregation.

[0183] S62: Perform semantic decomposition on the modal semantic vector;

[0184] For each modal semantic vector Semantic decomposition yields shared risk semantic representations and modality-specific semantic representations:

[0185]

[0186] In the formula, For modality Shared risk semantic representation; For modality Modality-specific semantic representation; To share the decomposition matrix; It is a unique decomposition matrix; the shared risk semantic representation dimension is The dimension of modality-specific semantic representation is .

[0187] To reduce the mutual interference between shared risk semantic representations and modality-specific semantic representations, decorrelation constraints are introduced:

[0188]

[0189] This constraint is used to ensure that the shared risk semantic representation and the modality-specific semantic representation are independent of each other in the inner product sense.

[0190] in, The decorrelation constraint loss is used to measure the correlation between the shared risk semantic representation and the modality-specific semantic representation; This is a sample index pair consisting of a regional cell index and a time window index; For modality In regional units With time window Shared risk semantic representation vector at the location; For modality In regional units With time window The modality-specific semantic representation vector at that location; For transpose operator; The inner product term of the shared risk semantic representation vector and the modality-specific semantic representation vector is used to characterize the correlation between the two in the sense of the inner product; by minimizing prompt and They are independent of each other in the sense of inner product, thus achieving the decorrelation between shared risk semantic representation and modality-specific semantic representation.

[0191] S63: Impose explicit consistency constraints at the shared risk semantic level;

[0192] Based on the modal quality indication obtained in step S4, denoted as Its value range is A higher value indicates higher modal quality. Quality-aware weights are constructed using modal masks:

[0193]

[0194] In the formula, For the sample modality Aggregate weights; For modal masks; This is a modal quality indicator.

[0195] A quality-weighted aggregation of the multimodal shared risk semantic representations of samples within the same time window region yields a unified shared risk semantic representation.

[0196]

[0197] In the formula, For the sample A unified shared risk semantic representation is used; the denominator is used to normalize the effective weights participating in the aggregation. Furthermore, a consistency constraint is introduced at the shared risk semantic level to ensure that the multimodal shared risk semantic representation of the same sample converges within the semantic space.

[0198]

[0199] In the formula, Represents any pairwise combination within the modality set; It is a 2-norm; this consistency constraint is used to enhance the consistent expression of cross-modal shared risk semantics.

[0200] S64: Introduce a prototype anchor point mechanism at the modality-specific level to achieve relative alignment;

[0201] For each risk type Learn the corresponding risk semantic prototype vector And use it as an alignment anchor point for modality-specific semantics; for samples The set of positive risk types is defined as:

[0202] In the formula, A set of risk types; For the sample The corresponding set of positive risk types; Risk type The risk of occurrence is indicated by the label.

[0203] Construct a prototype alignment loss between modality-specific representations and risk semantic prototypes:

[0204]

[0205] in, The prototype alignment loss is used to constrain the modality-specific representations of different modalities to remain similar to the prototypes of the same risk subclass, thereby enhancing cross-modal comparability. For modality In the sample The validity weight or mask coefficient at the point is used to characterize whether the modality observation is available and its contribution to training, and its value range is a non-negative real number; For modality In the sample The modality-specific representation vector at that location; Risk subclass The prototype vector is used to characterize the category center of the risk subclass in the unified representation space; This is a similarity function used to measure the degree of similarity between vectors; This is a temperature coefficient used to adjust the smoothness of the similarity distribution.

[0206] S65: Define the overall objective function for cross-modal alignment

[0207] By integrating shared consistency constraints, prototype anchor constraints, and semantic decoupling constraints, a cross-modal alignment overall objective function is constructed:

[0208]

[0209] In the formula, , , These are weighting coefficients used to balance the contributions of different constraint terms during training.

[0210] By minimizing This enables stable alignment of multimodal representations within a unified risk semantic space, and provides consistent and comparable input representations for subsequent adaptive multimodal fusion and risk prediction.

[0211] This step outputs a modality mask, a modality-shared risk semantic representation, a modality-specific representation, and a risk semantic prototype vector, which are then used by step S7 as fusion input and as a basis for risk correlation calculation.

[0212] S7 specifically includes the following steps:

[0213] S71: Determine the fusion input and construct fusion candidate representations;

[0214] Using the modality-shared risk semantic representation and the modality-specific semantic representation output from step S6 as fusion inputs, the modality set is defined as follows. For any mode Constructing fusion candidate representation vectors:

[0215]

[0216] In the formula, For modality Shared risk semantic representation; For modality Modality-specific semantic representation; For modality Fuse candidate representation vectors; For region cell indexing; Index for regional time windows.

[0217] S72: Calculate the risk relevance score;

[0218] To measure the correlation between modal evidence and risk type, the risk semantic prototype vector from step S6 is reused. For any mode Define risk relevance score:

[0219]

[0220] In the formula, A set of risk types; For similarity functions; For modality Risk relevance score;

[0221] S73: Generate gating weights using modal quality indicators and risk correlation scores;

[0222] Based on the modal quality indication quantity from step S4, denoted as... The modal mask in step S6 is denoted as Construct a gating fusion module to control modalities. Calculate the gating score:

[0223]

[0224] In the formula, This is the gate vector; This refers to the quality weighting coefficient. This refers to the relevance weighting coefficient. For gated bias; For modality Gating scoring;

[0225] The dynamic fusion weights are obtained by normalizing all available modes:

[0226]

[0227] In the formula, For modality In the sample The fusion weight, .

[0228] S74: Form a fusion risk representation vector while retaining modal weights;

[0229] The fusion risk representation vector of the regional time window samples is obtained by weighted summation based on dynamic fusion weights.

[0230]

[0231] In the formula, The fused risk representation vector of the regional time window samples is used for the spatiotemporal risk evolution modeling in step S8. As the explanatory power follows Output them together for early warning interpretation and manual review.

[0232] S75: Extract the index of key contribution segments for interpretation and verification;

[0233] To locate key pieces of evidence, for any modality The segment-level feature sequence within the time window is denoted as Constructing fragment contribution score:

[0234]

[0235] In the formula, For modality Segment rating vector; For the first The contribution weight of each segment; This represents the number of segments within the observation window.

[0236] Select the top contributor with the largest contribution weight Each fragment index serves as the set of key contribution fragment indices:

[0237]

[0238] In the formula, For modality A set of indexes of key contribution fragments; To get the maximum Operators for indexing elements.

[0239] This step outputs the fusion risk representation vector and simultaneously outputs the modality fusion weights and key contribution fragment index set, which are used in step S8 to construct the node-level spatiotemporal input sequence.

[0240] S8 specifically includes the following steps:

[0241] S81: Constructing a node-level spatiotemporal input sequence

[0242] For any prediction start time window Take the past The fused risk representation vector of each regional time window is used as the spatiotemporal input, and the node-level input sequence is defined as follows:

[0243]

[0244] In the formula, The input sequence tensor; For time windows The node feature matrix; This represents the number of regional units; For regional units In the time window The fusion risk representation vector; This represents the length of the observation window. These are the node feature terms arranged in chronological order in the input sequence;

[0245] S82: Constructing the spatial adjacency and normalization matrix of a simplified spatiotemporal graph network

[0246] The venue topology adjacency matrix constructed in step S1 And introduce self-loop matrices This yields the adjacency matrix with added self-loops:

[0247]

[0248] In the formula, It is a self-loop matrix, and the self-loop matrix is ​​the identity matrix; To add a self-loop adjacency matrix;

[0249] Calculate the degree matrix And construct the normalized adjacency matrix:

[0250]

[0251] In the formula, It is a degree matrix; These are the diagonal elements of the degree matrix; The normalized adjacency matrix is ​​used to characterize the connectivity and diffusion relationship between the audience seating area and the corridor, the convergence and bottleneck relationship between the turnstiles and escalators, and the intrusion sensitivity relationship of the restricted area boundary zone.

[0252] S83: The temporal convolution module extracts the difference between short-term bursts and long-term trends; a one-dimensional temporal convolution module is constructed for... Perform convolution along the time dimension to obtain the temporal convolution output:

[0253]

[0254] In the formula, Given an input sequence tensor, It is a one-dimensional temporal convolution operator; These are the temporal convolution parameters; The output tensor of the time convolution;

[0255] S84: Spatial graph convolution module characterizes connectivity diffusion and bottleneck propagation;

[0256] Perform graph convolution propagation on the temporal convolution output at each time step to obtain spatial augmentation features:

[0257]

[0258] In the formula, For time step The node feature matrix; These are the spatial convolution parameters; It is a non-linear activation function; For time step Spatial enhanced feature matrix;

[0259] S85: Construct a simplified spatiotemporal modeling network and output node-level spatiotemporal hidden states;

[0260] A simplified spatiotemporal modeling network is formed by using a cascaded structure of temporal convolution and spatial graph convolution to obtain the prediction starting time window. Node-level spacetime hidden states:

[0261]

[0262] In the formula, For time-dimensional aggregation operators; This is the aggregated spatiotemporal representation matrix; For regional units In the predicted starting time window The spatiotemporal hidden state vector;

[0263] S86: Set the risk prediction header and complete multi-step prediction output:

[0264] Lightweight forecasting heads are set up for both the probability of risk occurrence and the risk level. For the future... Each prediction window outputs the probability of risk occurrence:

[0265]

[0266] In the formula, This represents the probability vector of multi-step risk occurrence. and These are the parameters for the probability prediction head. For element-wise Sigmoid function; This refers to the number of risk types.

[0267] For the future Each prediction window outputs a probability distribution of risk levels:

[0268]

[0269] In the formula, This represents a multi-step risk level probability distribution. and For the level prediction head parameters; This is the normalization function along the rank dimension; This represents the total number of preset risk levels.

[0270] Will and The data is rearranged according to the risk type and prediction step dimensions to form structured output fields. These fields include regional units, prediction time windows, risk types, risk probabilities, and risk levels, which are used to connect to the subsequent unified prediction interface and early warning linkage module.

[0271] S9: Establish a simplified risk output and closed-loop update mechanism

[0272] S91: Structured Risk Output and List-Based Presentation. The prediction results from step S8 are organized into structured output records. Output fields include regional units, prediction time windows, risk types, risk probabilities, and risk levels. Outputs within the same prediction time window are presented in a list format by regional units, forming a regional unit-level risk list for simultaneous display on the command center screen, mobile devices, and workstations. The risk list also retains the modal weights and key contribution fragment indexes from the output of step S7, facilitating rapid on-site identification of evidence sources and manual verification.

[0273] S92: Tiered Early Warning Triggering and Control Action Generation. To achieve tiered early warning triggering, early warning thresholds are set for each regional unit and risk type. These thresholds are adaptively adjusted based on regional capacity, historical baseline, and real-time situational characteristics to reduce systemic bias in false alarms and missed alarms. Specifically, the thresholds are initially set to the historical baseline and are corrected online based on changes in capacity utilization and real-time situational characteristics. Furthermore, they are periodically fine-tuned based on statistical results of early warning hit rate, false alarm rate, and missed alarm rate. Early warning trigger determination uses a comparison between risk probability and the threshold, defining an early warning trigger indicator.

[0274]

[0275] In the formula, This is an indicator function that takes the value 1 when the condition inside the indicator function is true, and takes the value 0 otherwise. This indicates that an alert has been triggered. This indicates that it will not be triggered. The risk prediction model outputs the regional unit where the risk type occurs at the prediction time window. The predicted probability; For regional units In terms of risk type The corresponding early warning threshold is adaptively adjusted based on regional capacity, historical baseline, and real-time situation characteristics;

[0276] When an alert is triggered, the system generates control action suggestions based on the risk type and risk level, and pushes them out in order of priority. The control action suggestions are generated using a combination of a rule base and resource constraints. The rule base uses risk type as the primary key and risk level as the branch, providing standardized, executable handling strategies. Resource constraints include the availability of resources such as security personnel, fire fighting personnel, medical personnel, broadcast and screen control permissions, and turnstile and escalator control permissions. Control action suggestions must include at least the following fields: handling area, handling time window, risk type, risk level, set of suggested actions, set of required resources, estimated handling duration, and verification evidence index.

[0277] When crowd risks are triggered, recommended actions include diverting and guiding traffic, one-way passage, temporarily closing off access points, limiting or suspending turnstile access, reducing or disabling escalators, increasing the number of guides and setting up isolation zones, and providing broadcast and on-screen guidance. When security risks are triggered, recommended actions include providing nearby security reinforcements, monitoring key groups, isolating conflict areas, temporarily closing off restricted areas, and using video surveillance and on-site announcements. When fire risks are triggered, recommended actions include smoke verification and location, activating emergency broadcasts, providing evacuation route guidance, dispatching fire-fighting equipment and personnel, and temporarily isolating suspicious areas. When facility and equipment risks are triggered, recommended actions include downgrading turnstiles or escalators, switching to backup access routes, dispatching maintenance personnel, isolating the area around the malfunction point and indicating alternative routes.

[0278] To ensure the feasibility of actions, the system performs conflict checks and priority decisions on suggested actions. Conflict checks include conflicts between passageway closures and evacuation routes, turnstile flow control and entry tasks, and escalator shutdowns and vertical evacuation capacity. Priority decisions are made primarily based on risk level, followed by risk probability, with area capacity and crowd density as correction factors. The final coordinated execution sequence is then output and sent to the venue's integrated management system interface.

[0279] S93: Handling feedback records and closed-loop update records include early warning triggers, manual verification results, and handling results. Manual verification results include a verification risk occurrence label and a verification risk level label, and record the source and credibility level of the verification data. Verification results are written back as training feedback samples and used for periodic retraining and threshold updates, achieving closed-loop iterative optimization of the model and rules. Closed-loop updates include two types of update objects.

[0280] The first category involves threshold and rule updates. The system statistically analyzes the early warning hit rate, false alarm rate, and false negative rate under different regional units and risk types, and combines this with the handling cost and scope of impact to update the threshold and rule settings. Periodic fine-tuning is performed, and the priority of action sets in the rule base is adjusted to make the handling actions under the same risk level more consistent with the actual execution effect of the venue.

[0281] The second category is model updates. Verified samples are added to the feedback sample set, and retraining is performed at fixed intervals or based on event-triggered conditions. This updates the parameters of the multimodal encoder, fusion module, and spatiotemporal model, allowing the model to gradually adapt to the risk evolution patterns under different event types, population structures, and venue configurations. After retraining, key indicators are validated through regression. Once validation is successful, the model is released to the online inference platform, achieving co-evolution of the model and rules.

[0282] This invention outputs retained modal fusion weights and key contribution fragment indexes, facilitating rapid risk identification and improving the efficiency of manual review; the warning threshold is adjusted according to regional capacity, historical baseline, and real-time situational characteristics to reduce false alarms and missed alarms; when a warning is triggered, control action suggestions matching the risk type are generated, covering measures such as diversion guidance, gate flow restriction, escalator downgrading, security reinforcement, emergency broadcasting, and temporary isolation, thereby improving the targeting and efficiency of the response, reducing the probability of risk spread, and enhancing the overall management and control effect of the venue.

[0283] The preferred embodiments of this application have been described in detail above with reference to the accompanying drawings. Typical known structures and common knowledge techniques in the preferred embodiments have not been described in detail here. Those skilled in the art can improve and implement the technical solutions of this invention based on the guidance provided in these embodiments and their own capabilities. Some typical known structures, known methods or common knowledge techniques should not be obstacles for those skilled in the art to implement this application.

[0284] The scope of protection claimed in this application shall be determined by the contents of its claims, and the contents described in the invention description, specific embodiments and drawings shall be used to interpret the claims.

[0285] Within the scope of the technical concept of this application, several modifications can be made to the specific implementation of this application, and these modified implementations should also be considered within the protection scope of this application.

Claims

1. An AI-based comprehensive management and control method for large-scale sports events, characterized in that, Includes the following steps: S1: Define regional units and time window samples, divide regional units, and construct a venue topology map according to connectivity; The event process is discrete using a sampling step size Δτ, and an observation window P and a prediction window H are set to form a regional time window sample index set Ω. S2: Construct a risk labeling system, classify risks, and refine risk subcategories; define risks as the probability and level of risk for a regional unit within a future window, and complete mixed labeling of strong and weak labels by combining disposal records and verification results; S3: Collect and organize multimodal data, acquire video, audio, time-series data from people flow sensors, and text-based public opinion data, covering the audience seating area, corridor, entrances and exits, turnstiles and escalators, and key areas respectively; uniformly bind regional units and timestamps to form a multimodal input set oriented towards regional time window samples; S4: Perform preprocessing and alignment, unify the video frame rate and sample keyframes and complete the viewpoint to region mapping, denoise and segment the audio and align it with the timeline, clean and smooth the pedestrian flow data and resample it according to Δτ, deduplicate and merge the text, and extract risk-related words. Multimodal fragments are aggregated according to the observation window, and quality indicators of occlusion, noise, missing rate and confidence are generated; S5: Construct a transfer learning modality encoder, and establish video, audio, pedestrian flow time-series and text encoders to extract risk cues; initialize with a general pre-trained model and transfer-adapt it on venue data, and map it to the same-dimensional semantic vector space through a unified projection layer to provide a consistent interface for cross-modal alignment and fusion; S6: Decompositional cross-modal alignment: Each modal representation is decomposed into a shared risk semantic representation and a modality-specific semantic representation; the shared layer applies consistency and distribution alignment constraints to suppress modality shift, and the specific layer introduces risk type prototype anchors to achieve comparable alignment; positive and negative sample pairs and modality mask constraints are used to support robust alignment under weak supervision and missing modalities. S7: Adaptive multimodal fusion, which integrates shared risk semantic representation and modality-specific semantic representation, constructs a gated weighting module, and assigns dynamic weights to each modality by combining quality indicators and risk correlation; Output the fused risk representation vector, and retain the modal weights and key contribution fragment indices for interpretation and verification; S8: Establish a spatiotemporal risk evolution prediction model, with regional units as nodes, inheriting the venue topology adjacency matrix constructed in step S1 as the connection relationship between nodes, used to model the spatial diffusion and bottleneck convergence relationship, and TCN is used in the time dimension to characterize short-term outbreaks and long-term trends; set up a multi-step prediction head to output risk type, risk level and risk probability for multiple future time windows, forming a unified prediction interface. S9: Risk output and closed-loop update, displaying forecast results in a list format by regional unit and time window, triggering early warnings by risk level and providing handling suggestions; thresholds are adaptively adjusted according to capacity, baseline and situation. The verification and processing results are written back to the sample, and periodic retraining and threshold updates achieve closed-loop optimization. 2.The AI-based comprehensive management method for large-scale sports events according to claim 1, wherein, Step S1 specifically includes the following steps: S11: Divide the venue space into a set of non-overlapping and fully covered regional units, as shown in the following formula: , wherein, is a set of region units; is a first region unit; is a total number of region units; Assign a region attribute vector to each region unit: , wherein, is a region unit attribute vector; is a region capacity index, used to represent the crowd size or traffic capacity that the region unit can carry; is a region type code, used to represent that the region unit belongs to a core area of a venue, a spectator stand, a ring corridor, a gate point, an escalator point, or a forbidden zone boundary; is a region position vector, used to represent the position of the region unit in a venue plane coordinate system; S12: Based on the connectivity of venue passageways, the connection between entrances and exits and passageways, the connection between escalators going up and down, and the contact relationship between restricted areas and adjacent areas, construct a regional unit topology diagram as shown below: , wherein, is a region topology graph; is a set of edges; edge represents a region unit has a reachability connectivity relationship with a region unit between the region units defining a topological weight adjacency matrix as follows: , wherein, is an adjacency matrix; is a connectivity weight from region cell to region cell is a connectivity strength indicator between region cell and region cell .​ 3.The AI-based integrated management method for large-scale sports events according to claim 1, wherein, S2 specifically includes the following steps: S21: defining a set of risk sub-classes As follows: , wherein, is a risk subcategory collection; is a crowd and trampling; is a panic running; is a conflict fighting; is a forbidden zone intrusion; is a smoke; is a fire; is a gate failure; is an escalator failure; S22: According to the historical treatment records and the verification results, the event occurrence and the comprehensive severity in a future prediction window are counted to form time slot level event count , weighted severity , and weighted severity accumulation in the future prediction window ; S23: Define risk sub-class in units of regional time window samples The binary label of whether an event occurs within a future prediction window is as follows: , in, Risk subclass In regional units Time Index The corresponding risk occurrence label; For regional units In time index Risk subclasses occurring within the corresponding time slot Event count; For indicator functions, For the prediction window length, it represents the time index. The number of time slots to be covered in the future; For the time slot offset index within the prediction window, For regional cell indexes, For discrete-time indexing, Index for risk subclasses; S24: To unify the severities of events within the future prediction window to discrete risk levels, let the total number of risk levels be and set a monotonically increasing sequence of level thresholds for each risk subcategory and define the risk level labels as shown below: , wherein, is a risk level label; is a risk subcategory of a first level threshold; S25: For each region's time window samples, construct the following set of supervision labels: , wherein, is a set of supervised labels for regional time window samples; is a risk type label; is a risk occurrence label; is a risk level label; is an occurrence region label; is an occurrence time window label. 4.The AI-based comprehensive management method for large-scale sports events according to claim 1, wherein, S3 specifically includes the following steps: S31: Divide the venue into several regional units and construct a set of regional units: , In the formula, For a set of regional units, For the first Each regional unit This represents the number of regional units; S32: Establish a unified timestamp sequence to obtain the target sampling time: , In the formula, For the first Each target sampling time, At the starting time, To standardize the sampling step size, The aligned sample number; S33: Collect video data and bind it to the area unit and timestamp; Camera Output video frame sequence Construct the spatial mapping matrix from camera to region unit: , And form regional unit-level video observation records: , in, For a moment Regional Unit A collection of video observations; For camera At any moment The acquired video observation data is represented as image frames or video segments composed of consecutive frames; For the collection of cameras in the venue; Index the camera to satisfy ; A camera coverage relationship indicator used to characterize the camera. With regional units The correspondence between them is covered. For the first Each regional unit For regional cell indexes, For discrete-time indexing; S34: Acquire audio data and bind it to the region unit and timestamp; A collection of microphones was deployed in key areas both inside and outside the venue. ,pickup Output audio waveform The regional unit-level audio observation records are shown below: , In the formula, For regional units At any moment The audio observation set, This is the spatial mapping matrix from the pickup to the region unit; S35: Collect time-series data from the turnstile and escalator pedestrian flow sensors and bind it to the area unit and timestamp; A set of pedestrian flow sensors is installed at the turnstiles and escalators. Each sensor The output count and speed-related timing are denoted as: , In the formula, For a moment Inward throughput For a moment Outbound throughput For a moment Through rate, through the mapping from sensor to regional unit. Bind it to the corresponding region unit ; S36: Collect text-based public opinion data and bind it to regional units and timestamps; Collect text sets related to the event A single text record is denoted as: , In the formula, For the publication timestamp, For text content, To assess source credibility or source weight, text records are grouped into regional units based on geographic tags, key location terms, and event terms. ; S37: Form a multimodal input set for regional time window samples; Let the length of the observation time window be... Then The observation time window for the final moment is defined as: , And construct a multimodal input set of regional time window samples: , In the formula, For regional units During the observation time window A collection of video clips within. A collection of audio clips, A collection of time-series segments of human flow. A collection of text fragments These serve as multimodal input samples for subsequent models.

5. The AI-based comprehensive management and control method for large-scale sports events according to claim 1, characterized in that, S4 specifically includes the following steps: S41: Unified timestamp and resampling step size; S42: Video preprocessing and spatial mapping; S43: Audio preprocessing and time alignment; S44: Cleaning, smoothing and resampling of pedestrian flow sensor data; S45: Textual sentiment analysis deduplication and filtering, time merging, and extraction of risk-related keywords; S46: After alignment, aggregate multimodal fragments according to the observation time window; S47: Generate a quality indicator for each mode; Specifically, S47 includes: S471: The quality indicators corresponding to the degree of video occlusion are shown below: , In the formula, For video quality indicators, The attenuation coefficient is... For the occlusion ratio; S472: Audio noise intensity corresponding to quality indication: , In the formula, This is an indicator of audio quality. Noise intensity index; S473: Quality indicator corresponding to missing population data: , In the formula, As an indicator of the quality of human flow data, The missing rate; S474: Quality Indicator for Text Credibility: , In the formula, This is a text quality indicator. To observe the number of text entries within the time window. To assess the credibility or weight of the text source. This is a single text entry.

6. The AI-based comprehensive management and control method for large-scale sports events according to claim 1, characterized in that: S5 specifically includes the following steps: S51: Determine the input object and modality set. Based on the multimodal segments of the regional time window samples obtained in step S3, denot the video segment as... The audio clip is The time sequence of the flow of people is The text fragment is Simultaneously, based on the modal quality indicator obtained in step S4, the video quality indicator is denoted as... The audio quality indicator is The quality indicator of pedestrian flow is Text quality indicator is ; S52: Establish a video encoder and extract high-level video representation; Building a video encoder ,right Characterization is performed to obtain the high-level representation of the video: , In the formula, For regional units In the time window The high-level representation vector of the video; These are the parameters for the video encoder; the video encoder is used to extract cues of crowd density changes, motion patterns, and abnormal behavior. S53: Establish an audio encoder and extract high-level audio representation; Building an audio encoder ,right Characterization is performed to obtain a high-level audio representation: , In the formula, This is the high-level audio representation vector; These are the parameters for the audio encoder; the audio encoder is used to extract clues such as screams, sudden bursts of sound, and abnormal changes in overall sound pressure. S54: Establish a pedestrian flow temporal encoder and extract the high-level representation of pedestrian flow; Constructing a pedestrian flow time sequence encoder ,right Characterization was performed to obtain the representation of the upper levels of the pedestrian flow: , In the formula, The vector representing the upper layer of human flow; These are the parameters of the pedestrian flow time sequence encoder; the pedestrian flow time sequence encoder is used to extract dynamic clues such as turnstile and escalator throughput, queue accumulation and congestion reduction; S55: Establish a text encoder and extract high-level text representations; Building a text encoder ,right Characterization is performed to obtain a high-level representation of the text; , In the formula, A high-level representation vector of the text; These are the parameters for the text encoder; the text encoder is used to extract public sentiment polarity, topic clustering, and risk trigger signals. S56: Output semantic vectors of the same dimension through a unified projection layer; To ensure a consistent representation interface for the outputs of each modality, a projection layer is constructed for each modality to obtain semantic vectors of the same dimension: , In the formula, For modal identifiers, the set of modal identifiers is: ; For modality semantic vector; For modality Projection matrix; For modality Projection bias vector; It is a 2-norm; the dimension of the semantic vectors of each modality is unified as . ; S57: Initialize with a general pre-trained model and perform transfer learning and domain adaptation; Each modal encoder is initialized using common pre-trained parameters, denoted as . After transfer learning, the parameters are Optimize and update the venue data: , In the formula, For modality Learning rate; For about The gradient operator; For modality Transfer learning objective function; To enhance the controllability of domain adaptation, a risk-label-based auxiliary supervision head is introduced. Output the probability vector of risk occurrence: , in, For modality For regional units In the time window The probability prediction vector of risk occurrence; For modality The auxiliary supervisory head function is used to map the modal representation to the risk occurrence probability output; For regional units In the time window modality The fusion representation serves as the input features for the auxiliary supervision head; For modality Parameters of the auxiliary monitoring head; Modal identifier; The auxiliary supervision loss adopts the class-by-class binary cross-entropy form, as shown below: , in, For modality The loss of auxiliary supervision; For modal identification, used to refer to one of the following modalities: video modality, audio modality, crowd flow time sequence modality, or text modality; This is a sample index pair consisting of a regional cell index and a time window index; A collection of risk subclasses; For risk subclass indexes, satisfying ; For regional units In the time window Regarding risk subcategories The risk occurrence label takes a value of 0 or 1; Modal identifier In regional units With time window Risk subclass The predicted probability of occurrence; symbol Indicates the predicted quantity; For items corresponding to the "risk not occurred" label; This is the logarithm of the probability that the risk has not occurred; To suppress overfitting caused by venue noise and occlusion, a feature regularization term weighted by the quality indicator is introduced. , In the formula, Modal identifier Quality indicators; the lower the quality, the greater the penalty. Final mode The transfer learning objective function is defined as , In the formula, The regularization weight coefficients, through the above transfer learning and domain adaptation, enable each modal encoder to output semantic vectors of the same dimension in the venue scenario, providing a consistent representation interface for cross-modal alignment and subsequent fusion in step S6.

7. The AI-based comprehensive management and control method for large-scale sports events according to claim 1, characterized in that, S6 specifically includes the following steps: S61: Introduce a modality mask and determine the set of available modes; For scenarios with missing modalities, a modal mask is defined for each region's time window samples: , In the formula, Indicates sample modality Available Representing modes Missing or unavailable; S62: Perform semantic decomposition on the modal semantic vector; For each modal semantic vector Semantic decomposition yields shared risk semantic representations and modality-specific semantic representations: , In the formula, For modality Shared risk semantic representation; For modality Modality-specific semantic representation; To share the decomposition matrix; It is a unique decomposition matrix; To reduce the mutual interference between shared risk semantic representations and modality-specific semantic representations, decorrelation constraints are introduced: , in, The decorrelation constraint loss is used to measure the correlation between the shared risk semantic representation and the modality-specific semantic representation; This is a sample index pair consisting of a regional cell index and a time window index; For modality In regional units With time window Shared risk semantic representation vector at the location; For modality In regional units With time window The modality-specific semantic representation vector at that location; For transpose operator; The inner product term of the shared risk semantic representation vector and the modality-specific semantic representation vector; S63: Impose explicit consistency constraints at the shared risk semantic level; Based on the modal quality indication obtained in step S4, denoted as Its value range is A higher value indicates higher modal quality. Constructing quality-aware weights by combining modal masks: , In the formula, For the sample modality Aggregate weights; For modal masks; This is a modal quality indicator. A quality-weighted aggregation of the multimodal shared risk semantic representations of samples within the same time window region yields a unified shared risk semantic representation. , In the formula, For the sample A unified shared risk semantic representation is used; the denominator is used to normalize the effective weights participating in the aggregation, and a consistency constraint is further introduced at the shared risk semantic level to make the multimodal shared risk semantic representation of the same sample converge in the semantic space: , In the formula, Represents any pairwise combination within the modality set; It is a 2-norm; this consistency constraint is used to enhance the consistent expression of cross-modal shared risk semantics; S64: Introduce a prototype anchor point mechanism at the modality-specific level to achieve relative alignment; For each risk type Learn the corresponding risk semantic prototype vector And use it as an alignment anchor point for modality-specific semantics; For the sample The set of positive risk types is defined as: , In the formula, A set of risk types; For the sample The corresponding set of positive risk types; Risk type Risk occurrence label; Construct a prototype alignment loss between modality-specific representations and risk semantic prototypes: , in, The prototype alignment loss is used to constrain the modality-specific representations of different modalities to remain similar to the prototypes of the same risk subclass, thereby enhancing cross-modal comparability. For modality In the sample The validity weight or mask coefficient at the point is used to characterize whether the modality observation is available and its contribution to training, and its value range is a non-negative real number; For modality In the sample The modality-specific representation vector at that location; Risk subclass The prototype vector is used to characterize the category center of the risk subclass in the unified representation space; This is a similarity function used to measure the degree of similarity between vectors; This is a temperature coefficient used to adjust the smoothness of the similarity distribution; S65: Define the overall objective function for cross-modal alignment; By integrating shared consistency constraints, prototype anchor constraints, and semantic decoupling constraints, a cross-modal alignment overall objective function is constructed: , In the formula, , , These are weighting coefficients used to balance the contributions of different constraint terms during training.

8. The AI-based comprehensive management and control method for large-scale sports events according to claim 1, characterized in that: S7 specifically includes the following steps: S71: Determine the fusion input and construct fusion candidate representations; Using the modality-shared risk semantic representation and the modality-specific semantic representation output from step S6 as fusion inputs, the modality set is defined as follows. For any mode Constructing fusion candidate representation vectors: , In the formula, For modality Shared risk semantic representation; For modality Modality-specific semantic representation; For modality Fusion candidate representation vectors; For region cell indexing; Index for regional time windows; S72: Calculate the risk relevance score; To measure the correlation between modal evidence and risk type, the risk semantic prototype vector from step S6 is reused. For any mode Define risk relevance score: , In the formula, A set of risk types; For similarity functions; For modality Risk relevance score; S73: Generate gating weights using modal quality indicators and risk correlation scores; Based on the modal quality indication quantity from step S4, denoted as... The modal mask in step S6 is denoted as Construct a gating fusion module to control modalities. Calculate the gating score: , In the formula, This is the gate vector; This refers to the quality weighting coefficient. This refers to the relevance weighting coefficient. For gated bias; For modality Gating scoring; S74: Form a fusion risk representation vector while retaining modal weights; The fusion risk representation vector of the regional time window samples is obtained by weighted summation based on dynamic fusion weights. , In the formula, The fused risk representation vector of the regional time window samples is used for the spatiotemporal risk evolution modeling in step S8. As the explanatory power follows Output together for early warning interpretation and manual review; To locate key pieces of evidence, for any modality The segment-level feature sequence within the time window is denoted as Constructing fragment contribution score: , In the formula, For modality Segment rating vector; For the first The contribution weight of each segment; The number of segments within the observation window; Select the top contributor with the largest contribution weight Each fragment index serves as the set of key contribution fragment indices: , In the formula, For modality A set of indexes of key contribution fragments; To get the maximum Operators for indexing elements.

9. The AI-based comprehensive management and control method for large-scale sports events according to claim 1, characterized in that, S8 specifically includes the following steps: S81: Construct node-level spatiotemporal input sequences; For any prediction start time window Take the past The fused risk representation vector of each regional time window is used as the spatiotemporal input, and the node-level input sequence is defined as follows: , , In the formula, The input sequence tensor; For time windows The node feature matrix; This represents the number of regional units; For regional units In the time window The fusion risk representation vector; The length of the observation window. These are the node feature terms arranged in chronological order in the input sequence; S82: Constructing the spatial adjacency and normalization matrix of a simplified spatiotemporal graph network; The venue topology adjacency matrix constructed in step S1 And introduce self-loop matrices This yields the adjacency matrix with added self-loops: , In the formula, It is a self-loop matrix, and the self-loop matrix is ​​the identity matrix; To add a self-loop adjacency matrix; Calculate the degree matrix And construct the normalized adjacency matrix: , , In the formula, It is a degree matrix; These are the diagonal elements of the degree matrix; The normalized adjacency matrix is ​​used to characterize the connectivity and diffusion relationships between the seating areas and the corridor, the convergence and bottleneck relationships between turnstiles and escalators, and the intrusion sensitivity of the restricted area boundary zone.

10. The AI-based comprehensive management and control method for large-scale sports events according to claim 9, characterized in that, S8 also includes the following steps: S83: Temporal convolution module extracts the difference between short-term bursts and long-term trends; Construct a one-dimensional temporal convolution module for... Perform convolution along the time dimension to obtain the temporal convolution output: , In the formula, Given an input sequence tensor, It is a one-dimensional temporal convolution operator; These are the temporal convolution parameters; The output tensor of the time convolution; S84: Spatial graph convolution module characterizes connectivity diffusion and bottleneck propagation; Perform graph convolution propagation on the temporal convolution output at each time step to obtain spatial augmentation features: , In the formula, For time steps The node feature matrix; These are the spatial convolution parameters; It is a non-linear activation function; For time steps Spatial enhanced feature matrix; S85: Construct a simplified spatiotemporal modeling network and output node-level spatiotemporal hidden states; A simplified spatiotemporal modeling network is formed by using a cascaded structure of temporal convolution and spatial graph convolution to obtain the prediction starting time window. Node-level spacetime hidden states: , , In the formula, For time-dimensional aggregation operators; This is the aggregated spatiotemporal representation matrix; For regional units In the predicted starting time window The spatiotemporal hidden state vector; S86: Set the risk prediction header and complete multi-step prediction output: Lightweight forecasting heads are set up for both the probability of risk occurrence and the level of risk, to predict the future. Each prediction window outputs the probability of risk occurrence: , In the formula, This represents the probability vector of multi-step risk occurrence. and These are the parameters for the probability prediction head. For element-wise Sigmoid function; Number of risk types; For the future Each prediction window outputs a probability distribution of risk levels: , In the formula, This represents a multi-step risk level probability distribution. and For the level prediction head parameters; This is the normalization function along the rank dimension; This represents the total number of preset risk levels.