Crowd space-time anomaly detection method based on unsupervised-semi-supervised stacking under small sample condition
By employing an unsupervised-semi-supervised stacking method and utilizing techniques such as Bootstrap resampling and consensus matrix calculation, a crowd anomaly detection model is constructed. This model addresses the issues of scarce labeled samples and spatiotemporal heterogeneity, achieving high-precision and robust crowd anomaly detection.
Patent Information
- Application Number
- CN202510782219.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies for anomaly detection in population suffer from problems such as scarce labeled samples, class imbalance, spatiotemporal heterogeneity, and differences in data distribution, which lead to decreased detection model performance and insufficient generalization ability.
We employ an unsupervised-semi-supervised stacking approach, using Bootstrap resampling, unsupervised component learners, consensus matrix calculation, and PAC scoring to achieve adaptive parameter tuning. We then combine a space-time hybrid enhancement strategy and a focal loss function to construct a crowd anomaly detection model.
Achieving high-precision anomaly detection in small sample conditions reduces the need for labeled data, improves model robustness and cross-scenario generalization ability, and effectively alleviates the problem of scarce labeled samples.
Smart Images

Figure CN120997751A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of crowd monitoring, and particularly relates to a crowd spatio-temporal anomaly detection method based on unsupervised-semi-supervised stacking under a small sample condition. BACKGROUND
[0002] With the acceleration of urbanization and the increase of large-scale public activities, safety monitoring of crowd gathering places has become an important part of urban management and public security guarantee. Crowd spatio-temporal anomaly detection refers to identifying events deviating from normal crowd behavior patterns in time and space dimensions, such as crowd falls, abnormal gatherings, panic escapes, and stampede risks. Timely and accurate detection of such abnormal events is of great significance for public safety, crowd management optimization, and effective emergency intervention.
[0003] Existing spatio-temporal anomaly detection methods mainly include the following categories: 1) methods based on statistical models, such as time series analysis and spatial statistical analysis, which have good interpretability but limited effectiveness in dealing with complex nonlinear spatio-temporal patterns, making it difficult to adapt to dynamically changing anomaly patterns. 2) methods based on traditional machine learning, such as support vector machines, random forests, and clustering algorithms, which can handle certain complex data patterns but usually require manual feature design and are prone to dimensionality curse problems when dealing with high-dimensional spatio-temporal data. 3) methods based on deep learning, such as convolutional neural networks, recurrent neural networks, and graph neural networks, which can automatically learn complex data representations but require a large amount of labeled data for training, and their performance significantly decreases in scenarios with scarce labeled data. 4) methods based on semi-supervised learning, such as pseudo-labeling, consistency regularization, and MixMatch.
[0004] However, existing methods still have the following problems in practical applications:
[0005] First, the lack of labeled samples of crowd anomaly events severely hinders the performance of detection models. Anomaly events such as crowd falls and stampede accidents occur infrequently, and labeling these events in large-scale monitoring systems is both expensive and time-consuming. Traditional supervised learning algorithms rely heavily on a large amount of labeled data to accurately learn patterns and make predictions, and the lack of labeled instances leads to poor reliability of detection models.
[0006] Second, there is a serious class imbalance problem in crowd anomaly detection. In crowd spatio-temporal data sets, the number of normal behavior samples far exceeds that of anomaly event samples, leading the model to predict the majority class and resulting in a high false negative rate, causing important anomaly events to be missed.
[0007] Third, the spatiotemporal heterogeneity of crowd anomaly patterns increases the detection difficulty. The crowd behavior in different regions and time periods shows different abnormal patterns, and a single detection model is difficult to generalize to a diversified spatiotemporal environment. For example, the crowd flow pattern in the subway station peak period is significantly different from the crowd gathering pattern in the shopping mall leisure period.
[0008] Fourth, existing semi-supervised learning methods often assume that labeled and unlabeled data come from the same distribution, which is often not true in practical applications of crowd monitoring. The crowd data collected in different monitoring scenarios, time periods and environmental conditions may have distribution differences, affecting the generalization performance of the model. SUMMARY
[0009] The present application provides a crowd spatiotemporal anomaly detection method based on unsupervised-semi-supervised stacking under small sample conditions, which can realize high-precision crowd anomaly detection.
[0010] To solve the above technical problems, the present application provides the following technical solutions: a crowd spatiotemporal anomaly detection method based on unsupervised-semi-supervised stacking under small sample conditions, characterized in that it comprises the following steps:
[0011] S1, collect crowd detection data, Bootstrap resample the input labeled data and unlabeled data, and generate multiple data subsets;
[0012] S2, use an unsupervised component learner to extract meta-features in parallel;
[0013] S3, realize adaptive parameter tuning of the component learner through consensus matrix calculation and PAC score;
[0014] S4, fuse the meta-features with the original features to build a semi-supervised meta-learner;
[0015] S5, input the crowd detection data processed by steps S1 to S4, and output the corresponding label, build a crowd anomaly detection network, and use a spatial-temporal hybrid enhancement strategy and focal loss function to train the crowd anomaly detection model, and output the detection result.
[0016] Further, the aforementioned step S1 comprises the following sub-steps:
[0017] S1.1, for the collected crowd monitoring data, let X be the labeled sample set, Y be the corresponding label, and U be the unlabeled sample set; combine the labeled data set X and the unlabeled data set U into a mixed data set as follows:
[0018] M={X,U}
[0019] Remove the label and use unsupervised learning for feature enhancement;
[0020] S1.2, Bootstrap resampling is performed on the mixed dataset M, as follows:
[0021]
[0022] where K is the number of resampling times, and the outputs are and
[0023] S1.3, the mixed dataset M is resampled K times and H times, respectively, with resampling ratios ρ1 and ρ2, respectively, for data augmentation and tuning; the Bootstrap resampled dataset is represented as:
[0024] {m k ∣k=1,...,K}
[0025] {m h ∣h=1,...,H}
[0026] The size of the Bootstrap resampled sample is |m| = ρ|M|, where ρ is the resampling ratio;
[0027] m k The labeled and unlabeled samples in
[0028]
[0029] Further, the aforementioned step S2 includes the following sub-steps:
[0030] S2.1, a low-level representation of the crowd spatio-temporal data is extracted using an unsupervised component learner; the classification results of the unsupervised component detector are defined as follows:
[0031] D={d c :c=1,...,C},
[0032] S2.2, the classification results of the unsupervised component learner are represented as:
[0033]
[0034] represent a plurality of different unsupervised component learners;
[0035] S2.3, in the process of generating meta-features, the dataset is resampled K times and input into the fine-tuned component learner, and the classification results are used as meta-features; the final augmented dataset connection representation is:
[0036]
[0037] where the fine-tuning hyperparameters determined by a consensus-based tuning procedure.
[0038] Further, the aforementioned unsupervised component learners include: density clustering, natural break, outlier detection, K-Nearest Neighbors anomaly detection, and Isolation Forest.
[0039] Further, the aforementioned step S3 includes the following sub-steps:
[0040] S3.1, establishing the generalization error bounds of component learners and meta-learners in a two-stage structure, as follows:
[0041]
[0042] where the variability AV(f Q between component learners reduces the generalization error of meta-learners.
[0043] S3.2, using consensus clustering as an unsupervised method to evaluate the robustness of each component learner.
[0044] For the hth Bootstrap resampling dataset m h , define the connection matrix
[0045]
[0046] S3.3, based on C component learners, the mixed dataset M = {X, U} contains N = |M| samples, then the connection matrix is an N x N matrix.
[0047] Based on the connection matrix, the consensus matrix of component learners d c is calculated as follows: And the consensus matrix is used to calculate the consensus cumulative distribution function CDF dc , and the fuzzy clustering proportion PAC is used to measure the robustness of the clustering model:
[0048]
[0049] where,
[0050]
[0051] where is an indicator matrix, and the entry (i, j) is equal to 1 only when sample i and sample j are selected in the same Bootstrap resampling iteration, the hyperparameters x1 and x2 control the scale of the selected curve area, stable component learners produce CDF graphs with consistent flat shapes near the center, and learners with excellent generalization ability obtain low PAC scores.
[0052] S3.4, determining the optimal hyperparameters by minimizing the PAC value:
[0053]
[0054] The hyperparameters δ corresponding to the minimum PAC value are selected as the optimal learning model.
[0055] Further, the aforementioned step S4 is specifically: constructing a meta-detector as a high-level model, learning to adapt to different spatiotemporal anomaly patterns of different people, and optimizing the task-independent strategy by using the meta-features generated by the unsupervised component learner; in the consensus-based data augmentation process, both labeled and unlabeled data are augmented as the optimization process of the unsupervised component learner; extending the MixMatch framework to achieve more comprehensive anomaly detection, and taking this process as a meta-detector for anomaly re-identification; sharpening for entropy minimization, implementing spatial-temporal mixing enhancement ST-MixUp that creates smoother boundaries, and focal loss function for handling unbalanced class problems.
[0056] Further, the aforementioned implementation of ST-MixUp is specifically:
[0057] In combination with the Beta distribution λ ~ Beta(α, α), set λ' = max(λ, 1-λ) to determine the degree of fusion,
[0058] Integrate the weighted random sampling WRS module into ST-MixUp, and design it to generate new samples within clusters and produce ambiguous samples between adjacent samples from different clusters;
[0059] The WRS process is defined as:
[0060]
[0061] Where WRS is a process of selecting sample pairs as ST-MixUp targets based on abnormal spatial feature selection;
[0062] Calculate the distance between samples: first calculate the distance between samples in S = Shuffle(Concat(X, U)):
[0063] d ij =dist(S i ,S j )
[0064] Use the distance as the sampling weight, the closer the two samples are, the higher the probability of being selected,
[0065] Use the inverse transform sampling method to obtain the sampling pair: calculate the logit of the two samples as the logits tensor:
[0066]
[0067] The input logits tensor is converted to a probability distribution by a Softmax operation:
[0068] p i =Softmax(W i )
[0069] Applied to each row slice to determine the probability of each sample;
[0070] Randomly draw u ~ uniform(0, 1) as the input of the cumulative distribution function,
[0071] Find the smallest index j that satisfies the following formula:
[0072] F(p i (j-1))≤u<F(p ij )
[0073] The index is used as the selected sample index of sample i, where F(p ij ) is the cumulative probability of sample j;
[0074] Obtain a sample pair by inverse transformation sampling, denoted as (i, j);
[0075] ST-MixUp synthesizes new training samples by constructing convex combinations of these pairs:
[0076]
[0077] Further, the foregoing step S5 comprises: adopting a focal loss function as a preferred loss function for enhancing labeled samples, using a binary focal cross-entropy to solve the problem of class distribution difference, and assigning higher importance to a small number of abnormal instances;
[0078] The focal loss function is defined as follows:
[0079]
[0080] Wherein α is a balance factor, γ is a focusing parameter, is the predicted probability.
[0081] Further, the foregoing final expression of constructing the USemiS loss function:
[0082] S5.1, apply MixMatch processing:
[0083] X′,U′=MixMatch(X * ,U * ,T,K,α)
[0084] Wherein X *U * For the data processed by steps S1-S4, T is the temperature parameter, K is the enhancement times, and a is the Beta distribution parameter;
[0085] S5.2, calculate the focal loss function of the marked data:
[0086]
[0087] Where p meta is the prediction probability of the meta detector;
[0088] S5.3, calculate the unmarked data loss:
[0089]
[0090] Where q is the pseudo label, obtained by sharpening processing,
[0091] S5.4, construct the total loss function:
[0092] L=L x +lambda u L u
[0093] Where lambda u is the weight coefficient of the unmarked data loss.
[0094] Further, the aforementioned crowd spatio-temporal anomaly detection method based on unsupervised-semi-supervised stacking in a small sample situation, characterized in that, further comprising applying the crowd anomaly detection model to train station transportation hubs, large event venues, and public service venues to detect crowd abnormal behaviors.
[0095] Compared with the prior art, the beneficial technical effects of the above technical solutions of the present application are as follows:
[0096] 1. The internal structure of the data is extracted by the unsupervised component learner, and effective feature representation can be obtained without relying on a large amount of labeled data, which significantly reduces the demand for labeled data.
[0097] 2. High-precision crowd anomaly detection is achieved under extreme small sample conditions.
[0098] 3. The unsupervised component learner is used to fully mine the abnormal patterns in unmarked crowd data, effectively alleviating the problem of labeled sample scarcity;
[0099] 4. The PAC measurement mechanism based on consensus clustering is introduced to realize the unsupervised adaptive optimization of the component learner parameters, and the model robustness is improved;
[0100] 5. The ST-MixUp spatio-temporal mixing augmentation strategy is designed to generate diverse training samples between different monitoring areas and time periods, enhancing the model's cross-scene generalization ability. BRIEF DESCRIPTION OF DRAWINGS
[0101] Figure 1 The figure is a schematic diagram of the overall process of the method of the present application.
[0102] Figure 2 The figure is a general workflow diagram of the present application, in which (a) is a meta-learner framework based on MixMatch, (b) is a schematic diagram of data augmentation using component learners to generate meta-features, (c) is a schematic diagram of the consensus-based component learner optimization process, and (d) is a schematic diagram of ST-MixUp seeking to create smooth boundaries for the model.
[0103] Figure 3 The figure is a schematic diagram of the empirical collection and study of two spatio-temporal data sets in the present application, in which (a) is a schematic diagram of vehicle trajectory detection research, and (b) is a schematic diagram of crowd fall detection research.
[0104] Figure 4 The figure is a consensus matrix of component learners on two grids in the present application. DETAILED DESCRIPTION
[0105] In order to better understand the technical content of the present application, specific embodiments are described below with reference to the accompanying drawings.
[0106] Aspects of the present application are described in the present application with reference to the accompanying drawings, which show a number of illustrative embodiments. The embodiments of the present application are not limited to the drawings described. It should be understood that the present application is implemented by any one of the above-described concepts and embodiments, as well as the concepts and embodiments described in detail below, since the concepts and embodiments disclosed in the present application are not limited to any embodiment. In addition, some aspects disclosed in the present application can be used alone or in any suitable combination with other aspects disclosed in the present application.
[0107] Reference Figure 1 The present embodiment provides a crowd spatio-temporal anomaly detection method based on unsupervised-semi-supervised stacking in a small sample case, including a meta-learner framework based on MixMatch, component learner data augmentation, a consensus-based optimization process, and ST-MixUp boundary smoothing.
[0108] As Figure 3 As shown in (a) and (b), the present application is verified in two typical crowd monitoring scenarios: vehicle trajectory anomaly detection and crowd fall detection.
[0109] For the vehicle trajectory anomaly detection scenario, vehicle target data is collected through video monitoring, and YOLO-v5 and DeepSORT algorithms are used to identify and track vehicle objects, and extract the trajectory information of vehicles in time and space dimensions. According to the traffic rules, the time range and geographic coordinates of abnormal events are labeled to form a labeled dataset X, including normal trajectory and abnormal trajectory samples.
[0110] For the crowd fall detection scenario, sensors are used to monitor crowd movement and detect fallen personnel, and computer vision is used for human pose estimation, and sensor-generated point cloud data of crowd movement is used. Normal crowd behavior and fall abnormal events are labeled separately to build a crowd monitoring dataset.
[0111] Let the labeled dataset be X, containing labeled samples and their corresponding labels Y; the unlabeled dataset U contains a large amount of unlabeled crowd behavior data.
[0112] The specific steps are as follows:
[0113] S1, collect crowd detection data, Bootstrap resample the input labeled data and unlabeled data to generate multiple data subsets. Including the following sub-steps:
[0114] S1.1, for the collected crowd monitoring data, let X be the labeled sample set, corresponding to the label Y, and U be the unlabeled sample set; combine the labeled dataset X and the unlabeled dataset U into a mixed dataset as follows:
[0115] M = {X, U}
[0116] Remove the label and use unsupervised learning for feature enhancement;
[0117] S1.2, Bootstrap resample the mixed dataset M as follows:
[0118]
[0119] Where K is the number of resampling times, and the output are the resampled labeled and unlabeled datasets, respectively;
[0120] S1.3, resample the mixed dataset M K and H times, respectively, and set the resampling ratios ρ1 and ρ2, respectively, for data enhancement and tuning; the Bootstrap resampled dataset is represented as:
[0121] {m k | k = 1,..., K}
[0122] {m h | h = 1,..., H}
[0123] The size of the bootstrap resampled sample is |m| = p|M|, where p is the resampling ratio.
[0124] m k The labeled and unlabeled samples in are defined as:
[0125]
[0126] S2, Extract meta-features in parallel using unsupervised component learners. Five complementary unsupervised component learners are used to extract crowd abnormal features:
[0127] (1) DBSCAN clustering: a density-based clustering method to identify crowd density abnormal areas.
[0128] (2) Natural Breaks (JNB): According to the distribution characteristics of crowd flow speed, the threshold of break point is automatically determined to divide the normal and abnormal flow speed levels.
[0129] (3) Outlier detection: statistical method is used to detect individual trajectories deviating from normal distribution.
[0130] (4) KNN method: based on the spatial proximity of the crowd, local abnormal behavior patterns are detected.
[0131] (5) Isolation Forest (ISO): specifically used to identify crowd gathering and dispersion anomaly patterns.
[0132] For each resampled dataset, meta-features are extracted using component learners:
[0133]
[0134] where is the hyperparameter of component learner d c .
[0135] The specific steps are as follows:
[0136] S2.1, Extract low-level representation of crowd spatio-temporal data using unsupervised component learners; define the classification results of unsupervised component detectors as follows:
[0137] D = {d c : c = 1,..., C},
[0138] S2.2, The classification results of unsupervised component learners are represented as:
[0139]
[0140] represents multiple different unsupervised component learners;
[0141] S2.3, in the generation of meta-features process, the data set is resampled K times and input into the fine-tuned component learner, and the classification result is used as the meta-feature; the final enhanced data set connection representation is:
[0142]
[0143] Where the fine-tuning hyperparameters Determined by the consensus-based tuning process.
[0144] S3, the adaptive parameter tuning of the component learner is realized by consensus matrix calculation and PAC score.
[0145] As Figure 2 (a) to (d) and Figure 4 The present application proposes an unsupervised parameter tuning mechanism based on consensus clustering, which selects the optimal parameters by quantifying the stability of the component learner in multiple resampling.
[0146] (1) For the hth resampling data set m h , define the connection matrix
[0147]
[0148] Assuming that there are C component learners, and the mixed data set M={X, U} contains N=|M| samples, then the connection matrix is an N×N matrix.
[0149] (2) Calculate the consensus matrix CM c of the component learner d dc (i,j):
[0150]
[0151] Where is an indicator matrix, and the entry (i,j) is equal to 1 only when sample i and sample j are selected in the same bootstrap resampling iteration.
[0152] (3) Calculate the consensus cumulative distribution function CDF:
[0153]
[0154] (4) Calculate the fuzzy clustering proportion PAC measure:
[0155] PAC dc = CDF dc (x2)- CDF dc (x1)
[0156] Where the hyperparameters x1 and x2 control the scale of the selected curve area.
[0157] (5) Determine the optimal hyperparameter:
[0158] δ dc = argmin δ PAC dc
[0159] Select the hyperparameter δ corresponding to the minimum PAC value to obtain the optimal learner.
[0160] The consensus-based process evaluation component learns the robustness on the Bootstrap dataset by minimizing the fuzzy clustering proportion PAC score, so that the stable learner obtains greater weight in the meta-feature aggregation process.
[0161] S4, fuse meta-features with original features to construct a semi-supervised meta-learner.
[0162] Specific implementation process:
[0163] Expand the MixMatch framework to achieve more comprehensive anomaly detection, and use this process as a meta-detector for anomaly re-identification.
[0164] (1) ST-MixUp data augmentation process:
[0165] Combined with the Beta distribution λ~Beta(α,α), set λ'=max(λ,1-λ) to determine the fusion degree.
[0166] Integrate the weighted random sampling (WRS) module into ST-MixUp, designed to generate new samples within clusters and produce fuzzy samples between adjacent samples from different clusters.
[0167] The WRS process is defined as:
[0168]
[0169] Where WRS is a process for selecting sample pairs as ST-MixUp targets based on anomaly space features.
[0170] (2) Calculate the distance between samples: First, calculate the distance between samples in S=Shuffle(Concat(X,U)): d ij =dist(S i ,S j )
[0171] Use the distance as the sampling weight, the closer the two samples are, the higher the probability of being selected.
[0172] (3) Use inverse transform sampling method to obtain the sampling pair:
[0173] The log probabilities of the two samples are computed as a logits tensor:
[0174]
[0175] The logits tensor is converted to a probability distribution by the Softmax operation:
[0176] p i =Softmax(W i )
[0177] Applied to each row slice to determine the probability of each sample.
[0178] A random draw u ~ uniform(0, 1) is taken as input to the cumulative distribution function.
[0179] Find the smallest index j such that:
[0180] F(p i (j-1))≤u<F(p ij )
[0181] This index is used as the selected sample index for sample i, where F(p ij ) is the cumulative probability of sample j.
[0182] A sample pair is obtained by inverse transform sampling, denoted as (i, j).
[0183] ST-MixUp synthesizes new training samples by constructing convex combinations of these pairs:
[0184]
[0185] S5, the crowd detection data processed by steps S1 to S4 is input, and the corresponding label is output, a crowd anomaly detection network is constructed, a spatial and temporal mixing enhancement strategy and a focal loss function are used to train the crowd anomaly detection model, and a detection result is output.
[0186] Use binary focal cross-entropy to solve the problem of class distribution difference, and assign higher importance to minority abnormal instances;
[0187] (1) Define the focal loss function:
[0188]
[0189] Where α is the balance factor, γ is the focusing parameter, is the predicted probability.
[0190] (2) Construct the final expression of USemiS loss function:
[0191] Firstly, the MixMatch is applied to process:
[0192] X', U' = MixMatch(X * ,U * ,T,K, alpha)
[0193] Wherein X * ,U * is the data processed by steps S1-S4, T is a temperature parameter, K is the number of enhancement, and alpha is a Beta distribution parameter.
[0194] The focal loss function of the labeled data is calculated:
[0195]
[0196] Wherein p meta is the prediction probability of the meta detector.
[0197] The loss of unlabeled data is calculated:
[0198]
[0199] Wherein q is a pseudo label, which is obtained by sharpening processing.
[0200] The total loss function is constructed:
[0201] L = L x + lambda u L u
[0202] The experimental results show that in the vehicle trajectory anomaly detection task, the AUC of the present application under 100, 200, 500 labeled samples (corresponding to 0.4%, 0.8%, 2% annotation ratio) is 0.9539, 0.9714, 0.9719 respectively; in the crowd fall detection task, the corresponding AUC is 0.8125, 0.8784, 0.9224 respectively.
[0203] Although the present application has been described as above with preferred embodiments, it is not intended to limit the present application. Those skilled in the art can make various modifications and improvements without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application shall be subject to the definition of the claims.
Claims
1. A method for detecting spatiotemporal anomalies in crowds based on unsupervised-semi-supervised stacking in small sample cases, characterized in that, Includes the following steps: S1. Collect crowd detection data, and perform Bootstrap resampling on the input labeled and unlabeled data to generate multiple data subsets; S2. Extract meta-features in parallel using unsupervised component learners; S3. Adaptive parameter tuning of the component learner is achieved through consensus matrix calculation and PAC scoring; S4. Fuse the meta-features with the original features to construct a semi-supervised meta-learner; S5. Using the crowd detection data processed in steps S1 to S4 as input and the corresponding labels as output, construct a crowd anomaly detection network, and train the crowd anomaly detection model using a spatial-temporal hybrid enhancement strategy and a focal loss function, and output the detection results.
2. The method for detecting spatiotemporal anomalies in a small sample size based on unsupervised-semi-supervised stacking, as described in claim 1, is characterized in that... Step S1 includes the following sub-steps: S1.
1. For the collected population monitoring data, let X be the labeled sample set with corresponding label Y, and U be the unlabeled sample set; merge the labeled dataset X and the unlabeled dataset U into a mixed dataset, as shown in the following formula: M = {X, U} Remove labels and use unsupervised learning for feature enhancement; S1.2, Perform Bootstrap resampling on the mixed dataset M, as follows: Where K is the number of resampling operations, and the output is... These are the resampled labeled and unlabeled datasets, respectively. S1.
3. Resample the mixed dataset M K times and H times respectively, setting the resampling ratios to ρ1 and ρ2, respectively, for data augmentation and optimization; the Bootstrap resampled dataset is represented as: {m k ∣k=1,...,K} {m h ∣h=1,...,H} The size of the Bootstrap resampled sample is: |m|=ρ|M|, where ρ is the resampling ratio; m k Labeled and unlabeled samples in the dataset are defined as follows:
3. The method for detecting spatiotemporal anomalies in a small sample size based on unsupervised-semi-supervised stacking, as described in claim 1, is characterized in that... Step S2 includes the following sub-steps: S2.
1. Use an unsupervised component learner to extract low-level representations of the spatiotemporal data of the crowd; define the classification result of the unsupervised component detector as follows: D={d c :c=1,...,C}, S2.2 The classification results of the unsupervised component learner are represented as follows: This represents multiple different unsupervised component learners; S2.3 During the generation of meta-features, the dataset is resampled K times and input into the fine-tuned component learner, and the classification result is used as the meta-feature. The final augmented dataset connection representation is as follows: Among them, fine-tuning hyperparameters Determined through a consensus-based tuning process.
4. The method for detecting spatiotemporal anomalies in a small sample size based on unsupervised-semi-supervised stacking, as described in claim 1, is characterized in that... Unsupervised component learners include: density clustering, natural breakpoints, outlier detection, K-nearest neighbor anomaly detection, and isolated forest.
5. The method for detecting spatiotemporal anomalies in a small sample size based on unsupervised-semi-supervised stacking, as described in claim 1, is characterized in that... Step S3 includes the following sub-steps: S3.1 Establish the generalization error bounds for the component learners and meta-learners in the two-stage structure, as follows: The variability AV(f) among component learners Q (D) reduces the generalization error of the meta-learner. S3.
2. Consensus clustering is used as an unsupervised method to evaluate the robustness of the learners of each component. For the h-th resampled dataset m h Define the connection matrix S3.
3. Given a learner with C components and a mixed dataset M = {X, U} containing N = |M| samples, what is the connection matrix? It is an N×N matrix; Learner d based on connection matrix calculation component c consensus matrix And the consensus matrix is used to calculate the cumulative distribution function (CDF). dc The robustness of the clustering model was evaluated using the fuzzy clustering proportion PAC metric. in, in The entry (i,j) is equal to 1 if and only if sample i and sample j are selected in the same Bootstrap resampling iteration. The hyperparameters x1 and x2 control the scale of the selected curve region. The stable component learner produces a CDF map with a consistent flat shape near the center. The learner with excellent generalization ability obtains a low PAC score. S3.4 Determine the optimal hyperparameters by minimizing the PAC value: The optimal learner is represented by the hyperparameter δ that corresponds to the minimum PAC value.
6. The method for detecting spatiotemporal anomalies in a small sample size based on unsupervised-semi-supervised stacking, as described in claim 1, is characterized in that... Step S4 specifically involves: constructing a meta-detector as a high-level model to learn and adapt to spatiotemporal anomaly patterns in different population groups; optimizing task-independent strategies by utilizing meta-features generated by an unsupervised component learner; enhancing both labeled and unlabeled data during consensus-based data augmentation as a tuning process for the unsupervised component learner; extending the MixMatch framework to achieve more comprehensive anomaly detection, using this process as a meta-detector for anomaly re-identification; and implementing sharpening for entropy minimization, implementing ST-MixUp for creating smoother boundaries, and a focal loss function to handle imbalanced class problems.
7. The method for detecting spatiotemporal anomalies in a small sample size based on unsupervised-semi-supervised stacking, as described in claim 6, is characterized in that... Implementing ST-MixUp specifically involves: Combining the Beta distribution λ ~ Beta(α,α), we set λ′ = max(λ,1-λ) to determine the degree of fusion; we integrated the weighted random sampling (WRS) module into ST-MixUp, designed to generate new samples within clusters and fuzzy samples between adjacent samples from different clusters; The WRS process is defined as follows: WRS is a process of selecting sample pairs as ST-MixUp targets based on anomaly space features; Calculate the distance between samples: First, calculate the distance between samples in S = Shuffle(Concat(X,U)): d ij =dist(S i ,S j ) Using distance as the sampling weight, the closer two samples are, the higher the probability of being selected. The inverse transform sampling method is used to obtain sample pairs: the log probabilities of the two samples are calculated and used as the logits tensor. The input logits tensor is transformed into a probability distribution through the Softmax operation: p i =Softmax(W i ) Apply to each row slice to determine the probability of each sample; Randomly select u ~ uniform(0,1) as the input to the cumulative distribution function. Find the minimum index j that satisfies the following formula: F(p i (j-1))≤u<F(p ij ) This index is used as the selection sample index for sample i, where F(p) ij ) is the cumulative probability of sample j; The sampling pairs are obtained by inverse transformation sampling and are represented as (i,j); ST-MixUp synthesizes new training samples by constructing convex combinations of these pairs:
8. The method for detecting spatiotemporal anomalies in a small sample size based on unsupervised-semi-supervised stacking, as described in claim 6, is characterized in that... Step S5 includes: using the focal loss function as the preferred loss function for enhancing labeled samples, using binary focal cross-entropy to address the class distribution difference problem, and assigning higher importance to a few anomalous instances; Define the focal loss function as follows: Where α is the balance factor and γ is the focusing parameter. To predict probabilities.
9. A method for detecting spatiotemporal anomalies in a small sample size based on unsupervised-semi-supervised stacking, as described in claim 7, is characterized in that... Construct the final expression for the USemiS loss function: S5.1, Apply MixMatch processing: X′,U′=MixMatch(X * ,U * ,T,K,α) Where X * U * The data is processed through steps S1-S4, where T is the temperature parameter, K is the number of enhancements, and α is the Beta distribution parameter. S5.2 Calculate the focal loss function for the labeled data: Where p meta The predicted probability of the meta-detector; S5.3 Calculate the loss of unlabeled data: Where q is a pseudo-label, obtained through sharpening. S5.4 Constructing the total loss function: L=L x +λ u L u Where λ u The weighting coefficients are for the loss of unlabeled data.
10. The method for detecting spatiotemporal anomalies in a small sample size based on unsupervised-semi-supervised stacking, as described in claim 1, is characterized in that... It also includes applying crowd anomaly detection models to railway transportation hubs, large event venues, and public service venues to detect abnormal crowd behavior.