Risk calculation method and system for improving low-rank attention mechanism
By using an improved low-rank attention mechanism, combined with surveillance video and sensor data, a cross-modal attention map is generated and a real-time risk value is calculated. This solves the problems of existing technologies being unable to capture nonlinear changes in spatiotemporal data and having high computational complexity, thus achieving efficient and accurate risk monitoring.
Patent Information
- Application Number
- CN202511005600.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-05-26
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies cannot effectively capture the nonlinear changes in spatiotemporal data in public safety risk monitoring. They have high computational complexity, are difficult to deploy in real time, and low-rank attention mechanisms lack flexibility and cannot adapt to the complexity of different scenarios, resulting in poor monitoring performance.
An improved low-rank attention mechanism is adopted, which combines surveillance video and sensor data to generate a cross-modal attention map by dynamically adjusting the rank, nonlinear low-rank decomposition and multi-head low-rank attention mechanism, and performing feature fusion. The real-time risk value is calculated using a GRU network.
It improves the computational efficiency and accuracy of risk monitoring, reduces network deployment and maintenance costs, and enhances the system's adaptability and stability in complex scenarios.
Smart Images

Figure CN120953907A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for monitoring safety in public places, and more specifically to a risk calculation method and system for improving low-rank attention mechanisms. Background Technology
[0002] With the accelerating pace of urbanization, safety in public places is receiving increasing attention. Against this backdrop, real-time monitoring and early warning of dynamic risks such as stampedes, fires, and falling objects from heights have become crucial issues in the field of public safety.
[0003] Traditional risk monitoring methods typically rely on static rules or pre-defined sensor thresholds for judgment, but these methods struggle to effectively capture the nonlinear variations in spatiotemporal data. This means that in dynamic and complex scenarios, traditional methods cannot accurately reflect changes, resulting in insufficient accuracy and timeliness of warnings. Existing deep learning models, due to the high complexity of their attention mechanisms and massive computational demands, are limited in their application in real-time monitoring scenarios, particularly in terms of efficient deployment on edge devices. Risk monitoring in public places often requires the collaborative work of multiple data sources, such as video surveillance data and sensor data. However, existing models, when fusing these heterogeneous data, fail to fully explore and utilize the correlations between them, leading to underutilization of potential information and impacting monitoring effectiveness.
[0004] To address the aforementioned issues, low-rank attention mechanisms have emerged, which reduce computational complexity through matrix factorization, thereby improving the model's computational efficiency. However, existing low-rank attention methods typically pre-determine a fixed rank value. This fixed choice cannot adapt to the complexity of different scenarios, thus failing to flexibly address various monitoring needs. Due to the limitations of low-rank approximation, the model may lose some key spatiotemporal features and data relationships, thereby affecting the accuracy of monitoring and early warning.
[0005] Therefore, it is necessary to design a new method to address the problems of existing technologies in public safety risk monitoring, such as the inability to capture nonlinear changes in spatiotemporal data, the complexity of deep learning computation making real-time deployment difficult, and the lack of flexibility of low-rank attention mechanisms, while reducing the deployment cost of risk computing networks and improving the computational efficiency of risk monitoring. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide an improved risk calculation method and system for low-rank attention mechanisms.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: a risk calculation method for an improved low-rank attention mechanism, comprising:
[0008] Acquire monitoring video and sensor data of the scene to be monitored to obtain initial data;
[0009] The initial data is preprocessed to obtain the preprocessing result;
[0010] Feature extraction is performed on the preprocessing results to obtain the extraction results;
[0011] The extracted results are input into the improved MLA module to generate a cross-modal attention map; wherein, the improved MLA module integrates multimodal spatiotemporal features and temporal features through dynamic rank adjustment, nonlinear low-rank decomposition and multi-head low-rank attention mechanism;
[0012] Cross-modal fusion features are determined based on the cross-modal attention map;
[0013] The cross-modal fusion features are input into the GRU network to calculate the real-time risk value;
[0014] The corresponding early warning strategy is triggered based on the real-time risk value and the set threshold.
[0015] The further technical solution is as follows: the preprocessing of the initial data to obtain the preprocessing result includes:
[0016] Image frames are extracted from the initial data, human pose estimation is performed, group density and motion vectors are calculated, and the sensor data is normalized to obtain preprocessed results.
[0017] The further technical solution is as follows: The feature extraction of the preprocessing result to obtain the extraction result includes:
[0018] The spatiotemporal features in the preprocessed results are extracted using a 3D-CNN model, and the temporal features in the preprocessed results are extracted using a 1D-CNN model. The extracted results include both spatiotemporal features and temporal features.
[0019] The further technical solution is as follows: the step of inputting the extraction result into the improved MLA module to generate a cross-modal attention map includes:
[0020] The extraction results are input into the improved MLA module, where the attention values of the spatiotemporal features and the temporal features are calculated using a cross-modal attention calculation function, and then scaled using an exponential function to obtain a weighted value.
[0021] The weighted values are normalized to obtain the attention weights;
[0022] The spatiotemporal features and the temporal features are weighted and averaged according to the attention weights to obtain a cross-modal attention map.
[0023] The further technical solution is as follows: when the extraction result is input into the improved MLA module to generate a cross-modal attention map, a dynamic rank selection strategy and a nonlinear low-rank decomposition strategy are adopted to adjust according to the complexity and nonlinear relationship of the features, and a multi-head low-rank attention is adopted to generate a cross-modal attention map.
[0024] The further technical solution is as follows: the dynamic rank selection strategy includes dynamically adjusting the rank size in the low-rank decomposition according to the entropy value of the input features; the nonlinear low-rank decomposition strategy includes nonlinearly enhancing the traditional low-rank decomposition by introducing the ReLU activation function; the multi-head low-rank attention includes parallel computing of multiple low-rank attention heads and concatenating the results of the low-rank attention heads to generate a cross-modal attention map.
[0025] The further technical solution is as follows: determining the cross-modal fusion features based on the cross-modal attention map includes:
[0026] Entity-related information is detected based on the preprocessing results, and the vector is embedded through the embedding layer. Relevant features are determined through the cross-modal attention map, and the vector is concatenated with the relevant features to obtain cross-modal fusion features.
[0027] The further technical solution is as follows: inputting the cross-modal fusion features into the GRU network to calculate the real-time risk value includes:
[0028] The cross-modal fusion features are input into the GRU network to generate hidden layer states. These states are then processed through a fully connected layer and an activation function to obtain real-time risk values.
[0029] The further technical solution is as follows: the early warning strategy includes the processing strategies corresponding to the three early warning levels of attention, warning, and high risk.
[0030] This invention also provides an improved risk calculation system for low-rank attention mechanisms, comprising:
[0031] The data acquisition unit is used to acquire monitoring video and sensor data of the scene to be monitored in order to obtain initial data;
[0032] A preprocessing unit is used to preprocess the initial data to obtain a preprocessing result;
[0033] A feature extraction unit is used to extract features from the preprocessing results to obtain the extraction results;
[0034] An attention map generation unit is used to input the extraction results into the improved MLA module to generate a cross-modal attention map;
[0035] A feature determination unit is used to determine cross-modal fusion features based on the cross-modal attention map;
[0036] The calculation unit is used to input the cross-modal fusion features into the GRU network to calculate the real-time risk value;
[0037] The early warning unit is used to trigger the corresponding early warning strategy based on the real-time risk value and the set threshold.
[0038] The advantages of this invention compared to existing technologies are as follows: This invention optimizes the risk calculation process by improving the low-rank attention mechanism module and combining multimodal information from surveillance video and sensor data. First, through preprocessing and feature extraction, noise is effectively removed and key features are extracted. Then, an improved MLA module is used to generate a cross-modal attention map, and then cross-modal fusion features are calculated. Finally, the fused features are input into the GRU network to calculate the risk value in real time and trigger an early warning strategy based on a set threshold. This method reduces the dependence of traditional methods on complex network structures and large amounts of data processing through efficient data fusion and calculation, significantly improving the computational efficiency of risk monitoring while reducing the cost of network deployment and maintenance.
[0039] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 A flowchart illustrating the risk calculation method for the improved low-rank attention mechanism provided in this embodiment of the invention;
[0042] Figure 2 A schematic diagram of a sub-process of the risk calculation method for the improved low-rank attention mechanism provided in an embodiment of the present invention;
[0043] Figure 3 A schematic diagram illustrating the workflow of the improved MLA module provided in an embodiment of the present invention;
[0044] Figure 4 A flowchart illustrating the risk calculation method for the improved low-rank attention mechanism provided in this embodiment of the invention. Figure 2 ;
[0045] Figure 5 A schematic block diagram of a risk calculation system with an improved low-rank attention mechanism provided in an embodiment of the present invention;
[0046] Figure 6A schematic block diagram of the attention graph generation unit of the risk calculation system with an improved low-rank attention mechanism provided in an embodiment of the present invention;
[0047] Figure 7 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0049] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0050] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0051] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0052] Please see Figure 1 , Figure 1This is a schematic flowchart illustrating the risk calculation method based on an improved low-rank attention mechanism provided in an embodiment of the present invention. This improved low-rank attention mechanism risk calculation method is applied in a server. The server interacts with sensors, cameras, etc., and by introducing an improved low-rank attention mechanism, combined with multimodal data (such as surveillance video and sensor data), the accuracy and efficiency of risk calculation are effectively improved. By preprocessing the initial data, extracting features, performing cross-modal fusion, and calculating attention, and combining dynamic rank selection and nonlinear low-rank decomposition strategies, the information processing and feature fusion process are optimized. Furthermore, using a GRU network for real-time risk assessment effectively improves the system's response speed and prediction accuracy, and reduces the computational cost in traditional risk calculation methods. The improved multi-head low-rank attention mechanism not only reduces computational resource consumption but also enhances the system's adaptability and stability in complex scenarios, thereby improving the computational efficiency of risk monitoring while reducing deployment costs.
[0053] Figure 1 This is a flowchart illustrating the risk calculation method for the improved low-rank attention mechanism provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110 to S170.
[0054] S110. Acquire the monitoring video and sensor data of the scene to be monitored to obtain initial data.
[0055] In this embodiment, initial data refers to raw information acquired through various sensing devices (such as surveillance cameras and sensors). Specifically, this data includes the following two categories:
[0056] Surveillance video data primarily originates from cameras, which record real-time video streams within the monitored area. This video data is typically a continuous sequence of dynamic images, containing objects, pedestrians, vehicles, and other potentially dynamic elements within the scene.
[0057] From videos, especially in densely populated or complex dynamic scenes, many valuable spatiotemporal features can be extracted, such as the direction of crowd movement, density changes, sudden activities, or abnormal behaviors. This information can help identify potential risks (such as congestion, stampedes, etc.).
[0058] Sensor data includes numerical information collected from environmental monitoring devices such as temperature sensors, smoke sensors, and pressure sensors. Sensor data is an important way to capture changes in environmental conditions. For example, abnormal temperature fluctuations may indicate fire hazards, an increase in smoke concentration may be an early sign of a fire, and pressure changes may indicate certain structural risks.
[0059] This data is typically time-series, which helps assess dynamic changes in the environment, and when combined with surveillance video data, it can provide a more comprehensive risk assessment.
[0060] S120. The initial data is preprocessed to obtain the preprocessing result.
[0061] In this embodiment, image frames are extracted from the initial data, human pose estimation is performed, group density and motion vectors are calculated, and the sensor data is normalized to obtain preprocessing results.
[0062] In this embodiment of the invention, the preprocessing step is crucial. Its main purpose is to clean, extract features from, and standardize the initial data obtained from the monitoring system to make it suitable for subsequent risk calculation and model training. The specific preprocessing operations are divided into the following parts:
[0063] Image frames are extracted from surveillance video within a specific time range. These image frames will serve as the basis for subsequent processing. The extracted image frames are typically sampled at fixed time intervals to ensure the continuity and integrity of spatiotemporal information.
[0064] The extracted image frames are processed by a human pose estimation algorithm, with the goal of identifying and locating the poses and positions of pedestrians or other key objects in the image. This method captures each person's activity trajectory, stance, and direction of movement, which is crucial for calculating group density and motion vectors.
[0065] The number and distribution of pedestrians within an image region are estimated from image frames. Crowd density is calculated based on human pose estimation, by counting the number of pedestrians in each frame or by measuring the ratio of the area occupied by pedestrians to the total area in the image. Sudden changes in crowd density can indicate potential risks in different public safety scenarios, making this an important characteristic.
[0066] Motion vectors represent the direction and velocity of change of objects (such as pedestrians) in an image over time. By comparing differences between consecutive frames, the trajectory of each object can be calculated. Motion vectors not only help predict the future location of pedestrians but also help determine crowd mobility and aggregation trends. This is especially important in emergency situations, determining the potential risks of overcrowding, conflict, or stampedes.
[0067] Since sensor data often varies due to sensor measurement errors or environmental changes, standardization is necessary to ensure that all sensor input data are compared and processed at the same scale. Sensor data normalization adjusts data of different dimensions or magnitudes to the same range (e.g., between 0 and 1), preventing some sensor data from having an excessive impact on the model due to their large magnitude. Common normalization methods include min-max normalization or Z-score standardization.
[0068] After the above processing, all data has been cleaned and transformed, resulting in a unified preprocessed result. This result will be input into the subsequent risk calculation model. The final preprocessed result includes:
[0069] Image features: pedestrian density and motion vector information from video frames;
[0070] Sensor characteristics: Normalized environmental sensor data (such as temperature, pressure, smoke concentration, etc.);
[0071] Data integration: Ultimately, image data and sensor data will be fused through an improved low-rank attention mechanism (improved MLA module) to extract more spatiotemporal features and potential risk information;
[0072] Through this series of preprocessing steps, the system is able to extract useful features from the raw data, providing a solid foundation for subsequent risk analysis and prediction.
[0073] S130. Perform feature extraction on the preprocessing results to obtain the extraction results.
[0074] In this embodiment, the extraction result refers to the spatiotemporal features and temporal features extracted from video data and sensor data respectively using 3D-CNN and 1D-CNN models.
[0075] Specifically, a 3D-CNN model is used to extract spatiotemporal features from the preprocessed results, and a 1D-CNN model is used to extract temporal features from the preprocessed results. The extracted results include both spatiotemporal features and temporal features.
[0076] Step S130 mainly refers to feature extraction from the preprocessed results. Feature extraction is a crucial step in data processing, especially when dealing with multimodal data. How to extract effective features from complex raw data is essential for subsequent model building and prediction tasks.
[0077] The goal of feature extraction is to extract representations from the preprocessed results that express the spatiotemporal and temporal features of the data. Specifically, we need to extract spatiotemporal features from video data and temporal features from sensor data. These features provide the foundational information for subsequent model learning and risk prediction.
[0078] 3D Convolutional Neural Networks (3D-CNNs) are neural network models capable of performing convolutions simultaneously in the temporal and spatial dimensions, making them particularly suitable for processing data with spatiotemporal relationships, such as video data. In this embodiment, a 3D-CNN model is used to extract spatiotemporal features from video data. Specifically:
[0079] Input data: Video frame data, which typically includes multiple consecutive image frames representing a dynamic scene over a certain period of time.
[0080] Spatiotemporal Feature Extraction: By applying convolutional operations simultaneously in both time and space dimensions, 3D-CNNs can capture complex spatiotemporal features in videos, such as object motion and scene changes. For example, in public safety surveillance videos, key information such as changes in crowd density and movement trajectories can be extracted.
[0081] Output Features: The features output by a 3D-CNN model are typically a high-dimensional spatiotemporal feature matrix, representing various information captured from the video data in both time and space dimensions. This feature matrix plays a crucial role in subsequent attention calculations.
[0082] Temporal features refer to features related to a time series, describing the patterns of data change over time. In this embodiment, sensor data (such as temperature, smoke concentration, pressure, etc.) is time-series data; therefore, a 1D convolutional neural network (1D-CNN) is used for feature extraction. Specifically:
[0083] Input data: Sensor data sequences, typically sensor readings collected at multiple times, which reflect changes in the physical environment of the hazardous scenario (such as increased temperature, increased smoke concentration, etc.).
[0084] Temporal Feature Extraction: 1D-CNN models extract important features from temporal data through convolutional operations, enabling them to capture trends in sensor data over time. For example, fluctuations in temperature and smoke concentration may be related to the occurrence of emergencies such as fires.
[0085] Output Features: The 1D-CNN model outputs a time-series feature matrix, representing the underlying patterns and trends in the time-series data. This feature matrix is crucial for subsequent risk assessment and dynamic prediction.
[0086] In this way, the spatiotemporal features extracted by 3D-CNN and the temporal features extracted by 1D-CNN respectively represent different aspects of video and sensor data. In the subsequent multimodal feature fusion stage, these two types of features will be combined to capture more comprehensive spatiotemporal dynamic information. This fusion process typically relies on an improved low-rank attention mechanism (MLA, Matrix Low-Rank Attention), which can effectively learn and weight these features, thereby achieving more accurate risk quantification and prediction.
[0087] The goal of step S130 is to extract spatiotemporal features using a 3D-CNN model and temporal features using a 1D-CNN model, thereby providing rich feature information for subsequent model decisions. These features will be used as input and further processed and analyzed through a multimodal feature fusion module (such as an improved MLA module) to complete the risk prediction and early warning tasks.
[0088] S140. Input the extraction result into the improved MLA module to generate a cross-modal attention map.
[0089] In this embodiment, the generation of cross-modal attention maps is a core part of the entire risk calculation process. It involves effectively fusing information from different modalities (visual features and sensor features) through an improved low-rank attention mechanism and generating an attention map that describes the relationship between them. Cross-modal attention maps help the model focus on the important correlations between spatiotemporal and temporal features, thereby making risk predictions more accurate. Specifically, the improved MLA module fuses multimodal spatiotemporal and temporal features through dynamic rank adjustment, nonlinear low-rank decomposition, and a multi-head low-rank attention mechanism.
[0090] Weighted values refer to numerical values that assign different importance to different features or data during the calculation process, while attention weights represent the degree of attention paid to each modal information in cross-modal fusion, and are used to dynamically adjust the model's emphasis on each input.
[0091] In public safety scenarios, the spatiotemporal complexity of different risk events varies significantly (for example, the entropy value of sudden changes in population density is much higher than in normal scenarios), making it difficult to balance computational efficiency and feature representation capability with fixed rank. By dynamically adjusting the rank, it is possible to adapt to complex scenarios and reduce redundant computation.
[0092] In one embodiment, please refer to Figure 2 The above-mentioned step S140 may include steps S141 to S143.
[0093] S141. The extraction result is input into the improved MLA module, and the attention value of the spatiotemporal feature and the temporal feature is calculated by the cross-modal attention calculation function, and then scaled by the exponential function to obtain the weighted value.
[0094] In this embodiment, the features input to the improved MLA module are spatiotemporal features (from the visual modality, such as video data features extracted by 3D-CNN) and temporal features (from the sensor modality, such as temporal features of sensor data such as temperature, smoke, and pressure).
[0095] In the improved MLA module, cross-modal computation is first performed on spatiotemporal and temporal features. This process uses an attention mechanism to measure the correlation between two different modalities and calculates the attention value (weight) between them. The goal of cross-modal attention computation is to learn the importance of spatiotemporal and temporal features in specific risk scenarios.
[0096] To amplify the influence of important features, the calculated attention values are scaled using an exponential function. This process helps to highlight the weights of highly relevant features and allows the model to focus more intently on these key components.
[0097] S142. Normalize the weighted values to obtain attention weights.
[0098] In this embodiment, the weighted attention values are normalized to convert them into a uniform range (usually between 0 and 1), allowing the weights of each feature to be adjusted appropriately. After normalization, the attention weights better reflect the relative importance of features and facilitate subsequent weighted averaging operations.
[0099] The normalized attention weights represent the cross-modal relationship between spatiotemporal features and temporal features, which can guide the model to focus more on high-weight features and ignore low-weight parts when generating cross-modal attention maps.
[0100] S143. Perform a weighted average of the spatiotemporal features and the temporal features according to the attention weights to obtain a cross-modal attention map.
[0101] In this embodiment, a weighted averaging operation is used to weight spatiotemporal features and temporal features based on the calculated attention weights. This operation enables the model to adjust the contribution of features according to the attention weights, thereby obtaining a cross-modal attention map that integrates spatiotemporal and temporal features.
[0102] The final cross-modal attention map is obtained by weighted averaging of the feature maps. This map represents the correlation between visual and sensor modal features in a specific risk scenario, and can provide effective information support for subsequent risk prediction and decision-making.
[0103] Specifically, the spatiotemporal features F of the video data are extracted using 3D-CNN. v ∈R TxHxWxC Temporal features F were extracted from temperature, smoke, and pressure data using a 1D-CNN. s ∈R TxD .
[0104] Input the multimodal features into the improved MLA module to calculate cross-modal attention weights: Joint features F are generated through weighted fusion. fusion ∈R TxC .
[0105] Furthermore, when the extraction results are input into the improved MLA module to generate a cross-modal attention map, a dynamic rank selection strategy and a nonlinear low-rank decomposition strategy are adopted to adjust according to the complexity and nonlinear relationship of the features, and a multi-head low-rank attention is used to generate the cross-modal attention map.
[0106] The dynamic rank selection strategy includes dynamically adjusting the rank size in the low-rank decomposition based on the entropy value of the input features; the nonlinear low-rank decomposition strategy includes nonlinearly enhancing the traditional low-rank decomposition by introducing the ReLU activation function; the multi-head low-rank attention includes computing multiple low-rank attention heads in parallel and concatenating the results of the low-rank attention heads to generate a cross-modal attention map.
[0107] In each step of the above process, the dynamic rank selection strategy and the nonlinear low-rank decomposition strategy play a crucial role:
[0108] The rank k in the low-rank decomposition is dynamically adjusted based on the complexity of the input features (measured by the entropy value Entropy(X)). By choosing an appropriate rank, redundant computation can be reduced and computational efficiency improved while maintaining the expressive power of the features. This is based on the input features X∈R. nxd The input features are a concatenation of multimodal spatiotemporal features, including visual and sensor features, where the entropy value of the input features is dynamically adjusted to adjust the rank k. Where We is a learnable parameter, α is the Sigmoid function, and Entropy(X) is the Shannon entropy of the input features, used to measure data complexity.
[0109] By introducing the ReLU activation function, the traditional low-rank decomposition is enhanced with nonlinearity. This allows the model to better capture the nonlinear relationships in spatiotemporal and temporal features, improving its performance. The traditional linear decomposition W = AB is improved to a three-factor decomposition: W = A·ReLU(B)·C + D; where A ∈ RT nxk ,B∈R kxk ,C∈R kxn ,D∈R nxn As a learnable matrix, the residual term D retains high-frequency information.
[0110] Multiple low-rank attention heads are computed in parallel, and their results are concatenated. This method not only further improves the computational efficiency of the model but also enhances its expressive power, enabling cross-modal feature fusion from multiple perspectives. Multiple improved MLA heads are parallelized, and the outputs are concatenated and fused through a fully connected layer: MultiHead(Q,K,V)=Concat(head1,...,head) h W O The computational complexity of each head is O(n^2). 2 ) decreases to O(nk)(k << n).
[0111] In this implementation process, step S140 involves how to effectively fuse the spatiotemporal and temporal features of the visual and sensor modalities through an improved low-rank attention mechanism to generate a cross-modal attention map. This process includes steps such as cross-modal attention calculation, weight normalization, and weighted averaging of features. Combined with dynamic rank selection and nonlinear low-rank decomposition strategies, it ultimately achieves efficient multimodal feature fusion, providing more accurate feature input for subsequent risk calculation.
[0112] Please see Figure 3 The improved MLA module works as follows: Input data is dynamically selected via a router and guided to different Multi-Head Expert Models (MLAs) based on specific conditions. Each MLA is designed to be parallelized for efficient computation.
[0113] For each query (Q), key (K), and value (V) matrix in an MLA, low-rank processing is performed. This means that in engineering implementations, it is not necessary to store the complete Q, K, and V matrices, but only their ranks (i.e., the reduced-dimensional representation of the matrices), thereby reducing GPU memory consumption.
[0114] During inference, the QKV matrix is restored to a form compatible with multi-head attention (MHA) computation. This is achieved through matrix absorption, which transforms the low-rank representation back into its full matrix form.
[0115] This strategy of low-ranking and matrix absorption makes the memory-intensive tasks of MHA computation more focused on computation during inference, thereby improving overall computational efficiency and making better use of the hardware's computing power.
[0116] The outputs of multiple heads in each MLA are concatenated and fused through a fully connected layer. The goal of this step is to integrate the feature information from different heads to ultimately generate the model's output.
[0117] In summary, this architecture reduces memory requirements through low-rank processing and optimizes the computation process through matrix absorption, transforming the originally memory-intensive MHA computation into a computation-intensive one, which greatly improves inference efficiency.
[0118] In this embodiment, the rank k is adjusted using the Shannon entropy of the input feature. Shannon entropy measures the complexity of the data; therefore, input features with higher entropy values result in higher rank, thus capturing more complex information.
[0119] The specific calculation formula is as follows: Among them, W e Here, α is a learnable parameter, α is the Sigmoid function, and Entropy(X) is the Shannon entropy of the input feature, used to measure data complexity; the input feature X∈R nxd The input feature matrix X has n samples, each with d dimensions. Here, the input features are a concatenation of multimodal spatiotemporal features, including visual features and sensor features.
[0120] Entropy (X) is used to measure the complexity or uncertainty of data; the higher the entropy, the greater the complexity of the data. Here, Shannon entropy is used to quantify the complexity of features.
[0121] Dynamically adjusting the rank k: To select a suitable low-rank approximation, the rank k is first dynamically adjusted using the entropy value of the input features. This rank determines the number of features in the matrix factorization or the dimension after dimensionality reduction.
[0122] k max k min These are the minimum and maximum values of the rank, respectively. W e is a learnable parameter matrix whose rank is adjusted by multiplying it with the entropy value. σ is a sigmoid function that compresses the input to the interval [0, 1] for nonlinear transformation. Finally, the calculated k determines the rank value used for low-rank decomposition in subsequent operations.
[0123] In this way, the rank k is dynamically adjusted according to the complexity of the input features (such as entropy), thereby adapting to different scenarios.
[0124] The nonlinear low-rank decomposition strategy improves the traditional linear decomposition W = AB into a three-factor decomposition using the formula: W = A · ReLU(B) · C + D; where A ∈ R nxk ,B∈R kxk ,C∈R kxn ,D∈R nxn As a learnable matrix, the residual term D retains high-frequency information.
[0125] Linear decomposition typically uses W = A·B or W = A·B·C, but this module has made improvements.
[0126] The traditional decomposition method is improved by introducing a nonlinear activation function (ReLU). Specifically, matrix BB is processed by the ReLU function before being multiplied with other matrices.
[0127] Explanation of each matrix:
[0128] A∈R n×k Let represent an n×k matrix, where n is the number of samples and k is the dynamically selected rank. B∈R k×k Let C represent a k×k matrix. C∈R k×n Let D represent a k×n matrix. n×n It is a residual term matrix used to retain high-frequency information and ensure that the model has stronger expressive power.
[0129] The nonlinear low-rank decomposition here uses the ReLU activation function, which makes the decomposition more expressive and able to capture nonlinear relationships in the data.
[0130] The formula used for multi-head low-rank attention is:
[0131] MultiHead(Q,K,V)=Concat(head1,...,head h W O Multi-head attention mechanism: This is a common attention mechanism that aims to perform attention computations simultaneously in multiple subspaces to capture different feature representations. Specifically, multiple independent attention heads compute in parallel, and then their results are concatenated.
[0132] The computational complexity of traditional self-attention mechanisms is O(n^2). 2 ), where n is the length of the input sequence.
[0133] Here, by using low-rank approximation and parallel processing, the computational complexity is reduced to O(nk), where k is the rank and k << n. This significantly reduces the computational cost.
[0134] Meaning of each part:
[0135] Q, K, V: These are the query, key, and value matrices, respectively. They are typically derived from input features and used for self-attention calculation. head1, ..., head h : Represents different attention heads, which perform independent attention calculations using different weight matrices. Concat: Concatenates the results of all attention heads. W O It is a weight matrix of a fully connected layer, used to fuse the concatenated results.
[0136] By parallelizing multiple attention heads, information can be captured simultaneously in multiple subspaces, thereby enhancing the model's expressive power.
[0137] By dynamically adjusting the rank, nonlinear low-rank decomposition, and multi-head low-rank attention, we can effectively process and fuse multimodal features and reduce computational complexity.
[0138] By employing dynamic rank adjustment and nonlinear low-rank decomposition, the improved low-rank attention mechanism (MLA) more effectively balances computational efficiency and expressive power when handling complex spatiotemporal features. The dynamic rank selection module adjusts the rank based on the entropy value of the input data, enabling adaptive selection of appropriate rank in different scenarios. Nonlinear low-rank decomposition further enhances the model's expressive power, while the multi-head low-rank attention mechanism reduces computational complexity, allowing the overall model to efficiently handle large-scale spatiotemporal data.
[0139] S150. Determine cross-modal fusion features based on the cross-modal attention map.
[0140] In this embodiment, cross-modal fusion features refer to the weighted fusion of information from different modalities (such as vision, sensor data, etc.) to form a comprehensive feature representation.
[0141] Specifically, entity-related information is detected based on the preprocessing results, and the vector is embedded through the embedding layer. Relevant features are determined through the cross-modal attention map, and the vector is concatenated with the relevant features to obtain cross-modal fusion features.
[0142] The goal of this stage is to fuse information from multiple different modalities to form an effective cross-modal feature representation. The specific process is as follows:
[0143] Using object detection algorithms like YOLOv5, the first step is to detect objects that are associated with risk, such as pedestrians, fire sources, and objects hanging from heights. The YOLOv5 algorithm can not only detect the category of an object but also provide location information for each object (such as the coordinates of its bounding box).
[0144] The embedding layer encodes each detected entity category (e.g., pedestrian, fire source) and its corresponding location information into a vector. This vector contains information about the entity category and its spatial location encoding, ensuring that the model can understand the spatial relationships between different types of entities.
[0145] To effectively fuse information from different modalities (such as visual features and sensor data), a cross-modal attention mechanism is used. Through cross-modal attention maps, the model can dynamically learn the importance of different modalities and weightedly fuse information from them. For example, visual features may be more important for risk assessment in some situations, while sensor data (such as temperature and smoke concentration) may be more valuable in others.
[0146] Relevant features calculated through cross-modal attention maps (such as entity information from vision, location, temperature, smoke concentration, and other sensor data) are concatenated with entity vectors. The resulting feature vector is the "cross-modal fusion feature," which contains multimodal information and helps the model better understand different types of risk information.
[0147] S160. Input the cross-modal fusion features into the GRU network to calculate the real-time risk value.
[0148] In this embodiment, the real-time risk value refers to a numerical value that reflects the current risk level, calculated by analyzing cross-modal data at the current moment.
[0149] Specifically, the cross-modal fusion features are input into the GRU network to generate hidden layer states, which are then processed through a fully connected layer and an activation function to obtain real-time risk values.
[0150] In this step, the cross-modal fusion features are input into a GRU (Gated Recurrent Unit) network. The goal is to capture temporal dependencies through the GRU network to calculate the real-time risk value. The specific steps are as follows: Inputting cross-modal fusion features into the GRU network: The cross-modal fusion features obtained after step S150 are used as input and fed into the GRU network. GRU is a variant of Recurrent Neural Network (RNN) specifically designed for processing temporal data. Through a gating mechanism, it can effectively capture long-term dependencies in time-series data. The GRU network processes the input cross-modal fusion features to generate hidden states. These hidden states contain memories of historical information, helping the network understand the current risk state over time. The hidden states are further processed through a fully connected layer to map the features to the final risk output space. An activation function (such as the sigmoid function) is applied to the output of the fully connected layer, producing a value R.t ∈[0,1] represents the real-time risk value at time t. This value reflects the risk level at the current moment.
[0151] Specifically, a gated recurrent unit (GRU) is used to capture timing dependencies and output the risk probability: Among them, R t ∈[0,1] represents the risk level at time t.
[0152] Traditional risk calculation methods often employ simple splicing or weighted fusion when processing multimodal data, which makes it difficult to fully explore the complex relationships and deep information between different modal features, resulting in poor fusion effects and an inability to accurately reflect the actual risk situation.
[0153] The improved MLA module in this embodiment, through dynamic rank adjustment, nonlinear low-rank decomposition, and multi-head low-rank attention mechanisms, can adaptively adjust based on feature complexity and nonlinear relationships, more accurately capturing the complex correlations and complementary information between multimodal spatiotemporal features and temporal features. For example, the dynamic rank selection strategy can dynamically adjust the rank size in the low-rank decomposition according to the entropy value of the input features, flexibly adapting to the complexity of features in different scenarios; the nonlinear low-rank decomposition strategy introduces the ReLU activation function, enhancing the nonlinear expressive power of traditional low-rank decomposition, enabling the model to better handle complex nonlinear feature relationships; and the multi-head low-rank attention computes multiple low-rank attention heads in parallel and concatenates the results to generate a cross-modal attention map from multiple perspectives, more comprehensively reflecting the correlations between features of different modalities. Multimodal feature fusion achieved in this way, compared to existing technologies, can more accurately determine cross-modal fusion features, providing a richer and more accurate information foundation for subsequent risk calculation, thereby making the calculated real-time risk value more accurate and the triggering of early warning strategies more targeted and effective.
[0154] Moreover, previous risk calculation methods mostly focused on single image features or simple motion features when processing surveillance video data. They did not delve deeply enough into the rich spatiotemporal and temporal information contained in the video, and were not accurate or comprehensive enough in extracting key information such as human posture and group density, resulting in inaccurate and untimely assessment of scene risks.
[0155] This embodiment employs a 3D-CNN model to extract spatiotemporal features from the preprocessed results, simultaneously capturing changes and correlations in both spatial and temporal dimensions of the video. This provides a more comprehensive reflection of the spatiotemporal information, such as the motion trajectory and spatial distribution of objects or people within the video. Simultaneously, a 1D-CNN model is used to extract temporal features, focusing on mining the patterns of change in the data over time and accurately capturing subtle changes in temporal characteristics. Furthermore, the surveillance video undergoes human pose estimation, group density calculation, and motion vector generation. These operations enable more detailed analysis of key information such as the behavioral state of individuals, the degree of group aggregation, and overall movement trends, providing richer and more accurate feature descriptions for risk assessment. By comprehensively utilizing these technologies, this method can deeply mine valuable information from multiple dimensions and levels when processing surveillance video data. Compared to existing technologies, it significantly improves the perception and identification accuracy of risk factors, resulting in more reliable and forward-looking real-time risk values. This effectively enhances the accuracy and timeliness of risk warnings, better meeting the high requirements for risk monitoring and early warning in real-world scenarios.
[0156] S170. Trigger the corresponding early warning strategy based on the real-time risk value and the set threshold.
[0157] In this embodiment, the early warning strategy includes processing strategies corresponding to three early warning levels: attention, warning, and high risk.
[0158] In this embodiment, at this stage, based on the real-time risk value and a pre-set threshold, the system decides whether to issue an alert and determines the specific alert level. The specific steps are as follows:
[0159] The calculated real-time risk value is compared with a set threshold. Different threshold ranges are set according to different risk levels. When the real-time risk value exceeds a certain threshold, the system will trigger the corresponding early warning strategy.
[0160] Based on the real-time risk value, the system will output corresponding early warning strategies, typically including three levels:
[0161] Attention Level: When the risk value is in a low range, the system may issue an alert to remind relevant personnel to pay attention to potential risks, but no immediate action is required.
[0162] Warning Level: When the risk value reaches a medium level, the system will issue a warning to remind relevant personnel to pay attention and that certain preventive measures may need to be taken.
[0163] High-risk level: When the risk value exceeds the set high-risk threshold, the system will issue a high-risk warning and may require immediate emergency measures to prevent an accident from occurring.
[0164] Through the above process, the system can perform intelligent analysis based on real-time perceived information (such as sensor data such as temperature and smoke concentration) and visual information (such as the location and category of entities detected by YOLOv5), thereby making rapid risk judgments and corresponding early warnings.
[0165] Please see Figure 4 The first step in the entire risk assessment process is to identify risk scenario factors, which mainly involves real-time monitoring and collection of various information within the scenario through a sensing system. This step includes two aspects:
[0166] Static information acquisition typically includes physical environment data of the scene, such as the locations of buildings, roads, equipment, and people. These fixed and time-invariant elements are acquired through various sensors (such as cameras and environmental sensors). This static information provides context for subsequent dynamic changes, helping to understand the layout and structure of the scene.
[0167] Dynamic factor collection mainly involves time-related data that changes with the scene. Dynamic factors include crowd density, movement trajectories, speed changes, and personnel behavior, which can be obtained through technologies such as real-time video surveillance systems, sensors, and drones. This data can reflect the changing trends of risks.
[0168] The core objective of this step is to collect comprehensive scenario information to provide sufficient raw data support for subsequent analysis, calculation, and early warning.
[0169] Next, vectorization is performed. Risk factor vectorization involves processing the static and dynamic factors obtained from the perception system to facilitate subsequent calculations and model input. This step specifically consists of the following two parts:
[0170] By identifying various entities in the scene (such as walls, roads, and equipment), these entities are transformed into vector forms. These vectors can contain features such as the entity's category, size, location, and density, facilitating subsequent analysis.
[0171] For dynamic risk factors, such as human movement, speed, and density, the data collected by sensors needs to be processed. These dynamic factors are extracted using specific algorithms (such as human posture recognition and motion trajectory analysis) and transformed into numerical vectors. These vectors can effectively describe dynamic characteristics such as human behavior and movement trends, and help predict potential risk events.
[0172] By vectorizing static and dynamic factors and combining them with intelligent analysis methods, the system can identify potential risks in a scenario. For example, it can detect stampedes that may occur in densely populated areas, or safety accidents caused by equipment malfunctions in certain areas. Through analysis of these factors, the system can identify and assess potential risk events.
[0173] Next, pedestrian risk calculation is performed, which is a process of in-depth analysis and calculation of pedestrian risk in the scene using an improved low-rank attention mechanism (MLA) network. By integrating multimodal data (such as video surveillance data and sensor data), the system can assess the risk level in the current scene in real time.
[0174] The improved low-rank attention mechanism, by introducing cross-modal attention computation, can find key correlations between spatiotemporal and temporal features, and improve the model's computational accuracy and efficiency through dynamic strategy adjustment. This mechanism can not only handle static scenes, but also reflect the risks brought about by dynamic changes in real time, such as changes in pedestrian density and the evolution of congestion.
[0175] By inputting preprocessed and extracted features into the MLA module, the system can generate cross-modal attention maps and use the information in these maps to calculate real-time pedestrian flow risk values. These risk values reflect the risk level of potential safety events such as pedestrian congestion and accidents in the current scenario.
[0176] Finally, the system outputs the pedestrian flow risk. This risk output is based on a calculated risk value. The system performs a risk assessment and triggers the corresponding early warning mechanism. The risk assessment uses a tiered strategy; based on different risk levels, the system will output the following three early warning levels:
[0177] Attention: This level of risk is low. There may be some minor hazards or uncertainties, but they will not immediately threaten personnel safety. This level of alert may remind management to monitor the scene and take appropriate safety precautions.
[0178] Warning: This level indicates a medium risk. Certain factors in the scenario (such as increased crowd density, abnormal behavior, etc.) may lead to dangerous events. This level of warning requires management to take timely intervention measures, such as strengthening scene supervision and crowd control.
[0179] High Risk: A high risk level means that the current scenario already presents a high level of safety risk, and a serious safety incident may occur. In this case, emergency response measures must be activated immediately, people rapidly evacuated, and emergency measures taken to prevent an accident from happening.
[0180] These warning levels will directly influence subsequent decisions and action plans, helping managers to respond quickly and effectively, and reduce potential risks that could cause losses to personnel and property.
[0181] By acquiring static and dynamic information from risk scenarios, vectorizing risk factors and identifying potential risks, and combining advanced computing models to perform real-time calculations on pedestrian flow risks and output risk warnings, the efficiency and accuracy of risk management and emergency response can be effectively improved.
[0182] The improved low-rank attention mechanism-based risk calculation method described above optimizes the risk calculation process by improving the low-rank attention mechanism module and combining multimodal information from surveillance video and sensor data. First, through preprocessing and feature extraction, noise is effectively removed and key features are extracted. Then, an improved MLA module is used to generate a cross-modal attention map, followed by cross-modal fusion feature calculation. Finally, the fused features are input into the GRU network to calculate the risk value in real time and trigger an early warning strategy based on a set threshold. This method reduces the dependence of traditional methods on complex network structures and large amounts of data processing through efficient data fusion and calculation, significantly improving the computational efficiency of risk monitoring while reducing the cost of network deployment and maintenance.
[0183] Figure 5 This is a schematic block diagram of a risk calculation system 300 with an improved low-rank attention mechanism provided in an embodiment of the present invention. Figure 5 As shown, corresponding to the above-described improved low-rank attention mechanism risk calculation method, the present invention also provides an improved low-rank attention mechanism risk calculation system 300. This improved low-rank attention mechanism risk calculation system 300 includes a unit for executing the above-described improved low-rank attention mechanism risk calculation method, and the system can be configured in a server. Specifically, please refer to... Figure 6 The risk calculation system 300 with the improved low-rank attention mechanism includes a data acquisition unit 301, a preprocessing unit 302, a feature extraction unit 303, an attention map generation unit 304, a feature determination unit 305, a calculation unit 306, and an early warning unit 307.
[0184] The data acquisition unit 301 is used to acquire monitoring video and sensor data of the scene to be monitored to obtain initial data; the preprocessing unit 302 is used to preprocess the initial data to obtain preprocessing results; the feature extraction unit 303 is used to extract features from the preprocessing results to obtain extraction results; the attention map generation unit 304 is used to input the extraction results into the improved MLA module to generate a cross-modal attention map; the feature determination unit 305 is used to determine cross-modal fusion features based on the cross-modal attention map; the calculation unit 306 is used to input the cross-modal fusion features into the GRU network to calculate the real-time risk value; and the early warning unit 307 is used to trigger the corresponding early warning strategy based on the real-time risk value and the set threshold.
[0185] In one embodiment, the preprocessing unit 302 is used to extract image frames from the initial data, perform human pose estimation, calculate group density and motion vectors, and normalize the sensor data to obtain preprocessing results.
[0186] In one embodiment, the feature extraction unit 303 is used to extract spatiotemporal features from the preprocessing result using a 3D-CNN model and extract temporal features from the preprocessing result using a 1D-CNN model, wherein the extraction result includes spatiotemporal features and temporal features.
[0187] In one embodiment, such as Figure 6 As shown, the attention map generation unit 304 includes a weighted value determination subunit 3041, a processing subunit 3042, and a weighted average subunit 3043.
[0188] The weighting value determination subunit 3041 is used to input the extraction result into the improved MLA module, calculate the attention value of the spatiotemporal feature and the temporal feature through the cross-modal attention calculation function, and scale it through the exponential function to obtain the weighting value; the processing subunit 3042 is used to normalize the weighting value to obtain the attention weight; the weighted averaging subunit 3043 is used to perform a weighted average of the spatiotemporal feature and the temporal feature according to the attention weight to obtain the cross-modal attention map.
[0189] In one embodiment, the feature determination unit 305 is used to detect entity-related information based on the preprocessing result, and to determine relevant features by using the embedding layer as a vector and the cross-modal attention map, and to concatenate the vector with the relevant features to obtain cross-modal fusion features.
[0190] In one embodiment, the computing unit 306 is used to input the cross-modal fusion features into the GRU network to generate hidden layer states, which are then processed through a fully connected layer and an activation function to obtain real-time risk values.
[0191] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the risk calculation system 300 and each unit of the improved low-rank attention mechanism described above can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0192] The aforementioned improved low-rank attention mechanism risk calculation system 300 can be implemented as a computer program, which can, for example... Figure 7 It runs on the computer device shown.
[0193] Please see Figure 7 , Figure 7 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0194] See Figure 7 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0195] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform a risk calculation method based on an improved low-rank attention mechanism.
[0196] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0197] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a risk calculation method with an improved low-rank attention mechanism.
[0198] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0199] The processor 502 is used to run a computer program 5032 stored in the memory to perform the following steps:
[0200] The system acquires monitoring video and sensor data of the scene to be monitored to obtain initial data; preprocesses the initial data to obtain preprocessing results; extracts features from the preprocessing results to obtain extraction results; inputs the extraction results into an improved MLA module to generate a cross-modal attention map; determines cross-modal fusion features based on the cross-modal attention map; inputs the cross-modal fusion features into a GRU network to calculate a real-time risk value; and triggers a corresponding early warning strategy based on the real-time risk value and a set threshold.
[0201] The early warning strategy includes the handling strategies corresponding to three early warning levels: attention, warning, and high risk.
[0202] In one embodiment, when the processor 502 performs the step of preprocessing the initial data to obtain a preprocessing result, it specifically implements the following steps:
[0203] Image frames are extracted from the initial data, human pose estimation is performed, group density and motion vectors are calculated, and the sensor data is normalized to obtain preprocessed results.
[0204] In one embodiment, when the processor 502 performs feature extraction on the preprocessed result to obtain the extraction result, it specifically implements the following steps:
[0205] The spatiotemporal features in the preprocessed results are extracted using a 3D-CNN model, and the temporal features in the preprocessed results are extracted using a 1D-CNN model. The extracted results include both spatiotemporal features and temporal features.
[0206] In one embodiment, when the processor 502 implements the step of inputting the extracted result into the improved MLA module to generate a cross-modal attention map, the following steps are specifically implemented:
[0207] The extraction results are input into the improved MLA module, where the attention values of the spatiotemporal features and the temporal features are calculated using a cross-modal attention calculation function, and scaled using an exponential function to obtain weighted values. The weighted values are then normalized to obtain attention weights. Based on the attention weights, the spatiotemporal features and the temporal features are weighted and averaged to obtain a cross-modal attention map.
[0208] Specifically, when the extraction results are input into the improved MLA module to generate a cross-modal attention map, a dynamic rank selection strategy and a nonlinear low-rank decomposition strategy are adopted to adjust according to the complexity and nonlinear relationship of the features, and a multi-head low-rank attention is adopted to generate a cross-modal attention map.
[0209] The dynamic rank selection strategy includes dynamically adjusting the rank size in the low-rank decomposition based on the entropy value of the input features; the nonlinear low-rank decomposition strategy includes nonlinearly enhancing the traditional low-rank decomposition by introducing the ReLU activation function; the multi-head low-rank attention includes computing multiple low-rank attention heads in parallel and concatenating the results of the low-rank attention heads to generate a cross-modal attention map.
[0210] In one embodiment, when implementing the step of determining cross-modal fusion features based on the cross-modal attention map, the processor 502 specifically implements the following steps:
[0211] Entity-related information is detected based on the preprocessing results, and the vector is embedded through the embedding layer. Relevant features are determined through the cross-modal attention map, and the vector is concatenated with the relevant features to obtain cross-modal fusion features.
[0212] In one embodiment, when processor 502 implements the step of inputting the cross-modal fusion features into the GRU network to calculate the real-time risk value, it specifically implements the following steps:
[0213] The cross-modal fusion features are input into the GRU network to generate hidden layer states. These states are then processed through a fully connected layer and an activation function to obtain real-time risk values.
[0214] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0215] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0216] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform the following steps:
[0217] The system acquires monitoring video and sensor data of the scene to be monitored to obtain initial data; preprocesses the initial data to obtain preprocessing results; extracts features from the preprocessing results to obtain extraction results; inputs the extraction results into an improved MLA module to generate a cross-modal attention map; determines cross-modal fusion features based on the cross-modal attention map; inputs the cross-modal fusion features into a GRU network to calculate a real-time risk value; and triggers a corresponding early warning strategy based on the real-time risk value and a set threshold.
[0218] The early warning strategy includes the handling strategies corresponding to three early warning levels: attention, warning, and high risk.
[0219] In one embodiment, when the processor executes the computer program to perform the step of preprocessing the initial data to obtain a preprocessing result, it specifically implements the following steps:
[0220] Image frames are extracted from the initial data, human pose estimation is performed, group density and motion vectors are calculated, and the sensor data is normalized to obtain preprocessed results.
[0221] In one embodiment, when the processor executes the computer program to perform feature extraction on the preprocessed result to obtain the extraction result, it specifically implements the following steps:
[0222] The spatiotemporal features in the preprocessed results are extracted using a 3D-CNN model, and the temporal features in the preprocessed results are extracted using a 1D-CNN model. The extracted results include both spatiotemporal features and temporal features.
[0223] In one embodiment, when the processor executes the computer program to implement the step of inputting the extracted results into the improved MLA module to generate a cross-modal attention map, it specifically implements the following steps:
[0224] The extraction results are input into the improved MLA module, where the attention values of the spatiotemporal features and the temporal features are calculated using a cross-modal attention calculation function, and scaled using an exponential function to obtain weighted values. The weighted values are then normalized to obtain attention weights. Based on the attention weights, the spatiotemporal features and the temporal features are weighted and averaged to obtain a cross-modal attention map.
[0225] Specifically, when the extraction results are input into the improved MLA module to generate a cross-modal attention map, a dynamic rank selection strategy and a nonlinear low-rank decomposition strategy are adopted to adjust according to the complexity and nonlinear relationship of the features, and a multi-head low-rank attention is adopted to generate a cross-modal attention map.
[0226] The dynamic rank selection strategy includes dynamically adjusting the rank size in the low-rank decomposition based on the entropy value of the input features; the nonlinear low-rank decomposition strategy includes nonlinearly enhancing the traditional low-rank decomposition by introducing the ReLU activation function; the multi-head low-rank attention includes computing multiple low-rank attention heads in parallel and concatenating the results of the low-rank attention heads to generate a cross-modal attention map.
[0227] In one embodiment, when the processor executes the computer program to implement the step of determining cross-modal fusion features based on the cross-modal attention map, it specifically implements the following steps:
[0228] Entity-related information is detected based on the preprocessing results, and the vector is embedded through the embedding layer. Relevant features are determined through the cross-modal attention map, and the vector is concatenated with the relevant features to obtain cross-modal fusion features.
[0229] In one embodiment, when the processor executes the computer program to implement the step of inputting the cross-modal fusion features into the GRU network to calculate the real-time risk value, it specifically implements the following steps:
[0230] The cross-modal fusion features are input into the GRU network to generate hidden layer states. These states are then processed through a fully connected layer and an activation function to obtain real-time risk values.
[0231] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0232] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0233] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each unit is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0234] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the system of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0235] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0236] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A risk calculation method for an improved low-rank attention mechanism, characterized in that, include: Acquire monitoring video and sensor data of the scene to be monitored to obtain initial data; The initial data is preprocessed to obtain the preprocessing result; Feature extraction is performed on the preprocessing results to obtain the extraction results; The extracted results are input into the improved MLA module to generate a cross-modal attention map; wherein, the improved MLA module integrates multimodal spatiotemporal features and temporal features through dynamic rank adjustment, nonlinear low-rank decomposition and multi-head low-rank attention mechanism; Cross-modal fusion features are determined based on the cross-modal attention map; The cross-modal fusion features are input into the GRU network to calculate the real-time risk value; The corresponding early warning strategy is triggered based on the real-time risk value and the set threshold.
2. The risk calculation method for an improved low-rank attention mechanism according to claim 1, characterized in that, The preprocessing of the initial data to obtain the preprocessing result includes: Image frames are extracted from the initial data, human pose estimation is performed, group density and motion vectors are calculated, and the sensor data is normalized to obtain preprocessed results.
3. The risk calculation method for an improved low-rank attention mechanism according to claim 1, characterized in that, The step of extracting features from the preprocessed results to obtain the extraction results includes: The spatiotemporal features in the preprocessed results are extracted using a 3D-CNN model, and the temporal features in the preprocessed results are extracted using a 1D-CNN model. The extracted results include both spatiotemporal features and temporal features.
4. The risk calculation method for an improved low-rank attention mechanism according to claim 1, characterized in that, The step of inputting the extracted results into the improved MLA module to generate a cross-modal attention map includes: The extraction results are input into the improved MLA module, where the attention values of the spatiotemporal features and the temporal features are calculated using a cross-modal attention calculation function, and then scaled using an exponential function to obtain a weighted value. The weighted values are normalized to obtain the attention weights; The spatiotemporal features and the temporal features are weighted and averaged according to the attention weights to obtain a cross-modal attention map.
5. The risk calculation method for an improved low-rank attention mechanism according to claim 4, characterized in that, The extracted results are input into the improved MLA module to generate a cross-modal attention map. A dynamic rank selection strategy and a nonlinear low-rank decomposition strategy are used to adjust the cross-modal attention map according to the complexity and nonlinear relationship of the features. Multi-head low-rank attention is used to generate the cross-modal attention map.
6. The risk calculation method for an improved low-rank attention mechanism according to claim 5, characterized in that, The dynamic rank selection strategy includes dynamically adjusting the rank size in the low-rank decomposition based on the entropy value of the input features; the nonlinear low-rank decomposition strategy includes nonlinearly enhancing the traditional low-rank decomposition by introducing the ReLU activation function; the multi-head low-rank attention includes computing multiple low-rank attention heads in parallel and concatenating the results of the low-rank attention heads to generate a cross-modal attention map.
7. The risk calculation method for an improved low-rank attention mechanism according to claim 3, characterized in that, The step of determining cross-modal fusion features based on the cross-modal attention map includes: Entity-related information is detected based on the preprocessing results, and the vector is embedded through the embedding layer. Relevant features are determined through the cross-modal attention map, and the vector is concatenated with the relevant features to obtain cross-modal fusion features.
8. The risk calculation method for an improved low-rank attention mechanism according to claim 7, characterized in that, The step of inputting the cross-modal fusion features into the GRU network to calculate the real-time risk value includes: The cross-modal fusion features are input into the GRU network to generate hidden layer states. These states are then processed through a fully connected layer and an activation function to obtain real-time risk values.
9. The risk calculation method for an improved low-rank attention mechanism according to claim 1, characterized in that, The early warning strategy includes the handling strategies corresponding to three early warning levels: attention, warning, and high risk.
10. A risk calculation system with an improved low-rank attention mechanism, characterized in that, include: The data acquisition unit is used to acquire monitoring video and sensor data of the scene to be monitored in order to obtain initial data; A preprocessing unit is used to preprocess the initial data to obtain a preprocessing result; A feature extraction unit is used to extract features from the preprocessing results to obtain the extraction results; An attention map generation unit is used to input the extraction results into the improved MLA module to generate a cross-modal attention map; A feature determination unit is used to determine cross-modal fusion features based on the cross-modal attention map; The calculation unit is used to input the cross-modal fusion features into the GRU network to calculate the real-time risk value; The early warning unit is used to trigger the corresponding early warning strategy based on the real-time risk value and the set threshold.
Citation Information
Patent Citations
Multi-modal medical image classification method based on structure tensor decomposition neural network
CN117788933A
Log anomaly detection method based on efficient fine tuning of adaptive low-rank parameters
CN118260689A
Risk early warning method and device constructed based on dynamic rule, equipment and medium
CN119671294A