Museum monitoring visualization method and device based on AI vision and medium

By collecting and processing multimodal data to generate a global event risk matrix, the problem of insufficient dynamic risk perception in museum monitoring systems has been solved, realizing the transformation from passive alarm to proactive early warning and improving the adaptability and response efficiency of the monitoring system.

CN121482268APending Publication Date: 2026-02-06LIGHT OF THE EARTH MUSEUM OPERATIONS & MANAGEMENT (WUXI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511636795.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing museum monitoring systems suffer from insufficient dynamic risk perception across the entire domain and lagging monitoring response in complex scenarios. They are unable to effectively achieve deep clustering and behavioral feature fusion of multimodal perception data, have weak cross-camera behavioral modeling, and struggle to predict abnormal group events and perform adaptive rendering and priority scheduling.

Method used

Multimodal perception data is collected and preprocessed to generate comprehensive behavioral features. Spatiotemporal fusion and causal reasoning are performed through cross-camera behavior models to output a global event risk matrix. Adaptive visualization scene map generation and dynamic priority allocation are performed, response instruction sets are generated and multi-dimensional hierarchical scheduling is performed, and finally, a museum security management strategy is generated.

Benefits of technology

It achieves a deep understanding of social dynamics and event chains in the monitoring scenario, transforming passive alarms after the fact into proactive early warnings before the event, automatically discovering blind spots in security strategies and conducting root cause analysis, thereby improving the real-time performance and effectiveness of museum security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482268A_ABST
    Figure CN121482268A_ABST
Patent Text Reader

Abstract

The invention discloses a museum monitoring visualization method and device based on AI vision and a medium, and relates to the technical field of AI vision, and the method comprises the steps: carrying out the space-time semantic remapping and hierarchical rendering of a global event risk matrix, and generating a self-adaptive visual scene map; joint reasoning and dynamic priority distribution are carried out on the self-adaptive visual scene map and the global event risk matrix, and a response instruction set is output; executing the response instruction set, performing multi-dimensional hierarchical scheduling, generating alarm execution mapping, performing time sequence tracking and risk response on the alarm execution mapping, and generating intervention feedback data; and performing spatio-temporal behavior pattern analysis and abnormal behavior path reconstruction on the intervention feedback data to generate a museum safety management strategy. According to the method, the cross-lens behavior model is constructed, and the museum safety management strategy is synchronously generated, so that deep understanding of social dynamics and event chains in a monitoring scene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of AI vision technology, and in particular to a method, device and medium for visualizing museum monitoring based on AI vision. Background Technology

[0002] Museums, as important venues for cultural heritage protection and public education, need to address multiple risks in complex environments through security monitoring, such as artifact safety, visitor behavior management, and emergency response. With the evolution of security technology, visual surveillance has gradually developed from early analog closed-circuit television to digital and networked intelligent video surveillance. In recent years, breakthroughs in artificial intelligence technology, especially computer vision, have driven the intelligentization of surveillance technology. For example, deep learning-based object detection algorithms can achieve real-time recognition of people and objects, while behavior analysis models can perform preliminary classification of specific actions. Multimodal perception technology is beginning to integrate visual, sound, environmental sensors, and positioning data, improving the comprehensiveness of scene understanding through data fusion methods.

[0003] However, existing technologies still have shortcomings in complex scenarios such as museums. Current methods mostly rely on independent processing of visual data, failing to achieve deep clustering and behavioral feature fusion of multimodal perception data. Existing methods are weak in cross-camera behavior modeling, often employing simple trajectory association or rule-based reasoning based on overlapping view domains, lacking spatiotemporal semantic embedding and causal reasoning mechanisms. This limits the ability to dynamically assess risk events across the entire domain; for example, it cannot predict abnormal group events through cross-view domain target interaction modeling, and the visualization is often a static layer, making adaptive rendering and priority scheduling based on risk levels difficult, resulting in delayed monitoring responses. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides an AI-based visual method for museum monitoring to address the problems of insufficient perception of dynamic risks across the entire area and delayed monitoring response.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a museum monitoring visualization method based on AI vision, comprising,

[0008] Collect and preprocess multimodal sensing data, perform cluster analysis and behavior detection on the preprocessed multimodal sensing data, and generate comprehensive behavioral features;

[0009] The comprehensive behavior characteristics are input into a cross-lens behavior model, a time-space fusion layer performs time sequence synchronization and space semantic embedding mapping, a correlation reasoning layer performs cross-view target interaction modeling and causal reasoning, and a global event risk matrix is output;

[0010] The global event risk matrix is subjected to time-space semantic remapping and hierarchical rendering to generate an adaptive visual scene graph;

[0011] The adaptive visual scene graph and the global event risk matrix are subjected to joint reasoning and dynamic priority allocation to output a response instruction set;

[0012] The response instruction set is executed and subjected to multi-dimensional hierarchical scheduling to generate an alarm execution mapping, and the alarm execution mapping is subjected to time sequence tracking and risk response to generate intervention feedback data;

[0013] The intervention feedback data is subjected to time-space behavior pattern analysis and abnormal behavior path reconstruction to generate a museum safety management strategy.

[0014] As a preferred scheme of the museum monitoring visualization method based on AI vision, the multi-modal perception data includes visual data, environmental data, sound data and visitor position information.

[0015] The preprocessing includes data cleaning, noise removal, time synchronization, format standardization, data filling and normalization.

[0016] As a preferred scheme of the museum monitoring visualization method based on AI vision, the multi-modal perception data includes visual data, environmental data, sound data and visitor position information.

[0017] The visual data after preprocessing is subjected to target detection and behavior recognition to generate visual behavior characteristics.

[0018] The sound data after preprocessing is subjected to sound event detection and clustering analysis to generate sound behavior characteristics.

[0019] The environmental data after preprocessing is subjected to pattern analysis clustering to generate environmental behavior characteristics.

[0020] The visitor position information after preprocessing is subjected to trajectory analysis to generate visitor behavior pattern characteristics.

[0021] The visual behavior characteristics, sound behavior characteristics, environmental behavior characteristics and visitor behavior pattern characteristics are structured and integrated to generate comprehensive behavior characteristics.

[0022] As a preferred scheme of the museum monitoring visualization method based on AI vision, wherein: the comprehensive behavior characteristics are input into the cross-lens behavior model, the spatio-temporal fusion layer performs time sequence synchronization and spatial semantic embedding mapping, the correlation reasoning layer performs cross-view target interaction modeling and causal reasoning, and a global event risk matrix is output, the specific steps are as follows,

[0023] The spatio-temporal fusion layer is built using a spatio-temporal convolution network, and the correlation reasoning layer is built using a graph neural network.

[0024] The cross-lens behavior model is constructed through cross-layer feature fusion and hierarchical stacking of the spatio-temporal fusion layer and the correlation reasoning layer by using a skip connection.

[0025] The comprehensive behavior characteristics are input into the cross-lens behavior model, the spatio-temporal fusion layer performs time sequence synchronization and spatial semantic embedding mapping on the comprehensive behavior characteristics through spatio-temporal convolution to generate spatio-temporal behavior characteristics.

[0026] The correlation reasoning layer performs target interaction modeling and causal reasoning on the spatio-temporal behavior characteristics to generate a cross-view behavior pattern.

[0027] The cross-view behavior pattern is associated analyzed by using a causal reasoning algorithm to obtain an abnormal behavior pattern.

[0028] The abnormal behavior pattern is analyzed by using a multi-level risk assessment algorithm to output a global event risk matrix.

[0029] As a preferred scheme of the museum monitoring visualization method based on AI vision, wherein: the global event risk matrix is subjected to spatio-temporal semantic remapping and hierarchical rendering to generate an adaptive visualization scene atlas, the specific steps are as follows,

[0030] The global event risk matrix is subjected to spatio-temporal semantic remapping to generate a spatio-temporal embedding dataset.

[0031] The spatio-temporal embedding dataset is subjected to risk level evaluation and region association to output a visualization layer.

[0032] The visualization layer is subjected to hierarchical rendering to generate an adaptive visualization scene atlas.

[0033] As a preferred scheme of the museum monitoring visualization method based on AI vision, wherein: the adaptive visualization scene atlas and the global event risk matrix are subjected to joint reasoning and dynamic priority allocation to output a response instruction set, the specific steps are as follows,

[0034] The adaptive visualization scene atlas and the global event risk matrix are subjected to joint reasoning and cross-domain association analysis to obtain cross-domain association analysis data.

[0035] Pattern recognition is performed on the cross-domain correlation analysis data, and a risk assessment data set is outputted;

[0036] The cross-domain correlation analysis data and the risk assessment data set are fused by attention, a risk feature data set is generated, the risk feature data set is prioritized, and a response instruction set is outputted.

[0037] As a preferred scheme of the AI vision-based museum monitoring visualization method, the response instruction set is executed and multi-dimensional hierarchical scheduling is performed, an alarm execution mapping is generated, time sequence tracking and risk response are performed on the alarm execution mapping, and intervention feedback data is generated.

[0038] The response instruction set is task decoupled and scheduling ordered, and a task scheduling scheme is generated.

[0039] The task scheduling scheme is prioritized and resource scheduled and path planned, and an alarm execution mapping is generated.

[0040] The alarm execution mapping is time sequence tracked, a time sequence tracking report is generated, and risk response and priority adjustment are performed on the time sequence tracking report, and intervention feedback data is generated.

[0041] As a preferred scheme of the AI vision-based museum monitoring visualization method, the intervention feedback data is time-space behavior pattern analyzed and abnormal behavior path reconstructed, and a museum safety management strategy is generated.

[0042] The intervention feedback data is time-space behavior pattern analyzed by using a moving average method, and time-space pattern data is generated.

[0043] The time-space behavior pattern data is abnormal behavior path reconstructed by using a dynamic time warping algorithm and a path deviation analysis method, and abnormal behavior path data is generated.

[0044] The abnormal behavior path data is time-space pattern recognized and behavior correlation analyzed, and a museum safety management strategy is generated.

[0045] In a second aspect, the present application provides a computer device, comprising a memory and a processor, the memory storing a computer program, wherein the computer program is executed by the processor to implement any step of the AI vision-based museum monitoring visualization method according to the first aspect of the present application.

[0046] In a third aspect, the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement any step of the AI vision-based museum monitoring visualization method according to the first aspect of the present application.

[0047] The beneficial effects of this invention are as follows: by constructing a cross-camera behavior model and generating a global event risk matrix, a deep understanding of social dynamics and event chains in the monitoring scene is achieved, realizing a key shift from post-event "passive alarm" to pre-event "proactive early warning"; at the same time, museum security management strategies are generated, which can automatically discover blind spots and ineffective links in existing security strategies and realize root cause analysis. Attached Figure Description

[0048] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Fig. 1 This is a flowchart of a museum monitoring visualization method based on AI vision.

[0050] Fig. 2 A flowchart for outputting the global event risk matrix.

[0051] Fig. 3 A flowchart generated for visualizing the scene.

[0052] Fig. 4 A flowchart generated in response to a command. Detailed Implementation

[0053] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0054] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0055] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0056] Reference Figs. 1-4 This is one embodiment of the present invention, which provides a museum monitoring visualization method based on AI vision, including the following steps:

[0057] S1, collect multi-modal perception data and preprocess, conduct clustering analysis and behavior detection on the preprocessed multi-modal perception data, and generate comprehensive behavior features;

[0058] S1.1, the multi-modal perception data includes visual data, environmental data, sound data and visitor position information;

[0059] It should be noted that the visual data is collected in real time by the camera installed in the museum, including video stream and image information; environmental data is collected by temperature and humidity sensors, light sensors, reflecting changes in temperature, humidity and light in the museum; sound data is collected by microphone devices to monitor sound changes and abnormalities in the environment; visitor position information is obtained through access control and smartphone positioning functions to track the flow trajectory of visitors in the museum in real time.

[0060] S1.2, preprocessing includes data cleaning, noise removal, time synchronization, format standardization, data filling and normalization;

[0061] It should be noted that the multi-modal perception data is cleaned to remove incomplete or abnormal records, filtered and denoised by noise removal algorithm to remove environmental interference and sensor errors; data from different sensors are aligned and time synchronized to ensure consistency of various data on the time axis; data from different sources are converted to a unified format, missing values are filled by linear interpolation to ensure data integrity, and all data are converted to a unified scale to eliminate the bias caused by different data magnitudes.

[0062] S1.3, target detection and behavior recognition are performed on the preprocessed visual data to generate visual behavior features;

[0063] It should be noted that edge detection and image segmentation are performed on the preprocessed visual data to extract moving targets such as objects and people in the scene, the motion trajectory of the moving target is obtained by comparing consecutive frames, the motion trajectory of each moving target is matched for behavior, including judging whether the moving target is stationary, whether it contacts other objects or moves quickly, etc., the matching results are integrated to form visual behavior features.

[0064] S1.4, detect sound events in the preprocessed sound data and conduct clustering analysis to generate sound behavior features;

[0065] It should be noted that the audio signal in the preprocessed sound data is decomposed into multiple frequency components using short-time Fourier transform, and different sound events such as falling object sound, conversation sound, door opening sound, etc. are identified according to the energy threshold; the K-means clustering algorithm is applied to group the sound events according to the event type, and the duration, frequency and intensity of each event type are obtained. Sound information such as sound information is integrated to generate sound behavior characteristics for describing sound behavior patterns in museum environment.

[0066] It should also be noted that the energy threshold is defined according to the energy distribution of background noise and the energy peak value of sound event, and the value range is (-50dB, -30dB), wherein -50dB is the average noise level obtained by long-term monitoring and statistics, and is usually used to represent the background noise in a relatively quiet environment, such as air conditioning or low-intensity mechanical noise; -30dB is suitable for background noise in a relatively noisy environment, such as crowd conversation, equipment operation, etc., which usually represents more obvious environmental noise.

[0067] S1.5, mode analysis and clustering of preprocessed environment data to generate environment behavior characteristics;

[0068] It should be noted that by fixing the time length of the environment window, the temperature, humidity, light and other environmental variables in the preprocessed environment data are divided into hourly segments to generate multi-dimensional environment time series, and the K-means clustering algorithm is applied to cluster similar environment states (such as high temperature, low light, humidity change, etc.) in the multi-dimensional environment time series together to obtain typical environment patterns (such as "high temperature, high humidity, strong light" and "low temperature, low humidity, no light"). The typical environment patterns are integrated according to the environment characteristics (including temperature and humidity change amplitude, light intensity) to generate environment behavior characteristic data.

[0069] S1.6, trajectory analysis of preprocessed visitor location information to generate visitor behavior pattern characteristics;

[0070] It should be noted that the preprocessed visitor location information is connected by coordinate points to connect the visitor's position at each time into a trajectory, and a visitor path map in the museum is constructed; by analyzing the path density of the visitor path map, the hot spot area where the visitor stays and the main path where the visitor moves are identified, and the main path is analyzed in time sequence to obtain the moving speed, stay time and path change of the visitor. Different behavior patterns (such as long stay, fast flow or reverse movement, etc.) are identified to generate visitor behavior pattern characteristics.

[0071] S1.7, structured integration of visual behavior characteristics, sound behavior characteristics, environment behavior characteristics and visitor behavior pattern characteristics to generate comprehensive behavior characteristics;

[0072] It should be noted that the visual behavior features, sound behavior features, environment behavior features and visitor behavior pattern features are time-aligned and feature-standardized, the data ranges of various features are unified, and the unified features are spliced according to the feature dimensions to generate comprehensive behavior features.

[0073] S2, input the comprehensive behavior features into the cross-lens behavior model, the spatio-temporal fusion layer performs time sequence synchronization and spatial semantic embedding mapping, the correlation reasoning layer performs cross-view target interaction modeling and causal reasoning, and outputs a global event risk matrix;

[0074] S2.1, use a spatio-temporal convolution network to build a spatio-temporal fusion layer, and use a graph neural network to build a correlation reasoning layer;

[0075] It should be noted that in the PyTorch framework, the spatio-temporal convolution network is called through nn.Conv3d (three-dimensional convolution layer) and initialized, the size of the convolution kernel is set according to the local dependence of the comprehensive behavior features in the spatio-temporal dimension, the size of the expansion factor is set according to the spatio-temporal distribution of the comprehensive behavior features, the non-linear feature extraction capability is improved through the ReLU (Rectified Linear) activation function, the global pooling layer is connected after the spatio-temporal convolution network for dimension reduction processing, and the residual connection is used to prevent the gradient vanishing problem, and the spatio-temporal fusion layer is built;

[0076] The graph neural network is called by torch_geometric (a popular library) to build the correlation reasoning layer. In specific operations, the graph convolution network is called through the GCNConv (graph convolution network layer) function, the parameters of the graph convolution network are initialized and regularized to prevent overfitting, the graph convolution is performed through the graph neural network, the interaction of the cross-lens target (i.e. the target included in the comprehensive behavior features) is modeled, the target behavior pattern under multiple views is captured, and the target correlation in the target behavior pattern under different views is enhanced through the attention mechanism to improve the accuracy of cross-lens data fusion.

[0077] It should also be noted that in the time dimension, the comprehensive behavior features are aligned according to the time stamp, the comprehensive behavior features are divided by using a sliding time window with a fixed time length, the behavior change correlation distribution in different time periods is obtained by obtaining the Pearson correlation coefficient of the comprehensive behavior features between adjacent sliding time windows, if the behavior change correlation distribution is concentrated in a short time (for example, about 3 time steps), it means that the local dependence is strong, and the convolution kernel time scale is set to a small size (for example, set to 3); if the correlation remains stable in a longer period of time (for example, 5 to 7 time steps), it means that the behavior change period is longer, and the convolution kernel time scale is set to a larger size (for example, set to 7);

[0078] In the spatial dimension, spatial similarity is obtained by calculating the Euclidean distance between the comprehensive behavioral features of adjacent regions. When the spatial similarity is high (i.e., the behavioral changes are gradual between neighborhoods), the spatial size of the convolution kernel is set to local (e.g., 3×3 region); when the spatial difference is large, the range of the convolution kernel is expanded to a wider region (e.g., 5×5 region) to capture global features.

[0079] The expansion factor is determined based on the decay rate of behavioral change correlation distribution and spatial similarity with distance. If behavioral change correlation distribution and spatial similarity decay slowly with increasing distance, it indicates that distant information still has a significant impact, and the expansion factor should be set larger to expand the receptive field (e.g., set to 5). If behavioral change correlation distribution and spatial similarity decay rapidly with increasing distance, it indicates that local features dominate, and the expansion factor should be set smaller (e.g., set to 2).

[0080] The expression for calculating the Pearson correlation coefficient is:

[0081] ;

[0082] in, This represents the Pearson correlation coefficient. Indicates the first Data points in variables The observed values ​​in Indicates the first Data points in variables The observed values ​​in Representing variables The mean, Representing variables The mean.

[0083] The expression for calculating Euclidean distance is:

[0084] ;

[0085] in, Representing a spatial point With spatial point The Euclidean distance between them Indicates the number of feature dimensions. Representing a spatial point In the Feature values ​​in each dimension Representing a spatial point In the Feature values ​​in each dimension.

[0086] S2.2. By using skip connections, cross-layer feature fusion and hierarchical stacking are performed on the spatiotemporal fusion layer and the related inference layer to construct a cross-camera behavior model;

[0087] It should be pointed out that the outputs of the spatio-temporal fusion layer and the association reasoning layer are spliced by skip connection to generate multi-dimensional fusion features, the multi-dimensional fusion features are nonlinearly transformed to enhance complexity, the multi-dimensional fusion features after nonlinear transformation are weighted by attention mechanism to generate dynamic fusion weights, and the spatio-temporal fusion layer and the association reasoning layer are hierarchically stacked according to the dynamic fusion weights to construct the cross-shot behavior model.

[0088] The historical comprehensive behavior features are divided into a training set, a validation set and a test set, wherein the training set is used for training the cross-shot behavior model, the validation set is used for real-time adjustment of hyperparameters, and the test set is used for evaluating the performance of the cross-shot behavior model. In the training process, the Adam optimizer is used for gradient descent optimization, and the learning rate scheduler is set to dynamically adjust the learning rate to avoid fast convergence in the later training. The cross-shot behavior model is trained to minimize the cross-entropy loss function, helping the cross-shot behavior model to reduce prediction errors by maximizing the probability of correct classification; the training set is trained in batches to obtain training loss and update network weights; after each training round, the cross-shot behavior model is evaluated with the validation set to check the training loss and accuracy, and then the hyperparameters are adjusted; when the training loss on the validation set does not decrease significantly for a number of rounds (such as 10 rounds), the training is stopped, and the trained cross-shot behavior model is output.

[0089] S2.3, input the comprehensive behavior features into the cross-shot behavior model, and the spatio-temporal fusion layer maps the comprehensive behavior features through time sequence synchronization and spatial semantic embedding to generate spatio-temporal behavior features;

[0090] It should be pointed out that the comprehensive behavior features are input into the cross-shot behavior model, and the spatio-temporal fusion layer maps the comprehensive behavior features through multi-dimensional convolution kernel to ensure the consistency of the comprehensive behavior features in time dimension and space dimension; the comprehensive behavior features of different spatial regions are mapped to a unified semantic space through spatial semantic embedding to generate spatio-temporal behavior features.

[0091] S2.4, the association reasoning layer models target interaction and causal reasoning on the spatio-temporal behavior features to generate cross-view behavior patterns;

[0092] It should be pointed out that in the association reasoning layer, the spatio-temporal behavior features are input into the graph convolution network, and the related behavior data of the cross-shot targets are represented as a graph structure through graph convolution, wherein each cross-shot target is regarded as a graph node in the graph structure, and the interaction relationship between the cross-shot targets is regarded as an edge; the graph convolution captures the dependency relationship between different cross-shot targets in the time and space dimensions through information propagation between graph nodes to generate cross-view behavior patterns.

[0093] S2.5, correlation analysis of cross-view behavior patterns is performed by applying a causal reasoning algorithm to obtain abnormal behavior patterns;

[0094] It should be noted that the correlation analysis of cross-view behavior patterns is performed by applying a causal reasoning algorithm to identify the time causal relationship between the related behaviors of the cross-camera target. Specifically, by analyzing the delay relationship in the spatio-temporal behavior characteristics (obtained by detecting whether one time series data can predict the change of another time series data), it is determined whether a behavior will have a forward influence on other behaviors, thereby inferring the potential causal chain, and further identifying the abnormal behavior patterns in the cross-view behavior patterns, such as abnormal behavior sequence or unexpected cross-camera target interaction.

[0095] S2.6, risk analysis of abnormal behavior patterns is performed by a multi-level risk assessment algorithm to output a global event risk matrix;

[0096] It should be noted that the risk scores of each abnormal behavior pattern are generated by weighted sum of the characteristics such as occurrence frequency, duration and location of the abnormal behavior pattern. The risk scores are mapped with the corresponding time stamp and spatial location based on the spatio-temporal information of each abnormal behavior pattern (such as the time period and spatial location of the event occurrence), forming a risk sample set. The risk sample set is interpolated and smoothed in the spatio-temporal dimension, and the correlation between the risk scores is obtained by statistically analyzing the change amplitude and direction of the risk scores of adjacent spatial regions in the same time sequence. The risk sample set after interpolation and smoothing is weighted averaged according to the correlation between the risk scores, to reduce the influence of local fluctuations, and a continuous risk distribution field is generated. The risk distribution field is discretized to form a global event risk matrix.

[0097] S3, spatio-temporal semantic remapping and hierarchical rendering of the global event risk matrix are performed to generate an adaptive visualization scene graph;

[0098] S3.1, spatio-temporal semantic remapping of the global event risk matrix is performed to generate a spatio-temporal embedding data set;

[0099] It should be noted that the spatio-temporal information of the cross-view behavior patterns is extracted from the global event risk matrix, including time stamp, spatial location and related behavior information, and the spatio-temporal information is converted into a unified spatio-temporal representation format. Linear interpolation method is used to fill the missing spatio-temporal information in the spatio-temporal information, and the spatial location is encoded to ensure the consistency of the spatial information. The filled spatio-temporal information is mapped to a standardized spatio-temporal grid to generate a spatio-temporal embedding data set.

[0100] S3.2, risk level evaluation and regional correlation of the spatio-temporal embedding data set are performed to output a visualization layer;

[0101] It should be pointed out that the time stamp and spatial location of each event in the spatio-temporal embedded dataset are analyzed to extract the spatio-temporal coordinates of each event, including the specific time and spatial location of the event occurrence, combined with the preset rule constraints, to dynamically assign a risk score to each event, and at the same time, according to the risk score, the risk level is divided, according to the spatial location and risk score of the event, the regions with similar risk scores in different regions are spatially clustered, and the risk level corresponding to each risk region is color labeled, wherein the high-risk region is marked with red to highlight the warning, the medium-risk region is marked with yellow to prompt attention, and the low-risk region is marked with green to represent the safe state; the risk region after spatial clustering and color labeling is superimposed into a unified spatial coordinate system to form a visualization layer.

[0102] It should also be pointed out that the preset rule constraints are defined based on the statistical distribution and historical data of events in museum monitoring. The rule constraints include the occurrence time, spatial location, duration of the event, and the degree of association with other behaviors. For example, certain events (such as touching exhibits, entering restricted areas, etc.) occurring in a specific time and specific spatial region will be given a higher risk score; while those events occurring with less interference and deviating from the expected behavior pattern will have a lower risk score.

[0103] According to the risk score of each event, it is divided into three levels, including first-level risk events (high risk), second-level risk events (medium risk), and third-level risk events (low risk); specifically, the risk score range of first-level risk events is set to [0.8, 1.0], the first-level risk events are marked as the highest priority events, which are usually related to high-risk behaviors such as exhibit touching and area intrusion; the risk score range of second-level risk events is set to [0.5, 0.8], the second-level risk events are medium-risk events, such as long-term stay in certain areas without causing direct threats; the risk score range of third-level risk events is set to [0, 0.5], the third-level risk events are marked as low-priority events, representing normal behavior patterns or minor abnormalities; wherein the division of the risk score range of the first-level risk events is based on the top 20% high segment (about 0.8 or above) corresponding to the events that have caused actual security risks or major disturbances in history, meaning that the occurrence of the event has a very high threat, such as exhibit damage, area intrusion, etc. major risk behaviors; the division of the risk score range of the second-level risk events is based on the middle 30% (about 0.5-0.8) corresponding to behaviors with potential threats but not causing direct damage, covering behaviors that pose a certain threat to museum safety but do not directly lead to serious consequences, such as long-term stay in restricted areas or crowd congestion; the division of the risk score range of the third-level risk events is based on the last 50% (less than 0.5) corresponding to normal or minor abnormal behaviors, such as short-term stay or minor sound interference.

[0104] S3.3, render the visualization layer in a hierarchical manner to generate an adaptive visualization scene graph;

[0105] It should be noted that the risk level and spatial distribution information of each region are extracted from the visualization layer, different rendering is applied according to different risk levels, for high-risk areas, more intense colors or highlighting effects are used to highlight, so that workers can quickly identify high-risk areas; for medium-risk areas, gradient colors or medium-intensity colors are used to display, reminding workers to pay attention; while low-risk areas use relatively dull colors to avoid information overload; according to the risk level and region association, the layers are rendered and synthesized layer by layer, the transparency and depth between different layers are adjusted, and the overall scene display effect is optimized to ensure that the visualization information of each region is complete and clear, and an adaptive visualization scene graph is output.

[0106] S4, joint reasoning and dynamic priority allocation of adaptive visualization scene graph and global event risk matrix, output response instruction set;

[0107] S4.1, joint reasoning and cross-domain correlation analysis of adaptive visualization scene graph and global event risk matrix, obtain cross-domain correlation analysis data;

[0108] It should be noted that the spatio-temporal data in the adaptive visualization scene graph and the risk score in the risk matrix are aligned and synchronized to construct a joint data table; in the joint data table, the risk intensity matrix of the region is generated by weighted average according to the risk score of the event, the time dependence relationship between events in the region risk intensity matrix is detected combined with the spatio-temporal position and time stamp of the event, and the behavior pattern similarity between regions is identified through clustering analysis, the region matching is performed according to the region correlation and behavior pattern similarity, and the potential event correlation pattern in different regions is mined to generate cross-domain correlation analysis data.

[0109] S4.2, pattern recognition of cross-domain correlation analysis data, output risk assessment data set;

[0110] It should be noted that the clustering algorithm is used to divide the space and time of the cross-domain correlation analysis data, the same events are classified through similarity measurement (Euclidean distance), the behavior patterns appearing in the same time period or the same region are identified, the labels of the behavior patterns are further constructed through the time stamp, spatial position and risk score of the behavior patterns, and abnormal patterns such as unusual behavior stay, regional intrusion or other behavior deviating from the conventional pattern are identified; according to the abnormal pattern and the corresponding risk score, a risk assessment data set is generated.

[0111] S4.3, attention fusion of cross-domain correlation analysis data and risk assessment data set, generate risk feature data set, and prioritize risk feature data set, output response instruction set.

[0112] It should be noted that the Pearson correlation coefficient between the cross-domain correlation analysis data and the risk assessment data set is calculated to determine the correlation between each feature and the aligned event, thereby assigning an attention weight to each feature, weighting and fusing the corresponding features in the cross-domain correlation analysis data and the risk assessment data set under the same event according to the attention weight, generating a risk feature data set, multiplying the risk score of each risk feature vector in the risk feature data set by the attention weight to obtain a risk priority score, and prioritizing according to the risk priority score to output a response instruction set.

[0113] S5, execute the response instruction set and perform multi-dimensional hierarchical scheduling, generate an alarm execution mapping, perform time sequence tracking and risk response on the alarm execution mapping, and generate intervention feedback data;

[0114] S5.1, task decoupling and scheduling of the response instruction set, generating a task scheduling scheme;

[0115] It should be noted that the response instruction set is decoupled into different operations such as monitoring adjustment, alarm activation, security personnel scheduling, etc., and each task is assigned a priority by analyzing the dependency relationship and execution time of each task in the response instruction set; the execution order of the tasks is determined by using topological sorting method to ensure that each task is executed after the preconditions are completed; according to the timeliness of task execution and the related priority information of the task, the scheduling order is optimized to reduce delay and improve response efficiency, generating a task scheduling scheme containing all task execution orders and resource allocation.

[0116] S5.2, priority resource scheduling and path planning of the task scheduling scheme, generating an alarm execution mapping;

[0117] It should be noted that according to the priority and resource demand of each task in the task scheduling scheme, resource allocation is performed to ensure that high-priority tasks can obtain critical resources first, while ensuring the fairness and efficiency of resource allocation; path planning algorithm is used to minimize the conflict and delay between tasks to obtain the best execution path of the task, including physical path (such as the shortest path of security personnel or robot to the event area) and information path (such as the transmission and execution link of instructions in the control device); the best execution path and resource allocation result are combined to generate a complete alarm execution mapping, which contains the execution order, required resources and selected best execution path of each task.

[0118] S5.3, timing tracking is performed on the alarm execution mapping, a timing tracking report is generated, and risk response and priority adjustment are performed on the timing tracking report to generate intervention feedback data.

[0119] It should be noted that according to the task execution order and timestamp information in the alarm execution mapping, the dynamic time warping algorithm is applied to track the execution process of each task, to ensure that the task is completed on time and to identify any execution delay or abnormality, and to generate a timing tracking report, wherein the timing tracking report details the actual execution time, execution status and possible delay of each task; risk response is performed on the timing tracking report in combination with risk scoring and task priority, the priority of high-risk tasks is adjusted in time to ensure that important tasks are executed first, and intervention feedback data is generated, wherein the intervention feedback data includes the adjusted execution plan and task status update, which is used to guide task execution.

[0120] S6, spatio-temporal behavior pattern analysis and abnormal behavior path reconstruction are performed on the intervention feedback data to generate a museum security management strategy.

[0121] S6.1, moving average method is used to analyze the spatio-temporal behavior pattern of the intervention feedback data, and spatio-temporal pattern data is generated;

[0122] It should be noted that a sliding window of fixed time length is set, the timestamp and spatial position of each event are extracted from the intervention feedback data and input into the sliding window, the mean value of the relevant data in the sliding window is obtained, short-term fluctuations are smoothed out and long-term trends are retained, thereby reducing noise, and behavior pattern data is generated; cluster analysis is performed on the behavior pattern data to identify behavior patterns that repeatedly occur at a certain spatial position within a certain time period (such as long-term stay near a certain exhibit, generating a potential behavior pattern of "high attention stay"), and spatio-temporal pattern data is generated.

[0123] S6.2, abnormal behavior path reconstruction is performed on the spatio-temporal behavior pattern data by dynamic time warping algorithm and path deviation analysis method, and abnormal behavior path data is generated;

[0124] It should be noted that the time series and spatial path information of each behavior are extracted from the spatio-temporal behavior pattern data, the behavior paths are integrated, the dynamic time warping algorithm is used to measure the similarity between the behavior paths by calculating the Pearson correlation coefficient, and the behavior paths in different time periods are aligned, the aligned spatio-temporal path data is generated, the deviation of the spatio-temporal path data from the historical normal behavior pattern is compared, and the obvious deviation part is identified as abnormal behavior path data.

[0125] S6.3, spatio-temporal pattern recognition and behavior correlation analysis are performed on the abnormal behavior path data to generate a museum security management strategy.

[0126] It should be noted that by performing spatio-temporal pattern recognition on the abnormal behavior path data, the time stamp, spatial position and path density of each abnormal behavior path are extracted, the dynamic time warping algorithm is applied to align the abnormal behavior path data of different time periods, the consistency of the spatio-temporal features is ensured and the error is reduced; the aligned abnormal behavior path data is classified by K-means, different types of abnormal behavior types are identified according to the Pearson correlation coefficient between the paths, the abnormal behavior types are matched with the risk scores, the risk score distribution characteristics (including the mean, variance and fluctuation trend of the risk score) associated with the abnormal behavior types are counted, and the risk score distribution characteristics are weighted and corrected according to the occurrence frequency, duration, spatial position of each type of abnormal behavior type, and the spatial overlap and time co-occurrence rate with the historical high-risk events (such as entering restricted areas, touching exhibits and interfering with equipment, etc.) in the museum historical data, and output the abnormal behavior risk level table; the abnormal behavior type, risk score distribution characteristics and abnormal behavior risk level table are integrated to develop a museum safety management strategy, including a monitoring optimization scheme for high-risk behavior types, a patrol scheduling adjustment and a real-time early warning strategy.

[0127] The embodiment also provides a computer device suitable for the AI vision-based museum monitoring visualization method, which comprises a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the AI vision-based museum monitoring visualization method proposed in the above embodiment.

[0128] The computer device can be a terminal, which comprises a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, operator network, NFC (near field communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0129] The embodiment also provides a storage medium on which a computer program is stored, the program being executed by a processor to implement the method for realizing museum monitoring visualization based on AI vision proposed in the above embodiment; the storage medium can be realized by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk.

[0130] In summary, the application achieves a deep understanding of social dynamics and event chains in the monitoring scene by constructing a cross-lens behavior model and generating a global event risk matrix, and achieves a key transition from post-event "passive alarm" to pre-event "active early warning"; a museum security management strategy is simultaneously generated, which can automatically find the blind spots and invalid links of the existing security strategy, and realizes traceable root cause analysis.

[0131] It should be noted that the above embodiments are only used to illustrate the technical solutions of the application rather than limit the application. Although the application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the application, and all of them should be covered in the scope of the claims of the application.

Claims

1. A museum monitoring visualization method based on AI vision, characterized in that: include, Collect and preprocess multimodal sensing data, perform cluster analysis and behavior detection on the preprocessed multimodal sensing data, and generate comprehensive behavioral features; The comprehensive behavioral features are input into the cross-camera behavior model. The spatiotemporal fusion layer performs temporal synchronization and spatial semantic embedding mapping, and the correlation reasoning layer performs cross-visual target interaction modeling and causal reasoning, outputting a global event risk matrix. Perform spatiotemporal semantic remapping and hierarchical rendering on the global event risk matrix to generate an adaptive visual scene map; Perform joint reasoning and dynamic priority allocation on the adaptive visualized scene map and the global event risk matrix, and output a response instruction set; Execute response instruction set and perform multi-dimensional hierarchical scheduling, generate alarm execution mapping, perform time-series tracking and risk response on alarm execution mapping, and generate intervention feedback data; Spatiotemporal behavioral pattern analysis and abnormal behavior path reconstruction are performed on intervention feedback data to generate museum safety management strategies.

2. The museum monitoring visualization method based on AI vision as described in claim 1, characterized in that: The multimodal perception data includes visual data, environmental data, sound data, and visitor location information; The preprocessing includes data cleaning, noise removal, time synchronization, format standardization, data filling, and normalization.

3. The AI-based vision-based museum monitoring visualization method as described in claim 2, characterized in that: The process of performing cluster analysis and behavior detection on the preprocessed multimodal sensing data to generate comprehensive behavioral features involves the following steps: Target detection and behavior recognition are performed on the preprocessed visual data to generate visual behavior features; Detect sound events in the preprocessed sound data and perform cluster analysis to generate sound behavior features; Pattern analysis and clustering are performed on the preprocessed environmental data to generate environmental behavior characteristics; Trajectory analysis is performed on the preprocessed visitor location information to generate visitor behavior pattern characteristics; Visual behavioral features, auditory behavioral features, environmental behavioral features, and visitor behavioral pattern features are structurally integrated to generate comprehensive behavioral features.

4. The AI-based vision-based museum monitoring visualization method as described in claim 3, characterized in that: The process involves inputting comprehensive behavioral features into a cross-camera behavior model, performing temporal synchronization and spatial semantic embedding mapping in a spatiotemporal fusion layer, and conducting cross-view target interaction modeling and causal inference in an association reasoning layer to output a global event risk matrix. The specific steps are as follows: A spatiotemporal fusion layer is built using a spatiotemporal convolutional network, and an association reasoning layer is built using a graph neural network. By using skip connections, cross-layer feature fusion and hierarchical stacking are performed on the spatiotemporal fusion layer and the correlation inference layer to construct a cross-shot behavior model; The comprehensive behavioral features are input into the cross-camera behavior model. The spatiotemporal fusion layer performs temporal synchronization and spatial semantic embedding mapping on the comprehensive behavioral features through spatiotemporal convolution to generate spatiotemporal behavioral features. The correlation reasoning layer performs target interaction modeling and causal reasoning on spatiotemporal behavioral features to generate cross-perspective behavioral patterns. Causal reasoning algorithms are applied to perform correlation analysis on cross-perspective behavioral patterns to obtain abnormal behavioral patterns; A multi-level risk assessment algorithm is used to analyze the risks of abnormal behavior patterns and output a global event risk matrix.

5. The AI-based vision-based museum monitoring visualization method as described in claim 4, characterized in that: The specific steps for performing spatiotemporal semantic remapping and hierarchical rendering on the global event risk matrix to generate an adaptive visual scene map are as follows. Perform spatiotemporal semantic remapping on the global event risk matrix to generate a spatiotemporal embedded dataset; Risk level assessment and regional association are performed on the spatiotemporal embedded dataset, and a visualization layer is output. Perform hierarchical rendering on the visualization layers to generate an adaptive visualization scene map.

6. The AI-based vision-based museum monitoring visualization method as described in claim 5, characterized in that: The steps for jointly reasoning and dynamically prioritizing the adaptive visualized scene map and the global event risk matrix, and outputting a response instruction set, are as follows: By performing joint reasoning and cross-domain correlation analysis on the adaptive visualization scene map and the global event risk matrix, cross-domain correlation analysis data can be obtained. Perform pattern recognition on cross-domain correlation analysis data and output a risk assessment dataset; Attention fusion is performed between cross-domain correlation analysis data and risk assessment dataset to generate risk feature dataset. The risk feature dataset is then prioritized and a response instruction set is output.

7. The AI-based vision-based museum monitoring visualization method as described in claim 6, characterized in that: The execution response instruction set is executed and multi-dimensional hierarchical scheduling is performed to generate alarm execution maps. These alarm execution maps are then subjected to time-series tracking and risk response to generate intervention feedback data. The specific steps are as follows: The response instruction set is decoupled and sorted for task scheduling to generate a task scheduling scheme. Perform priority resource scheduling and path planning on the task scheduling scheme, and generate alarm execution mapping; Perform time-series tracking on alarm execution mapping, generate time-series tracking reports, and adjust the risk response and priority of time-series tracking reports to generate intervention feedback data.

8. The AI-based vision-based museum monitoring visualization method as described in claim 7, characterized in that: The process of performing spatiotemporal behavioral pattern analysis and abnormal behavior path reconstruction on intervention feedback data to generate museum security management strategies involves the following specific steps. The moving average method was used to analyze the spatiotemporal behavioral patterns of the intervention feedback data, generating spatiotemporal pattern data. The abnormal behavior path data is reconstructed from the spatiotemporal behavior pattern data by using the dynamic time warping algorithm and the path deviation analysis method to generate abnormal behavior path data. Spatiotemporal pattern recognition and behavioral correlation analysis are performed on abnormal behavior path data to generate museum security management strategies.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the AI ​​vision-based museum monitoring visualization method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the AI ​​vision-based museum monitoring visualization method according to any one of claims 1 to 8.