Real-time monitoring system and method for safety of smart tourist attraction based on big data analysis

By building a security perception trust index and combining big data analysis and neural networks, the problem of mismatch between image risk recognition and coding adaptability in the existing smart cultural and tourism scenic area monitoring system was solved, achieving high accuracy and stability of the monitoring system.

CN120688000AInactive Publication Date: 2025-09-23NANJING HONGPO INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510751000.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing smart cultural and tourism scenic area security monitoring system fails to effectively combine the risk identification results of image content with the communication link status for adaptive adjustment, resulting in frequent false alarms and missed alarms, affecting the accuracy and real-time performance of the monitoring system.

Method used

Through a method based on big data analysis, using convolutional neural networks and particle swarm optimization algorithms, a security perception trust index is constructed, which dynamically integrates the image content risk recognition strength and coding adaptability to achieve real-time evaluation and coding transmission processing of monitoring images.

Benefits of technology

It improves the safety judgment accuracy and response stability of the monitoring system, ensures the consistency of image quality and risk identification, effectively suppresses false alarms and missed alarms, and improves the real-time and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688000A_ABST
    Figure CN120688000A_ABST
Patent Text Reader

Abstract

The invention discloses a big data analysis-based intelligent travel scenic spot safety real-time monitoring system and method, and relates to the technical field of scenic spot monitoring. According to the big data analysis-based safety real-time monitoring method for the smart text tourist attraction, monitoring image data and communication state data of each region of the text tourist attraction are synchronously acquired and analyzed, an image semantic bearing strength index and a link stability index are extracted respectively, and an image coding adaptation index is calculated based on a fusion result of the two indexes; guiding a coding transmission strategy of the image data, and carrying out coding transmission processing; and performing security assessment analysis on the monitoring image data of each region after the coding transmission processing to obtain a corresponding security perception credible index. According to the invention, corresponding security alarm is performed on each region of the text tourist attraction after the coding transmission processing based on the security perception credible index; therefore, the risk of misjudgment caused by poor image quality is suppressed, and the safety judgment accuracy and response stability of scenic spot monitoring are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of scenic area monitoring technology, and specifically to a smart cultural and tourism scenic area safety real-time monitoring system and method based on big data analysis. Background Art

[0002] With the continuous advancement of smart cultural tourism construction, scenic area security monitoring systems play an important role in improving tourist safety and optimizing management efficiency. At present, many scenic spots have deployed video surveillance cameras to obtain real-time image data from various areas, and identify security incidents through manual supervision or preliminary image processing technology. However, due to factors such as the complexity of the network environment, the variability of image content, and the singleness of the processing mechanism, the existing system still has many deficiencies in the process of information acquisition, transmission and judgment. For example, the existing technology usually adopts fixed encoding parameters or preset compression standards, and fails to adaptively adjust the codability of image content in combination with the current communication link status. It is easy to have problems such as excessive image compression, distortion of key content or congestion of high-quality image bandwidth, which affects the accuracy and real-time performance of the back-end analysis system.

[0003] Prior art, such as the patent application with announcement number: CN116528036B, discloses a real-time security monitoring system and method for scenic spots based on smart cultural tourism. The system includes an area division module, a shooting angle acquisition module, a shooting duration calculation module and a shooting module. The area division module is used to partition the monitoring area of ​​the rotatable surveillance camera, dividing the monitoring area into multiple shooting areas; the shooting angle acquisition module is used to obtain the shooting angle of each shooting area; the shooting duration calculation module is used to obtain the shooting duration of each shooting angle according to a preset calculation strategy; the shooting module is used to control the rotatable camera to shoot according to the shooting angle and shooting duration to obtain a monitoring video. The present invention also provides a method corresponding to the above system. The present invention greatly reduces the probability of the problem of the shooting time of the shooting area being too long or too short, and effectively improves the quality of safety monitoring of the scenic area.

[0004] Based on the above solution, it is found that the limitations of the existing technology include at least the following problems: the existing technology lacks a mechanism for fusion modeling between the risk identification results of image content and the codability of the image under the current communication link conditions. In actual monitoring applications, the content complexity of the image determines the security risk information it may carry, while the communication link status and the encoding adaptability of the image itself determine whether this information can be accurately extracted, compressed and transmitted to the back-end for analysis. However, a modeling structure that reflects the degree of match between the image expression ability and the content risk intensity has not been constructed. This may easily lead to false alarms and false positives even if the monitoring image identifies high-risk targets. If the image itself lacks clarity and readability due to excessive compression, noise interference or poor communication quality, it may still easily trigger false alarms. Conversely, when the image quality is good but the scene complexity is high and the model's risk identification ability is limited, it is also easy to ignore potential risks, resulting in missed reports. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a real-time security monitoring system and method for smart cultural and tourism scenic spots based on big data analysis, which solves the problems that the existing technology does not model the matching relationship between image risks and coding adaptability, is prone to false alarms and missed alarms, and is difficult to ensure the credibility of alarms.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a real-time security monitoring method for a smart cultural and tourism scenic area based on big data analysis, comprising the following steps: real-time acquisition of monitoring image data of several areas of the cultural and tourism scenic area, and input into a pre-trained content recognition model for comprehensive analysis to obtain the image semantic carrying strength index of each area of ​​the cultural and tourism scenic area; and acquisition of current communication status data of each area of ​​the cultural and tourism scenic area, and feature analysis to obtain the link stability index of each area of ​​the cultural and tourism scenic area; comprehensive analysis of the image semantic carrying strength index and link stability index of each area of ​​the cultural and tourism scenic area to obtain the image coding adaptation index of each area of ​​the cultural and tourism scenic area, and encoding and transmitting processing; security assessment analysis of the monitoring image data of each area of ​​the cultural and tourism scenic area after encoding and transmission processing to obtain the security perception trust index of each area of ​​the cultural and tourism scenic area after encoding and transmission processing; and corresponding security alarms for each area of ​​the cultural and tourism scenic area after encoding and transmission processing based on the security perception trust index.

[0007] Furthermore, the monitoring image data specifically includes the pixel value and two-dimensional coordinates of each pixel point, and the specific steps for obtaining the image semantic carrying strength index of each area of ​​the cultural and tourism scenic area are as follows: the monitoring image data of each area of ​​the cultural and tourism scenic area is input into the pre-trained content recognition model for evaluation and analysis to obtain the image content evaluation set of each area of ​​the cultural and tourism scenic area, including the semantic target density index, occlusion risk index, visual field structure hierarchy index, and high-risk component significance index; the monitoring frame content evaluation set of each area of ​​the cultural and tourism scenic area is comprehensively analyzed to obtain the image semantic carrying strength index of each area of ​​the cultural and tourism scenic area.

[0008] Furthermore, the content recognition model is specifically a convolutional neural network, which includes an input layer, a feature extraction layer, a semantic compression layer, a feature aggregation layer, and a prediction output layer. The specific steps for obtaining the image content evaluation set of each area of ​​the cultural and tourism scenic area are as follows: in the input layer of the convolutional neural network, the monitoring image data of each area of ​​the cultural and tourism scenic area is received and preprocessed; in the feature extraction layer of the convolutional neural network, the preprocessed monitoring image data of each area of ​​the cultural and tourism scenic area is feature encoded to obtain the initial image feature map of each area of ​​the cultural and tourism scenic area; in the semantic compression layer of the convolutional neural network, the initial image feature map of each area of ​​the cultural and tourism scenic area is convolutionally compressed to obtain the semantic feature map of each area of ​​the cultural and tourism scenic area; in the feature aggregation layer of the convolutional neural network, the semantic feature map of each area of ​​the cultural and tourism scenic area is globally averaged pooled to obtain the image structure expression vector of each area of ​​the cultural and tourism scenic area; in the prediction output layer of the convolutional neural network, the image structure expression vector of each area of ​​the cultural and tourism scenic area is predicted to obtain the semantic target density index, occlusion risk index, visual field structure hierarchy index, and high-risk component significance index of each area of ​​the cultural and tourism scenic area.

[0009] Furthermore, the specific formula for calculating the image semantic carrying strength index of a certain area of ​​a cultural and tourism scenic spot is as follows: Among them, TxC is the image semantic carrying intensity index of a certain area of ​​the cultural and tourism scenic area, YmB is the semantic target density index of a certain area of ​​the cultural and tourism scenic area, α1 is the target density adjustment coefficient stored in the database, SyG is the visual field structure hierarchy index of a certain area of ​​the cultural and tourism scenic area, α2 is the structure hierarchy adjustment coefficient stored in the database, GwX is the high-risk component significance index of a certain area of ​​the cultural and tourism scenic area, α3 is the high-risk adjustment coefficient stored in the database, ZdF is the occlusion risk index of a certain area of ​​the cultural and tourism scenic area, and α4 is the occlusion adjustment coefficient stored in the database.

[0010] Furthermore, the current communication status data includes link pileup value, frequency domain interference intensity index, signal-to-noise ratio value, communication thermal interference index, carrier frequency drift index, and power feedback residual value. The specific steps for obtaining the link stability index of each area of ​​the cultural and tourism scenic area are as follows: comprehensively analyze the current communication status data of each area of ​​the cultural and tourism scenic area to obtain a link evaluation set for each area of ​​the cultural and tourism scenic area, including a communication interference sensitivity index and a link smoothness index; comprehensively analyze the link evaluation set of each area of ​​the cultural and tourism scenic area based on the particle swarm optimization algorithm to obtain the link stability index of each area of ​​the cultural and tourism scenic area.

[0011] Furthermore, the specific steps for obtaining the link evaluation set of each area of ​​the cultural and tourism scenic area are as follows: based on the particle swarm optimization algorithm, a comprehensive analysis is performed on the frequency domain interference intensity index, communication thermal interference index, carrier frequency drift index, and power feedback residual value of each area of ​​the cultural and tourism scenic area to obtain the communication interference sensitivity index of each area of ​​the cultural and tourism scenic area; based on the particle swarm optimization algorithm, a comprehensive analysis is performed on the link stacking value and signal-to-noise ratio value of each area of ​​the cultural and tourism scenic area to obtain the link smoothness index of each area of ​​the cultural and tourism scenic area.

[0012] Furthermore, the specific steps for obtaining the security perception trust index of each area of ​​the cultural and tourism scenic area after the coding and transmission processing are as follows: input the monitoring image data of each area of ​​the cultural and tourism scenic area after the coding and transmission processing into the pre-trained security recognition model for evaluation and analysis, and obtain the security assessment set of each area of ​​the cultural and tourism scenic area after the coding and transmission processing, including the abnormal behavior index, the illegal intrusion index, and the scene component risk index; based on the entropy weight method, a comprehensive analysis is performed on the security assessment set of each area of ​​the cultural and tourism scenic area after the coding and transmission processing, and the risk identification strength index of each area of ​​the cultural and tourism scenic area after the coding and transmission processing is obtained; the image coding adaptation index of each area of ​​the cultural and tourism scenic area is read, and combined with the risk identification strength index of each area of ​​the cultural and tourism scenic area after the coding and transmission processing, a comprehensive analysis is performed to obtain the security perception trust index of each area of ​​the cultural and tourism scenic area after the coding and transmission processing.

[0013] Furthermore, the security identification model is specifically a risk identification neural network, which includes an image input layer, a unified extraction layer, a task classification layer, and a regression output layer. The specific steps for obtaining the security assessment set of each area of ​​the cultural and tourism scenic area after coding and transmission processing are as follows: in the image input layer of the risk identification neural network, the monitoring image data of each area of ​​the cultural and tourism scenic area after coding and transmission processing is received and preprocessed; in the unified extraction layer of the risk identification neural network, the monitoring image data of each area of ​​the cultural and tourism scenic area after preprocessing and coding and transmission processing is subjected to joint feature extraction processing to obtain the shared feature vector of each area of ​​the cultural and tourism scenic area after coding and transmission processing; in the task classification layer of the risk identification neural network, the shared feature vector of each area of ​​the cultural and tourism scenic area after coding and transmission processing is subjected to classification extraction processing to obtain the risk vector set of each area of ​​the cultural and tourism scenic area after coding and transmission processing; in the regression output layer of the risk identification neural network, the risk vector set of each area of ​​the cultural and tourism scenic area after coding and transmission processing is subjected to regression mapping processing respectively to obtain the abnormal behavior index, illegal intrusion index, and scene component risk index of each area of ​​the cultural and tourism scenic area after coding and transmission processing.

[0014] Furthermore, the specific formula for calculating the security perception trust index of a certain area of ​​the cultural and tourism scenic area after the coding transmission processing is as follows: TxC=TsY η1 *[1-tanh(η2*FxS)]*[1-exp(-η3*|TsY-FxS|)]; where TxC is the security perception credibility index of a certain area of ​​the cultural and tourism scenic area after coding and transmission processing, TsY is the image coding adaptation index of a certain area of ​​the cultural and tourism scenic area, η1 is the image coding adjustment coefficient stored in the database, FxS is the risk identification intensity index of a certain area of ​​the cultural and tourism scenic area after coding and transmission processing, η2 is the risk adjustment coefficient stored in the database, and η3 is the difference adjustment coefficient stored in the database.

[0015] The smart real-time security monitoring system for cultural and tourism scenic spots based on big data analysis includes: a data acquisition and analysis module, which is used to acquire monitoring image data of several areas of the cultural and tourism scenic spot in real time, and input it into a pre-trained content recognition model for comprehensive analysis to obtain the image semantic carrying strength index of each area of ​​the cultural and tourism scenic spot; a feature analysis module, which is used to acquire the current communication status data of each area of ​​the cultural and tourism scenic spot, and perform feature analysis to obtain the link stability index of each area of ​​the cultural and tourism scenic spot; a data coding and transmission module, which is used to perform comprehensive analysis on the image semantic carrying strength index and link stability index of each area of ​​the cultural and tourism scenic spot, obtain the image coding adaptation index of each area of ​​the cultural and tourism scenic spot, and perform coding and transmission processing; a security assessment and analysis module, which is used to perform security assessment and analysis on the monitoring image data of each area of ​​the cultural and tourism scenic spot after coding and transmission processing, and obtain the security perception trust index of each area of ​​the cultural and tourism scenic spot after coding and transmission processing; a security alarm feedback module, which is used to issue corresponding security alarms to each area of ​​the cultural and tourism scenic spot after coding and transmission processing based on the security perception trust index.

[0016] The present invention has the following beneficial effects:

[0017] (1) This method of real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis realizes the dynamic fusion modeling of the matching degree between the image content risk recognition intensity index and the image coding adaptation index by constructing a security perception trust index. This method not only extracts the risk features in the current monitoring image based on the neural network, but also quantitatively models the codability of the image under the current communication link to form a coding adaptation index. Based on the two, a trust index model is constructed to comprehensively evaluate whether the current image has effective expression capabilities and consistency with the real risks. It is finally used to drive the alarm trigger mechanism, thereby effectively suppressing the risk of misjudgment caused by poor image quality, and then ensuring that the response is triggered under the premise that the image expression is complete and the recognition result is trustworthy, and significantly improving the safety judgment accuracy and response stability of scenic spot monitoring.

[0018] (2) This method of real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis dynamically evaluates the optimal coding and transmission strategy of each frame of monitoring image under the current link environment through a dual-factor fusion calculation based on the image semantic carrying strength index and the communication link stability index. The image semantic carrying strength index extracts the semantic density, occlusion distribution, structural hierarchy and component significance of the image content through a deep convolutional neural network, while the communication link stability index obtains a comprehensive evaluation result based on parameters such as link stacking, frequency domain interference, and signal-to-noise ratio through particle swarm optimization. The combination of the two constitutes a unified evaluation of image readability and transmittability, thereby driving the subsequent use of high compression ratio coding or low-loss lossless coding mode, and effectively alleviating link pressure, thereby preventing high-value images from being damaged during transmission or irrelevant images from occupying resources.

[0019] (3) This method of real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis realizes the specialized identification and collaborative analysis of different safety risk factors in scenic spot monitoring images by constructing a safety risk identification network with a multi-model processing path with a clear structural division of labor. It also integrates multiple types of feature extraction strategies for personnel behavior, regional intrusion and component status within a unified architecture, and can extract respective features and output risk assessment results in a targeted manner according to different risk types in the image. At the same time, it forms an efficient and collaborative task division mechanism within the model, thereby ensuring that all types of risk information can be accurately identified, independently modeled and uniformly integrated. In addition, it can continuously and stably output high-reliability risk indicators in actual scenarios with high image complexity and many interference factors, thereby enhancing the targetedness, stability and interpretability of the identification results.

[0020] (4) The real-time safety monitoring system for smart cultural and tourism scenic spots based on big data analysis has built a modular monitoring system with a clear structure and clear division of tasks. The modules are connected in series in an orderly manner based on parameter output and logical coupling to form a complete state perception, processing, and response closed loop, thus having good real-time processing capabilities and state linkage. In particular, the feature analysis module can dynamically generate coding adaptation parameters based on the current communication link status, drive the coding strategy selection, and directly affect the processing results to the subsequent risk assessment and alarm judgment process, realizing the logical linkage of image value, transmission conditions, and risk credibility in the monitoring process, thereby effectively improving the real-time, robustness, and adaptability of the system in actual deployment.

[0021] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a flow chart of the real-time safety monitoring method for smart cultural and tourism scenic spots based on big data analysis of the present invention.

[0023] Figure 2 This is a schematic diagram of the semantic target density index area sequence of cultural and tourism scenic spots in the real-time safety monitoring method of smart cultural and tourism scenic spots based on big data analysis of the present invention.

[0024] Figure 3 This is a schematic diagram of the regional sequence of the visual field structure hierarchy index of a cultural tourism scenic spot in the real-time safety monitoring method of a smart cultural tourism scenic spot based on big data analysis of the present invention.

[0025] Figure 4 This is a schematic diagram of the significant index area sequence of high-risk components in the cultural and tourism scenic area in the smart cultural and tourism scenic area safety real-time monitoring method based on big data analysis of the present invention.

[0026] Figure 5 This is a schematic diagram of the occlusion risk index area sequence of a cultural and tourism scenic spot in the real-time safety monitoring method of a smart cultural and tourism scenic spot based on big data analysis of the present invention.

[0027] Figure 6 This is a flowchart of the specific steps of obtaining the security perception trust index of each area of ​​the cultural and tourism scenic area after encoding and transmission processing in the real-time security monitoring method of the smart cultural and tourism scenic area based on big data analysis of the present invention.

[0028] Figure 7 This is a block diagram of the real-time safety monitoring system for smart cultural and tourism scenic spots based on big data analysis of the present invention. DETAILED DESCRIPTION

[0029] See also Figure 1 , an embodiment of the present invention provides a technical solution: a smart real-time security monitoring method for cultural and tourism scenic spots based on big data analysis, comprising the following steps: real-time acquisition of monitoring image data (of the current monitoring frame) of several areas of the cultural and tourism scenic spot (a sequence of monitoring image frames continuously acquired by the current monitoring frame deployed in each area of ​​the cultural and tourism scenic spot. Because it is a real-time analysis, the image data of the current monitoring frame is analyzed here, and the analysis logic of each monitoring frame is consistent), and input it into a pre-trained content recognition model for comprehensive analysis to obtain the image semantic carrying strength index of each area of ​​the cultural and tourism scenic spot; and obtain the current communication status data of each area of ​​the cultural and tourism scenic spot (the communication link where the monitoring acquisition node of the current monitoring frame is located), and perform feature analysis to obtain the link stability index of each area of ​​the cultural and tourism scenic spot, and comprehensively analyze the image semantic carrying strength index and link stability index of each area of ​​the cultural and tourism scenic spot. Analyze and obtain the image coding adaptation index of each area of ​​the cultural and tourism scenic area, and encode and transmit the data; perform security assessment analysis on the monitoring image data of each area of ​​the cultural and tourism scenic area after the encoding and transmission processing, and obtain the security perception trust index of each area of ​​the cultural and tourism scenic area after the encoding and transmission processing; based on the security perception trust index, issue a corresponding security alarm for each area of ​​the cultural and tourism scenic area after the encoding and transmission processing, and the specific steps are: judge and analyze the security perception trust index of each area of ​​the cultural and tourism scenic area after the encoding and transmission processing with the preset security perception trust index threshold; if the security perception trust index of each area of ​​the cultural and tourism scenic area after the encoding and transmission processing is higher than the preset security perception trust index threshold, no security alarm is issued for the area; if the security perception trust index of each area of ​​the cultural and tourism scenic area after the encoding and transmission processing is lower than or equal to the preset security perception trust index threshold, a security alarm is issued for the area.

[0030] The specific formula for calculating the image coding adaptation index of a certain area of ​​a cultural and tourism scenic spot is as follows: Among them, TsY is the image coding adaptation index of a certain area of ​​the cultural and tourism scenic area, TxC is the image semantic carrying strength index of a certain area of ​​the cultural and tourism scenic area, φ1 is the image carrying adjustment coefficient stored in the database, LwD is the link stability index of a certain area of ​​the cultural and tourism scenic area, φ2 is the link adjustment coefficient stored in the database, and φ3 is the interaction adjustment coefficient stored in the database.

[0031] It should be explained that φ1, φ2, and φ3 can be obtained through the following steps: using historical data, combined with the image semantic carrying strength index and the link stability index, to conduct statistical regression analysis, quantify the specific impact of each factor on the image coding adaptation index, and thus fit the initial weight value. Secondly, using the sensitivity analysis method, adjust the value range of each coefficient, observe its impact on the image coding adaptation evaluation results, ensure the stability and rationality of the model, and based on regional characteristics and actual conditions, correct and optimize the preliminary fitted coefficients, and finally determine the coefficient value applicable to the specific area.

[0032] The specific steps of encoding and transmitting the monitoring image data of each area of ​​the cultural and tourism scenic area based on the image coding adaptation index are as follows: comparing and analyzing the image coding adaptation index of each area of ​​the cultural and tourism scenic area with the preset image coding adaptation index threshold;

[0033] If the image coding adaptability index of each area of ​​the cultural and tourism scenic area is lower than or equal to the preset image coding adaptability index threshold, a first coding transmission measure is adopted for the monitoring image data of the area, which is specifically to perform high compression ratio coding processing to reduce the image data volume and transmit it;

[0034] If the image coding adaptation index of each area of ​​the cultural and tourism scenic area is higher than the preset image coding adaptation index threshold, the second coding transmission measure will be taken for the monitoring image data of the area, which is specifically: using low compression ratio encoding or lossless compression encoding to retain the image detail structure, and directly calling the scheduling controller for real-time transmission resource allocation.

[0035] Specifically, the monitoring image data includes the pixel value and two-dimensional coordinates of each pixel point. The specific steps for obtaining the image semantic carrying strength index of each area of ​​the cultural and tourism scenic area are as follows: the monitoring image data of each area of ​​the cultural and tourism scenic area is input into the pre-trained content recognition model for evaluation and analysis to obtain the image content evaluation set of each area of ​​the cultural and tourism scenic area, including the semantic target density index, occlusion risk index, visual field structure hierarchy index, and high-risk component significance index; the monitoring frame content evaluation set of each area of ​​the cultural and tourism scenic area is comprehensively analyzed to obtain the image semantic carrying strength index of each area of ​​the cultural and tourism scenic area.

[0036] The content recognition model is specifically a convolutional neural network, which includes an input layer, a feature extraction layer, a semantic compression layer, a feature aggregation layer, and a prediction output layer. The specific steps for obtaining the image content evaluation set for each area of ​​the cultural and tourism scenic area are as follows: In the input layer of the convolutional neural network, the monitoring image data of each area of ​​the cultural and tourism scenic area (i.e., the pixel value and two-dimensional coordinates of each pixel point) are received and preprocessed;

[0037] In the feature extraction layer of the convolutional neural network, feature encoding processing is performed on the monitoring image data of each area of ​​the pre-processed cultural and tourism scenic area (using multiple convolution kernels with a receptive field of 3×3, and extracting edge gradient changes, corner point responses and texture density features in the image in a sliding window manner; for example: it can extract repetitive boundaries or local changes from areas such as tourists' clothing textures, floor tile lines, and building edges; batch normalization processing is performed on the output results of each group of convolution channels to stabilize the distribution range of the characteristic values ​​of each channel, and the ReLU activation function operation is performed to compress all negative values ​​to 0, retain positive responses, and enhance nonlinear modeling capabilities; for example: in areas where the background light is too dark, ReLU helps retain the activation information of the edges of buildings or signs. Subsequently, multiple deep convolution kernels with a receptive field of 5×5 are used to further process the above-mentioned activated feature map to extract the contours and spatial distribution information of large areas. For example, the overall structure of a group of tourists, ground patterns, and scenic area guide signs can be extracted in a square scene. Through the above-mentioned multi-level stacked convolution operation, the multi-scale structural expression of the image is gradually extracted, and finally the multi-channel image initial feature map is output as the expression result of the low-level image structure. The result includes the basic perception of the current monitoring frame in terms of texture, edge, contrast change, contour direction, etc., and the initial image feature map of each area of ​​the cultural and tourism scenic area is obtained.

[0038] In the semantic compression layer of the convolutional neural network, the initial feature map of the image of each area of ​​the cultural and tourism scenic area is subjected to convolution compression processing (downsampling is performed using a 3×3 convolution kernel with a step size of 2 to reduce the spatial size of the feature map and achieve regional aggregation; for example, in an aisle scene with multiple people standing, multiple separate human body areas are merged into a crowd hot zone; then, batch normalization and ReLU activation are performed on the downsampling results to unify the distribution of response values ​​of each channel and enhance the response capabilities of object edges, occluded areas, and deep texture structures in the scene; for example: in a scene where people and vehicles meet, the recognition response of the part of the front of the vehicle that is occluded by people can be enhanced; then, a 1×1 or 3×3 convolution kernel is used to expand the number of channels to map each spatial position of the feature map to a richer semantic dimension. Degree, such as forming response channels for people, guardrails, exits, light boxes and other components respectively, for example: one channel specifically responds to the parallel line structure of the guardrail, and another channel responds to the brightness boundary of the background electronic billboard; after the above operations are repeatedly stacked, a feature transmission path with continuous spatial compression and continuous semantic enhancement is formed. The final output semantic feature map includes but is not limited to: object boundary preservation: such as the outline of people, the boundary line between the top of the vehicle and the background; regional aggregation mode: such as the fusion of the crowd gathering area or the area with dense guide signs in the scenic area; occlusion overlapping structure: such as the front person occluding the rear vehicle, the fence occluding the small items below; spatial hierarchical structure: such as the sense of separation of the three-layer space of foreground people, middle ground road surface, and background building, etc.), to obtain the semantic feature map of each area of ​​the cultural and tourism scenic area;

[0039] In the feature aggregation layer of the convolutional neural network, global average pooling processing is performed on the semantic feature map of each area of ​​the cultural and tourism scenic area (global average pooling processing is performed at the channel level, that is, pixel mean statistics are performed on the two-dimensional response area within each channel, and the average response intensity of the channel in the entire image is output; multiple channel pooling values ​​are combined to form an image structure expression vector, which represents the global structural response level of the image under dimensions such as semantic density, structural complexity, occlusion pattern and component existence; for example: a high value of a certain channel indicates a high density of people gathering areas in the image; a low value of another channel indicates that there are no obvious component boundaries in the frame, such as railings / signs; the combination of pooling results reflects the structural semantic portrait of the entire frame image), and the image structure expression vector of each area of ​​the cultural and tourism scenic area is obtained;

[0040] In the prediction output layer of the convolutional neural network, the image structure expression vector of each area of ​​the cultural and tourism scenic area is predicted (the following four types of structure-related feature sub-vectors are extracted from the image structure expression vector respectively: sub-vectors related to the number of people and the density of the target area are used to calculate the semantic target density index; sub-vectors related to boundary fuzziness, occlusion distribution, and non-continuous area response are used to calculate the occlusion risk index; sub-vectors related to the upper and lower depth of field structure and the change of far and near spatial frequency are used to calculate the visual field structure hierarchy index; sub-vectors related to the structural response of railings, warning signs, entrance and exit guide components, etc. are used to calculate the high-risk component significance index; linear weighting processing is performed on each group of the above sub-vectors, and the S The image is normalized by the igmoid function, and the final output is the following four image structure semantic indices: semantic target density index: indicates the spatial distribution density of the number of semantic targets, such as people, cars, and objects, appearing in the current monitoring frame; occlusion risk index: evaluates the occlusion relationship between objects in the image and its impact on the integrity of the scene; field of view structure hierarchy index: indicates the boundary hierarchy between the foreground, middle, and background at different depth levels in the image; high-risk component saliency index: indicates the saliency response strength of the image containing potential high-risk suggestive components, such as railings, elevator buttons, emergency exits, etc.). The semantic target density index, occlusion risk index, field of view structure hierarchy index, and high-risk component saliency index are obtained for each area of ​​the cultural and tourism scenic area.

[0041] Among them, the input layer is used to receive and preprocess the surveillance image and complete normalization and size adaptation.

[0042] The feature extraction layer is used to extract edge, texture and basic semantic features in the image.

[0043] The semantic compression layer is used to compress the spatial structure and enhance the semantic expression, extracting mid- and high-order semantic features.

[0044] The feature aggregation layer is used to compress the semantic feature map into a unified structure vector and extract the global response features.

[0045] The prediction output layer is used to output semantic object density, occlusion, hierarchy and component index based on the structure vector.

[0046] And the pre-training steps of the convolutional neural network are as follows:

[0047] A labeled surveillance image dataset is obtained, which contains video frame images from different areas of cultural and tourism scenic spots and their corresponding structural annotation information. The annotations include annotations of the distribution areas of people in the image (for semantic density), annotations of occlusion relationship areas (such as the overlapping areas of front and back objects), annotations of scene depth partitions (such as the division of front, middle and background areas), and annotations of key component areas (such as railings, exits, warning signs, etc.); and the dataset is divided into a monitoring training set and a monitoring verification set for model learning and performance evaluation.

[0048] Initialize the convolutional neural network and all weight parameters. Use the Kaiming method to randomly initialize the convolutional layer weights to adapt to the ReLU activation distribution. Load some basic models pre-trained on general scene images (such as ResNet base layer or general image encoder) as feature extraction layer initialization to improve convergence speed and structural generalization ability.

[0049] Training is performed based on the monitored training set, with the number of training rounds (e.g., 100 rounds) and batch size set. Images in the monitored training set are used as input samples, and the following steps are performed in each round of training: Forward propagation: the image is input into the convolutional neural network, and the semantic target density index, occlusion risk index, field of view structure hierarchy index, and high-risk component saliency index are predicted respectively; Loss calculation: the loss function between the predicted value and the labeled value of each task branch is calculated separately, including the L2 loss of the density index, the cross entropy loss of the occlusion index, the KL divergence of the structure hierarchy index, and the smoothed L1 loss of the component index, and the weighted fusion of each loss term is performed to obtain the total loss; Back propagation and parameter update: gradient back propagation is performed based on the total loss, and the network parameters are updated using the AdamW optimizer. Strategies such as learning rate cosine decay, L2 regularization, and gradient clipping are combined to improve the model robustness and prevent gradient explosion during training.

[0050] After each round of training, full forward reasoning is performed using the monitored validation set images, and four types of structural indices are output respectively. The performance of the four types of structural indices on the validation set is evaluated, including the mean absolute error (MAE) of the semantic target density index, the IoU and accuracy of the occlusion risk prediction, and the correlation index of the field of view level index (such as Pearson or R 2 ), the F1-score of the component significance index, and record the changing trends of loss and performance indicators with training rounds to assist in judging the convergence of the model; when the performance does not improve or the loss value does not decrease after several consecutive rounds of verification, execute Early Stopping to terminate the training early.

[0051] After training is complete, save the model weight parameter file that performs best in the convergence phase and export it to a deployable model format.

[0052] The specific formula for calculating the image semantic carrying strength index of a certain area in a cultural and tourism scenic area is as follows: Among them, TxC is the image semantic carrying intensity index of a certain area of ​​the cultural and tourism scenic area, YmB is the semantic target density index of a certain area of ​​the cultural and tourism scenic area, α1 is the target density adjustment coefficient stored in the database, SyG is the visual field structure hierarchy index of a certain area of ​​the cultural and tourism scenic area, α2 is the structure hierarchy adjustment coefficient stored in the database, GwX is the high-risk component significance index of a certain area of ​​the cultural and tourism scenic area, α3 is the high-risk adjustment coefficient stored in the database, ZdF is the occlusion risk index of a certain area of ​​the cultural and tourism scenic area, and α4 is the occlusion adjustment coefficient stored in the database.

[0053] It should be explained that α1, α2, α3, and α4 can be obtained through the following steps: based on the historical data of the region, the initial influence weights of each variable (such as the semantic target density index, the visual field structure hierarchy index, the high-risk component significance index, and the occlusion risk index) on the image semantic carrying intensity index are determined through statistical regression analysis. Then, the value range of the coefficient is adjusted using the sensitivity analysis method to evaluate the stability and applicability of these parameters to the formula output. Next, the weights are further fitted through model optimization (such as machine learning algorithms) to ensure that the formula can accurately reflect the actual intensity of the image semantic carrying.

[0054] The specific implementation example of calculating the image semantic carrying strength index of a certain area of ​​the cultural and tourism scenic area is as follows. The existing data are as follows: semantic target density index, visual field structure level index, high-risk component significance index, and occlusion risk index of three areas (randomly selected) of the cultural and tourism scenic area, as shown in Table 1 and Figure 2-5 As shown:

[0055] Table 1 Example of regional sequence data of image content evaluation set of cultural and tourism scenic spots

[0056]

[0057] The target density adjustment coefficient α1 stored in the database is approximately: 0.534;

[0058] The structural level adjustment coefficient α2 stored in the database is approximately: 0.671;

[0059] The high-risk adjustment coefficient α3 stored in the database is approximately: 0.243;

[0060] The occlusion adjustment coefficient α4 stored in the database is approximately: 1.245;

[0061] Substituting the data in Table 1 and the above adjustment coefficients into the specific formula for calculating the image semantic carrying strength index of a certain area in a cultural and tourism scenic area, we obtain:

[0062] The image semantic carrying strength index of a certain area of ​​the cultural and tourism scenic spot = (tanh(0.864 0.543 ×(1+0.7230.671 )×(exp(0.357))×0.243)) / (1+ln(1+0.357×1.245))≈0.463;

[0063] The image semantic carrying strength index of a certain area of ​​the cultural and tourism scenic spot = (tanh(0.532 0.543 ×(1+0.415 0.671 )×(exp(0.689))×0.243)) / (1+ln(1+0.426×1.245))≈0.343;

[0064] The image semantic carrying strength index of a certain area of ​​the cultural and tourism scenic spot = (tanh(0.872 0.543 ×(1+0.549 0.671 )×(exp(0.743))×0.243)) / (1+ln(1+0.182×1.245))≈0.547.

[0065] In this implementation, a convolutional neural network is used to extract multidimensional structural features such as target density, occlusion risk, structural hierarchy and component significance in the image, and a parameterized formula construction mechanism is introduced to map each type of semantic feature into an adjustable index contribution item through an adjustment factor. Secondly, combined with regional historical data, statistical regression analysis and sensitivity assessment methods are used to dynamically obtain the optimal value range of each adjustment parameter, and the stability and generalization ability of the output are improved through model fitting, thereby ensuring that the semantic carrying index has the ability to express structural separability and adjustable weights, thereby giving the model the ability to adaptively model the complexity of image structures under different scenarios and regional conditions, thereby effectively improving the system's response accuracy and structural adaptability to subsequent coding strategies and risk judgment mechanisms.

[0066] Specifically, the current communication status data includes link accumulation value, frequency domain interference intensity index, signal-to-noise ratio value, communication thermal interference index, carrier frequency drift index, and power feedback residual value. The specific steps for obtaining the link stability index of each area of ​​the cultural and tourism scenic area are as follows: the current communication status data of each area of ​​the cultural and tourism scenic area are comprehensively analyzed to obtain a link evaluation set of each area of ​​the cultural and tourism scenic area, including a communication interference sensitivity index and a link smoothness index; based on the particle swarm optimization algorithm, a comprehensive analysis is performed on the link evaluation set of each area of ​​the cultural and tourism scenic area (that is, weighted processing, and in the process of weighted processing, the weighted coefficients corresponding to the communication interference sensitivity index and the link smoothness index are obtained through the particle swarm optimization algorithm, which is as follows: first initialize a number of particles, each particle represents a set of possible weights The weight coefficient combination corresponds to the communication interference sensitivity index and link smoothness index to be weighted; then the objective function is set, which is used to evaluate the degree of fit between the link stability valuation result and the historical link real-time feedback result of a certain set of weight combinations in the current scenario; the position and velocity of the particle swarm are updated during the iterative optimization process, so that the particles are continuously close to the solution that makes the objective function optimal in the search space; in each iteration, the current individual optimal position and the global optimal position are recorded, and all particles are guided to gradually converge in the search space; when the convergence condition is met or the maximum number of iterations is reached, a set of optimized weight coefficients are finally obtained, which are used as the weight coefficients corresponding to the communication interference sensitivity index and the link smoothness index to obtain the link stability index of each area in the cultural and tourism scenic area.

[0067] Among them, the link backlog value is the total amount of data queued in the link send buffer when the current monitoring frame is sent to the link scheduler. It can be obtained by directly reading the queue depth register corresponding to the send buffer when the current frame arrives at the scheduler. The register records the total number of bytes of the data units that have not been completed. This value can directly reflect the current link load pressure.

[0068] The frequency domain interference strength index is the total power value of external interference other than the target signal in the communication frequency band used by the current monitoring frame. It can be obtained by performing a fast Fourier transform on the received intermediate frequency signal in the current frame pilot stage, excluding the center frequency point, and then accumulating the signal power at the remaining frequency points. The larger the value, the stronger the interference.

[0069] The signal-to-noise ratio value is the ratio of the target signal strength of the previous monitoring frame at the receiving end to the background noise power, which can be directly read through the register inside the communication chip.

[0070] The communication thermal interference index is the degree of increase in intermediate frequency thermal noise caused by environmental heat sources (such as glass curtain walls, bare rock reflection, and ground thermal radiation). It can be achieved by setting an infrared thermopile sensor in the communication direction to collect the infrared radiation temperature value in the current frame transmission direction in real time. At the same time, the noise power in the intermediate frequency path is obtained, and the infrared radiation temperature value and noise power are standardized. Based on the standardized processing results, weighted processing is performed to obtain the communication thermal interference index.

[0071] The carrier frequency drift index is the instantaneous drift value of the local phase-locked oscillator (LO) relative to the reference frequency in the current frame, reflecting the quality of communication synchronization. It can feedback the current local oscillator frequency compensation value through the phase-locked control loop, and the compensation value can be directly read through the frequency control word register.

[0072] The power feedback residual value is the instantaneous difference between the transmission control set power and the actual output power of the power amplifier (PA), reflecting the control accuracy of the communication transmission path. It can be achieved by having the main control chip write the set transmission power value into the power control register before the current monitoring frame enters the transmission state. At the same time, the voltage detection circuit arranged at the output end of the PA samples the actual output power, and sends the sampled analog signal to the analog-to-digital converter (ADC) to convert it into the corresponding digital power value; the controller performs real-time difference calculation on the set power value and the sampled power value to obtain the power feedback residual value of the current frame.

[0073] The specific steps for obtaining the link evaluation set of each area of ​​the cultural and tourism scenic area are as follows: Based on the particle swarm optimization algorithm, a comprehensive analysis is performed on the frequency domain interference intensity index, communication thermal interference index, carrier frequency drift index, and power feedback residual value of each area of ​​the cultural and tourism scenic area (the frequency domain interference intensity index, communication thermal interference index, carrier frequency drift index, and power feedback residual value are first standardized, and weighted based on the standardized processing results. In the weighted processing process, the frequency domain interference intensity index, communication thermal interference index, carrier frequency drift index, and power feedback residual value corresponding to the particle swarm optimization algorithm are obtained, and compared with the communication interference sensitivity index and link smoothness index. The corresponding weighting coefficients are obtained through the particle swarm optimization algorithm (the logic is consistent with that obtained by the particle swarm optimization algorithm), and the communication interference sensitivity index of each area in the cultural and tourism scenic area is obtained; based on the particle swarm optimization algorithm, the link stacking value and the signal-to-noise ratio value of each area in the cultural and tourism scenic area are comprehensively analyzed (the link stacking value and the signal-to-noise ratio value are first standardized, and weighted processing is performed based on the standardized processing results. In the weighted processing process, the link stacking value and the signal-to-noise ratio value corresponding to the particle swarm optimization algorithm are obtained, and the logic is consistent with that of the weighting coefficients corresponding to the communication interference sensitivity index and the link smoothness index obtained through the particle swarm optimization algorithm), and the link smoothness index of each area in the cultural and tourism scenic area is obtained.

[0074] In this implementation scheme, a link feature set based on multiple physical communication status indicators such as link accumulation, interference intensity, thermal noise drift, and signal-to-noise ratio is constructed, and the communication interference sensitivity index and link smoothness index are weightedly modeled based on the particle swarm optimization algorithm to form a communication link stability index, thereby effectively improving the accuracy and dynamic adaptability of link status valuation. Secondly, the introduction of the particle swarm algorithm enables each weight factor to be iteratively optimized based on the historical communication actual feedback results, thereby dynamically adjusting the proportion of the influence of different communication parameters on link stability. Finally, this step has a keen response capability to unconventional communication anomalies such as frequency domain interference and carrier drift, thereby effectively enhancing the stable monitoring and transmission control capabilities under extreme communication conditions such as low signal-to-noise ratio, strong thermal interference, and frequent electromagnetic disturbances, thereby providing a stable and efficient data foundation for coding adaptive adjustment and subsequent reliable acquisition of image information.

[0075] Specifically, if Figure 6As shown in FIG, the specific steps for obtaining the security perception trust index of each area of ​​the cultural and tourism scenic area after the coding and transmission processing are as follows: the monitoring image data of each area of ​​the cultural and tourism scenic area after the coding and transmission processing is input into the pre-trained security recognition model for evaluation and analysis, and the security assessment set of each area of ​​the cultural and tourism scenic area after the coding and transmission processing is obtained, including the abnormal behavior index, the illegal intrusion index, and the scene component risk index; based on the entropy weight method, a comprehensive analysis is performed on the security assessment set of each area of ​​the cultural and tourism scenic area after the coding and transmission processing (that is, weighted processing, and in the process of weighted processing, the weighting coefficients corresponding to the abnormal behavior index, the illegal intrusion index, and the scene component risk index are obtained by the entropy weight method, which is specifically as follows: constructing a number of original data matrices of security assessment sets corresponding to cultural and tourism scenic area areas, each column in the matrix corresponds to the abnormal behavior index, the illegal intrusion index and the scene component risk index, and each row corresponds to the actual assessment data of a monitoring area; secondly, each column of the data in the above original data matrix is ​​dimensionlessly standardized, and the risk index value of each area is mapped to 0-1 using interval normalization. interval, to ensure the comparability between different types of indexes and avoid the interference of dimensional differences on the weight calculation results; then, the entropy value of each column of indicators is calculated based on the standardized matrix, and the proportion of the standardized value of each area in the corresponding indicator column is calculated, and based on the proportion value, the entropy value of the column indicator is calculated according to the information entropy formula. The entropy value reflects the information discreteness of the current indicator. The higher the discreteness, the greater the amount of information; further, the information utility value of each column of indicators is calculated based on the above entropy value, that is, the entropy value of the indicator is subtracted from 1 to measure the proportion of effective information contained in the indicator in the overall evaluation; then, the three information utility values ​​are normalized separately to obtain the weighting coefficients of the abnormal behavior index, the illegal intrusion index and the scene component risk index), and the risk identification strength index of each area of ​​the cultural and tourism scenic area after the coding and transmission processing is obtained; the image coding adaptation index of each area of ​​the cultural and tourism scenic area is read, and combined with the risk identification strength index of each area of ​​the cultural and tourism scenic area after the coding and transmission processing, a comprehensive analysis is performed to obtain the security perception trust index of each area of ​​the cultural and tourism scenic area after the coding and transmission processing.

[0076] The security identification model is specifically a risk identification neural network (through YOLOv5s, ResNet-50, Patch CNN, etc.), the risk identification neural network includes an image input layer, a unified extraction layer, a task classification layer, and a regression output layer. The specific steps of obtaining the safety assessment set of each area of ​​the cultural and tourism scenic area after the coding and transmission processing are as follows: in the image input layer of the risk identification neural network, the monitoring image data of each area of ​​the cultural and tourism scenic area after the coding and transmission processing is received and preprocessed; in the unified extraction layer of the risk identification neural network, the preprocessed monitoring image data of each area of ​​the cultural and tourism scenic area after the coding and transmission processing is subjected to joint feature extraction processing to obtain the shared feature vector of each area of ​​the cultural and tourism scenic area after the coding and transmission processing; in the task classification layer of the risk identification neural network, the shared feature vector of each area of ​​the cultural and tourism scenic area after the coding and transmission processing is subjected to classification extraction processing to obtain the risk vector set of each area of ​​the cultural and tourism scenic area after the coding and transmission processing (including behavior feature sub-vectors, intrusion feature sub-vectors, and component feature sub-vectors); in the regression output layer of the risk identification neural network, the risk vector set of each area of ​​the cultural and tourism scenic area after the coding and transmission processing is subjected to regression mapping processing respectively to obtain the abnormal behavior index, illegal intrusion index, and scene component risk index of each area of ​​the cultural and tourism scenic area after the coding and transmission processing.

[0077] The specific process of feature extraction is as follows:

[0078] A convolutional network based on the ResNet-50 network structure is used to perform layer-by-layer convolution operations on the input surveillance image data. Specifically, the first layer uses a 7×7 convolution kernel and a stride of 2 to extract preliminary edge and brightness features of the image. Then, a maximum pooling layer (3×3, stride 2) is used for spatial downsampling. Then, three residual block groups are sequentially passed through, each consisting of two 3×3 convolution layers and a batch normalization layer (BatchNorm) and residual connection operations to obtain a multi-scale spatial texture feature map. The output basic feature map is used to support subsequent regional target and posture structure perception tasks.

[0079] Based on the feature map, a YOLOv5s network detection head structure is integrated to detect human targets and structural targets (such as railings, doors, escalators, etc.) in the image. YOLOv5s uses a three-layer convolutional downsampling structure of different scales and a feature pyramid (FPN) module to output a set of candidate boxes and category confidence at multiple scales. The detection results are mapped to a weight map of the same size as the feature map by generating a spatial attention map. The weighted pixel-by-pixel multiplication operation is performed with the backbone feature map, allowing the network to explicitly focus on target areas with safety risks.

[0080] Then, for the detected human target area, a lightweight pose estimation network BlazePose model structure is introduced. Its internal key point detection branch generates a two-dimensional heat map of 33 key points (one heat map for each key point). The pose heat map is fused with the backbone feature map through channel splicing. The channel dimension is compressed and semantically aligned through two layers of 3×3 convolution, ReLU activation function and BatchNorm convolution unit to obtain a semantic feature map that integrates the pose structure information. This structure is used to enhance the network's ability to recognize changes in human behavior, standing state, running / wandering, and other posture changes.

[0081] Next, a regional mask corresponding to the image is introduced to represent the management boundary information of each area in the cultural and tourism scenic area. A set of 1×1 convolutional layers is used to encode the mask into a spatial vector graph with channel dimensions. This is then fused with the posture semantic feature map output by the previous stage through a dot product operation to enhance the spatial distribution sensitivity of each region. The regional mask is embedded using the SE attention mechanism (Squeeze-and-Excitation) structure, which first performs spatial compression and channel excitation on the mask, and then performs weighted fusion. This enables the network to identify the corresponding area and helps to determine whether there is any regional intrusion.

[0082] Then, for the multi-channel feature map of the fused region and posture, the detected component region image blocks (candidate boxes marked as components by YOLOv5) are selected and structural state perception operations are performed in sequence: first, the classic edge-responsive convolution kernel Sobel operator (horizontal + vertical) is used to extract the contour intensity distribution in the image; then a Patch-based CNN structure (two layers of 3×3 convolution and pooling) is introduced to analyze the degree of local texture change in the region; finally, a difference analysis module is used to compare the edge disconnection, texture gradient mutation rate, and occlusion rate between the historical component image features and the current image features to form a component state feature map for risk index modeling;

[0083] Finally, the multi-channel feature map that integrates multi-source features such as target detection, posture recognition, region boundary mask and component state response is input into the global average pooling layer to generate a one-dimensional feature vector of fixed dimension, namely the shared feature vector.

[0084] The specific process of classification extraction is as follows:

[0085] Three groups of parallel channel selection sub-networks are set up in the task classification layer, corresponding to the tasks of abnormal behavior recognition of personnel, illegal intrusion recognition of regions and risk status recognition of components respectively; each group of channel selection sub-networks contains a group of channel weight generation structures, which are used to extract key semantic channels related to the current task from the shared feature vector, and take the shared feature vector as input, and pass through the first fully connected layer (output dimension is half of the length of the shared vector, activation function is ReLU) and the second fully connected layer (output dimension is the same as the shared vector, activation function is Sigmoid) in sequence to generate the channel weight vector of the corresponding task; the channel weight vector is used to characterize the importance of each channel dimension under the current task; the channel weight vector generated in the channel selection sub-network of each task ... The quantity is element-wise multiplied with the shared feature vector to obtain the original sub-vector related to the task; then, the original sub-vector is input into a feature compression network containing a two-layer fully connected structure (the first layer is 128 dimensions and the second layer is 64 dimensions, both using ReLU activation), and a fixed-length task sub-vector is output. The personnel abnormal behavior task sub-vector mainly reflects the channel dimensions related to the personnel target position, posture key point state, and motion trajectory pattern in the shared features; the regional illegal intrusion task sub-vector mainly reflects the channel dimensions related to spatial layout features such as the monitoring area boundary mask, residence time pattern, and trajectory offset; the component risk status task sub-vector reflects the channel dimensions related to the physical state of the component such as static target texture deformation, edge fracture rate, and occlusion probability.

[0086] The specific process of regression mapping is as follows:

[0087] For each task subvector, a set of independent fully connected regression subnetworks is set up in the regression output layer. Each regression subnetwork includes: the first layer is a 128-dimensional fully connected layer with the ReLU activation function to enhance nonlinear mapping capabilities; the second layer is a 32-dimensional fully connected layer with the ReLU activation function; the third layer is a 1-dimensional output layer with the Sigmoid activation function to normalize the results to the [0, 1] range and output the final risk index value;

[0088] Index output description: The regression sub-network with the behavior feature sub-vector as input outputs the abnormal behavior index, which is used to characterize the abnormal probability of human behavior in the current area;

[0089] The regression sub-network, which takes the intrusion feature sub-vector as input, outputs the violation intrusion index, which is used to characterize the degree of violation of personnel behavior relative to the management boundary.

[0090] The regression subnetwork, which takes the component feature subvector as input, outputs the scene component risk index, which is used to quantify the risk level of damage, obstruction or unsafe state of the physical components in the current area.

[0091] The image input layer is used to receive the surveillance image data after encoding and transmission processing;

[0092] The unified extraction layer is used to extract the shared feature vectors of fused objects, postures, regions, and components;

[0093] The task classification layer is used to extract sub-vectors of behavior, intrusion and component risks from the shared feature vector;

[0094] The regression output layer is used to perform regression calculations on each sub-vector and output the corresponding risk index.

[0095] The pre-training process of the risk identification neural network is as follows:

[0096] A structured surveillance image dataset is obtained, which contains surveillance image frame data from multiple areas of cultural and tourism scenic spots and their corresponding structured annotation information; and the annotation information includes: personnel behavior type and behavior area annotation (used to identify behavior patterns such as running, wandering, and abnormal staying); area boundary and control area mask annotation (used to identify cross-border intrusion risks); category and boundary annotation of component areas in the scene (such as railings, doors, passageways, etc., for component risk identification); event risk level annotation in the area (marking the safety level of special components or complex environment areas); and the structured surveillance image dataset is divided into two parts: a structured training set and a structured validation set.

[0097] The risk identification neural network parameters are initialized, and the Kaiming initialization method is used to randomly initialize the weight parameters of all convolutional neural network layers to adapt to the distribution characteristics of the ReLU activation function; among them, the basic image encoding structure in the network loads the pre-trained ResNet-50 backbone convolutional layer weights to improve feature extraction efficiency and structural generalization ability.

[0098] Training is performed based on a structured training set, with the number of training rounds (e.g., 100 rounds) and batch size (e.g., 32). The monitored images in the training set are used as input samples, and the following training steps are performed: Forward propagation processing: The input images are sequentially input into the risk identification neural network to predict and generate three risk index results: abnormal behavior index, illegal intrusion index, and component risk index; Multi-task loss calculation processing: The loss function is calculated for the predicted results and the true labeled values ​​respectively: The abnormal behavior index uses the L2 regression loss function to measure the behavior pattern prediction error; the illegal intrusion index uses the cross entropy loss function to measure the classification accuracy of regional judgment; the component risk index uses the smooth L1 loss function to improve the regression accuracy of edge status and occlusion risk; All loss terms are weighted and fused to form a total loss function, which serves as the basis for backpropagation; Gradient update and training optimization strategy: The AdamW optimizer is used to update the model parameters, and the learning rate cosine decay scheduling strategy is introduced to control the learning rate change during training; At the same time, L2 regularization is combined to suppress overfitting, and the gradient clipping strategy is used to limit the gradient norm to prevent the gradient explosion problem during training.

[0099] After each round of training, forward reasoning is performed using a structured validation set, and validation metrics are calculated for the three output indices, including the mean absolute error (MAE) of the abnormal behavior index; the accuracy and intersection over union (IoU) of the illegal intrusion index; and the F1-score and root mean square error (RMSE) of the component risk index. The loss reduction and performance indicator change trends during each round of training are recorded to determine the model convergence status. If the validation metrics do not improve or the total loss stops decreasing for several consecutive rounds, the Early Stopping strategy is automatically triggered to terminate training early.

[0100] Finally, a set of model parameter weight files with the best performance during the training process is saved and exported as a deployable model structure file for subsequent integration and call by the cultural and tourism scenic area monitoring system.

[0101] The specific formula for calculating the security perception trust index of a certain area of ​​a cultural and tourism scenic area after coding transmission processing is as follows: TxC=TsY η1 *[1-tanh(η2*FxS)]*[1-exp(-η3*|TsY-FxS|)]; where TxC is the security perception credibility index of a certain area of ​​the cultural and tourism scenic area after coding and transmission processing, TsY is the image coding adaptation index of a certain area of ​​the cultural and tourism scenic area, η1 is the image coding adjustment coefficient stored in the database, FxS is the risk identification intensity index of a certain area of ​​the cultural and tourism scenic area after coding and transmission processing, η2 is the risk adjustment coefficient stored in the database, and η3 is the difference adjustment coefficient stored in the database.

[0102] It should be explained that the term [1-exp(-η3*|TsY-FxS|)] in the formula is used to adjust the impact of the numerical deviation between the risk identification strength index and the image coding adaptation index on the security perception credibility index, so as to avoid the security perception credibility index being too high or too low.

[0103] η1, η2, and η3 can be obtained through the following steps: Based on historical data, the initial impact weights of each variable (image coding adaptation index, risk identification intensity index) on the security perception credibility index are determined through statistical regression analysis. Then, the value range of the coefficient is adjusted using the sensitivity analysis method to evaluate the stability and applicability of these parameters to the formula output. Next, the weights are further fitted through model optimization (such as machine learning algorithms or multi-objective optimization) to ensure that the formula can accurately reflect the safety status of the actual scenic area. The coefficients are fine-tuned based on the characteristics of different regions to ensure that they are suitable for specific scenic area safety assessment needs.

[0104] In this implementation plan, a credibility modeling path for security monitoring tasks in cultural and tourism scenic spots is constructed by integrating a multi-source neural network structure and an exponential weighting strategy. A risk identification neural network that integrates YOLOv5s target detection, BlazePose pose estimation, and Patch CNN component analysis is used to support the independent quantification of abnormal behavior, illegal intrusion, and scene component risks, thereby significantly improving the fine-grained recognition capabilities of behavior, region, and component dimensions. Subsequently, the entropy weight method is used to evaluate the information weight of each recognition result, and a weighting coefficient is automatically generated based on the degree of discreteness of the indicator in the global data. This achieves dynamic weight allocation without manual intervention, thereby improving the objectivity and adaptability of the fusion analysis. Finally, by constructing a security perception credibility index model, the image coding adaptation index and the risk identification strength index are weighted and fused, thereby achieving a dynamic match between risk identification and communication expression capabilities, thereby ensuring the ability to output security warnings with high credibility even under adverse conditions such as severe image compression and link jitter.

[0105] See also Figure 7, an embodiment of the present invention provides a technical solution: a smart real-time security monitoring system for cultural and tourism scenic spots based on big data analysis, including: a data acquisition and analysis module, used to acquire monitoring image data of several areas of the cultural and tourism scenic spot in real time, and input it into a pre-trained content recognition model for comprehensive analysis to obtain the image semantic carrying strength index of each area of ​​the cultural and tourism scenic spot; a feature analysis module, used to acquire the current communication status data of each area of ​​the cultural and tourism scenic spot, and perform feature analysis to obtain the link stability index of each area of ​​the cultural and tourism scenic spot; a data coding and transmission module, used to perform comprehensive analysis on the image semantic carrying strength index and link stability index of each area of ​​the cultural and tourism scenic spot, obtain the image coding adaptation index of each area of ​​the cultural and tourism scenic spot, and perform coding and transmission processing; a security assessment and analysis module, used to perform security assessment and analysis on the monitoring image data of each area of ​​the cultural and tourism scenic spot after coding and transmission processing, and obtain the security perception trust index of each area of ​​the cultural and tourism scenic spot after coding and transmission processing; a security alarm feedback module, used to issue corresponding security alarms to each area of ​​the cultural and tourism scenic spot after coding and transmission processing based on the security perception trust index.

[0106] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0107] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A real-time safety monitoring method for smart cultural and tourism scenic spots based on big data analysis, characterized in that: The following steps are involved: Acquire surveillance image data from several areas of a cultural and tourism scenic spot in real time, input it into a pre-trained content recognition model for comprehensive analysis, and obtain the image semantic carrying strength index of each area of ​​the cultural and tourism scenic spot; The current communication status data of each area of ​​the cultural and tourism scenic area is obtained, and characteristic analysis is performed to obtain the link stability index of each area of ​​the cultural and tourism scenic area; Comprehensively analyze the image semantic carrying strength index and link stability index of each area in the cultural and tourism scenic area to obtain the image coding adaptability index of each area in the cultural and tourism scenic area, and then encode and transmit it; Conduct security assessment and analysis on the surveillance image data of each area in the cultural and tourism scenic area after encoding and transmission processing, and obtain the security perception trust index of each area in the cultural and tourism scenic area after encoding and transmission processing; Based on the security perception trust index, corresponding security alerts are issued to each area of ​​the cultural and tourism scenic area after the encoding transmission processing.

2. The method for real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis according to claim 1 is characterized in that: The monitoring image data specifically includes the pixel value and two-dimensional coordinates of each pixel point. The specific steps for obtaining the image semantic carrying strength index of each area of ​​the cultural and tourism scenic area are as follows: The surveillance image data of each area of ​​the cultural and tourism scenic area is input into the pre-trained content recognition model for evaluation and analysis, and an image content evaluation set of each area of ​​the cultural and tourism scenic area is obtained, including the semantic target density index, occlusion risk index, visual field structure hierarchy index, and high-risk component prominence index. A comprehensive analysis is performed on the monitoring frame content evaluation set of each area of ​​the cultural and tourism scenic area to obtain the image semantic carrying strength index of each area of ​​the cultural and tourism scenic area.

3. The method for real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis according to claim 2 is characterized in that: The content recognition model is specifically a convolutional neural network, which includes an input layer, a feature extraction layer, a semantic compression layer, a feature aggregation layer, and a prediction output layer. The specific steps of obtaining the image content evaluation set for each area of ​​the cultural and tourism scenic area are as follows: In the input layer of the convolutional neural network, the monitoring image data of each area of ​​the cultural and tourism scenic area is received and preprocessed; In the feature extraction layer of the convolutional neural network, feature encoding is performed on the pre-processed monitoring image data of each area of ​​the cultural and tourism scenic area to obtain the initial image feature map of each area of ​​the cultural and tourism scenic area; In the semantic compression layer of the convolutional neural network, the initial image feature map of each area of ​​the cultural and tourism scenic area is subjected to convolution compression processing to obtain the semantic feature map of each area of ​​the cultural and tourism scenic area; In the feature aggregation layer of the convolutional neural network, the semantic feature map of each area of ​​the cultural and tourism scenic area is globally averaged and pooled to obtain the image structure expression vector of each area of ​​the cultural and tourism scenic area. In the prediction output layer of the convolutional neural network, the image structure expression vector of each area of ​​the cultural and tourism scenic area is predicted and processed to obtain the semantic target density index, occlusion risk index, visual field structure hierarchy index, and high-risk component significance index of each area of ​​the cultural and tourism scenic area.

4. The method for real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis according to claim 3 is characterized in that: The specific formula for calculating the image semantic carrying strength index of a certain area in a cultural and tourism scenic area is as follows: Among them, TxC, YmB, SyG, GwX, and ZdF are the image semantic carrying strength index, semantic target density index, visual field structure hierarchy index, high-risk component significance index, and occlusion risk index of a certain area of ​​the cultural and tourism scenic area, respectively. YmB is the semantic target density index of a certain area of ​​the cultural and tourism scenic area. α1, α2, α3, and α4 are the target density adjustment coefficient, structure hierarchy adjustment coefficient, high-risk adjustment coefficient, and occlusion adjustment coefficient stored in the database, respectively.

5. The method for real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis according to claim 1 is characterized in that: The current communication status data includes link accumulation value, frequency domain interference intensity index, signal-to-noise ratio value, communication thermal interference index, carrier frequency drift index, and power feedback residual value. The specific steps for obtaining the link stability index of each area in the cultural and tourism scenic area are as follows: Comprehensively analyze the current communication status data of each area in the cultural and tourism scenic area to obtain a link evaluation set for each area in the cultural and tourism scenic area, including the communication interference sensitivity index and link smoothness index; Based on the particle swarm optimization algorithm, a comprehensive analysis of the link evaluation set of each area in the cultural and tourism scenic area is performed to obtain the link stability index of each area in the cultural and tourism scenic area.

6. The method for real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis according to claim 5 is characterized in that: The specific steps to obtain the link evaluation set for each area of ​​the cultural and tourism scenic area are as follows: Based on the particle swarm optimization algorithm, a comprehensive analysis of the frequency domain interference intensity index, communication thermal interference index, carrier frequency drift index, and power feedback residual value of each area in the cultural and tourism scenic area was performed to obtain the communication interference sensitivity index of each area in the cultural and tourism scenic area; Based on the particle swarm optimization algorithm, a comprehensive analysis of the link accumulation value and signal-to-noise ratio value of each area in the cultural and tourism scenic area is performed to obtain the link smoothness index of each area in the cultural and tourism scenic area.

7. The method for real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis according to claim 2 is characterized in that: The specific steps for obtaining the security perception trust index of each area in the cultural and tourism scenic area after encoding and transmission processing are as follows: The surveillance image data of each area of ​​the cultural and tourism scenic area after encoding and transmission processing is input into the pre-trained security recognition model for evaluation and analysis, and a security assessment set of each area of ​​the cultural and tourism scenic area after encoding and transmission processing is obtained, including abnormal behavior index, illegal intrusion index, and scene component risk index; Based on the entropy weight method, a comprehensive analysis is conducted on the safety assessment set of each area in the cultural and tourism scenic area after the coding and transmission processing, and the risk identification intensity index of each area in the cultural and tourism scenic area after the coding and transmission processing is obtained; The image coding adaptation index of each area of ​​the cultural and tourism scenic area is read, and a comprehensive analysis is performed on the risk identification intensity index of each area of ​​the cultural and tourism scenic area after coding and transmission processing to obtain the safety perception trust index of each area of ​​the cultural and tourism scenic area after coding and transmission processing.

8. The method for real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis according to claim 7 is characterized in that: The safety identification model is specifically a risk identification neural network, which includes an image input layer, a unified extraction layer, a task classification layer, and a regression output layer. The specific steps of obtaining the safety assessment set of each area of ​​the cultural and tourism scenic area after encoding and transmission processing are as follows: In the image input layer of the risk identification neural network, the monitoring image data of each area of ​​the cultural and tourism scenic area after encoding and transmission processing is received and preprocessed; In the unified extraction layer of the risk identification neural network, joint feature extraction is performed on the monitoring image data of each area of ​​the cultural and tourism scenic area after pre-processing and encoding and transmission processing, and a shared feature vector of each area of ​​the cultural and tourism scenic area after encoding and transmission processing is obtained; In the task classification layer of the risk identification neural network, the shared feature vectors of each area of ​​the cultural and tourism scenic area after the encoding and transmission processing are classified and extracted to obtain the risk vector set of each area of ​​the cultural and tourism scenic area after the encoding and transmission processing; In the regression output layer of the risk identification neural network, the risk vector set of each area of ​​the cultural and tourism scenic area after the coding and transmission processing is subjected to regression mapping processing to obtain the abnormal behavior index, illegal intrusion index and scene component risk index of each area of ​​the cultural and tourism scenic area after the coding and transmission processing.

9. The method for real-time safety monitoring of smart cultural and tourism scenic spots based on big data analysis according to claim 7 is characterized in that: The specific formula for calculating the security perception trust index of a certain area of ​​a cultural and tourism scenic spot after coding and transmission processing is as follows: Among them, TxC and FxS are the security perception credibility index and risk identification intensity index of a certain area of ​​the cultural and tourism scenic area after coding transmission processing, TsY is the image coding adaptation index of a certain area of ​​the cultural and tourism scenic area, and η1, η2, and η3 are the image coding adjustment coefficient, risk adjustment coefficient, and difference adjustment coefficient stored in the database, respectively.

10. A smart real-time safety monitoring system for cultural and tourism scenic spots based on big data analysis, applying the smart real-time safety monitoring method for cultural and tourism scenic spots based on big data analysis according to any one of claims 1 to 9, characterized in that: include: The data acquisition and analysis module is used to obtain monitoring image data of several areas of the cultural and tourism scenic area in real time, and input it into the pre-trained content recognition model for comprehensive analysis to obtain the image semantic carrying strength index of each area of ​​the cultural and tourism scenic area; The feature analysis module is used to obtain the current communication status data of each area in the cultural and tourism scenic area and perform feature analysis to obtain the link stability index of each area in the cultural and tourism scenic area; The data encoding and transmission module is used to comprehensively analyze the image semantic carrying strength index and link stability index of each area of ​​the cultural and tourism scenic area, obtain the image encoding adaptation index of each area of ​​the cultural and tourism scenic area, and perform encoding and transmission processing; A security assessment and analysis module is used to perform security assessment and analysis on the surveillance image data of each area of ​​the cultural and tourism scenic area after the encoding and transmission processing, and obtain a security perception trust index for each area of ​​the cultural and tourism scenic area after the encoding and transmission processing; The security alarm feedback module is used to issue corresponding security alarms to each area of ​​the cultural and tourism scenic area after the encoding transmission processing based on the security perception trust index.

Citation Information

Patent Citations

  • A Real-time Safety Monitoring System and Method for Scenic Areas Based on Smart Cultural Tourism

    CN116528036B