Big data analysis method and system based on cloud service platform
By formatting the satellite image data and deep learning feature extraction on the cloud business platform, and adjusting the detection parameters in combination with the historical database matching, the problems of poor environmental adaptability and insufficient trend analysis in the existing technology are solved, and high-precision abnormality recognition and continuous monitoring are achieved.
Patent Information
- Application Number
- CN202510315407.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, feature extraction is based on a fixed feature template, and it is difficult to adapt to changes under different environmental conditions, resulting in a decrease in the accuracy of abnormal identification in complex scenarios. Anomaly detection mainly relies on current data, making it difficult to conduct long-term trend analysis, and it is easy to misjudgment that short-term fluctuations are abnormal.
The big data analysis method based on the cloud business platform is adopted, and by collecting satellite image data, formatting and quality screening, feature extraction is used using deep learning, matching the abnormal detection results with the historical database, adjusting the detection parameters, and continuously monitoring.
It improves the adaptability to complex environments, reduces false detection and missed detection, and realizes high-precision abnormality identification and long-term trend analysis of complex environments, ensuring the continuity and accuracy of monitoring results.
Smart Images

Figure CN120259902A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image understanding, and particularly to a big data analysis method and system based on a cloud business platform. Background Art
[0002] Image understanding is the core technology in the field of computer vision, focusing on extracting high-level semantic information from image or video data, enabling the computer to automatically analyze, interpret, and reason about visual content. The environmental big data analysis method is an application based on image understanding technology. By processing and analyzing satellite remote sensing images, surveillance videos, and data obtained from sensors, it identifies environmental changes, abnormal activities, and ecological trends.
[0003] In the prior art, feature extraction is based on a fixed feature template, which is difficult to adapt to changes under different environmental conditions, and the feature expression ability is limited, reducing the abnormal recognition accuracy in complex scenarios. Abnormal detection mainly relies on current data, making it difficult to conduct long-term trend analysis, resulting in short-term fluctuations being easily misjudged as abnormalities, while potential environmental change trends are difficult to identify. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art, and to propose a big data analysis method and system based on a cloud business platform.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions. A big data analysis method based on a cloud business platform includes the following steps:
[0006] Collect satellite image data, integrate and format the satellite image data to generate a unified format image set; perform quality screening on the unified format image set, remove images with unqualified quality, and obtain a screened image set;
[0007] Based on the screened image set, use deep learning for feature extraction, extract environmental visual features from the images to generate image feature data, and based on the image feature data, analyze the features to identify abnormal environmental activities and obtain an abnormal detection result;
[0008] Match the abnormal detection result with a known abnormal activity database to verify the abnormal type and risk level, generate an abnormal activity result, and based on the abnormal activity result, adjust the detection parameters to generate adjusted detection parameters;
[0009] Use the adjusted detection parameters to continuously monitor the satellite image data and generate an environmental real-time monitoring result.
[0010] Preferably, the step of obtaining the unified format image set is:
[0011] Collect satellite image data, remove data items with inconsistent time, and perform color correction on the remaining data to obtain synchronized satellite image data;
[0012] Based on the synchronized satellite image data, adjust all images to a unified resolution setting and unify the color mode to obtain standardized satellite image data;
[0013] According to the standardized satellite image data, convert each image file to a unified file format to obtain a unified format image set.
[0014] Preferably, the steps for obtaining the filtered image set are as follows:
[0015] Extract the exposure level, contrast, and image integrity of each image from the unified format image set to obtain an image quality metadata list;
[0016] Based on the image quality metadata list, calculate the comprehensive quality score for each image. The calculation formula is:
[0017]
[0018] where Q represents the comprehensive quality score, E represents the exposure level, DC represents the contrast, I represents the image integrity, and XP represents the pixel density;
[0019] Remove images according to the comprehensive quality score to obtain a filtered image set.
[0020] Preferably, the steps for obtaining the image feature data are as follows:
[0021] Load the filtered image set and initialize the deep learning model, including setting the network layers and parameters, to obtain an initialized deep learning model;
[0022] Extract the environmental visual features of each image, including texture, shape, and color, through the initialized deep learning model to obtain an extracted feature data set;
[0023] According to the extracted feature data set, apply a pooling layer and an activation function to compress the features to obtain image feature data.
[0024] Preferably, the steps for obtaining the anomaly detection result are as follows:
[0025] Load the image feature data, parse the shape features, color features, and texture features of each image, construct a feature vector matrix, and perform normalization processing on all feature vector matrices to obtain a normalized feature vector matrix;
[0026] Based on the normalized feature vector matrix, calculate the environmental anomaly score. The expression is:
[0027]
[0028] where D i is the environmental anomaly score of the i-th image, represents the shape feature value of the i-th image on the j-th feature dimension, R i,j represents the mean shape feature on the j-th feature dimension in the normal samples, V i,j represents the shape change range of the i-th image on the j-th feature dimension, U i,j represents the local gradient change of the shape of the i-th image on the j-th feature dimension, Z i,j represents the color feature value of the i-th image on the j-th feature dimension, L i,j represents the mean color feature on the j-th feature dimension in the normal samples, G i,j represents the color change range of the i-th image on the j-th feature dimension, W i,j represents the local gradient change of the color of the i-th image on the j-th feature dimension, and xm is the total number of features;
[0029] Classify and judge the images according to the environmental anomaly scores to obtain the anomaly detection results.
[0030] Preferably, the steps for obtaining the abnormal activity results are as follows:
[0031] Query the anomaly detection results based on the anomaly detection results, and match them with the records in the known abnormal activity database to obtain the matched abnormal activity records;
[0032] Verify the anomaly types and risk levels of each record according to the matched abnormal activity records to obtain the abnormal activity results.
[0033] Preferably, the steps for obtaining the adjusted detection parameters are as follows:
[0034] Calculate the detection parameter adjustment value based on the abnormal activity results. The expression is:
[0035]
[0036] where P i is the i-th detection parameter adjustment value, X i,j represents the change amplitude of the i-th abnormal event on the j-th feature dimension, Y i,j represents the occurrence duration of the i-th abnormal event on the j-th feature dimension, T i,j represents the current detection threshold of the i-th abnormal event on the j-th feature dimension, B i,jrepresents the historical mean threshold of the i-th abnormal event in the j-th feature dimension, C i,j represents the number of frequent changes of the i-th abnormal event in the j-th feature dimension, N i represents the historical cumulative detection times of the i-th abnormal event, and m is the total number of features;
[0037] Adjust the detection threshold according to the adjusted value of the detection parameter to generate the adjusted detection parameter.
[0038] Preferably, the step of obtaining the real-time monitoring result is as follows:
[0039] Load the adjusted detection parameter, perform feature analysis on the image, extract the environmental change features in the image, mark the potential abnormal areas, and generate the environmental real-time monitoring result.
[0040] The present invention provides a big data analysis system, including:
[0041] A data acquisition module that receives the image data transmitted by the satellite, performs unified format processing on these data, and obtains a unified format image set;
[0042] An image screening module that screens images from the unified format image set, eliminates images with a resolution lower than the preset standard, and obtains an image set;
[0043] A feature extraction module that uses deep learning technology to extract environmental visual features from the image set, performs the conversion from image to features, and obtains an image feature data set; converts the image feature data set into the format required for feature analysis to obtain environmental feature description data;
[0044] An anomaly detection module that, based on the environmental feature description data, identifies features that do not conform to the normal pattern, obtains an anomaly detection result, compares the anomaly detection result with the anomaly activity database, verifies the anomaly type and risk level, adjusts and optimizes the detection logic, and generates an anomaly activity recognition result;
[0045] A monitoring optimization module that adjusts the monitoring parameters according to the anomaly activity recognition result, monitors the satellite image using the new parameters, and obtains the environmental real-time monitoring result.
[0046] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0047] In the present invention, deep learning is used for feature extraction to make the mining of environmental visual information more refined. It can identify various environmental features and generate high-dimensional feature data, improving the adaptability to complex environments. Anomaly detection combines real-time analysis with a historical anomaly database for matching. It not only verifies the anomaly type but also conducts risk assessment by combining past environmental change patterns, enabling anomaly recognition to not only rely on current data but also possess trend analysis capabilities. Parameter adjustment dynamically optimizes the detection mechanism based on the anomaly detection results. It can adapt the detection strategy according to different environmental conditions, reducing false detections and missed detections and improving the stability of long-term monitoring. Continuous monitoring analyzes image data in real time based on the adjusted detection parameters and dynamically optimizes the monitoring strategy in combination with historical trends, ensuring the adaptability of the environmental monitoring system to complex environments and making the monitoring results more continuous and accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0049] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0050] Please refer to Figure 1 , the present invention provides a technical solution, a big data analysis method based on a cloud business platform, including the following steps:
[0051] Collect satellite image data, integrate and format the satellite image data to generate a unified format image set; perform quality screening on the unified format image set, remove images with unqualified quality, and obtain a screened image set;
[0052] Based on the screened image set, use deep learning for feature extraction, extract environmental visual features from the images to generate image feature data, and based on the image feature data, analyze the features to identify abnormal environmental activities and obtain anomaly detection results;
[0053] Match the anomaly detection results with a known abnormal activity database to verify the anomaly type and risk level, generate abnormal activity results, and based on the abnormal activity results, adjust the detection parameters to generate adjusted detection parameters;
[0054] Use the adjusted detection parameters to continuously monitor the satellite image data and generate real-time environmental monitoring results.
[0055] The steps for obtaining the unified format image set are as follows:
[0056] Collect satellite image data, remove data items with inconsistent time, and perform color correction on the remaining data to obtain synchronized satellite image data;
[0057] Based on the synchronized satellite image data, adjust all images to a unified resolution setting and unify the color mode to obtain standardized satellite image data;
[0058] According to the standardized satellite image data, convert each image file to a unified file format to obtain an image set with a unified format.
[0059] Specifically, based on the collected original satellite image data, first read the shooting timestamp of each image and disassemble and extract the time identifier, lighting condition, and shooting geographical location information. Then, compare each of these time identifiers one by one with a record containing the Coordinated Universal Time deviation. If it is found that some time identifiers exceed the expected observation period by more than 2 minutes, they are determined to be time inconsistent and directly excluded. To determine this 2-minute judgment value, first calculate the statistical results for adjacent observation cycles within a week. For example, when the average deviation of the observation time within a week is 30 seconds and the variance is 10 seconds, add the average deviation to three times the variance (30 seconds + 3×10 seconds = 60 seconds) to obtain the initial judgment reference value. Then, multiply this reference value by 2 minutes as the exclusion threshold according to the current observation requirements. After excluding the time anomaly data, continue to perform color correction on the remaining images. During the correction, calculate the average red component, average green component, and average blue component of each image and compare them with a pre-established standard color comparison table. If the offset of the red, green, or blue component of an image is greater than 5%, gradient correction is initiated. This 5% is obtained by summarizing and analyzing the color deviation distribution range of one thousand normally shot samples (4% + 1% redundancy). After completing the color correction, centrally process all the data that has passed the time consistency check and color correction to obtain synchronized satellite image data.
[0060] Based on the synchronized satellite image data, first view the original resolution of each image and disassemble and extract the number of horizontal pixels and the number of vertical pixels. Then, check an observation reference table containing common resolution standards to determine whether interpolation scaling is required. For example, set the target resolution to 2048×2048 pixels and calculate the difference ratio between the current image and the target resolution. If the difference ratio exceeds 10%, perform resolution interpolation processing. The reason for choosing 10% as the threshold is that in historical samples, the minimum interpolation error ratio that occurred after various resolution scalings was approximately 8%. Adding 2% as a redundancy, the final setting of 10% is obtained. Subsequently, when unifying the color mode, it is necessary to extract the color space parameters of each image (including gamut, chromaticity coordinates, and white point information) and compare them with the parameters of a standard sRGB color space. By calculating the ΔE value to evaluate the difference between the two, if the ΔE value is greater than 2, perform color mode conversion. The determination of the threshold 2 is based on the statistical analysis of a large number of calibrated image samples, and the average ΔE is approximately 1.5. Leaving a redundancy of 0.5, 2 is obtained. After all images are unified in resolution and color mode, they are merged into a new output set to obtain standardized satellite image data.
[0061] According to the standardized satellite image data, it is necessary to detect the file format of each image one by one and disassemble and extract the file header information and encoding method. Then, define the final conversion format according to the general format description required by the user. For example, when uniformly using the lossless compression TIFF format, it is necessary to determine whether the bit depth of each image reaches 8 bits. If the bit depth is insufficient, use the methods of interpolation and grayscale expansion to enhance the corresponding level. The value of 8 bits comes from the statistical results of the lowest bit depth standard for high-fidelity images. If there are higher color level requirements, 16 bits can also be empirically set as the threshold. Subsequently, decide whether to adjust the compression ratio according to specific needs. To determine the compression ratio, multiple trial compressions can be first performed on ten representative images. If it is found that the color deviation of the compressed image does not exceed 3% and the loss of texture details does not exceed 2%, the compression ratio is fixed at approximately 5:1. If it exceeds the above threshold, gradually reduce the compression ratio. After completing the format conversion of all images, they are uniformly output in the specified format to obtain an image set with a unified format.
[0062] The steps to obtain the filtered image set are as follows:
[0063] Extract the exposure level, contrast, and image integrity of each image from the image set with a unified format to obtain a list of image quality metadata;
[0064] Based on the list of image quality metadata, calculate the comprehensive quality score of each image. The calculation formula is:
[0065]
[0066] Among them, Q represents the comprehensive quality score, E represents the exposure level, DC represents the contrast, I represents the image integrity, and XP represents the pixel density;
[0067] Remove images according to the comprehensive quality score to obtain the filtered image set.
[0068] Specifically, extract the exposure level, contrast, and image integrity of each image from the uniformly formatted image set. Combine the content of a file containing the reference range of the shooting light environment to disassemble and extract the shooting light information of each image. When extracting, first read the luminance channel data of each pixel line by line and calculate the average luminance of the whole image as the reference value of the exposure level. Then, according to a pre-organized luminance range comparison record, for example, divide the luminance into three intervals: 0 to 50, 50 to 150, and 150 to 255. Check which interval this image falls into item by item and mark the exposure level as the corresponding values of low, medium, and high accordingly. Subsequently, when calculating the contrast, it is necessary to first count the gray distribution of all pixels in the image. The interval of the gray distribution can be refined into three ranges: 0 to 50, 51 to 120, and 121 to 255. Compare each range with the number of pixels in the original image data. If the gray concentration is higher than 60%, set the contrast value of this image to be above 1.2. If the gray is evenly distributed in both intervals, the contrast value is set between 1.0 and 1.2. If the gray is evenly dispersed throughout the range, the contrast value is less than 1.0. Next, before extracting the image integrity, check whether there are damaged or missing row and column pixels in the image. If the number of damaged pixels is more than 50, it is determined that the integrity is low. If it is between 10 and 50, the integrity is considered medium. If it is less than 10, it is recorded as good integrity. The determination of relevant thresholds is set based on the retrieved image defect probability distribution. For example, by detecting 500 conventional images and counting the frequency distribution of damaged pixels, it is found that the sample number with damaged pixels concentrated between 30 and 40 is the largest. Therefore, in the final setting, 50 pixel damages are used as the lower limit of integrity, and 10 pixel damages are used as the higher standard of integrity. The integrity value of each image is disassembled into a numerical form according to the above standard and recorded. After completing the extraction of the exposure level, contrast, and image integrity, organize these values into an image quality metadata list to obtain the image quality metadata list.
[0069] The advantage of the formula is that by applying different coefficient weights to the exposure level, contrast, and image integrity in the numerator part, it can take into account brightness, visual clarity, and defect conditions in image quality evaluation, and introduce the combination of pixel density and logarithmic function in the denominator part to balance the additional quality impact brought by high resolution.
[0070] The steps to obtain E are as follows:
[0071] When measuring each image, its grayscale value is recorded point by point and weighted averaged to obtain the overall brightness index. Then, referring to the established brightness standard range of 0 to 255, the index is converted to a normalized interval between 0 and 1. The specific calculation formula can be written as: Where M represents the total number of pixels in the image. The grayscale value of each pixel is obtained through the digital output of the acquisition instrument, and the value range is between 0 and 255. If the average grayscale value of an image is 130, it is substituted into the calculation to obtain
[0072] The steps to obtain DC are:
[0073] Contrast is usually achieved by traversing the pixel grayscale and calculating the brightness extreme difference. To obtain a more accurate value, the peak-to-valley difference of the grayscale distribution can be counted first, and then the result can be calibrated based on the proportion of the overall distribution peak. G means dividing the grayscale range from 0 to 255 into several small intervals. The max (grayscale) and min (grayscale) can be directly obtained by traversing all pixels. k is the number of pixels in the kth interval, the grayscale peak distribution coefficient k It is the value obtained by quantifying the contrast influencing factors of each interval in the previous shooting environment record. For example, when the grayscale interval of 10 to 20 accounts for 30% of the total pixel volume and the corresponding contrast influence coefficient is set to 1.1, the pixel number k Take 0.3×M, grayscale peak distribution coefficient k =1.1, if the grayscale of the entire image is evenly distributed in four to five main intervals, the contrast value can be obtained by accumulating the calculations of each interval.
[0074] The steps to obtain I are:
[0075] Image integrity is mainly achieved by detecting whether pixels are damaged, whether rows and columns are aligned, and whether there is resolution loss. Where M still represents the total number of pixels in the image. The total number of bad pixels can be checked pixel by pixel when decoding the image. If non-compliant data appears, it will be counted once. The number of defective rows and columns indicates the number of rows and columns that are severely offset or lost. α is a conversion factor that quantifies the row and column damage to between 0 and 1. The conversion factor can be set according to the scanning and recognition accuracy of the equipment. For example, in daily monitoring, if a defective row is detected, the corresponding integrity decreases by 0.02, then α can be set to 50. In this way, each missing row or column will produce a 0.02 deduction in the integrity value. Taking an image with M of 3 million pixels as an example, if the total number of bad pixels detected is 45 and the number of defective rows and columns is 1, then
[0076] The steps to obtain XP are as follows:
[0077] The pixel density can be measured by calculating the number of pixels per unit area. Let where β represents the actual area covered by the sensor, usually in square millimeters, which can be directly obtained by reading the specification parameters of the imaging device. For example, for a sensor with a size of 4.5mm × 3.0mm, the covered area is 13.5 square millimeters. If the total number of actual pixels in the image is 3 million, then
[0078] Calculation process:
[0079] Now, taking E = 0.51, DC = 1.12, I = 0.98, and XP = 222222 as actual examples, substitute them step by step:
[0080] First step, calculate the numerator first:
[0081] 0.3E 2 = 0.3×(0.51) 2 ≈0.3×0.2601 = 0.07803
[0082] 0.5DC 2 = 0.5×(1.12) 2 ≈0.5×1.2544 = 0.6272
[0083] 0.2I 2 = 0.2×(0.98) 2 ≈0.2×0.9604 = 0.19208
[0084] 0.07803 + 0.6272 + 0.19208 = 0.89731
[0085]
[0086] Second step, calculate the denominator:
[0087] 1 + ln(1 + XP) = 1 + ln(1 + 222222) = 1 + ln(222223) ≈ 1 + 12.307 ≈ 13.307
[0088] Third step, divide the numerator by the denominator:
[0089]
[0090] The results show that when the Q value is around 0.07, it can be used as a quantitative reference for measuring the comprehensive quality of the image. In a large number of image statistics, if the Q value is greater than 0.2, it means that its exposure, contrast and integrity are all good. On the contrary, if the Q value is lower than 0.05, it indicates that there are at least problems such as too low exposure, unclear contrast or serious pixel defects. In the subsequent steps, the images can be screened or excluded according to this Q value.
[0091] Remove images according to the comprehensive quality score. First, sort the previously calculated comprehensive quality scores, and analyze whether the score is less than 0.05 or greater than 0.2 one by one by comparing with a list of score thresholds obtained from the statistical analysis of measured image samples. The list of score thresholds comes from the range of measured scores of 3,000 satellite images. It is statistically found that images with scores lower than 0.05 often have insufficient brightness or serious pixel damage during subsequent processing. Therefore, 0.05 is used as the lower limit for exclusion. If the Q value is between 0.05 and 0.2, it is marked as average quality, and if it is greater than 0.2, it is marked as high quality. Extract the images with Q values less than 0.05 for removal. When removing, directly mark them in the index list and exclude these images from the subsequent analysis process. Then retain the images with Q values between 0.05 and 0.2 and greater than 0.2, and record the indexes of the retained images. If there are a large number of Q values in the image set concentrated near the 0.05 threshold, the specific components of E, DC, and I of these images can be further refined and compared, and whether a second judgment is needed can be confirmed by checking the exposure deviation or the distribution of bad points. For some images that still do not meet the usage requirements after manual review, they are also removed in this step. Finally, all qualified images are integrated to obtain the screened image set.
[0092] The steps for obtaining the image feature data are as follows:
[0093] Load the screened image set and initialize the deep learning model, including setting the network layers and parameters to obtain the initialized deep learning model;
[0094] Extract the environmental visual features of the images one by one through the initialized deep learning model, including texture, shape and color, to obtain the extracted feature data set;
[0095] According to the extracted feature data set, apply the pooling layer and activation function to compress the features to obtain the image feature data.
[0096] Specifically, load the recorded filtered image set, first read the index number corresponding to each image and its pixel distribution and size information one by one, and when extracting this information, first confirm whether the filtered image set contains image identifiers of the same shooting batch or the same observation period, and then refer to a record containing neural network pre-training parameters and network structure configuration to disassemble and extract directly usable values such as network layers, convolution kernel size, and activation method. If it is found that the convolution kernel size specified in the record is listed as 3×3 and the number of network layers is 8, then compare the number of network layers with the number of images to be processed later. If the total number of images is greater than 5000 and a convolution operation of approximately 0.02 seconds is required at each layer, then the total time is calculated and further adjusted according to the device performance and observation time limit. When continuing to clarify the learning rate and batch size, it is necessary to refer to the data results accumulated in the previous multiple rounds of testing processes. The learning rate will be set to 0.001 to 0.005 according to the convergence of each batch of data. The specific value can be found in the fluctuation range in a record containing multiple training time and error curves, and the error is selected when the error drops to 0. .01 is taken as the final learning rate. For example, the batch size is directly related to the memory capacity. If the current memory allows 16 images to be loaded per batch, the batch size is set to 16. If overload occurs, it is adjusted to 8. After completing the setting of the learning rate and batch size, continue to set the number of training rounds and the number of convolution channels. For example, from a sample for starry sky image processing, it is found that the error can be relatively stable after 12 rounds of training. At the same time, the number of convolution channels is statistically found to have an acceptable processing speed when the initial setting is 32 channels. If more sophisticated feature extraction is required, it is increased to 64 channels. After all parameters are disassembled and determined, the screened image set is divided into two parts: training and verification. On this basis, the images are loaded in turn and their grayscale distribution, color components and texture changes are preliminarily regularized. When setting the regularization, if the average deviation of the pixel value is within 0.1, no additional adjustment is made. If it exceeds 0.1, it is normalized to the range of 0 to 1 to avoid affecting the stability of subsequent training. All initialization operations are completed and the above-configured network structure is called to create a deep learning model, and finally an initialized deep learning model is obtained.
[0097] When extracting the environmental visual features of an image one by one through the initialized deep learning model obtained previously, first retrieve the texture distribution in the first convolutional operation of the model. The specific method is to divide each image into several 10×10 sub-regions when reading each image and calculate the average gray gradient value in each sub-region. When the average gray gradient value of some sub-regions is greater than 0.8, it is marked as a high-texture change area. This 0.8 value is the average value obtained from the texture gradient distribution results of thousands of satellite samples plus an offset of 0.1. Subsequently, summarize all the sub-region markings and pass the extracted texture information backward. Then, identify the image shape in the second layer. First, detect the edge pixels and calculate the connected regions. If the aspect ratio of the connected region is greater than 2 and the number of pixels exceeds 1000, it is regarded as a long strip region. These two values, 2 and 1000, are also obtained through statistical analysis of the previous shape distribution data. Among them, 2 means that when the aspect ratio exceeds this value in the target scene, it has an obvious stretching feature, and 1000 represents that it can be basically regarded as an identifiable block at a resolution of 1024×1024. For color feature extraction, it will be completed in the third layer. First, decode the red, green, and blue components of each image and compare the pixel mean and variance of each component with a color feature reference range. When the proportion of the mean value of the red component between 0.2 and 0.4 and the mean value of the green component between 0.3 and 0.5 exceeds 80%, it can be determined that the color distribution is relatively uniform, otherwise it is marked as a color deviation area. Finally, perform multiple convolutional and normalization operations in the subsequent hidden layers of the network to stabilize these texture, shape, and color feature values. During the process, if it is found that the activation values extracted for some features are too high or too low, the learning rate can be further adjusted and the training can be restarted. Use dozens of images with known comparison results as the validation set for multiple rounds of iteration to gradually converge the model parameters. Stop training when the number of training rounds is sufficient and the validation error is less than 0.005. Extract the complete environmental visual features such as texture, shape, and color and summarize them into a feature dataset, and then obtain the extracted feature dataset.
[0098] According to the previously obtained feature dataset, when applying the pooling layer and activation function for feature compression, first downsample the dimensions of each feature vector in the pooling layer. The originally high-resolution feature map is evenly split into sub-blocks of 4×4 or 8×8. If the mean value of all feature values within a sub-block is higher than 0.6 and the variance is lower than 0.1, the entire block of data is regarded as a homogeneous region. If the variance is greater than 0.3, it is regarded as a heterogeneous region. These values are determined with reference to the common texture differences and color distribution differences in the feature statistics of multiple previous environmental images. After completing the pooling operation, the extracted pooling output is sent to the activation function for re-mapping. At this time, if a function based on a fixed threshold is used, the threshold needs to be set between 0.2 and 0.8. Through preliminary experimental observations, when the activation threshold is too high, the features tend to be 0 or 1, affecting the retention of subtle differences. If the activation threshold is too low, most feature values are concentrated in a narrow range and are difficult to distinguish. Therefore, it is fixed at 0.5 and allowed to float up and down by 0.1 to cope with different scenarios. Through multiple rounds of iteration, confirm the correspondence between the activated features and the textures and shapes in the actual environment on the training set and further compress the activated vector data. Then, through re-normalization processing, all feature values are maintained in the interval of 0 to 1. If any feature value is detected to exceed 1, it is restricted within 1 to prevent subsequent calculation errors. After the entire compression process is completed, the final compressed environmental visual feature set will be output and summarized for recording. Through this set, target localization, abnormal area discrimination, or other operations can be performed in subsequent steps, and finally the image feature data is obtained.
[0099] The steps to obtain the anomaly detection result are as follows:
[0100] Load the image feature data, parse the shape features, color features, and texture features of each image, construct a feature vector matrix, and perform normalization processing on all feature vector matrices to obtain a normalized feature vector matrix;
[0101] Based on the normalized feature vector matrix, calculate the environmental anomaly score, and the expression is:
[0102]
[0103] where D i is the environmental anomaly score of the i-th image, represents the shape feature value of the i-th image in the j-th feature dimension, R i,j represents the mean value of the shape features in the j-th feature dimension in the normal samples, V i,j represents the shape change range of the i-th image in the j-th feature dimension, U i,j represents the local gradient change of the shape of the i-th image in the j-th feature dimension, Z i,j represents the color feature value of the i-th image in the j-th feature dimension, Li,j represents the mean color feature in the j-th feature dimension of the normal sample, G i,j represents the color change range of the i-th image in the j-th feature dimension, W i,j represents the local gradient change of the color of the i-th image in the j-th feature dimension, where xm is the total number of features;
[0104] Based on the environmental anomaly score, classify and judge the images to obtain the anomaly detection result.
[0105] Specifically, after loading the image feature data obtained previously, it is necessary to individually retrieve the specific values of the shape features, color features, and texture features in each image. Most of these values are derived from the analysis of the pixel distribution in the image. When analyzing, the edge contour information of each image will be read in and the main contour and several sub-contours will be separated from it. Then, the corresponding length values, width values, and radian change values will be disassembled and extracted by referring to a record containing standard shape elements. Next, in the color feature processing, the brightness ratios and main color tone distributions of the red, green, and blue components will be measured respectively. If the brightness ratio of a certain component is higher than 50% of the entire image, the image will be marked as a color deviation scene. Then, in the texture feature processing, the local gradient statistical method will be used to traverse each pixel, collect the gray gradient distribution of each local area and summarize the recorded results. If it is found that the gradient difference is concentrated and greater than 0.2, it is regarded as a relatively obvious texture change. This 0.2 value is the average gray difference threshold statistically obtained by referring to 300 environmental images with representative texture distributions. Then, all the shape, color, and texture data that have been analyzed will be combined into a feature vector matrix with the same dimensional structure, and the consistency between the number of rows of the matrix and the number of images and the integrity of each column component will be checked. If there are blank or out-of-limit abnormal values, they will be corrected by interpolation or elimination methods. Finally, when uniformly normalizing these feature vector matrices, the maximum-minimum normalization method can be used to map the range of each component to between 0 and 1, or zero-mean normalization processing can be performed when the mean and standard deviation are known. When normalizing, if a component appears outside the range of 1 or below 0, it will be truncated to a reasonable interval by the clipping method. After completing the normalization of all feature vector matrices, these matrices will be recorded to obtain the normalized feature vector matrix.
[0106] The benefit of the formula is that it uses two major feature dimensions of shape and color to measure the environmental differences, statistically calculates the deviation based on the absolute difference in the numerator part, and introduces the change range and local gradient change of shape and color in the denominator part. By adding 1 and taking the square root, it suppresses the influence of excessive or too small noise, enabling the overall features and local differences to be considered when detecting anomalies.
[0107] parameter ET i,j The acquisition steps are as follows:
[0108] First, identify the j-th shape element of the i-th image in the image. After inputting the image into a pre-trained contour detection model, several potential contours are obtained. Then, the lengths, aspect ratios, and endpoint curvatures corresponding to the contours are statistically analyzed, so as to quantify these shape information and superimpose them into a comprehensive value. This comprehensive value can make ET i,j = α1×contour length + α2×aspect ratio + α3×curvature, where the values of α1, α2, and α3 are obtained by fitting the measurement data for different shapes in the early stage. For example, by measuring the contours of 500 common surface buildings, the length range is 100 to 500 pixels, the aspect ratio is 1.0 to 3.0, and the curvature is 0.01 to 0.2. The empirically derived weight values are α1 = 0.5, α2 = 1.5, and α3 = 2.0. Finally, when the contour length is 300, the aspect ratio is 2.2, and the curvature is 0.05, ET i,j = 0.5×300 + 1.5×2.2 + 2.0×0.05 = 150 + 3.3 + 0.1 = 153.4.
[0109] Parameter R i,j The acquisition steps are as follows:
[0110] This is the mean value of the j-th shape feature corresponding to the normal sample. To obtain R i,j , it is necessary to first extract the same shape elements from the targets of the same category or similar scenes in the image dataset under normal conditions, accumulate the corresponding quantitative indicators such as shape length, aspect ratio, and curvature, and calculate the arithmetic mean. If the shape elements extracted from 300 normal samples are {x1, x2,..., x 300}}, then it can be defined as When the shape elements in each batch of normal images are statistically analyzed, a series of means corresponding to different j can be obtained. Taking the mean result of about 150.0 of the j-th element in a certain scene in 300 normal images as an example, then R i,j = 150.0.
[0111] Parameter V i,j The acquisition steps are as follows:
[0112] This is the change range of the i-th image in the j-th shape feature dimension. It is necessary to first calculate the difference between the maximum and minimum values of the basic data such as the extracted shape length, aspect ratio, and curvature, and quantify the difference between them into an interval measure. For example, it can be denoted as range = max(set of element values) - min(set of element values). Then, according to the environmental requirements or image resolution, introduce a normalization factor β1 to convert this range to the range of 0 to 1. The process can be written as The value of β1 can be set by analyzing the shape ranges of hundreds of images. For example, if the detected minimum possible length is about 50 pixels and the maximum possible length is about 800 pixels, taking β1 = 750 makes V i,j fall between 0 and 1. Assuming that in a certain detection, max(set of element values) = 350 and min(set of element values) = 60, then range = 290,
[0113] Parameter U i,j is obtained as follows:
[0114] This is the local gradient change value of the i-th image in the j-th shape feature dimension. It is necessary to traverse the possible local regions in the image and perform differential statistics on the shape edge changes within the local regions. For example, taking a first-order difference of the edge coordinate function f(x) to obtain Then taking the absolute value, accumulating, and calculating the mean value, and normalizing the result to the range of 0 to 1 for direct comparison with other feature values. Let The mean value of is denoted as m grad , then where γ1 can be selected as a fixed value from multiple gradient tests on the contours of common buildings. For example, after measuring the average value of the gradient range as 10 for 200 local samples, γ1 can be set to 10.0, making U i,j become a floating value between 0 and 1. When the average value of the edge coordinate gradient is detected to be 4 at a certain location,
[0115] Parameter Z i,j is obtained as follows:
[0116] This is the value of the i-th image in the j-th color feature dimension. It is necessary to quantitatively analyze the chromaticity distribution of the image based on the color component decomposition results. For example, separately counting the total pixel brightness of the red, green, and blue channels, then dividing by the corresponding number of pixels to obtain the average brightness, and then mapping it to between 0 and 1 according to on-site experience or normalization of the color values. When the original average brightness of a certain channel is 120 and the maximum value of this channel can reach 255,
[0117] Parameter L i,j is obtained as follows:
[0118] This is the j-th color feature mean value corresponding to the normal sample. Similar to the shape mean value calculation method, it is necessary to first extract the average brightness or chromaticity values of the same channels one by one from the reference normal color distribution samples, and record these data uniformly as {y1, y2,..., y n}, and then let If n = 400 images are accumulated for a common scenario, and the average brightness of each image's color channels is distributed between 0.4 and 0.6, then an average value, such as approximately 0.5, can be obtained by summing and dividing by 400. All subsequent images can reference this mean value for comparison.
[0119] Parameter G i,j The acquisition steps are as follows:
[0120] This is the color change range of the i-th image in the j-th color feature dimension. It is necessary to similarly count the brightest and darkest pixels in this channel, normalize the difference to colorRange, and then map it to the range of 0 to 1 in combination with the normalization factor β2 used in the scenario, which can be written as β2 is obtained by analyzing the distribution of the brightest and darkest pixel values of a large number of images. For example, in daily monitoring, the darkest may be a brightness of 0 and the brightest is 255, so β2 = 255 can be set. If the brightest value is 220 and the darkest value is 40 in a certain image, then colorRange = 180.
[0121] Parameter W i,j The acquisition steps are as follows:
[0122] This is the local color gradient change value of the i-th image in the j-th color feature dimension. It is necessary to statistically analyze the gradient distribution of the color channels at the pixel level, calculate the difference in the channel brightness for each local area, and calculate the mean value of | brightness|, and then compare it with a reference coefficient γ2 to normalize it to the range of 0 to 1. When the brightness mean is approximately d, it can be set as γ2 comes from the centralized estimation of the color gradients of a large number of normal images before. For example, if the statistical distribution range of the brightness gradient is between 0 and 20, then γ2 = 20 can be set. If the mean gradient of a local area is actually detected to be 5, then
[0123] The acquisition steps of parameter xm are as follows:
[0124] This is the total number of features, which is used to mark the total number of feature dimensions divided by each image under the two major categories of shape and color. How many dimensions the shape features or color features are specifically divided into generally depends on the actual on-site requirements and the detail of the acquisition. For example, for the shape, it can be divided into three dimensions: contour length, aspect ratio, and curvature. For the color, it can be divided into three channels of red, green, and blue plus extended dimensions such as saturation or brightness. If there are a total of 6 dimensions, then xm = 6.
[0125] Calculation process:
[0126] Now, for example, take xm = 2 and select the following parameters:
[0127] ETi,1 = 1.20, ET i,2 = 1.30, R i,1 = 1.00, R i,2 = 1.10, V i,1 = 0.30, V i,2
[0128] = 0.25, U i,1 = 0.05, U i,2 = 0.04, Z i,1 = 0.70, Z i,2 = 0.72, L i,1 = 0.65, L i,2
[0129] = 0.70, G i,1 = 0.20, G i,2 = 0.22, W i,1 = 0.05, W i,2 = 0.06
[0130] Step 1, calculate the numerator of the shape part:
[0131]
[0132] Step 2, calculate the denominator of the shape part:
[0133]
[0134] Step 3, obtain the result of the shape part:
[0135]
[0136] Step 4, calculate the numerator of the color part:
[0137]
[0138] Step 5, calculate the denominator of the color part:
[0139]
[0140] Step 6, obtain the result of the color part:
[0141]
[0142] Step 7, add the shape part and the color part:
[0143] D i = 0.3125 + 0.0566 = 0.3691
[0144] The result shows that when D iWhen it is approximately 0.37, it can be compared with the previously statistically normal sample threshold. For example, if the normal sample threshold is set to 0.50, then 0.37 being less than 0.50 means that the difference between this image and the normal image is not significant. Subsequently, this value can be compared with the anomaly scores of more images for comparative analysis to determine which images have more prominent environmental differences.
[0145] Based on the environmental anomaly scores obtained previously, it is necessary to retrieve and compare the score values item by item in the index corresponding to each image. First, disassemble and extract the anomaly judgment criteria from a record containing conventional threshold configurations. For example, when it is statistically found that the score range of a large number of normal samples is between 0.1 and 0.5, 0.5 can be set as the anomaly judgment boundary. The threshold can also be fine-tuned to around 0.6 by referring to the score distribution in some extreme scenarios. Then, through a loop, compare whether the D i value of each image exceeds the set threshold. If it is found that the D i value of a certain image reaches or is greater than 0.6, then mark it as an abnormal image. If it is between 0.3 and 0.6, it is determined as an image to be observed, and further manual inspection or additional monitoring of relevant shape and color features is required in the future. If it is lower than 0.3, it is classified as a basically normal sample. The reason for taking 0.3 as the reference for the lower range is based on a relatively stable distribution lower limit formed when counting thousands of images that are manually confirmed to be normal and have extremely low differences. For images between 0.3 and 0.6, the floating conditions of the specific shape and color components can be checked again to see if they exceed certain specific thresholds. For example, check whether the shape change range is greater than 0.4 or the color gradient change is greater than 0.5. These more refined bases can be found in the features obtained previously. When certain features are confirmed to be significantly abnormal, they can also be added to the anomaly list. Finally, after uniformly recording all the judgment results, the anomaly detection results can be obtained.
[0146] The steps to obtain the abnormal activity results are as follows:
[0147] Based on the anomaly detection results, query the anomaly detection results and match them with the records in the known abnormal activity database to obtain the matched abnormal activity records;
[0148] According to the matched abnormal activity records, verify the anomaly type and risk level of each record to obtain the abnormal activity results.
[0149] Specifically, based on the anomaly detection results obtained previously, first read the numbers of each anomaly record and their corresponding environmental anomaly scores, shape feature labels, and color feature labels. Then, compare these numbers and feature information item by item in a database record containing anomaly activity types and geographical area identifiers. When disassembling and extracting the matching fields of this database, it is necessary to first clarify the range of feature values for each anomaly activity. For example, divide the environmental anomaly scores into three intervals: 0.3 to 0.5, 0.5 to 0.8, and 0.8 to 1.0; divide the shape feature labels into three categories: polygon edges, linear extensions, and block distributions; divide the color feature labels into several categories such as water area blue, vegetation green, and soil brown. Then, query the scores and labels in the current anomaly detection results according to the same or similar intervals and categories. If a record with the same category and a matching score interval is found, it is considered a preliminary match. If no exactly matching record is retrieved in the first query step, float the score up and down by 0.1 and perform a fuzzy match again. The reason for choosing 0.1 is that, referring to the environmental anomaly score distributions of a large amount of historical data, it is found that floating up and down by 0.1 can cover most of the error ranges when comparing other scenarios. When multiple potentially matching database records are queried, select the one closest to the score of the current anomaly detection result for the final match, and organize this matching information to form a matched anomaly activity record.
[0150] According to the matched anomaly activity records obtained previously, read the anomaly types and possible risk level elements indicated in the records item by item. When disassembling and extracting each event label, first check the historical occurrence frequency and corresponding influence range of the event. For example, in the database, divide the occurrence frequency into three levels: 1 to 5 times, 6 to 10 times, and 11 times and above; divide the influence range into three levels: local, regional, and large-scale. Then, compare the feature labels in the record with the geographical coordinates of the anomaly detection results to determine which level the spatial range where the event is located belongs to. If it is found that an event has occurred more than 6 times and is in the regional range in combination with the occurrence frequency, it can be classified into the medium or higher level in terms of risk level. If it has occurred more than 10 times and is in the large-scale level, it is marked as a higher risk. The specific risk level classification needs to read a previously compiled risk coefficient comparison table. For example, in the comparison table, the medium risk coefficient is set in the range of 0.3 to 0.6, the high risk coefficient is set in the range of 0.6 to 1.0, and below 0.3 is recorded as low risk. If the risk coefficient of the corresponding record is around 0.65, it is finally determined as a high risk and the event is marked in the high risk category. If the risk coefficient is 0.25, it is regarded as a low risk category. After completing the confirmation of the types and risk levels of all matched anomaly activity records, the anomaly activity results are obtained.
[0151] The steps for obtaining the adjusted detection parameters are as follows:
[0152] Based on the abnormal activity results, calculate the adjustment value of the detection parameter, and the expression is:
[0153]
[0154] Wherein, P i is the adjustment value of the i-th detection parameter, X i,j represents the change amplitude of the i-th abnormal event in the j-th feature dimension, Y i,j represents the occurrence duration of the i-th abnormal event in the j-th feature dimension, T i,j represents the current detection threshold of the i-th abnormal event in the j-th feature dimension, B i,j represents the historical mean threshold of the i-th abnormal event in the j-th feature dimension, C i,j represents the number of frequent changes of the i-th abnormal event in the j-th feature dimension, N i represents the historical cumulative detection times of the i-th abnormal event, and m is the total number of features;
[0155] According to the adjustment value of the detection parameter, adjust the detection threshold to generate the adjusted detection parameter.
[0156] Specifically, the benefit of the formula is that it introduces multiple factors such as the deformation amplitude, occurrence duration, difference between the current detection threshold and the historical mean threshold, and the number of frequent changes and historical cumulative detection times. Each factor comes from the multi-dimensional measurement of abnormal events, and the square root operation form of the sum of squared differences is used in the denominator part, and the logarithmic function is combined at the end to comprehensively consider multiple frequent changes, so that the final adjustment value of the detection parameter can more comprehensively balance the current abnormal intensity and historical reference, thus helping to more precisely reflect the influence of each dimension in the subsequent threshold adjustment.
[0157] X i,j The acquisition steps are as follows:
[0158] The change amplitude of the i-th abnormal event in the j-th feature dimension is mainly obtained by comparing the quantization differences corresponding to the normal environment and the current event environment. If the feature dimension corresponds to the water pollution concentration, it is necessary to first calculate the average concentration within a period of time as the normal baseline, then take the difference between the currently detected concentration and this baseline and take the absolute value, denoted as diff. Then, to make the values adapt to the unified dimension, a normalization factor α can be introduced to compress diff to between 0 and 1, so that α can be obtained from the difference between the upper and lower limits of the concentration obtained from long-term monitoring. For example, if the concentration is usually between 0mg / L and 300mg / L, then α can be taken as 300. If the concentration detected in this event deviates from the baseline by 120mg / L, then
[0159] Y i,j The acquisition steps are as follows:
[0160] This is the occurrence duration of the i-th abnormal event in the j-th feature dimension. It is necessary to record the start time and end time of the event first, calculate the time difference, then convert the time difference to a unified unit (such as hours or days), and finally normalize it to the interval from 0 to 1. The method can be written as β is the statistical value of the longest duration for this feature dimension. For example, if the highest duration found in multiple historical records is 72 hours, then β can be taken as 72. When the duration of this event in a certain dimension is 24 hours, then
[0161] T i,j The acquisition steps of
[0162] It represents the current detection threshold of the i-th abnormal event in the j-th feature dimension, which can be read from the threshold set by the system during real-time monitoring. For example, in water quality monitoring, a record will indicate that the abnormal threshold of the PH value is set in the range of 6.0 to 9.0, the dissolved oxygen threshold is 6 mg / L, the temperature threshold is 35 °C, etc. If the current threshold for a certain dimension has been determined to be 35, and this dimension points to the temperature feature, then T i,j That is, it is recorded as 35. For the convenience of subsequent comparison or normalization in the calculation process, this value can also be mapped to the range from 0 to 1 according to the actual situation. For example, if considering the interval distribution of the temperature from 0 °C to 50 °C, then
[0163] B i,j The acquisition steps of
[0164] It represents the historical mean threshold of the i-th abnormal event in the j-th feature dimension, which mainly comes from the mean value of the data accumulated through long-term monitoring. For example, in a record of one year or longer, continuous monitoring is carried out for the temperature dimension of a certain area and its average abnormal alarm critical value is statistically calculated, and then this value is encapsulated as a reference threshold. If the historical mean threshold is statistically calculated to be 30 °C, denoted as base = 30, and if the temperature range from 0 °C to 50 °C is still used for normalization, then
[0165] C i,j The acquisition steps of
[0166] This is the number of frequent changes of the i-th abnormal event in the j-th feature dimension. It is necessary to record the frequency of fluctuations or overlimits in this dimension during the monitoring process. Once a significant exceedance or fall below the established threshold occurs, it is recorded as 1 time, and the total number of times is accumulated until the final total is obtained. For example, in water body monitoring, it is found that the pollution index has flipped up and down many times continuously for three days, and 4 times are statistically counted every day. Then the total number of frequent changes within three days can reach 12 times. Then, for subsequent calculations, appropriate normalization can be carried out. For example, if the maximum number of frequent changes within the observation interval is denoted as γ, then If γ is 20 and the actual number of frequent changes this time is 12, then
[0167] N i is obtained as follows:
[0168] This is the historical cumulative detection times of the i-th abnormal event. By querying the previous monitoring records, check how many observation cycles the event has been repeatedly detected in the past and sum up this value. If the event has 5 detection records in dimensions such as temperature and pH value in one year, then N i = 5.
[0169] The steps to obtain m are as follows:
[0170] This is the total number of features. It is necessary to confirm how many quantifiable feature dimensions are included in the analysis of abnormal events in the previous text. For example, in environmental monitoring, each of temperature, pH value, chromaticity, turbidity, dissolved oxygen, etc. can be regarded as a feature, and the sum is m. When performing the calculation of (P i ), it only needs to correspond to all the previously defined dimensions. If 5 features are finally listed, then m = 5.
[0171] Calculation process:
[0172] The i-th abnormal event contains 3 feature dimensions, that is, m = 3, and the following values are obtained previously:
[0173] X i,1 = 0.40, X i,2 = 0.30, X i,3 = 0.50
[0174] Y i,1 = 0.33, Y i,2 = 0.20, Y i,3 = 0.25
[0175] T i,1 = 0.70, T i,2 = 0.60, T i,3 = 0.55
[0176] B i,1 = 0.60, B i,2 = 0.50, B i,3 = 0.52
[0177] C i,1 = 0.60, C i,2 = 0.40, C i,3 = 0.50
[0178] N i = 5
[0179] First step, calculate the numerator first
[0180] (0.40×0.33)+(0.30×0.20)+(0.50×0.25)=0.132+0.06+0.125=0.317
[0181] Second step, calculate in the denominator
[0182] (0.70 - 0.60) 2 +(0.60 - 0.50) 2 +(0.55 - 0.52) 2 =0.10 2 +0.10 2 +0.03 2
[0183] =0.01+0.01+0.0009=0.0209
[0184]
[0185] Third step, calculate the logarithmic term
[0186]
[0187] N i +1=5+1=6
[0188]
[0189] ln(1 + 0.25)=ln(1.25)≈0.2231
[0190] Fourth step, combine the numerator and denominator, and then add the logarithmic term:
[0191]
[0192] The result shows that when the detection parameter adjustment value P i ≈0.537, it can be compared with a reference range set based on historical monitoring experience. If the reference range stipulates that P i is significantly increased to the next detection threshold only when it is greater than 0.7, and only slightly increased between 0.4 and 0.7, and remains unchanged when it is less than 0.4, then 0.537 will fall within the range of 0.4 to 0.7, indicating that this adjustment has a medium impact on the threshold, and the threshold of temperature or pollution concentration can be specifically increased by a certain proportion, and then refined according to the various dimensional situations obtained previously.
[0193] According to the detection parameter adjustment value obtained above, read a record containing the threshold update rules and parameter mapping instructions, and compare the interval of the detection parameter adjustment value one by one with the entries therein. If the record marks the interval of 0.3 to 0.5 as a small increase threshold, the interval of 0.5 to 0.7 as a moderate increase threshold, and the interval above 0.7 as a large increase threshold, first confirm which interval the current detection parameter adjustment value falls into. For example, if the detection parameter adjustment value is around 0.55, it belongs to the moderate increase threshold range, and then disassemble and extract the corresponding increase ratio in the record, such as indicating that when the detection parameter adjustment value is 0. 5 to 0.7, the temperature threshold increment can be set between 1°C and 3°C, and the concentration threshold increment can be set between 10mg / L and 20mg / L. The specific values of these increments come from the reliable range obtained after multiple threshold adjustments within a year, and the actual adjustment range list of each abnormal event is cross-analyzed with the current monitoring needs. For example, if the area has low temperatures all year round or water pollution is not serious, the lower limit of 1°C or 10mg / L is tended to be taken. Otherwise, the upper limit of 3°C or 20mg / L can be selected as the threshold update value for this time. The completed adjustment results are summarized to obtain the adjusted detection parameters.
[0194] The steps to obtain real-time monitoring results are as follows:
[0195] Load the adjusted detection parameters, perform feature analysis on the image, extract the environmental change features in the image, mark potential abnormal areas, and generate real-time environmental monitoring results.
[0196] Specifically, load the adjusted detection parameters obtained above, first read the latest temperature threshold, pollution concentration threshold and other similar thresholds arranged by feature dimension, then read the current image one by one and count the numerical distribution of each image in terms of temperature, color or physical shape mark, if the difference between the average value of temperature distribution and the new threshold is less than 1°C, keep the normal label, if it exceeds the threshold, mark the image as a potential abnormal area, and perform the same detection on the pollution concentration, if the concentration mean detected by concentrated distribution has exceeded the new threshold, it is recorded as suspicious, if it falls within 0.2 times of the threshold, it can be recorded as critical to be checked, the 0.2 times is the compensation ratio obtained after summarizing the error fluctuation in a large amount of data, after analyzing all features in this way, mark the suspicious or abnormal image area at the corresponding position, for example, further check the texture and color difference of the marked pixel set, if it is confirmed that the difference also appears frequently in the original warning standard, then increase the abnormal weight, if the difference is small, only prompt manual review, after all images are marked, the real-time environmental monitoring results can be summarized.
[0197] The present invention provides a big data analysis system, comprising:
[0198] The data acquisition module receives the image data transmitted by the satellite, uniformly processes these data to obtain a unified format image set;
[0199] The image screening module screens images from the unified format image set, eliminates images with a resolution lower than the preset standard, and obtains an image set;
[0200] The feature extraction module uses deep learning technology to extract environmental visual features from the image set, performs the conversion from images to features to obtain an image feature data set; converts the image feature data set into the format required for feature analysis to obtain environmental feature description data;
[0201] The anomaly detection module, based on the environmental feature description data, identifies features that do not conform to the normal pattern to obtain an anomaly detection result. The anomaly detection result is compared with the anomaly activity database to verify the anomaly type and risk level, adjust and optimize the detection logic, and generate an anomaly activity recognition result;
[0202] The monitoring optimization module adjusts the monitoring parameters according to the anomaly activity recognition result, and uses the new parameters to monitor the satellite images to obtain the environmental real-time monitoring result.
[0203] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A big data analysis method based on a cloud business platform, characterized in that, It includes the following steps: Collect satellite image data, integrate and format the satellite image data to generate a unified format image set; perform quality screening on the unified format image set, remove images with unqualified quality, and obtain a screened image set; Based on the screened image set, use deep learning for feature extraction, extract environmental visual features from the images to generate image feature data, and based on the image feature data, analyze the features to identify abnormal environmental activities and obtain an anomaly detection result; Match the anomaly detection result with a known abnormal activity database to verify the anomaly type and risk level, generate an abnormal activity result, and based on the abnormal activity result, adjust the detection parameters to generate adjusted detection parameters; Use the adjusted detection parameters to continuously monitor the satellite image data and generate a real-time environmental monitoring result.
2. The big data analysis method based on a cloud business platform according to claim 1, wherein The steps for obtaining the unified format image set are as follows: Collect satellite image data, remove data items with inconsistent times, and perform color correction on the remaining data to obtain synchronized satellite image data; Based on the synchronized satellite image data, adjust all images to a unified resolution setting and unify the color mode to obtain standardized satellite image data; According to the standardized satellite image data, convert each image file to a unified file format to obtain a unified format image set.
3. The big data analysis method based on a cloud business platform according to claim 1, characterized in that The steps for obtaining the screened image set are as follows: Extract the exposure level, contrast, and image integrity of each image from the unified format image set to obtain an image quality metadata list; Based on the image quality metadata list, calculate the comprehensive quality score of each image, and the calculation formula is: where Q represents the comprehensive quality score, E represents the exposure level, DC represents the contrast, I represents the image integrity, and XP represents the pixel density; Remove images according to the comprehensive quality score to obtain a screened image set.
4. The big data analysis method based on a cloud business platform according to claim 1, characterized in that The steps for obtaining the image feature data are as follows: Load the screened image set and initialize the deep learning model, including setting the network layers and parameters, to obtain an initialized deep learning model; Extract the environmental visual features of the images one by one through the initialized deep learning model, including texture, shape, and color, to obtain an extracted feature data set; According to the extracted feature data set, apply a pooling layer and an activation function to compress the features to obtain image feature data.
5. The big data analysis method based on a cloud business platform according to claim 1, characterized in that, The steps for obtaining the anomaly detection result are as follows: Load the image feature data, analyze the shape features, color features, and texture features of each image, construct a feature vector matrix, and perform normalization processing on all feature vector matrices to obtain a normalized feature vector matrix; Based on the normalized feature vector matrix, calculate the environmental anomaly score, and the expression is: Among them, D i is the environmental anomaly score of the i-th image, ET i,j represents the shape feature value of the i-th image in the j-th feature dimension, R i,j represents the mean of the shape features in the j-th feature dimension in the normal samples, V i,j represents the shape change range of the i-th image in the j-th feature dimension, U i,j represents the local gradient change of the shape of the i-th image in the j-th feature dimension, Z i,j represents the color feature value of the i-th image in the j-th feature dimension, L i,j represents the mean of the color features in the j-th feature dimension in the normal samples, G i,j represents the color change range of the i-th image in the j-th feature dimension, W i,j represents the local gradient change of the color of the i-th image in the j-th feature dimension, and xm is the total number of features; According to the environmental anomaly score, perform a classification judgment on the images to obtain an anomaly detection result.
6. The big data analysis method based on the cloud business platform according to claim 1, characterized in that The steps for obtaining the abnormal activity result are as follows: Based on the anomaly detection result, query the anomaly detection result and match it with the records in the known abnormal activity database to obtain a matched abnormal activity record; Verify the exception type and risk level of each record according to the matched exception activity record to obtain the exception activity result.
7. The big data analysis method based on a cloud business platform according to claim 1, characterized in that The steps for obtaining the adjusted detection parameters are as follows: Based on the exception activity result, calculate the detection parameter adjustment value, and the expression is: Among them, P i is the adjustment value of the i-th detection parameter, X i,j represents the change amplitude of the i-th abnormal event in the j-th feature dimension, Y i,j represents the occurrence duration of the i-th abnormal event in the j-th feature dimension, T i,j represents the current detection threshold of the i-th abnormal event in the j-th feature dimension, B i,j represents the historical mean threshold of the i-th abnormal event in the j-th feature dimension, C i,j represents the number of frequent changes of the i-th abnormal event in the j-th feature dimension, N i represents the historical cumulative detection times of the i-th abnormal event, and m is the total number of features; Adjust the detection threshold according to the detection parameter adjustment value to generate the adjusted detection parameters.
8. The big data analysis method based on a cloud service platform according to claim 1, wherein The steps for obtaining the real-time monitoring result are as follows: Load the adjusted detection parameters, perform feature analysis on the image, extract the environmental change features in the image, and mark the potential abnormal areas to generate the environmental real-time monitoring result.
9. The big data analysis system for the big data analysis method based on the cloud business platform according to any one of claims 1-8, characterized in that, Include: A data acquisition module that receives the image data transmitted by the satellite, performs unified format processing on these data to obtain a unified format image set; An image screening module that screens images from the unified format image set, eliminates images with a resolution lower than the preset standard to obtain an image set; A feature extraction module that uses deep learning technology to extract environmental visual features from the image set, performs the conversion from image to feature to obtain an image feature data set; converts the image feature data set into the format required for feature analysis to obtain environmental feature description data; An exception detection module that, based on the environmental feature description data, identifies features that do not conform to the normal pattern to obtain an exception detection result. The exception detection result is compared with the exception activity database to verify the exception type and risk level, adjust and optimize the detection logic, and generate an exception activity recognition result; A monitoring optimization module that adjusts the monitoring parameters according to the exception activity recognition result, and uses the new parameters to monitor the satellite image to obtain the environmental real-time monitoring result.
Citation Information
Cited By
Multi-physics coupling slope monitoring method and system based on unmanned aerial vehicle technology
CN120894718A