A safety monitoring system and monitoring method based on machine vision

By combining generalized Fourier-Bessel transform and adaptive nonlinear filtering, deep network and spatiotemporal graph analysis with higher-order polynomial nonlinear superposition, the problems of noise interference, complex fault pattern recognition and multimodal data fusion in the prior art are solved, and more accurate and flexible server hardware security monitoring and early warning response are achieved.

CN119299332BActive Publication Date: 2025-06-17KUNMING INST OF BOTANY CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411404721.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2025-06-17
Estimated Expiration
2044-10-09

AI Technical Summary

Technical Problem

The existing machine vision-based server hardware security monitoring system is difficult to effectively separate and suppress noise interference, resulting in false alarms or missed reports; it is difficult to identify complex hardware failure modes or potential hazardous maintenance operations; it is not possible to fully utilize multiple sensor data for multimodal fusion, resulting in insufficient comprehensive risk assessment; the early warning mechanism lacks flexibility, lag in response speed or insufficient accurate response measures.

Method used

Multi-frequency domain decomposition and noise suppression are used to generate reconstructed images after denoising. Then, the global feature representation is extracted through a deep network of higher-order polynomial nonlinear superposition, a spatiotemporal graph is constructed for behavioral analysis, and a fusion process of fuzzy logic and convolutional integral is combined with multiple sensor data, and risk assessment and early warning strategies are carried out.

Benefits of technology

Effectively separate and suppress noise interference, improve detection accuracy and robustness; be able to accurately detect complex hardware failure modes and potential hazardous operations, enhance risk assessment capabilities for server operation status; through multimodal data fusion, comprehensive risk assessment, improve the flexibility and response speed of the early warning mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119299332B_ABST
    Figure CN119299332B_ABST
Patent Text Reader

Abstract

The present invention relates to the fields of network security monitoring and image processing, and particularly to a security monitoring system and a monitoring method based on machine vision. It includes: collecting visual image data of server hardware in real time, performing multi-frequency domain decomposition on it, filtering the signals obtained in different frequency domains and recombining them to generate a denoised reconstructed image; obtaining a global feature representation based on a deep network, constructing a spatio-temporal graph and performing behavior analysis to obtain a comprehensive spatio-temporal feature representation; after fusing the comprehensive spatio-temporal feature representation with sensor data, performing risk assessment to determine a warning strategy. It solves the problems that noise interference cannot be effectively separated and suppressed, resulting in false alarms or missed alarms in detection results; it is difficult to effectively identify complex abnormal behaviors or dangerous operations; it fails to perform multi-modal fusion with sensor data, resulting in incomplete risk assessment; the warning mechanism lacks flexibility, the response speed lags or the response measures are not precise enough.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of network security monitoring and image processing, and particularly to a security monitoring system and a monitoring method based on machine vision. Background Art

[0002] With the rapid development of data center and cloud computing technologies, the security monitoring of server hardware has become increasingly important. In a data center, the monitoring of server hardware is a crucial link to ensure the stable operation of the system. Previously, the monitoring of server hardware mainly relied on manual inspections and basic hardware monitoring tools, which had problems such as untimely monitoring and high false alarm rates. With the development of technology, some server hardware security monitoring systems based on machine vision have been applied to actual scenarios.

[0003] However, the existing server hardware security monitoring systems based on machine vision still have some deficiencies. First, when processing images in the complex environment inside a server cabinet, these systems often have difficulty effectively separating and suppressing interference from various noise sources such as server device operations and changes in internal cabinet lighting, resulting in the hardware status detection results being easily affected by noise, leading to false alarms or missed alarms. Second, in terms of server hardware anomaly detection and maintenance operation analysis, the existing systems are difficult to effectively identify complex hardware failure modes or potential dangerous maintenance operations, especially in scenarios with multiple servers and high-density deployments, where the performance of the system often fails to meet the actual requirements. Third, many systems do not fully utilize data from different sensors inside the server for multimodal fusion, resulting in an incomplete risk assessment of the server operating state and potentially overlooking some potential hardware failures or security hazards. The early warning mechanisms of existing systems often lack flexibility, resulting in a lag in the response speed or inaccurate response measures to server abnormal states in actual applications, ultimately affecting the security operation effect of the data center. Summary of the Invention

[0004] The present invention provides a security monitoring system and a monitoring method based on machine vision to solve the problems that it is impossible to effectively separate and suppress interference from various noise sources such as device operations and ambient light changes, resulting in the detection results being easily affected by noise, leading to false alarms or missed alarms; it is difficult to effectively identify complex abnormal behaviors or dangerous operations, and not fully utilize data from different sensors for multimodal fusion, resulting in an incomplete risk assessment and potentially overlooking potential risk factors in the environment; the early warning mechanism lacks flexibility, which may lead to a lag in the response speed or inaccurate response measures in actual applications, ultimately affecting security assurance.

[0005] A security monitoring system and a monitoring method based on machine vision of the present invention specifically include the following technical solutions:

[0006] A safety monitoring method based on machine vision, comprising the following steps:

[0007] S1. Real-time collect the visual image data of the server hardware, perform multi-frequency domain decomposition on the visual image data through the generalized Fourier-Bessel transform to obtain signals in different frequency domains; perform adaptive non-linear filtering processing on the signals in different frequency domains to obtain the frequency domain signals after filtering processing; based on the frequency domain signals after filtering processing, generate a reconstructed image after denoising;

[0008] S2. Based on the reconstructed image after denoising, obtain a global feature representation through a deep network with high-order polynomial non-linear superposition; construct a spatio-temporal graph based on the global feature representation and perform behavior analysis to obtain a comprehensive spatio-temporal feature representation; perform fusion processing on the comprehensive spatio-temporal feature representation and sensor data to obtain fused fuzzy decision data; based on the fused fuzzy decision data, perform risk assessment and generate a warning strategy.

[0009] Preferably, the S1 specifically includes:

[0010] Perform adaptive non-linear filtering processing on the signals in different frequency domains by introducing non-linear transformation to obtain the frequency domain signals after filtering processing.

[0011] Preferably, the S1 specifically includes:

[0012] Recombine the frequency domain signals after filtering processing through weighted fusion to obtain a reconstructed image after denoising.

[0013] Preferably, the S2 specifically includes:

[0014] Based on the reconstructed image after denoising, introduce a deep network with high-order polynomial non-linear superposition, extract the feature representation of each layer of the deep network, and perform weighted fusion on the feature representations of all layers of the deep network to obtain a global feature representation.

[0015] Preferably, the S2 specifically includes:

[0016] Based on the global feature representation, construct a spatio-temporal graph; introduce the joint transformation of high-order Bessel functions and Laplace operators to analyze the spatio-temporal graph to obtain a spatio-temporal feature representation; based on the spatio-temporal feature representation, perform a convolution operation to aggregate the spatio-temporal feature information at different times to obtain a comprehensive spatio-temporal feature representation.

[0017] Preferably, the S2 specifically includes:

[0018] Perform fusion processing on the comprehensive spatio-temporal feature representation and sensor data through fuzzy logic and convolution integral to obtain fused fuzzy decision data.

[0019] Preferably, the S2 specifically includes:

[0020] Based on the fused fuzzy decision-making data, risk assessment is carried out to obtain the risk assessment result; based on the risk assessment result, an early warning strategy is generated through non-linear mapping and a fuzzy inference model.

[0021] A safety monitoring system based on machine vision includes the following parts:

[0022] An image acquisition module, a data preprocessing module, an object detection module, a behavior analysis module, an early warning response module, and a database;

[0023] The image acquisition module captures the visual image data of the server hardware in real time and transmits the captured visual image data to the data preprocessing module;

[0024] The data preprocessing module performs multi-frequency domain decomposition on the visual image data through the generalized Fourier-Bessel transform to obtain signals in different frequency domains; performs adaptive non-linear filtering processing on the signals in different frequency domains to obtain the filtered frequency domain signals; based on the filtered frequency domain signals, reconstructs the image through weighted fusion to generate a denoised reconstructed image; transmits the denoised reconstructed image to the object detection module and the database;

[0025] The object detection module, based on a deep network with high-order polynomial non-linear superposition, extracts the feature representation of each layer of the deep network from the denoised reconstructed image, and performs weighted fusion on the feature representations of all layers of the deep network to obtain a global feature representation; transmits the global feature representation to the behavior analysis module and the database;

[0026] The behavior analysis module, based on the global feature representation, constructs a spatio-temporal graph using a high-order Bessel transform network, extracts spatio-temporal feature representations, aggregates spatio-temporal features through convolution operations to form a comprehensive spatio-temporal feature representation; transmits the comprehensive spatio-temporal feature representation to the early warning response module and the database;

[0027] The early warning response module fuses the comprehensive spatio-temporal feature representation with sensor data to obtain fused fuzzy decision-making data; based on the fused fuzzy decision-making data, conducts risk assessment to obtain the risk assessment result; based on the risk assessment result, generates an early warning strategy through non-linear mapping and a fuzzy inference model;

[0028] The database is used to store the data transmitted by the data preprocessing module, the object detection module, and the behavior analysis module.

[0029] The beneficial effects of the technical solution of the present invention are:

[0030] 1. By combining the generalized Fourier-Bessel transform and adaptive nonlinear filtering, environmental noise is effectively separated and suppressed while key visual image features are retained. The adaptive nonlinear filtering process not only considers the characteristics of local noise but also enhances the contrast of visual image signals through nonlinear transformation, thereby improving the accuracy and robustness of visual image processing and ensuring the reliability of subsequent target detection.

[0031] 2. A deep network using high-order polynomial nonlinear superposition is used to perform multi-level feature extraction and fusion on the denoised reconstructed image, effectively capturing complex features, especially high-order features, in the denoised reconstructed image. The superposition of multi-level features and the generation of global features enable the machine vision system to accurately detect potential dangerous targets or abnormal device states.

[0032] 3. By constructing a spatio-temporal graph and combining the joint transform of high-order Bessel functions and Laplace operators, the machine vision system can accurately capture the behavior patterns and interrelationships of target objects in the spatio-temporal dimension. The aggregation of spatio-temporal features further enhances the machine vision system's ability to analyze target behavior and identify abnormal operations, and can effectively detect potential safety risks.

[0033] 4. This invention not only relies on the data of the machine vision system but also combines data from multiple sensors (such as temperature, pressure, vibration, etc.). These multi-modal data are fused through fuzzy logic and convolution integration to form unified risk assessment data, improving the machine vision system's comprehensive judgment ability of the environmental state and making the risk assessment more comprehensive and accurate.

[0034] 5. Based on the fused fuzzy decision-making data, the machine vision system realizes the dynamic assessment of risks through second-order derivative analysis and triggers different levels of early warning response mechanisms according to the risk assessment results. The machine vision system sets multiple early warning thresholds and can flexibly control different early warning levels, from low-level warnings (such as sound alarms) to high-level warnings (such as automatic shutdown), ensuring that the machine vision system can make timely and accurate responses in various risk situations. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a structural diagram of a safety monitoring system based on machine vision according to the present invention;

[0036] Figure 2 It is a flowchart of a safety monitoring method based on machine vision according to the present invention;

[0037] Figure 3 It is a topological diagram of the server vision monitoring scenario according to the present invention;

[0038] Figure 4It is a design diagram of a safety monitoring system architecture based on machine vision according to the present invention. Detailed implementation manners

[0039] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0041] The following specifically describes the specific solutions of a safety monitoring system and a monitoring method based on machine vision provided by the present invention in conjunction with the accompanying drawings.

[0042] Referring to the attached Figure 1 , which shows a structural diagram of a safety monitoring system based on machine vision provided by an embodiment of the present invention. The system includes the following parts:

[0043] An image acquisition module, a data preprocessing module, a target detection module, a behavior analysis module, an early warning response module, and a database;

[0044] The image acquisition module captures visual image data of the server hardware in real time through a camera and transmits the captured image data to the data preprocessing module;

[0045] The data preprocessing module performs multi-frequency domain decomposition on the visual image data by using the generalized Fourier-Bessel transform to obtain signals in different frequency domains; suppresses noise and enhances features of the signals in different frequency domains through an adaptive nonlinear filter, and finally reconstructs the image data through weighted fusion to obtain a denoised reconstructed image; transmits the denoised reconstructed image to the target detection module and the database;

[0046] The target detection module extracts multi-layer features from the denoised reconstructed image based on a deep network with high-order polynomial nonlinear superposition, generates a global feature representation for identifying potential dangerous targets or abnormal device states; transmits the global feature representation to the behavior analysis module and the database;

[0047] The behavior analysis module, based on the global feature representation, constructs a spatio-temporal graph using a high-order Bessel transform network, extracts the spatio-temporal feature representation, aggregates the spatio-temporal features through convolution operations to form a comprehensive spatio-temporal feature representation, which is used to analyze the behavior patterns of the target and identify dangerous behaviors or abnormal operations; transmits the comprehensive spatio-temporal feature representation to the early warning response module and the database;

[0048] The early warning response module fuses and processes the comprehensive spatio-temporal feature representation with other sensor data to obtain the fused fuzzy decision-making data; based on the fused fuzzy decision-making data, conducts risk assessment to obtain the risk assessment result; based on the risk assessment result, generates an early warning signal through a non-linear mapping and a fuzzy inference model to trigger corresponding safety response measures. The machine vision system sets multiple early warning thresholds, and each early warning threshold corresponds to a different level of early warning or response mechanism.

[0049] The database is used to store the data transmitted by the data preprocessing module, the target detection module, and the behavior analysis module.

[0050] Refer to the appendix Figure 2 which shows a flowchart of a machine vision-based safety monitoring method provided by an embodiment of the present invention. The method includes the following steps:

[0051] S1. Real-time collect the visual image data of the server hardware, perform multi-frequency domain decomposition on the visual image data through the generalized Fourier-Bessel transform to obtain signals in different frequency domains; perform adaptive non-linear filtering processing on the signals in different frequency domains to obtain the filtered frequency domain signals; based on the filtered frequency domain signals, generate a denoised reconstructed image;

[0052] Refer to the appendix Figure 3 and the appendix Figure 4 The client captures the visual image data of the server hardware through the image acquisition module and transmits it to the server side. The server side includes a data preprocessing module, a target detection module, a behavior analysis module, and an early warning response module. The captured visual image data is stored in a standard format and transmitted to the data preprocessing module in real time through a high-speed data transmission interface (such as GigE or USB3.0) and stored and managed in the server cabinet to ensure the integrity and availability of the data at any time.

[0053] The collected visual image data contains rich environmental information, but at the same time contains a large amount of complex noise. The noise may come from equipment operation, environmental light changes, and other uncertain factors. Therefore, preprocessing is required to improve the accuracy of subsequent analysis. The data preprocessing module in the server-side elastic computing device uses the generalized Fourier-Bessel transform to perform multi-frequency domain decomposition on the visual image data. The generalized Fourier-Bessel transform formula is:

[0054]

[0055] Among them, F i (x, y) represents the value of the visual image signal in the i-th frequency domain, which is the frequency domain component after the generalized Fourier-Bessel transform of the input visual image data, that is, the signal in the i-th frequency domain; I(x, y) is the gray value of the input visual image data at the spatial coordinates (x, y); m, n are the spatial position offsets of the visual image data; ψi(m, n) represents the basis function of the i-th frequency domain; is a complex exponential term, representing the contribution of different frequency components in the Fourier transform, ω m,n is the angular frequency, and j is the ordinal unit; J v (ar) is the Bessel function, representing the amplitude distribution in the frequency domain, v is the order of the Bessel function, α is the parameter for adjusting the frequency amplitude, and r is the spatial distance

[0056] After the multi-frequency domain decomposition of the visual image data is completed, the data preprocessing module also performs adaptive non-linear filtering on the signals F i (x, y) in different frequency domains to suppress noise and enhance key features. The adaptive non-linear filter not only considers the characteristics of local noise but also introduces non-linear transformation to further optimize the processing effect. The specific formula for adaptive non-linear filtering is:

[0057]

[0058] Among them, F′ i (x, y) represents the visual image signal in the i-th frequency domain after adaptive non-linear filtering processing, that is, the frequency domain signal after filtering processing; represents the estimated value of the noise of the visual image signal at the position (x, y) in the i-th frequency domain, which is obtained by statistically analyzing the gray value change of the local area of the visual image signal; λ i is the adjustment parameter of the adaptive non-linear filter, which is used to control the filtering intensity to ensure that useful signals are retained while suppressing noise; the non-linear function tanh(γ·F i (x, y)) further enhances the contrast of the visual image signal in the i-th frequency domain, where γ is the non-linear coefficient, which is used to adjust the enhancement strength of the visual image signal in the i-th frequency domain; ω is the frequency component in the frequency domain, which is used to describe the distribution of the visual image signal in the i-th frequency domain in the frequency domain.

[0059] Recombine the frequency domain signals after filtering processing, and form a denoised reconstructed image by weighted fusion of the results of different frequency domains. Its formula is:

[0060]

[0061] where \(I'(x, y)\) is the value of the denoised reconstructed image at the spatial coordinates \((x, y)\); \(\alpha\) i is the weighting coefficient of the \(i\)-th frequency domain, representing the importance of the \(i\)-th frequency domain in the final image reconstruction; represents the total number of frequency domains; \(\cos(\beta\omega)\) is the cosine function used for frequency domain fusion, \(\beta\) is the parameter for adjusting the frequency of the cosine function, controlling the role of the frequency domain components in the reconstructed image, and \(\omega\) is the frequency component in the frequency domain. The above-obtained denoised reconstructed image \(I'(x, y)\) has effectively suppressed the noise while retaining important image features.

[0062] S2. Based on the denoised reconstructed image, obtain the global feature representation through a deep network with high-order polynomial non-linear superposition; construct a spatio-temporal graph based on the global feature representation and perform behavior analysis to obtain the comprehensive spatio-temporal feature representation; fuse the comprehensive spatio-temporal feature representation with the sensor data to obtain the fused fuzzy decision-making data; based on the fused fuzzy decision-making data, conduct risk assessment and generate a warning strategy.

[0063] The target detection module takes the denoised reconstructed image as input, uses a deep network with high-order polynomial non-linear superposition for target detection, and accurately identifies potential dangerous targets or abnormal device states from the denoised reconstructed image.

[0064] Specifically, extract multi-layer feature representations through a deep network with high-order polynomial non-linear superposition, and its formula is as follows:

[0065]

[0066] where \(G\) l is the feature representation of the \(l\)-th layer of the deep network; \(Y\) l represents the input feature of the \(l\)-th layer of the deep network; \(Y_1 = I'(x, y)\) represents the input denoised reconstructed image, which is the input feature of the first layer of the deep network; after multi-layer non-linear transformation the generated feature representation will be used as the input feature of the \((l + 1)\)-th layer of the deep network; is the \(q\)-th power of the input feature of the \(l\)-th layer of the deep network, \(V\) l is the weight matrix of the \(l\)-th layer of the deep network, \(b\) l is the bias term of the \(l\)-th layer of the deep network, \(q\) represents the order of the polynomial, and by calculating the \(q\)-th derivative of a specific dimension in the input feature of the \(l\)-th layer of the deep network it is possible to further capture the high-order features in the denoised reconstructed image, \(z\) is a specific dimension (such as spatial coordinates) in the input feature of the \(l\)-th layer of the deep network; \(p\) is the highest order of the polynomial, that is, the depth of the extracted features, which determines the non-linear expression ability of the deep network; is the superposition coefficient that controls the influence of the feature representation of the previous layer of the deep network on the feature representation of the current layer of the deep network. The feature representation G of each layer of the deep network l will be gradually superimposed in multiple levels to form a more expressive global feature representation.

[0067] The generation of the global feature representation is achieved through the following formula:

[0068]

[0069] where G global is the global feature representation; δ l is the fusion weight of the feature representation of each layer of the deep network; L represents the total number of layers of the deep network with high-order polynomial non-linear superposition. By performing weighted fusion on the feature representations of all layers of the deep network and integrating in the time dimension r, a global feature representation that synthesizes the information of each layer of the deep network is obtained.

[0070] The behavior analysis module is based on the global feature representation G global to construct the spatio-temporal graph G. The purpose of constructing the spatio-temporal graph is to capture the mutual relationships and behavior patterns of the target object in time and space through the structured representation of the graph, and identify dangerous behaviors or abnormal operations. The specific implementation steps are as follows:

[0071] Each detection target in the global feature representation G global , that is, the feature representation of each layer of the deep network, serves as a node of the spatio-temporal graph. The feature vector of the node comes from a specific subset in the global feature representation and is used to represent the spatial position, speed, shape features, etc. of the detection target; the edges between the nodes represent the spatio-temporal relationships between different detection targets, and the features of the edges are calculated from the feature vectors of the two nodes. Common calculation methods include Euclidean distance, cosine similarity, etc. The weight of the edge depends on the physical distance and relative speed between the detection targets, that is, with the natural logarithm as the base, and the exponent is the Euclidean distance between the two nodes divided by the adjustment parameter used to control the decay speed of the edge weight; the spatio-temporal graphs at different times are connected through the time dimension to form a time series graph. The nodes at each moment are connected to their nodes at the next moment through edges to ensure spatio-temporal continuity.

[0072] After constructing the spatio-temporal graph G, in order to better capture the spatio-temporal features, the joint transformation of the high-order Bessel function and the Laplace operator is introduced, and the formula is as follows:

[0073]

[0074] H0 = σ(W0·G global + b0)

[0075] where Ht+1 is the spatio-temporal feature representation at the next moment, which combines the second-order derivative features of the spatio-temporal graph and the spatio-temporal feature representation H at the current moment t ; H0 is the initial spatio-temporal feature representation; is the Laplace operator, which is used to extract the local second-order derivative features of the spatio-temporal graph and enhance the representation of spatio-temporal dependence; J k is the k-th order Bessel function; represents the highest order of the Bessel function; is the weight corresponding to the k-th order Bessel function, which controls the contribution of different-order spatio-temporal features in the spatio-temporal graph; λ and are the adjustment coefficient and the frequency parameter respectively, which are used to adjust the amplitude and frequency of the change of spatio-temporal features in the time dimension; cosh is the hyperbolic cosine function, which further enhances the expression ability of spatio-temporal features; W0 is the weight matrix, which is used to perform a linear transformation on the global feature representation; b0 is the bias vector, which is added to the result of the linear transformation to adjust the center position of the spatio-temporal feature representation and avoid over-reliance on zero values.

[0076] After the extracted spatio-temporal feature representations are aggregated through convolution operations, a comprehensive spatio-temporal feature representation is formed, and its formula is:

[0077]

[0078] where, H final is the comprehensive spatio-temporal feature representation; is the weighted coefficient in the time dimension; W c is the convolution kernel, which aggregates the spatio-temporal feature information at different moments through convolution operations on the spatio-temporal feature representation, and finally forms a comprehensive spatio-temporal feature representation; T is the length of the time series, that is, the total number of time steps.

[0079] The comprehensive spatio-temporal feature representation serves as a comprehensive video stream that can reflect the current risk state. The server processes the processed video stream through the front end for display. Operators can view real-time data, adjust the parameters of the security monitoring system, and view historical alarm records through the configuration page.

[0080] The early warning response module takes the comprehensive spatio-temporal feature representation generated by the behavior analysis module as input, and performs fusion processing with the data collected by other sensors (such as temperature sensors, pressure sensors, vibration sensors, humidity sensors, etc.). The data of other sensors are used to supplement the image data of the machine vision system, providing more-dimensional data to help the security monitoring system more accurately evaluate the environmental state and potential security risks. The fusion processing is achieved through fuzzy logic and convolution integral, and its formula is:

[0081]

[0082] where, Df is the fused fuzzy decision-making data, serving as the basic data for risk assessment; u is the index of other sensor data; N is the total number of other sensor data; is the membership function of the fuzzy set, indicating that according to the calculated membership degree, is the eigenvalue extracted from S u where S u is the u-th sensor data matrix, obtained by fusing through convolution operation with comprehensive spatio-temporal features; dω is the integral of the frequency variable, representing the comprehensive processing of sensor data in the frequency domain, and the integral operation ensures that all frequency components in the sensor data are effectively processed.

[0083] Risk assessment is performed on the fused fuzzy decision-making data, and its formula is as follows:

[0084]

[0085] where R(t) is the risk level at the current time t, representing the risk value evaluated according to the input data at the current time, that is, the risk assessment result; M is the total number of risk assessment factors; c is the index of the risk assessment factor; represents the second derivative of the fused fuzzy decision-making data in the time dimension, reflecting the dynamic change of the fused fuzzy decision-making data over time; is the weight parameter in fuzzy inference; R(t - 1) is the risk level at the previous time.

[0086] According to the risk assessment result, the warning strategy is determined through a non-linear mapping and a fuzzy inference model, and the formula is as follows:

[0087]

[0088] where D output (t) is the decision output signal at time t; θ is the decision weight, indicating the influence of the risk assessment result in the warning decision; τ is the normalization coefficient, used to adjust the scale of the risk level R(t) to ensure the smoothness of the decision output signal. The finally output D output (t) is the warning signal of the alarm system, used to trigger corresponding safety measures, such as alarms, notifications, or automatic shutdowns.

[0089] The alarm system sets multiple warning thresholds, and each warning threshold corresponds to a different level of warning or response mechanism. When D 。utputWhen (t) reaches or exceeds a certain warning threshold, the alarm system will trigger the corresponding warning mechanism. According to the triggered warning mechanism, the alarm system decides which warning measures to initiate. For example, the warning levels are divided into three levels: low, medium, and high, corresponding to different risk levels. Low-level warning: may trigger an audible alarm or display a warning message on the monitoring screen to prompt the operator's attention; Medium-level warning: the alarm system may send text messages or emails to notify the security management personnel to further investigate potential risks; High-level warning: the alarm system may immediately execute emergency measures such as automatic shutdown and safety isolation to prevent possible major accidents.

[0090] Thus, flexible control of different warning levels is achieved, ensuring that the machine vision system can make timely and accurate responses in various risk situations, thereby guaranteeing the security of the network environment.

[0091] In summary, a security monitoring system and monitoring method based on machine vision are completed.

[0092] The sequence of the invention embodiments is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0093] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.

[0094] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. A safety monitoring method based on machine vision, characterized in that: The following steps are involved: S1, real-time collection of visual image data from server hardware, multi-frequency domain decomposition of the visual image data through generalized Fourier-Bessel transform, to obtain signals in different frequency domains; Performing adaptive nonlinear filtering on signals in different frequency domains to obtain filtered frequency domain signals; generating a denoised reconstructed image based on the filtered frequency domain signals; S2. Based on the denoised reconstructed image, a deep network with high-order polynomial nonlinear superposition is used to extract the feature representation of each layer of the deep network, and the feature representation of all layers of the deep network is weightedly fused to obtain the global feature representation; Based on the global feature representation, a spatiotemporal graph is constructed and analyzed to obtain the spatiotemporal feature representation. Based on the spatiotemporal feature representation, a convolution operation is performed to aggregate the spatiotemporal feature information at different times to obtain a comprehensive spatiotemporal feature representation. The comprehensive spatiotemporal feature representation is fused with the sensor data to obtain fused fuzzy decision data. Based on the fused fuzzy decision data, risk assessment is performed to generate an early warning strategy.

2. A safety monitoring method based on machine vision according to claim 1, characterized in that: In S1, performing adaptive nonlinear filtering on signals in different frequency domains to obtain filtered frequency domain signals specifically includes: By introducing nonlinear transformation, adaptive nonlinear filtering is performed on signals in different frequency domains to obtain frequency domain signals after filtering.

3. A safety monitoring method based on machine vision according to claim 2, characterized in that: In S1, generating a denoised reconstructed image based on the filtered frequency domain signal specifically includes: The filtered frequency domain signals are recombined through weighted fusion to obtain the denoised reconstructed image.

4. The method for safety monitoring based on machine vision according to claim 1, characterized in that: In S2, the spatiotemporal graph is analyzed to obtain a spatiotemporal feature representation, which specifically includes: The high-order Bessel function and Laplace operator joint transformation are introduced to analyze the space-time graph and obtain the space-time feature representation.

5. The method for safety monitoring based on machine vision according to claim 4, characterized in that: In S2, the comprehensive spatiotemporal feature representation is fused with the sensor data to obtain fused fuzzy decision data, which specifically includes: Through fuzzy logic and convolution integral, the comprehensive spatiotemporal feature representation and sensor data are fused to obtain the fused fuzzy decision data.

6. A safety monitoring method based on machine vision according to claim 5, characterized in that: In S2, based on the fused fuzzy decision data, risk assessment is performed and an early warning strategy is generated, which specifically includes: Based on the fused fuzzy decision data, risk assessment is performed to obtain risk assessment results; based on the risk assessment results, an early warning strategy is generated through nonlinear mapping and fuzzy reasoning models.

7. A machine vision-based safety monitoring system, applied to the machine vision-based safety monitoring method as claimed in claim 1, characterized in that: Includes the following parts: Image acquisition module, data preprocessing module, target detection module, behavior analysis module, early warning response module and database; An image acquisition module captures the visual image data of the server hardware in real time and transmits the collected visual image data to the data preprocessing module; The data preprocessing module performs multi-frequency domain decomposition on the visual image data through generalized Fourier-Bessel transform to obtain signals in different frequency domains; Adaptively perform nonlinear filtering on signals in different frequency domains to obtain frequency domain signals after filtering; reconstruct images through weighted fusion based on the frequency domain signals after filtering to generate denoised reconstructed images; transmit the denoised reconstructed images to the target detection module and the database; The target detection module is based on a deep network with high-order polynomial nonlinear superposition. It extracts the feature representation of each layer of the deep network from the denoised reconstructed image, and performs weighted fusion on the feature representation of all layers of the deep network to obtain a global feature representation. The global feature representation is transmitted to the behavior analysis module and the database. The behavior analysis module uses a high-order Bessel transform network to construct a spatiotemporal graph based on the global feature representation, extracts the spatiotemporal feature representation, aggregates the spatiotemporal features through convolution operations, and forms a comprehensive spatiotemporal feature representation; the comprehensive spatiotemporal feature representation is transmitted to the early warning response module and the database; A database, used to store data transmitted by the data preprocessing module, the target detection module and the behavior analysis module; The early warning response module fuses the comprehensive spatiotemporal feature representation with the sensor data to obtain the fused fuzzy decision data; Based on the fused fuzzy decision data, risk assessment is carried out to obtain risk assessment results; based on the risk assessment results, an early warning strategy is generated through nonlinear mapping and fuzzy reasoning models.

Citation Information

Patent Citations

  • Multi-person abnormal behavior detection and recognition method based on machine vision

    CN109522793A

  • Automobile instrument automatic detection device and method based on machine vision

    CN110044405A