An Automatic Detection Method for Building Equipment Status Based on Image Recognition

By constructing an automatic detection method for the condition of building equipment based on image recognition, and utilizing continuous video acquisition and an improved ARDiff judgment model, the problem of difficulty in identifying minute structural changes in existing technologies is solved, and accurate detection and efficient condition judgment of early anomalies of building equipment are achieved.

CN122493395APending Publication Date: 2026-07-31XIONGAN WANWEI OPTOELECTRONICS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIONGAN WANWEI OPTOELECTRONICS TECHNOLOGY CO LTD
Filing Date
2026-05-13
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing methods for monitoring the condition of building equipment are difficult to accurately extract information on minute structural changes, especially in the early stages of anomalies, and are easily affected by changes in lighting and camera shake.

Method used

By constructing an automatic detection method for the status of building equipment based on image recognition, and utilizing continuous video acquisition, visual reference system construction, subpixel-level motion estimation, frequency domain decomposition, and an improved ARDiff decision model, micro-vibration, micro-displacement, and micro-deformation features are extracted to generate composite feature image sequences, and then region division and status determination are performed.

Benefits of technology

It enables accurate identification of early anomalies in building equipment, improves the automation level and resistance to light interference in detection, enhances the accuracy of condition determination, and reduces reliance on manual inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493395A_ABST
    Figure CN122493395A_ABST
Patent Text Reader

Abstract

This invention discloses an automatic detection method for the status of building equipment based on image recognition, comprising the following steps: obtaining an image frame sequence; performing visual reference system construction processing to obtain a standardized image sequence; performing sub-pixel level motion estimation to construct a microstructure change signal sequence; performing frequency domain decomposition processing to generate a microstructure change enhancement sequence; constructing a change feature map and fusing the change feature map with the corresponding image in the standardized image sequence to form a composite feature image sequence; performing region division to construct a feature vector of equipment operating status; performing analysis and processing using an improved ARDiff judgment model to output the current operating status category of the equipment; generating corresponding status identification information and automatically outputting an alarm signal when an abnormal status is detected, thereby improving the ability to identify micro-vibrations, micro-displacements, and micro-deformations of building equipment and realizing automatic detection of the operating status of building equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent monitoring of building equipment, and in particular to an automatic detection method for the status of building equipment based on image recognition. Background Technology

[0002] As buildings continue to expand in scale, the number of elevators, ventilation equipment, water supply and drainage equipment, air conditioning units, pumps, and various electromechanical equipment in building operation and management is constantly increasing. The operating status of building equipment directly affects the safety, stability, and maintenance costs of the building. Existing methods for monitoring the condition of building equipment mainly include manual inspection, sensor detection, and appearance inspection based on ordinary image recognition. Manual inspection usually relies on maintenance personnel to periodically observe the appearance, operating sounds, vibration, or indication status of equipment on-site. This method suffers from long detection cycles, strong subjectivity, and difficulty in timely detection of early anomalies. Although sensor detection can obtain operating parameters such as vibration, temperature, and current, it requires the installation of additional sensors on the equipment, resulting in high deployment costs and significant difficulties in modifying existing building equipment. Ordinary image recognition methods typically acquire images of the equipment through camera devices and then use target detection, image classification, or defect recognition models to determine whether there are obvious anomalies. However, these methods mostly rely on visible appearance features such as cracks, damage, leaks, corrosion, and foreign object obstruction, and are insufficient in recognizing subtle changes such as micro-vibrations, micro-displacements, and micro-deformations that occur in the early stages of equipment operation.

[0003] Existing image recognition-based methods for building equipment condition detection typically extract features and classify conditions directly from acquired images, lacking a dedicated processing mechanism for microstructural changes in continuous video data. When building equipment is in an early stage of anomaly, its appearance may not yet show obvious defects, with only minute positional changes occurring at structural connection points, fixed installation points, edge intersections, or the outer contour of the equipment. These changes are difficult to directly reflect in a single frame image and are easily masked by noise, lighting variations, and camera shake during ordinary inter-frame differencing or conventional image classification. Therefore, existing technologies struggle to accurately extract minute structural changes during the operation of building equipment from image data, resulting in insufficient response to early anomalies in equipment condition detection. Summary of the Invention

[0004] One objective of this invention is to propose an automatic detection method for the status of building equipment based on image recognition. This invention fully utilizes continuous video acquisition, visual reference system construction, sub-pixel level motion estimation, frequency domain decomposition, microstructure change amplification, and an improved ARDiff judgment model. It describes in detail the detection method for extracting micro-vibration, micro-displacement, and micro-deformation features from building equipment image frame sequences and automatically judging the operating status of the equipment. It has the advantages of strong early anomaly recognition capability, strong resistance to light interference, high degree of automation in detection, high accuracy in status judgment, and reduced reliance on manual inspection.

[0005] An automatic detection method for the status of building equipment based on image recognition according to an embodiment of the present invention includes the following steps:

[0006] The system continuously collects data on the operation of building equipment, automatically acquires continuous video data of the equipment's operation, and extracts image frame sequences.

[0007] Visual reference system construction processing is performed on the image frame sequence to extract stable feature points in the building equipment structure, and geometric alignment and illumination normalization processing are performed on each frame in the image frame sequence to obtain a standardized image sequence.

[0008] Subpixel-level motion estimation is performed on standardized image sequences to extract displacement information of each pixel over time, and a microstructure change signal sequence in the time dimension is constructed.

[0009] Frequency domain decomposition is performed on the microstructure change signal sequence to filter out the target change component, and amplitude amplification is performed on the target change component to generate a microstructure change enhanced sequence.

[0010] A change feature map is constructed based on the microstructure change enhancement sequence, and the change feature map is fused with the corresponding image in the normalized image sequence to form a composite feature image sequence;

[0011] The composite feature image sequence is divided into regions, and the change amplitude features, change direction features, and temporal continuity features of each region are extracted to construct the equipment operating status feature vector;

[0012] The feature vector of the equipment's operating status is input into the improved ARDiff decision model for analysis and processing, and the current operating status category of the equipment is output.

[0013] Based on the operating status category, corresponding status identification information is generated, and an alarm signal is automatically output when an abnormal status is detected.

[0014] Optionally, an image acquisition device is fixedly installed in the target monitoring area of ​​the building equipment. The image acquisition device continuously acquires video of the building equipment operation process at a preset acquisition frame rate. The acquired continuous video data is analyzed frame by frame, and each frame of image data is extracted from the continuous video data and arranged in the order of acquisition time to generate an image frame sequence.

[0015] Optionally, obtaining the standardized image sequence specifically includes:

[0016] Extract stable feature points from the building equipment structure in each frame of the image frame sequence, and construct the feature point set corresponding to each frame image;

[0017] The first frame in the image frame sequence is determined as the reference image, and the set of feature points corresponding to the reference image is used as the reference feature point set.

[0018] A spatial reference coordinate system is established based on the set of reference feature points to form a unified visual reference benchmark.

[0019] The feature point set corresponding to each frame in the image frame sequence is matched with the reference feature point set to establish a one-to-one correspondence between the feature points and obtain a set of matched feature point pairs.

[0020] The geometric transformation matrix of each frame in the image frame sequence relative to the reference image is calculated based on the set of matching feature point pairs;

[0021] Based on the geometric transformation matrix, coordinate mapping processing is performed on each frame of the image frame sequence, and combined with a unified visual reference datum, a geometrically aligned image sequence is obtained;

[0022] Illumination statistics are performed on each frame of the geometrically aligned image sequence to calculate the average pixel grayscale value and the pixel grayscale dispersion of each frame. Illumination normalization is then performed on each frame of the geometrically aligned image sequence to obtain a standardized image sequence.

[0023] Optionally, the construction of the microstructure change signal sequence specifically includes:

[0024] Read the normalized image sequence and determine the current normalized image and the next normalized image according to the frame number order;

[0025] In the current standardized image, determine the pixel to be estimated, and in the next standardized image, determine the local search region corresponding to the pixel to be estimated;

[0026] The local image block containing the pixel to be estimated is matched with the candidate image block in the local search area, and the initial displacement range of the pixel to be estimated between two adjacent normalized images is determined from the matching results.

[0027] Perform subpixel-level motion estimation within the initial displacement range to obtain subpixel-level displacement information of the pixel to be estimated between two adjacent normalized frames;

[0028] The sub-pixel displacement information of the same pixel to be estimated between two adjacent standardized images is summarized according to the frame number sequence to construct the time displacement signal corresponding to the pixel to be estimated. The time displacement signals corresponding to each pixel to be estimated are summarized according to the pixel spatial position to construct a microstructure change signal sequence in the time dimension.

[0029] Optionally, the generation of the microstructure change enhancement sequence specifically includes:

[0030] Frequency domain decomposition is performed on the microstructure change signal sequence to convert the displacement change signal in the time domain into displacement frequency components in the frequency domain, thus obtaining transverse frequency domain components and longitudinal frequency domain components.

[0031] Based on the frequency range of structural response corresponding to micro-vibration, micro-displacement, or micro-deformation during the operation of building equipment, target change components are selected from the transverse and longitudinal frequency domain components.

[0032] The amplitude of the target change component is amplified by multiplying the amplitude of the target change component by an amplification factor to obtain the amplified target change component.

[0033] The magnified target change components are recombined according to their corresponding pixel spatial locations and frame numbers to generate a microstructure change enhancement sequence.

[0034] Optionally, the formation of the composite feature image sequence specifically includes:

[0035] Read the microstructure change enhancement sequence, and extract the lateral and longitudinal enhancement displacement information corresponding to each pixel to be estimated according to the pixel spatial location and frame number;

[0036] Based on the horizontal and vertical enhanced displacement information, the change intensity value of each pixel to be estimated is calculated. The change intensity value of each pixel to be estimated is written into the image coordinate plane according to the corresponding pixel spatial position to construct a change feature mapping map.

[0037] Read the corresponding image from the standardized image sequence according to the frame number, and align the change feature map with the corresponding image from the standardized image sequence in terms of spatial position;

[0038] Channel fusion processing is performed on the spatially aligned change feature map and the corresponding image in the normalized image sequence. The corresponding image in the normalized image sequence is used as the structural appearance channel, and the change feature map is used as the microstructure change channel to form a composite feature image. The composite feature images are arranged in order of frame number to form a composite feature image sequence.

[0039] Optionally, the construction of the device operating state feature vector specifically includes:

[0040] Read the composite feature image sequence, obtain each frame of composite feature image in order of frame number, and divide each frame of composite feature image into multiple detection regions;

[0041] Read the structural appearance channel data and microstructure change channel data in each detection area, and extract the change intensity value corresponding to each pixel in the detection area;

[0042] The variation amplitude characteristics of each detection area are calculated based on the change intensity value;

[0043] Based on the lateral and longitudinal enhanced displacement information of each pixel in the microstructure change channel data, the change direction features of each detection area are extracted.

[0044] The temporal continuity features of the detection region are extracted by statistically analyzing the change amplitude and change direction features of the same detection region in the composite feature image of consecutive frames according to the frame number sequence.

[0045] The variation amplitude features, variation direction features, and time continuity features corresponding to each detection area are combined according to the detection area number to construct the equipment operation status feature vector.

[0046] Optionally, the output of the current operating status category of the device specifically includes:

[0047] The equipment operating status feature vector is input into the improved ARDiff decision model for analysis and processing. The improved ARDiff decision model includes a microstructure enhancement feature spectroscopy module, a regional correlation diffusion modeling module, a time inversion difference determination module, and a state classification output module.

[0048] In the microstructure enhancement feature spectralization module, the change amplitude features, change direction features, and time continuity features in the device operation state feature vector are subjected to phase-consistent spectralization to generate the microstructure enhancement feature spectrum;

[0049] In the regional correlation diffusion modeling module, a regional correlation diffusion map is constructed based on the spatial adjacency relationship, structural connection relationship and microstructure change transmission direction between each detection area in the building equipment structure. The microstructure enhancement feature spectrum is then input into the regional correlation diffusion map for structural propagation processing to obtain the regional correlation diffusion features.

[0050] Anomaly propagation suppression processing is performed on the regional association diffusion characteristics. The anomaly propagation suppression processing includes identifying the diffusion characteristics of isolated mutations within a single detection region, and performing consistency verification on the diffusion characteristics of isolated mutations based on the regional association diffusion characteristics of adjacent detection regions to obtain the verified regional association diffusion characteristics.

[0051] In the time inversion difference determination module, time inversion denoising and reconstruction processing is performed on the verified regional correlation diffusion features. The normal state prediction features of building equipment under normal conditions are reconstructed in reverse order of frame number. The difference between the normal state prediction features and the verified regional correlation diffusion features is compared to obtain the state deviation features.

[0052] In the state classification output module, the state deviation features are subjected to evidence reliability gating processing. Based on the state deviation features after evidence reliability gating processing, a state deviation degree index is generated, and the current operating state category of the equipment is output according to the state deviation degree index.

[0053] Optionally, corresponding status identifier information is generated based on the current operating status category of the device. When the operating status category is normal, a normal status identifier is generated, and the normal status identifier, corresponding frame number, detection time period, and device number are stored. When the operating status category is warning, a warning status identifier is generated, and the warning status identifier, corresponding frame number, detection time period, device number, and status deviation index are stored. When the operating status category is abnormal, an abnormal status identifier is generated, and an alarm signal is automatically output.

[0054] The beneficial effects of this invention are:

[0055] This invention constructs a unified visual reference system based on stable feature points and performs geometric alignment and illumination normalization on image frame sequences. This ensures that images acquired at different times maintain consistency in spatial location and illumination conditions, effectively eliminating interference caused by camera shake, changes in shooting angle, and fluctuations in ambient lighting. Compared to existing methods that directly analyze raw images, this invention significantly improves the consistency and comparability of image sequences during the preprocessing stage, providing a stable data foundation for subsequent microstructure change extraction and thus enhancing the reliability of the overall detection process.

[0056] This invention performs sub-pixel-level motion estimation on standardized image sequences, refining the extraction of minute displacement changes of pixels between consecutive frames and constructing a microstructural change signal sequence in the time dimension. This allows for the quantification of micro-vibrations, micro-displacements, and micro-deformations that are difficult to observe in a single frame. Furthermore, by performing frequency domain decomposition on the microstructural change signal sequence and selecting target change components related to the operating characteristics of building equipment, while simultaneously amplifying their amplitude, the expression intensity of weak changes at the signal level is enhanced, making early anomaly features more prominent and effectively compensating for the shortcomings of existing technologies in identifying subtle structural changes.

[0057] This invention fuses enhanced microstructural change information with equipment appearance structure information to form a composite feature image sequence. It then jointly extracts features related to the magnitude, direction, and temporal continuity of these changes at the spatial region level to construct a feature vector representing the equipment's operating state. This approach not only considers changes at a single pixel level but also integrates changes within and between regions, resulting in a more comprehensive and accurate description of the equipment's state. Compared to methods that rely solely on a single feature for judgment, this invention can more fully reflect the complex change patterns during equipment operation.

[0058] This invention improves the ARDiff decision model by introducing processing mechanisms such as microstructure enhancement feature spectroscopy, regional correlation diffusion modeling, time inversion difference determination, and state-level output. This allows for multi-level modeling and analysis of equipment operating state characteristics, enabling the detection results to not only reflect the current state but also the degree of deviation from the normal operating state. This processing method enhances the ability to detect abnormal changes, allowing equipment to be identified at an early stage of abnormal trends, thereby improving the sensitivity and accuracy of state detection.

[0059] This invention enables automatic detection of the operating status of building equipment based solely on image data without relying on additional sensors. It significantly improves the identification of minute changes, anti-interference capabilities, and accuracy of status determination, providing a more reliable technical means for the intelligent operation and maintenance of building equipment. Attached Figure Description

[0060] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0061] Figure 1 This is an overall flowchart of an automatic detection method for the status of building equipment based on image recognition proposed in this invention;

[0062] Figure 2This is a schematic diagram illustrating the construction of a microstructure change signal sequence in an image recognition-based automatic detection method for building equipment status proposed in this invention.

[0063] Figure 3 This is a schematic diagram illustrating the construction of an improved ARDiff decision model for an automatic detection method of building equipment status based on image recognition proposed in this invention. Detailed Implementation

[0064] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0065] refer to Figures 1-3 An automatic detection method for the status of building equipment based on image recognition includes the following steps:

[0066] The system continuously collects data on the operation of building equipment, automatically acquires continuous video data of the equipment's operation, and extracts image frame sequences.

[0067] Visual reference system construction processing is performed on the image frame sequence to extract stable feature points in the building equipment structure, and geometric alignment and illumination normalization processing are performed on each frame in the image frame sequence to obtain a standardized image sequence.

[0068] Subpixel-level motion estimation is performed on standardized image sequences to extract displacement information of each pixel over time, and a microstructure change signal sequence in the time dimension is constructed.

[0069] Frequency domain decomposition is performed on the microstructure change signal sequence to filter out the target change component, and amplitude amplification is performed on the target change component to generate a microstructure change enhanced sequence.

[0070] A change feature map is constructed based on the microstructure change enhancement sequence, and the change feature map is fused with the corresponding image in the normalized image sequence to form a composite feature image sequence;

[0071] The composite feature image sequence is divided into regions, and the change amplitude features, change direction features, and temporal continuity features of each region are extracted to construct the equipment operating status feature vector;

[0072] The feature vector of the equipment's operating status is input into the improved ARDiff decision model for analysis and processing, and the current operating status category of the equipment is output.

[0073] Based on the operating status category, corresponding status identification information is generated, and an alarm signal is automatically output when an abnormal status is detected.

[0074] In this embodiment, an image acquisition device is fixedly installed in the target monitoring area of ​​the building equipment. The image acquisition device is an industrial camera or a high-definition video camera with video acquisition function. The image acquisition device continuously acquires video of the building equipment operation process at a preset acquisition frame rate of 20 to 60 frames per second. The acquired continuous video data is analyzed frame by frame, and each frame of image data is extracted from the continuous video data and arranged in the order of acquisition time to generate an image frame sequence.

[0075] In this embodiment, obtaining the standardized image sequence specifically includes:

[0076] Extract stable feature points from the building equipment structure in each frame of the image frame sequence, and construct the feature point set corresponding to each frame image;

[0077] The process of constructing the feature point set is as follows: Each frame in the image frame sequence is read frame by frame. The target area of ​​the building equipment in the current frame image is located. Stable feature points in the equipment structure are extracted within the target area. Stable feature points include corner points of the equipment's outer contour, structural connection points, fixed installation points, edge intersection points, and pixels with significant texture changes that maintain positional correspondence across multiple consecutive frames. Each stable feature point is assigned a corresponding position number according to its spatial location in the current frame image, and the horizontal and vertical coordinates and corresponding frame number of each stable feature point are recorded. All stable feature points in the current frame image are summarized according to their position numbers to form the feature point set corresponding to the frame image.

[0078] The first frame in the image frame sequence is determined as the reference image, and the set of feature points corresponding to the reference image is used as the reference feature point set.

[0079] A spatial reference coordinate system is established based on the set of reference feature points to form a unified visual reference benchmark.

[0080] The process of forming a unified visual reference benchmark is as follows: Based on the two-dimensional pixel coordinates of each feature point in the reference feature point set, combined with the camera's intrinsic parameters and shooting geometry, these two-dimensional pixel points are mapped to three-dimensional space through geometric calibration or projection transformation methods to obtain the corresponding three-dimensional space coordinates; with the three-dimensional space coordinate points as a reference, a three-dimensional space reference coordinate system is established, and the origin, coordinate axis direction and scale unit of the coordinate system are determined to form a unified visual reference benchmark.

[0081] The feature point set corresponding to each frame in the image frame sequence is matched with the reference feature point set to establish a one-to-one correspondence between the feature points and obtain a set of matched feature point pairs.

[0082] The process of obtaining the matching feature point pair set is as follows: Using each stable feature point in the reference feature point set as the matching reference, the feature point set corresponding to the current frame image in the image frame sequence is read. Based on the point number, spatial proximity, and local image feature description information, the stable feature points in the current frame image are matched point-by-point with the stable feature points in the reference image. For stable feature points with the same point number and whose horizontal and vertical coordinate offsets are both within a preset matching threshold range, they are determined as candidate matching feature point pairs. The candidate matching feature point pairs are then subjected to consistency screening, eliminating those that are inconsistent with the spatial relationship of the building equipment structure, and retaining those that meet the criteria of consistent point number, consistent spatial proximity, and matching local feature description information. The retained candidate matching feature point pairs are then summarized according to their corresponding frame numbers to obtain the matching feature point pair set of the current frame image relative to the reference image.

[0083] The geometric transformation matrix of each frame in the image frame sequence relative to the reference image is calculated based on the set of matching feature point pairs;

[0084] The calculation process of the geometric transformation matrix is ​​as follows: Using the set of matching feature point pairs between the current frame image and the reference image as input, a coordinate correspondence is established between the coordinates of the current frame feature points in each matching feature point pair and the coordinates of the reference image feature points. Based on all coordinate correspondences, the affine transformation parameters from the current frame image to the reference image are solved. These affine transformation parameters include horizontal scale parameters, horizontal rotation and shearing parameters, horizontal translation parameters, vertical rotation and shearing parameters, vertical scale parameters, and vertical translation parameters. The horizontal translation parameter is the integer part of the horizontal coordinate between the stable feature points in the current frame image after scale and rotation / shear transformations and the corresponding stable feature points in the reference image. The volume position difference is calculated as follows: the vertical translation parameter is the overall position difference between the stable feature point in the current frame image and the corresponding stable feature point in the reference image in the vertical direction after scaling and rotation shearing transformations; the horizontal and vertical scale parameters are calculated based on the relative distance changes between the matching feature point pairs; the horizontal and vertical rotation shearing parameters are calculated based on the changes in the orientation angles between the matching feature point pairs and the influence of coordinate axis intersections; the horizontal scale parameter, horizontal rotation shearing parameter, horizontal translation parameter, vertical rotation shearing parameter, vertical scale parameter, and vertical translation parameter are combined into the geometric transformation matrix corresponding to the current frame image according to the two-dimensional affine transformation relationship.

[0085] Based on the geometric transformation matrix, coordinate mapping processing is performed on each frame of the image frame sequence, and combined with a unified visual reference datum, a geometrically aligned image sequence is obtained;

[0086] The process of obtaining the geometrically aligned image sequence is as follows: For each pixel in the current frame image, the original coordinates of the pixel are transformed to the target coordinate position under a unified visual reference using the geometric transformation matrix corresponding to the current frame image. The transformation process involves: taking the original x-coordinate and original y-coordinate of the pixel to be mapped in the current frame image, and combining the original x-coordinate, original y-coordinate, and a constant term to form a homogeneous coordinate vector; multiplying the geometric transformation matrix corresponding to the current frame image with the homogeneous coordinate vector to obtain the mapped homogeneous coordinate vector; and normalizing the mapped homogeneous coordinate vector to obtain the pixel's position under the unified visual reference. The target x-coordinate and target y-coordinate are used to form the target coordinate position of the pixel in the geometrically aligned image. When the target coordinate position falls within the target coordinate range, the pixel value of the pixel is written into the corresponding target coordinate position. When multiple pixels are mapped to the same target coordinate position, the pixel closest to the target coordinate position is selected. When there is no corresponding pixel value at the target coordinate position, the pixel value at the target coordinate position is supplemented by neighbor interpolation. The geometrically aligned image corresponding to the current frame image is obtained. The geometrically aligned images corresponding to each frame image are arranged according to the frame number order of the image frame sequence to obtain the geometrically aligned image sequence.

[0087] Illumination statistics are performed on each frame of the geometrically aligned image sequence to calculate the average pixel grayscale value and the pixel grayscale dispersion of each frame. Illumination normalization is then performed on each frame of the geometrically aligned image sequence to obtain a standardized image sequence.

[0088] The process of obtaining the standardized image sequence is as follows: For the current frame image in the geometrically aligned image sequence, perform pixel grayscale statistics, read the grayscale value of each pixel in the current frame image, sum all the grayscale values ​​and divide by the total number of pixels to obtain the average pixel grayscale value of the current frame image; calculate the difference between the grayscale value of each pixel in the current frame image and the average pixel grayscale value, square all the differences and sum them, divide by the total number of pixels, and take the square root of the result to obtain the pixel grayscale dispersion of the current frame image; using the average pixel grayscale value and pixel grayscale dispersion of the reference image as the illumination reference standard, perform illumination normalization processing on each pixel in the current frame image to obtain the brightness normalization value, multiply the brightness normalization value by the pixel grayscale dispersion of the reference image, and add the average pixel grayscale value of the reference image to obtain the normalized grayscale value of the pixel; reassemble the normalized grayscale values ​​of all pixels in the current frame image according to the original pixel coordinates to form a standardized image, and arrange the standardized images according to the frame number order to obtain the standardized image sequence.

[0089] In this embodiment, the construction of the microstructure change signal sequence specifically includes:

[0090] Read the normalized image sequence and determine the current normalized image and the next normalized image according to the frame number order. Specifically, read the frame numbers corresponding to each normalized image in the normalized image sequence in ascending order, determine the normalized image corresponding to the currently read frame number as the current normalized image, and determine the normalized image whose frame number is adjacent to the current normalized image and whose frame number is one bit greater than the current frame number as the next normalized image.

[0091] In the current standardized image, determine the pixel to be estimated, and in the next standardized image, determine the local search region corresponding to the pixel to be estimated;

[0092] The process of determining the pixel to be estimated and the local search region is as follows: Read all pixels within the target area of ​​the building equipment in the current standardized image; determine the pixels located at the edge of the building equipment structure, the neighborhood of stable feature points, the structural connection area, and the preset detection area as the pixels to be estimated; record the horizontal and vertical coordinates and the corresponding frame number of the pixels to be estimated in the current standardized image; based on the horizontal and vertical coordinates of the pixels to be estimated, in the next standardized image, using the same coordinate position as the search center, expand horizontally and vertically according to the preset search radius to form a rectangular area; determine the rectangular area as the local search region corresponding to the pixel to be estimated.

[0093] The local image block containing the pixel to be estimated is matched with the candidate image block in the local search area, and the initial displacement range of the pixel to be estimated between two adjacent normalized images is determined from the matching results.

[0094] The determination of the initial displacement range is as follows: A local image patch is extracted centered on the pixel to be estimated in the current standardized image. Within the local search area of ​​the next standardized image, candidate image patches of the same size as the local image patch are extracted one by one according to their pixel positions. The texture similarity value between the local image patch and each candidate image patch is calculated, and the candidate image patches are sorted in descending order of texture similarity value. The candidate image patch with the highest matching degree among the sorted results is selected, and the difference in horizontal and vertical coordinates between the center point of the candidate image patch and the pixel to be estimated is determined as the initial displacement center. Based on the matching results of the adjacent candidate image patches around the candidate image patch with the highest matching degree, a continuous candidate displacement interval around the initial displacement center is determined, and this continuous candidate displacement interval is used as the initial displacement range of the pixel to be estimated between two adjacent standardized images.

[0095] Perform subpixel-level motion estimation within the initial displacement range to obtain subpixel-level displacement information of the pixel to be estimated between two adjacent normalized frames;

[0096] The subpixel-level displacement information is obtained as follows: Within the initial displacement range, taking the integer pixel displacement position corresponding to the candidate image block with the highest matching degree as the center, the matching results corresponding to the horizontal adjacent position, vertical adjacent position, and diagonal adjacent position are read to construct a local matching surface around the integer pixel displacement position; a quadratic surface fitting is performed on the local matching surface to determine the subpixel coordinate position corresponding to the extreme value of the matching degree; the horizontal difference between the subpixel coordinate position and the original coordinate of the pixel to be estimated is determined as the horizontal displacement, and the vertical difference between the subpixel coordinate position and the original coordinate of the pixel to be estimated is determined as the vertical displacement; the horizontal displacement and the vertical displacement together constitute the subpixel-level displacement information of the pixel to be estimated between two adjacent normalized images;

[0097] The sub-pixel displacement information of the same pixel to be estimated between two adjacent standardized images is summarized according to the frame number order to construct the time displacement signal corresponding to the pixel to be estimated. The time displacement signals corresponding to each pixel to be estimated are summarized according to the pixel spatial position to construct the microstructure change signal sequence in the time dimension.

[0098] The construction process of the microstructure change signal sequence is as follows: Following the frame number order of the standardized image sequence, sub-pixel-level displacement information of the same pixel to be estimated is read between each pair of adjacent standardized images. The lateral and longitudinal displacements of the pixel to be estimated under different frame numbers are arranged sequentially to form the time displacement signal corresponding to the pixel to be estimated. The time displacement signal characterizes the minute motion changes of the pixel to be estimated over continuous time. Time displacement signal construction processing is performed on all pixels to be estimated within the target area of ​​the building equipment to obtain the time displacement signal corresponding to each pixel. Multiple time displacement signals are written to their corresponding spatial positions according to the pixel spatial positions of each pixel in the standardized image. The spatial positions and time displacement signals are jointly organized to form a microstructure change signal sequence in the time dimension, including pixel spatial position, lateral displacement change, longitudinal displacement change, and frame number order.

[0099] In this embodiment, the generation of the microstructure change enhancement sequence specifically includes:

[0100] Frequency domain decomposition is performed on the microstructure change signal sequence to convert the displacement change signal in the time domain into displacement frequency components in the frequency domain, thus obtaining transverse frequency domain components and longitudinal frequency domain components.

[0101] The process of obtaining the horizontal and vertical frequency domain components is as follows: The time displacement signal corresponding to each pixel to be estimated is read from the microstructure change signal sequence according to its spatial position. The time displacement signal is then split into a horizontal displacement time series and a vertical displacement time series. A discrete frequency domain transformation is performed on the horizontal displacement time series according to the frame number order to obtain the horizontal displacement amplitude and phase at different frequency points. The horizontal displacement amplitude and phase are then used together as the horizontal frequency domain component. Similarly, a discrete frequency domain transformation is performed on the vertical displacement time series according to the frame number order to obtain the vertical displacement amplitude and phase at different frequency points. The vertical displacement amplitude and phase are then used together as the vertical frequency domain component.

[0102] Based on the frequency range of structural response corresponding to micro-vibration, micro-displacement, or micro-deformation during the operation of building equipment, target change components are selected from the transverse and longitudinal frequency domain components.

[0103] The specific screening process for the target variation components is as follows: read the displacement amplitude and displacement phase corresponding to each frequency point in the transverse and longitudinal frequency components, compare each frequency point with the structural response frequency range corresponding to micro-vibration, micro-displacement, or micro-deformation during the operation of the building equipment, and retain the transverse and longitudinal frequency components that are within the structural response frequency range; perform amplitude continuity and phase continuity judgment on the retained transverse and longitudinal frequency components, and eliminate frequency components with isolated frequencies, abrupt amplitude changes, and discontinuous phases. The transverse and longitudinal frequency components that satisfy the structural response frequency range, amplitude continuity, and phase continuity after retention are determined as the target variation components.

[0104] The target variation component is amplified by multiplying its amplitude by an amplification factor to obtain the amplified target variation component. The amplification factor is determined based on the original amplitude of the target variation component, the image noise level, and the allowable displacement range of the building equipment structure.

[0105] The magnified target change components are recombined according to their corresponding pixel spatial positions and frame numbers to generate a microstructure change enhancement sequence.

[0106] The generation process of the microstructure change enhancement sequence is as follows: The amplified target change components are read, and according to the corresponding pixel spatial positions of the pixels to be estimated, the amplified lateral and longitudinal target change components are written into the positions corresponding to the pixels to be estimated; Time-domain reconstruction processing is performed on the written frequency domain components according to the frame number order, restoring the amplified displacement changes in the frequency domain to time-varying lateral and longitudinal enhanced displacement signals; The lateral and longitudinal enhanced displacement signals corresponding to the same pixel to be estimated are combined according to the frame number order to form the microstructure change enhancement signal for that estimated pixel; Microstructure change enhancement signals are generated for all pixels to be estimated within the target area of ​​the building equipment, and then summarized according to the pixel spatial positions of each pixel to be estimated to obtain the microstructure change enhancement sequence.

[0107] In this embodiment, the formation of the composite feature image sequence specifically includes:

[0108] Read the microstructure change enhancement sequence, and extract the lateral and longitudinal enhancement displacement information corresponding to each pixel to be estimated according to the pixel spatial location and frame number;

[0109] The extraction process of lateral and longitudinal enhancement displacement information is as follows: each microstructure change enhancement signal in the microstructure change enhancement sequence is read; the corresponding pixel to be estimated is determined according to the pixel spatial position recorded in the microstructure change enhancement signal; the corresponding time position is determined according to the frame number recorded in the microstructure change enhancement signal; at the time position, the enhancement displacement data corresponding to the pixel to be estimated is read; the displacement component that changes along the horizontal axis of the image in the enhancement displacement data is determined as the lateral enhancement displacement information; and the displacement component that changes along the vertical axis of the image in the enhancement displacement data is determined as the longitudinal enhancement displacement information.

[0110] Based on the horizontal and vertical enhanced displacement information, the change intensity value of each pixel to be estimated is calculated. The change intensity value of each pixel to be estimated is written into the image coordinate plane according to the corresponding pixel spatial position to construct a change feature mapping map.

[0111] The construction process of the change feature map is as follows: Read the horizontal and vertical enhancement displacement information corresponding to each pixel to be estimated under the same frame number; synthesize the horizontal and vertical enhancement displacement information of each pixel to be estimated to obtain the change intensity value of the pixel to be estimated; the change intensity value represents the microstructural change amplitude of the pixel to be estimated under the current frame number, and is obtained by taking the square root of the sum of the squares of the horizontal and vertical enhancement displacement information; according to the pixel spatial position of each pixel to be estimated in the standardized image sequence, write the change intensity value corresponding to each pixel to be estimated into the corresponding image coordinate position; for image coordinate positions where no pixel to be estimated exists, interpolation is performed to fill in the position based on the change intensity values ​​of adjacent pixels to be estimated; arrange all change intensity values ​​according to the image coordinate positions to form a change feature map with the same spatial size as the corresponding image in the standardized image sequence.

[0112] Read the corresponding image from the standardized image sequence according to the frame number, and align the change feature map with the corresponding image from the standardized image sequence in terms of spatial position;

[0113] Channel fusion processing is performed on the spatially aligned change feature map and the corresponding image in the normalized image sequence. The corresponding image in the normalized image sequence is used as the structural appearance channel, and the change feature map is used as the microstructure change channel to form a composite feature image. The composite feature images are arranged in order of frame number to form a composite feature image sequence.

[0114] The formation process of the composite feature image sequence is as follows: Read the corresponding image from the standardized image sequence according to the frame number, and read the change feature map under the same frame number; adjust the corresponding image and the change feature map to have the same image size and the same pixel coordinate range; use the pixel grayscale value of the corresponding image as the structural appearance channel data, and use the change intensity value of each pixel position in the change feature map as the microstructure change channel data, and perform channel superposition according to the same image size and the same pixel coordinate range to generate a composite feature image that simultaneously contains the structural appearance channel and the microstructure change channel; perform the above channel superposition process on each frame image in the standardized image sequence to obtain the composite feature image corresponding to each frame image; arrange the composite feature images according to the frame number order to form a composite feature image sequence.

[0115] In this embodiment, the construction of the device operating status feature vector specifically includes:

[0116] Read the composite feature image sequence, obtain each frame of composite feature image in order of frame number, and divide each frame of composite feature image into multiple detection regions;

[0117] Read the structural appearance channel data and microstructure change channel data in each detection area, and extract the change intensity value corresponding to each pixel in the detection area;

[0118] The extraction process of the change intensity value is as follows: All pixels falling within the detection area in the composite feature image are read according to the region boundary of the detection area. For each pixel, its structural appearance channel data and microstructure change channel data are read simultaneously. The microstructure change channel data stores the change intensity value corresponding to the pixel. The value of the pixel in the microstructure change channel is determined as the change intensity value corresponding to the pixel, and the pixel spatial position, frame number, and detection area number corresponding to the change intensity value are recorded. The above reading and recording process is repeated for all pixels within the detection area to obtain the change intensity value corresponding to each pixel within the detection area.

[0119] The change amplitude feature of each detection area is calculated based on the change intensity value. The change amplitude feature is the average change intensity value obtained by dividing the sum of the change intensity values ​​of all pixels in the detection area by the number of pixels in the detection area.

[0120] Based on the lateral and longitudinal enhanced displacement information of each pixel in the microstructure change channel data, the change direction features of each detection area are extracted.

[0121] The extraction process of change direction features is as follows: All pixels falling within the detection area from the microstructure change channel data are read according to the region boundary of the detection area. For each pixel, its lateral and longitudinal enhancement displacement information are read, and the displacement direction of the pixel is determined based on these information. The displacement directions of all pixels within the same detection area are statistically analyzed according to their angles to obtain the displacement distribution in different directions within the detection area. The direction with the highest proportion in the displacement distribution is selected as the main change direction of the detection area. The lateral change direction feature is determined based on the overall distribution of lateral enhancement displacement information within the detection area, and the longitudinal change direction feature is determined based on the overall distribution of longitudinal enhancement displacement information within the detection area. The main change direction, the lateral change direction feature, and the longitudinal change direction feature are combined as the change direction feature of the detection area.

[0122] The temporal continuity features of the detection region are extracted by statistically analyzing the change amplitude and change direction features of the same detection region in the composite feature image of consecutive frames according to the frame number sequence.

[0123] The extraction process of temporal continuity features is as follows: The amplitude and direction features of the same detection region in consecutive frame composite feature images are read sequentially according to frame number. The amplitude features are arranged according to frame number to form a temporal sequence of amplitude, and the direction features are arranged according to frame number to form a temporal sequence of direction. Continuity statistics are performed on the amplitude time series to extract the increasing / decreasing trend, duration, and fluctuation of amplitude between adjacent frames. Directional consistency statistics are performed on the direction time series to extract the degree of direction retention, directional shift, and duration of directional change between adjacent frames. The increasing / decreasing trend, duration, fluctuation, retention, shift, and duration of directional change are used as the temporal continuity features of the detection region.

[0124] The change amplitude features, change direction features, and time continuity features corresponding to each detection area are combined according to the detection area number to construct the equipment operation status feature vector;

[0125] The specific process of constructing the equipment operation status feature vector is as follows: read the change amplitude feature, change direction feature, and time continuity feature corresponding to each detection area according to the detection area number; arrange the change amplitude feature, change direction feature, and time continuity feature of the same detection area in a fixed order to form the area status feature sub-vector corresponding to the detection area; and concatenate the area status feature sub-vectors corresponding to all detection areas according to the detection area number order to obtain the equipment operation status feature vector representing the overall operation status of the building equipment.

[0126] In this embodiment, the output of the current operating status category of the device specifically includes:

[0127] The device operating status feature vector is input into the improved ARDiff judgment model for analysis and processing. The improved ARDiff judgment model includes a microstructure enhancement feature spectralization module, a regional correlation diffusion modeling module, a time inversion difference judgment module, and a state classification output module. The improvement of the improved ARDiff judgment model is as follows: The traditional ARDiff judgment model uses a single image feature as input and performs noise perturbation, feature reconstruction, or classification judgment on the single image feature through the diffusion process. The improved ARDiff judgment model inputs the device operating status feature vector into the microstructure enhancement feature spectralization module, the regional correlation diffusion modeling module, the time inversion difference judgment module, and the state classification output module in sequence. The microstructure enhancement feature spectrum is obtained through phase consistency spectralization processing. The spatial adjacency relationship, structural connection relationship, and microstructure change transmission direction between each detection area are modeled through the regional correlation diffusion map. The consistency of the diffusion features of isolated mutations is verified through abnormal propagation suppression processing. The normal state prediction feature is generated by combining time inversion denoising and reconstruction processing. The state deviation feature is obtained based on the difference between the normal state prediction feature and the verified regional correlation diffusion feature. The current operating status category of the device is output by using evidence reliability gating processing.

[0128] The improved ARDiff decision model is structured as follows: It sequentially connects a microstructure enhancement feature spectralization module, a regional correlation diffusion modeling module, a time-inversion difference determination module, and a state-level output module. The equipment operating state feature vector serves as the model input, with its data format being a feature sequence arranged by detection area number. Each detection area corresponds to a change amplitude feature, a change direction feature, and a time continuity feature. The microstructure enhancement feature spectralization module converts the equipment operating state feature vector into a microstructure enhancement feature spectrum and transmits this spectrum to the regional correlation diffusion modeling module according to the detection area number. The regional correlation diffusion modeling module performs structural propagation processing on the microstructure enhancement feature spectrum based on the regional correlation diffusion map, outputting regional correlation diffusion features arranged by detection area number and frame number. The time-inversion difference determination module receives the regional correlation diffusion features, generates normal state prediction features, compares the normal state prediction features with the regional correlation diffusion features, and outputs state deviation features arranged by detection area number and frame number. The state-level output module receives the state deviation features, calculates the state deviation degree index, and outputs the current operating state category of the equipment.

[0129] The training process of the improved ARDiff decision model is as follows: Training data comes from composite feature image sequences collected during the continuous operation of building equipment, and equipment operating status feature vectors constructed from these composite feature image sequences. The training data is labeled according to the collection time period and equipment operating status, including normal, warning, and abnormal states. For each training sample, the corresponding detection region number, frame number, change amplitude feature, change direction feature, and temporal continuity feature are recorded. The loss function of the improved ARDiff decision model includes the reconstruction difference loss between the normal state prediction features and the input region correlation diffusion features, and the loss between the state deviation features and the corresponding operating status labels. The improved ARDiff decision model incorporates classification loss between frames and temporal continuity constraint loss between state deviation features in adjacent frames. Training parameters include the number of training rounds, batch size, learning rate, state classification threshold update interval, and model parameter update step size. During training, the device operating state feature vector is input into the improved ARDiff decision model, which sequentially performs microstructure enhancement feature spectralization, regional correlation diffusion modeling, temporal inversion difference determination, and state classification output. The model parameters are then updated in reverse based on the loss function. When the decrease in the loss function is less than the set convergence threshold in several consecutive training rounds, and the output results of the operating state category of the verification samples remain stable, the improved ARDiff decision model is considered to have reached convergence.

[0130] In the microstructure enhancement feature spectralization module, the change amplitude features, change direction features, and time continuity features in the device operation state feature vector are subjected to phase-consistent spectralization to generate the microstructure enhancement feature spectrum;

[0131] The generation of the microstructure enhancement feature spectrum is as follows: The device operating status feature vector is read in the microstructure enhancement feature spectralization module. The amplitude of change, direction of change, and temporal continuity features corresponding to each detection area are extracted according to the detection area number. The amplitude of change features of the same detection area are used as the spectral intensity benchmark, the direction of change features as the spectral direction constraint, and the temporal continuity features as the phase-preserving constraint between consecutive frames. The amplitude of change features are arranged temporally according to the frame number to form a amplitude of change time series. The decomposition relationship of the amplitude of change time series in the horizontal and vertical directions is determined based on the direction of change features. Frequency components are extracted from the amplitude of change time series, retaining frequency components with consistent trends between consecutive frames and removing discrete jump components inconsistent with the temporal continuity features. The retained frequency components, direction of change constraints, and phase-preserving constraints are combined according to the corresponding detection area number to generate a microstructure enhancement feature spectrum characterizing the distribution relationship of microstructure changes in building equipment in spatial region, direction of change, and temporal phase.

[0132] In the regional correlation diffusion modeling module, a regional correlation diffusion map is constructed based on the spatial adjacency relationship, structural connection relationship and microstructure change transmission direction between each detection area in the building equipment structure. The microstructure enhancement feature spectrum is then input into the regional correlation diffusion map for structural propagation processing to obtain the regional correlation diffusion features.

[0133] The process of obtaining the regional correlation diffusion features is as follows: A regional correlation diffusion map is established according to the spatial location, structural connection relationship, and microstructure change transmission direction of each detection area within the building equipment structure. Each detection area is treated as a regional node. Diffusion edges are established between adjacent or structurally connected detection areas, and the correlation weight of the diffusion edges is determined based on the distance between the two detection areas, the connection strength, and the consistency of the change direction. The microstructure enhancement feature spectrum is input into the corresponding regional node according to the detection sub-region number, so that each regional node carries the corresponding microstructure enhancement feature spectrum. Structural propagation processing is performed on the microstructure enhancement feature spectrum of each regional node along the diffusion edges, so that the features of the current detection area are fused with the features of adjacent detection areas according to the correlation weight. After repeating the structural propagation processing on all detection areas, a regional correlation diffusion feature is obtained that simultaneously contains information on microstructure changes in the current area, information on transmission between adjacent areas, and information on structural connection relationships.

[0134] Anomaly propagation suppression processing is performed on the regional association diffusion characteristics. The anomaly propagation suppression processing includes identifying the diffusion characteristics of isolated mutations within a single detection region, and performing consistency verification on the diffusion characteristics of isolated mutations based on the regional association diffusion characteristics of adjacent detection regions to obtain the verified regional association diffusion characteristics.

[0135] The verification of the region-related diffusion features is obtained as follows: The region-related diffusion features corresponding to each detection region are read. The change amplitude of the diffusion features in the detection region is calculated according to the frame number, and the change amplitude is compared with the change amplitude of the diffusion features in adjacent detection regions under the same frame number. When the region-related diffusion features of a detection sub-region abruptly change within a single frame or a short period of consecutive frames, and the abrupt change does not form a corresponding change transmission relationship in adjacent detection regions, the diffusion feature corresponding to the abrupt change is determined as an isolated abrupt change diffusion feature. Consistency verification is performed on the isolated abrupt change diffusion features, including spatial consistency verification, directional consistency verification, and temporal consistency verification. When the isolated abrupt change diffusion features do not satisfy the spatial correlation relationship, microstructure change transmission direction, and continuous frame change trend of adjacent detection regions, the isolated abrupt change diffusion features are suppressed. When the isolated abrupt change diffusion features satisfy the spatial correlation relationship, microstructure change transmission direction, and continuous frame change trend of adjacent detection regions, the isolated abrupt change diffusion features are retained. The region-related diffusion features of each detection region after suppression are recombined according to the detection region number to obtain the verified region-related diffusion features.

[0136] In the time inversion difference determination module, time inversion denoising and reconstruction processing is performed on the verified regional correlation diffusion features. The normal state prediction features of building equipment under normal conditions are reconstructed in reverse order of frame number. The difference between the normal state prediction features and the verified regional correlation diffusion features is compared to obtain the state deviation features.

[0137] The state deviation feature is obtained as follows: The verified regional correlation diffusion features are read sequentially according to frame number, and the regional correlation diffusion feature corresponding to the last frame is used as the inversion starting feature. Temporal inversion denoising and reconstruction processing is performed frame by frame in reverse order of frame number. In each inversion process, the verified regional correlation diffusion feature of the current frame, the inversion reconstruction feature of the next frame, and the temporal continuity feature of the corresponding detection area are read. Diffusion disturbance components that do not satisfy the continuous change trend in the current frame are removed, and diffusion feature components that conform to the microstructural change law of normal operation of building equipment are retained to generate the normal state prediction feature corresponding to the current frame. The above inversion reconstruction process is repeated for all frame numbers to obtain the normal state prediction feature corresponding to each frame. The normal state prediction feature under the same frame number and the same detection area is compared item by item with the verified regional correlation diffusion feature, and the difference in change amplitude, change direction, and temporal continuity are calculated respectively. The difference in change amplitude, change direction, and temporal continuity are combined according to the corresponding frame number and detection area number to form a state deviation feature characterizing the degree of deviation of the actual microstructural change of the building equipment from the normal state prediction feature.

[0138] In the state classification output module, the state deviation features are subjected to evidence reliability gating processing. Based on the state deviation features after evidence reliability gating processing, a state deviation degree index is generated, and the current operating state category of the equipment is output according to the state deviation degree index.

[0139] The output of the current operating status category of the equipment is as follows: Evidence reliability gating is applied to the status deviation features. Evidence reliability gating refers to the selection and retention or weakening of status deviation features based on the consecutive occurrence frequency of the features, the consistency of adjacent areas in the corresponding detection sub-region, and the stability of the change direction. The status deviation features retained after evidence reliability gating are summarized according to the detection area number and frame sequence number to calculate the status deviation degree index. The status deviation degree index characterizes the overall deviation of the current operating status of the building equipment from the normal operating status. When the status deviation degree index is less than the preset first status classification threshold, the output of the current operating status category of the equipment is normal. When the status deviation degree index is greater than or equal to the preset first status classification threshold and less than the preset second status classification threshold, the output of the current operating status category of the equipment is warning status. When the status deviation degree index is greater than or equal to the preset second status classification threshold, the output of the current operating status category of the equipment is abnormal status.

[0140] In this embodiment, corresponding status identifier information is generated according to the current operating status category of the device. When the operating status category is normal, a normal status identifier is generated, and the normal status identifier, corresponding frame number, detection time period, and device number are stored. When the operating status category is warning, a warning status identifier is generated, and the warning status identifier, corresponding frame number, detection time period, device number, and status deviation index are stored. When the operating status category is abnormal, an abnormal status identifier is generated, and an alarm signal is automatically output.

[0141] Example 1: A circulating water pump in an underground equipment room of a building is used as the detection object. This circulating water pump operates continuously for a long time, and there are pipes, supports, valves, and fixed bases around the equipment. During operation, it is easily affected by factors such as water flow impact, bearing wear, loose fixing bolts, and slight displacement of the pump body. In the early stage of anomalies, the circulating water pump usually does not show obvious damage to its appearance, and it is difficult to directly observe the slight vibrations and displacements at the pump's outer contour, structural connection points, or fixed installation points during manual inspection. Only when the anomaly further expands will it manifest as obvious noise, severe vibration, or equipment shutdown. Therefore, this scenario can illustrate the problem that existing manual inspections and ordinary image recognition methods are insufficient to detect early anomalies in building equipment in a timely manner.

[0142] In this scenario, an image acquisition device is fixedly installed directly in front of the circulating water pump to continuously capture the pump's operation, automatically obtaining continuous video data of the building equipment's operation and extracting image frame sequences from this data. After the image frame sequences undergo visual reference system construction processing, the system extracts stable feature points at locations such as the pump's outer contour corners, the connection points between the pump body and the base, the mounting points of fixing bolts, and the intersection points of pipeline edges. Based on these stable feature points, the system performs geometric alignment and illumination normalization on each frame to obtain a standardized image sequence. Through this processing, even if there are slight flickering lights in the equipment room or slight vibrations in the image acquisition device due to environmental vibrations, each frame can still be analyzed under a unified visual reference standard.

[0143] The system performs sub-pixel-level motion estimation on standardized image sequences. Within the target area of ​​the building equipment, it identifies pixels to be estimated and calculates the lateral and longitudinal displacements of these pixels between adjacent frames, constructing a microstructure change signal sequence over time. This microstructure change signal sequence no longer represents only the appearance information in a single frame image but also the minute motion changes at different structural positions of the circulating water pump during continuous operation. Subsequently, the system performs frequency domain decomposition on the microstructure change signal sequence, filtering out target change components related to the pump's micro-vibrations, micro-displacements, or micro-deformations. These target change components are then amplified to generate an enhanced microstructure change sequence.

[0144] In this embodiment, the microstructure change enhancement sequence is further converted into a change feature map and fused with the corresponding image in the standardized image sequence to form a composite feature image sequence. The composite feature image sequence simultaneously contains both the structural appearance information and microstructure change information of the circulating water pump. The system divides the detection area according to the equipment's outer contour region, structural connection region, fixed installation region, and edge intersection region, extracting change amplitude features, change direction features, and temporal continuity features respectively to construct an equipment operating state feature vector. This equipment operating state feature vector is input into the improved ARDiff judgment model. The microstructure enhancement feature spectrum is generated by the microstructure enhancement feature spectralization module, the structural correlation between each detection region is analyzed by the regional correlation diffusion modeling module, the state deviation features are obtained by the time inversion difference judgment module, and finally, the current operating state category of the equipment is output by the state classification output module.

[0145] To verify the practical effectiveness of this method, operational data was continuously collected from the same circulating water pump for 30 days. The first 20 days represented normal operation. From day 21 to day 25, a fixed installation point was manually loosened to simulate slight bearing wear, creating an early warning state sample. From day 26 to day 30, the loosening was further increased to create an abnormal state sample. Comparison methods included manual inspection, standard image classification, and image detection without microstructural change enhancement. Each method was tested on the same collected data, and the early anomaly detection rate, false alarm rate, and average alarm lead time were statistically analyzed.

[0146] Table 1 Comparison of Automatic Status Detection Effects of Circulating Water Pumps in the Network

[0147] Manual inspection 42.3% 31.8% 6.5% 0.8 days Common image classification methods 61.7% 22.4% 9.8% 1.4 days Image detection methods without microstructural enhancement 73.6% 14.2% 7.1% 2.6 days Method of the present invention 91.8% 4.7% 3.9% 5.3 days

[0148] As shown in Table 1, manual inspection is weak in early anomaly identification. This is mainly because early loosening of fixed installation points and slight bearing wear do not immediately produce obvious visual changes; manual inspection can only detect anomalies after significant increases in equipment vibration or noise. While ordinary image classification methods can identify some visual anomalies, they primarily rely on salient visual features in equipment images and have limited ability to identify minute displacements at pump connection points and slight periodic vibrations in edge areas. Therefore, the anomaly false negative rate remains high.

[0149] Image detection methods without microstructural change enhancement show an improvement over conventional image classification methods, indicating that inter-frame variation information in continuous image sequences can reflect some changes in device status. However, because this method does not perform frequency domain decomposition and amplitude amplification on the microstructural change signal, weak changes are still easily affected by illumination fluctuations, acquisition noise, and slight camera shake. Therefore, the early anomaly detection rate is still lower than that of the method of this invention.

[0150] This invention reduces interference from inter-frame geometric offset and illumination changes by constructing a visual reference system. It extracts minute displacement information of the pixel to be estimated between consecutive frames through sub-pixel-level motion estimation, highlights target variation components through frequency domain decomposition and amplitude amplification, and analyzes the feature vector of equipment operating status using an improved ARDiff judgment model. This allows circulating water pumps to be identified even before anomalies develop into obvious visual defects. Experimental results show that the early anomaly identification rate of this invention reaches 91.8%, the anomaly false negative rate is reduced to 4.7%, and the average alarm lead time reaches 5.3 days. It can detect micro-vibration, micro-displacement, and micro-deformation anomalies in building equipment operation earlier, thus demonstrating the application value of this method in the automatic detection of building equipment status.

[0151] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An automatic detection method for the status of building equipment based on image recognition, characterized in that, Includes the following steps: The system continuously collects data on the operation of building equipment, automatically acquires continuous video data of the equipment's operation, and extracts image frame sequences. Visual reference system construction processing is performed on the image frame sequence to extract stable feature points in the building equipment structure, and geometric alignment and illumination normalization processing are performed on each frame in the image frame sequence to obtain a standardized image sequence. Subpixel-level motion estimation is performed on standardized image sequences to extract displacement information of each pixel over time, and a microstructure change signal sequence in the time dimension is constructed. Frequency domain decomposition is performed on the microstructure change signal sequence to filter out the target change component, and amplitude amplification is performed on the target change component to generate a microstructure change enhanced sequence. A change feature map is constructed based on the microstructure change enhancement sequence, and the change feature map is fused with the corresponding image in the normalized image sequence to form a composite feature image sequence; The composite feature image sequence is divided into regions, and the change amplitude features, change direction features, and temporal continuity features of each region are extracted to construct the equipment operating status feature vector; The feature vector of the equipment's operating status is input into the improved ARDiff decision model for analysis and processing, and the current operating status category of the equipment is output. Based on the operating status category, corresponding status identification information is generated, and an alarm signal is automatically output when an abnormal status is detected.

2. The automatic detection method for building equipment status based on image recognition according to claim 1, characterized in that, An image acquisition device is fixedly installed in the target monitoring area of ​​the building equipment. The image acquisition device continuously acquires video of the building equipment operation process at a preset acquisition frame rate. The acquired continuous video data is analyzed frame by frame, and each frame of image data is extracted from the continuous video data and arranged in the order of acquisition time to generate an image frame sequence.

3. The automatic detection method for building equipment status based on image recognition according to claim 1, characterized in that, The acquisition of the standardized image sequence specifically includes: Extract stable feature points from the building equipment structure in each frame of the image frame sequence, and construct the feature point set corresponding to each frame image; The first frame in the image frame sequence is determined as the reference image, and the set of feature points corresponding to the reference image is used as the reference feature point set. A spatial reference coordinate system is established based on the set of reference feature points to form a unified visual reference benchmark. The feature point set corresponding to each frame in the image frame sequence is matched with the reference feature point set to establish a one-to-one correspondence between the feature points and obtain a set of matched feature point pairs. The geometric transformation matrix of each frame in the image frame sequence relative to the reference image is calculated based on the set of matching feature point pairs; Based on the geometric transformation matrix, coordinate mapping processing is performed on each frame of the image frame sequence, and combined with a unified visual reference datum, a geometrically aligned image sequence is obtained; Illumination statistics are performed on each frame of the geometrically aligned image sequence to calculate the average pixel grayscale value and the pixel grayscale dispersion of each frame. Illumination normalization is then performed on each frame of the geometrically aligned image sequence to obtain a standardized image sequence.

4. The automatic detection method for building equipment status based on image recognition according to claim 1, characterized in that, The construction of the microstructure change signal sequence specifically includes: Read the normalized image sequence and determine the current normalized image and the next normalized image according to the frame number order; In the current standardized image, determine the pixel to be estimated, and in the next standardized image, determine the local search region corresponding to the pixel to be estimated; The local image block containing the pixel to be estimated is matched with the candidate image block in the local search area, and the initial displacement range of the pixel to be estimated between two adjacent normalized images is determined from the matching results. Perform subpixel-level motion estimation within the initial displacement range to obtain subpixel-level displacement information of the pixel to be estimated between two adjacent normalized frames; The sub-pixel displacement information of the same pixel to be estimated between two adjacent standardized images is summarized according to the frame number sequence to construct the time displacement signal corresponding to the pixel to be estimated. The time displacement signals corresponding to each pixel to be estimated are summarized according to the pixel spatial position to construct a microstructure change signal sequence in the time dimension.

5. The automatic detection method for building equipment status based on image recognition according to claim 1, characterized in that, The generation of the microstructure change enhancement sequence specifically includes: Frequency domain decomposition is performed on the microstructure change signal sequence to convert the displacement change signal in the time domain into displacement frequency components in the frequency domain, thus obtaining transverse frequency domain components and longitudinal frequency domain components. Based on the frequency range of structural response corresponding to micro-vibration, micro-displacement, or micro-deformation during the operation of building equipment, target change components are selected from the transverse and longitudinal frequency domain components. The amplitude of the target change component is amplified by multiplying the amplitude of the target change component by an amplification factor to obtain the amplified target change component. The magnified target change components are recombined according to their corresponding pixel spatial locations and frame numbers to generate a microstructure change enhancement sequence.

6. The automatic detection method for building equipment status based on image recognition according to claim 1, characterized in that, The formation of the composite feature image sequence specifically includes: Read the microstructure change enhancement sequence, and extract the lateral and longitudinal enhancement displacement information corresponding to each pixel to be estimated according to the pixel spatial location and frame number; Based on the horizontal and vertical enhanced displacement information, the change intensity value of each pixel to be estimated is calculated. The change intensity value of each pixel to be estimated is written into the image coordinate plane according to the corresponding pixel spatial position to construct a change feature mapping map. Read the corresponding image from the standardized image sequence according to the frame number, and align the change feature map with the corresponding image from the standardized image sequence in terms of spatial position; Channel fusion processing is performed on the spatially aligned change feature map and the corresponding image in the normalized image sequence. The corresponding image in the normalized image sequence is used as the structural appearance channel, and the change feature map is used as the microstructure change channel to form a composite feature image. The composite feature images are arranged in order of frame number to form a composite feature image sequence.

7. The automatic detection method for building equipment status based on image recognition according to claim 1, characterized in that, The construction of the device operating status feature vector specifically includes: Read the composite feature image sequence, obtain each frame of composite feature image in order of frame number, and divide each frame of composite feature image into multiple detection regions; In each detection area, read the structural appearance channel data and microstructure change channel data, and extract the change intensity value corresponding to each pixel in the detection area; The variation amplitude characteristics of each detection area are calculated based on the change intensity value; Based on the lateral and longitudinal enhanced displacement information of each pixel in the microstructure change channel data, the change direction features of each detection area are extracted. The temporal continuity features of the detection region are extracted by statistically analyzing the change amplitude and change direction features of the same detection region in the composite feature image of consecutive frames according to the frame number sequence. The variation amplitude features, variation direction features, and time continuity features corresponding to each detection area are combined according to the detection area number to construct the equipment operation status feature vector.

8. The automatic detection method for building equipment status based on image recognition according to claim 1, characterized in that, The output of the current operating status category of the device specifically includes: The equipment operating status feature vector is input into the improved ARDiff decision model for analysis and processing. The improved ARDiff decision model includes a microstructure enhancement feature spectroscopy module, a regional correlation diffusion modeling module, a time inversion difference determination module, and a state classification output module. In the microstructure enhancement feature spectralization module, the change amplitude features, change direction features, and time continuity features in the device operation state feature vector are subjected to phase-consistent spectralization to generate the microstructure enhancement feature spectrum; In the regional correlation diffusion modeling module, a regional correlation diffusion map is constructed based on the spatial adjacency relationship, structural connection relationship and microstructure change transmission direction between each detection area in the building equipment structure. The microstructure enhancement feature spectrum is then input into the regional correlation diffusion map for structural propagation processing to obtain the regional correlation diffusion features. Anomaly propagation suppression processing is performed on the regional association diffusion characteristics. The anomaly propagation suppression processing includes identifying the diffusion characteristics of isolated mutations within a single detection region, and performing consistency verification on the diffusion characteristics of isolated mutations based on the regional association diffusion characteristics of adjacent detection regions to obtain the verified regional association diffusion characteristics. In the time inversion difference determination module, time inversion denoising and reconstruction processing is performed on the verified regional correlation diffusion features. The normal state prediction features of building equipment under normal conditions are reconstructed in reverse order of frame number. The difference between the normal state prediction features and the verified regional correlation diffusion features is compared to obtain the state deviation features. In the state classification output module, the state deviation features are subjected to evidence reliability gating processing. Based on the state deviation features after evidence reliability gating processing, a state deviation degree index is generated, and the current operating state category of the equipment is output according to the state deviation degree index.

9. The automatic detection method for building equipment status based on image recognition according to claim 1, characterized in that, Based on the current operating status category of the device, corresponding status identification information is generated. When the operating status category is normal, a normal status identifier is generated, and the normal status identifier, corresponding frame number, detection time period, and device number are stored. When the operating status category is warning, a warning status identifier is generated, and the warning status identifier, corresponding frame number, detection time period, device number, and status deviation index are stored. When the operating status category is abnormal, an abnormal status identifier is generated, and an alarm signal is automatically output.