An audio and video device automatic patrol method for realizing fault reporting

By constructing a bit consumption matrix sequence and generating a residual feature vector sequence, the problem of difficulty in identifying faults in static scenes of video surveillance equipment in the existing technology is solved, and accurate detection of equipment freeze and loop playback faults is achieved.

CN121967677BActive Publication Date: 2026-07-31ZHEJIANG HARMAN AV TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG HARMAN AV TECH CO LTD
Filing Date
2026-03-31
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing video surveillance equipment inspection technologies struggle to accurately identify equipment malfunctions in long-term static scenarios. In particular, weak thermal noise is often misinterpreted as valid motion under efficient encoding mechanisms. Furthermore, some equipment malfunctions manifest as short-term loop playback in the buffer zone, making it difficult for existing detection methods to distinguish between normal static images and equipment malfunctions.

Method used

By acquiring multiple monitoring video frames when the video stream enters a macroscopic static state, a bit consumption matrix sequence is constructed, and a residual feature vector sequence is generated based on the centroid offset trajectory. The sequence fluctuation characteristics of the residual feature vector sequence are used to identify equipment freeze or loop playback faults.

Benefits of technology

It enables accurate identification of audio and video equipment malfunctions in static scenes, avoids misjudgment due to thermal noise, and improves the accuracy and reliability of fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967677B_ABST
    Figure CN121967677B_ABST
Patent Text Reader

Abstract

This application discloses an automated inspection method for audio-visual equipment to achieve fault reporting, relating to the field of monitoring equipment maintenance technology. The method includes: when a video stream is detected to enter a macroscopic static state, acquiring multiple monitoring video frames of the target device within a preset sampling period; constructing a bit consumption matrix sequence based on the bit consumption of each monitoring video frame; generating a residual feature vector sequence based on the centroid offset trajectory of each matrix in the bit consumption matrix sequence, the residual feature vector sequence being used to characterize the degree of centroid change when the screen display information of each monitoring video frame is refreshed; and reporting a device freeze fault signal or a loop playback fault signal if the sequence fluctuation characteristics of the residual feature vector sequence meet preset conditions. This application achieves the technical effect of accurately identifying fault conditions in audio-visual equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of monitoring equipment maintenance technology, specifically to an automated inspection method for audio and video equipment that enables fault reporting. Background Technology

[0002] Video surveillance systems are widely used in security, unmanned data centers, and industrial monitoring. To ensure the continuity and integrity of monitoring data, real-time monitoring of the operating status of front-end cameras and transmission links is necessary. Existing video surveillance equipment inspection technologies mainly rely on macro bitrate or image pixel difference for motion detection. However, in long corridors, unmanned data centers, and other long-term static scenarios, conventional detection algorithms face a dilemma: if the detection threshold is set too high, it cannot identify image freezes caused by equipment crashes; if the threshold is set too low, the inherent photoelectric and thermal noise of the sensor will be misjudged as valid motion. Especially under efficient encoding mechanisms such as H.264 / H.265, weak thermal noise is often filtered out because it is below the quantization step size, resulting in "normal static images" and "data interruptions caused by crashes" both appearing as extremely low or zero bit consumption in the bitstream data, making them difficult to distinguish. In addition, some equipment failures manifest as short-term loop playback in the buffer, with the macro bitrate and pixel difference still showing normal fluctuations, making them highly concealed. Existing amplitude-based detection methods are unable to accurately identify audio and video equipment failures. Summary of the Invention

[0003] To address the technical problem that amplitude-based detection methods in related technologies are unable to accurately identify fault conditions in audio and video equipment, this application provides an automated inspection method for audio and video equipment that enables fault reporting.

[0004] The specific technical solution adopted is as follows: When the video stream is detected to enter a macroscopic static state, multiple monitoring video frames of the target device within a preset sampling period are acquired. Based on the bit consumption of each monitored video frame, construct a sequence of bit consumption matrices; Based on the centroid offset trajectory of each matrix in the bit consumption matrix sequence, a residual feature vector sequence is generated. The residual feature vector sequence is used to characterize the degree of centroid change when the screen display information of each monitoring video frame is refreshed. If the sequence fluctuation characteristics of the residual feature vector sequence meet the preset conditions, report the equipment freeze fault signal or play the fault signal in a loop.

[0005] In one possible implementation of this application, a residual feature vector sequence is generated based on the centroid offset trajectory of each matrix in the bit consumption matrix sequence, including: Iterate through each bit consumption matrix in the bit consumption matrix sequence; Based on the bit consumption matrix, calculate the number of single-frame encoded bits for each monitored video frame; Based on the single-frame encoding bit count, each monitoring video frame is filtered to obtain multiple active frames. Based on the bit consumption matrix and the number of coded bits per frame, the centroid coordinates of each active frame are calculated. Based on the coordinate difference between each element in the bit consumption matrix and the centroid coordinate, the residual feature vector of each active frame is calculated, and the residual feature vectors of each active frame are integrated into a residual feature vector sequence.

[0006] In one possible implementation of this application, based on the single-frame coding bit count, each monitored video frame is filtered to obtain multiple active frames, including: Based on the number of encoded bits per frame, determine the lower limit of the number of encoded bits for each monitored video frame; Compare the number of coded bits per frame with the lower limit of the number of coded bits; If the number of encoded bits in a single frame is greater than the lower limit of the number of encoded bits, then the monitoring video frame corresponding to the number of encoded bits in a single frame will be taken as the active frame to be analyzed. If the number of encoded bits per frame is less than or equal to the lower limit of the encoding amount, the current monitored video frame is determined to be a dead frame.

[0007] In one possible implementation of this application, the residual feature vector of each active frame is calculated based on the coordinate difference between each element in the bit consumption matrix and the centroid coordinate, including: The spatial dispersion of each active frame is calculated based on the coordinate difference between each element in the bit consumption matrix and the centroid coordinate. Based on the centroid coordinates and spatial dispersion, the residual feature vector of each active frame is calculated.

[0008] In one possible implementation of this application, the residual feature vector of each active frame is calculated based on the centroid coordinates and spatial dispersion, including: Based on the centroid coordinates and spatial dispersion, determine the mean centroid and the mean dispersion. Based on the mean value of the centroid, the centroid coordinates are debiased and normalized to obtain the horizontal and vertical centroid components. Based on the mean of the dispersion, the spatial dispersion is debiased and normalized to obtain the dispersion components. By combining the horizontal centroid component, the vertical centroid component, and the dispersion component, the residual feature vector of each active frame is obtained.

[0009] In one possible implementation of this application, if the sequence fluctuation characteristics of the residual feature vector sequence meet a preset condition, a device freeze fault signal or a loop playback fault signal is reported, including: Calculate the Euclidean norm of each residual eigenvector in the residual eigenvector sequence to obtain the eigenvalue of each residual eigenvector; Based on the feature magnitude, the sequence fluctuation characteristics are determined, including the average amplitude of the feature fluctuations and the variance of the feature sequence. If the variance of the residual feature vector sequence is less than the preset noise floor threshold and the average amplitude of the feature fluctuation meets the preset freeze judgment condition, then a device freeze fault signal is reported. If the variance of the residual feature vector sequence is greater than or equal to the preset noise floor threshold, and the discrete autocorrelation coefficient corresponding to the feature magnitude satisfies the preset cyclic judgment condition, then a device freeze fault signal is reported.

[0010] In one possible implementation of this application, determining the sequence fluctuation characteristics based on the characteristic modulus includes: The average amplitude of characteristic fluctuations is calculated based on the average characteristic modulus of each residual eigenvector. The variance of the feature sequence is calculated based on the squared difference of the feature magnitude and the average amplitude of the feature fluctuation.

[0011] In one possible implementation of this application, if the variance of the feature sequence of the residual feature vector sequence is less than a preset noise floor threshold, and the average amplitude of the feature fluctuation meets a preset freeze determination condition, then a device freeze fault signal is reported, including: If the variance of the residual feature vector sequence is less than the preset noise floor threshold, then the feature sequence of the residual feature vector sequence is determined to be in a state of no fluctuation. The average amplitude of the characteristic fluctuation is compared with the preset freeze determination threshold; If the comparison results show that the average amplitude of the feature fluctuation is greater than or equal to the preset freeze judgment threshold, then the video frame monitored by the current target device is determined to be in a normal static state. If the comparison results show that the average amplitude of the characteristic fluctuation is less than the preset freeze judgment threshold, it is determined that the current target device has no effective coding operation and is in a completely frozen fault state, and the device freeze fault signal is reported.

[0012] In one possible implementation of this application, if the variance of the feature sequence of the residual feature vector sequence is greater than or equal to a preset noise floor threshold, and the discrete autocorrelation coefficient corresponding to the feature magnitude satisfies a preset cyclic judgment condition, then a device freeze fault signal is reported, including: If the variance of the residual feature vector sequence is greater than or equal to the preset noise floor threshold, then the feature sequence of the residual feature vector sequence is determined to be in a fluctuating state. Based on the feature magnitude and the average amplitude of feature fluctuations, the discrete autocorrelation coefficients of the residual feature vector sequence under different lag steps are calculated, and the maximum value of the correlation coefficient among the discrete autocorrelation coefficients is determined. If the maximum value of the correlation coefficient is less than or equal to the preset cyclic judgment threshold, the target device is determined to be in a normal state. If the maximum value of the correlation coefficient is greater than the preset loop judgment threshold, the target device is determined to be in a buffer loop playback fault state, and a device freeze fault signal is reported.

[0013] In one possible implementation of this application, a bit consumption matrix sequence is constructed based on the bit consumption of each monitored video frame, including: Determine the total number of bits occupied by multiple macroblocks in each monitoring video frame, and use the total number of bits as the bit consumption of the monitoring video frame; Based on the bit consumption of each monitored video frame, a sequence of bit consumption matrices is constructed, with each monitored video frame corresponding to a bit consumption matrix.

[0014] This application has, but is not limited to, the following technical effects: When the video stream is detected to enter a macroscopic static state, multiple monitoring video frames of the target device within a preset sampling period are acquired. Based on the bit consumption of each monitoring video frame, a bit consumption matrix sequence is constructed. Based on the centroid offset trajectory of each matrix in the bit consumption matrix sequence, a residual feature vector sequence is generated. When the sequence fluctuation characteristics of the residual feature vector sequence meet preset conditions, a device freeze fault signal or a loop playback fault signal is reported. In this application, a bit consumption matrix sequence is constructed based on the bit consumption of each monitoring video frame, the bit consumption of each video frame is quantified, and a residual feature vector sequence is generated based on the centroid offset trajectory of each matrix in the bit consumption matrix sequence. By capturing the micro-centroid changes in OSD (On-Screen Display) refresh, and then, when the sequence fluctuation characteristics of the residual feature vector sequence meet preset conditions, a device freeze fault signal or a loop playback fault signal is reported, thereby accurately identifying the fault condition of audio and video devices where the OSD has stopped refreshing. Attached Figure Description

[0015] Figure 1 A flowchart illustrating the first embodiment of the automated inspection method for fault reporting audio and video equipment according to this application; Figure 2 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of this application. Detailed Implementation

[0016] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0017] This application provides an automated inspection method for audio-visual equipment that implements fault reporting. In the first embodiment of this application's automated inspection method for audio-visual equipment that implements fault reporting, refer to... Figure 1 ,include: Step S10: When the video stream is detected to enter a macroscopic static state, acquire multiple monitoring video frames of the target device within a preset sampling period.

[0018] As an example, the automated inspection method for audio and video equipment that enables fault reporting can be applied to an automated inspection device for audio and video equipment that enables fault reporting. The automated inspection device for audio and video equipment that enables fault reporting belongs to an automated inspection system for audio and video equipment that enables fault reporting. The automated inspection system for audio and video equipment that enables fault reporting belongs to an automated inspection device for audio and video equipment that enables fault reporting.

[0019] As an example, in security monitoring scenarios, both static environments and frozen screens caused by device malfunctions exhibit extremely low data fluctuations at the macro-level bitrate, making it impossible to distinguish between them directly using bitrate magnitude. However, the micro-feature analysis required for subsequent steps relies on a continuous and stable data sample. Therefore, using the variance of the macro-level bitrate as a trigger condition can effectively locate static segments and initiate batch sampling, thus providing a complete statistical set for subsequent offline calculations. By monitoring macro-level bitrate fluctuations, static scenes in the monitoring video are filtered out, and a sampling batch buffer queue including multiple monitoring video frames is established. This method is suitable for video monitoring scenarios with periodic OSD (On-Screen Display) refreshes. The second-level refresh of OSD (such as timestamps) is a physical heartbeat characteristic of "device viability."

[0020] As an example, the target device can be a video surveillance device, and the preset sampling period can be 3s, 5s, etc., without any specific limitation.

[0021] As an example, during the acquisition of surveillance video frames, this system initially operates in streaming monitoring mode. The system is set to a duration of... The system establishes a macro-level bitrate fluctuation monitoring window (e.g., 5 seconds). Within this window, the system reads the encoded byte size of each P-frame in the video stream in real time and calculates its variance. The system also presets a macro-level stillness detection threshold. This threshold can be set based on historical statistical data, for example, taking 1 / 10 of the bitrate variance in normal motion scenarios. When the variance value within the monitoring window is lower than... When the system determines that the current video stream has entered a macroscopic static state, it immediately initializes a fixed-size K-frame (e.g., The sampling batch buffer queue (corresponding to approximately 3 seconds of video data) At this point, the system switches from streaming monitoring to batch sampling, and begins to fill the queue with subsequent eligible monitoring video frames. It should be noted that here... The total number of samples in a single detection cycle is defined. Once the queue is full, the collection of data for that batch is considered complete, and the system will stop writing and lock the data until the subsequent fault judgment module completes processing and releases the queue.

[0022] As an example, because I-frames (keyframes) in the H.264 / H.265 video coding standard primarily use intra-frame predictive coding to record the texture information of the image, their data volume is usually much larger than that of P-frames, and they contain a large number of DC components related to the image content. P-frames (predictive frames), on the other hand, use inter-frame predictive coding to record the residual changes in the image. In static scenes, the residuals of P-frames are mainly driven by OSD (On-Screen Display) refresh, sensor thermal noise, or slight changes in ambient light. To accurately capture these subtle dynamic changes, the interference from I-frames must be removed during sampling, retaining only P-frames to ensure the homogeneity of the sample data.

[0023] Specifically, for video streams entering the sampling stage, the system parses the header information of their Network Extraction Unit (NALU) to identify the frame type.

[0024] For H.264 encoding, parse the nal_unit_type field in the NALU Header. If the value is 5 (IDR image) or 1 (non-IDR slice) and the slice_type is I, then it is determined to be an I-frame.

[0025] For H.265 encoding, if the nal_unit_type field in the NALU Header is parsed and its value is 19-20 (IDR) or 16-21 (IRAP), it is determined to be an I-frame.

[0026] If the current frame is identified as an I-frame, the system discards the frame data directly, without occupying queue space. If the current frame is identified as a P-frame, the system retains it and stores it in the sampling batch buffer queue. .

[0027] To ensure the continuity of subsequent matrix calculations, the system assigns a unique intra-batch sequence index to each P-frame stored in the queue. ,in, The range of values ​​is strictly defined as follows: This index This represents the logical storage location of the data in the buffer queue, decoupled from the original playback timestamp (PTS) of the video stream. When indexing... Accumulate to When the queue is full, the sampling process ends.

[0028] As an example, the automated inspection method for audio-visual equipment that implements fault reporting can be applied to an automated inspection system for audio-visual equipment that implements fault reporting. This system consists of a closed-loop processing architecture composed of three logical functional modules: a data acquisition and preprocessing module (M1), a moment feature construction and bias removal module (M2), and a parallel fault decision module (M3). The processing flow and functions of each module are as follows: M1: Data Acquisition and Preprocessing Module M1.1 Macro Trigger Unit: Real-time monitoring of P-frame (predicted frame) bitstream variance, identification of the starting point of physical stillness, and triggering batch sampling.

[0029] M1.2 Frame Type Filtering Unit: Removes texture interference from I-frames (keyframes), retains only P-frames that reflect residual information, and maintains the sampling batch buffer queue.

[0030] M1.3 Spatial Mapping Unit: Extracts the bit consumption value of each macroblock and constructs a spatially aligned bit consumption matrix sequence. ).

[0031] M2: Moment feature construction and bias removal module; M2.1 Validity Determination Unit: Identifies "dead zone frames" (full Skip frames), performs adaptive threshold determination and numerical completion to prevent calculation errors.

[0032] M2.2 Moment Calculation Unit: Calculates the centroid coordinates (first moment, capturing OSD position) and spatial dispersion (second moment, characterizing distribution pattern) of bit consumption frame by frame.

[0033] M2.3 Bias Reduction and Normalization Unit: Calculates the batch mean as a baseline, decenters the features, and performs relative scaling normalization according to resolution, generating a residual feature vector sequence. ).

[0034] M3: Parallel fault decision module; M3.1 Statistical Analysis Unit: Calculating the variance of characteristic sequences ( ) and the average amplitude of characteristic fluctuations ( ).

[0035] M3.2 Diversion Decision Unit: Static branch ( ): Determine a complete freeze fault based on the average amplitude of characteristic fluctuations.

[0036] Dynamic branching ( ): Determining loop playback faults based on the autocorrelation of feature sequences.

[0037] M3.3 Status Output and Reset Unit: Outputs the final device status (normal / frozen / cycle) and resets the system.

[0038] Step S20: Construct a bit consumption matrix sequence based on the bit consumption of each monitored video frame.

[0039] As an example, since the video bitstream is a one-dimensional linear bitstream at the transport layer, it cannot directly reflect the distribution pattern of the encoded resources in the two-dimensional image space. In order to analyze the microscopic behavior of the encoder when dealing with OSD refresh and local thermal noise, it is necessary to map the linear data back to the two-dimensional physical coordinate system of the image and construct a data carrier with spatial topology. Based on this, a bit consumption matrix sequence is constructed according to the bit consumption of each monitoring video frame.

[0040] As an example, the bit consumption matrix sequence can be composed of... A sequence of bit consumption matrices, where each bit consumption matrix corresponds to the number of encoding bits occupied by each surveillance video frame.

[0041] Step S20 includes: Determine the total number of bits occupied by multiple macroblocks in each monitoring video frame, and use the total number of bits as the bit consumption of the monitoring video frame; Based on the bit consumption of each monitored video frame, a sequence of bit consumption matrices is constructed, with each monitored video frame corresponding to a bit consumption matrix.

[0042] As an example, each monitored video frame includes multiple macroblocks. In this embodiment, a macroblock refers to the smallest coding unit in a video coding standard, such as the H.264 standard. A macroblock of pixels; in the H.265 / HEVC standard, it corresponds to a Code Tree Unit (CTU), whose size is typically [size missing]. Pixel.

[0043] Specifically, the system obtains the number of macroblock grid rows of the image based on the video coding parameters. And macroblock grid column number For the buffer queue with index , The system analyzes the slice data of the monitored video frames, extracts the actual number of bits occupied by each macroblock in the compressed bitstream, and then assigns the data to the image's [frame number]. line, number Column (where, The total number of bits occupied by a macroblock is defined as the macroblock bit consumption value, denoted as . The sum of the following three parts, including but not limited to: Macroblock header information: contains information such as macroblock type and encoding mode; Motion vector difference: Represents the positional offset of the current macroblock relative to the reference frame; Transform coefficient residuals: Characterize the pixel residual data after prediction.

[0044] For each monitored video frame index The system constructs a dimension as A two-dimensional matrix, with corresponding positions Fill in the matrix.

[0045] After iterating through all the queues After each surveillance video frame, the system generates a file containing... The time series of a matrix is ​​defined as the bit consumption matrix sequence. ,Right now . It is a static three-dimensional data tensor (time × height × width).

[0046] Step S30: Based on the centroid offset trajectory of each matrix in the bit consumption matrix sequence, generate a residual feature vector sequence. The residual feature vector sequence is used to characterize the degree of centroid change when the screen display information of each monitoring video frame is refreshed.

[0047] As an example, this step aims to use statistical moment analysis to transform discrete macroblock bit consumption data into continuous spatial morphological features. In this embodiment, an OSD (On-Screen Display) refresh mechanism is introduced as a physical anchor point for device survival. In security monitoring screens, even if the background is completely still, the OSD area (such as a timestamp) will refresh at a fixed frequency (usually 1Hz). This refresh causes periodic bit consumption in the macroblock at the corresponding position of the encoder. By calculating the centroid of the bit consumption matrix, the spatial displacement of the encoding hotspot caused by OSD refresh can be accurately captured, thereby distinguishing whether the device is alive / operating normally at the microscopic level.

[0048] As an example, a residual feature vector sequence can be a time series composed of multiple feature vectors, which serves as a complete digital profile of the micro-state of the target device for subsequent fault mode identification.

[0049] Step S30 includes steps S31 to S35: Step S31: Traverse each bit consumption matrix in the bit consumption matrix sequence.

[0050] Step S32: Calculate the number of single-frame encoded bits for each monitored video frame based on the bit consumption matrix.

[0051] As an example, in the constructed bit consumption matrix sequence, due to the quantization dead-zone mechanism of the video encoder, when the residual amplitude of the input signal is lower than the quantization step size, the encoder will output "dead-zone frames" (full-skip frames) that only contain slice header information and have no actual macroblock residual data. If such video frames are directly used in subsequent moment calculations, the denominator will approach zero, causing instability in numerical calculations. In addition, under extremely high-quality encoding or pure black scenes, the entire batch may consist of dead-zone frames. Therefore, validity determination and numerical completion must be performed before feature calculation to ultimately determine multiple active frames for subsequent analysis and calculation.

[0052] As an example, the system traverses the bit consumption matrix sequence. Each bit consumption matrix in (in ), the system calculates the first Number of coded bits per frame The calculation formula is as follows:

[0053] in, and These represent the total number of rows and columns of the macroblock grid in each bit consumption matrix.

[0054] Step S33: Based on the single-frame encoding bit quantity, each monitoring video frame is filtered to obtain multiple active frames.

[0055] As an example, in order to adapt to different video encoding levels and resolutions, the system sets an adaptive lower limit for the effective encoding amount. The number of active frames can be determined by calculating the number of single-frame encoded bits and then comparing the number of single-frame encoded bits corresponding to each monitoring video frame with the lower limit of the encoding amount.

[0056] Step S33 includes: Based on the number of encoded bits per frame, the lower limit of the encoding amount for each monitored video frame is determined.

[0057] In one embodiment, the lower limit of the coding amount is set to 5% of the average bitrate of the current sampling batch or the theoretical minimum overhead of the slice head. The calculation method for the lower limit of the coding amount is as follows:

[0058] In another embodiment, the theoretical minimum overhead of the H.264 / H.265 standard slice header is... It can be set to 40 bits, where K refers to the number of surveillance video frames collected.

[0059] Compare the number of coded bits per frame with the lower limit of the number of coded bits.

[0060] If the number of encoded bits in a single frame is greater than the lower limit of the number of encoded bits, then the monitoring video frame corresponding to the number of encoded bits in a single frame is taken as the active frame to be analyzed.

[0061] If the number of encoded bits per frame is less than or equal to the lower limit of the encoding amount, the current monitored video frame is determined to be a dead frame.

[0062] As an example, the system is based on the number of bits encoded in a single frame. Lower limit of coding size The following logic applies to the relationship: Decision logic: If The current frame is determined to be an active frame, and subsequent feature calculations are performed.

[0063] Complete the logic: If If the current frame is determined to be a dead frame, then no further moment calculations are performed, and its corresponding centroid offset (residual) is directly set to 0.

[0064] Full batch silent fallback: If after traversing the entire batch, all frames are found to be dead frames (i.e., ... For all monitored video frames (Established) indicates that the screen is in a state of extreme stillness, and even OSD refresh has not triggered effective encoding (e.g., OSD is turned off or the screen is completely black); at this time, the system directly outputs a sequence of residual feature vectors containing all zero features. The current state is marked as "silent state," directly triggering the subsequent freeze decision process and skipping intermediate calculation steps. Where k=1 is a dead frame (or an I-frame placeholder), the feature value is initialized to the geometric center of the image (…). =N / 2, =M / 2) and the dispersion is 0, instead of forward indexing.

[0065] Step S34: Based on the bit consumption matrix and the number of coded bits per frame, calculate the centroid coordinates of each active frame.

[0066] As an example, since OSD information (such as timestamps) is usually located in the corner of the image, its digital jumps per second constitute a local coding hotspot. The distribution of sensor thermal noise on the imaging plane is usually random across the entire image. Calculating the first-order moment of origin (centroid) can compress the two-dimensional bit consumption distribution into a single coordinate point. When the OSD refreshes, this centroid coordinate will shift significantly towards the location of the OSD. When the OSD does not refresh (only background noise), the centroid will randomly oscillate around the center of the image. Therefore, the centroid coordinate is a key feature for capturing the "heartbeat" of the OSD.

[0067] Specifically, for video frames determined to be active frames The system is based on a bit consumption matrix. Calculate the centroid coordinates of bit consumption, including the horizontal centroid coordinates. and vertical centroid coordinates .

[0068] The formula for calculating the horizontal centroid coordinates is:

[0069] The formula for calculating the vertical centroid coordinates is:

[0070] in, Column indexes representing macroblocks, The row index representing the macroblock. The weight of the corresponding position is the value of any element in the bit consumption matrix. The centroid coordinates reflect the weighted average position of the encoded resources on the image plane at the current moment.

[0071] Step S35: Based on the coordinate difference between each element in the bit consumption matrix and the centroid coordinate, calculate the residual feature vector of each active frame, and integrate the residual feature vectors of each active frame into a residual feature vector sequence.

[0072] As an example, through the above calculations, the originally complex two-dimensional bit distribution is reduced to a time-varying trajectory signal. A normal OSD refresh will cause this track to exhibit obvious periodic pulses or displacements, while a device freeze will cause the track to disappear or remain stationary for a long time.

[0073] Furthermore, a second-order central moment is introduced to describe the spatial morphology of hotspots. Through bias removal and normalization, the influence of individual device differences (such as resolution and installation background) on the features is eliminated, constructing a normalized residual feature vector corresponding to each active frame. Then, the residual feature vectors corresponding to the monitoring video frames of each active frame are integrated to obtain a result from... A time series composed of residual eigenvectors is defined as a residual eigenvector sequence. ,in, This represents the residual feature vector of the first active frame, and so on. This sequence serves as a complete digital profile of the device's micro-state for subsequent fault mode identification.

[0074] Step S35 includes: The spatial dispersion of each active frame is calculated based on the coordinate difference between each element in the bit consumption matrix and the centroid coordinate.

[0075] As an example, while a single centroid coordinate can capture the positional offset of the OSD, in some extreme cases (e.g., uniformly distributed thermal noise across the entire screen and OSD jumps concentrated at a single point may have the same geometric centroid), the centroid alone cannot distinguish the type of noise. Therefore, utilizing the centroid... Using the origin as a reference, spatial dispersion is calculated. Spatial dispersion is used to characterize the topological looseness (i.e., radius of inertia) of the coding hotspots in space.

[0076] As an example, taking the k-th frame as an example, spatial dispersion The calculation formula is as follows:

[0077] This formula calculates the average of the weighted Euclidean distances of all macroblocks relative to the current centroid. The meanings of the letters have been explained above and will not be repeated here.

[0078] Physical meaning: When the image only contains background thermal noise, the noise distribution is relatively uniform. The values ​​are relatively large; when an OSD refresh (single-point jump) occurs, the encoding hotspots are highly concentrated in the OSD area. The values ​​will decrease significantly. This "breathing effect" of dispersion is another dimension of evidence for the normal operation of the equipment.

[0079] Based on the centroid coordinates and spatial dispersion, the residual feature vector of each active frame is calculated.

[0080] As an example, in actual monitoring scenarios, differences in camera installation angle, lighting conditions, and background texture complexity will introduce a fixed DC bias component to the encoding distribution. For instance, if the upper half of the image contains trees with complex textures and the lower half contains a flat road surface, the centroid coordinates will be biased upwards for a long time. In order to extract the AC component that purely reflects the micro-dynamic behavior of OSD refresh and thermal noise, and to avoid the complexity of pre-training the background model, this step performs debiasing processing on the centroid coordinates and spatial dispersion to obtain the residual feature vector of each active frame.

[0081] Furthermore, the resolution differences between different monitoring devices are enormous (from CIF to 4K), resulting in a large number of macroblock grids. The differences are so great that directly using absolute coordinates will lead to inconsistent eigenvalue dimensions, making subsequent decision thresholds unusable. Therefore, this application adopts a normalization method based on relative proportions.

[0082] The step of calculating the residual feature vector for each active frame based on the centroid coordinates and spatial dispersion includes: Based on the centroid coordinates and spatial dispersion, the mean centroid and the mean dispersion are determined.

[0083] As an example, when completing the task... video frames After calculation (including completed values), the system first calculates the batch background baseline value for this batch, including the mean centroid value. and mean of dispersion The calculation method is as follows:

[0084] The background baseline value for this batch represents the static "DC component" of the background texture within the current batch.

[0085] Based on the mean centroid, the centroid coordinates are debiased and normalized to obtain the horizontal and vertical centroid components.

[0086] Based on the mean of the dispersion, the spatial dispersion is debiased and normalized to obtain the dispersion components.

[0087] By combining the horizontal centroid component, the vertical centroid component, and the dispersion component, the residual feature vector of each active frame is obtained.

[0088] As an example, the system iterates through each frame in the batch again to construct the first... Residual feature vector of a frame This vector is formed by coupling three components: a decentralized and relatively normalized horizontal centroid, a vertical centroid, and a spatial dispersion. Specifically:

[0089] Formula explanation: Debiasing (numerator): Subtracting the mean eliminates the fixed positional bias introduced by the background texture, retaining only the fluctuation relative to the center of the background.

[0090] Relative normalization (denominator): centroid component divided by the number of grid columns and number of rows Convert absolute coordinates to percentages relative to the image's length and width (the theoretical range of values ​​is within...). between).

[0091] Dispersion component divided by (i.e., the square of the diagonal length), converting the area moment into a dimensionless scale relative to the total area of ​​the image.

[0092] Through the above processing It becomes a dimensionless vector decoupled from the device resolution. Regardless of whether the device is 720P or 4K, the relative fluctuation caused by OSD refresh will be on the same order of magnitude, thus making the subsequent fault judgment threshold widely applicable.

[0093] Step S40: If the sequence fluctuation characteristics of the residual feature vector sequence meet the preset conditions, report the equipment freeze fault signal or loop the fault signal.

[0094] As an example, the preset conditions include preset freeze judgment conditions and preset loop judgment conditions. The two judgment conditions represent two branches of fault existence. Based on the variance-based parallel decision architecture, the faults existing in the device are determined. This architecture uses the "activity" (variance) of the feature sequence to divide the processing flow into "static branches" and "dynamic branches" to detect two completely different fault modes, namely complete freeze and loop playback, thereby solving the logical deadlock problem of weak loop faults being misjudged as freeze.

[0095] Step S40 includes steps S41 to S44: Step S41: Calculate the Euclidean norm of each residual eigenvector in the residual eigenvector sequence to obtain the eigenvalue of each residual eigenvector. As an example, for a sequence of feature vectors If the sequence has been marked as "silent" (i.e., a full batch dead zone) in the previous steps, the statistical result is set to zero and the process jumps to the quiescent branch. Otherwise, the system needs to calculate the overall activity and average amplitude of the sequence.

[0096] Specifically, the system first calculates the residual feature vector for each residual in the sequence. The L2 norm (Euclidean norm) is denoted as the eigenmode. :

[0097] in, These represent the three-dimensional components of the residual eigenvector.

[0098] Step S42: Based on the feature magnitude, determine the sequence fluctuation characteristics, which include the average amplitude of the feature fluctuations and the variance of the feature sequence. Step S42 includes: The average amplitude of characteristic fluctuations is calculated based on the average characteristic modulus of each residual eigenvector.

[0099] As an example, the average amplitude of characteristic fluctuations The average distance of a sequence from zero is used to characterize the sequence's deviation from the zero point and is calculated as follows:

[0100] The variance of the feature sequence is calculated based on the squared difference of the feature magnitude and the average amplitude of the feature fluctuation.

[0101] As an example, the variance of the feature sequence The method used to characterize the activity level of a sequence over time (i.e., whether the OSD is fluctuating) is as follows:

[0102] Step S43: If the variance of the feature sequence of the residual feature vector sequence is less than the preset noise floor threshold and the average amplitude of the feature fluctuation meets the preset freeze judgment condition, then report the equipment freeze fault signal. As an example, the system sets a very small preset noise floor threshold. (For example, This threshold is used to distinguish whether a feature sequence is in a "dormant / micro-fluctuation" state or an "active / significantly fluctuating" state. The system is based on... and The relationship enters different decision branches, when When the system enters the static branch (for a complete freeze fault), it indicates that the feature sequence has almost no fluctuations and the OSD refresh feature disappears. At this time, the system mainly relies on the amplitude to confirm whether it is frozen.

[0103] Step S43 includes: If the variance of the residual feature vector sequence is less than the preset noise floor threshold, then the feature sequence of the residual feature vector sequence is determined to be in a state of no fluctuation.

[0104] The average amplitude of the characteristic fluctuation is compared with the preset freeze determination threshold; If the comparison results show that the average amplitude of the characteristic fluctuation is greater than or equal to the preset freeze judgment threshold, then the video frame monitored by the current target device is determined to be in a normal static state.

[0105] As an example, the system checks the average amplitude of characteristic fluctuations. Compared with the preset freeze determination threshold The relationship is due to the relative normalization of the average amplitude of the characteristic fluctuations. It can be set to a general experience value (e.g.) ).

[0106] like This indicates that although the fluctuation is extremely small, there is a non-zero DC component (which may be a residual image from a still image under extremely high-quality encoding), and the system judges it as a "normal still state".

[0107] If the comparison results show that the average amplitude of the characteristic fluctuation is less than the preset freeze judgment threshold, it is determined that the current target device has no effective coding operation and is in a completely frozen fault state, and the device freeze fault signal is reported.

[0108] As an example, if It may have been marked as "silent state", indicating that the device has no valid coding activity (no OSD, no background noise). The system determines that the device is in a "complete freeze fault" state and generates a level 1 alarm signal / device freeze fault signal.

[0109] Step S44: If the variance of the feature sequence of the residual feature vector sequence is greater than or equal to the preset noise floor threshold, and the discrete autocorrelation coefficient corresponding to the feature magnitude satisfies the preset cyclic judgment condition, then the equipment freeze fault signal is reported.

[0110] As an example, when When this occurs, it indicates that there are significant fluctuations in the feature sequence (OSD is refreshing or there is strong noise). At this time, the system needs to further analyze whether this fluctuation is random (normal) or mechanically periodic (fault).

[0111] As an example, the preset cycle determination condition can be that the maximum value of the discrete autocorrelation coefficient is greater than the preset cycle determination threshold.

[0112] Step S44 includes: If the variance of the residual feature vector sequence is greater than or equal to the preset noise floor threshold, then the feature sequence of the residual feature vector sequence is determined to be in a fluctuating state.

[0113] Based on the characteristic modulus and the average amplitude of characteristic fluctuations, the discrete autocorrelation coefficients of the residual eigenvector sequence under different lag steps are calculated, and the maximum value of the correlation coefficient among the discrete autocorrelation coefficients is determined.

[0114] As an example, the system calculates the residual eigenvector sequence at different lag steps. ( Discrete autocorrelation coefficient under ) To prevent the denominator from being zero during the calculation (even though at this time...). However, protection is still needed under extreme floating-point operations. A very small factor to prevent division by zero is introduced into the denominator. (For example Specifically, the discrete autocorrelation coefficient The calculation method can be:

[0115] in, This represents the eigenvalue of the k-th residual eigenvector. Indicates the average amplitude of characteristic fluctuations. Indicates the first The feature magnitude of each residual eigenvector.

[0116] System search The maximum value, which is also the maximum value of the correlation coefficient. and with the preset cyclic judgment threshold (For example, 0.8, 0.95, etc., can also be based on the historical maximum autocorrelation coefficient of normal equipment under the same coding configuration.) Settings, for example, = +Δ) are compared.

[0117] If the maximum value of the correlation coefficient is less than or equal to the preset cyclic judgment threshold, the target device is determined to be in a normal state. If the maximum value of the correlation coefficient is greater than the preset loop judgment threshold, the target device is determined to be in a buffer loop playback fault state, and a device freeze fault signal is reported.

[0118] As an example, if This indicates that the fluctuation is random (consistent with normal OSD refresh or noise characteristics), and the system determines that the device is in a static (normal) state in the real environment.

[0119] As an example, if This indicates that the feature sequence has extremely strong periodic repetition, which does not conform to the random characteristics of OSD refresh or thermal noise. The system determines that the device is in a "buffer loop playback fault" state and generates a secondary alarm signal.

[0120] After completing the above branch decision, the system outputs the final device status conclusion and records it in the operation and maintenance log. Subsequently, the system releases the sampling batch buffer queue established in the aforementioned steps. Clear intermediate data in memory (such as bit consumption matrix sequence). and residual eigenvector sequence The system then resets the monitoring state machine to the reflux monitoring state, awaiting the next triggering of macroscopic static conditions, thus completing a full automated inspection loop. This mechanism ensures the system can operate stably for a long time and will not fail due to memory leaks.

[0121] This application provides an automated inspection method for audio-visual equipment to achieve fault reporting. When the video stream is detected to enter a macroscopic static state, multiple monitoring video frames of the target device within a preset sampling period are acquired. Based on the bit consumption of each monitoring video frame, a bit consumption matrix sequence is constructed. Based on the centroid offset trajectory of each matrix in the bit consumption matrix sequence, a residual feature vector sequence is generated. When the sequence fluctuation characteristics of the residual feature vector sequence meet preset conditions, a device freeze fault signal or a loop playback fault signal is reported. In this application, a bit consumption matrix sequence is constructed based on the bit consumption of each monitoring video frame, the bit consumption of each video frame is quantified, and a residual feature vector sequence is generated based on the centroid offset trajectory of each matrix in the bit consumption matrix sequence. By capturing the micro-centroid changes in OSD (On-Screen Display) refresh, and then, when the sequence fluctuation characteristics of the residual feature vector sequence meet preset conditions, a device freeze fault signal or a loop playback fault signal is reported, thereby accurately identifying the fault condition of audio-visual equipment where the OSD has stopped refreshing.

[0122] Reference Figure 2 , Figure 2 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of this application.

[0123] like Figure 2 As shown, the automated inspection device for audio-visual equipment that implements fault reporting may include: a processor 1001, a memory 1003, and a communication bus 1002. The communication bus 1002 is used to realize the connection and communication between the processor 1001 and the memory 1003.

[0124] Optionally, the automated inspection device for audio and video equipment that enables fault reporting may also include a user interface, a network interface, a camera, RF (Radio Frequency) circuitry, sensors, a WiFi module, etc. The user interface may include a display screen and an input submodule such as a keyboard; optionally, the user interface may also include standard wired or wireless interfaces. The network interface may include standard wired or wireless interfaces (such as a Wi-Fi interface).

[0125] Those skilled in the art will understand that Figure 2 The structure of the automated inspection equipment for audio and video devices that realizes fault reporting shown in the figure does not constitute a limitation on the automated inspection equipment for audio and video devices that realizes fault reporting. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0126] like Figure 2As shown, the memory 1003, serving as a storage medium, may include an operating system, a network communication module, and an automated inspection program for audio-visual equipment that implements fault reporting. The operating system is a program that manages and controls the hardware and software resources of the automated inspection equipment for audio-visual equipment that implements fault reporting, supporting the operation of the automated inspection program for audio-visual equipment that implements fault reporting, as well as other software and / or programs. The network communication module is used to enable communication between the various components within the memory 1003, as well as communication with other hardware and software in the automated inspection system for audio-visual equipment that implements fault reporting.

[0127] exist Figure 2 In the automated inspection device for audio and video equipment that implements fault reporting, the processor 1001 is used to execute the automated inspection program for audio and video equipment that implements fault reporting stored in the memory 1003, and implement the steps of the automated inspection method for audio and video equipment that implements fault reporting as described above.

[0128] The specific implementation method of the automated inspection device for audio and video equipment that realizes fault reporting in this application is basically the same as the embodiments of the automated inspection method for audio and video equipment that realizes fault reporting described above, and will not be repeated here.

[0129] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0130] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0132] The above are merely preferred embodiments of this application and do not limit the scope of this application. Any equivalent structural or procedural transformations made based on the description and drawings of this application, or direct or indirect applications in other related technical fields, are similarly included within the scope of protection of this application.

[0133] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0134] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. An audio and video device automatic patrol inspection method for realizing fault reporting, characterized in that, The method includes: When the video stream is detected to enter a macroscopic static state, multiple monitoring video frames of the target device within a preset sampling period are acquired; wherein, the multiple monitoring video frames are P frames retained after removing I frames; Based on the bit consumption of each of the monitoring video frames, a bit consumption matrix sequence is constructed, including: determining the total number of bits occupied by multiple macroblocks in each of the monitoring video frames, and using the total number of bits as the bit consumption of the monitoring video frame; based on the bit consumption of each of the monitoring video frames, a bit consumption matrix sequence is constructed, with each monitoring video frame corresponding to a bit consumption matrix. Based on the centroid offset trajectory of each matrix in the bit consumption matrix sequence, a residual feature vector sequence is generated. The residual feature vector sequence is used to characterize the degree of centroid change when the screen display information of each monitoring video frame is refreshed; the screen display information includes a timestamp. If the sequence fluctuation characteristics of the residual feature vector sequence meet preset conditions, a device freeze fault signal or a loop playback fault signal is reported, including: calculating the Euclidean norm of each residual feature vector in the residual feature vector sequence to obtain the feature magnitude of each residual feature vector; determining the sequence fluctuation characteristics based on the feature magnitude, the sequence fluctuation characteristics including the average amplitude of feature fluctuations and the variance of feature sequences; if the variance of the feature sequences of the residual feature vector sequence is less than a preset noise floor threshold, and the average amplitude of feature fluctuations meets a preset freeze judgment condition, then a device freeze fault signal is reported; if the variance of the feature sequences of the residual feature vector sequence is greater than or equal to a preset noise floor threshold, and the discrete autocorrelation coefficient corresponding to the feature magnitude meets a preset loop judgment condition, then a device freeze fault signal is reported. The step of generating a residual feature vector sequence based on the centroid offset trajectory of each matrix in the bit consumption matrix sequence includes: traversing each bit consumption matrix in the bit consumption matrix sequence; calculating the single-frame encoded bit quantity of each monitoring video frame based on the bit consumption matrix; filtering each monitoring video frame based on the single-frame encoded bit quantity to obtain multiple active frames; calculating the centroid coordinates of each active frame based on the bit consumption matrix and the single-frame encoded bit quantity; calculating the residual feature vector of each active frame based on the coordinate difference between each element in the bit consumption matrix and the centroid coordinates; and integrating the residual feature vectors of each active frame into a residual feature vector sequence. wherein the barycentric coordinates comprise horizontal barycentric coordinates and vertical barycentric coordinates whose formulae are as follows: wherein, is the number of bits for single frame encoding, M is the number of macroblock grid rows, N is the number of macroblock grid columns, j represents the column index of the macroblock, i represents the row index of the macroblock, is the bit weight for the corresponding position; k is the intra-batch sequence index; The formula for calculating the number of coded bits per frame is as follows: The formula for calculating the residual eigenvector is as follows: in, For residual eigenvectors, For spatial dispersion, The average of the horizontal centroids. The average value of the vertical centroid. The mean of the dispersion; Among them, spatial dispersion Horizontal centroid mean Mean of vertical centroid Mean of dispersion The calculation formula is as follows: Where K is the number of surveillance video frames.

2. The automated inspection method for audio and video equipment to achieve fault reporting as described in claim 1, characterized in that, Based on the single-frame encoded bit count, the monitoring video frames are filtered to obtain multiple active frames, including: Based on the single-frame encoding bit quantity, a lower limit for the encoding quantity of each monitoring video frame is determined; Compare the single-frame coded bit count with the lower limit of the coded bit count; If the number of single-frame encoded bits is greater than the lower limit of the number of encoded bits, then the monitoring video frame corresponding to the number of single-frame encoded bits is taken as the active frame to be analyzed. If the number of encoded bits per frame is less than or equal to the lower limit of the encoding amount, then the current monitored video frame is determined to be a dead frame.

3. The automated inspection method for audio and video equipment to achieve fault reporting as described in claim 1, characterized in that, The step of determining the sequence fluctuation characteristics based on the feature magnitude includes: The average amplitude of characteristic fluctuation is calculated based on the average characteristic magnitude of each residual characteristic vector. The variance of the feature sequence is calculated based on the feature magnitude and the squared difference of the average amplitude of the feature fluctuation.

4. The automated inspection method for audio-visual equipment to achieve fault reporting as described in claim 1, characterized in that, If the variance of the residual feature vector sequence is less than a preset noise floor threshold, and the average amplitude of the feature fluctuation meets a preset freeze determination condition, then a device freeze fault signal is reported, including: If the variance of the feature sequence of the residual feature vector sequence is less than the preset noise floor threshold, then the feature sequence of the residual feature vector sequence is determined to be in a state of no fluctuation. The average amplitude of the characteristic fluctuation is compared with a preset freeze determination threshold; If the comparison result shows that the average amplitude of the characteristic fluctuation is greater than or equal to the preset freeze determination threshold, then it is determined that the video frame monitored by the current target device is in a normal static state. If the comparison result shows that the average amplitude of the characteristic fluctuation is less than the preset freeze judgment threshold, it is determined that the current target device has no effective coding operation and is in a completely frozen fault state, and a device freeze fault signal is reported.

5. The automated inspection method for audio and video equipment to achieve fault reporting as described in claim 1, characterized in that, If the variance of the residual feature vector sequence is greater than or equal to a preset noise floor threshold, and the discrete autocorrelation coefficient corresponding to the feature magnitude satisfies a preset cyclic judgment condition, then a device freeze fault signal is reported, including: If the variance of the feature sequence of the residual feature vector sequence is greater than or equal to a preset noise floor threshold, then the feature sequence of the residual feature vector sequence is determined to be in a fluctuating state. Based on the feature magnitude and the average amplitude of the feature fluctuation, the discrete autocorrelation coefficients of the residual feature vector sequence under different lag steps are calculated, and the maximum value of the correlation coefficient among the discrete autocorrelation coefficients is determined. If the maximum value of the correlation coefficient is less than or equal to the preset cyclic judgment threshold, the target device is determined to be in a normal state. If the maximum value of the correlation coefficient is greater than the preset loop judgment threshold, the target device is determined to be in a buffer loop playback fault state, and a device freeze fault signal is reported.