An audio automatic recognition switching method and system applied to a monitoring camera
By constructing a coupled perturbation hidden state model and particle filter estimation, and combining the changes in ground potential and power amplifier supply voltage, an audio silence flag is generated. This solves the problem of accurate identification in complex environments for automatic audio switching methods of surveillance cameras, and achieves more stable audio switching and higher sound quality performance.
Patent Information
- Application Number
- CN202511803980.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2045-12-03
AI Technical Summary
Existing automatic audio switching methods for surveillance cameras struggle to accurately identify appropriate switching times in complex environments with mechanical and electrical disturbances from the pan-tilt unit. This makes it difficult to suppress the impact of mechanical noise, electrical noise, and switching transients on audio monitoring and system stability.
By constructing a coupled disturbance hidden state model, combining particle filter estimation and variational marginalization calculation, and comprehensively considering the gimbal movement, ground potential change and power amplifier supply voltage, disturbance characteristic values and avoidance time windows are generated. Then, a silent trigger signal is generated using audio buffer and zero crossover state, and a switching control signal is output.
In real-world monitoring scenarios with multiple sources of interference, it significantly reduces noise and false triggering during audio switching, improving sound quality and system stability.
Smart Images

Figure CN121260182B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and more specifically, to a method and system for automatic audio recognition and switching applied to surveillance cameras. Background Technology
[0002] With the widespread adoption of video surveillance systems, pan-tilt-zoom (PTZ) cameras are widely used in urban security, building management, and industrial sites. To ensure sound coverage from different directions, distances, and with different audio pickup devices, surveillance systems often employ multiple audio channels, automatically switching the appropriate pickup path based on the PTZ's rotation. Typically, a switching command is issued when the audio signal is detected to be in a silent phase or the PTZ rotation is low, minimizing the impact of switching on the monitoring audio quality.
[0003] In existing technologies, automatic audio switching often relies on simple energy thresholds, zero-crossing count statistics, or envelope detection as the basis for determining silence. Some solutions incorporate thresholds for gimbal angular velocity or angular acceleration to avoid switching during violent rotation. On the electrical side, most systems only roughly determine whether the system is in a stable state by monitoring whether the power amplifier supply voltage is within the allowable range. However, in real-world engineering environments, transient disturbances from gimbal motor drives, gimbal mechanism movements, and power and ground references can collectively affect the audio link, making it difficult to accurately and promptly identify the appropriate switching moment using only the aforementioned simple thresholds.
[0004] On the one hand, the acceleration and deceleration of the gimbal transmits mechanical vibrations and structural noise to the microphone through the gimbal support structure and camera housing. Simply using angular velocity or angular acceleration below a certain fixed threshold as a "stability" criterion fails to reflect the complex nonlinear relationship between mechanical disturbances and acoustic coupling, potentially leading to incorrect switching during periods when significant coupling noise still exists. On the other hand, changes in gimbal drive, power load, and the ambient electromagnetic environment can easily cause transient fluctuations in the power amplifier's supply voltage and ground potential. When the switching command falls during these transient moments, audio output often exhibits popping sounds, glitches, and other interference. However, traditional methods that only monitor DC voltage amplitude struggle to identify these transient electrical disturbances in a timely manner. Furthermore, relying solely on silent detection of the continuous audio stream while ignoring the audio buffer frame structure and waveform zero-crossing positions can easily cause switching operations to occur in the middle of the audio frame or near non-zero-crossing points, further amplifying perceived abrupt changes in sound.
[0005] Therefore, in complex operating environments with both mechanical and electrical disturbances from the pan-tilt unit, existing automatic audio switching methods for surveillance cameras still struggle to accurately identify appropriate audio switching opportunities and effectively suppress the combined impact of mechanical noise, electrical noise, and switching transients on audio monitoring and system stability. To address this, it is necessary to propose a new automatic audio switching method to solve the problem of how to more accurately identify appropriate audio switching opportunities in surveillance cameras under complex operating environments with both mechanical and electrical disturbances from the pan-tilt unit, thereby minimizing the impact on sound quality and system stability during automatic audio switching. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this application provides a method and system for automatic audio recognition and switching applied to surveillance cameras.
[0007] In a first aspect, this application provides a method for automatic audio recognition and switching applied to surveillance cameras, including:
[0008] Obtain the angular acceleration sequence characterizing the motion state of the gimbal; construct a coupled perturbation hidden state model based on the angular acceleration sequence, wherein the hidden state includes the coupling strength; establish a nonlinear observation model based on the coupled perturbation hidden state model;
[0009] Based on the nonlinear observation model, particle filtering estimation is performed on the hidden state. In each particle update process, variational marginalization calculation is performed on the continuous variables in the hidden state to obtain the posterior distribution of the hidden state.
[0010] Within a preset prediction window, the statistics of the coupling strength are determined based on the posterior distribution, a disturbance feature value is generated, and the avoidance time window of the gimbal movement is determined based on the disturbance feature value.
[0011] Collect ground potential change signals and power amplifier supply voltage, calculate the rate of change of the power amplifier supply voltage, and generate an electrical stability flag based on the ground potential change signals and the rate of change of the power amplifier supply voltage.
[0012] The system detects the frame alignment flag of the audio buffer and the zero-crossing state of the current audio waveform, and generates an audio silence flag based on the detection results; it also generates a silence trigger signal based on the audio silence flag.
[0013] When the current time point is not within the avoidance time window and the electrical stability flag and the audio silence flag are valid, output a switching control signal.
[0014] Optionally, the hidden state further includes a coupling direction angle and an extended scale parameter, and the method further includes:
[0015] The coupling direction angle and the extended scale parameter are defined in a two-dimensional direction vector space;
[0016] An elliptic parametric equation is established in the two-dimensional direction vector space, where the major axis of the ellipse corresponds to the coupling direction angle, and the minor axis corresponds to the orthogonal direction of the coupling direction angle.
[0017] The scale ratio and rotation angle of the ellipse are calculated by fitting the variance ratio of the angular acceleration sequence in the major axis direction and the minor axis direction.
[0018] The coupled perturbation hidden state model, including directional constraint terms, is generated based on the obtained ellipse parameters.
[0019] Optionally, establishing the nonlinear observation model includes:
[0020] The angular acceleration sequence is directionally projected along the coupling direction angle and its orthogonal direction to obtain a first direction projection sequence and a second direction projection sequence.
[0021] The adaptive time window and bandwidth are determined according to the extended scale parameters, and the sub-band energy measurement with the same scale is calculated for the first direction projection sequence and the second direction projection sequence.
[0022] To address impulse noise and ground potential abrupt changes, a robust noise model based on a skewed t-distribution is established, generating observation residuals.
[0023] A dependency structure model is performed on the edge distributions of the first direction projection sequence and the second direction projection sequence using a connection function to obtain the joint observation distribution;
[0024] Based on the joint observation distribution and the observation residuals, the parameter set and observation mapping of the nonlinear observation model are determined.
[0025] Optionally, obtaining the posterior distribution of the hidden state includes:
[0026] After detecting the silence trigger signal, the instantaneous value of angular acceleration corresponding to the gimbal attitude is read, and the acoustic spatial drift is calculated based on the instantaneous value of angular acceleration.
[0027] The prior mean of the continuous variables in the particle set is corrected based on the acoustic spatial drift, thereby limiting the particle update range;
[0028] During the detection period when the electrical stability indicator is valid, variational marginalization calculation is performed on the continuous variable to generate a marginalized distribution of the continuous variable;
[0029] When the silence trigger signal is detected to change from valid to invalid, the distribution of the continuous variable is used as the initial input for particle propagation to update the posterior distribution of the hidden state.
[0030] Optionally, the generation of the marginalized continuous variable distribution includes:
[0031] Based on the acoustic spatial drift and gimbal attitude, a preset acoustic coupling calibration table is invoked to obtain the prior interval of coupling strength and the coupling noise baseline under the corresponding attitude.
[0032] Under the constraint of the prior interval of the coupling strength, a variational distribution family for characterizing the continuous variable is selected, and the mean vector and covariance matrix of the continuous variable are iteratively updated within the variational distribution family;
[0033] In each iteration update, the residual noise power measured during the effective period of the silent trigger signal is compared with the coupling noise power predicted by the variational distribution. When the deviation between the two exceeds a first preset threshold, the corresponding particle weight is reduced and the prior interval of the coupling strength is tightened.
[0034] When the variational optimization convergence condition is met, the final variational distribution is taken as the marginalized continuous variable distribution.
[0035] Optionally, the generated perturbation feature values include:
[0036] During the assembly stage of the surveillance camera, the low-order resonant frequency band of the camera housing and mounting bracket, as well as the unit impact response function related to the motion coupling of the pan-tilt unit, are obtained in advance through simulation analysis, and the low-order resonant frequency band and the unit impact response function are stored in the structural mode library.
[0037] Within the preset prediction time window, a time series estimate of coupling strength is generated based on the posterior distribution, and the time series estimate of coupling strength is convolved with the unit impulse response function to obtain the structural vibration response sequence at the microphone installation location.
[0038] Based on at least one of the peak amplitude or root mean square value of the structural vibration response sequence within the preset prediction time window, a statistical quantity characterizing the amplitude of structural vibration is obtained, and the statistical quantity is used as the disturbance characteristic value.
[0039] Optionally, the statistical measures used to characterize the amplitude of structural vibration include:
[0040] For each candidate switching moment, within a local time window centered on that candidate switching moment, the structural vibration response sequence is decomposed into frequency bands, and the ratio of energy falling into the low-order resonance frequency band to the total energy is calculated. The obtained energy ratio is used as the structural resonance risk index for the candidate switching moment.
[0041] The structural resonance risk index is compared with a preset resonance threshold. When the structural resonance risk index is higher than the preset resonance threshold, the corresponding time period is marked as the avoidance time window of the gimbal movement.
[0042] Optionally, the generation of the audio silence flag includes:
[0043] During the deployment phase of the surveillance camera, audio buffer data is collected to characterize the audio buffer structure. Based on the starting position of the audio frame in the audio buffer data, a reference range for the frame start position is determined, and the reference range is stored as a frame header template.
[0044] During operation, for real-time audio frames written to the audio buffer, the frame alignment flag of the audio frame is determined based on whether the time information and sampling position of the audio frame fall within the allowable offset range given by the frame alignment template.
[0045] For audio frames with valid frame alignment flags, the zero-crossing feature and energy feature of the current audio waveform are extracted in its starting neighborhood. When the zero-crossing feature and the energy feature meet the preset silence determination conditions, the audio frame is marked as a silence candidate frame.
[0046] Based on the distribution of silent candidate frames within a preset time window, determine whether the audio state of the corresponding time period is silent, and generate the audio silence flag when it is determined to be silent.
[0047] Optionally, marking the audio frame as a silent candidate frame includes:
[0048] During the deployment phase of the surveillance camera, background audio buffer data is collected for different pan-tilt positions. Based on the zero-crossing characteristics and amplitude characteristics of the silent segment audio frames under each pan-tilt position, a corresponding background noise template is generated, and the association between each pan-tilt position and the background noise template is stored in the silent judgment library.
[0049] During operation, the corresponding background noise template is retrieved from the silence determination library according to the current gimbal attitude. For audio frames with valid frame alignment, zero-crossing features and amplitude features are extracted in their starting neighborhood, and the deviation metric between the zero-crossing features and amplitude features and the background noise template is calculated.
[0050] When the deviation metric meets the preset silence threshold condition, the audio frame is determined as the silence candidate frame.
[0051] Secondly, this application provides an automatic audio recognition and switching system for surveillance cameras, comprising:
[0052] The module is used to acquire the angular acceleration sequence characterizing the motion state of the gimbal; construct a coupled perturbation hidden state model based on the angular acceleration sequence, wherein the hidden state includes the coupling strength; and establish a nonlinear observation model based on the coupled perturbation hidden state model.
[0053] The calculation module performs particle filtering estimation on the hidden state based on the nonlinear observation model. During each particle update, it performs variational marginalization calculation on the continuous variables in the hidden state to obtain the posterior distribution of the hidden state. Within a preset prediction window, it determines the statistics of the coupling strength based on the posterior distribution, generates perturbation feature values, and determines the avoidance time window of the gimbal movement based on the perturbation feature values.
[0054] The processing module is used to collect ground potential change signals and power amplifier supply voltage, calculate the rate of change of the power amplifier supply voltage, and generate an electrical stability flag based on the ground potential change signal and the rate of change of the power amplifier supply voltage; detect the frame alignment flag of the audio buffer and the zero crossover state of the current audio waveform, and generate an audio mute flag based on the detection results; and generate a mute trigger signal based on the audio mute flag.
[0055] The output module is used to output a switching control signal when the current time point is not within the avoidance time window and the electrical stability flag and the audio silence flag are valid.
[0056] Compared with existing technologies, this application introduces an estimation mechanism that combines coupled perturbation hidden state modeling with particle filtering and variational marginalization. Instead of simply approximating the gimbal motion state with a threshold of gimbal angular velocity or angular acceleration, it constructs a hidden state to characterize the coupling strength from the angular acceleration sequence and obtains the posterior distribution of the hidden state using a nonlinear observation model. Within a preset prediction window, it generates perturbation feature values based on the statistics of coupling strength and determines the avoidance time window of the gimbal motion. This allows for a more accurate and robust characterization of the actual impact of gimbal mechanical perturbation on the audio link.
[0057] This application comprehensively considers both the ground potential change signal and the rate of change of the power amplifier supply voltage to generate an electrical stability indicator. It is no longer limited to a rough judgment of the steady-state amplitude of the supply voltage, and can more sensitively identify power supply transient fluctuations and ground reference disturbances, avoiding audio switching during periods of electrical instability and reducing the probability of pops and electrical noise. Furthermore, this application uses the frame alignment flag of the audio buffer and the zero-crossing state of the current audio waveform to generate an audio silence flag on the audio side, and generates a silence trigger signal based on this. This aligns the silence determination and switching time with the buffer frame boundary and the near-zero segment of the waveform, reducing the auditory impact caused by waveform truncation and phase abrupt changes.
[0058] By combining the pan-tilt-zoom (PTZ) movement avoidance time window, electrical stability indicator, and audio silence indicator as prerequisites for the switching control signal output, this application enables audio switching actions to occur during periods of low PTZ mechanical disturbance, stable electrical environment, and audio silence. This significantly reduces noise and false triggering during audio switching in real monitoring scenarios with multiple sources of interference, thereby improving sound quality and overall system stability. Attached Figure Description
[0059] Figure 1 A flowchart illustrating an automatic audio recognition and switching method for a surveillance camera, provided as an embodiment of this application;
[0060] Figure 2 A flowchart illustrating a method for obtaining the posterior distribution of hidden states, provided in an embodiment of this application;
[0061] Figure 3 This application provides a schematic diagram of an audio output path and channel switching structure for a surveillance camera.
[0062] Figure 4 A schematic diagram of an automatic audio recognition and switching system for a surveillance camera provided in this application embodiment;
[0063] Figure 5 This application provides a schematic diagram of an audio power amplifier circuit and speaker connection structure.
[0064] Figure 6 A schematic diagram of an audio channel selection switch circuit structure provided in this application embodiment;
[0065] Figure 7 This is a schematic diagram of an external speaker insertion detection circuit provided in an embodiment of this application. Detailed Implementation
[0066] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0067] See Figure 1 The diagram shows a flowchart of an automatic audio recognition and switching method for a surveillance camera provided in an embodiment of this application, including steps S101 to S105, wherein:
[0068] S101: Obtain the angular acceleration sequence characterizing the motion state of the gimbal; construct a coupled perturbation hidden state model based on the angular acceleration sequence, wherein the hidden state includes the coupling strength; establish a nonlinear observation model based on the coupled perturbation hidden state model;
[0069] S102: Based on the nonlinear observation model, particle filtering estimation is performed on the hidden state. During each particle update, variational marginalization calculation is performed on the continuous variables in the hidden state to obtain the posterior distribution of the hidden state.
[0070] S103: Within a preset prediction window, determine the statistics of the coupling strength based on the posterior distribution, generate disturbance feature values, and determine the avoidance time window for the gimbal movement based on the disturbance feature values; collect ground potential change signals and power amplifier supply voltage, calculate the rate of change of the power amplifier supply voltage, and generate an electrical stability flag based on the ground potential change signals and the rate of change of the power amplifier supply voltage.
[0071] S104: Detect the frame alignment flag of the audio buffer and the zero-crossing state of the current audio waveform, and generate an audio silence flag based on the detection result; generate a silence trigger signal based on the audio silence flag;
[0072] S105: When the current time point is not within the avoidance time window and the electrical stability flag and the audio silence flag are valid, output a switching control signal.
[0073] Regarding the above S101:
[0074] In one embodiment, the surveillance camera used to execute the audio automatic recognition and switching method of this application includes: a pan-tilt mechanism, a motor and its driving circuit for driving the pan-tilt, a sensing unit for collecting pan-tilt attitude information, an audio acquisition unit, and a main control processor. The main control processor can be a microprocessor with floating-point operation capabilities, a DSP, or a SoC with an AI acceleration unit, used to run the coupling disturbance estimation and audio switching control algorithm of this application.
[0075] The main control processor acquires the angular acceleration sequence characterizing the gimbal's motion state. Specifically, the horizontal rotation axis and pitch axis of the gimbal can be equipped with angle encoders, gyroscope sensors, or inertial measurement units with three-axis accelerometers and three-axis gyroscopes. The main control processor periodically reads the angle or angular velocity data of each axis at a sampling period Ts, and obtains the corresponding angular acceleration sequence through differential operations and digital filtering.
[0076] For example, the difference between adjacent angular velocities can be calculated at each sampling time k, and a first- or second-order low-pass filter can be used to suppress measurement noise, thereby obtaining discrete-time angular acceleration samples β(k), which are then continuously arranged within the monitoring period to form an angular acceleration sequence {β(k)}. In the scenario of a dual-axis gimbal, the main control processor can concatenate the angular accelerations of the horizontal and pitch axes into a two-dimensional vector sequence in a fixed order, which can be processed separately or used as a unified multi-dimensional input in the subsequent modeling stage.
[0077] The main control processor constructs a coupled perturbation hidden state model based on angular acceleration sequences to characterize the perturbation intensity of the gimbal's mechanical motion coupled to the audio link through the structural path. In this embodiment, the coupled perturbation between the gimbal and the audio link is abstracted as a hidden state quantity that changes slowly over time. The hidden state includes at least a coupling strength parameter to characterize the degree of coupling.
[0078] For example, the hidden state vector x(k) at discrete time k can be defined to contain at least a coupling strength component α(k), where α(k) characterizes the disturbance gain of the unit gimbal motion excitation transmitted to the microphone mounting position under the current attitude and load conditions. Considering that the gimbal motion state and structural contact conditions change slowly over time, the hidden state evolution can be modeled as a first-order Markov process, i.e., x(k+1) depends on x(k) and the process noise term w(k), where the process noise characterizes the unmodeled disturbance and modeling error. The main control processor can empirically model the hidden state evolution relationship based on the gimbal mechanism's stiffness, damping characteristics, and drive strategy. In simple implementations, a linear or piecewise linear model can be used, while in complex implementations, a nonlinear state transition model including nonlinear saturation and friction terms can be used.
[0079] In this embodiment, to ensure the hidden state model better reflects the actual relationship between gimbal mechanical disturbances and acoustic disturbances, the main control processor can collect a set of representative gimbal motion trajectories and corresponding audio disturbance data during the equipment factory calibration phase or the on-site self-calibration phase. Regression analysis or least squares methods are then used to estimate the hidden state evolution parameters, ensuring that the changing trend of the coupling strength component α(k) in the hidden state matches the actually observed changes in audio noise energy. The coupled disturbance hidden state model established in this way is no longer merely a threshold judgment of the magnitude of angular acceleration, but can comprehensively reflect the combined impact of gimbal mechanical disturbances on the audio link using lower-dimensional state variables.
[0080] The main control processor establishes a nonlinear observation model based on the coupled perturbation hidden state model, which is used to map the directly measurable angular acceleration sequence to the observation of the hidden state. In one embodiment, the main control processor uses the angular acceleration sequence and its statistical characteristics within a short time window as the observation input, such as calculating the root mean square value, peak value, rate of change, or other characteristic quantities reflecting the intensity of motion within a preset time window, and combining them to form the observation vector z(k). The nonlinear observation model is used to characterize the functional relationship between the hidden state vector x(k) and the observation vector z(k), which can be obtained through empirical modeling, physical modeling, or a combination of both.
[0081] For example, based on the dynamic relationship of the gimbal structure, the nonlinear function form between the coupling strength component α(k) and the angular acceleration characteristics can be derived. Alternatively, the observation function h(·) can be fitted using data-driven methods such as neural networks and kernel regression, so that, given the hidden state, the deviation between the observed estimate output by the model and the actual observed vector is as small as possible in a statistical sense.
[0082] In addition, in a simplified implementation, the observation model can be in the form of a polynomial with saturation terms, using a finite number of parameters to approximate the nonlinear relationship between angular acceleration characteristics and coupling strength, thereby reducing computational complexity. In another implementation, the observation model can be segmented according to different gimbal attitudes and rotation modes, using different parameter sets for different working intervals to improve the overall fitting accuracy.
[0083] Regardless of the form used, the main control processor uses the nonlinear observation model as the observation equation for subsequent particle filter estimation. This allows the corresponding observation prediction value to be calculated given the hidden state sample, and compared with the actual observation vector extracted from the angular acceleration sequence to form the observation residual, providing a basis for evaluating the confidence of each particle.
[0084] Optionally, based on the above embodiments, in order to further characterize the directional coupling characteristics of the gimbal motion to the audio link and enable the system to accurately estimate the disturbance intensity under different rotation directions and attitude combinations, in one optional embodiment, the coupled disturbance hidden state not only includes the coupling intensity parameter, but also further includes the coupling direction angle used to characterize the main coupling direction and the extended scale parameter used to characterize the degree of directional extension. By introducing elliptic parameterization constraints in the two-dimensional direction vector space, the ability to describe the anisotropy of the gimbal motion direction is improved.
[0085] Specifically, in this embodiment, a coupling direction angle component and an extended scale parameter component are added to the aforementioned hidden state vector x(k). The coupling direction angle is used to characterize the dominant direction of the gimbal motion on the audio disturbance at the current moment, and the extended scale parameter is used to characterize the extent of the disturbance distribution along the dominant direction and its orthogonal direction. In this embodiment, the horizontal rotation axis and pitch axis of the gimbal are used as the basis of the two-dimensional direction vector space. The main control processor assembles the angular acceleration components of the gimbal on these two axes into a two-dimensional angular acceleration vector, which is used as the basis for directional analysis.
[0086] In this two-dimensional direction vector space, the main control processor establishes the elliptic parametric equation with the coupling direction angle and the scaling parameter as independent variables. Specifically, the unit circle can be stretched in both the length and width directions according to different proportions of the scaling parameter to obtain an elliptical profile representing the disturbance distribution contour. A rotation operation is then performed to align the major axis of the ellipse with the coupling direction angle, and the minor axis with the orthogonal direction of the coupling direction angle. Thus, the major axis of the ellipse corresponds to the primary coupling direction caused by the gimbal motion, and the minor axis corresponds to the secondary coupling direction orthogonal to it. The scaling parameter controls the stretch ratio in these two directions, thereby reflecting the strength differences of the disturbance in different directions.
[0087] To determine the aforementioned ellipse parameters, the main control processor performs directional statistics on the angular acceleration sequence within a sliding time window. Specifically, within each time window, the two-dimensional angular acceleration vector is first projected onto the major and minor axes based on candidate coupling direction angles, resulting in projection sequences for the major and minor axes. Then, the variance or energy statistics of the two projection sequences are calculated, and their ratio is used as the basis for evaluating the rationality of the current directional hypothesis. In this embodiment, the direction angle that makes the variance ratio between the major and minor axes most significant can be found through traversal or optimization search. This direction angle is used as an estimate of the coupling direction angle. Simultaneously, the scale ratio and overall scaling factor of the ellipse are determined based on the variance ratio and the overall energy level to obtain the corresponding extended scale parameters.
[0088] In practical implementation, the main control processor can use least-squares fitting or other numerical optimization methods to fit the distribution of angular acceleration in the two-dimensional direction vector space into a set of elliptical parameters, such that the ellipse can statistically encompass most of the angular acceleration sample points. The major axis direction of the fitted ellipse serves as the estimate of the current coupling direction angle, while the ratio of the major to minor axis lengths and the absolute length correspond to the parts representing directional differences and the part representing the overall disturbance amplitude in the extended scale parameters, respectively. By periodically updating these elliptical parameters, the main control processor can track the impact of changes in the gimbal's motion mode on the perturbation directionality.
[0089] After obtaining the ellipse parameters, the main control processor generates a coupled perturbation hidden state model, including directional constraint terms, based on these parameters. Specifically, constraint terms related to the major and minor axes of the ellipse can be introduced into the hidden state evolution equations. This makes the coupling strength parameter more likely to increase under gimbal motion aligned with the major axis, while suppressing it under motion aligned with or significantly deviating from the minor axis. Alternatively, the range of variation of the coupling direction angle can be limited in the hidden state prior distribution, allowing it to evolve smoothly within adjacent time windows and avoiding unconstrained jumps. In the observation model, the main control processor can also preferentially utilize the projection features along the coupling direction angle and its orthogonal direction to construct the observation vector, enabling the particle filter to fully utilize the statistical information of angular acceleration in the main coupling direction during the update process.
[0090] In this way, by introducing coupling direction angle and extended scale parameters into the hidden state, the coupled perturbation hidden state model can not only describe the magnitude of the perturbation intensity but also the distribution characteristics of the perturbation in different directions. This helps to improve the estimation accuracy and stability of coupled perturbations under complex gimbal motion conditions, thereby providing more detailed directional information for determining the subsequent avoidance time window. Those skilled in the art can adjust the specific implementation of the dimension of the direction vector space and the elliptic parameters according to the number of degrees of freedom and structural layout of the gimbal; all of these are optional implementations of this application.
[0091] Furthermore, in order to enable the observation model to make fuller use of the energy distribution characteristics of the gimbal motion in the main coupling direction and its orthogonal direction, and to maintain good robustness in the presence of impact noise and ground potential abrupt changes, in one optional implementation, the process of establishing the nonlinear observation model may include steps such as directional projection, sub-band energy measurement, robust noise modeling, and dependent structure modeling.
[0092] In one specific embodiment, the main control processor determines a set of orthogonal bases in the two-dimensional direction vector space using the coupling direction angles in the aforementioned hidden state during each observation update cycle. Specifically, taking the coordinate system formed by the horizontal and pitch axes of the gimbal as a reference, the two-dimensional angular acceleration vector representing the current motion state of the gimbal is projected along the principal direction corresponding to the coupling direction angle and the secondary direction orthogonal to it, respectively, to obtain the first direction projection value and the second direction projection value. By repeating this projection operation at multiple consecutive sampling times, the first direction projection sequence and the second direction projection sequence can be formed respectively, thereby separating the components of the gimbal motion in the principal coupling direction and the orthogonal direction, which facilitates subsequent analysis of directional energy characteristics.
[0093] Considering the potential differences in disturbance duration and spectral range under various operating conditions, in this embodiment, the main control processor adaptively determines the time analysis window length and bandwidth according to the extended scale parameter in the hidden state. A larger extended scale parameter indicates a slower, narrower bandwidth and more time-frequency distribution of the disturbance. In this case, the main control processor can select a relatively longer time window and a narrower bandwidth to improve its ability to resolve narrowband vibrations or slowly varying disturbances. Conversely, a smaller extended scale parameter allows for a shorter time window and a wider bandwidth to enhance the response speed to rapidly changing or broadband disturbances. Based on the aforementioned adaptive time window and bandwidth, the main control processor performs filtering or time-frequency analysis on the first and second direction projection sequences, respectively, calculating multi-scale, bandwidth-consistent sub-band energy measurements within their respective analysis windows, thereby obtaining a set of observations characterizing the energy distribution in the primary and secondary directions.
[0094] In real-world monitoring environments, factors such as mechanical shocks during pan-tilt rotation, relay activation, motor startup, and poor grounding contact can easily introduce abrupt amplitude changes and heavy-tailed noise into angular acceleration characteristics and subband energy measurements. Observation models that only use the Gaussian noise assumption often fail to accurately reflect these anomalies. To improve the robustness of the observation model to impact noise and ground potential abrupt changes, in this embodiment, the main control processor establishes a robust noise model based on a skewed t-distribution to address the difference between the subband energy measurement and the energy value predicted from the hidden state.
[0095] Specifically, the difference between the observed and predicted quantities is regarded as the observation residual, and it is assumed that the residual follows a skewed t-distribution with adjustable skewness and heavy tails. By estimating the degrees of freedom, location, and scale of the skewed t-distribution based on historical residual samples during the calibration phase or operation, the noise model can better encompass the asymmetric shocks and intermittent large-amplitude disturbances that occur in the actual observation, thereby reducing the interference of abnormal samples on the latent state estimation during subsequent observation updates.
[0096] Since perturbations in the main coupling direction and orthogonal direction are often correlated, modeling the edge distributions of the two directions alone cannot reflect their dependency structure. Therefore, this embodiment further introduces a connection function to model the dependency structure of the sub-band energy edge distributions in both directions. Specifically, the main control processor estimates the respective edge distributions based on the sub-band energy measurements corresponding to the projection sequences of the first and second directions, respectively. An empirical distribution or a suitable parameterized distribution can be used for fitting. Based on this, a connection function is introduced as a function to characterize the dependency relationship between random variables given the edge distributions. By selecting or fitting an appropriate form of connection function, the constructed joint observation distribution can reflect the correlation characteristics of the primary and secondary directional energies under actual perturbation conditions, such as the cooperative behavior of simultaneous increases or complementary behavior of one high and one low under certain operating conditions.
[0097] After comprehensively considering the observation residuals generated by the skewed t-distribution noise model and the joint observation distribution constructed by the connection function, the main control processor determines the parameter set and observation mapping form of the nonlinear observation model through maximum likelihood, variational inference or other parameter estimation methods.
[0098] Specifically, the coupling strength, coupling direction angle, and extended scale parameters in the hidden state can be mapped to predicted values and distribution parameters of energy measurements in each sub-band, forming a nonlinear mapping relationship from the hidden state space to the observation space. This, combined with the joint probability density obtained from robust noise modeling and dependent structure modeling, forms a likelihood function for particle filter observation updates. During each particle update, the main control processor calculates the weight of each particle based on the observation residual between this joint observation distribution and the actual observations, and performs resampling. This allows for stable estimation of the posterior distribution of the hidden state even in complex environments with directional differences, impulse noise, and sudden changes in ground potential, providing reliable observational support for determining subsequent perturbation eigenvalues and avoidance time windows.
[0099] This improves the robustness and characterization of the observation model against impact noise and related disturbances, making the mapping between the gimbal motion state and audio disturbances more consistent with the statistical characteristics of actual engineering scenarios.
[0100] Regarding S102 above:
[0101] Specifically, in the initialization phase of step S102, the main control processor first generates a set of particles to approximate the hidden state distribution based on prior information obtained from factory calibration or the estimation results from the last stable operating time. Each particle carries a set of candidate values for the hidden state variables, including coupling strength and optional directional and scale parameters, and is assigned the same initial weight. In practical applications, the number of particles can be selected according to hardware computing power and real-time requirements. For example, on embedded platforms with limited resources, tens to one or two hundred particles can be selected. When computing resources are sufficient, the number of particles can be appropriately increased to improve estimation accuracy.
[0102] Within each audio processing cycle, the main control processor performs time updates on the particle ensemble based on the hidden state evolution relationship. Specifically, for particles with larger weights in the previous cycle, the main control processor predicts the hidden state changes at the current moment according to the coupled perturbation hidden state model. During the prediction process, process noise reflecting modeling errors and external perturbations is superimposed to ensure that the particles do not all shrink to a single point but expand within a reasonable range to cover possible changes in the hidden state. After this processing, the particle ensemble provides a set of prior candidate values for the hidden state at the current moment.
[0103] Subsequently, the main control processor uses the aforementioned nonlinear observation model and the real-time acquired angular acceleration observations to perform observation updates on the particle ensemble. For each particle, the main control processor substitutes the hidden state carried by the particle into the nonlinear observation model to obtain the corresponding observation prediction value, and then compares it with the actual observation value calculated based on the angular acceleration sequence at the current time to obtain the observation residual.
[0104] For example, the observation residual can be characterized by the difference between the predicted value and the actual observed value, or the sum of squares of the differences. The smaller the residual, the better the hidden state corresponding to the particle can explain the current observation, and the particle should be given a higher weight; the larger the residual, the lower the weight accordingly.
[0105] In this embodiment, to avoid excessive computational burden caused by numerically integrating all continuous variables in the hidden state one by one, the main control processor uses variational marginalization to process continuous variables while updating observations. Specifically, the continuous variables in the hidden state are divided into key variables that require fine characterization (e.g., coupling strength and its directly related scale) and auxiliary variables that can be approximated by statistical properties. For each particle, given the observation data and candidate values of the key variables, the main control processor iteratively adjusts a set of parameters used to describe the posterior distribution of the auxiliary continuous variables. For example, this set of parameters can correspond to the mean vector and correlation features of the auxiliary variables. The goal of the iteration is to make the statistical properties of the observation prediction calculated based on this set of parameters as close as possible to the statistical properties of the actual observations. In engineering implementation, the computational load of each variational update cycle can be controlled by limiting the maximum number of iterations or setting a threshold for parameter changes.
[0106] After variational marginalization, the master processor "absorbs" the auxiliary continuous variables that explicitly appear in the observation update into the observation likelihood in their variational distribution form, thereby obtaining the effective observation probability that depends only on the key hidden state variables and observation data, which is used to correct the particle weights. In this way, the weight corresponding to each particle no longer simply depends on a fixed set of continuous variable values, but comprehensively considers the overall contribution of the auxiliary continuous variables to the observation under their posterior distribution, which statistically improves the stability and robustness of the weight calculation. For particles with extremely small weights, the master processor can perform a resampling operation at appropriate intervals to eliminate them and replicate particles with larger weights to prevent particle degradation.
[0107] In this way, at the end of each sampling period, the main control processor obtains the particle representation of the posterior distribution of the hidden state, that is, a set of particles carrying the candidate values of the hidden state and their corresponding weights. This set of particles can be used to calculate various statistics of coupling strength, such as weighted average, variance, or quantiles, and can also serve as the starting point for the prediction of the hidden state in subsequent periods.
[0108] Optional, see Figure 2 The flowchart below illustrates a method for obtaining the posterior distribution of hidden states, as provided in an embodiment of this application, including steps S201 to S204, wherein:
[0109] S201: After detecting the silence trigger signal, read the instantaneous value of angular acceleration corresponding to the gimbal attitude, and calculate the acoustic spatial drift based on the instantaneous value of angular acceleration;
[0110] S202: Correct the prior mean of the continuous variables in the particle set according to the acoustic spatial drift amount, and limit the particle update range;
[0111] S203: During the detection period when the electrical stability indicator is valid, perform variational marginalization calculation on the continuous variable to generate a marginalized distribution of the continuous variable;
[0112] S204: When the silence trigger signal is detected to change from valid to invalid, the distribution of the continuous variable is used as the initial input for particle propagation, and the posterior distribution of the hidden state is updated.
[0113] In a preferred embodiment, in order to calibrate the hidden state of the coupled disturbance using the audio silence period and make the posterior distribution of the hidden state closer to the current real working condition, the process of obtaining the posterior distribution of the hidden state may include the following.
[0114] Upon detecting a silence trigger signal, the main control processor first reads the angle information corresponding to the current gimbal attitude and the instantaneous value of angular acceleration at that moment. The gimbal attitude here can include two components: horizontal rotation angle and pitch angle. To obtain the acoustic spatial drift, the main control processor can pre-establish the correspondence between gimbal attitude and sound field distribution during the device's factory manufacturing or installation and commissioning phases. For example, test sound sources can be played at multiple typical attitude angles, and the frequency response gain or main direction of arrival at the microphone can be recorded. This data can then be compiled into a calibration table. During operation, the main control processor uses the current attitude to look up or interpolate the gain change or main direction of arrival offset relative to the reference attitude in the calibration table, and defines this change as the acoustic spatial drift, used to characterize the degree of microphone offset from the target sound field.
[0115] The main control processor corrects the prior mean of continuous variables in the particle set based on the acoustic spatial drift and limits the particle update range.
[0116] Specifically, if the acoustic space drift is small, it indicates that the current posture is close to the calibration reference posture. The main control processor can align the prior mean of the coupling strength with the calibration value under the reference posture and tighten the particle value range to a small interval near the calibration value, for example, limiting it to a certain preset percentage range above and below the reference value. If the acoustic space drift is large, it indicates that the current posture differs significantly from the reference posture. The main control processor can appropriately widen the particle value range, but still focus on the trend of coupling strength change obtained from calibration to avoid particles spreading to regions that are significantly inconsistent with the current posture. In this way, the particle search space is guided to a more likely hidden state region during the silent period, improving the convergence speed and robustness of subsequent estimations.
[0117] During detection periods where the silence trigger signal is valid and the electrical stability flag is valid, the system considers the current time period's speech signal and electrical disturbances to be at low levels. At this time, the observed data is dominated by background noise and structural perturbations, making it suitable as a calibration window for hidden state continuous variables. During these detection periods, the main control processor performs variational marginalization calculations based on the current particle set and observed data.
[0118] For example, the posterior distribution of auxiliary continuous variables other than coupling strength can be approximated as a type of parameterized distribution, such as a one-dimensional distribution or a low-dimensional joint distribution with mean and variance as parameters. These parameters are then slightly updated based on the observation residuals in each silent detection period, so that the predicted noise statistics calculated based on this distribution gradually approach the actual observed residual noise statistics.
[0119] To control computational load, the main control processor can limit the number of iterations within each detection cycle, or terminate the update early when parameter changes fall below a preset threshold. After several silent detection cycles, a set of continuous variable distributions that converge under the current pose and environmental conditions is obtained, i.e., the marginalized continuous variable distribution.
[0120] When the silence trigger signal is detected to change from valid to invalid, that is, the silence calibration phase ends and the system is about to enter the voice activity or switching sensitive phase, the main control processor uses the above-mentioned marginalized continuous variable distribution as the initial input for particle propagation.
[0121] One approach is to regenerate a new set of particles based on the mean and dispersion of the distribution, for example, by randomly sampling continuous variable values within a certain range around the mean, while assigning the same initial weight to each particle.
[0122] Another approach is to directly adjust the center position of continuous variables in the existing particles using the mean of the distribution, and adjust the spacing between particles using the dispersion of the distribution. After this processing, the posterior distribution of the hidden state corresponding to the particle set has been updated in conjunction with the calibration information during the silent period, serving as the starting state for coupling perturbation estimation in subsequent processing cycles. This allows the system to continue tracking the perturbation of the audio link by the gimbal motion based on a more accurate and stable hidden state estimation after a silent calibration.
[0123] Optionally, the generation of the marginalized continuous variable distribution includes:
[0124] Based on the acoustic spatial drift and gimbal attitude, a preset acoustic coupling calibration table is invoked to obtain the prior interval of coupling strength and the coupling noise baseline under the corresponding attitude.
[0125] Under the constraint of the prior interval of the coupling strength, a variational distribution family for characterizing the continuous variable is selected, and the mean vector and covariance matrix of the continuous variable are iteratively updated within the variational distribution family;
[0126] In each iteration update, the residual noise power measured during the effective period of the silent trigger signal is compared with the coupling noise power predicted by the variational distribution. When the deviation between the two exceeds a first preset threshold, the corresponding particle weight is reduced and the prior interval of the coupling strength is tightened.
[0127] When the variational optimization convergence condition is met, the final variational distribution is taken as the marginalized continuous variable distribution.
[0128] To make the distribution of continuous variables after marginalization closer to the actual engineering scenario, when the silent trigger signal is valid, the main control processor first reads the coupling parameters under the corresponding attitude from the preset acoustic coupling calibration table based on the current gimbal horizontal angle, pitch angle and acoustic space drift.
[0129] The calibration table can be established during the factory or installation and commissioning phase. For example, the gimbal can be divided into several attitude grids with a horizontal range of 30° and a pitch range of 15°. At each attitude point, the audio source is turned off, and only ambient noise and the gimbal are kept in standby mode. Audio data is continuously collected for no less than 10 seconds. The mean and statistical dispersion of the background noise power are calculated for the audio data of each attitude, and then converted into the typical value and allowable fluctuation range of the coupling strength under that attitude through the structural model. The main control processor can set this range as the a priori interval of coupling strength, for example, with the typical value as the center and extending several decibels above and below it to correspond to the coupling strength gain interval, while using the mean noise power under the corresponding attitude as the coupling noise baseline.
[0130] After obtaining the prior interval of coupling strength, the main control processor selects a variational distribution family to characterize the continuous variable under the constraints of this interval. In one implementation, the continuous variable includes at least the coupling strength and a scale parameter related to the structural path or frequency band. The main control processor can approximate this with a bivariate distribution using the mean vector and covariance matrix as parameters, initially setting the mean at the middle of the prior interval and initializing the covariance to cover most of the prior interval, for example, ensuring that 95% of the probability mass falls between the calibrated upper and lower limits. In this way, the initial variational distribution is both constrained by the calibrated prior and retains some adjustment space.
[0131] During the detection period when the silence trigger signal remains valid, the main control processor iteratively updates the aforementioned variational distribution parameters. In each iteration, the main control processor extracts several sets of continuous variable samples based on the current variational distribution, estimates the corresponding coupled noise power through a structural coupling model, averages these estimates over a time window to obtain the predicted coupled noise power, and simultaneously calculates the residual noise power based on the actual acquired audio signal within the same silence window. The main control processor compares the differences between the two. For example, it can use the ratio of the difference between the residual noise power and the predicted noise power to the predicted value as a deviation measure. When this ratio is lower than a preset first threshold (e.g., 20%), the current variational distribution is considered reasonable, and only the mean and covariance are slightly adjusted to make the predicted noise power gradually approach the measured value. When the deviation measure exceeds the threshold, it indicates that the values of some continuous variables do not match the actual coupling level under the current attitude. On the one hand, the main control processor reduces the particle weights of those samples that correspond to significantly higher or lower predicted noise, for example, by attenuating them by a certain proportion to reduce their influence in the posterior estimation. On the other hand, it appropriately tightens the prior interval of coupling strength, for example, by reducing the interval width by a certain contraction coefficient and shifting it to the vicinity of the coupling strength corresponding to the measured noise power, so that subsequent iterations are more concentrated in a reasonable range.
[0132] With multiple silent detection cycles, if the deviation between the predicted value of coupled noise and the residual noise power is lower than the first threshold in several consecutive iterations, and the changes in the mean vector and covariance matrix of the continuous variables are lower than the preset convergence threshold (e.g., the relative change is within a few percentage points), then the main control processor considers the variational optimization to have converged. At this time, the final variational distribution is used as the marginalized distribution of the continuous variables.
[0133] In this way, the distribution reflects both the prior interval of coupling strength given in the calibration phase and the statistical characteristics of residual noise in the current silent window. It can serve as the basis for subsequent particle initialization and hidden state propagation, enabling the system to estimate coupling perturbations in the new round of operation with a continuous variable starting point that is more in line with the current attitude and environment, thereby improving the accuracy and stability of the avoidance time window calculation.
[0134] Regarding the above S103:
[0135] In one embodiment, after obtaining the posterior distribution of the hidden state at the current moment, the main control processor performs forward prediction of the gimbal motion disturbance within a preset prediction window, and generates an avoidance time window and an electrical stability flag by combining the changes in ground potential and power amplifier supply voltage.
[0136] In one specific embodiment, the preset prediction window can be set to a time length of tens to hundreds of milliseconds, for example, it can be selected to cover the time range of several subsequent audio buffer frames. Starting from the posterior distribution represented by the current particle set, the main control processor, without introducing new observations, propagates forward for each particle step by step within the prediction window according to the aforementioned hidden state evolution relationship, thereby obtaining the hidden state prediction distribution at each prediction time.
[0137] For each prediction time, the main control processor extracts the corresponding coupling strength sample from the particle set and calculates the statistics of the coupling strength according to the particle weights, such as the weighted average, weighted variance, and high quantile value.
[0138] To summarize the above statistical results into disturbance characteristic values for judgment, in this embodiment, the main control processor can combine the weighted average of the coupling strength with statistics representing uncertainty to characterize the overall risk level of the gimbal disturbance at the predicted time. For example, when the weighted average of the coupling strength is high and the high quantile value is significantly higher than the normal range obtained during the calibration phase, the disturbance risk corresponding to the predicted time can be considered to be relatively high.
[0139] The main control processor concatenates the disturbance characteristic values at each prediction moment within the prediction window to form a time series, and compares it with a preset disturbance threshold. When there are time intervals where the disturbance characteristic values are consistently higher than the threshold, these time intervals are marked as avoidance time windows for gimbal movement. If necessary, the main control processor can also add a certain safety margin before and after these intervals to cope with modeling errors and actual response lags, thus obtaining the final set of avoidance time windows used for switching control.
[0140] On the electrical side, to identify transient disturbances in the power amplifier and its power supply path, in this embodiment, the main control processor acquires ground potential change signals and power amplifier supply voltage through a detection circuit connected to a ground reference point and the power amplifier power supply terminal. The detection circuit may include sampling resistors, filtering networks, etc., connected to the main control processor's built-in analog-to-digital converter module. The main control processor samples the signals at a sampling period coordinated with audio processing or an integer multiple thereof. Within each detection period, the main control processor calculates the rate of change of the power amplifier supply voltage over a short time, for example, by comparing the average supply voltage in the current detection period with the average value of the previous detection period to indicate the speed of voltage change. Simultaneously, the amplitude of the ground potential change signal is evaluated; ground potential fluctuations can be reflected by calculating the offset relative to the system reference ground or the energy level within a certain frequency band.
[0141] In one implementation, the main control processor sets thresholds for the voltage change rate and ground potential fluctuation. For example, when the power amplifier supply voltage change rate is lower than a predetermined threshold for multiple consecutive detection cycles, and the ground potential deviation amplitude is within the normal operating range or lower than a preset disturbance threshold, the current electrical environment is considered stable, and the main control processor sets the electrical stability flag to valid. When a sudden increase in the supply voltage change rate is detected in a short period of time, or when there is a significant jump or abnormal fluctuation in the ground potential, the main control processor sets the electrical stability flag to invalid and maintains this state for a set buffer time to avoid performing audio switching before the electrical transient has completely disappeared.
[0142] This allows subsequent switching control to simultaneously avoid periods of high mechanical disturbance and periods of unstable electrical environment, providing reliable timing constraints for the output of switching control signals in subsequent steps.
[0143] Optionally, the generated perturbation feature values include:
[0144] During the assembly stage of the surveillance camera, the low-order resonant frequency band of the camera housing and mounting bracket, as well as the unit impact response function related to the motion coupling of the pan-tilt unit, are obtained in advance through simulation analysis, and the low-order resonant frequency band and the unit impact response function are stored in the structural mode library.
[0145] Within the preset prediction time window, a time series estimate of coupling strength is generated based on the posterior distribution, and the time series estimate of coupling strength is convolved with the unit impulse response function to obtain the structural vibration response sequence at the microphone installation location.
[0146] Based on at least one of the peak amplitude or root mean square value of the structural vibration response sequence within the preset prediction time window, a statistical quantity characterizing the amplitude of structural vibration is obtained, and the statistical quantity is used as the disturbance characteristic value.
[0147] In order to make the disturbance feature values more closely reflect the actual structural vibration of the camera, in an optional implementation, the process of generating disturbance feature values can be predicted by combining structural modal analysis and unit impact response.
[0148] During the assembly phase of a surveillance camera, the main control processor or upper-level engineering tools can obtain the low-order resonance characteristics of the camera housing and mounting bracket through finite element simulation and / or experimental modal testing.
[0149] Specifically, a structural model including the gimbal base, rotary joint, camera housing, and mounting bracket can be established. Pulse or frequency sweep excitations are applied to the gimbal in typical working postures to analyze the main resonant frequency bands and corresponding mode shapes within the tens to hundreds of hertz range. Alternatively, short-term mechanical impacts can be applied to the gimbal drive motor or structural connection points on the prototype using an accelerometer or laser vibrometer, recording the acceleration or displacement response near the microphone mounting position to obtain the unit impact response function related to the gimbal's motion coupling. The low-order resonant frequency bands (e.g., narrow bands near a certain bending mode of the housing) and the unit impact response curves obtained from the analysis are organized and stored in a structural mode library, and associated with the specific camera model, mounting method, and microphone position.
[0150] During camera operation, after obtaining the posterior distribution of the hidden state at the current moment through step S102, the main control processor performs forward prediction of the coupling strength within a preset prediction time window, forming a temporal estimate of the coupling strength. This temporal estimate can be understood as the "gain sequence" of the unit gimbal motion excitation coupled to the microphone position through the structure over a future period. The main control processor retrieves a unit impulse response function from the structural modal library that matches the current installation configuration and gimbal attitude, and performs a convolution operation between the temporal estimate of the coupling strength and the unit impulse response function to obtain the structural vibration response sequence at the microphone installation position.
[0151] In this way, without the need for additional sensors, structural vibrations near the microphone can be predicted proactively based on the gimbal's motion state and structural modal characteristics.
[0152] In order to transform the above structural vibration response sequence into disturbance feature values that can be used for comparison and threshold determination, in this embodiment, the main control processor performs statistical calculations on the structural vibration response within a preset prediction time window.
[0153] For example, the peak amplitude of vibration displacement or acceleration can be calculated throughout the entire prediction window to characterize the structural vibration intensity at the most unfavorable moment; alternatively, the root mean square value of the response within the prediction window can be calculated to characterize the average vibration energy over a period of time. In some scenarios, the system can simultaneously employ both peak and root mean square statistics, using weighted or optimized methods to construct disturbance characteristic values, thus taking into account both instantaneous impact and overall vibration level.
[0154] In practical applications, the system can set a set of reference ranges for the above statistics at the factory or during the commissioning phase, based on the structural vibration level and the user's acceptable sound quality indicators.
[0155] For example, when the predicted peak statistic is significantly close to or exceeds the empirical upper limit of the corresponding resonant mode in the structural mode library, the main control processor can assess the predicted segment as having a high risk of structural vibration and map the corresponding statistic to a larger disturbance characteristic value; when both the peak value and root mean square are at a low level, the disturbance characteristic value is correspondingly smaller. In the subsequent determination of the avoidance time window, the main control processor uses these disturbance characteristic values as the basis for judgment, marking the time period when the characteristic value exceeds the preset vibration threshold as the interval where audio switching should be avoided. This ensures that the avoidance time window not only considers the change in the gimbal angular acceleration itself, but also comprehensively considers the impact of structural resonance and actual casing vibration on the microphone.
[0156] Based on the above implementation method of generating structural vibration response sequences by combining structural modal libraries, in order to more specifically identify moments approaching structural resonance, in an optional embodiment, the statistical quantity characterizing the structural vibration amplitude can be further refined into the "resonance band energy ratio" around the candidate switching moment, and used as a structural resonance risk indicator to mark the avoidance time window of the gimbal movement.
[0157] Specifically, within a preset prediction window, the main control processor determines several candidate switching moments based on the time granularity of the system design. For example, the prediction window can be divided into several time steps that are equal to or slightly smaller than the length of the audio buffer frame, with the center sampling point of each time step serving as a candidate switching moment; alternatively, the frame alignment position and zero-crossing position on the audio side can be combined to select prediction moments aligned with these positions as candidate switching moments.
[0158] For each candidate switching moment, the main control processor symmetrically selects a local time window before and after it, for example, a length of several milliseconds to tens of milliseconds, to extract a local segment of the structural vibration response sequence.
[0159] Within each local time window, the main control processor performs frequency band decomposition on the structural vibration response sequence. In specific implementation, a set of digital bandpass filters or filter banks can be used to perform multi-subband decomposition on the response signal, which includes at least one or more narrowband filters covering the low-order resonant frequency band and a broadband filter covering the entire operating frequency band.
[0160] The center frequency and bandwidth of the low-order resonant band can be determined through modal analysis during the assembly stage. For example, a narrowband filter can be set in the 100Hz–150Hz frequency band for a specific bending mode. The overall operating frequency band can cover a range from tens of hertz to thousands of hertz. The main control processor calculates the sub-band energy falling into the low-order resonant band within the local time window, as well as the total energy across all sub-bands. The ratio of the low-order resonant band energy to the total energy is used as an indicator of structural resonance risk at the candidate switching moment.
[0161] In addition, for cases with multiple low-order modes, the energy of each resonant frequency band can be accumulated and then compared, or multiple energy ratios can be calculated separately and combined into a comprehensive risk index according to preset weights.
[0162] To map the structural resonance risk index into a criterion for judging disturbance characteristics, in this embodiment, the main control processor sets a preset resonance threshold for the energy ratio. For example, when the energy ratio is significantly greater than a certain benchmark value (such as the energy proportion of the resonance frequency band exceeding half of the total energy, the specific value of which can be determined during calibration and debugging), it indicates that the structural vibration within the current local time window is mainly concentrated on the low-order resonance frequency band, and the switching operation at this moment is more likely to trigger mechanical noise and popping sounds caused by the amplification of structural resonance.
[0163] After calculating the structural resonance risk index for each candidate switching moment, the main control processor marks the moment that exceeds the preset resonance threshold and the time period corresponding to the local time window as the avoidance time window for the gimbal movement. In the case where multiple consecutive candidate moments are higher than the threshold, these local time windows can be merged into a continuous avoidance interval, and a certain time margin can be added before and after the interval to improve the protection against resonance effects.
[0164] In this way, the statistical quantity obtained to characterize the amplitude of structural vibration is no longer just the overall vibration peak value or root mean square value, but further introduces the structural resonance risk index of "low-order resonance frequency band energy ratio", which enables the system to perform more precise avoidance control for the sensitive period of structural resonance, avoid performing audio switching when the structural vibration energy is concentrated in the resonance frequency band, and thus further reduce the switching noise risk in complex mechanical vibration environment.
[0165] Regarding S104 and S105 above:
[0166] In one embodiment, in order to identify a suitable switching opportunity on the audio side, the present application detects the frame structure and waveform morphology of the audio buffer in step S104, generates an audio silence flag accordingly, and further forms a silence trigger signal that can be immediately used for filter calibration and switching control.
[0167] Specifically, the raw audio data output by the audio acquisition unit is written to the audio buffer with a fixed frame length. The frame length can be configured from a few milliseconds to tens of milliseconds, depending on the system design, for example, using a sampling length of 256 points or 512 points. Each time a new frame is written to the buffer, the main control processor or audio driver module marks the start position of the frame, forming a frame alignment flag to indicate whether the current time point is at the start boundary of the audio frame. In actual implementation, the frame alignment flag can be generated by the audio codec's interrupt signal, DMA transfer completion flag, or a software counter, as long as it can reliably distinguish between the "frame start time" and the "frame middle time".
[0168] When a frame is detected to be aligned at a certain time point, the main control processor reads a certain number of audio sampling points in the starting neighborhood of that frame, such as reading several samples before and after the frame, to obtain a short waveform segment. On one hand, the main control processor calculates the local energy or root mean square value of this waveform segment and compares it with a preset silence energy threshold. When the local energy is lower than the threshold, it is considered that the speech or other significant sound source components in the current frame are weak. On the other hand, the main control processor detects the zero-crossing state of the current audio waveform, such as counting the number of waveform symbol changes in the starting neighborhood, or determining whether the waveform crosses from a positive value to a negative value or from a negative value to a positive value near the candidate switching time. When the local energy is low and the zero-crossing state meets preset conditions, such as at least one zero-crossing in a short period of time, or the waveform amplitude fluctuates near zero, the frame is marked as a silent candidate frame.
[0169] The main control processor summarizes the detection results of consecutive frames on the timeline. When several consecutive silent candidate frames are detected within a preset time window, and no significant energy surge or zero-crossing anomaly occurs during this period, the audio state of the corresponding time period is confirmed as silent, and the audio silence flag is set. The silence flag can be designed to remain valid during the silence period and be cleared when a new frame is detected that no longer meets the silence conditions. At the same time, the main control processor generates a silence trigger signal at the moment when the silence flag switches from invalid to valid. This signal serves as an edge trigger event for the start of the silence interval, driving the aforementioned hidden state filtering calibration process and timing alignment in subsequent decision logic. This ensures that silence calibration and switching control are coordinated with the buffer frame boundaries and waveform zero-crossing positions.
[0170] In step S105, to avoid performing audio channel switching during times of significant gimbal mechanical disturbance, unstable electrical environment, or non-silent audio, this application combines multiple judgment conditions to control the output of the switching command. Specifically, the main control processor maintains the correspondence between the current system time and a set of avoidance time windows in each scheduling cycle. When the system time is within any avoidance time window, it is considered that the risk of gimbal mechanical disturbance is high, and the switching operation is not allowed. At the same time, the main control processor checks the status of the electrical stability flag. If the electrical stability flag is invalid, it indicates that the recent power amplifier supply voltage change rate or ground potential fluctuation has exceeded the normal range, and switching is also prohibited. Only when the current time point is not within any avoidance time window, the electrical stability flag is valid, and the audio silence flag is valid, does the main control processor consider that the current time point simultaneously meets the comprehensive conditions of "low mechanical disturbance, stable electrical environment, and audio silence".
[0171] If all three conditions are met, the main control processor outputs a switching control signal at the nearest frame header boundary to control the audio channel selection switch, mixing matrix, or digital routing module to perform channel switching, or to update the audio source binding relationship associated with a certain video feed in the monitoring system. In practical engineering, the switching control signal can be a control level for an analog switch chip, a configuration command for a digital audio interface, or a control message for a host network video recording device. In this embodiment, if any condition is not met, the main control processor keeps the current audio channel unchanged, thereby avoiding triggering switching at unfavorable times and minimizing the impact of mechanical vibration, electrical disturbances, and audio waveform abrupt changes on monitoring sound quality and system stability.
[0172] Optionally, in order to reliably identify the frame start position even in the presence of clock jitter or buffer queue offset, and to improve the reliability of silence determination, during the deployment phase of the surveillance camera, the main control processor collects audio buffer data to characterize the audio buffer structure under the condition of system no load or only stable background noise.
[0173] Specifically, with a fixed sampling rate (e.g., 16kHz, 32kHz) and a stable frame length configuration (e.g., 256 or 512 points per frame), the audio acquisition and buffer writing logic can be continuously run to record the starting position indices of multiple audio frames over a period of time. The main control processor statistically analyzes the distribution of these frame starting positions on the sampling count axis, for example, by calculating the value of the buffer write pointer at each interrupt or DMA completion, and the offset of that value relative to the internal time base or system clock.
[0174] By statistically analyzing a large number of samples, a reference range for the frame start position can be obtained. For example, it can be found that the vast majority of frame start positions are concentrated around a certain count value, and the jitter range does not exceed a certain number of sampling points. The main control processor stores this reference range and the allowable offset width as a frame header template, which is used to characterize the approximate position range of the audio buffer frame header under normal operating conditions of this type of camera.
[0175] During operation, when a new audio frame is written to the buffer, the main control processor compares it with the reference range in the frame alignment template based on the time information and sampling position of the audio frame. If the sampling index or timestamp at the beginning of the current frame falls within the allowable offset range given by the template, the frame is considered an "aligned frame," and the frame alignment flag is set to valid. If the starting position deviates too much from the template, the frame is considered to have a boundary misalignment due to abnormal interruption, buffer squeezing, or other reasons, and the frame alignment flag is set to invalid, thereby avoiding silent judgment and switching operations on these unreliable frame boundaries.
[0176] For audio frames with valid frame alignment, the main control processor extracts the zero-crossing and energy features of the current audio waveform within its initial neighborhood. For example, a short window can be formed by taking several sampling points before and after the start of the frame. Within this window, the local energy or root mean square value is calculated, and zero-crossing related features such as the number of waveform symbol changes and symbol duration are statistically analyzed. When the local energy is below the silence energy threshold, and the zero-crossing features fall within a preset silence interval (e.g., a moderate number of zero-crossings, and the waveform oscillates slightly near zero without significant unilateral shift), the main control processor marks the frame as a silent candidate frame. If the energy increases significantly or the zero-crossing features are abnormal, such as maintaining the same symbol for a long time or having a significant non-zero shift, the frame is not considered a silent candidate frame.
[0177] The main control processor performs statistical analysis on the distribution of silent candidate frames along the timeline. Specifically, within a preset time window, such as covering the duration of several to dozens of audio frames, it counts the proportion of silent candidate frames and whether the candidate frames appear consecutively in time. When the number of silent candidate frames within this time window reaches a preset proportion, such as exceeding half or reaching a certain absolute frame number, and the interval between these candidate frames does not exceed the set maximum interval, the main control processor determines the audio state of the corresponding time period as silent and sets the audio silence flag to valid within that time period. Conversely, if the number of silent candidate frames is insufficient, the distribution is too sparse, or the audio silence is frequently interrupted by high-energy frames, the audio silence flag remains invalid.
[0178] In this way, the silent interval can still be reliably identified even when there is slight jitter in the buffer structure, making the timing of subsequent silent trigger signals and switching control more stable and controllable.
[0179] Based on the above implementation of generating audio silence markers based on frame alignment templates and zero crossover features, in order to further adapt to the changes in background noise patterns under different gimbal postures, in an optional embodiment, the process of marking audio frames as silence candidate frames can be combined with posture-related background noise templates.
[0180] During the deployment phase of the surveillance cameras, the main control processor, under relatively stable on-site conditions and without voice input, collects background audio buffer data for several typical pan-tilt-zoom (PTZ) postures. Specifically, the horizontal and vertical angles of the PTZ can be divided into several posture intervals with a certain step size, such as 30° for horizontal and 15° for vertical. At each posture position, the PTZ remains stationary, continuously collecting audio data for several seconds to tens of seconds, and buffering and framing it according to the same frame length and sampling rate as during operation. For the silent audio frames in each posture, the main control processor extracts zero-crossing features and amplitude features in the frame's starting neighborhood. For example, it counts the number of zero-crossings within the window, the distribution of adjacent zero-crossing intervals, the peak and root mean square values of the waveform envelope, and the short-time energy distribution. These features are then statistically analyzed over time to obtain statistical descriptions such as the mean, fluctuation range, or histogram of each feature. Based on the above statistical results, the main control processor generates a corresponding background noise template for each posture interval, recording the zero-crossing mode and amplitude mode of the typical silent background in that posture. Subsequently, the association between the gimbal attitude range and the corresponding background noise template is stored in the silent judgment library so that it can be called according to attitude during operation.
[0181] When the surveillance camera is operating normally, the main control processor acquires the current pan-tilt attitude information in real time, places it within a pre-defined attitude range, and retrieves the background noise template corresponding to that attitude from the silence determination library. For audio frames that are determined to have a valid frame alignment flag in step S104, the main control processor also extracts zero-crossing features and amplitude features in its initial neighborhood to form the feature vector of the current frame. Then, the main control processor compares this feature vector with the typical silence features recorded in the background noise template to calculate a deviation metric. For example, the difference between each feature of the current frame and the mean of the template features can be calculated separately, and the ratio of the difference to the allowable fluctuation range of the template can be normalized. Then, the deviations of each feature can be weighted and summed according to preset weights to obtain a comprehensive deviation index. The smaller the comprehensive deviation, the higher the feature matching degree between the current frame and the background silence noise under that attitude.
[0182] In one implementation, the main control processor can set a preset silence threshold condition for the aforementioned deviation index based on the on-site debugging results. For example, when the overall deviation is lower than a certain set value and the local energy does not significantly exceed the background noise baseline, the temporal shape of the current frame is considered sufficiently close to the background silence template under that posture, and the audio frame is then identified as a silence candidate frame. If the deviation significantly exceeds the threshold, it is considered that the current frame may contain speech or other abnormal sound sources, and it is not included in the statistics as a silence candidate frame. By introducing the posture-related background noise template into the silence candidate frame determination, the system can maintain the reliability of silence recognition even when the gimbal turns to different directions and the background noise spectrum changes with the direction.
[0183] For example, see Figure 3 This application provides a schematic diagram of an audio output path and channel switching structure for a surveillance camera.
[0184] The surveillance camera in this embodiment includes a main control SoC, an audio power amplifier chip circuit, a switching chip, an external speaker, and a camera speaker. The main control SoC outputs an audio signal to the audio power amplifier chip circuit through the audio signal output terminal, and controls the conduction path of the switching chip through GPIO to select the target output channel between the external speaker and the camera speaker. The main control SoC also obtains the insertion and removal status of the external speaker through the ADC insertion detection channel. The switching control signal generated in step S105 can be specifically manifested as the control level of the GPIO, thereby realizing the automatic switching of the audio channel.
[0185] For example, see Figure 5 This is a schematic diagram of an audio amplifier circuit and speaker connection structure provided in an embodiment of this application. The audio amplifier circuit in the surveillance camera includes a DC blocking capacitor C36, an input resistor R109, a feedback resistor R108, an audio amplifier chip U15, and speakers G1 and G2 connected to interface J12. The audio output terminal AC_HPOUT of the main control SoC is connected to the input terminal of the audio amplifier chip U15 through the DC blocking capacitor C36 and the input resistor R109 in series. U15 can be a mono audio amplifier chip, such as MS8002D. The output terminal of U15 is connected to the socket J12 through a differential capacitor C180, voltage stabilizing and filtering capacitors C177, C459, and C178, and protective TVS diodes D17 and D18. The corresponding pins G1 and G2 are further connected to the camera speaker or an external speaker column.
[0186] SPEAK_EN is the power amplifier enable control signal output by the main control SoC. When the main control SoC does not detect the insertion of an external speaker column and needs to play audio through the camera speaker, it sets SPEAK_EN to a high level, putting U15 into operation. At this time, the audio signal, after having its DC component removed by the DC blocking capacitor C36, is applied to the input terminal of U15. After being amplified by the power amplifier chip, it generates an output voltage at the J12 terminal sufficient to drive the speaker, and the speaker plays the corresponding sound. In another embodiment, SPEAK_EN can also be used as one of the specific forms of the switching control signal in step S105. Under the condition that the gimbal avoidance time window, electrical stability flag, and audio silence flag are all valid, the main control SoC controls the conduction state of SPEAK_EN and the upstream switching chip, thereby switching the audio channel between the camera speaker and the external speaker column.
[0187] For example, see Figure 6 This is a schematic diagram of an audio channel selection switch circuit provided in an embodiment of this application.
[0188] The audio channel selection circuit of the surveillance camera includes a switching chip U45. The switching chip U45 can be a dual-channel analog switching device, such as the BCT4157, with two audio input terminals B0 and B1, one output terminal A (exposed as HPOUT in this embodiment), as well as a selection control terminal SEL and a power supply terminal VCC. The main control SoC outputs audio signals for different purposes on the AC_HPOUT and LINE_HPOUT pins, respectively, and these two signals are connected to terminals B0 and B1 of U45. Output terminal A of U45 is connected to the camera speaker or an external speaker column through a subsequent power amplifier circuit, thereby realizing the switching of the audio path.
[0189] HPOUT_SEL is the GPIO control signal output by the main control SoC, connected to the SEL pin of U45, used to control the conduction path of the switching chip. When the main control SoC does not detect the insertion of an external speaker and needs to play audio through the camera's speaker, HPOUT_SEL is set to the first level (e.g., low level). At this time, U45 connects the AC_HPOUT channel to the HPOUT terminal, and the audio signal corresponding to AC_HPOUT is driven by the power amplifier circuit to play through the camera speaker. When the main control SoC detects that the external speaker has been inserted through the aforementioned insertion detection path, HPOUT_SEL is switched to the second level (e.g., high level). U45 connects the LINE_HPOUT channel to the HPOUT terminal, and the corresponding audio signal is sent to the external speaker for playback, thereby realizing automatic switching between the built-in speaker and the external speaker.
[0190] In the overall scheme of this application, the switching control signal in step S105 can be specifically manifested as an update action of the HPOUT_SEL control level: only when the current time point is not within the avoidance time window of the gimbal movement, and the electrical stability flag and the audio mute flag are valid, will the main control SoC allow the change of the HPOUT_SEL state, thereby updating the conduction path of the switching chip U45. By applying the "switching control signal" at the algorithm level to the channel selection circuit shown in the figure, this embodiment provides a specific hardware implementation of the audio automatic recognition and switching method of this application.
[0191] For example, Figure 7 This is a schematic diagram of an external speaker insertion detection circuit provided in an embodiment of this application. The surveillance camera also includes a detection circuit for detecting the insertion status of the external speaker. J15 is the external speaker interface, with terminals G1 and G2 connected to both ends of the external speaker, respectively; D204 is a surge protector used for overvoltage clamping protection of the external speaker line during lightning strikes or surges; C405 is a voltage-stabilizing filter capacitor used to smooth the voltage at the insertion detection node. R35 and R36, together with the internal resistance of the external speaker, form a voltage divider network. AUDIO_ADC is the analog-to-digital conversion input channel of the main control SoC, used to sample the voltage at this voltage divider node.
[0192] During deployment or debugging, typical voltage values for both speaker insertion and non-speaker insertion states can be determined based on circuit parameters. For example, in one specific implementation, when no external speaker is connected, the voltage divider network consists only of R35 and R36, and the steady-state voltage sampled by the AUDIO_ADC is approximately 1V. When an external speaker is inserted, its equivalent internal resistance is incorporated into the voltage divider network, changing the total resistance relationship and causing the voltage at the voltage divider node to drop to approximately 0.4V. The main control SoC periodically reads the voltage value of the AUDIO_ADC. When the voltage is detected to be consistently stable near the first level, it is determined that no external speaker is currently inserted; when the voltage is detected to be stable near the second level, it is determined that the external speaker has been inserted.
[0193] In the implementation of the method of this application, the main control SoC can use the above-mentioned insertion detection result as one of the prerequisites for switching control: for example, when it is determined that the external speaker has been inserted and the gimbal avoidance time window, electrical stability flag and audio silence flag are all valid, the audio output path is switched to the external speaker by controlling the selection signal of the aforementioned switch chip; when no external speaker insertion is detected, the output is maintained or switched to the camera body speaker output. In this way, the insertion detection circuit provides hardware support for the existence and connectivity of the external speaker in the automatic audio recognition and switching method of this application.
[0194] Based on the same inventive concept, this application also provides an audio automatic recognition and switching system for a surveillance camera, corresponding to an audio automatic recognition and switching method for a surveillance camera. Since the principle of the system in this application is similar to the audio automatic recognition and switching method for a surveillance camera described above, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be described again.
[0195] See Figure 4 This is a schematic diagram of an automatic audio recognition and switching system for a surveillance camera, provided in an embodiment of this application. The system includes:
[0196] Module 10 is used to acquire an angular acceleration sequence characterizing the motion state of the gimbal; construct a coupled perturbation hidden state model based on the angular acceleration sequence, wherein the hidden state includes the coupling strength; and establish a nonlinear observation model based on the coupled perturbation hidden state model.
[0197] The calculation module 20 performs particle filtering estimation on the hidden state based on the nonlinear observation model. During each particle update, it performs variational marginalization calculation on the continuous variables in the hidden state to obtain the posterior distribution of the hidden state. Within a preset prediction window, it determines the statistics of the coupling strength based on the posterior distribution, generates perturbation feature values, and determines the avoidance time window of the gimbal movement based on the perturbation feature values.
[0198] The processing module 30 is used to collect ground potential change signals and power amplifier supply voltage, calculate the rate of change of the power amplifier supply voltage, and generate an electrical stability flag based on the ground potential change signal and the rate of change of the power amplifier supply voltage; detect the frame alignment flag of the audio buffer and the zero crossover state of the current audio waveform, and generate an audio mute flag based on the detection results; and generate a mute trigger signal based on the audio mute flag.
[0199] Output module 40 is used to output a switching control signal when the current time point is not within the avoidance time window and the electrical stability flag and the audio silence flag are valid.
[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for automatic audio recognition and switching applied to surveillance cameras, characterized in that, include: Obtain the angular acceleration sequence characterizing the motion state of the gimbal; A coupled perturbation hidden state model is constructed based on the angular acceleration sequence. The hidden state includes the coupling strength, which is used to characterize the perturbation gain of a unit gimbal motion excitation transmitted to the microphone mounting position under the current attitude and load conditions. A nonlinear observation model is established based on the coupled perturbation hidden state model. Based on the nonlinear observation model, particle filtering estimation is performed on the hidden state. In each particle update process, variational marginalization calculation is performed on the continuous variables in the hidden state to obtain the posterior distribution of the hidden state. Within a preset prediction window, the statistics of the coupling strength are determined based on the posterior distribution, a disturbance feature value is generated, and the avoidance time window of the gimbal movement is determined based on the disturbance feature value. Collect ground potential change signals and power amplifier supply voltage, calculate the rate of change of the power amplifier supply voltage, and generate an electrical stability flag based on the ground potential change signals and the rate of change of the power amplifier supply voltage. Detect the frame alignment flag of the audio buffer and the zero-crossing state of the current audio waveform, and generate an audio silence flag based on the detection results; Generate a silence trigger signal based on the audio silence flag; When the current time point is not within the avoidance time window and the electrical stability flag and the audio silence flag are valid, output a switching control signal.
2. The method for automatic audio recognition and switching applied to a surveillance camera according to claim 1, characterized in that, The hidden state also includes a coupling direction angle and an extended scale parameter, and the method further includes: The coupling direction angle and the extended scale parameter are defined in a two-dimensional direction vector space; An elliptic parametric equation is established in the two-dimensional direction vector space, where the major axis of the ellipse corresponds to the coupling direction angle, and the minor axis corresponds to the orthogonal direction of the coupling direction angle. The scale ratio and rotation angle of the ellipse are calculated by fitting the variance ratio of the angular acceleration sequence in the major axis direction and the minor axis direction. The coupled perturbation hidden state model, including directional constraint terms, is generated based on the obtained ellipse parameters.
3. The method for automatic audio recognition and switching applied to a surveillance camera according to claim 2, characterized in that, The establishment of the nonlinear observation model includes: The angular acceleration sequence is directionally projected along the coupling direction angle and its orthogonal direction to obtain a first direction projection sequence and a second direction projection sequence. The adaptive time window and bandwidth are determined according to the extended scale parameters, and the sub-band energy measurement with the same scale is calculated for the first direction projection sequence and the second direction projection sequence. To address impulse noise and ground potential abrupt changes, a robust noise model based on a skewed t-distribution is established, generating observation residuals. A dependency structure model is performed on the edge distributions of the first direction projection sequence and the second direction projection sequence using a connection function to obtain the joint observation distribution; Based on the joint observation distribution and the observation residuals, the parameter set and observation mapping of the nonlinear observation model are determined.
4. The method for automatic audio recognition and switching applied to a surveillance camera according to claim 1, characterized in that, The posterior distribution of the hidden state is obtained as follows: After detecting the silence trigger signal, the instantaneous value of angular acceleration corresponding to the gimbal attitude is read, and the acoustic spatial drift is calculated based on the instantaneous value of angular acceleration. The prior mean of the continuous variables in the particle set is corrected based on the acoustic spatial drift, thereby limiting the particle update range; During the detection period when the electrical stability indicator is valid, variational marginalization calculation is performed on the continuous variable to generate a marginalized distribution of the continuous variable; When the silence trigger signal is detected to change from valid to invalid, the distribution of the continuous variable is used as the initial input for particle propagation to update the posterior distribution of the hidden state.
5. The method for automatic audio recognition and switching applied to a surveillance camera according to claim 4, characterized in that, The generated marginalized continuous variable distribution includes: Based on the acoustic spatial drift and gimbal attitude, a preset acoustic coupling calibration table is invoked to obtain the prior interval of coupling strength and the coupling noise baseline under the corresponding attitude. Under the constraint of the prior interval of the coupling strength, a variational distribution family for characterizing the continuous variable is selected, and the variational distribution for characterizing the continuous variable is obtained by iteratively updating within the variational distribution family. The mean vector and covariance matrix of the variational distribution are iteratively updated. In each iteration update, the residual noise power measured during the effective period of the silent trigger signal is compared with the coupling noise power predicted by the variational distribution obtained in the current iteration. When the deviation between the two exceeds the first preset threshold, the weight of the corresponding particle is reduced and the prior interval of the coupling strength is tightened. When the variational optimization convergence condition is met, the final variational distribution is taken as the marginalized continuous variable distribution.
6. The method for automatic audio recognition and switching applied to a surveillance camera according to claim 1, characterized in that, The generated perturbation feature values include: During the assembly stage of the surveillance camera, the low-order resonant frequency band of the camera housing and mounting bracket, as well as the unit impact response function related to the motion coupling of the pan-tilt unit, are obtained in advance through simulation analysis, and the low-order resonant frequency band and the unit impact response function are stored in the structural mode library. Within the preset prediction time window, a time series estimate of coupling strength is generated based on the posterior distribution, and the time series estimate of coupling strength is convolved with the unit impulse response function to obtain the structural vibration response sequence at the microphone installation location. Based on at least one of the peak amplitude or root mean square value of the structural vibration response sequence within the preset prediction time window, a statistical quantity characterizing the amplitude of structural vibration is obtained, and the statistical quantity is used as the disturbance characteristic value.
7. The method for automatic audio recognition and switching applied to a surveillance camera according to claim 6, characterized in that, The statistical measures used to characterize the amplitude of structural vibration include: For each candidate switching moment, within a local time window centered on that candidate switching moment, the structural vibration response sequence is decomposed into frequency bands, and the ratio of energy falling into the low-order resonance frequency band to the total energy is calculated. The obtained energy ratio is used as the structural resonance risk index for the candidate switching moment. The structural resonance risk index is compared with a preset resonance threshold. When the structural resonance risk index is higher than the preset resonance threshold, the corresponding time period is marked as the avoidance time window of the gimbal movement.
8. The method for automatic audio recognition and switching applied to a surveillance camera according to claim 1, characterized in that, The generated audio silence flag includes: During the deployment phase of the surveillance camera, audio buffer data is collected to characterize the audio buffer structure. Based on the starting position of the audio frame in the audio buffer data, a reference range for the frame start position is determined, and the reference range is stored as a frame header template. During operation, for real-time audio frames written to the audio buffer, the frame alignment flag of the audio frame is determined based on whether the time information and sampling position of the audio frame fall within the allowable offset range given by the frame alignment template. For audio frames with valid frame alignment flags, the zero-crossing feature and energy feature of the current audio waveform are extracted in its starting neighborhood. When the zero-crossing feature and the energy feature meet the preset silence determination conditions, the audio frame is marked as a silence candidate frame. Based on the distribution of silent candidate frames within a preset time window, determine whether the audio state of the corresponding time period is silent, and generate the audio silence flag when it is determined to be silent.
9. The method for automatic audio recognition and switching applied to a surveillance camera according to claim 8, characterized in that, The step of marking the audio frame as a silent candidate frame includes: During the deployment phase of the surveillance camera, background audio buffer data is collected for different pan-tilt positions. Based on the zero-crossing characteristics and amplitude characteristics of the silent segment audio frames under each pan-tilt position, a corresponding background noise template is generated, and the association between each pan-tilt position and the background noise template is stored in the silent judgment library. During operation, the corresponding background noise template is retrieved from the silence determination library according to the current gimbal attitude. For audio frames with valid frame alignment, zero-crossing features and amplitude features are extracted in their starting neighborhood, and the deviation metric between the zero-crossing features and amplitude features and the background noise template is calculated. When the deviation metric meets the preset silence threshold condition, the audio frame is determined as the silence candidate frame.
10. An automatic audio recognition and switching system for surveillance cameras, characterized in that, The system includes: A construction module is used to obtain an angular acceleration sequence characterizing the motion state of the gimbal; a coupled perturbation hidden state model is constructed based on the angular acceleration sequence, wherein the hidden state includes the coupling strength, which is used to characterize the perturbation gain of a unit gimbal motion excitation transmitted to the microphone mounting position under the current attitude and load conditions; and a nonlinear observation model is established based on the coupled perturbation hidden state model. The calculation module performs particle filtering estimation on the hidden state based on the nonlinear observation model. During each particle update, it performs variational marginalization calculation on the continuous variables in the hidden state to obtain the posterior distribution of the hidden state. Within a preset prediction window, it determines the statistics of the coupling strength based on the posterior distribution, generates perturbation feature values, and determines the avoidance time window of the gimbal movement based on the perturbation feature values. The processing module is used to collect ground potential change signals and power amplifier supply voltage, calculate the rate of change of the power amplifier supply voltage, and generate an electrical stability flag based on the ground potential change signal and the rate of change of the power amplifier supply voltage; detect the frame alignment flag of the audio buffer and the zero crossover state of the current audio waveform, and generate an audio mute flag based on the detection results; and generate a mute trigger signal based on the audio mute flag. The output module is used to output a switching control signal when the current time point is not within the avoidance time window and the electrical stability flag and the audio silence flag are valid.
Citation Information
Patent Citations
Automatic switching system and method of audio signals
CN107948873A
Microphone and noise elimination device for microphone channel switching
CN113891218A