Interrogation psychological sensitive point detection method based on timing attention mechanism

By employing a temporal attention mechanism-based approach, and utilizing short-time Fourier transform and affine invariance metric to correct the geometric shift of the spatial covariance matrix, the problem of temporal drift in interrogation psychological testing is solved, enabling accurate localization and adaptive enhancement of psychologically sensitive points.

CN121502298BActive Publication Date: 2026-05-01BEIJING TONGFANG SHENHUO UNITED SCI & TECH DEV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING TONGFANG SHENHUO UNITED SCI & TECH DEV
Filing Date
2025-11-12
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing interrogation psychological testing methods cannot effectively correct the geometric distribution shift of the spatial covariance matrix when the acoustic environment is unstable, leading to temporal drift and misjudgment of psychologically sensitive points.

Method used

The spatial covariance matrix time series is calculated by short-time Fourier transform, the relative displacement weighting kernel is extracted, the curvature density index of the two-flow section is constructed using affine invariant metric, gating weights are generated, geometric alignment is performed, and the psychologically sensitive moments of interrogation are output in combination with adaptive threshold.

Benefits of technology

It achieves accurate positioning of psychologically sensitive points during interrogation in a dynamic acoustic environment, effectively filters out interference from non-psychological factors, and improves the accuracy and adaptability of positioning during psychologically sensitive moments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502298B_ABST
    Figure CN121502298B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of attention mechanism, and discloses an interrogation psychological sensitive point detection method based on a timing attention mechanism, wherein a microphone signal is processed in a local time window through a short-time Fourier transform to generate a spatial covariance matrix time sequence under a symmetric positive definite manifold, and a relative displacement weighted kernel reflecting inter-frame time lag correlation is extracted from a timing attention model output; then, based on an affine invariant metric, a bisection vector is constructed to calculate a bisection cross-sectional curvature density to quantize the geometric inconsistency of the two; subsequently, the weighted kernel is adjusted by the index gate, the geometric aligned predicted spatial covariance matrix is obtained through cut space summation and exponential mapping, the actual and predicted matrix residuals are calculated by measuring the distance of the affine invariant metric, a single time sequence score is generated by energy proportion weighting, and finally, an adaptive threshold is constructed by using the median absolute deviation of the score and a threshold coefficient to screen sensitive moments that meet the conditions, thereby effectively improving the detection accuracy and scene adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

A Method for Detecting Psychological Sensitivity Points in Interrogation Based on Temporal Attention Mechanism Technical Field

[0001] This invention relates to the field of attention mechanism technology, and more specifically, to a method for detecting psychologically sensitive points during interrogation based on a temporal attention mechanism. Background Technology

[0002] When interrogation speech propagates in an enclosed space, it is affected by changes in walls, tabletops, and the position of the personnel. The room's impact response drifts slowly over time, causing the spatial covariance matrix of the signal received by the microphone array to continuously change. Modern interrogation speech understanding models typically utilize temporal attention mechanisms to aggregate information within local time windows to capture semantic cues and emotional responses. This method can effectively locate psychologically sensitive moments when the acoustic environment is stable.

[0003] Existing interrogation psychological testing methods typically rely on the short-term stability of acoustic propagation, employing covariance smoothing or weighted averaging in Euclidean space. When subtle changes occur in the sound source location or the speaker's posture, the room's impact response changes over time, shifting the geometric distribution of the spatial covariance matrix, while the attention mechanism still aggregates neighborhood information through a fixed window. However, this approach ignores the sequential difference between acoustic transmission and attention aggregation, easily leading to temporal localization errors. This causes the model to focus on non-psychologically triggered moments during abrupt response changes, interruptions, or posture shifts, resulting in a misalignment between the detection results and the actual psychological response.

[0004] Under conditions of dynamic changes in room impact response, the direction of movement of the spatial covariance matrix on a symmetric positive definite manifold is inconsistent with the convergence direction of the temporal attention mechanism, and the two are no longer geometrically commutative. This phenomenon causes statistical shifts in the local temporal neighborhood, leading to the propagation of attention weights along incorrect paths and resulting in temporal drift of psychologically sensitive points. Existing methods lack a unified mechanism to observe and correct this directional bias at the manifold geometry level. Summary of the Invention

[0005] This invention provides a method for detecting psychologically sensitive points during interrogation based on a temporal attention mechanism, which solves the technical problems mentioned in the background art.

[0006] This invention provides a method for detecting psychologically sensitive points during interrogation based on a temporal attention mechanism, comprising:

[0007] Calculate the time series of the spatial covariance matrix within a local time window using short-time Fourier transform;

[0008] Extract the relative displacement weighting kernel from the attention matrix output by the temporal attention model;

[0009] Based on the affine invariant metric, a first tangent vector corresponding to the time-varying room impact response and a second tangent vector corresponding to the attention weighting are constructed at the center frame, and the curvature density index of the two-flow section spanned by the first tangent vector and the second tangent vector is calculated.

[0010] Based on the curvature density index of the dual-flow section and the gating coefficient, the gating weight is generated, and the relative displacement weighted kernel is gating. After summing in the tangent space, the geometrically aligned prediction space covariance matrix is ​​obtained through exponential mapping.

[0011] The residuals between the actual spatial covariance matrix and the geometrically aligned predicted spatial covariance matrix are calculated using geodesic distance with affine invariant metric, and the residuals are weighted according to energy ratio to obtain a single time series score.

[0012] An adaptive threshold is constructed using the median absolute deviation of a single time series score and a threshold coefficient. The psychologically sensitive moments during interrogation are then output based on the adaptive threshold.

[0013] The beneficial effects of this invention include: by placing the spatial covariance matrix on a symmetric positive definite manifold, constructing a double tangent vector using an affine invariant metric, and calculating the curvature density index of the double manifold cross section, the inconsistency between temporal attention aggregation and spatial signal geometric changes is quantified. Then, by combining the index with a dynamically gated relative displacement weighted kernel, tangent space summation, and exponential mapping, a geometrically aligned prediction matrix is ​​obtained. Finally, an adaptive threshold is generated based on the statistical characteristics of a single time series score. This effectively filters interference caused by non-psychological factors in interrogation scenarios, avoids temporal drift and misjudgment of psychologically sensitive points, and significantly improves the accuracy of locating psychologically sensitive moments and the adaptability to different interrogation acoustic scenarios. Attached Figure Description

[0014] Figure 1 is a flowchart of the interrogation psychological sensitivity detection method based on the temporal attention mechanism of the present invention. Detailed Implementation

[0015] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.

[0016] As shown in Figure 1, the interrogation psychological sensitivity detection method based on the temporal attention mechanism includes:

[0017] Calculate the time series of the spatial covariance matrix within a local time window using short-time Fourier transform;

[0018] Extract the relative displacement weighting kernel from the attention matrix output by the temporal attention model;

[0019] Based on the affine invariant metric, a first tangent vector corresponding to the time-varying room impact response and a second tangent vector corresponding to the attention weighting are constructed at the center frame, and the curvature density index of the two-flow section spanned by the first tangent vector and the second tangent vector is calculated.

[0020] Based on the curvature density index of the dual-flow section and the gating coefficient, the gating weight is generated, and the relative displacement weighted kernel is gating. After summing in the tangent space, the geometrically aligned prediction space covariance matrix is ​​obtained through exponential mapping.

[0021] The residuals between the actual spatial covariance matrix and the geometrically aligned predicted spatial covariance matrix are calculated using geodesic distance with affine invariant metric, and the residuals are weighted according to energy ratio to obtain a single time series score.

[0022] An adaptive threshold is constructed using the median absolute deviation of a single time series score and a threshold coefficient. The psychologically sensitive moments during interrogation are then output based on the adaptive threshold.

[0023] In one embodiment of the present invention, the spatial covariance matrix time series is calculated within a local time window using short-time Fourier transform, including:

[0024] S21. For the discrete-time signal of each microphone channel, select continuous discrete-time sample segments according to the preset window length, apply a window function to the discrete-time sample segments and perform Fourier operation to obtain the time-frequency signal of each microphone channel corresponding to different frames and different frequency points; combine the time-frequency signals of all microphone channels under the same frame and the same frequency point to form a multi-channel time-frequency vector.

[0025] S22, set the center frame and the number of frames in the window radius, determine the number of frames contained in the local time window based on the number of frames in the window radius, which is twice the number of frames in the window radius plus one; using the center frame as the reference, select all frames from the center frame minus the number of frames in the window radius to the center frame plus the number of frames in the window radius to form the local time window;

[0026] S23, for each center frame and each frequency point, within the corresponding local time window, perform conjugate transpose operation on the multi-channel time-frequency vector of each frame and then multiply it with the multi-channel time-frequency vector itself to obtain the vector outer product corresponding to each frame.

[0027] S24, average the vector outer product of all frames within the local time window to obtain the spatial covariance matrix corresponding to the center frame and the frequency point;

[0028] S25, change the center frame sequentially, and repeat S22 to S24 to obtain the time series of the spatial covariance matrix that varies with the center frame.

[0029] For the discrete-time signal received by each microphone, a continuous discrete-time sample segment is extracted according to a pre-set window length (e.g., 20 to 30 milliseconds to adapt to the voice signal). To reduce spectral leakage, a window function (e.g., Hanning window, Hamming window) is applied to the sample segment, and then Fourier operation (e.g., short-time Fourier transform) is performed to obtain the time-frequency signal (complex value, including amplitude and phase information) of each microphone at different time frames and different frequency points. Subsequently, for the same time frame and the same frequency point, the time-frequency signals of all microphone channels are integrated into a vector, namely the multi-channel time-frequency vector, which can reflect the spatial distribution of signals from multiple microphones at that time-frequency point.

[0030] First, determine the center frame (i.e., the target frame for which the spatial covariance matrix needs to be calculated) and the number of frames within the window radius (e.g., 4 or 8 frames). The total number of frames within the local time window can be determined based on the number of frames within the window radius. Using the center frame as a reference, select all time frames within the preset local time window to form the local time window. The purpose of the local time window is to statistically analyze the spatial correlation of the signal within a local time range, balancing temporal resolution (avoiding excessively large windows that lead to blurred temporal information) with statistical stability (avoiding excessively small windows that lead to significant noise interference).

[0031] For each center frame and each frequency point, within the corresponding local time window, the target operation is performed on the multi-channel time-frequency vector of each frame, as follows:

[0032] First, perform the conjugate transpose of the vector (a standard operation for processing complex signals).

[0033] Multiply the conjugate transpose of the vector with the original multi-channel time-frequency vector (i.e., the vector outer product operation).

[0034] Through target operation, the vector outer product of each frame at that frequency point can be obtained. The vector outer product is the basic unit for calculating the spatial covariance matrix, and its elements reflect the correlation of signals from different microphone channels at that time and frequency point.

[0035] Calculate the arithmetic mean of the vector outer products of all frames within the local time window at the current frequency point; this mean is the spatial covariance matrix corresponding to the current center frame and the current frequency point. The dimension of the spatial covariance matrix is ​​the same as the number of microphones, and its element values ​​reflect the spatial correlation strength between signals from different microphone channels. Due to the non-negative signal energy and correlation characteristics, this matrix belongs to a symmetric positive definite manifold.

[0036] Keeping the window radius and number of frames constant, the center frame is moved forward sequentially (e.g., by one frame each time). For each new center frame, the operations S22 (defining the local time window corresponding to the new center frame), S23 (calculating the vector cross product of each frame within the window), and S24 (averaging to obtain the spatial covariance matrix) are repeated. Through this process, a series of spatial covariance matrices varying with the center frame (i.e., time) can be obtained, i.e., the spatial covariance matrix time series. The spatial covariance matrix time series can dynamically capture the impact of room impact response changes over time (e.g., speaker posture micro-movements) on the spatial correlation of multi-microphone signals.

[0037] In one embodiment of the present invention, extracting the relative displacement weighting kernel from the attention matrix output by the temporal attention model includes:

[0038] The set of relative displacement values ​​is determined based on the number of frames at the window radius; the range of the relative displacement value set is from negative to positive window radius frames.

[0039] For each relative displacement in the set of relative displacement values, an equivalence discrimination function is set; the equivalence discrimination function takes the value of one when the input value is zero and takes the value of zero when the input value is not zero.

[0040] For all elements in the temporal attention matrix whose row and column time frame positions both belong to a local time window, the difference between the row and column time frame positions is equal to the current relative displacement is filtered out using an equality discrimination function. The filtered elements are then summed to obtain the unnormalized cumulative relative displacement corresponding to that relative displacement.

[0041] The normalization factor is obtained by summing all elements in the temporal attention matrix whose row and column time frame positions both belong to the local time window.

[0042] Divide the cumulative unnormalized relative displacement corresponding to each relative displacement by the normalization factor to obtain the weight corresponding to each relative displacement. All relative displacements and their corresponding weights together constitute the relative displacement weighting kernel.

[0043] Specifically, the range of relative displacement values ​​is determined by the number of frames in the window radius: from negative to positive. For example, if the window radius is 4 frames, the set of relative displacement values ​​includes all integers from -4 to 4. This range precisely covers the position difference between any two frames within the local time window (based on the center frame, extended by the window radius in frames to the left and right).

[0044] The purpose of the equality judgment function is to determine whether the difference between the row time frame position and the column time frame position is equal to the relative displacement to be calculated. The rule is: when the input frame position difference equals the target relative displacement (i.e., the input value is zero, here the input value refers to the frame position difference minus the target relative displacement), the function outputs 1; if it is not equal (the input value is non-zero), it outputs 0. This function can quickly identify elements in the temporal attention matrix that match the current relative displacement.

[0045] First, the statistical scope is limited: only elements in the temporal attention matrix whose row and column time frame positions both belong to the local time window are considered (to avoid including attention weights from irrelevant frames outside the window and ensure data relevance). Next, elements are filtered using an equality discrimination function: only elements where row frame position - column frame position = current relative displacement are retained. Finally, the filtered elements are summed, and the result is the unnormalized cumulative relative displacement. The unnormalized cumulative relative displacement reflects the sum of attention weights for all frame pairs separated by the current relative displacement within the local time window, but scaling has not yet been performed.

[0046] The summation is performed on all elements in the temporal attention matrix whose row and column time frame positions both belong to the local time window. This sum is the normalization factor, which reflects the total attention weight of all frame pairs within the local time window. The normalization factor can scale the unnormalized cumulative values ​​of different relative displacements to the same scale.

[0047] The unnormalized cumulative value corresponding to each relative displacement is divided by the normalization factor to obtain the normalized weight corresponding to that relative displacement (the sum of the weights of all displacements is 1). Finally, the relative displacements are mapped one-to-one with their corresponding normalized weights to form a relative displacement weighting kernel. The function of this weighting kernel is to transform the two-dimensional frame pair attention matrix into a one-dimensional displacement-weight relationship, so that it can match the time delay characteristics of acoustic propagation (such as the time delay of room impact response) on the same dimension (displacement / time delay).

[0048] In one embodiment of the present invention, based on an affine invariant metric, a first tangent vector corresponding to the time-varying room impact response and a second tangent vector corresponding to the attention weighting are constructed at the center frame, including:

[0049] S41, for each discrete frequency point, take the first symmetric matrix and the second symmetric matrix with the same dimension as the spatial covariance matrix;

[0050] Invert the first spatial covariance matrix at the center frame, and then multiply it sequentially with the first symmetric matrix, the inverted first spatial covariance matrix at the center frame, and the second symmetric matrix.

[0051] S42, find the trace of the result of the multiplication operation to form an affine invariant metric;

[0052] S43, take the first spatial covariance matrix at the center frame and any spatial covariance matrix in the neighborhood of the center frame; perform a square root operation on the first spatial covariance matrix at the center frame, and then invert the square root matrix;

[0053] S44, multiply the inverted matrix with the spatial covariance matrix in the neighborhood and the square root of the inverted matrix in turn; take the logarithm of the result of the multiplication, and then multiply the logarithm with the square root of the first spatial covariance matrix at the center frame to form a logarithmic mapping.

[0054] S45, sets the time step frame number as a positive integer;

[0055] S46, take the first spatial covariance matrix at the center frame and the second spatial covariance matrix after the center frame is delayed by the time step frame number; process the first spatial covariance matrix and the second spatial covariance matrix through the logarithmic mapping in S44 to obtain the fusion mapping result; divide the fusion mapping result by the time step frame number to obtain the first tangent vector corresponding to the time-varying room impact response.

[0056] S47, take the relative displacement weighting kernel, and for each relative displacement, take the third spatial covariance matrix after the center frame is delayed by the relative displacement;

[0057] S48, the third space covariance matrix is ​​processed by the logarithmic mapping of S44 to obtain the mapping result corresponding to each relative displacement; the mapping result corresponding to each relative displacement is multiplied by the weight of the relative displacement in the relative displacement weighting kernel; the sum of all multiplication results is obtained to obtain the second tangent vector corresponding to the attention weighting.

[0058] For each discrete frequency point, two symmetric matrices with the same dimension as the spatial covariance matrix are selected (due to the manifold tangent vector being a symmetric matrix, dimension matching is required). The inverse of the first spatial covariance matrix of the center frame is then multiplied sequentially by the first symmetric matrix, the inverted first spatial covariance matrix of the center frame, and the second symmetric matrix. This operation achieves scale invariance through matrix inverse transformation, ensuring that the metric is not affected by the overall scaling of the spatial covariance matrix. The trace of the multiplication result (the sum of the elements on the main diagonal of the matrix) is then taken to transform the matrix into a scalar. This scalar is the inner product of the two symmetric matrices under an affine invariant metric, ultimately forming a directly usable affine invariant metric.

[0059] Take the first spatial covariance matrix of the central frame (the reference matrix) and any spatial covariance matrix in the neighborhood (the matrix to be compared). Perform a square root operation on the central frame matrix (to obtain the positive definite square root), and then invert the square root matrix. The purpose is to center the neighborhood matrix to near the identity matrix, reducing the complexity of subsequent logarithmic calculations. Multiply the inverse square root matrix sequentially with the neighborhood matrix and the inverted square root matrix (to complete the centering transformation and eliminate the scaling effect of the reference matrix). Take the logarithm of the result (to transform the nonlinear difference after centering into a linear vector). Then multiply the logarithm with the square root matrix of the central frame (to map the linear vector back to the original tangent space). Finally, a logarithmic mapping is formed, realizing the transformation from manifold matrix difference to tangent space vector.

[0060] Set the time step frame size as a positive integer (usually 1 frame). A time step frame size that is too large will miss subtle time variations in the RIR, while a time step size that is too small will introduce noise; a positive integer setting ensures the discreteness of the time interval and is consistent with the time unit of the frame sequence.

[0061] Take the first spatial covariance matrix of the center frame (current state) and the second spatial covariance matrix after the time step delay (time-varying state); through the logarithmic mapping of S44, transform the difference between the two matrices into the fusion mapping result in the tangent space (linearized expression of nonlinear difference); divide the fusion mapping result by the number of time steps to obtain the change vector of the spatial covariance matrix per unit time, i.e., the first tangent vector.

[0062] A relative displacement weighting kernel is selected. For each displacement, a third spatial covariance matrix (the matrix to be aggregated in the neighborhood) is chosen after delaying the center frame by that displacement. Through the logarithmic mapping of S44, the difference between each third spatial covariance matrix and the center frame matrix is ​​transformed into a tangent space vector (the difference vector of a single displacement). Each difference vector is multiplied by the attention weight of the corresponding displacement (reflecting the difference in attention contribution of different displacements). The summation of all weighted difference vectors yields the total difference vector of the attention aggregation direction, i.e., the second tangent vector, which reflects the aggregation driving effect of attention weighting on the matrix.

[0063] In one embodiment of the present invention, calculating the curvature density index of the two-flow cross-section spanned by the first tangent vector and the second tangent vector includes:

[0064] Calculate the first inner product of the second tangent vector and the affine invariant metric, the second inner product of the first tangent vector and the affine invariant metric, and the fused inner product of the second tangent vector and the first tangent vector, respectively.

[0065] Multiply the first inner product by the second inner product, then subtract the square of the merged inner product, and take the square root of the result to obtain the area element;

[0066] Take the first spatial covariance matrix at the center frame and invert it. Then multiply the inverted matrix with the second tangent vector and the first tangent vector respectively to obtain the first product matrix and the second product matrix.

[0067] Perform commutation operations on the first and second product matrices to obtain the commutation operation results;

[0068] Calculate the Frobenius norm of the result of the variability operation, and multiply the square of the Frobenius norm by one-quarter to obtain the numerator;

[0069] Use the square of the area element as the denominator;

[0070] The curvature of the cross section is obtained by dividing the numerator by the denominator and taking the opposite number.

[0071] The cross-sectional curvature is multiplied by the area element and the opposite number is taken to obtain the curvature density index of the dual-flow cross-section.

[0072] First inner product: Calculates the inner product of the second tangent vector under the affine invariant metric, quantifying the geometric length of the second tangent vector.

[0073] Second inner product: Calculate the inner product of the first tangent vector under the affine invariant metric, quantifying the geometric length of the first tangent vector.

[0074] The third inner product: calculates the inner product of the second tangent vector and the first tangent vector under the affine invariant metric, quantifying the angle relationship between the two tangent vectors (i.e. the degree of correlation of the driving effect).

[0075] First, multiply the first inner product with the second inner product (representing the product of the lengths of the two vectors), then subtract the square of the fused inner product (to correct for the reduction in span caused by the included angle). The scalar obtained by taking the square root of the result is the area element. The larger the area element, the more dispersed the two tangent vectors are in the tangent space, and the larger the spatial span.

[0076] Take the first spatial covariance matrix at the center frame and invert it (obtain the inverse matrix to eliminate the scale bias of the original matrix). Multiply the inverse matrix with the second tangent vector and the first tangent vector respectively to obtain the first product matrix (corresponding to the corrected second tangent vector) and the second product matrix (corresponding to the corrected first tangent vector).

[0077] Perform a commutation operation on the first and second product matrices. The result of the commutation operation is a matrix. If the resulting matrix is ​​close to zero, it means that the two driving forces can be approximately commutated (strong consistency); if the resulting matrix is ​​highly non-zero, it means that the two driving forces are difficult to commutate (strong inconsistency).

[0078] First, calculate the Frobenius norm of the result of the change operation (by taking the square root of the sum of the squares of all elements of the matrix, the matrix is ​​transformed into a scalar that reflects its overall size), and then multiply the square of the norm by one-quarter to facilitate the calculation of the cross-sectional curvature.

[0079] Divide the numerator by the denominator and then take the opposite of the result (since the cross-sectional curvature of a symmetric positive definite manifold is inherently non-positive, taking the opposite value more intuitively reflects the degree of curvature). The resulting scalar is the cross-sectional curvature. The larger the absolute value of the cross-sectional curvature, the more significant the curvature of the plane spanned by the two tangent vectors, and the stronger the inconsistency of the driving force.

[0080] Multiplying the cross-sectional curvature by the area element and then taking the negative of the result yields a scalar, which is the curvature density index of the two-flow cross-section. The larger the value of this curvature density index, the stronger the geometric inconsistency between the driving forces corresponding to the two tangent vectors; a curvature density index value close to zero indicates that the two driving forces can be approximately synergistic, with strong geometric consistency.

[0081] In one embodiment of the present invention, a gating weight is generated based on the curvature density index of the dual-flow section and the gating coefficient. The relative displacement weighted kernel is then gating, and after tangent space summation, an exponential mapping is used to obtain a geometrically aligned prediction space covariance matrix, including:

[0082] Set the gating coefficient for positive real numbers;

[0083] For each relative displacement in the set of relative displacement values, the curvature density index of the dual-flow section corresponding to the frame after the relative displacement is delayed in the center frame is taken as the target curvature density index of the dual-flow section.

[0084] Multiply the gating coefficient by the curvature density index of the target bifluid cross section, take the negative of the product as the exponent of the exponential function, calculate the result of the exponential function, and obtain the gating weight corresponding to the relative displacement.

[0085] Multiply the weight of each relative displacement in the relative displacement weighting kernel by the corresponding gating weight to obtain the weighted product of the corresponding relative displacement; sum the weighted products of all relative displacements to obtain the gating normalization factor.

[0086] Divide the weight product of each relative displacement by the gating normalization factor to obtain the gating weight for each relative displacement. All relative displacements and their corresponding gating weights together constitute the gating weighted kernel of relative displacements.

[0087] For each relative displacement, the first spatial covariance matrix at the center frame and the third spatial covariance matrix after the center frame is delayed by the relative displacement are taken. The first spatial covariance matrix and the third spatial covariance matrix are processed by S44 logarithmic mapping to obtain the logarithmic mapping result corresponding to the relative displacement.

[0088] The logarithmic mapping result corresponding to each relative displacement is multiplied by the weight of that relative displacement in the gated relative displacement weighted kernel, and the sum of all multiplication results is obtained to obtain the tangent space summation quantity;

[0089] Take the first spatial covariance matrix at the center frame, perform a square root operation on the first spatial covariance matrix to obtain the square root matrix; invert the square root matrix to obtain the square root inverse matrix; multiply the tangent space summator on the left and right by the square root inverse matrix to obtain the intermediate matrix; perform a matrix exponent operation on the intermediate matrix to obtain the exponent matrix; multiply the square root matrix on the left and right by the exponent matrix to obtain the geometrically aligned prediction spatial covariance matrix.

[0090] The gating coefficient is set to a positive real number. A positive real number ensures that after multiplying with the index, the gating logic of suppressing high indexes and retaining low indexes can be achieved through exponential operation. The size of the coefficient determines the gating sensitivity. The larger the coefficient, the stronger the suppression of high indexes (strong geometric inconsistency), and vice versa.

[0091] For each relative displacement in the set of relative displacement values, the curvature density index of the dual-flow section corresponding to the frame after delaying that relative displacement in the center frame is selected as the target index. Since different relative displacements correspond to different neighboring frames, the geometric inconsistencies (index values) of each frame are different. Therefore, a target index needs to be matched separately for each displacement to achieve precise gating.

[0092] The gating coefficient is multiplied by the target index, and the negative of the product is taken as the exponent of the exponential function. The result of the exponential function is the gating weight of the relative displacement. The exponential function property allows a high target index (strong inconsistency) to correspond to a low gating weight (suppressing the contribution of the displacement), and a low target index (strong consistency) to correspond to a high gating weight (preserving the contribution of the displacement), thus achieving adaptive filtering of unstable neighborhoods.

[0093] Weighted product: The attention weight of each relative displacement in the original weighted kernel is multiplied by the corresponding gating weight to obtain the weighted product. The weighted product combines the requirements of attention aggregation and geometric consistency.

[0094] Gating normalization factor: The normalization factor is obtained by summing the weights of all relative displacements. The normalization factor ensures that the sum of the weights after subsequent gating is 1.

[0095] Divide the weight product of each relative displacement by the gating normalization factor to obtain the gated weight of that displacement; map all relative displacements to the gated weights to form a gated relative displacement weighting kernel. The gated relative displacement weighting kernel has filtered out geometrically unstable displacements and only retains the contribution of stable displacements that are effective for aggregation.

[0096] For each relative displacement, the first spatial covariance matrix (reference matrix) at the center frame and the third spatial covariance matrix (neighborhood matrix) after delaying the displacement at the center frame are taken. The two matrices are processed by the logarithmic mapping defined by S44. The nonlinear matrix difference on the manifold is transformed into a linear vector in the tangent space (logarithmic mapping result).

[0097] The logarithmic mapping result (tangent space vector) corresponding to each relative displacement is multiplied by the weight of that displacement in the gated weighted kernel. The sum of all the product results is obtained as a linear aggregation of the contributions of each displacement in the stable neighborhood, retaining only geometrically consistent neighborhood information.

[0098] Matrix preprocessing: The square root operation is performed on the first spatial covariance matrix of the center frame to obtain the square root matrix, and then the inverse of the square root matrix is ​​obtained. The purpose is to center the summation of the tangent space to the vicinity of the identity matrix to ensure the rationality of the exponential mapping.

[0099] Intermediate matrix generation: Multiply the summation of the tangent space on the left and on the right by the square root inverse matrix to obtain the intermediate matrix, thus completing the centering transformation and adapting to the operational requirements of the exponential mapping;

[0100] Exponential mapping and scale restoration: Perform matrix exponentiation on the intermediate matrix (mapping the tangent space vector back to the manifold) to obtain the exponential matrix; multiply the square root matrix on the left and right by the exponential matrix to restore the scale of the original matrix, and finally obtain the geometrically aligned prediction space covariance matrix. This matrix is ​​geometrically consistent with the actual matrix on the manifold and has no aggregation bias.

[0101] In one embodiment of the present invention, the residual between the actual spatial covariance matrix and the geometrically aligned predicted spatial covariance matrix is ​​calculated using geodesic distance with an affine invariant metric, and a single time series score is obtained by weighting the residual by energy proportion, including:

[0102] Take the geometrically aligned predicted spatial covariance matrix and its corresponding actual spatial covariance matrix in the spatial covariance matrix time series;

[0103] Perform a square root operation on the actual spatial covariance matrix, then invert the square root matrix to obtain the square root inverse matrix; multiply the square root inverse matrix on the left by the prediction spatial covariance matrix and on the right by the square root inverse matrix to obtain the intermediate matrix; perform a matrix logarithm operation on the intermediate matrix to obtain the logarithmic matrix;

[0104] Calculate the Frobenius norm of the logarithmic matrix to obtain the geodesic distance residuals;

[0105] For each discrete frequency point, within each frame of the local time window, calculate the L2 norm of the multi-channel time-frequency vector of that frame; sum the L2 norms of all frames within the local time window to obtain the energy value of that frequency point; sum the energy values ​​of all discrete frequency points to obtain the total energy value; divide the energy value of each frequency point by the total energy value to obtain the energy ratio of that frequency point.

[0106] Multiply the geodetic distance residual of each discrete frequency point by the energy ratio of that frequency point to obtain the residual weight value of each frequency point; sum the residual weight values ​​of all discrete frequency points to obtain the single time series score of the corresponding center frame.

[0107] The geometrically aligned predicted spatial covariance matrix obtained by exponential mapping, and the actual spatial covariance matrix corresponding to the geometrically aligned predicted spatial covariance matrix in the spatial covariance matrix time series, are both taken. Both correspond to the same center frame and frequency point, ensuring the spatiotemporal consistency of difference calculation.

[0108] Preprocessing: The square root operation is performed on the actual spatial covariance matrix, and then the inverse of the square root matrix is ​​obtained to obtain the square root inverse matrix. The purpose is to eliminate the scale effect of the actual matrix, center the prediction matrix to the vicinity of the identity matrix, and simplify the difference calculation.

[0109] Centering transformation: Multiply the square root inverse matrix on the left by the prediction space covariance matrix and on the right by the square root inverse matrix to obtain the intermediate matrix, thus centering the prediction matrix and making it consistent with the scale benchmark of the actual matrix.

[0110] Linear transformation: Perform a matrix logarithm operation on the intermediate matrix to obtain a logarithmic matrix, which transforms the nonlinear matrix differences on the manifold into a logarithmic matrix in linear vector form, thus realizing a quantifiable expression of the differences.

[0111] The Frobenius norm of the logarithmic matrix is ​​calculated (by taking the square root of the sum of the squares of all elements of the matrix, the matrix is ​​transformed into a scalar that reflects the magnitude of its overall difference). The resulting scalar is the geodesic distance residual. The larger the residual value, the more significant the manifold difference between the predicted matrix and the actual matrix; the smaller the residual value, the better the geometric alignment and the more accurate the prediction.

[0112] Single-frame energy: For each discrete frequency point, within each frame of the local time window, calculate the L2 norm of the multi-channel time-frequency vector of that frame (the square of the L2 norm is the vector energy, which directly reflects the signal strength of that frequency point in that frame).

[0113] Total frequency energy: The frequency energy value of a frequency point is obtained by summing the L2 norms of all frames within a local time window, reflecting the overall energy contribution of that frequency point within the window;

[0114] Total Energy and Proportion: Sum the energy values ​​of all discrete frequency points to obtain the total energy value; then divide the energy value of each frequency point by the total energy value to obtain the energy proportion of the frequency point. The larger the proportion, the stronger the information contribution of that frequency point, and the higher the residual weight should be assigned.

[0115] Residual weighting: The geodesic distance residual at each discrete frequency point is multiplied by the energy ratio of that frequency point to obtain the residual weighting value. The residual of a frequency point with higher energy has a greater contribution to the weighting value.

[0116] Temporal integration: The residuals of all frequency points are weighted and summed to obtain a single time series score for the corresponding center frame. Each center frame corresponds to only one score, forming a score sequence that changes over time. The score sequence that changes over time reflects the intensity of the difference between the prediction and the actual situation at different times.

[0117] In one embodiment of the present invention, an adaptive threshold is constructed using the median absolute deviation of a single time series score and a threshold coefficient. The psychologically sensitive moments during interrogation are output based on the adaptive threshold, including:

[0118] Obtain the individual time series scores corresponding to all time frames and form a single time series score set;

[0119] Sort all individual time series scores in the single time series score set by numerical value, and take the middle value after sorting as the median;

[0120] Calculate the absolute difference between each individual time series score and the median to obtain all absolute differences; sort all absolute differences by numerical value and take the value in the middle position after sorting as the absolute deviation of the median;

[0121] Set the consistency adjustment factor to 1.4826;

[0122] Set the threshold coefficient for positive real number types;

[0123] Multiply the threshold coefficient by the consistency adjustment coefficient to obtain the first product;

[0124] Multiply the first product by the absolute deviation of the median to obtain the second product;

[0125] Add the median to the second product to obtain the adaptive threshold;

[0126] Set the peak neighborhood radius in frames to a positive integer.

[0127] The peak neighborhood relative displacement set is determined based on the peak neighborhood radius frame number. The value range of this peak neighborhood relative displacement set is from negative peak neighborhood radius frame number to positive peak neighborhood radius frame number, and does not include zero.

[0128] First criterion: Whether the score of a single time series in each time frame is greater than or equal to the adaptive threshold;

[0129] The second criterion is whether the single time series score of each time frame is greater than or equal to the single time series score of each frame after the relative displacement in the relative displacement set of the neighborhood of the delay peak of that time frame.

[0130] The time frame that simultaneously satisfies the first and second criteria is determined as the psychologically sensitive moment during interrogation.

[0131] Obtain the single time series score corresponding to each time frame, and integrate all scores into a single time series score set, which covers the score data of the entire interrogation period.

[0132] Sort all scores in the score set in ascending order of value, and take the middle value after sorting as the median. If the total number of frames is odd, take the median value directly; if it is even, take the average of the two middle values ​​to ensure that the benchmark value can represent the normal score level of most frames.

[0133] First, calculate the absolute difference between each score and the median (to eliminate the problem of positive and negative biases canceling each other out), obtaining all absolute differences. Then, sort all the absolute differences by value and take the middle value as the median absolute deviation. The larger the median absolute deviation, the more significant the fluctuation in the score data, and vice versa.

[0134] The consistency adjustment coefficient is fixed at 1.4826. 1.4826 is a statistically accepted constant that can convert MAD into a dispersion index comparable to the standard deviation of a normal distribution, thus avoiding threshold scale confusion caused by differences in data distribution.

[0135] Set a threshold coefficient of positive real number type. The larger the threshold coefficient, the higher the final threshold (the stricter the scoring requirements, the lower the false detection rate but the possibility of missed detection); the smaller the coefficient, the lower the threshold (the sensitivity is improved but false detection is possible), to adapt to the detection needs of different scenarios.

[0136] First, multiply the threshold coefficient by the consistency adjustment coefficient to obtain the first product (achieving a combination of sensitivity and scale).

[0137] Multiply the first product by the absolute deviation of the median to obtain the second product (which determines the dynamic adjustment range of the threshold; the greater the dispersion, the greater the adjustment range).

[0138] Finally, the median is added to the second product to obtain the adaptive threshold. The adaptive threshold changes dynamically with the data distribution, covering normal score fluctuations while capturing significant abnormal scores.

[0139] Set the peak neighborhood radius frame number (e.g., 1 or 2 frames) as a positive integer to define the range of the local neighborhood; determine the peak neighborhood relative displacement set based on the radius: the value range is from the negative radius frame number to the positive radius frame number, and does not include zero, that is, it only includes the neighboring frames around the current frame (e.g., when the radius is 1, the set is {-1,1}), and does not include the current frame itself, to provide the neighborhood range for subsequent local maxima determination.

[0140] First determination: Determine whether the current frame score is greater than or equal to the adaptive threshold, ensuring that the frame score is significantly higher than the normal level, and excluding normal fluctuations;

[0141] Second determination: Determine whether the current frame score is greater than or equal to the frame scores corresponding to all displacements in the neighborhood displacement set, ensuring that the frame is the highest score (local peak) in the local neighborhood, and excluding non-peak frames in continuous high score regions;

[0142] Ultimately, the time frame that simultaneously meets both judgment conditions is identified as the psychologically sensitive moment during interrogation.

[0143] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.

Claims

1. A method for detecting psychologically sensitive points during interrogation based on a temporal attention mechanism, characterized in that, include: The spatial covariance matrix time series is calculated within a local time window using short-time Fourier transform; the relative displacement weighting kernel is extracted from the attention matrix output by the temporal attention model. Based on affine invariance metrics, a first tangent vector corresponding to the time-varying room impact response and a second tangent vector corresponding to attention weighting are constructed at the center frame. The curvature density index of the two-flow section spanned by the first and second tangent vectors is calculated. Based on the curvature density index of the two-flow section and the gating coefficient, a gating weight is generated to gating the relative displacement weighting kernel. After summing in the tangent space, the geometrically aligned prediction space covariance matrix is ​​obtained through exponential mapping. The residuals between the actual spatial covariance matrix and the geometrically aligned predicted spatial covariance matrix are calculated using geodesic distance with affine invariant metric, and the residuals are weighted according to energy ratio to obtain a single time series score. An adaptive threshold is constructed using the median absolute deviation of a single time series score and a threshold coefficient. The psychologically sensitive moments during interrogation are then output based on the adaptive threshold.

2. The interrogation psychological sensitivity detection method based on temporal attention mechanism according to claim 1, characterized in that, The spatial covariance matrix time series is calculated within a local time window using short-time Fourier transform, including: S21, for the discrete-time signal of each microphone channel, selecting continuous discrete-time sample segments according to a preset window length, applying a window function to the discrete-time sample segments, and performing Fourier operation to obtain the time-frequency signals of each microphone channel corresponding to different frames and different frequency points; combining the time-frequency signals of all microphone channels under the same frame and the same frequency point to form a multi-channel time-frequency vector; S22, setting the center frame and the number of frames within the window radius, and determining the number of frames contained within the local time window based on the number of frames within the window radius, which is twice the number of frames within the window radius plus one; Using the center frame as a reference, select all frames from the center frame minus the window radius frame number to the center frame plus the window radius frame number to form a local time window; S23, for each center frame and each frequency point, within the corresponding local time window, perform conjugate transpose operation on the multi-channel time-frequency vector of each frame and multiply it with the multi-channel time-frequency vector itself to obtain the vector outer product corresponding to each frame; S24, calculate the average of the vector outer products of all frames within the local time window to obtain the spatial covariance matrix corresponding to the center frame and the frequency point; S25, change the center frame sequentially and repeat S22 to S24 to obtain the time series of the spatial covariance matrix that varies with the center frame.

3. The interrogation psychological sensitivity point detection method based on temporal attention mechanism according to claim 2, characterized in that, Extracting relative displacement weighting kernels from the attention matrix output by the temporal attention model includes: determining the set of relative displacement values ​​based on the window radius frame number; wherein the range of the relative displacement value set is from negative to positive window radius frame number; setting an equivalence discrimination function for each relative displacement in the set of relative displacement values; wherein the equivalence discrimination function takes a value of 1 when the input value is zero and a value of 0 when the input value is not zero; and for all elements in the temporal attention matrix whose row and column time frame positions both belong to the local time window, using equivalence... The discriminant function filters elements whose difference between the row time frame position and the column time frame position equals the current relative displacement. The filtered elements are summed to obtain the unnormalized cumulative relative displacement corresponding to the relative displacement. All elements in the temporal attention matrix whose row time frame position and column time frame position belong to the local time window are summed to obtain the normalization factor. The unnormalized cumulative relative displacement corresponding to each relative displacement is divided by the normalization factor to obtain the weight corresponding to each relative displacement. All relative displacements and their corresponding weights together constitute the relative displacement weighting kernel.

4. The interrogation psychological sensitivity detection method based on temporal attention mechanism according to claim 3, characterized in that, Based on the affine invariant metric, a first tangent vector corresponding to the time-varying room impact response and a second tangent vector corresponding to the attention weighting are constructed at the center frame, including: S41, for each discrete frequency point, taking a first symmetric matrix and a second symmetric matrix with the same dimension as the spatial covariance matrix; inverting the first spatial covariance matrix at the center frame, and performing multiplication operations sequentially with the first symmetric matrix, the inverted first spatial covariance matrix at the center frame, and the second symmetric matrix; S42, taking the trace of the result of the multiplication operation to form the affine invariant metric; S43, taking the first spatial covariance matrix at the center frame and any spatial covariance matrix in the neighborhood of the center frame; performing a square root operation on the first spatial covariance matrix at the center frame, and then inverting the square root matrix; S44, performing multiplication operations sequentially with the spatial covariance matrix in the neighborhood and the square root matrix of the inverted matrix; taking the matrix logarithm of the result of the multiplication operation, and then multiplying the matrix logarithm with the square root matrix... S45: Multiply the first spatial covariance matrix at the center frame to form a logarithmic mapping; S46: Set the time step frame number as a positive integer; S47: Take the first spatial covariance matrix at the center frame and the second spatial covariance matrix after delaying the center frame by the time step frame number; Process the first and second spatial covariance matrices through the logarithmic mapping in S44 to obtain the fused mapping result; Divide the fused mapping result by the time step frame number to obtain the first tangent vector corresponding to the time-varying room impact response; S48: Take the relative displacement weighting kernel, and for each relative displacement, take the third spatial covariance matrix after delaying the relative displacement in the center frame; S49: Process the third spatial covariance matrix through the logarithmic mapping in S44 to obtain the mapping result corresponding to each relative displacement; Multiply the mapping result corresponding to each relative displacement with the weight of the relative displacement in the relative displacement weighting kernel; Summate all the multiplication results to obtain the second tangent vector corresponding to the attention weighting.

5. The interrogation psychological sensitivity detection method based on temporal attention mechanism according to claim 4, characterized in that, The calculation of the curvature density index of the two-stream cross-section spanned by the first and second tangent vectors includes: calculating the first inner product of the second tangent vector and the affine invariant metric, the second inner product of the first tangent vector and the affine invariant metric, and the fused inner product of the second tangent vector and the first tangent vector; multiplying the first inner product by the second inner product, then subtracting the square of the fused inner product, and taking the square root of the result to obtain the area element; taking the first spatial covariance matrix at the center frame and inverting it, then multiplying the inverted matrix by the second tangent vector and the first tangent vector to obtain the first product matrix and the second product matrix; performing a commutator operation on the first product matrix and the second product matrix to obtain the commutator operation result; calculating the Frobenius norm of the commutator operation result, multiplying the square of the Frobenius norm by one-quarter to obtain the numerator; using the square of the area element as the denominator; dividing the numerator by the denominator and taking the opposite number to obtain the cross-sectional curvature; multiplying the cross-sectional curvature by the area element and taking the opposite number to obtain the curvature density index of the two-stream cross-section.

6. The interrogation psychological sensitivity detection method based on temporal attention mechanism according to claim 5, characterized in that, Gating weights are generated based on the dual-flow cross-section curvature density index and gating coefficients. The relative displacement weighting kernel is then gated. After tangent space summation, an exponential mapping is used to obtain the geometrically aligned prediction space covariance matrix. This process includes: setting positive real-valued gating coefficients; for each relative displacement in the set of relative displacement values, taking the dual-flow cross-section curvature density index corresponding to the frame after delaying that relative displacement as the target dual-flow cross-section curvature density index; multiplying the gating coefficients by the target dual-flow cross-section curvature density index, taking the negative of the product as the exponent of the exponential function, and calculating the result of the exponential function to obtain the gating weight corresponding to that relative displacement; multiplying the weight of each relative displacement in the relative displacement weighting kernel by the corresponding gating weight to obtain the weight product of the corresponding relative displacement; summing the weight products of all relative displacements to obtain the gating normalization factor; dividing the weight product of each relative displacement by the gating normalization factor to obtain the gating weight corresponding to each relative displacement. All relative displacements and their corresponding gating weights are then calculated. The gating weights together constitute the gating relative displacement weighting kernel. For each relative displacement, the first spatial covariance matrix at the center frame and the third spatial covariance matrix after delaying the relative displacement at the center frame are taken. The first and third spatial covariance matrices are processed by S44 logarithmic mapping to obtain the logarithmic mapping result corresponding to the relative displacement. The logarithmic mapping result corresponding to each relative displacement is multiplied by the weight of the relative displacement in the gating relative displacement weighting kernel, and all multiplication results are summed to obtain the tangent space summation. The first spatial covariance matrix at the center frame is taken, and the square root operation is performed on the first spatial covariance matrix to obtain the square root matrix. The square root matrix is ​​inverted to obtain the square root inverse matrix. The tangent space summation is multiplied left and right by the square root inverse matrix to obtain the intermediate matrix. The intermediate matrix is ​​subjected to matrix exponentiation to obtain the exponent matrix. The square root matrix is ​​multiplied left and right by the exponent matrix to obtain the geometrically aligned prediction spatial covariance matrix.

7. The interrogation psychological sensitivity detection method based on temporal attention mechanism according to claim 6, characterized in that, The residuals between the actual spatial covariance matrix and the geometrically aligned predicted spatial covariance matrix are calculated using geodesic distance with affine invariant metric. The residuals are then weighted according to energy ratios to obtain a single time series score. This process includes: obtaining the geometrically aligned predicted spatial covariance matrix and its corresponding actual spatial covariance matrix in the spatial covariance matrix time series; performing a square root operation on the actual spatial covariance matrix, then inverting the square root matrix to obtain the square root inverse matrix; left-multiplying the square root inverse matrix by the predicted spatial covariance matrix and right-multiplying it by the square root inverse matrix to obtain an intermediate matrix; performing a matrix logarithm operation on the intermediate matrix to obtain a logarithmic matrix; and calculating the F-value of the logarithmic matrix. The Robenius norm is used to obtain the geodesic distance residual. For each discrete frequency point, the L2 norm of the multi-channel time-frequency vector of each frame within the local time window is calculated. The L2 norms of all frames within the local time window are summed to obtain the energy value of that frequency point. The energy values ​​of all discrete frequency points are summed to obtain the total energy value. The energy value of each frequency point is divided by the total energy value to obtain the energy proportion of that frequency point. The geodesic distance residual of each discrete frequency point is multiplied by the energy proportion of that frequency point to obtain the residual weight value of each frequency point. The residual weight values ​​of all discrete frequency points are summed to obtain the single time series score of the corresponding center frame.

8. The interrogation psychological sensitivity detection method based on temporal attention mechanism according to claim 7, characterized in that, An adaptive threshold is constructed using the median absolute deviation of individual time-series scores and a threshold coefficient. Based on this adaptive threshold, the psychologically sensitive moments during interrogation are output. The process includes: acquiring individual time-series scores for all time frames to form a set of individual time-series scores; sorting all individual time-series scores in the set by numerical value and taking the middle value as the median; calculating the absolute difference between each individual time-series score and the median, obtaining all absolute differences; sorting all absolute differences by numerical value and taking the middle value as the median absolute deviation; setting a consistency adjustment coefficient of 1.4826; setting a threshold coefficient of positive real numbers; and multiplying the threshold coefficient by the consistency adjustment coefficient. Obtain the first product; multiply the first product by the absolute deviation of the median to obtain the second product; add the median to the second product to obtain the adaptive threshold; set the peak neighborhood radius frame number as a positive integer; determine the peak neighborhood relative displacement set based on the peak neighborhood radius frame number, the value range of which is from the negative peak neighborhood radius frame number to the positive peak neighborhood radius frame number, and does not include zero; first determination: whether the single time series score of each time frame is greater than or equal to the adaptive threshold; second determination: whether the single time series score of each time frame is greater than or equal to the single time series score corresponding to each frame after the relative displacement in the peak neighborhood relative displacement set of that time frame; determine the time frames that simultaneously satisfy the first and second determinations as psychologically sensitive moments for interrogation.

Citation Information

Patent Citations

  • Multi-dimensional time sequence anomaly detection method based on time variable double-attention mechanism

    CN118378139A