Frightening and rapid alarm method based on multi-mode voice

By using multi-channel audio signal processing and a temporal scenario deterioration prediction model, the problems of high false alarm rate and high false alarm rate in existing technologies have been solved. This enables proactive risk warning and dynamic deterrence in complex acoustic environments, improving the timeliness and proactivity of personal safety protection.

CN121483280APending Publication Date: 2026-02-06ZHONGAO ENVIRONMENTAL TECHNOLOGY (JIANGSU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511626543.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing personal safety protection technologies have high false alarm and false alarm rates in complex acoustic environments, cannot effectively cope with stressful situations that are not accompanied by significant acoustic energy bursts, and lack the ability to model and predict the dynamic processes of dangerous situations, resulting in the inability to provide timely and proactive warnings and deterrence.

Method used

By acquiring multi-channel audio signals, separating the voice streams of the device holder and the conversation partner, extracting multi-dimensional dialogue dynamics features, generating real-time situation deterioration scores using a temporal situation deterioration prediction model, and executing graded response strategies, including sound source direction finding and dynamic deterrence operations.

Benefits of technology

It reduces false alarm and false alarm rates in complex acoustic environments, provides proactive risk warnings, enhances the initiative and timeliness of personal safety protection, and effectively deters inappropriate behavior through dynamic deterrence operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121483280A_ABST
    Figure CN121483280A_ABST
Patent Text Reader

Abstract

The invention discloses a frightening and rapid alarm method based on multi-modal voice, and the method comprises the steps: collecting a multi-channel audio signal in an environment, and carrying out the processing and separation of independent voice streams of an equipment holder and a dialogue party; extracting multi-dimensional dialogue dynamics features including rhythm layer features from the separated voice stream according to a preset time step length to form a time sequence feature vector sequence; inputting the sequence into a pre-trained time sequence situation deterioration prediction model, and generating a real-time situation deterioration score representing the current situation danger level; according to the real-time situation deterioration score, a preset hierarchical response strategy is executed, and the strategy comprises dynamic frightening or automatic alarm operation; according to the method, through time sequence modeling of dialogue dynamics, conversion from static event detection to dynamic situation prediction is realized, risks can be pre-warned in a prospective manner, and initiative, accuracy and effectiveness of personal safety protection are remarkably improved in combination with multi-dimensional feature analysis and a hierarchical response mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent voice control technology, and in particular to a deterrence and rapid alarm method based on multimodal voice. Background Technology

[0002] With the rapid development of Internet of Things (IoT) and Artificial Intelligence (AI) technologies, personal safety protection technologies are evolving from traditional passive, physically triggered alarms to proactive, intelligent risk warnings based on multimodal perception. Current mainstream personal safety solutions (such as emergency applications or emergency contact functions) generally rely on users' proactive physical intervention when encountering danger. However, in many coercive situations, victims often lose the ability to actively seek help due to physical limitations or extreme psychological stress (e.g., high fear or anxiety), thus delaying crucial intervention opportunities. To address this pain point, academia and industry have begun exploring the use of wearable devices or smart terminals to continuously monitor users' physiological and environmental data, aiming to achieve automatic alarms without intervention.

[0003] Despite this, current acoustic analysis-based automatic alarm technologies still have significant limitations. On the one hand, existing technologies mainly rely on threshold judgments of single acoustic features, such as triggering alarms by detecting specific high-energy, high-fundamental-frequency sound events like screams and cries. Such methods are highly dependent on the signal-to-noise ratio (SNR), making it difficult to control both false alarm and false negative rates in complex acoustic environments (such as streets and public places), and they cannot effectively address coercive situations without significant acoustic energy bursts, such as "silent violence" or covert threats. On the other hand, while some studies have introduced emotion recognition, most remain at the level of static emotion classification of isolated utterances (such as anger and fear), failing to effectively utilize contextual information. This approach ignores the fact that dangerous situations are often a dynamic evolutionary process, with key clues hidden in the interaction patterns of the two parties, the shifts in power dynamics, and the subtle discrepancies between acoustic performance and semantic content. Consequently, there is a general lack of modeling and prediction capabilities for the dynamic process of "situation deterioration," making it impossible to provide forward-looking warnings and deterrence before danger escalates into substantial harm. Therefore, there are technical bottlenecks that urgently need to be addressed in terms of the timeliness and proactivity of personal safety protection. Summary of the Invention

[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0005] In view of the aforementioned existing problems, this invention is proposed. Therefore, this invention provides a deterrence and rapid alarm method based on multimodal voice to solve the problems mentioned in the background art.

[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a method for deterrence and rapid alarm based on multimodal voice, comprising: Acquire multi-channel audio signals from the environment and process the multi-channel audio signals to separate multiple independent audio streams, each containing at least the device holder's voice stream and at least one other party's voice stream. For the voice stream of the device holder and the voice stream of the dialogue party, multi-dimensional dialogue dynamic features are extracted at a preset time step to form a temporal feature vector sequence. The multi-dimensional dialogue dynamic features include at least prosodic layer features that reflect the rhythm of dialogue interaction. The sequence of temporal feature vectors is input into a pre-trained temporal situation deterioration prediction model to generate a real-time situation deterioration score that characterizes the degree of danger of the current situation. Based on the real-time situation deterioration score, a preset graded response strategy is executed, which includes deterrence or alarm operations.

[0007] As a preferred embodiment of the multimodal voice-based deterrence and rapid alarm method of the present invention, the multi-channel audio signal is processed, including: The multi-channel audio signal is enhanced by spatial filtering to suppress ambient noise, and the enhanced multi-channel audio signal is processed by a deep neural network to separate overlapping speech. Furthermore, by comparing the separated overlapping voiceprint features with the pre-stored device owner voiceprint features, the device owner's voice stream and the dialogue party's voice stream are identified and marked.

[0008] As a preferred embodiment of the multimodal speech-based deterrence and rapid alarm method described in this invention, the multi-dimensional dialogue dynamics features further include physiological acoustic layer features, which are used to quantify the physiological tension state of the speaker.

[0009] In a preferred embodiment of the multimodal speech-based deterrence and rapid alarm method of the present invention, the physiological acoustic layer features are extracted through the speech stream of the device holder and / or the speech stream of the conversation partner, and the physiological acoustic layer features include: The parameter characterizing the fundamental frequency stability is obtained by normalizing the fundamental frequency of the corresponding speech stream through continuous variation. And parameters that reflect the distribution of high-frequency harmonic energy in the corresponding speech stream.

[0010] As a preferred embodiment of the multimodal speech-based deterrence and rapid alarm method described in this invention, the prosodic layer features include: Used to measure the interruption rate or the frequency of speech turn switching in a conversation; In addition, a parameter characterizing the degree of cooperation between the two parties in the dialogue is obtained by calculating the cross-correlation between the device holder's speech stream and the speech stream of the other party in the dialogue.

[0011] As a preferred embodiment of the multimodal speech-based deterrence and rapid alarm method described in this invention, the multi-dimensional dialogue dynamics feature further includes a semantic-acoustic inconsistency feature. This semantic-acoustic inconsistency feature is calculated using the device holder's speech stream or the dialogue party's speech stream, and its calculation method is as follows: The first and second emotional state distributions are inferred from the text content and acoustic properties of the speech stream, respectively, and a metric is calculated to measure the difference between the first and second emotional state distributions.

[0012] As a preferred embodiment of the multimodal speech-based deterrence and rapid alarm method described in this invention, the temporal context deterioration prediction model is a temporal model with a hierarchical attention mechanism, and this model is configured as follows: The features within local time segments of the temporal feature vector sequence are weighted and aggregated to generate a segment-level representation; The multiple fragment-level representations are weighted and aggregated again to generate a vector representing the global dialogue history, and the real-time context deterioration score is generated based on this vector of global dialogue history.

[0013] As a preferred embodiment of the multimodal voice-based deterrence and rapid alarm method described in this invention, the method further includes: When the real-time situation deterioration score exceeds a first preset threshold, the sound source direction finding step is activated. This step uses the collected multi-channel audio signals to estimate and track the sound source location of the speech stream of the dialogue party. Furthermore, the changing trend of the sound source location is used as supplementary information to adjust the real-time situation deterioration score or influence the decision-making of the graded response strategy.

[0014] As a preferred embodiment of the multimodal speech-based deterrence and rapid alarm method of the present invention, the sound source direction finding step includes: Based on the phase information of the cross-correlation function between multi-channel audio signals, preliminary observations of the sound source's location are estimated. Furthermore, the Kalman filter is used to process the continuous observations of the sound source's azimuth, outputting a smooth and robust azimuth tracking result.

[0015] As a preferred embodiment of the multimodal voice-based deterrence and rapid alarm method described in this invention, the hierarchical response strategy includes: When the real-time situation deterioration score exceeds the second preset threshold, a dynamic deterrence operation is triggered. The dynamic deterrence operation includes: capturing a segment of the conversational speaker's voice, using a voice conversion technology that changes the timbre while keeping the voice content unchanged, and combining it with a preset warning text to generate and play a targeted voice warning.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By constructing multi-dimensional dialogue dynamics features and using a temporal model to model the situation evolution process, it is possible to predict the trend of danger occurrence, rather than just detecting isolated danger signals (such as screams), thus achieving proactive risk warning and improving the initiative and timeliness of personal safety protection. 2. By introducing multi-channel audio processing technology (beamforming and deep learning source separation), the interference of environmental noise and human voice overlap is effectively suppressed. At the same time, the fusion analysis of multi-dimensional features can more comprehensively and accurately assess the situation compared with traditional methods that rely on a single threshold, significantly reducing the false alarm rate and false negative rate. 3. This invention designs a graded response strategy that matches the level of danger, especially the dynamic deterrence operation. By using the other party's voice to change the timbre and generate a warning, it has a stronger psychological deterrent and is more targeted than a simple alarm sound. It can effectively stop inappropriate behavior without escalating the situation. At the same time, it combines sound source direction finding to track the dynamics of physical space, making the risk assessment more three-dimensional and reliable. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating the overall process of a multimodal voice-based deterrence and rapid alarm method according to an embodiment of the present invention. Detailed Implementation

[0018] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0019] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0020] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0021] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0022] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0023] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0024] Example 1 Reference Figure 1This is the first embodiment of the present invention, which provides a deterrence and rapid alarm method based on multimodal voice, including: S1. Acquire multi-channel audio signals from the environment and process the multi-channel audio signals to separate multiple independent audio streams, each containing at least the device owner's voice stream and at least one other party's voice stream. It should be noted that this step aims to accurately and stably separate the speech of each party participating in the dialogue from a complex acoustic environment, so as to provide high-quality, distinguishable input information for subsequent dialogue dynamics analysis; Furthermore, the present invention operates on a smart terminal device (such as a smartphone, smartwatch, etc.) equipped with a microphone array to capture acoustic data, the microphone array containing at least two microphones (number of channels). It can also simultaneously acquire acoustic signals from the environment to form multi-channel audio signals; Furthermore, to effectively suppress the interference of environmental noise and room reverberation on subsequent speech processing, it is necessary to preprocess the acquired raw multi-channel audio signals. In one feasible implementation of the present invention, a non-directional noise suppression algorithm, such as spectral subtraction or Wiener filtering, can be used to perform preliminary noise reduction on all channels. Subsequently, a microphone array is used to estimate the sound source direction of the pre-denoised signal to locate active speakers in the environment. For example, the sound source direction can be estimated by calculating the generalized cross-correlation function between the multi-channel signals. After determining the approximate directions of the device holder and at least one interlocutor, the present invention employs adaptive Minimum Variance Distortionless Response (MVDR) beamforming technology. At this point, one direction can be used as the target direction, and other directions as interference directions. The weight vector of the beamformer is dynamically adjusted to enhance the speech signals from different speaker directions and suppress noise and interference from other directions. The optimal weight vector of the MVDR beamformer is... It can be calculated using the following formula: in, This represents the complex weight vector used for weighted summation of multi-channel signals. The noise covariance matrix is ​​used to describe the correlation of noise signals across different microphone channels. This matrix is ​​typically estimated and updated by detecting signals during gaps in speech activity (i.e., segments where no one is speaking). Indicates the direction of the target speaker. The steering vector represents the phase delay relationship generated in each channel when a plane wave signal from this direction arrives at the microphone array; The superscript indicates the conjugate transpose operation. Superscripts indicate the inverse operation of a matrix; It should be noted that by applying the optimal weight vector of the MVDR beamformer to the original multi-channel audio signal, one or more enhanced audio signals with significantly improved signal-to-noise ratios, pointing to different speakers, can be obtained, providing high-quality input for subsequent separation and recognition. Furthermore, since the enhanced audio signal may still contain speech overlap caused by the device holder and the other party speaking at the same time (i.e., the "cocktail party effect"), in order to solve this problem, this invention uses a deep neural network model to perform blind source separation (BSS) on the overlapping speech. Specifically, a deep separation network based on time-frequency domain masks can be used. In one feasible implementation of the present invention, the deep separation network based on time-frequency domain masks uses a convolutional neural network with a U-Net structure. The workflow of this network is as follows: (i) The main audio signal after beamforming enhancement is converted to the time-frequency domain by short-time Fourier transform (STFT) to obtain a complex spectrum; (ii) Input the complex spectrum into a pre-trained deep separation network, which predicts a corresponding time-frequency mask for each sound source to be separated (e.g., device holder and conversation partner). The mask is a matrix of the same size as the spectrum, with element values ​​between 0 and 1, indicating the probability that the time-frequency unit belongs to a specific sound source. (iii) The original mixed spectrogram is multiplied element-wise with each predicted mask to obtain the estimated spectrogram of each sound source. Then, the inverse short-time Fourier transform (iSTFT) is performed on all the estimated spectrograms to reconstruct the clean and independent speech signals that are separated in the time domain. At the same time, in order to solve the permutation ambiguity problem inherent in blind source separation, the deep separation network preferably adopts the permutation invariant training (PIT) strategy for training. It should be noted that after separating multiple independent voice signals through the above steps, it is still necessary to determine which one is the voice of the device holder. In one feasible implementation of the present invention, voiceprint recognition technology is used to distinguish the voice of the device holder, as follows: Device owners need to register their voiceprints on the device beforehand. During this process, the device guides the user to read a specified text and extracts a unique voiceprint feature vector from their speech, which can uniquely identify them—that is, the device owner's voiceprint feature. This vector can be extracted using advanced deep learning models such as x-vector or d-vector and is securely stored locally on the device. For each speech stream separated from the data, a real-time voiceprint feature vector is extracted using the same voiceprint model as during registration. Subsequently, the identity is determined by calculating the similarity between each real-time voiceprint vector and the pre-stored voiceprint vector of the holder, where the similarity is measured using cosine similarity. in, This represents the dot product operation of vectors. The Euclidean norm (or L2 norm) of a vector is used to represent the vector. Specifically, the higher the similarity score, the greater the likelihood that the two voiceprints belong to the same person. Therefore, the voice stream with the highest similarity score that exceeds the preset judgment threshold can be marked as the "device owner voice stream", while other voice streams are marked as the "speaker voice stream". S2. Extract multi-dimensional dialogue dynamics features from the device holder's voice stream and the dialogue party's voice stream at a preset time step to form a temporal feature vector sequence. The multi-dimensional dialogue dynamics features include at least prosodic layer features that reflect the rhythm of dialogue interaction. It should be noted that this step is a crucial step in realizing the transformation from raw speech stream to contextual understanding. The core idea of ​​this step is to no longer treat speech as merely an independent acoustic event, but rather to place it within the dynamic interactive framework of dialogue. This involves extracting deep features that reflect the relationship, intentions, and emotional state evolution of the two parties in the dialogue, and organizing these features into a vector sequence that evolves over time, providing structured input for subsequent temporal model prediction. In this process, feature extraction is performed at a fixed time step (e.g., 1 second or 2 seconds), and within each time step, the speech stream for that time period is analyzed to generate a feature vector for that time period. Furthermore, in one implementable process of the present invention, the multi-dimensional dialogue dynamics features are composed of the following features, which collectively characterize the dialogue context from different perspectives: Physiological acoustic layer features: These features aim to quantify changes in acoustic parameters that are difficult to fake, caused by changes in the physiological state of the speaker (especially autonomic nervous system activity related to stress, tension, and fear). They are obtained by extracting parameters that characterize fundamental frequency stability and parameters that reflect the distribution of high-frequency harmonic energy. Specifically, the parameter characterizing fundamental frequency stability is set as a fundamental frequency, which is the frequency of vocal cord vibration. It is a sensitive indicator of emotions and stress. Because a person's ability to control the fine motor skills of the throat muscles decreases when under high stress, involuntary micro-vibrations in the fundamental frequency may occur. Based on this, this invention captures this instability through a continuously varying normalized calculation. Specifically, firstly, the fundamental frequency value is extracted from each speech frame to form a fundamental frequency profile sequence. Subsequently, the first-order difference of the local fundamental frequency profile is calculated and normalized to obtain the normalized standard deviation of the fundamental frequency change rate. : in, It is the first The fundamental frequency value of each audio frame. It is the absolute difference in fundamental frequencies between adjacent audio frames. It is all The average value, It is the average fundamental frequency of this speech segment. It is the number of audio frames within the audio segment; It should be noted that the normalized standard deviation of the fundamental frequency change rate is usually positively correlated with physiological arousal and stress levels; Specifically, regarding parameters reflecting the distribution of high-frequency harmonic energy, when a person feels tense or attempts to forcefully speak, the vocal cords close more abruptly and forcefully, resulting in richer high-frequency harmonic components in the sound spectrum. Based on this, the spectral tilt can be quantified by calculating the spectral tilt, which describes the rate at which spectral energy decays with increasing frequency. Under stress, the glottal closure is more rapid, resulting in relatively richer high-frequency energy and a reduced spectral tilt (or a gentler slope). In one feasible implementation of this invention, the spectral tilt can be obtained through linear predictive coding (LPC) analysis of the speech signal. That is, a set of linear prediction coefficients is obtained through LPC analysis, and the spectral tilt can be approximated by the first reflection coefficient or directly obtained through the autocorrelation function. and The calculation yielded: in, It is the first value of the short-time autocorrelation function of the speech signal, a spectral tilt that is closer to 0. A flatter spectrum indicates a stronger state of tension; It should be noted that this spectral tilt can be extracted separately from the device holder's voice stream and / or the voice stream of the other party in the conversation, thereby allowing a comparison of the tension levels between the two parties; Prosodic layer features: This feature focuses on the rhythm of the dialogue, that is, the speech rate, pitch, volume and rhythmic patterns beyond individual words, and pays special attention to the interaction between the two parties in these dimensions. These features are key to judging whether the dialogue is collaborative or adversarial. Similarly, they are obtained by extracting parameters to measure the dominance of the dialogue and parameters to characterize the degree of collaboration between the two parties. Specifically, the parameter used to measure conversational dominance includes two sub-parameters: interruption rate and speaking turn switching frequency. Specifically, interruption rate refers to the frequency with which one party (such as the speaker) forcibly inserts their own speech before the other party (such as the device owner) has finished speaking within a time window. It should be emphasized that a high interruption rate, especially a one-way high interruption rate, is a typical example of dialogue dominance and aggressive behavior. Specifically, the interruption rate calculation relies on Voice Activity Detection (VAD) of the device owner's voice stream and the dialogue party's voice stream. First, VAD is performed on the two voice streams in parallel to obtain their respective voice activity timestamp sequences (including the start and end times of each voice segment). Then, within a sliding time window, the start timestamps of the dialogue party's voice are traversed. If the timestamp falls within any voice activity interval of the device owner (i.e., later than the start time of the interval but earlier than its end time), it is recorded as an interruption event. The interruption rate is the number of interruption events that occur per unit time. Specifically, the frequency of turn-taking refers to the frequency at which the two parties in a conversation exchange speaking power. An abnormally high (interruption) or abnormally low (one party is silent or the other is talking incessantly) frequency of turn-taking may indicate an abnormal conversation state. Specifically, regarding parameters characterizing the degree of cooperation between dialogue participants, healthy dialogue typically exhibits a certain degree of acoustic coupling or prosodic convergence, meaning that the two participants unconsciously imitate and converge with each other in terms of speech rate, pitch range, or volume. Based on this, this invention quantifies this degree of cooperation by calculating the cross-correlation between specific acoustic feature profiles. That is, the energy profiles or reference profiles of the device holder's speech stream and the dialogue participant's speech stream are extracted respectively, and the cross-correlation function of these two time-series profiles within a certain delay range is calculated: in, and These are the acoustic feature profiles of the device owner and the other party in the dialogue, respectively. It's a time delay. This represents the expected value (usually replaced by a time average). It should be noted that the peak size and peak position of the cross-correlation function can be used as features. A significant peak with near-zero delay indicates strong acoustic synchronization between the two sides, which is a sign of high cooperation. Conversely, a flat cross-correlation function or one without obvious peaks means that the two sides are out of rhythm, which may indicate antagonism or indifference. Semantic-acoustic inconsistency feature: This feature aims to capture dangerous situations such as "hidden meanings" or "smiling while harboring malice," where the speaker's acoustically expressed emotion contradicts the literal meaning of their utterance; specifically, this feature is calculated as follows: First, automatic speech recognition (ASR) is performed on the speech stream (from the device owner or the speaker) to obtain its text content. Then, this text is input into a pre-trained text sentiment analysis model (such as a model based on BERT or other Transformer architectures) to output a first sentiment state distribution. The distribution of this emotional state These represent the probabilities that the text content is positive, neutral, or negative, respectively. Simultaneously, the acoustic features of the speech stream (such as MFCC, pitch, energy, etc.) are input into a pre-trained speech emotion recognition (SER) model, which outputs a second emotion state distribution. The distribution of this emotional state These represent the probabilities of perceiving a positive, neutral, or negative emotion from a sound. By employing KL divergence, the difference between the distributions of sentiment probabilities (positive, neutral, negative) in the text content and the audio is calculated, yielding the following: It should be noted that a high KL divergence value means that there is a significant mismatch between the emotion expressed in the text and the emotion conveyed acoustically (for example, saying "I'm fine" in a very angry tone), which is considered an extremely important danger signal in coercive situations. Furthermore, by calculating all the above features within each preset time step, a high-dimensional feature vector can be generated. Arranging all the feature vectors in chronological order forms a temporal feature vector sequence, which fully records the dynamic evolution of the dialogue from the beginning to the current moment. Furthermore, when there are multiple parties in the dialogue, the solution of the present invention can adopt one of the following strategies for processing: Strategy 1: Merge all non-device owner voice streams into a single "dialogue group" voice stream and extract overall dialogue dynamics features from it to assess the overall adversarial atmosphere; Strategy 2: Extract features from each speaker's voice stream individually and combine them with their interactions with the device owner to calculate their respective threat contribution. The final real-time context deterioration score can be the maximum value or a weighted sum of individual threat contributions. It should be noted that in one feasible process of the present invention, strategy 2 is preferred, and the dialogue party that poses the greatest threat to the device holder (e.g., the highest interruption rate, the most intense acoustic physiological characteristics) is taken as the main analysis object. S3. Input the sequence of temporal feature vectors into a pre-trained temporal situation deterioration prediction model to generate a real-time situation deterioration score that represents the degree of danger of the current situation. It should be noted that this step receives the temporal feature vector sequence generated in step S2, which contains rich dialogue dynamics information, and uses an advanced deep learning model to perform real-time and dynamic quantitative assessment of the danger level of the situation. In this process, it is necessary not only to identify whether the situation is dangerous at the current moment, but also to predict the deterioration trend of the situation by learning from historical evolution. Preferably, the temporal context deterioration prediction model used in this invention is a temporal model with a hierarchical attention mechanism (Hierarchical Attention Network, HAN). This model structure is particularly suitable for processing long sequence data such as dialogues because it can simulate the human cognitive process: first focus on local key points at the phrase or sentence level, and then integrate these local information to form a global understanding of the entire text or dialogue. This hierarchical information processing method enables the model to effectively capture the multi-scale dynamics of dialogue, from short-term interaction patterns to long-term evolution trends. Specifically, the configuration and workflow of this model are as follows: Assume that the time-series feature vector sequence generated in step S2 is ,in, It is the first Each time step 3D feature vectors It is the total length of the vector sequence; S301, Context coding of local time segments: The entire temporal feature vector sequence is divided into multiple continuous, non-overlapping (or partially overlapping) local time segments; that is, each... Each consecutive time step is divided into a segment, the first... Each segment can be represented as ; For each segment The model uses a bidirectional gated recurrent unit (Bi-GRU) to perform context encoding on its internal feature vectors. Because Bi-GRU can process sequences from both the forward and backward directions simultaneously, it generates a hidden state for each time step feature vector that contains its left and right context information. For fragments any eigenvector in The calculation process of Bi-GRU is as follows: Forward GRU: ; Backward GRU: ; Hidden splicing state: ; in, and These are the forward and backward GRUs at time steps. The hidden state, This indicates a vector concatenation operation; It should be noted that, after Bi-GRU processing, in this way, It was encoded. Complete local context information within its segment; S302. Generate fragment-level representations through a local attention mechanism: Obtain the context encoding for all time steps within each segment. Then, by applying an attention mechanism to the model to dynamically assign weights to these encodings, a fragment-level representation vector that can represent the core information of the fragment is generated. This allows the model to automatically identify and focus on the "highlight moments" within a segment that are most likely to indicate danger; the calculation process for local attention in this process is as follows: First, each hidden state is computed using a small feedforward neural network. Importance score : in, It is a weight matrix. It is a bias vector. It is the transpose of a context vector. It is an activation function; It should be noted that the weight matrix, bias vector, and context vector are all learned through model training; Subsequently, the Softmax function is used to normalize these scores to obtain the attention weights. : in, This indicates that when constructing a fragment representation, the first... The proportion of features that should occupy each time step; Finally, by performing a weighted summation of all hidden states within a fragment, we obtain the fragment-level representation vector. : 303. Generate a vector of the global dialogue history through a global attention mechanism: After the above steps, the original temporal feature vector sequence is transformed into a shorter, higher-level segment-level representation sequence that reflects the evolution of the dialogue over a larger time scale. The model then applies a similar encoding and attention process to this sequence of segments again: Use another Bi-GRU to represent the fragment-level vector Encode to obtain the global context representation of each segment. ; A global attention mechanism is applied to compute the representation of each segment. The importance of [the subject / object] and its attention weight. It should be emphasized that this process is exactly the same as the local attention mechanism, but the object of operation is the fragment-level representation vector; By representing all global contexts By performing a weighted summation, a unique vector representing the global dialogue history is obtained. : It should be noted that this global dialogue history vector encapsulates all the key information from the micro to the macro, and from the local to the global, throughout the entire dialogue history. 304. Generate real-time situation deterioration scores: Finally, the vector input of the global dialogue history is fed into a simple feedforward neural network (typically one or two fully connected layers followed by a sigmoid activation function) to regress a continuous value between 0 and 1, namely the real-time context degradation score: in, and These are the weights and biases of the output layer. It is the Sigmoid function; Specifically, the higher the real-time situation deterioration score, the greater the likelihood that the model predicts the current situation will develop in a dangerous direction; It should be noted that the model learns all of the above parameters by pre-training on a labeled dataset containing a large number of dialogues that gradually evolve from normal to dangerous. Furthermore, the various pre-trained models in this invention, such as the deep separation network, speaker recognition model, text sentiment analysis model, speech emotion recognition (SER) model, and temporal context deterioration prediction model, can all be trained and constructed using standard deep learning processes. Taking the temporal context deterioration prediction model, the core of this invention, as an example, its training process includes: Construct a large-scale dialogue audio database that includes real or simulated dialogue scenarios of varying intensity, ranging from casual conversation and arguments to explicit coercion and bullying. Label the data, at least in the time dimension, labeling the continuous change curve of the "situation deterioration score" or labeling the danger level at key time points. For all audio files in the database, extract the temporal feature vector sequence according to step S2; The feature sequence and its corresponding labeled score / rank are used as input and target. The loss function, such as mean squared error (MSE), is used to optimize the parameters (such as the weights and biases of each layer) of the model (such as the aforementioned hierarchical attention network) through backpropagation algorithm until the model's performance on the validation set converges. It should be noted that for other models, such as deep separable networks and SER models, the above training process is also followed, and training is performed on the corresponding public datasets (such as LibriSpeech, IEMOCAP, etc.) or self-built datasets. S4. Based on the real-time situation deterioration score, execute the preset graded response strategy, which includes deterrence or alarm operations. It should be noted that this step aims to transform the quantified and abstract situation deterioration score in step S3 into specific and targeted safety intervention actions. The core idea of ​​this step is that response measures should not be singular and rigid, but should adopt a progressive strategy from low intervention to high intervention, from covert to overt, and from deterrence to seeking help, according to the gradual increase in the level of danger, in order to maximize the safety protection effect while minimizing unnecessary interference. Furthermore, the preset graded response strategy is triggered based on a comparison of the real-time situation deterioration score with a series of preset thresholds, wherein the range of the preset thresholds is as follows: These preset thresholds define different levels of danger; Furthermore, a series of preset thresholds can be determined as follows: The predictive model of this invention is tested on a validation set containing a large amount of dialogue data labeled with different risk levels. By plotting the Receiver Operating Characteristic (ROC) curve or the precision-recall curve, the optimal operating point is selected as the threshold based on different tolerances for false positives and false negatives in different application scenarios. Taking the first and second preset thresholds as examples, to ensure that no potential risks are missed, the first preset threshold can be selected at a point with a high recall rate; to ensure the accuracy of the deterrent operation, the second preset threshold can be selected at a point where precision and recall are relatively balanced. It should be emphasized that these preset thresholds can also be made available to users, allowing them to customize settings according to their personal sensitivities. Specifically, the tiered response strategy includes the following levels: Level 1: Potential risk monitoring (situation deterioration score is greater than the first preset threshold); When the real-time situation deterioration score exceeds the first preset threshold (e.g., 0.4) for the first time, it is determined that the current situation has entered the potential risk stage. At this time, there may be signs of disharmony in the conversation, but it has not yet reached a clear level of danger. In order to avoid premature intervention that may alert the other party or cause false alarms and embarrassment, a sound source direction finding step will be activated. Specifically, using the raw multi-channel audio signals acquired in step S1, the location of the sound source in the speech stream of the dialogue parties is estimated and tracked in real time, and the physical space dynamics are obtained. This sound source location estimation can be based on the Generalized Cross-Correlation (GCC) method, and preferably uses the Phase Transform (PHAT) weighted GCC-PHAT algorithm, because this algorithm has good robustness to reverberation environments. Therefore, for any pair of microphones... and their cross-correlation function It can be represented as: in, and These are the Fourier transforms of the two microphone signals, yes Conjugate; Specifically, by finding Delay difference to achieve the maximum value Based on the geometry of the microphone array, preliminary observations of the direction of sound arrival can be calculated. ; Furthermore, since single estimations may contain noise and jitter, this invention utilizes a Kalman filter to process continuous preliminary observation sequences. In processing this data, it's important to explain that the Kalman filter is a powerful tool for smoothing and predicting time-series data. Through an iterative prediction-update process, it combines state transition model-based predictions with noisy observations to output a smoother and more robust azimuth tracking result. ; Specifically, by continuously analyzing the trend of changes in the location of the sound source, if the location of the sound source of the dialogue party rapidly and continuously approaches the device holder in a short period of time (i.e., the rate of change of azimuth angle is fast and the distance estimate decreases), it indicates a strong aggressive signal in physical space. Then, the trend of such changes in the location of the sound source (angular velocity, approach velocity) is used as a new feature and input into a decision logic mechanism. This mechanism significantly improves the ability to predict physical attack intentions by dynamically adjusting the real-time situation deterioration score (e.g., applying a gain factor positively correlated with the approach velocity) or directly using it as an additional decision basis to trigger the next level of response. Specifically, in one feasible implementation of the present invention, the quantification of the changing trend of the sound source orientation is used to adjust the real-time situational degradation score. This can be achieved through a gain model, for example, by calculating the angular velocity of the sound source orientation of the dialogue party. and radial approximation velocity The new situation worsened the score. It can be updated using the following formula: in, and It is the preset gain coefficient. This indicates that only the approximation case (where the velocity is positive) is considered. The function ensures that the updated score remains within the interval [0, 1]. It should be noted that this update enables the situation deterioration score to incorporate spatial dynamic information, making it more sensitive to dangerous situations with intentions of physical approach. Level Two: Proactive Intervention and Dynamic Deterrence (Situation Deterioration Score is greater than the second preset threshold); When the real-time situation deterioration score rises further and exceeds the second preset threshold (e.g., 0.6) on top of the first preset threshold, it is determined that the current situation has entered a clear dangerous stage and requires immediate proactive intervention. At this time, dynamic deterrence operation will be triggered. It should be noted that the deterrent operation is characterized by its "dynamic" and "targeted" nature, aiming to break the perpetrator's psychological expectations in an unexpected way and create a deterrent effect of "being monitored by a third party." Specifically, once the situation deterioration score exceeds a second preset threshold, the system immediately captures the most recent speech segment from the speaker and uses an advanced voice conversion technology to convert the timbre (i.e., who is speaking) of the speech segment into a preset, authoritative, or intimidating timbre, while completely preserving the linguistic content (what was said), rhythm, and emotion (how it was said). For example, a serious, composed adult male or official announcement-style timbre can be preset. This process can be achieved using deep learning-based models (such as StarGAN-VC, AutoVC, etc.) by decoupling and recombining different attributes such as content, timbre, and rhythm in the speech. At this point, a preset warning text (e.g., "Warning: Your words and actions have been recorded. Please stop your inappropriate behavior immediately") is intelligently spliced ​​or fused with the timbre-converted speech segment to generate a highly targeted voice warning, which is then clearly played on the smart terminal device. Level 3: Emergency Alarm and Information Reporting (Situation deterioration score exceeds the third preset threshold); If, after the deterrent action, the situation deterioration score does not decrease significantly within a certain period of time but instead continues to rise, exceeding the highest third preset threshold (e.g., 0.9), or if the situation deterioration score directly spikes above this threshold, the situation is determined to have entered an extremely dangerous state, the deterrent action is ineffective, and external assistance must be sought immediately. The highest level of alarm action will be automatically triggered, which may include, but is not limited to, the following methods: Automatically dials the user's preset emergency contact number or the local emergency number (such as 110). Send help text messages or instant messages to multiple preset contacts, including precise geographical location (GPS location), real-time audio recordings of the scene, and a text description stating "the situation is extremely dangerous"; Obtain permissions from the current smart terminal device and forcibly adjust the volume in the device's permissions to the highest level to attract the attention of people in the surrounding area; It should be noted that this strategy not only intelligently assesses risks, but also designs psychologically grounded interventions that match the risk level, achieving proactivity, timeliness, and effectiveness in ensuring user safety.

[0025] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0026] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0027] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0028] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0029] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0030] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A deterrence and rapid alarm method based on multimodal voice, characterized in that, include: Acquire multi-channel audio signals from the environment and process the multi-channel audio signals to separate multiple independent audio streams, each containing at least the device holder's voice stream and at least one other party's voice stream. For the voice stream of the device holder and the voice stream of the dialogue party, multi-dimensional dialogue dynamic features are extracted at a preset time step to form a temporal feature vector sequence. The multi-dimensional dialogue dynamic features include at least prosodic layer features that reflect the rhythm of dialogue interaction. The sequence of temporal feature vectors is input into a pre-trained temporal situation deterioration prediction model to generate a real-time situation deterioration score that characterizes the degree of danger of the current situation. Based on the real-time situation deterioration score, a preset graded response strategy is executed, which includes deterrence or alarm operations.

2. The deterrence and rapid alarm method based on multimodal voice as described in claim 1, characterized in that, Processing the multi-channel audio signal includes: The multi-channel audio signal is enhanced by spatial filtering to suppress ambient noise, and the enhanced multi-channel audio signal is processed by a deep neural network to separate overlapping speech. Furthermore, by comparing the separated overlapping voiceprint features with the pre-stored device owner voiceprint features, the device owner's voice stream and the dialogue party's voice stream are identified and marked.

3. The deterrence and rapid alarm method based on multimodal voice as described in claim 1 or 2, characterized in that, The multidimensional dialogue dynamics features also include physiological acoustic layer features, which are used to quantify the physiological tension state of the speaker.

4. The deterrence and rapid alarm method based on multimodal voice as described in claim 3, characterized in that, The physiological acoustic layer features are extracted through the speech stream of the device holder and / or the speech stream of the other party in the dialogue, and the physiological acoustic layer features include: The parameter characterizing the fundamental frequency stability is obtained by normalizing the fundamental frequency of the corresponding speech stream through continuous variation. And parameters that reflect the distribution of high-frequency harmonic energy in the corresponding speech stream.

5. The deterrence and rapid alarm method based on multimodal voice as described in claim 1 or 2, characterized in that, The prosodic layer features include: Used to measure the interruption rate or the frequency of speech turn switching in a conversation; In addition, a parameter characterizing the degree of cooperation between the two parties in the dialogue is obtained by calculating the cross-correlation between the device holder's speech stream and the speech stream of the other party in the dialogue.

6. The deterrence and rapid alarm method based on multimodal voice as described in claim 1, characterized in that, The multi-dimensional dialogue dynamics feature also includes a semantic-acoustic inconsistency feature, which is calculated using the device holder's speech stream or the dialogue party's speech stream. The calculation method is as follows: The first and second emotional state distributions are inferred from the text content and acoustic properties of the speech stream, respectively, and a metric is calculated to measure the difference between the first and second emotional state distributions.

7. The deterrence and rapid alarm method based on multimodal voice as described in claim 1, characterized in that, The temporal context deterioration prediction model is a temporal model with a hierarchical attention mechanism, which is configured as follows: The features within local time segments of the temporal feature vector sequence are weighted and aggregated to generate a segment-level representation; The multiple fragment-level representations are weighted and aggregated again to generate a vector representing the global dialogue history, and the real-time context deterioration score is generated based on this vector of global dialogue history.

8. The deterrence and rapid alarm method based on multimodal voice as described in claim 1, characterized in that, The method also includes: When the real-time situation deterioration score exceeds a first preset threshold, the sound source direction finding step is activated. This step uses the collected multi-channel audio signals to estimate and track the sound source location of the speech stream of the dialogue party. Furthermore, the changing trend of the sound source location is used as supplementary information to adjust the real-time situation deterioration score or influence the decision-making of the graded response strategy.

9. The deterrence and rapid alarm method based on multimodal voice as described in claim 8, characterized in that, The sound source direction finding step includes: Based on the phase information of the cross-correlation function between multi-channel audio signals, preliminary observations of the sound source's location are estimated. Furthermore, the Kalman filter is used to process the continuous observations of the sound source's azimuth, outputting a smooth and robust azimuth tracking result.

10. The deterrence and rapid alarm method based on multimodal voice as described in claim 1, characterized in that, Tiered response strategies include: When the real-time situation deterioration score exceeds the second preset threshold, a dynamic deterrence operation is triggered. The dynamic deterrence operation includes: capturing a segment of the conversational speaker's voice, using a voice conversion technology that changes the timbre while keeping the voice content unchanged, and combining it with a preset warning text to generate and play a targeted voice warning.